Advanced30 min
Scalable Oversight
How to supervise a system on tasks humans cannot evaluate directly: debate, recursive reward modelling, weak-to-strong generalisation and their open problems.
Not yet written
This lesson is on the syllabus but has no text yet
The full curriculum is published up front so you can see the whole route and its dependencies. Lessons are being written in curriculum order.
What it will cover
- scalable oversight
- debate
- recursive reward modelling
- weak-to-strong