Advanced30 min
Preference Tuning in Practice
Building preference pairs, running DPO end to end, beta tuning, and diagnosing reward over-optimisation.
Assumes you know
Not yet written
This lesson is on the syllabus but has no text yet
The full curriculum is published up front so you can see the whole route and its dependencies. Lessons are being written in curriculum order.
What it will cover
- DPO in practice
- preference pairs
- beta
- over-optimisation