Skip to content
VibeFormer
Advanced30 min

What RLHF Does and Does Not Fix

An honest account of preference-based alignment: what it demonstrably achieves, and the failure modes it cannot address — sycophancy, annotator bias and reward-model overoptimisation.

Not yet written

This lesson is on the syllabus but has no text yet

The full curriculum is published up front so you can see the whole route and its dependencies. Lessons are being written in curriculum order.

What it will cover

  • RLHF
  • sycophancy
  • annotator bias
  • KL penalty
  • overoptimisation