Preference tuning: RLHF, DPO and related methods
Understand how models are tuned on human preferences, from reward models to direct preference losses
A taste of a lesson
During RLHF our reward keeps climbing but human raters now prefer the older checkpoint. What is going on?
That is the classic signature of overoptimisation. The policy has found patterns the reward model likes but people do not, so reward and real quality have separated. First, look at what changed: compare response length and tone between checkpoints, since longer, more agreeable answers are a frequent exploit. Then check the KL divergence from the reference; if it grew a lot, the policy has moved far from the data the reward model understands. Options include a stronger KL penalty, stopping earlier, or retraining the reward model on pairs from the new policy. Quick check: what would you plot to confirm length is the exploit?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Describe preference data and train a reward model with a pairwise loss
- Explain the RL objective with a KL penalty and why the penalty matters
- Recognise reward hacking and overoptimisation in training results
- Explain how DPO optimises preferences directly and what beta controls
- Evaluate preference tuned models while controlling for length and style
Lesson plan
- 1 Preference data Understand chosen and rejected pairs, how they are collected and what makes them useful. Start
- 2 Reward models Train a model that scores responses using a pairwise preference loss. Start
- 3 RL with a KL penalty Explain the objective of maximising reward while staying close to a reference model. Start
- 4 Reward hacking and overoptimisation Spot when the policy is gaming the reward rather than improving. Start
- 5 Direct preference methods Understand how DPO turns preferences into a supervised style loss. Start
- 6 Verifiable rewards and evaluation Use automatic checkers where possible and evaluate the tuned model fairly. Start
Try asking
About this tutor
For advanced learners who know supervised fine tuning and want to understand the preference tuning stage that shapes how assistants behave. You will build up the classic pipeline of preference data, a reward model and reinforcement learning with a penalty that keeps the model close to its reference. Then you study the main failure, reward hacking, and how direct preference methods such as DPO reach a similar goal with a simple loss and no separate reward model. The final lessons cover annotation quality and disagreement, rewards from automatic checkers for tasks with verifiable answers, and how to evaluate the result without being misled by length or style.
Reviews
4.7
3 ratingsSample
- Felix B.Sample
The lesson on overoptimisation described our run almost exactly. Plotting length against reward made the problem obvious to the whole team.
- Sakura N.Sample
Clear derivation of the DPO loss from the RL objective. Dense in places, but the tutor slowed down whenever I asked.
- Mohammed R.Sample
I appreciated that variants were treated as a family to test, not a leaderboard. The annotation quality lesson was more useful than I expected.
About the teacher
Fine tuning with judgment: when to do it, how to do it well, and how to know it worked
9 tutors 428 lessons taught Sample
I teach fine tuning and post training: choosing between prompting, retrieval and tuning, building datasets, parameter efficient methods, instruction and preference tuning, and evaluating the result. My background is in applied machine learning projects where the expensive mistake was usually tuning a model before anyone had defined what better meant. That is why I start every topic with the evaluation...
See Neha's profile and tutorsMore like this
Other tutors on the same or nearby topics.