Model Graded Evals and Their Pitfalls
Use language models as graders without fooling yourself: biases, validation against people and safeguards.
A taste of a lesson
Our model judge prefers version B in 70 percent of pairwise comparisons. Can we ship B?
Not on that number alone. First check position bias: rerun every comparison with A and B swapped and count only verdicts that agree in both orders. Then check length: if B's answers are much longer, the judge may simply prefer verbosity. Finally, compare the judge with people: have humans judge forty of the same pairs and measure agreement. If B still wins after those checks, the result is much more convincing. Was B shown first in most comparisons?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Choose between rubric, reference based, pairwise and grounded judge designs
- Detect position, length, self preference and leniency biases with experiments
- Protect judges from instructions hidden in graded outputs
- Validate a judge against human labels with agreement measures
- Monitor judges over time and avoid optimising for the judge
Lesson plan
- 1 Judge designs Pick the grading design that fits your task. Start
- 2 Known biases Recognise the systematic errors model judges make. Start
- 3 Testing for bias Run small experiments that reveal your judge's biases. Start
- 4 Hardening the judge Reduce bias and resist manipulation. Start
- 5 Validating against people Measure how well the judge agrees with human graders. Start
- 6 Living with a judge Keep model grading honest as the system evolves. Start
Try asking
About this tutor
For experienced builders who want to scale evaluation with a model acting as judge and need to know when its grades can be trusted. You learn the main designs (single output grading against a rubric, reference based grading, pairwise comparison), the known biases (position, length, self preference, leniency, sensitivity to wording) and attacks such as instructions hidden in the graded output. Most importantly, you learn to validate a judge against human labels with agreement measures, monitor it over time, and avoid optimising your system to please the judge rather than users.
Reviews
4.3
4 ratingsSample
- Linnea H.Sample
Strong content but very dense. I needed to repeat two lessons. Worth it if you already run evals.
- Nomvula D.Sample
The validation lesson gave us a real process: human labels, kappa, confusion matrix. Our judge prompt improved a lot after studying disagreements.
- Paulo J.Sample
Rigorous and sceptical in a healthy way. The Goodhart warning matched exactly what had happened to our scores.
- Erik S.Sample
Swapping positions halved our judge's apparent preference for the new version. Embarrassing, but much better to learn it here than after launch.
About the teacher
Evaluation for LLM products: eval sets, graders, regression tests, hallucination checks and live experiments
9 tutors 294 lessons taught Sample
I teach people how to know whether their LLM feature is any good, which is harder than building it. I started in software testing and quality work, moved into data analysis, and now help teams measure systems whose outputs vary from run to run. My lessons are about method: write down what good looks like, collect examples, grade them in...
See Gonzalo's profile and tutorsMore like this
Other tutors on the same or nearby topics.