Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Model Graded Evals and Their Pitfalls

Use language models as graders without fooling yourself: biases, validation against people and safeguards.

By Gonzalo Ibarra Advanced Evaluation and testing 4.3(4) 60 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $11 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Model Graded Evals and Their Pitfalls AI tutor following Gonzalo Ibarra's plan
Student:

Our model judge prefers version B in 70 percent of pairwise comparisons. Can we ship B?

Tutor:

Not on that number alone. First check position bias: rerun every comparison with A and B swapped and count only verdicts that agree in both orders. Then check length: if B's answers are much longer, the judge may simply prefer verbosity. Finally, compare the judge with people: have humans judge forty of the same pairs and measure agreement. If B still wins after those checks, the result is much more convincing. Was B shown first in most comparisons?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Choose between rubric, reference based, pairwise and grounded judge designs
  • Detect position, length, self preference and leniency biases with experiments
  • Protect judges from instructions hidden in graded outputs
  • Validate a judge against human labels with agreement measures
  • Monitor judges over time and avoid optimising for the judge

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Judge designs Pick the grading design that fits your task. Start
  2. 2 Known biases Recognise the systematic errors model judges make. Start
  3. 3 Testing for bias Run small experiments that reveal your judge's biases. Start
  4. 4 Hardening the judge Reduce bias and resist manipulation. Start
  5. 5 Validating against people Measure how well the judge agrees with human graders. Start
  6. 6 Living with a judge Keep model grading honest as the system evolves. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For experienced builders who want to scale evaluation with a model acting as judge and need to know when its grades can be trusted. You learn the main designs (single output grading against a rubric, reference based grading, pairwise comparison), the known biases (position, length, self preference, leniency, sensitivity to wording) and attacks such as instructions hidden in the graded output. Most importantly, you learn to validate a judge against human labels with agreement measures, monitor it over time, and avoid optimising your system to please the judge rather than users.

Reviews

4.3

4 ratingsSample

  • Linnea H.Sample

    Strong content but very dense. I needed to repeat two lessons. Worth it if you already run evals.

  • Nomvula D.Sample

    The validation lesson gave us a real process: human labels, kappa, confusion matrix. Our judge prompt improved a lot after studying disagreements.

  • Paulo J.Sample

    Rigorous and sceptical in a healthy way. The Goodhart warning matched exactly what had happened to our scores.

  • Erik S.Sample

    Swapping positions halved our judge's apparent preference for the new version. Embarrassing, but much better to learn it here than after launch.

About the teacher

Gonzalo Ibarra

Evaluation for LLM products: eval sets, graders, regression tests, hallucination checks and live experiments

9 tutors 4.5(17) 294 lessons taught Sample

I teach people how to know whether their LLM feature is any good, which is harder than building it. I started in software testing and quality work, moved into data analysis, and now help teams measure systems whose outputs vary from run to run. My lessons are about method: write down what good looks like, collect examples, grade them in...

See Gonzalo's profile and tutors