Golden Answers and Reference Checks
Write reference answers experts agree on and compare model output to them with the right matching method.
A taste of a lesson
My reference says '30 days' and the model says 'one month'. My check marks it wrong. Is it wrong?
That depends on the policy, which is exactly why matching needs thought. If the policy says 30 days, one month is not precisely the same and could mislead in February, so marking it wrong may be fair. If you decide they are equivalent, normalise both to days before comparing. Either way, write the rule into your guidelines so every case is judged the same. What does the source document actually say?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Decide between a single golden answer and a required facts list
- Write and review reference answers with subject experts
- Choose fair matching methods for text, numbers, lists and structured fields
- Recognise the limits of fuzzy and similarity scores
- Keep references current and report results per field or fact
Lesson plan
- 1 When golden answers fit Decide which tasks suit single references and which need fact lists. Start
- 2 Writing references with experts Produce reference answers that are correct and agreed. Start
- 3 Matching text, numbers and dates Compare simple answers without false failures. Start
- 4 Lists and structured output Score lists and JSON fields fairly. Start
- 5 Limits of similarity scores Know when fuzzy and embedding similarity mislead. Start
- 6 Keeping references current Maintain references as facts and policies change. Start
Try asking
About this tutor
For beginners building evaluation for tasks with definite answers: extracting fields, answering factual questions from documents, classification and calculations. You learn when a single golden answer fits and when a list of required facts works better, how to write references with subject experts, and how to compare outputs fairly: normalised exact match, numeric tolerance, set comparison for lists, field by field comparison for structured output, and the limits of fuzzy and similarity scores. The tutor also covers keeping references current when facts change and resolving disagreements between experts.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Evaluation for LLM products: eval sets, graders, regression tests, hallucination checks and live experiments
9 tutors 294 lessons taught Sample
I teach people how to know whether their LLM feature is any good, which is harder than building it. I started in software testing and quality work, moved into data analysis, and now help teams measure systems whose outputs vary from run to run. My lessons are about method: write down what good looks like, collect examples, grade them in...
See Gonzalo's profile and tutorsMore like this
Other tutors on the same or nearby topics.