Teacher since June 2026
Gonzalo Ibarra
Evaluation for LLM products: eval sets, graders, regression tests, hallucination checks and live experiments
9
tutors built
4.5Sample
average from 17 reviews
294Sample
lessons taught by their tutors
About Gonzalo
I teach people how to know whether their LLM feature is any good, which is harder than building it. I started in software testing and quality work, moved into data analysis, and now help teams measure systems whose outputs vary from run to run. My lessons are about method: write down what good looks like, collect examples, grade them in a way you can defend, and keep checking as prompts and models change. I am wary of single scores and of graders nobody has checked, and I try to make evaluation feel like a habit rather than a big project.
Knows about
Tutors by Gonzalo
9 tutors
Model Graded Evals and Their Pitfalls
Use language models as graders without fooling yourself: biases, validation against people and safeguards.60 lessonsSampleGonzalo Ibarra$11Build Your First Eval SetBuild Your First Eval Set
Create a small, honest set of test cases for your LLM feature and use it to judge every change.52 lessonsSampleGonzalo IbarraFreeRed Teaming and Prompt Injection TestingRed Teaming and Prompt Injection Testing
Test your LLM app against jailbreaks, prompt injection and data leaks with a repeatable attack suite.49 lessonsSampleGonzalo Ibarra$12Rubric Grading for Open Ended OutputRubric Grading for Open Ended Output
Design rubrics that make grading emails, summaries and explanations consistent, fair and repeatable.49 lessonsSampleGonzalo Ibarra$6Measuring Hallucination in LLM AppsMeasuring Hallucination in LLM Apps
Measure how often your app states unsupported things, with claim level checks you can repeat and trust.42 lessonsSampleGonzalo Ibarra$8Regression Testing Your PromptsRegression Testing Your Prompts
Catch quality drops before users do by running an eval suite on every prompt, model or setting change.42 lessonsSampleGonzalo Ibarra$7Golden Answers and Reference ChecksGolden Answers and Reference Checks
Write reference answers experts agree on and compare model output to them with the right matching method.Gonzalo Ibarra$5Human Review That ScalesHuman Review That Scales
Set up human review of AI outputs that is consistent, affordable and feeds back into better systems.Gonzalo Ibarra$5A/B Testing LLM Features in ProductionA/B Testing LLM Features in Production
Run fair online experiments on prompts, models and LLM features and read the results without fooling yourself.Gonzalo Ibarra$10Recent reviews
What students said about Gonzalo's tutors.
- Ola N.Sample
The faithfulness versus factuality distinction cleared up weeks of confused discussion. Validating the model judge took effort but was necessary.
- Hassan B.Sample
Diff reports with outputs side by side changed our review meetings. We caught a prompt edit that fixed tone but broke refusals in one category.
- Beatriz A.Sample
Good for domain experts like me who are not engineers. Kappa was explained simply enough to use.
- Mia K.Sample
Clear that prompt patches are not real fixes. We reduced tool permissions instead and the attack success rate dropped across categories.
- Imani W.Sample
Claim level labelling gave us a number we could defend to leadership, with sample size and all. The unanswerable question set was the most revealing part.
- Ana P.Sample
Had a working eval spreadsheet after the first session. It immediately showed my new prompt broke the decline behaviour I thought was fine.