Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Regression Testing Your Prompts

Catch quality drops before users do by running an eval suite on every prompt, model or setting change.

By Gonzalo Ibarra Intermediate Evaluation and testing 4.5(2) 42 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $7 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Regression Testing Your Prompts AI tutor following Gonzalo Ibarra's plan
Student:

My suite passed 88 percent yesterday and 85 percent today with no changes. Is something broken?

Tutor:

Maybe not. With variable outputs, a few points of movement between runs can be normal noise, especially on a small suite. Check three things: whether you run each case once or several times, which specific cases flipped, and whether any must pass case failed. If the flips are scattered borderline cases, it is likely noise; run important cases several times and use pass rates. If one category flipped together, look for a provider update. How many cases does your suite have?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Build a regression suite from eval cases, past bugs and critical behaviours
  • Combine deterministic checks with graded quality checks
  • Handle variable outputs with repeated runs and pass rate thresholds
  • Read diff reports and review every flipped case
  • Schedule smoke, full and periodic runs within budget

Lesson plan

6 lessons. Pick one to start there.

  1. 1 What regresses and why List the changes that can lower quality in an LLM feature. Start
  2. 2 Building the suite Assemble cases from evals, past bugs and critical behaviours. Start
  3. 3 Checks that fit the output Pair deterministic checks with graded ones. Start
  4. 4 Living with variability Judge variable outputs with repeated runs and thresholds. Start
  5. 5 Diff reports and review See exactly which cases changed and decide if the change is acceptable. Start
  6. 6 When to run what Run the right subset at the right time within budget. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For developers maintaining LLM features where prompts get edited, models get upgraded and settings get tweaked, and where fixing one case often breaks another. You learn to build a regression suite from your eval set and past bugs, combine deterministic checks with graded ones, deal with non deterministic outputs through repeated runs and pass rate thresholds, produce diff reports that show which cases flipped, and run the right subset at the right time: quick checks on every change, the full suite before release and scheduled runs to detect provider side changes.

Reviews

4.5

2 ratingsSample

  • Hassan B.Sample

    Diff reports with outputs side by side changed our review meetings. We caught a prompt edit that fixed tone but broke refusals in one category.

  • Greta P.Sample

    Repeated runs with pass rate thresholds solved our flaky suite problem. The scheduled weekly run caught a provider change once already.

About the teacher

Gonzalo Ibarra

Evaluation for LLM products: eval sets, graders, regression tests, hallucination checks and live experiments

9 tutors 4.5(17) 294 lessons taught Sample

I teach people how to know whether their LLM feature is any good, which is harder than building it. I started in software testing and quality work, moved into data analysis, and now help teams measure systems whose outputs vary from run to run. My lessons are about method: write down what good looks like, collect examples, grade them in...

See Gonzalo's profile and tutors