Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Testing Prompt Changes With Small Test Sets

Stop guessing whether a prompt edit helped: compare versions on a small, fixed set of real inputs.

Your first 3 replies from the tutor are free. Then you can continue the lesson for $10 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Testing Prompt Changes With Small Test Sets AI tutor following Darius Feld's plan
Student:

I tweaked our ticket summary prompt and the new one looked much better on the ticket I tried. Can I roll it out?

Tutor:

Not on one ticket. Outputs vary between runs, so a single good result may be luck. Pull fifteen recent tickets, including a couple of messy ones, and write four yes or no checks, for example: mentions the customer's actual problem, states the next step, under 80 words, no invented order details. Run both versions on all fifteen and count passes. If the new one fixes several failures and breaks none, roll it out. What would your four checks be?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Build a fixed test set of real and edge case inputs
  • Write a short yes or no checklist that defines a good output
  • Compare prompt versions fairly, including repeat runs and blind scoring
  • Use model graded scoring with spot checks and known biases in mind

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Why one run proves little Understand run to run variation and why single comparisons mislead. Start
  2. 2 Building the test set Assemble 10 to 20 representative inputs, including edge cases. Start
  3. 3 Defining good with a checklist Turn quality into a few consistent yes or no criteria. Start
  4. 4 Running a fair comparison Compare versions in a way that isolates the effect of one change. Start
  5. 5 Using a model as a grader Save time with model graded checks while controlling their biases. Start
  6. 6 Reading results honestly Draw conclusions that match the size of the evidence. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

Most people judge a prompt change by trying it once and liking the result. Because outputs vary from run to run, that tells you very little. This tutor teaches a lightweight testing habit for everyday power users and team leads who maintain shared prompts: build a small test set of real inputs, define what good looks like as a short checklist, run old and new versions side by side, and record results in a simple sheet. You also learn the traps, including random variation, comparing too many changes at once, and letting a model judge outputs without checking its biases. No coding required.

Reviews

Students can review a tutor after a paid lesson. Nobody has yet.

About the teacher

Darius Feld

Prompt workflows for heavy users: chaining, standing instructions, long documents and reasoning models

9 tutors 4.5(15) 269 lessons taught Sample

I work with people who already use AI assistants every day and want more dependable results. My background is in operations and process design, which taught me to treat a prompt like a small procedure: inputs, steps, checks and a clear output. I teach chaining, reusable instructions, long document work and how to test whether a prompt change actually helped....

See Darius's profile and tutors