A/B Testing LLM Features in Production
Run fair online experiments on prompts, models and LLM features and read the results without fooling yourself.
A taste of a lesson
After three days our new model shows a significant lift in thumbs up. Can we stop and ship?
I would not yet. Stopping as soon as a result looks significant inflates false positives, three days misses weekly patterns, and novelty can boost early reactions. Thumbs are also a thin signal: few users click them, and those who do are not typical. Check your pre registered primary metric, such as resolution without escalation, and guardrails like latency and cost, then run to the planned end date. What primary metric did you set before launch?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Decide when an online experiment is needed after offline evals
- Randomise by user and write a plan with primary and guardrail metrics
- Interpret thumbs, edits, regenerations and escalations carefully
- Plan sample size and avoid peeking, novelty and multiple comparison traps
- Use shadow testing, gradual ramps and interleaving where they fit
Lesson plan
- 1 Offline versus online Understand what online experiments add to offline evaluation. Start
- 2 Designing the experiment Write a pre registered plan with randomisation and decision rules. Start
- 3 Metrics for LLM features Pick and define signals that reflect real value. Start
- 4 Sample size and duration Plan how many users and how long the experiment needs. Start
- 5 Reading results honestly Avoid the common traps when analysing outcomes. Start
- 6 Rolling out safely Ramp changes and choose alternatives where they fit better. Start
Try asking
About this tutor
For experienced engineers, data scientists and product managers who have passed offline evaluation and now need to know whether a change helps real users. You learn when an online experiment is needed, how to randomise at the user level, choose a primary metric tied to user value and guardrail metrics for latency, cost and safety, interpret LLM specific signals such as thumbs, edits, regenerations and escalations, plan sample size, avoid peeking and novelty effects, and ramp changes safely. You also learn alternatives for ranking features, such as interleaving, and when shadow testing is the better first step.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Evaluation for LLM products: eval sets, graders, regression tests, hallucination checks and live experiments
9 tutors 294 lessons taught Sample
I teach people how to know whether their LLM feature is any good, which is harder than building it. I started in software testing and quality work, moved into data analysis, and now help teams measure systems whose outputs vary from run to run. My lessons are about method: write down what good looks like, collect examples, grade them in...
See Gonzalo's profile and tutorsMore like this
Other tutors on the same or nearby topics.