Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

How AI labs test models for safety

Understand evaluations, red teaming and staged release, and how to read a model's safety report

Your first 3 replies from the tutor are free. Then you can continue the lesson for $10 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

How AI labs test models for safety AI tutor following Bao Tran's plan
Student:

If a lab says its model failed to show dangerous capabilities in testing, can we trust that?

Tutor:

It is useful evidence, but weigh it carefully. Ask how hard they tried: elicitation effort matters, and weak prompting or no fine tuning can understate capability. Ask who tested: only internal teams, or independent experts and government institutes with real access? Were the thresholds defined before testing? What was not tested? And remember the basic limit: a negative result shows the tests did not find the capability, not that it is absent. Exercise: write three questions you would ask the lab about its elicitation methods.

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Distinguish capability, propensity and safeguard evaluations
  • Explain red teaming, elicitation and why under elicitation understates risk
  • Describe how capability thresholds and staged release shape deployment decisions
  • Read a system card critically and identify methodological gaps

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Kinds of evaluation Distinguish what different safety evaluations measure. Start
  2. 2 Red teaming and elicitation Understand adversarial testing and why effort matters. Start
  3. 3 Agentic and long horizon tests See how autonomous behaviour is evaluated. Start
  4. 4 Thresholds and release decisions Understand how results feed into governance. Start
  5. 5 Reading a system card Critique a safety report methodically. Start
  6. 6 The limits of testing Understand what evaluation can and cannot establish. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For policy readers, journalists, researchers, procurement staff and technical learners who want to know what safety testing actually involves. You study the main kinds of evaluation: capability evaluations for dangerous domains, propensity and behavioural tests, red teaming by internal and external experts, automated adversarial testing, and evaluations of agentic behaviour. You learn how results feed into release decisions through frameworks with capability thresholds, and how staged release and monitoring work. You practise reading a system card critically and consider the limits of testing: you can show a risk exists, but not prove its absence.

Reviews

Students can review a tutor after a paid lesson. Nobody has yet.

About the teacher

Bao Tran

I teach the big debates in AI, from alignment and safety testing to copyright, fairly and clearly

9 tutors 4.5(19) 312 lessons taught Sample

I teach the questions people argue about: whether advanced AI could be dangerous, how labs test models, why models flatter us, who owns AI output, and how power over AI is distributed. I present each position in its strongest form before we examine it, and I separate evidence from speculation out loud. My background is in research communication and running...

See Bao's profile and tutors