Evaluation and testing
Measure whether an AI feature works, with test sets, metrics and reviews.
21tutors
12teachers
1free to start
$4 to $14per paid lesson
Evaluation and testing tutors
21 tutors
Evaluating Retrieval Quality
Evaluating Retrieval Quality
Measure whether your retrieval finds the right passages with real queries, relevance labels and the right metrics.75 lessonsSampleFumiko Arai$11Evaluating an AI Feature Before LaunchEvaluating an AI Feature Before Launch
Build an evaluation set, grade outputs reliably and set launch gates your team agrees on70 lessonsSampleSanjana Rao$14Reviewing AI Support Answers for QualityReviewing AI Support Answers for Quality
Build a sampling and rubric process to judge AI support answers and catch failures early67 lessonsSampleTariq Haddad$12Evaluating a model after fine tuningEvaluating a model after fine tuning
Prove a tuned model is actually better, on your task and everywhere else it matters65 lessonsSampleNeha Varadan$8Logging and Tracing LLM CallsLogging and Tracing LLM Calls
Record every model call and pipeline step so you can explain cost, slowness and bad answers.65 lessonsSampleFarid Haddad$8Unit Tests for Python AI CodeUnit Tests for Python AI Code
Test the deterministic parts of LLM apps properly, with fake model clients instead of flaky live calls.61 lessonsSampleFelix Brandt$6Model Graded Evals and Their PitfallsModel Graded Evals and Their Pitfalls
Use language models as graders without fooling yourself: biases, validation against people and safeguards.60 lessonsSampleGonzalo Ibarra$11Build Your First Eval SetBuild Your First Eval Set
Create a small, honest set of test cases for your LLM feature and use it to judge every change.52 lessonsSampleGonzalo IbarraFreeCost and Latency Budgets for LLM FeaturesCost and Latency Budgets for LLM Features
Set cost and speed targets for an LLM feature, measure them honestly and trade them against quality.52 lessonsSampleFarid Haddad$8Red Teaming and Prompt Injection TestingRed Teaming and Prompt Injection Testing
Test your LLM app against jailbreaks, prompt injection and data leaks with a repeatable attack suite.49 lessonsSampleGonzalo Ibarra$12Rubric Grading for Open Ended OutputRubric Grading for Open Ended Output
Design rubrics that make grading emails, summaries and explanations consistent, fair and repeatable.49 lessonsSampleGonzalo Ibarra$6Measuring Hallucination in LLM AppsMeasuring Hallucination in LLM Apps
Measure how often your app states unsupported things, with claim level checks you can repeat and trust.42 lessonsSampleGonzalo Ibarra$8Teachers who teach Evaluation and testing
They wrote the lesson plans these tutors follow.
Tariq Haddad
Risk, vendors, data and support: the unglamorous work that makes AI safe to rely onvendor trials, data readiness, AI risk and compliance basics9 tutorsSampleNeha Varadan
Fine tuning with judgment: when to do it, how to do it well, and how to know it workedfine tuning strategy, dataset preparation, LoRA and parameter efficient methods9 tutorsSampleValentin Moreau
Lecturer and learning designer for universities, workplace training and coaching practicesHigher education teaching, assessment design, instructional design9 tutorsSampleSanjana Rao
Product management for AI features: deciding, specifying, testing and pricing them wellAI product management, product requirements, UX of AI features9 tutorsSampleFelix Brandt
Numerical Python and code quality for data and AI projects that have outgrown a single notebookNumPy, pandas, plotting9 tutorsSampleFarid Haddad
Keeps LLM features fast, affordable and available as traffic grows: cost, caching, retries and observabilitytoken counting, cost control, prompt caching9 tutorsSampleMore in Building with AI
- Python for AI The Python you need for AI work: data, notebooks, libraries and small scripts. 21 tutors
- Building with LLM APIs Call language models from your own code: requests, streaming, tools and costs. 30 tutors
- RAG and search Connect a model to your own documents with embeddings, search and retrieval. 21 tutors
- AI agents Models that plan, use tools and act in steps, and how to keep them on track. 31 tutors
- AI automation Hand repetitive work to AI with workflows, triggers and simple integrations. 25 tutors
- AI coding assistants Write, read and debug code faster with an AI pair programmer. 30 tutors