Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.

Find a tutor

657 tutors in 31 topics, built by 75 teachers. Each one follows a lesson plan its teacher wrote.

Filters

Clear
Evaluating Retrieval Quality

Evaluating Retrieval Quality

Measure whether your retrieval finds the right passages with real queries, relevance labels and the right metrics.AdvancedEvaluation and testing4.7(3)75 lessonsSample
Fumiko Arai$11
Evaluating an AI Feature Before Launch

Evaluating an AI Feature Before Launch

Build an evaluation set, grade outputs reliably and set launch gates your team agrees onAdvancedAI for product and operations4.7(3)70 lessonsSample
Sanjana Rao$14
Reviewing AI Support Answers for Quality

Reviewing AI Support Answers for Quality

Build a sampling and rubric process to judge AI support answers and catch failures earlyAdvancedAI for product and operations4.7(3)67 lessonsSample
Tariq Haddad$12
Evaluating a model after fine tuning

Evaluating a model after fine tuning

Prove a tuned model is actually better, on your task and everywhere else it mattersIntermediateEvaluation and testing4.7(3)65 lessonsSample
Neha Varadan$8
Logging and Tracing LLM Calls

Logging and Tracing LLM Calls

Record every model call and pipeline step so you can explain cost, slowness and bad answers.IntermediateBuilding with LLM APIs4.0(3)65 lessonsSample
Farid Haddad$8
Unit Tests for Python AI Code

Unit Tests for Python AI Code

Test the deterministic parts of LLM apps properly, with fake model clients instead of flaky live calls.IntermediateEvaluation and testing4.3(3)61 lessonsSample
Felix Brandt$6
Model Graded Evals and Their Pitfalls

Model Graded Evals and Their Pitfalls

Use language models as graders without fooling yourself: biases, validation against people and safeguards.AdvancedEvaluation and testing4.3(4)60 lessonsSample
Gonzalo Ibarra$11
Build Your First Eval Set

Build Your First Eval Set

Create a small, honest set of test cases for your LLM feature and use it to judge every change.BeginnerEvaluation and testing4.7(3)52 lessonsSample
Gonzalo IbarraFree
Cost and Latency Budgets for LLM Features

Cost and Latency Budgets for LLM Features

Set cost and speed targets for an LLM feature, measure them honestly and trade them against quality.All levelsBuilding with LLM APIs4.5(2)52 lessonsSample
Farid Haddad$8
Red Teaming and Prompt Injection Testing

Red Teaming and Prompt Injection Testing

Test your LLM app against jailbreaks, prompt injection and data leaks with a repeatable attack suite.AdvancedEvaluation and testing4.7(3)49 lessonsSample
Gonzalo Ibarra$12
Rubric Grading for Open Ended Output

Rubric Grading for Open Ended Output

Design rubrics that make grading emails, summaries and explanations consistent, fair and repeatable.IntermediateEvaluation and testing4.3(3)49 lessonsSample
Gonzalo Ibarra$6
Measuring Hallucination in LLM Apps

Measuring Hallucination in LLM Apps

Measure how often your app states unsupported things, with claim level checks you can repeat and trust.IntermediateEvaluation and testing4.5(2)42 lessonsSample
Gonzalo Ibarra$8