Teacher since May 2025
Farid Haddad
Keeps LLM features fast, affordable and available as traffic grows: cost, caching, retries and observability
9
tutors built
4.5Sample
average from 19 reviews
347Sample
lessons taught by their tutors
About Farid
Most of my working life has been on platform and reliability teams, and these days I spend it on LLM features: the bills that surprise people, the 429 errors on launch day, the logs nobody can read. I teach the operational side of building with model APIs. We estimate costs before writing code, add retries that do not make outages worse, cache what can be cached, and log enough to answer why something cost so much or ran so slowly. I like numbers you can check yourself, so most lessons include a small calculation or a log we read together.
Knows about
Tutors by Farid
9 tutors
Logging and Tracing LLM Calls
Record every model call and pipeline step so you can explain cost, slowness and bad answers.65 lessonsSampleFarid Haddad$8Fallbacks for Provider Outages and ErrorsFallbacks for Provider Outages and Errors
Keep LLM features working through outages, overloads and slowdowns with deadlines, breakers and fallbacks.60 lessonsSampleFarid Haddad$11Token Counting and Cost ControlToken Counting and Cost Control
Count tokens, predict what a feature will cost and cut spending without hurting answer quality.57 lessonsSampleFarid Haddad$4Rate Limits, Retries and BackoffRate Limits, Retries and Backoff
Handle 429s and overloads gracefully with backoff, jitter, client side pacing and retries that never cause storms.53 lessonsSampleFarid Haddad$7Cost and Latency Budgets for LLM FeaturesCost and Latency Budgets for LLM Features
Set cost and speed targets for an LLM feature, measure them honestly and trade them against quality.52 lessonsSampleFarid Haddad$8API Key Safety and Secrets HandlingAPI Key Safety and Secrets Handling
Keep model API keys out of code, repos, browsers and logs, and know exactly what to do if one leaks.32 lessonsSampleFarid HaddadFreePrompt Caching for Lower Cost and LatencyPrompt Caching for Lower Cost and Latency
Structure prompts so repeated prefixes are cached, then measure the savings in cost and response time.28 lessonsSampleFarid Haddad$7Batch Jobs for Large LLM WorkloadsBatch Jobs for Large LLM Workloads
Run thousands of model requests as batch jobs that are cheaper, resumable and easy to check.Farid Haddad$7LLM APIs on a Solo Developer BudgetLLM APIs on a Solo Developer Budget
Build and launch a side project on model APIs without a surprise bill, abuse or an overbuilt setup.Farid Haddad$4Recent reviews
What students said about Farid's tutors.
- Ken T.Sample
Advanced and dense, as advertised. The runbook lesson was short but useful.
- Anika R.Sample
The breakdown of where tokens go showed our examples block was bigger than everything else combined. Trimmed it and quality held on our test set.
- Tobias E.Sample
Clear explanation of why our German and Polish users cost more. Practical and not preachy.
- Mohammed E.Sample
We had a second provider wired up but had never evaluated it. Our parser broke on its output format in the first fault injection test. Found it in staging instead of production.
- Signe L.Sample
Circuit breaker explanation with real thresholds and probes was excellent. The point that degraded modes can beat a weaker model changed our design.
- Daniel K.Sample
As a product lead I appreciated budgets per active user rather than per token. The engineering parts were optional for me but still readable.