Fallbacks for Provider Outages and Errors
Keep LLM features working through outages, overloads and slowdowns with deadlines, breakers and fallbacks.
A taste of a lesson
We want to fail over to a second provider during outages. Is it enough to swap the API call?
Swapping the call is the easy part. The second model will respond differently: formatting, tool calling, refusals and length can all change, so prompts and parsers may need provider specific variants. Run your eval set on the fallback and know its quality gap before an outage. Also check its data handling terms meet your obligations, add a circuit breaker so you switch quickly, and test the failover by injecting errors in staging. What does your feature need the output to look like?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Classify LLM dependency failures and their user impact
- Set per attempt timeouts and overall deadlines from latency budgets
- Implement circuit breakers that stop calls to a failing service
- Design and evaluate fallback paths across models, providers and degraded modes
- Test failover with fault injection and monitor fallback use
Lesson plan
- 1 How LLM dependencies fail List the failure modes of model APIs and what each means for users. Start
- 2 Timeouts and deadlines Derive attempt timeouts and request deadlines from your latency budget. Start
- 3 Circuit breakers Stop sending traffic to a dependency that is clearly failing. Start
- 4 Choosing fallback paths Pick fallbacks that match the feature's value and risk. Start
- 5 Making fallbacks actually work Evaluate and adapt prompts and parsers for every fallback model. Start
- 6 Testing and monitoring failover Prove the design works before a real incident does. Start
Try asking
About this tutor
For experienced developers responsible for LLM features that must stay up when a provider has a bad day. You learn to classify failures (errors, overloads, timeouts, slow responses, quality regressions, model retirements), set per attempt timeouts and overall deadlines, use circuit breakers so you stop hammering a failing service, and design fallbacks: another model, another provider, a cached or simpler answer, a queue for later, or an honest message. The tutor is frank about the costs: fallback models behave differently and need their own evaluation, prompts may need adapting, and every extra provider brings its own data terms.
Reviews
4.7
3 ratingsSample
- Ken T.Sample
Advanced and dense, as advertised. The runbook lesson was short but useful.
- Mohammed E.Sample
We had a second provider wired up but had never evaluated it. Our parser broke on its output format in the first fault injection test. Found it in staging instead of production.
- Signe L.Sample
Circuit breaker explanation with real thresholds and probes was excellent. The point that degraded modes can beat a weaker model changed our design.
About the teacher
Keeps LLM features fast, affordable and available as traffic grows: cost, caching, retries and observability
9 tutors 347 lessons taught Sample
Most of my working life has been on platform and reliability teams, and these days I spend it on LLM features: the bills that surprise people, the 429 errors on launch day, the logs nobody can read. I teach the operational side of building with model APIs. We estimate costs before writing code, add retries that do not make outages...
See Farid's profile and tutorsMore like this
Other tutors on the same or nearby topics.