Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Long Running Agent Tasks That Survive Failure

Build agent tasks that run for hours, checkpoint progress and resume cleanly after crashes or timeouts

By Hiroshi Tanabe Advanced AI agents 4.7(3) 35 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $13 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Long Running Agent Tasks That Survive Failure AI tutor following Hiroshi Tanabe's plan
Student:

My agent processes 2,000 support tickets overnight in one conversation. Last night it stopped at ticket 1,400. What should I change first?

Tutor:

Stop using one conversation for the whole batch. Make each ticket a unit: a fresh, short context with the ticket and your instructions, and the result written to a table with a status. Before processing, the job checks the table and skips tickets already marked done, so a restart resumes at 1,401. Add retries with backoff for rate limits, and send tickets that fail three times to a review list. One question for you: does processing a ticket change anything outside your system, such as sending an email? That decides how careful the resume logic must be.

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Split long agent jobs into resumable units with clear inputs and outputs
  • Keep progress and state outside the model with checkpoints after each unit
  • Design idempotent actions and retry rules that never duplicate side effects
  • Set cost, time and step budgets with progress reports and escalation points
  • Plan verification by automatic checks and sampled human review

Lesson plan

6 lessons. Pick one to start there.

  1. 1 How long runs fail List the failure types that only appear when agent tasks run for a long time. Start
  2. 2 Units of work and external state Design a job as independent units with progress recorded outside the model. Start
  3. 3 Checkpoints, idempotency and resume Make restarts continue where the job stopped without repeating side effects. Start
  4. 4 Retries, backoff and dead letters Handle transient and persistent failures without loops or lost items. Start
  5. 5 Budgets, progress and escalation Keep humans informed and in control during long runs. Start
  6. 6 Verifying a long run Check the quality of hundreds of outputs without reading every one. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For engineers whose agents do work that takes far longer than a chat turn: migrating many files, processing large document sets, multi stage research or overnight jobs. Long runs meet failures that short runs never see: rate limits, timeouts, crashed workers, context overflow and partial results. You will learn to design tasks as resumable units, checkpoint state outside the model, make actions idempotent, handle retries and backoff, report progress to humans, and decide when the agent should stop and ask. We finish with a review of a long run design from start to recovery.

Reviews

4.7

3 ratingsSample

  • Grace N.Sample

    The 'what if it crashes right here' habit is now in every design review I run. Our document migration resumed cleanly after a deploy restart the week after I finished.

  • Pavel K.Sample

    Strong on state and idempotency. The verification lesson could use one more worked example, but I changed our sampling plan because of it.

  • Nadia F.Sample

    Treating long agent runs as batch jobs with a model inside was the key reframe for me.

About the teacher

Hiroshi Tanabe

I teach how AI agents are built: the loop, the tools, the memory, and when a plain workflow is the better choice

9 tutors 4.5(18) 310 lessons taught Sample

I build and teach the inner workings of AI agents. Most of my working life has been spent on backend systems, so I approach agents the way I approach any distributed system: what runs, in what order, what can fail, and what it costs. I like to start every topic with a drawing of the loop on a whiteboard and...

See Hiroshi's profile and tutors