Long Running Agent Tasks That Survive Failure
Build agent tasks that run for hours, checkpoint progress and resume cleanly after crashes or timeouts
A taste of a lesson
My agent processes 2,000 support tickets overnight in one conversation. Last night it stopped at ticket 1,400. What should I change first?
Stop using one conversation for the whole batch. Make each ticket a unit: a fresh, short context with the ticket and your instructions, and the result written to a table with a status. Before processing, the job checks the table and skips tickets already marked done, so a restart resumes at 1,401. Add retries with backoff for rate limits, and send tickets that fail three times to a review list. One question for you: does processing a ticket change anything outside your system, such as sending an email? That decides how careful the resume logic must be.
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Split long agent jobs into resumable units with clear inputs and outputs
- Keep progress and state outside the model with checkpoints after each unit
- Design idempotent actions and retry rules that never duplicate side effects
- Set cost, time and step budgets with progress reports and escalation points
- Plan verification by automatic checks and sampled human review
Lesson plan
- 1 How long runs fail List the failure types that only appear when agent tasks run for a long time. Start
- 2 Units of work and external state Design a job as independent units with progress recorded outside the model. Start
- 3 Checkpoints, idempotency and resume Make restarts continue where the job stopped without repeating side effects. Start
- 4 Retries, backoff and dead letters Handle transient and persistent failures without loops or lost items. Start
- 5 Budgets, progress and escalation Keep humans informed and in control during long runs. Start
- 6 Verifying a long run Check the quality of hundreds of outputs without reading every one. Start
Try asking
About this tutor
For engineers whose agents do work that takes far longer than a chat turn: migrating many files, processing large document sets, multi stage research or overnight jobs. Long runs meet failures that short runs never see: rate limits, timeouts, crashed workers, context overflow and partial results. You will learn to design tasks as resumable units, checkpoint state outside the model, make actions idempotent, handle retries and backoff, report progress to humans, and decide when the agent should stop and ask. We finish with a review of a long run design from start to recovery.
Reviews
4.7
3 ratingsSample
- Grace N.Sample
The 'what if it crashes right here' habit is now in every design review I run. Our document migration resumed cleanly after a deploy restart the week after I finished.
- Pavel K.Sample
Strong on state and idempotency. The verification lesson could use one more worked example, but I changed our sampling plan because of it.
- Nadia F.Sample
Treating long agent runs as batch jobs with a model inside was the key reframe for me.
About the teacher
I teach how AI agents are built: the loop, the tools, the memory, and when a plain workflow is the better choice
9 tutors 310 lessons taught Sample
I build and teach the inner workings of AI agents. Most of my working life has been spent on backend systems, so I approach agents the way I approach any distributed system: what runs, in what order, what can fail, and what it costs. I like to start every topic with a drawing of the loop on a whiteboard and...
See Hiroshi's profile and tutorsMore like this
Other tutors on the same or nearby topics.