Skip to content
Satya Prakash Solanki

Most agent failure handling I review starts life as a try block around the model call with a retry loop inside it. It works in testing. In production it retries things that should never be retried, waits far too long for things that are already dead, and turns a provider blip into a cost spike.

Failure handling for agents uses the same building blocks as any distributed system. What changes is that the agent’s outputs are non-deterministic, its tools often have side effects, and every retry costs tokens. This article covers the patterns I use and, just as important, when not to use them.

Start by classifying the failure

The right response depends on what failed. Before choosing a pattern, the handler needs to know which of these it is looking at:

Class Examples Retry?
Transient infrastructure Provider 5xx, rate limit, connection reset Yes, with backoff and jitter
Timeout Model or tool exceeds deadline Only if the operation is idempotent
Malformed output Invalid JSON, missing field, unknown tool name A bounded repair retry, then fall back
Permanent error Auth failure, 4xx validation error, policy refusal No. Fall back or escalate
Side-effecting tool failure Payment, ticket creation, email send Only with an idempotency key
Logical failure Loop, wrong plan, contradictory state No. Stop, checkpoint, escalate

The last row is the one teams miss. A loop is not a transient error, and retrying the step that loops simply buys another lap.

Timeouts: every call, and the whole task

Every model call and tool call needs a deadline. So does the task as a whole. Without a task deadline, three steps that each stay inside their own timeout can still produce an answer long after the user has given up.

I propagate a single task deadline down to each call and use whichever is smaller: the per-call timeout or the time remaining. That makes the late steps of a slow task fail fast rather than starting expensive work they cannot finish.

Timeouts also need a decision about what happens to the work that was in flight. A timed-out tool call may still complete on the server. If the tool is not idempotent, retrying after a timeout is how you get duplicate orders.

Retries: backoff, jitter and a hard cap

When a retry is appropriate, use exponential backoff with jitter and a small maximum attempt count. Backoff gives the dependency time to recover; jitter stops many clients retrying in lockstep. The AWS Builders’ Library article on timeouts, retries and backoff with jitter explains the reasoning well.

retry.py
import asyncio
import random
RETRYABLE = (TimeoutError, ConnectionError, RateLimited, ProviderUnavailable)
async def with_retry(fn, *, attempts: int = 3, base: float = 0.5, cap: float = 8.0):
for attempt in range(1, attempts + 1):
try:
return await fn()
except RETRYABLE:
if attempt == attempts:
raise
# Full jitter: random sleep between 0 and the exponential ceiling.
await asyncio.sleep(random.uniform(0, min(cap, base * 2 ** (attempt - 1))))

Two details matter. The retryable exceptions are listed explicitly, so permanent errors are never retried by accident. And attempts is small. If three attempts have failed, a fourth rarely succeeds and always costs.

Repairing malformed output

Malformed output is a special case. Retrying the same prompt often produces the same mistake. A repair retry includes the validation error in the next attempt so the model can correct it:

repair.py
from pydantic import BaseModel, ValidationError
async def call_structured(llm, prompt: str, schema: type[BaseModel], max_repairs: int = 1):
message = prompt
for _ in range(max_repairs + 1):
raw = await llm(message)
try:
return schema.model_validate_json(raw)
except ValidationError as exc:
message = (
f"{prompt}\n\nYour previous answer was invalid: {exc.errors()[:3]}\n"
"Return only JSON that matches the schema."
)
raise MalformedOutput(schema.__name__)

I keep max_repairs at one or two. Beyond that, the problem is usually the prompt, the schema or the model, and more retries just hide it.

When retries make things worse

Retries are the most common reliability control and the most common cause of reliability incidents. Three failure patterns come up repeatedly.

Non-idempotent tools. An agent calls “create refund”, the call times out, the agent retries, and the customer gets two refunds. The fix is an idempotency key derived from the task and step, passed to the tool and enforced by the downstream system. If the downstream system cannot enforce it, the tool is marked non-retryable and failures go to compensation or a human.

idempotent_tool.py
import hashlib
def idempotency_key(task_id: str, step: int, tool: str, args: dict) -> str:
payload = f"{task_id}:{step}:{tool}:{sorted(args.items())}"
return hashlib.sha256(payload.encode()).hexdigest()
async def create_refund(api, task_id: str, step: int, order_id: str, amount: float):
key = idempotency_key(task_id, step, "create_refund", {"order_id": order_id, "amount": amount})
return await api.post("/refunds", json={"order_id": order_id, "amount": amount},
headers={"Idempotency-Key": key})

Retry storms. Retries multiply across layers. If the HTTP client retries three times, the tool wrapper retries three times and the agent framework retries the step three times, one failing call becomes 27. Multiply that across every concurrent task during a provider incident and you are now part of the outage. Retry at one layer only, and make the others fail fast.

Cost explosions. A retry on a long-context model call repeats the full prompt. A repair loop that never converges, or an agent that keeps re-planning, burns tokens with nothing to show. Cost needs a budget like time does, checked before each step.

Circuit breakers

When a dependency is clearly down, retrying every request against it wastes each task’s budget and adds load to a service trying to recover. A circuit breaker tracks recent failures per dependency and, past a threshold, short-circuits calls straight to the fallback for a cool-down period. After the cool-down it lets a trial call through to test recovery. The circuit breaker pattern is well established.

breaker.py
import time
class CircuitBreaker:
def __init__(self, threshold: int = 5, cooldown_s: float = 30.0):
self.threshold, self.cooldown_s = threshold, cooldown_s
self.failures, self.opened_at = 0, None
def allow(self) -> bool:
if self.opened_at is None:
return True
# Half-open: allow a trial call once the cool-down has passed.
return time.monotonic() - self.opened_at >= self.cooldown_s
def record(self, ok: bool) -> None:
if ok:
self.failures, self.opened_at = 0, None
else:
self.failures += 1
if self.failures >= self.threshold:
self.opened_at = time.monotonic()

The agent-specific decision is what counts as a failure. A provider that returns 200 with malformed output every time is down for your purposes, even though its status page is green. I count validation failures after repair as breaker failures.

Step budgets and loop detection

Agents that plan and act in a loop need a hard cap on steps. LangGraph enforces this with recursion_limit in the run config and raises GraphRecursionError when the graph exceeds it. That is the backstop. I add a cheaper, earlier check: if the same agent calls the same tool with the same arguments twice in a row, the step is almost certainly looping, and the orchestrator should stop and escalate rather than wait for the limit.

Step budgets sit alongside time, token and cost budgets. Each is checked before a step starts, and exceeding any of them ends the task through the fallback path.

Fallbacks: design the order

A fallback is only useful if it is chosen deliberately and tested. I write the fallback order down per task type:

  1. 01Primary call
  2. 02Retry (if retryable)
  3. 03Fallback model or tool
  4. 04Cached or partial answer
  5. 05Human escalation
Figure 1. A typical fallback order. Each step is bounded, and each is exercised in testing.

A fallback model must pass the same regression suite as the primary, or it is an untested system that only runs during incidents. A cached answer must be labelled as such. A human escalation needs a monitored queue and the full trace attached, or it is just a slower way of dropping the task.

Checkpoint and resume

For long-running or multi-agent workflows, the most useful recovery pattern is not a retry at all. It is resuming from the last good state.

LangGraph persists graph state at each step through a checkpointer, keyed by a thread_id. If a worker crashes or a deployment restarts mid-task, invoking the graph again on the same thread continues from the latest checkpoint instead of starting over and repeating side effects. The LangGraph persistence documentation covers the in-memory saver for development and database-backed savers such as Postgres for production.

resume.py
from langgraph.checkpoint.postgres import PostgresSaver
from langgraph.errors import GraphRecursionError
with PostgresSaver.from_conn_string(DB_URL) as checkpointer:
checkpointer.setup()
graph = builder.compile(checkpointer=checkpointer)
config = {"configurable": {"thread_id": task_id}, "recursion_limit": 25}
try:
result = graph.invoke(inputs, config)
except GraphRecursionError:
escalate(task_id, reason="step_budget_exceeded")
except TransientInfraError:
# Resume from the last checkpoint rather than restarting the task.
result = graph.invoke(None, config)

Checkpointing does not remove the need for idempotent tools. A step that performed its side effect and then crashed before the checkpoint was written will run again on resume. Idempotency keys and checkpoints work together.

Where a side effect cannot be undone by retrying, define a compensating action instead: cancel the booking, void the refund, close the ticket. The compensating transaction pattern is the saga idea, and it maps cleanly onto agent workflows.

Make every handled failure visible

Handled failures are still failures. If the retry succeeded, the user never knew, but the dependency is degrading and your cost went up. Every handler should emit the failure class, attempt count and whether a fallback was used as attributes on the trace span. That is what turns failure handling from a silent safety net into something you can measure, which is the subject of Measuring AI reliability.

Further reading

AI reliability engineering

Try “evaluation”, “red-teaming”, “governance” or “agents”.