Reliability engineering for multi-agent workflows
Overview
A single agent that calls two tools is already a distributed system. Put three or four agents behind a supervisor, give them shared state and a dozen tools, and you have every failure mode of a microservice estate plus a new one: the components themselves are non-deterministic.
This is a reference architecture for engineering reliability into that kind of workflow. It is not a client story. It describes the patterns I reach for when a team asks “is this agent ready for production?” and the honest answer is “we do not know yet, because we have not defined what ready means”.
The core idea is simple. Reliability is a property you design, instrument and measure. If the only evidence you have is that the demo worked, you are hoping, not engineering. In my own practice I work on agent reliability engineering: failure-mode analysis, fallback and retry design, consistency and determinism testing, and SLOs for autonomous agents. The production side of that work is described in Reliability for AI agents in production. This piece goes one level deeper into the patterns inside the workflow itself.
Problem statement
Consider a typical multi-agent workflow: a supervisor agent receives a request, routes it to a research agent (retrieval and web tools), an action agent (internal APIs that change state) and a reviewer agent that checks the result before it goes back to the user.
Each step can fail, and the failures compound. If each of five steps succeeds 97% of the time independently, the whole task succeeds roughly 86% of the time. That arithmetic is the first thing I show teams, because it explains why a workflow made of “pretty good” agents feels unreliable in production.
The failures also look different from traditional software:
- The model returns valid prose but invalid JSON.
- A tool times out, the agent retries, and the downstream API processes the request twice.
- Two agents work from different versions of the shared state and produce contradictory answers.
- The supervisor keeps delegating to the same worker because neither recognises the task is finished.
- The context window fills with tool output and the model quietly drops the instructions it needed.
None of these raises a stack trace by default. Most of them return HTTP 200.
Engineering objectives
I frame the objectives as things we can test, not aspirations:
- Every failure mode is named. Each one has a detector, a handling policy and an owner.
- Every call is bounded. No model call, tool call or agent loop can run without a time, token, cost and step budget.
- Every output is validated. Non-deterministic outputs pass through a schema and a semantic check before anything acts on them.
- Every side effect is safe to retry or explicitly not retried. Tools are idempotent, or they are flagged and handled differently.
- Degradation is designed. When something fails, the system chooses a known fallback rather than improvising.
- Reliability is expressed as SLOs. Task success rate, latency, cost and escalation rate have targets and error budgets, and a release can be blocked on them.
Solution architecture
The architecture separates the workflow from the reliability controls around it. The agents stay focused on the task. The controls live in a thin layer that every model and tool call passes through, and in the telemetry pipeline behind it.
01 Orchestration
- Supervisor agent
- Worker agents
- Checkpointed state graph
- Step and loop limits
02 Guarded execution
- Timeouts and budgets
- Retry with backoff and jitter
- Schema validation
- Circuit breakers
- Fallback chain
03 Tools and models
- Idempotent tool adapters
- Compensating actions
- Primary and fallback models
- Response cache
04 Observability
- OpenTelemetry spans
- Token and cost telemetry
- Failure-mode tagging
- SLO dashboards
05 Assurance
- Regression eval suite
- Repeated-run consistency
- Error budget policy
- Release gate
Three design decisions shape the rest:
- The state graph is the source of truth. Agents read from and write to a checkpointed state object rather than passing free text to each other. In LangGraph this is the graph state persisted by a checkpointer, keyed by a thread identifier, which also gives us resume after failure.
- Guarded execution is a library, not a convention. If timeouts and retries are left to each agent author, they will be inconsistent. A single wrapper enforces them.
- Failure modes are first-class telemetry. Every handled failure emits a span attribute naming the failure mode. Without that, incident analysis becomes log archaeology.
Technical approach
A failure-mode taxonomy
The taxonomy is the backbone. I keep it short enough that engineers actually use the labels.
| Failure mode | What it looks like | Primary control |
|---|---|---|
| Model error | Provider 5xx, rate limit, refusal, empty completion | Retry with backoff, model fallback |
| Tool failure | Tool raises, returns error payload, or partial data | Retry if idempotent, else compensate or escalate |
| Timeout | Model or tool exceeds its deadline | Per-call timeout, overall task deadline |
| Malformed output | Invalid JSON, missing fields, wrong enum, hallucinated tool name | Structured output, schema validation, bounded repair retry |
| Loop | Supervisor re-delegates, agent repeats the same tool call | Step budget, repeated-action detector |
| Context overflow | Instructions truncated, tool output crowds out the task | Token budget per step, summarisation, trimming policy |
| State divergence | Agents act on stale or conflicting state | Single state graph, versioned writes, reviewer check |
The labels matter because they become the dimensions on your dashboards. “Agent failed” is not actionable. “Malformed output from the action agent after the model upgrade” is.
Share of handled failures by mode (illustrative)
| Malformed output | 31% | |
|---|---|---|
| Tool failure | 24% | |
| Timeout | 17% | |
| Model error | 12% | |
| Loop | 8% | |
| Context overflow | 5% | |
| State divergence | 3% |
The point of this chart is the habit, not the numbers. Once failures are tagged, a breakdown like this tells you where engineering effort pays back. In many workflows the largest bucket is not the dramatic one.
Handling non-deterministic outputs
The model will eventually return something you did not expect. The defence has three layers:
- Ask for structure. Use the provider’s structured output or tool-calling mode so the model is constrained to a schema where possible.
- Validate anyway. Parse into a typed model (Pydantic in Python) and add semantic checks the schema cannot express, such as “the referenced record ID exists” or “the amount is within the user’s limit”.
- Repair with a bound. On validation failure, retry once or twice with the validation error included in the prompt. After that, stop and fall back. Unbounded repair loops are a common source of both cost and latency incidents.
Timeouts, retries and budgets
Every call gets a deadline. Every task gets an overall budget across four dimensions: wall-clock time, tokens, cost and steps. The task budget is checked before each step, so a workflow that is already over budget does not start new work.
Retries use exponential backoff with jitter, which spreads retries out so that many clients recovering from the same outage do not hit the provider in lockstep. The AWS Builders’ Library article on timeouts, retries and backoff with jitter is still the clearest explanation of why.
Retries are only safe for operations that are idempotent. For tools with side effects I require an idempotency key derived from the task and step, so a retried “create ticket” call does not create two tickets. Where a tool cannot be made idempotent, it is marked non-retryable and failure goes straight to compensation or escalation.
A guarded call wrapper
This is the shape of the wrapper every agent uses. It is deliberately small.
import asyncioimport randomimport timefrom dataclasses import dataclass, fieldfrom typing import Awaitable, Callable, TypeVar
from pydantic import BaseModel, ValidationError
T = TypeVar("T", bound=BaseModel)
class BudgetExceeded(Exception): ...class NonRetryable(Exception): ...
@dataclassclass TaskBudget: deadline: float # absolute monotonic time max_tokens: int max_cost_usd: float max_steps: int tokens: int = 0 cost_usd: float = 0.0 steps: int = 0
def check(self) -> None: if time.monotonic() > self.deadline: raise BudgetExceeded("time") if self.tokens > self.max_tokens: raise BudgetExceeded("tokens") if self.cost_usd > self.max_cost_usd: raise BudgetExceeded("cost") if self.steps >= self.max_steps: raise BudgetExceeded("steps")
@dataclassclass CallResult: value: BaseModel | None failure_mode: str | None = None attempts: int = 0 used_fallback: bool = False meta: dict = field(default_factory=dict)
async def guarded_call( call: Callable[[str | None], Awaitable[str]], # takes optional repair hint schema: type[T], budget: TaskBudget, *, timeout_s: float = 20.0, max_attempts: int = 3, base_delay_s: float = 0.5, retryable: bool = True, fallback: Callable[[], Awaitable[T | None]] | None = None,) -> CallResult: repair_hint = None failure_mode = None for attempt in range(1, max_attempts + 1): budget.check() budget.steps += 1 remaining = budget.deadline - time.monotonic() try: raw = await asyncio.wait_for(call(repair_hint), timeout=min(timeout_s, remaining)) return CallResult(schema.model_validate_json(raw), attempts=attempt) except asyncio.TimeoutError: failure_mode = "timeout" except ValidationError as exc: failure_mode = "malformed_output" repair_hint = f"Previous output failed validation: {exc.errors()[:3]}" except NonRetryable: failure_mode = "tool_failure_non_retryable" break except Exception: failure_mode = "model_or_tool_error" if not retryable or attempt == max_attempts: break # Full jitter: sleep a random amount up to the exponential cap. cap = base_delay_s * (2 ** (attempt - 1)) await asyncio.sleep(random.uniform(0, cap))
if fallback is not None: value = await fallback() if value is not None: return CallResult(value, failure_mode, attempt, used_fallback=True) return CallResult(None, failure_mode, attempt)The timeout is the smaller of the per-call timeout and the remaining task deadline, so a late step cannot overrun the task. Validation errors feed back as a repair hint, still bounded by the attempt count. The result always carries a failure_mode for the tracing layer, and the wrapper returns rather than raises on ordinary failures, which keeps the decision about what to do next in the orchestrator.
Fallbacks and graceful degradation
When retries are exhausted, the fallback chain decides what the user gets. I design it explicitly, per task type:
- 01Primary model or tool
- 02Bounded retry with jitter
- 03Fallback model
- 04Cached or partial answer
- 05Escalate to a human
- Model fallback. A second model, possibly from another provider, for when the primary is unavailable. It must be in the regression suite too. An untested fallback is a second outage waiting to happen.
- Cached answer. For read-only, slowly changing questions, a recent validated answer is often better than no answer. Label it as cached.
- Partial result. Return what succeeded and say clearly what did not.
- Degrade to a human. For anything with side effects or low confidence, hand the task to a person with the full trace attached. Escalation is a success of the reliability design, not a failure of it, as long as it is measured.
Model and tool availability is handled by circuit breakers. After a run of failures to one dependency, the breaker opens and calls go straight to the fallback for a cool-down period, rather than every task spending its retry budget on a dependency that is known to be down. The circuit breaker pattern is well documented. The agent-specific part is choosing what “failure” means, because a 200 response with malformed output should count.
Resilience in multi-agent orchestration
Multi-agent workflows add coordination failures on top of call failures. The patterns I use:
- Supervisor with explicit termination. The supervisor must emit a structured “done” or “escalate” decision. A step budget and a repeated-action detector (same agent, same tool, same arguments, twice in a row) catch loops early.
- Checkpoints at every step. With a LangGraph checkpointer, each super-step is persisted, so a crashed worker or a redeploy resumes from the last good state rather than restarting the task and repeating side effects. See the LangGraph persistence documentation.
- Idempotent tools and compensating actions. Every state-changing tool has either an idempotency key or a defined compensating action (cancel the booking, close the ticket). This is the saga idea applied to agents; the compensating transaction pattern describes it well.
- Reviewer as a gate, not a suggestion. The reviewer agent checks outputs against the state, not against its own opinion. If the reviewer and the action agent disagree, the task escalates.
Context management and state consistency
Context overflow is a reliability problem, not just a cost problem. Each step gets a token budget. Tool outputs are trimmed or summarised before they enter the shared state, and the system instructions are pinned so they cannot be pushed out. Agents read structured fields from the state rather than re-reading the whole conversation.
State divergence is prevented by having one writer per field and versioning writes. If an agent tries to act on a state version older than the current one, the write is rejected and the step re-runs on fresh state.
Validation strategy
Reliability claims need evidence before release and in production.
Before release, the regression suite runs each scenario several times, not once, because a single pass on a non-deterministic system proves very little. Scenarios include the happy path and injected faults: provider timeouts, tool errors, malformed responses, a slow dependency that trips the circuit breaker, and a crash mid-task to prove checkpoint resume. I treat fault injection for agents the way I treated negative testing in traditional QA: the system is only as good as its behaviour on the bad path.
The release gate compares the candidate against the current baseline on task success rate across repeated runs, consistency, p95 latency, cost per task and escalation rate. A regression beyond the agreed tolerance blocks the release. The measurement side of this is covered in Measuring AI reliability.
In production, every model and tool call is a span with GenAI attributes following the OpenTelemetry GenAI semantic conventions, plus the failure-mode tag from the wrapper. Troubleshooting starts from the trace: which step failed, with which mode, after how many attempts, and did the fallback hold. Incident analysis then asks whether the failure mode was known, whether its control worked, and whether the taxonomy needs a new entry.
Metrics and measurement
Reliability becomes manageable when it is expressed as service level indicators with targets. The SLO concepts follow the Google SRE book chapter on service level objectives. The values below are example targets for illustration, not measured results; real targets depend on the task and its risk.
| SLI | Definition | Example SLO (illustrative) |
|---|---|---|
| Task success rate | Tasks completed and passing the reviewer and validation, over all tasks | 95% over 28 days |
| p95 task latency | 95th percentile wall-clock time from request to final answer | 20 s |
| Cost per task | Mean token cost per completed task | Within 1.2x of baseline |
| Escalation rate | Tasks handed to a human, over all tasks | At most 5% |
| Fallback rate | Tasks that used any fallback, over all tasks | At most 10% |
| Malformed output rate | Calls failing validation after repair, over all model calls | At most 0.5% |
The error budget is the gap between the SLO and perfection. With a 95% task success SLO, 5% of tasks may fail in the window. When the budget is being spent faster than expected, the team prioritises reliability work over features, following an error budget policy agreed in advance.
Two measurement details matter for agents. First, escalation is counted separately from failure. A task correctly handed to a human is not a silent failure, but a rising escalation rate is an early warning. Second, cost is an SLI. A retry storm or a repair loop shows up in cost per task before it shows up anywhere else.
Challenges and trade-offs
Retries versus latency and cost. Every retry improves the chance of success and worsens p95 latency and cost. Three attempts with jitter is a reasonable default for idempotent calls; for expensive model calls I often allow two.
Validation strictness versus usefulness. Strict schemas catch malformed output but also reject answers that were fine. Tune them on real traffic, and track the validation failure rate so you can see when a model change shifts it.
Fallback models versus consistency. A fallback model keeps the system up but can behave differently. Users may notice the change in tone or quality. Test the fallback on the same suite and decide which tasks it is allowed to serve.
Checkpointing versus performance and privacy. Persisting every step adds write latency and stores intermediate data, which may include personal data. Retention policies and redaction need to be part of the design.
Escalation versus autonomy. Escalating too readily makes the agent pointless; too rarely makes it dangerous. The escalation rate SLO makes that trade-off visible and negotiable rather than accidental.
Outcome and lessons learned
The outcome of this reference architecture is not a number. It is a workflow where every failure has a name, a bound, a handler and a metric, and where release decisions use evidence rather than demos.
Two lessons stand out. Start with the taxonomy, because naming failure modes changes how a team talks about reliability and the names become dashboard dimensions. And centralise guarded execution, because one wrapper beats twenty well-intentioned conventions.
The patterns for retries, timeouts and recovery are expanded in Agent failure handling patterns.
Further reading
- Reliability for AI agents in production
- Agent failure handling patterns
- Measuring AI reliability
- Google SRE book: Service Level Objectives
- Google SRE workbook: Error budget policy
- OpenTelemetry semantic conventions for generative AI
- LangGraph persistence and checkpointers
- AWS Builders’ Library: Timeouts, retries and backoff with jitter