Skip to content
Satya Prakash Solanki

Reliability engineering for multi-agent workflows

Overview

A single agent that calls two tools is already a distributed system. Put three or four agents behind a supervisor, give them shared state and a dozen tools, and you have every failure mode of a microservice estate plus a new one: the components themselves are non-deterministic.

This is a reference architecture for engineering reliability into that kind of workflow. It is not a client story. It describes the patterns I reach for when a team asks “is this agent ready for production?” and the honest answer is “we do not know yet, because we have not defined what ready means”.

The core idea is simple. Reliability is a property you design, instrument and measure. If the only evidence you have is that the demo worked, you are hoping, not engineering. In my own practice I work on agent reliability engineering: failure-mode analysis, fallback and retry design, consistency and determinism testing, and SLOs for autonomous agents. The production side of that work is described in Reliability for AI agents in production. This piece goes one level deeper into the patterns inside the workflow itself.

Problem statement

Consider a typical multi-agent workflow: a supervisor agent receives a request, routes it to a research agent (retrieval and web tools), an action agent (internal APIs that change state) and a reviewer agent that checks the result before it goes back to the user.

Each step can fail, and the failures compound. If each of five steps succeeds 97% of the time independently, the whole task succeeds roughly 86% of the time. That arithmetic is the first thing I show teams, because it explains why a workflow made of “pretty good” agents feels unreliable in production.

The failures also look different from traditional software:

  • The model returns valid prose but invalid JSON.
  • A tool times out, the agent retries, and the downstream API processes the request twice.
  • Two agents work from different versions of the shared state and produce contradictory answers.
  • The supervisor keeps delegating to the same worker because neither recognises the task is finished.
  • The context window fills with tool output and the model quietly drops the instructions it needed.

None of these raises a stack trace by default. Most of them return HTTP 200.

Engineering objectives

I frame the objectives as things we can test, not aspirations:

  1. Every failure mode is named. Each one has a detector, a handling policy and an owner.
  2. Every call is bounded. No model call, tool call or agent loop can run without a time, token, cost and step budget.
  3. Every output is validated. Non-deterministic outputs pass through a schema and a semantic check before anything acts on them.
  4. Every side effect is safe to retry or explicitly not retried. Tools are idempotent, or they are flagged and handled differently.
  5. Degradation is designed. When something fails, the system chooses a known fallback rather than improvising.
  6. Reliability is expressed as SLOs. Task success rate, latency, cost and escalation rate have targets and error budgets, and a release can be blocked on them.

Solution architecture

The architecture separates the workflow from the reliability controls around it. The agents stay focused on the task. The controls live in a thin layer that every model and tool call passes through, and in the telemetry pipeline behind it.

01 Orchestration

  • Supervisor agent
  • Worker agents
  • Checkpointed state graph
  • Step and loop limits

02 Guarded execution

  • Timeouts and budgets
  • Retry with backoff and jitter
  • Schema validation
  • Circuit breakers
  • Fallback chain

03 Tools and models

  • Idempotent tool adapters
  • Compensating actions
  • Primary and fallback models
  • Response cache

04 Observability

  • OpenTelemetry spans
  • Token and cost telemetry
  • Failure-mode tagging
  • SLO dashboards

05 Assurance

  • Regression eval suite
  • Repeated-run consistency
  • Error budget policy
  • Release gate
Figure 1. Reference architecture. Reliability controls wrap every model and tool call; telemetry feeds SLOs and the release gate.

Three design decisions shape the rest:

  • The state graph is the source of truth. Agents read from and write to a checkpointed state object rather than passing free text to each other. In LangGraph this is the graph state persisted by a checkpointer, keyed by a thread identifier, which also gives us resume after failure.
  • Guarded execution is a library, not a convention. If timeouts and retries are left to each agent author, they will be inconsistent. A single wrapper enforces them.
  • Failure modes are first-class telemetry. Every handled failure emits a span attribute naming the failure mode. Without that, incident analysis becomes log archaeology.

Technical approach

A failure-mode taxonomy

The taxonomy is the backbone. I keep it short enough that engineers actually use the labels.

Failure mode What it looks like Primary control
Model error Provider 5xx, rate limit, refusal, empty completion Retry with backoff, model fallback
Tool failure Tool raises, returns error payload, or partial data Retry if idempotent, else compensate or escalate
Timeout Model or tool exceeds its deadline Per-call timeout, overall task deadline
Malformed output Invalid JSON, missing fields, wrong enum, hallucinated tool name Structured output, schema validation, bounded repair retry
Loop Supervisor re-delegates, agent repeats the same tool call Step budget, repeated-action detector
Context overflow Instructions truncated, tool output crowds out the task Token budget per step, summarisation, trimming policy
State divergence Agents act on stale or conflicting state Single state graph, versioned writes, reviewer check

The labels matter because they become the dimensions on your dashboards. “Agent failed” is not actionable. “Malformed output from the action agent after the model upgrade” is.

Share of handled failures by mode (illustrative)

Share of handled failures by mode (illustrative)
Malformed output31%
Tool failure24%
Timeout17%
Model error12%
Loop8%
Context overflow5%
State divergence3%
Figure 2. Illustrative data, not a measured result. It shows the shape of a failure-mode breakdown, not real proportions.

The point of this chart is the habit, not the numbers. Once failures are tagged, a breakdown like this tells you where engineering effort pays back. In many workflows the largest bucket is not the dramatic one.

Handling non-deterministic outputs

The model will eventually return something you did not expect. The defence has three layers:

  1. Ask for structure. Use the provider’s structured output or tool-calling mode so the model is constrained to a schema where possible.
  2. Validate anyway. Parse into a typed model (Pydantic in Python) and add semantic checks the schema cannot express, such as “the referenced record ID exists” or “the amount is within the user’s limit”.
  3. Repair with a bound. On validation failure, retry once or twice with the validation error included in the prompt. After that, stop and fall back. Unbounded repair loops are a common source of both cost and latency incidents.

Timeouts, retries and budgets

Every call gets a deadline. Every task gets an overall budget across four dimensions: wall-clock time, tokens, cost and steps. The task budget is checked before each step, so a workflow that is already over budget does not start new work.

Retries use exponential backoff with jitter, which spreads retries out so that many clients recovering from the same outage do not hit the provider in lockstep. The AWS Builders’ Library article on timeouts, retries and backoff with jitter is still the clearest explanation of why.

Retries are only safe for operations that are idempotent. For tools with side effects I require an idempotency key derived from the task and step, so a retried “create ticket” call does not create two tickets. Where a tool cannot be made idempotent, it is marked non-retryable and failure goes straight to compensation or escalation.

A guarded call wrapper

This is the shape of the wrapper every agent uses. It is deliberately small.

guarded_call.py
import asyncio
import random
import time
from dataclasses import dataclass, field
from typing import Awaitable, Callable, TypeVar
from pydantic import BaseModel, ValidationError
T = TypeVar("T", bound=BaseModel)
class BudgetExceeded(Exception): ...
class NonRetryable(Exception): ...
@dataclass
class TaskBudget:
deadline: float # absolute monotonic time
max_tokens: int
max_cost_usd: float
max_steps: int
tokens: int = 0
cost_usd: float = 0.0
steps: int = 0
def check(self) -> None:
if time.monotonic() > self.deadline:
raise BudgetExceeded("time")
if self.tokens > self.max_tokens:
raise BudgetExceeded("tokens")
if self.cost_usd > self.max_cost_usd:
raise BudgetExceeded("cost")
if self.steps >= self.max_steps:
raise BudgetExceeded("steps")
@dataclass
class CallResult:
value: BaseModel | None
failure_mode: str | None = None
attempts: int = 0
used_fallback: bool = False
meta: dict = field(default_factory=dict)
async def guarded_call(
call: Callable[[str | None], Awaitable[str]], # takes optional repair hint
schema: type[T],
budget: TaskBudget,
*,
timeout_s: float = 20.0,
max_attempts: int = 3,
base_delay_s: float = 0.5,
retryable: bool = True,
fallback: Callable[[], Awaitable[T | None]] | None = None,
) -> CallResult:
repair_hint = None
failure_mode = None
for attempt in range(1, max_attempts + 1):
budget.check()
budget.steps += 1
remaining = budget.deadline - time.monotonic()
try:
raw = await asyncio.wait_for(call(repair_hint), timeout=min(timeout_s, remaining))
return CallResult(schema.model_validate_json(raw), attempts=attempt)
except asyncio.TimeoutError:
failure_mode = "timeout"
except ValidationError as exc:
failure_mode = "malformed_output"
repair_hint = f"Previous output failed validation: {exc.errors()[:3]}"
except NonRetryable:
failure_mode = "tool_failure_non_retryable"
break
except Exception:
failure_mode = "model_or_tool_error"
if not retryable or attempt == max_attempts:
break
# Full jitter: sleep a random amount up to the exponential cap.
cap = base_delay_s * (2 ** (attempt - 1))
await asyncio.sleep(random.uniform(0, cap))
if fallback is not None:
value = await fallback()
if value is not None:
return CallResult(value, failure_mode, attempt, used_fallback=True)
return CallResult(None, failure_mode, attempt)

The timeout is the smaller of the per-call timeout and the remaining task deadline, so a late step cannot overrun the task. Validation errors feed back as a repair hint, still bounded by the attempt count. The result always carries a failure_mode for the tracing layer, and the wrapper returns rather than raises on ordinary failures, which keeps the decision about what to do next in the orchestrator.

Fallbacks and graceful degradation

When retries are exhausted, the fallback chain decides what the user gets. I design it explicitly, per task type:

  1. 01Primary model or tool
  2. 02Bounded retry with jitter
  3. 03Fallback model
  4. 04Cached or partial answer
  5. 05Escalate to a human
Figure 3. The failure-handling path. Each stage is bounded; the last stage is a designed outcome, not an accident.
  • Model fallback. A second model, possibly from another provider, for when the primary is unavailable. It must be in the regression suite too. An untested fallback is a second outage waiting to happen.
  • Cached answer. For read-only, slowly changing questions, a recent validated answer is often better than no answer. Label it as cached.
  • Partial result. Return what succeeded and say clearly what did not.
  • Degrade to a human. For anything with side effects or low confidence, hand the task to a person with the full trace attached. Escalation is a success of the reliability design, not a failure of it, as long as it is measured.

Model and tool availability is handled by circuit breakers. After a run of failures to one dependency, the breaker opens and calls go straight to the fallback for a cool-down period, rather than every task spending its retry budget on a dependency that is known to be down. The circuit breaker pattern is well documented. The agent-specific part is choosing what “failure” means, because a 200 response with malformed output should count.

Resilience in multi-agent orchestration

Multi-agent workflows add coordination failures on top of call failures. The patterns I use:

  • Supervisor with explicit termination. The supervisor must emit a structured “done” or “escalate” decision. A step budget and a repeated-action detector (same agent, same tool, same arguments, twice in a row) catch loops early.
  • Checkpoints at every step. With a LangGraph checkpointer, each super-step is persisted, so a crashed worker or a redeploy resumes from the last good state rather than restarting the task and repeating side effects. See the LangGraph persistence documentation.
  • Idempotent tools and compensating actions. Every state-changing tool has either an idempotency key or a defined compensating action (cancel the booking, close the ticket). This is the saga idea applied to agents; the compensating transaction pattern describes it well.
  • Reviewer as a gate, not a suggestion. The reviewer agent checks outputs against the state, not against its own opinion. If the reviewer and the action agent disagree, the task escalates.

Context management and state consistency

Context overflow is a reliability problem, not just a cost problem. Each step gets a token budget. Tool outputs are trimmed or summarised before they enter the shared state, and the system instructions are pinned so they cannot be pushed out. Agents read structured fields from the state rather than re-reading the whole conversation.

State divergence is prevented by having one writer per field and versioning writes. If an agent tries to act on a state version older than the current one, the write is rejected and the step re-runs on fresh state.

Validation strategy

Reliability claims need evidence before release and in production.

Before release, the regression suite runs each scenario several times, not once, because a single pass on a non-deterministic system proves very little. Scenarios include the happy path and injected faults: provider timeouts, tool errors, malformed responses, a slow dependency that trips the circuit breaker, and a crash mid-task to prove checkpoint resume. I treat fault injection for agents the way I treated negative testing in traditional QA: the system is only as good as its behaviour on the bad path.

The release gate compares the candidate against the current baseline on task success rate across repeated runs, consistency, p95 latency, cost per task and escalation rate. A regression beyond the agreed tolerance blocks the release. The measurement side of this is covered in Measuring AI reliability.

In production, every model and tool call is a span with GenAI attributes following the OpenTelemetry GenAI semantic conventions, plus the failure-mode tag from the wrapper. Troubleshooting starts from the trace: which step failed, with which mode, after how many attempts, and did the fallback hold. Incident analysis then asks whether the failure mode was known, whether its control worked, and whether the taxonomy needs a new entry.

Metrics and measurement

Reliability becomes manageable when it is expressed as service level indicators with targets. The SLO concepts follow the Google SRE book chapter on service level objectives. The values below are example targets for illustration, not measured results; real targets depend on the task and its risk.

SLI Definition Example SLO (illustrative)
Task success rate Tasks completed and passing the reviewer and validation, over all tasks 95% over 28 days
p95 task latency 95th percentile wall-clock time from request to final answer 20 s
Cost per task Mean token cost per completed task Within 1.2x of baseline
Escalation rate Tasks handed to a human, over all tasks At most 5%
Fallback rate Tasks that used any fallback, over all tasks At most 10%
Malformed output rate Calls failing validation after repair, over all model calls At most 0.5%

The error budget is the gap between the SLO and perfection. With a 95% task success SLO, 5% of tasks may fail in the window. When the budget is being spent faster than expected, the team prioritises reliability work over features, following an error budget policy agreed in advance.

Two measurement details matter for agents. First, escalation is counted separately from failure. A task correctly handed to a human is not a silent failure, but a rising escalation rate is an early warning. Second, cost is an SLI. A retry storm or a repair loop shows up in cost per task before it shows up anywhere else.

Challenges and trade-offs

Retries versus latency and cost. Every retry improves the chance of success and worsens p95 latency and cost. Three attempts with jitter is a reasonable default for idempotent calls; for expensive model calls I often allow two.

Validation strictness versus usefulness. Strict schemas catch malformed output but also reject answers that were fine. Tune them on real traffic, and track the validation failure rate so you can see when a model change shifts it.

Fallback models versus consistency. A fallback model keeps the system up but can behave differently. Users may notice the change in tone or quality. Test the fallback on the same suite and decide which tasks it is allowed to serve.

Checkpointing versus performance and privacy. Persisting every step adds write latency and stores intermediate data, which may include personal data. Retention policies and redaction need to be part of the design.

Escalation versus autonomy. Escalating too readily makes the agent pointless; too rarely makes it dangerous. The escalation rate SLO makes that trade-off visible and negotiable rather than accidental.

Outcome and lessons learned

The outcome of this reference architecture is not a number. It is a workflow where every failure has a name, a bound, a handler and a metric, and where release decisions use evidence rather than demos.

Two lessons stand out. Start with the taxonomy, because naming failure modes changes how a team talks about reliability and the names become dashboard dimensions. And centralise guarded execution, because one wrapper beats twenty well-intentioned conventions.

The patterns for retries, timeouts and recovery are expanded in Agent failure handling patterns.

Further reading

Try “evaluation”, “red-teaming”, “governance” or “agents”.