Reliability for AI agents in production
Overview
Agents behave differently from traditional software. The same request can take a different path on the next run, call a different tool, or produce a subtly different answer. Pre-release testing is necessary, but it cannot tell you how an agent behaves on the thousandth real conversation.
I lead product engineering for a platform built to answer that question: is this agent still reliable, right now, in production? This page describes the engineering thinking behind it. The product itself is confidential, so the description stays at the level of architecture and method rather than product features.
Problem statement
Teams running agents in production needed two things that pull in different directions:
- Fast signals when something goes wrong in a live agent run, such as a failed tool call, a policy breach or a latency spike.
- Deep evaluation of quality over time, such as groundedness, consistency and drift, which is too expensive to run on every request in the hot path.
Doing both in one pipeline either slows the live signal or starves the deep evaluation. A slow alert is an alert that arrives after the user has already seen the failure. A starved evaluation is one that quietly samples less and less until it stops telling you anything.
There was a second, less obvious problem. “Reliable” meant different things to different people. Engineers thought in terms of errors and timeouts. Product owners thought in terms of correct answers. Neither view alone describes an agent that completes the right task, with the right tools, within an acceptable time and cost.
Engineering objectives
The design was shaped by a small number of objectives:
- Separate the hot path from the deep path. Live detection and quality evaluation must not compete for the same resources or fail together.
- Make checks pluggable. New failure modes appear as agents evolve. Adding a check should not mean touching transport, storage or alerting.
- Trace everything at step level. Every agent run should be reconstructable: which model, which tool, which inputs, how long, how many tokens.
- Express reliability as objectives, not opinions. Define service level objectives for autonomous agents so “good enough to keep running” has a shared, measurable meaning.
- Feed results back into release decisions. Production evidence should inform what ships next, not just what gets paged.
Solution architecture
The core decision was a two-plane architecture on Kafka. Both planes consume the same agent telemetry stream, but each scales, deploys and fails independently.
01 Instrument
- OpenTelemetry spans per agent step
- Token, cost and latency attributes
- Tool call inputs and outcomes
02 Transport
- Kafka telemetry topics
- Shared schema for both planes
03 Real-time plane
- Lightweight detectors
- Policy and error checks
- Alerting
04 Batch plane
- Quality evaluators
- Drift and regression analysis
- SLO reporting
The real-time plane runs checks that are cheap and deterministic: did a tool call fail, did the run exceed a latency budget, did an output trip a policy rule. These need to fire within seconds.
The batch plane runs checks that are expensive or need context across many runs: groundedness scoring, consistency across repeated prompts, drift in behaviour over days. These can tolerate minutes or hours of delay in exchange for depth.
Checks in both planes are packaged as detectors inside a common pipeline framework. A detector declares what telemetry it needs and what finding it emits. The framework handles consumption, retries, storage and routing.
Technical approach
Tracing by default
Agent and span-level traces are captured with OpenTelemetry. Each step of an agent run becomes a span with attributes for model, tool, token usage, cost and latency. Aligning attribute names with the OpenTelemetry GenAI semantic conventions keeps the data portable between tools, which matters when evaluating observability products such as Langfuse, Arize and the NVIDIA NeMo Agent Toolkit side by side.
Detectors as a contract
The value of a detector framework is the contract, not any single detector. A simplified, illustrative shape looks like this:
from dataclasses import dataclassfrom typing import Protocol
@dataclassclass Finding: run_id: str detector: str severity: str # e.g. "info", "warn", "critical" detail: str
class Detector(Protocol): name: str plane: str # "realtime" or "batch"
def applies_to(self, span: dict) -> bool: ... def evaluate(self, span: dict) -> Finding | None: ...Because each detector is small and declares its plane, a new failure mode can be added, tested in isolation and promoted from batch to real time once it is cheap enough.
Reliability engineering beyond monitoring
Monitoring tells you something broke. Reliability engineering asks why, and what the agent should do next time. The practice includes:
- Failure-mode analysis of agent runs: tool errors, malformed tool arguments, loops, premature termination, ungrounded answers.
- Fallback and retry design: which failures are safe to retry, which should fall back to a simpler path, and which must stop and hand over to a person.
- Consistency and determinism testing: running the same input repeatedly to measure how much behaviour varies, and whether the variance matters.
- SLOs for autonomous agents: objectives defined per agent and per task type, following the approach in the Google SRE book.
Validation strategy
A reliability platform has to be trusted before its findings are acted on. Validation works at three levels:
- Detector tests. Each detector is tested against recorded traces with known failures and known clean runs, so both missed failures and false alarms are visible.
- Replay. Captured telemetry can be replayed through both planes to confirm that a change to a detector or to the pipeline does not change findings unexpectedly.
- Plane isolation tests. Deliberately slowing or failing the batch plane confirms that real-time alerting continues unaffected. This is the property the architecture exists to protect, so it is tested directly.
Metrics and measurement
Metrics tracked in this kind of platform fall into three groups:
| Group | Examples |
|---|---|
| Run health | Task success rate, tool call error rate, retry and fallback rate, loop or step-limit terminations |
| Performance and cost | End-to-end and per-step latency, tokens per run, cost per successful task |
| Quality over time | Groundedness and consistency scores, drift against a baseline, regression after model or prompt changes |
SLOs are set on a small subset of these per agent, with error budgets that tell product owners when reliability work should take priority over new features. The platform also tracks its own health: consumer lag on each plane, detector failures and the time from event to finding.
Challenges and trade-offs
- Latency versus depth. Every check has to choose a plane. A typical trade-off is accepting a delayed groundedness score in exchange for not adding model calls to the hot path.
- Sampling. Batch evaluation of every run is rarely affordable. Sampling strategies need to keep rare but serious failures visible rather than averaging them away.
- Non-determinism. A single failed run may be noise. Detectors that alert on one bad output generate fatigue; detectors that wait for patterns can be slow. Thresholds need tuning per agent.
- Telemetry volume and sensitivity. Full prompts and tool payloads are valuable for debugging but can contain sensitive data, so what is captured, redacted and retained has to be decided up front.
- Schema discipline. Two planes sharing one stream only works if the telemetry schema is stable and versioned.
Outcome and lessons learned
The platform gives teams a live view of agent health and a slower, deeper view of quality, without either compromising the other. Detectors turn reliability expectations into explicit, testable rules, and SLOs give product owners a shared definition of “good enough to keep running”.
The main lessons:
- Separating real-time and batch evaluation is an architectural decision, and it is far cheaper to make early.
- Reliability for agents is mostly about behaviour, not uptime. The interesting failures return HTTP 200.
- Continuous evaluation is most useful when it feeds back into release decisions, closing the loop between production and pre-release testing.