Evaluation and observability for agentic workflows
Overview
An agentic workflow is not one model call. It is a plan, a sequence of tool calls, possibly a hand-off between agents, some retrieval, and a final answer. Any of those steps can be wrong while the final answer still looks plausible. That is what makes agents hard to test: the surface looks fine and the failure is three spans deep.
This page describes a reference architecture for an evaluation and observability platform for agentic systems. It covers what to measure at each step, how to trace an agent run so every score can be tied back to the step that produced it, how to run evaluation continuously across model, prompt and configuration changes, and how to turn the results into release gates.
It is a reference architecture, not a client story. It draws on patterns from my own practice, notably separating real-time and batch evaluation into two planes and enforcing evaluation results as release gates. Thresholds and numbers in the examples are illustrative.
Problem statement
Teams usually start with a demo that worked and a dashboard of latency and token counts. Neither answers the questions that matter before a release:
- Did the agent choose the right tools, with the right arguments, in a sensible order?
- Was the final answer correct, relevant and grounded in what the tools actually returned?
- When two agents collaborated, did the hand-off carry the right context, or did one agent quietly redo or undo the other’s work?
- Is this build better or worse than the last one, and on which kinds of task?
- When it fails in production, can we find the failing step without replaying the whole conversation by hand?
Agents make these questions hard. Outputs are probabilistic, so one run proves little. Trajectories vary, so there is often more than one correct path. And quality depends on the prompt, tool schemas, retrieval index, orchestration code and model version. A change to any of them is a change to the product.
Engineering objectives
I set the objectives as testable statements, the same way I would for any other system under test:
- Every score is traceable. Any evaluation result links to the trace, span and dataset item that produced it.
- Evaluate per layer, not just the end. Retrieval, plan, tool use, final answer and safety each get their own signals.
- Metrics follow risk. The metric set and thresholds for a use case are chosen from its risk tier, not from whatever a library offers by default.
- Regression is the default question. Every model, prompt, tool or configuration change runs the same suites against a versioned baseline.
- Production feeds evaluation. Sampled live traffic is scored online, and failures flow back into the offline datasets.
- Fast signals stay fast. Expensive judge-based scoring never sits in the request path.
Solution architecture
The platform has four groups of components: instrumentation, transport, evaluation and decision.
01 Instrument
- OpenTelemetry SDK in agent runtime
- GenAI span and attribute conventions
- Tool and retrieval wrappers
02 Transport
- Collector with sampling
- Event stream (Kafka topics)
- Trace store
03 Evaluate
- Real-time detectors
- Batch scorers and judges
- Offline suite runner
04 Decide
- Release gate in CI
- Alerts and SLOs
- Review queue and datasets
Tracing model
Everything starts with a consistent span model. I follow the OpenTelemetry semantic conventions for generative AI, which define spans for agent invocation, tool execution and model inference. The conventions are still marked as in development and are now maintained in a dedicated repository, so I pin the version I instrument against and treat upgrades as a change like any other.
A single agent run produces a tree like this:
invoke_agent support_agent gen_ai.operation.name=invoke_agent├── chat gpt-x gen_ai.usage.input_tokens / output_tokens├── execute_tool search_orders gen_ai.tool.name, gen_ai.tool.call.id│ └── (downstream HTTP/DB spans)├── chat gpt-x├── invoke_agent refunds_agent hand-off to a second agent│ ├── execute_tool get_refund_policy│ └── chat gpt-x└── chat gpt-x final answerThe span names follow the convention formats: invoke_agent {gen_ai.agent.name}, execute_tool {gen_ai.tool.name} and {gen_ai.operation.name} {gen_ai.request.model} for inference. Three practical rules make the tree useful for evaluation:
- Correlate by conversation.
gen_ai.conversation.idties multi-turn sessions together so a judge can see context, not just one turn. - Record tool arguments and results deliberately.
gen_ai.tool.call.argumentsandgen_ai.tool.call.resultare what tool-use scorers need, but they can contain personal data. I capture them behind a redaction step and an explicit opt-in per tool. - Add evaluation context as your own attributes. Dataset item id, suite version and prompt version go on the root span under a private namespace, so offline runs and live traffic share one schema.
Two evaluation planes
Fast signals and deep evaluation pull in different directions. A failed tool call or a policy breach should alert in seconds. Groundedness or multi-agent coordination quality needs a judge model, retrieved context and sometimes the whole conversation, which is too slow and too expensive for the hot path.
The pattern I use is two planes over one stream. Both consume the same trace events from the event bus, but they scale and fail independently.
| Plane | Runs on | Typical checks | Latency budget |
|---|---|---|---|
| Real-time | Every trace, streaming | Tool errors, schema violations, loop detection, guardrail hits, latency and cost budgets | Seconds |
| Batch | Sampled traces and full offline suites | Groundedness, correctness against reference, plan quality, coordination, judge-based rubrics | Minutes to hours |
Checks are packaged as detectors behind a common interface, so a new failure mode is a new detector, not a change to transport or storage. The same scorer code runs on live samples and on golden datasets; if offline and online used different scorers, their numbers could not be compared.
Technical approach
What to measure, layer by layer
| Layer | Question | Example metrics | Scoring method |
|---|---|---|---|
| Retrieval | Did we fetch the right evidence? | Context precision, context recall | Reference-based, LLM-assisted |
| Plan | Was the plan sensible and followed? | Plan quality, plan adherence, step efficiency | Judge with rubric |
| Tool use | Right tool, right arguments, right order? | Tool selection accuracy, argument correctness, sequence match, tool error rate | Mostly deterministic |
| Coordination | Did agents hand off cleanly? | Hand-off completeness, redundant work, conflicting actions | Trace analysis plus judge |
| Final answer | Is it correct, relevant and grounded? | Correctness, answer relevancy, faithfulness, hallucination rate | Reference-based and judge |
| Workflow | Did the user’s goal get done? | Task completion, goal accuracy | Reference-based where possible |
| Safety | Did it stay within bounds? | Policy violation rate, PII leakage, prompt-injection resistance | Classifiers, Promptfoo suites |
| Operations | Is it affordable and fast? | p50/p95 latency, tokens per task, cost per task, tool calls per task | Span arithmetic |
DeepEval separates trajectory metrics (task completion, step efficiency, plan adherence and plan quality) from component metrics (tool correctness, argument correctness), alongside RAG metrics such as faithfulness and contextual precision. RAGAS offers tool call accuracy, tool call F1 and agent goal accuracy for agentic workflows. I wrap them behind one scorer interface so the gate does not care which library produced a number.
Hallucination detection deserves a specific note. For agents, I split it into two checks: claims in the answer that are not supported by any tool result or retrieved context (a faithfulness problem), and claims that a tool was called or an action was taken when the trace shows it was not. The second kind is cheap to detect from spans and surprisingly common.
Choosing metrics by risk
Not every use case needs every metric. I select from the table above using the use case’s risk tier, which comes from the governance process rather than from the engineering team’s preference.
| Risk tier | Example use case | Required metrics | Human review |
|---|---|---|---|
| Low | Internal drafting assistant | Relevancy, safety baseline, cost | Spot checks |
| Medium | Customer-facing Q&A with tools | Above plus faithfulness, tool correctness, task completion | Sampled weekly |
| High | Agent that changes records or money | Above plus argument correctness, sequence match, idempotency checks, adversarial suite | Every failed gate, calibrated judges |
The tier also sets how strict the gate is. A low-tier assistant can ship with a small regression on relevancy if cost drops; a high-tier agent cannot ship with any regression on argument correctness.
Datasets, golden sets and regression suites
Datasets are versioned artefacts with an owner, like code. I keep four kinds:
- Golden set. Curated tasks with reference answers, expected tool calls and acceptable alternative trajectories. Small, high quality, reviewed by domain experts.
- Regression set. Every production failure that was triaged as a real defect becomes a case here, with the trace that exposed it.
- Adversarial set. Prompt injection, jailbreak and tool-misuse probes, maintained alongside the security test suites.
- Benchmark slices. Stratified samples by intent, language and tool, so a score can be reported per slice and not only as an average.
The suite configuration ties datasets, metrics and thresholds together:
suite: support-agentrisk_tier: mediumdatasets: - name: golden/support-v7 # pinned version - name: regression/support-2026q3 - name: adversarial/injection-v3runs_per_item: 3 # agents are stochastic; score the distributionjudge: model: judge-model-pinned rubric_version: faithfulness-r4 calibration_set: human-labels/support-v2metrics: - id: task_completion threshold: 0.85 # example value max_regression: 0.02 - id: tool_correctness threshold: 0.95 max_regression: 0.0 - id: argument_correctness threshold: 0.92 max_regression: 0.01 - id: faithfulness threshold: 0.90 max_regression: 0.02 - id: injection_resistance threshold: 1.0 # any failure blocksbudgets: p95_latency_s: 12 cost_per_task_usd: 0.08slices: [intent, language, tool]The release gate
The gate compares a candidate run against the pinned baseline. Two details matter more than the thresholds themselves. First, agents are stochastic, so I run each item more than once and compare distributions, not single scores. Second, a regression is only real if it is larger than the noise, so the gate uses a bootstrap interval on the difference rather than a raw comparison of means.
from dataclasses import dataclassimport randomimport statistics
@dataclassclass MetricRule: metric_id: str threshold: float # absolute floor max_regression: float # allowed drop vs baseline
@dataclassclass GateResult: metric_id: str passed: bool candidate_mean: float delta_low: float # lower bound of the bootstrap interval on (candidate - baseline) reason: str
def bootstrap_delta_low(candidate, baseline, n=2000, alpha=0.05, seed=7): """Lower bound of a (1 - alpha) interval on mean(candidate) - mean(baseline).""" rng = random.Random(seed) deltas = [] for _ in range(n): c = [rng.choice(candidate) for _ in candidate] b = [rng.choice(baseline) for _ in baseline] deltas.append(statistics.fmean(c) - statistics.fmean(b)) deltas.sort() return deltas[int(alpha / 2 * n)]
def evaluate_gate(rules, candidate_scores, baseline_scores): """candidate_scores / baseline_scores: metric_id -> list of per-run scores in [0, 1].""" results = [] for rule in rules: cand = candidate_scores.get(rule.metric_id, []) base = baseline_scores.get(rule.metric_id, []) if not cand: results.append(GateResult(rule.metric_id, False, 0.0, 0.0, "no scores: treat as failure")) continue mean = statistics.fmean(cand) if mean < rule.threshold: results.append(GateResult(rule.metric_id, False, mean, 0.0, f"mean {mean:.3f} below floor {rule.threshold}")) continue low = bootstrap_delta_low(cand, base) if base else 0.0 if low < -rule.max_regression: results.append(GateResult(rule.metric_id, False, mean, low, f"regression beyond {rule.max_regression} is plausible")) continue results.append(GateResult(rule.metric_id, True, mean, low, "ok")) return results
if __name__ == "__main__": rules = [MetricRule("tool_correctness", 0.95, 0.0), MetricRule("faithfulness", 0.90, 0.02)] # Illustrative scores, not measured results. candidate = {"tool_correctness": [1, 1, 1, 0, 1, 1, 1, 1, 1, 1] * 5, "faithfulness": [0.93, 0.88, 0.95, 0.91, 0.97] * 10} baseline = {"tool_correctness": [1] * 50, "faithfulness": [0.92, 0.90, 0.94, 0.89, 0.96] * 10} for r in evaluate_gate(rules, candidate, baseline): print(r) if not all(r.passed for r in evaluate_gate(rules, candidate, baseline)): raise SystemExit(1)Missing scores fail the gate. An evaluation that silently did not run is the most common way a gate turns into decoration.
Continuous evaluation and feedback
- 01Change: model, prompt, tool or config
- 02Offline suites in CI
- 03Gate against baseline
- 04Canary with online sampling
- 05Batch scoring of live traces
- 06Triage and add to regression set
Offline evaluation checks a candidate on known tasks. Online evaluation checks whether the known tasks still resemble reality. I sample live traces into the batch plane at a rate set per risk tier, with the sampler biased towards traces the real-time plane flagged (tool errors, long trajectories, guardrail hits, negative user feedback). Explicit feedback such as a thumbs-down is a weak label on its own, so it routes the trace to review rather than counting as a failure directly.
Triaged failures go back into the regression set with their trace. Over time it becomes the most valuable dataset you own, because it describes how your agent actually fails.
Validation strategy
An evaluation platform is itself a system under test. I validate it in four ways.
Judge calibration. Every LLM-as-a-judge metric is checked against a human-labelled calibration set before it can gate a release, and again whenever the judge model or rubric changes. I track agreement with human labels and, more usefully, the judge’s false-pass rate: how often it passes an answer that humans failed.
Seeded defects. I keep a small set of deliberately broken agent builds: a tool schema with a renamed argument, a prompt that skips a confirmation step, a retriever pointed at a stale index. Each must fail the relevant gate. If one passes, the metric or threshold is wrong.
Determinism checks on scorers. Deterministic scorers must return identical results on replay. Judge-based scorers are run several times on the same item to measure their own variance, which then informs how many runs per item the suite needs.
Trace completeness. A detector checks that every agent run has the expected span structure: a root invoke_agent span, a tool span for every tool call the model requested, and token usage on every inference span. Missing spans make every downstream score unreliable, so trace completeness is monitored like any other SLO.
Metrics and measurement
The platform reports two families of numbers: quality scores from the evaluation planes, and operational telemetry derived from spans. The OpenTelemetry conventions define metrics such as gen_ai.client.operation.duration, gen_ai.invoke_agent.duration and gen_ai.execute_tool.duration, and token usage is recorded on inference spans through gen_ai.usage.input_tokens and gen_ai.usage.output_tokens. Cost is then a join between token counts and a versioned price table, never a number the agent reports about itself.
| Metric | Definition | Plane | Example gate (illustrative) |
|---|---|---|---|
| Task completion | Share of tasks where the goal was achieved | Batch, offline | at least 0.85 |
| Tool correctness | Expected tools called, judged per task | Real-time and offline | at least 0.95, no regression |
| Argument correctness | Tool arguments valid and semantically right | Offline | at least 0.92 |
| Faithfulness | Share of answer claims supported by tool results or context | Batch, offline | at least 0.90 |
| Hallucinated actions | Claimed actions with no matching tool span | Real-time | zero tolerated in high tier |
| Hand-off completeness | Required context present in agent-to-agent hand-off | Batch | at least 0.90 |
| p95 latency per task | From root span duration | Real-time | within budget |
| Cost per task | Tokens times price table, summed over the trace | Real-time | within budget |
Every metric is reported per slice as well as overall. An average that holds steady while one intent collapses is the classic way a regression ships.
Challenges and trade-offs
Judge cost and latency. Judge-based scoring is the most informative and the most expensive part of the platform. Keeping it in the batch plane, sampling by risk, and caching scores by input hash keep the cost bounded. The trade-off is that some quality regressions are seen minutes or hours after they start, not seconds.
Multiple valid trajectories. Exact trajectory matching punishes an agent for finding a better path. I use exact matching only where order genuinely matters (for example, check before write), set-based matching elsewhere, and a judge for plan quality when the golden set lists alternatives.
Privacy versus evaluability. Tool arguments and results are what make tool-use evaluation possible, and they are also where personal data lives. Redaction before export, per-tool opt-in and short retention for raw payloads are the compromise, even though some scorers are slightly weaker on redacted text.
Convention churn. The GenAI semantic conventions are still in development, and attribute names have changed before. Pinning a version and mapping it in the collector keeps scorers stable when instrumentation libraries move.
Gate fatigue. Too many tight thresholds produce a gate that fails for noise, and teams learn to override it. Fewer metrics, chosen by risk and compared statistically, keep it credible.
Tool choice. I have evaluated tracing and evaluation tooling such as Langfuse, Arize and the NVIDIA NeMo Agent Toolkit. The lesson is less about which tool and more about the boundary: keep the span schema and the scorer interface yours, and treat vendors as interchangeable back ends.
Outcome and lessons learned
The value of this architecture is that every question in the problem statement gets an answer backed by a trace. A failing release points to a metric, a slice, a dataset item and a span. A production incident becomes a regression case with the evidence attached.
Further reading
- Beyond accuracy: evaluating agentic AI layer by layer: the framework behind the metric choices on this page.
- Testing tool-calling reliability in LLM applications: a pytest-style harness for tool selection, arguments and sequencing.
- OpenTelemetry semantic conventions for generative AI and the GenAI conventions repository.
- DeepEval metrics introduction: agentic, RAG and custom metrics.
- RAGAS documentation and its agent and tool-use metrics.
- NIST AI Risk Management Framework: for tying evaluation evidence to risk tiers.