Skip to content
Satya Prakash Solanki

Evaluation and observability for agentic workflows

Overview

An agentic workflow is not one model call. It is a plan, a sequence of tool calls, possibly a hand-off between agents, some retrieval, and a final answer. Any of those steps can be wrong while the final answer still looks plausible. That is what makes agents hard to test: the surface looks fine and the failure is three spans deep.

This page describes a reference architecture for an evaluation and observability platform for agentic systems. It covers what to measure at each step, how to trace an agent run so every score can be tied back to the step that produced it, how to run evaluation continuously across model, prompt and configuration changes, and how to turn the results into release gates.

It is a reference architecture, not a client story. It draws on patterns from my own practice, notably separating real-time and batch evaluation into two planes and enforcing evaluation results as release gates. Thresholds and numbers in the examples are illustrative.

Problem statement

Teams usually start with a demo that worked and a dashboard of latency and token counts. Neither answers the questions that matter before a release:

  • Did the agent choose the right tools, with the right arguments, in a sensible order?
  • Was the final answer correct, relevant and grounded in what the tools actually returned?
  • When two agents collaborated, did the hand-off carry the right context, or did one agent quietly redo or undo the other’s work?
  • Is this build better or worse than the last one, and on which kinds of task?
  • When it fails in production, can we find the failing step without replaying the whole conversation by hand?

Agents make these questions hard. Outputs are probabilistic, so one run proves little. Trajectories vary, so there is often more than one correct path. And quality depends on the prompt, tool schemas, retrieval index, orchestration code and model version. A change to any of them is a change to the product.

Engineering objectives

I set the objectives as testable statements, the same way I would for any other system under test:

  1. Every score is traceable. Any evaluation result links to the trace, span and dataset item that produced it.
  2. Evaluate per layer, not just the end. Retrieval, plan, tool use, final answer and safety each get their own signals.
  3. Metrics follow risk. The metric set and thresholds for a use case are chosen from its risk tier, not from whatever a library offers by default.
  4. Regression is the default question. Every model, prompt, tool or configuration change runs the same suites against a versioned baseline.
  5. Production feeds evaluation. Sampled live traffic is scored online, and failures flow back into the offline datasets.
  6. Fast signals stay fast. Expensive judge-based scoring never sits in the request path.

Solution architecture

The platform has four groups of components: instrumentation, transport, evaluation and decision.

01 Instrument

  • OpenTelemetry SDK in agent runtime
  • GenAI span and attribute conventions
  • Tool and retrieval wrappers

02 Transport

  • Collector with sampling
  • Event stream (Kafka topics)
  • Trace store

03 Evaluate

  • Real-time detectors
  • Batch scorers and judges
  • Offline suite runner

04 Decide

  • Release gate in CI
  • Alerts and SLOs
  • Review queue and datasets
Figure 1. Reference architecture. Instrumentation emits standard spans; two evaluation planes consume the same stream; decisions are made in CI, in alerting and in human review.

Tracing model

Everything starts with a consistent span model. I follow the OpenTelemetry semantic conventions for generative AI, which define spans for agent invocation, tool execution and model inference. The conventions are still marked as in development and are now maintained in a dedicated repository, so I pin the version I instrument against and treat upgrades as a change like any other.

A single agent run produces a tree like this:

invoke_agent support_agent gen_ai.operation.name=invoke_agent
├── chat gpt-x gen_ai.usage.input_tokens / output_tokens
├── execute_tool search_orders gen_ai.tool.name, gen_ai.tool.call.id
│ └── (downstream HTTP/DB spans)
├── chat gpt-x
├── invoke_agent refunds_agent hand-off to a second agent
│ ├── execute_tool get_refund_policy
│ └── chat gpt-x
└── chat gpt-x final answer

The span names follow the convention formats: invoke_agent {gen_ai.agent.name}, execute_tool {gen_ai.tool.name} and {gen_ai.operation.name} {gen_ai.request.model} for inference. Three practical rules make the tree useful for evaluation:

  • Correlate by conversation. gen_ai.conversation.id ties multi-turn sessions together so a judge can see context, not just one turn.
  • Record tool arguments and results deliberately. gen_ai.tool.call.arguments and gen_ai.tool.call.result are what tool-use scorers need, but they can contain personal data. I capture them behind a redaction step and an explicit opt-in per tool.
  • Add evaluation context as your own attributes. Dataset item id, suite version and prompt version go on the root span under a private namespace, so offline runs and live traffic share one schema.

Two evaluation planes

Fast signals and deep evaluation pull in different directions. A failed tool call or a policy breach should alert in seconds. Groundedness or multi-agent coordination quality needs a judge model, retrieved context and sometimes the whole conversation, which is too slow and too expensive for the hot path.

The pattern I use is two planes over one stream. Both consume the same trace events from the event bus, but they scale and fail independently.

Plane Runs on Typical checks Latency budget
Real-time Every trace, streaming Tool errors, schema violations, loop detection, guardrail hits, latency and cost budgets Seconds
Batch Sampled traces and full offline suites Groundedness, correctness against reference, plan quality, coordination, judge-based rubrics Minutes to hours

Checks are packaged as detectors behind a common interface, so a new failure mode is a new detector, not a change to transport or storage. The same scorer code runs on live samples and on golden datasets; if offline and online used different scorers, their numbers could not be compared.

Technical approach

What to measure, layer by layer

Layer Question Example metrics Scoring method
Retrieval Did we fetch the right evidence? Context precision, context recall Reference-based, LLM-assisted
Plan Was the plan sensible and followed? Plan quality, plan adherence, step efficiency Judge with rubric
Tool use Right tool, right arguments, right order? Tool selection accuracy, argument correctness, sequence match, tool error rate Mostly deterministic
Coordination Did agents hand off cleanly? Hand-off completeness, redundant work, conflicting actions Trace analysis plus judge
Final answer Is it correct, relevant and grounded? Correctness, answer relevancy, faithfulness, hallucination rate Reference-based and judge
Workflow Did the user’s goal get done? Task completion, goal accuracy Reference-based where possible
Safety Did it stay within bounds? Policy violation rate, PII leakage, prompt-injection resistance Classifiers, Promptfoo suites
Operations Is it affordable and fast? p50/p95 latency, tokens per task, cost per task, tool calls per task Span arithmetic

DeepEval separates trajectory metrics (task completion, step efficiency, plan adherence and plan quality) from component metrics (tool correctness, argument correctness), alongside RAG metrics such as faithfulness and contextual precision. RAGAS offers tool call accuracy, tool call F1 and agent goal accuracy for agentic workflows. I wrap them behind one scorer interface so the gate does not care which library produced a number.

Hallucination detection deserves a specific note. For agents, I split it into two checks: claims in the answer that are not supported by any tool result or retrieved context (a faithfulness problem), and claims that a tool was called or an action was taken when the trace shows it was not. The second kind is cheap to detect from spans and surprisingly common.

Choosing metrics by risk

Not every use case needs every metric. I select from the table above using the use case’s risk tier, which comes from the governance process rather than from the engineering team’s preference.

Risk tier Example use case Required metrics Human review
Low Internal drafting assistant Relevancy, safety baseline, cost Spot checks
Medium Customer-facing Q&A with tools Above plus faithfulness, tool correctness, task completion Sampled weekly
High Agent that changes records or money Above plus argument correctness, sequence match, idempotency checks, adversarial suite Every failed gate, calibrated judges

The tier also sets how strict the gate is. A low-tier assistant can ship with a small regression on relevancy if cost drops; a high-tier agent cannot ship with any regression on argument correctness.

Datasets, golden sets and regression suites

Datasets are versioned artefacts with an owner, like code. I keep four kinds:

  • Golden set. Curated tasks with reference answers, expected tool calls and acceptable alternative trajectories. Small, high quality, reviewed by domain experts.
  • Regression set. Every production failure that was triaged as a real defect becomes a case here, with the trace that exposed it.
  • Adversarial set. Prompt injection, jailbreak and tool-misuse probes, maintained alongside the security test suites.
  • Benchmark slices. Stratified samples by intent, language and tool, so a score can be reported per slice and not only as an average.

The suite configuration ties datasets, metrics and thresholds together:

suites/support-agent.yaml
suite: support-agent
risk_tier: medium
datasets:
- name: golden/support-v7 # pinned version
- name: regression/support-2026q3
- name: adversarial/injection-v3
runs_per_item: 3 # agents are stochastic; score the distribution
judge:
model: judge-model-pinned
rubric_version: faithfulness-r4
calibration_set: human-labels/support-v2
metrics:
- id: task_completion
threshold: 0.85 # example value
max_regression: 0.02
- id: tool_correctness
threshold: 0.95
max_regression: 0.0
- id: argument_correctness
threshold: 0.92
max_regression: 0.01
- id: faithfulness
threshold: 0.90
max_regression: 0.02
- id: injection_resistance
threshold: 1.0 # any failure blocks
budgets:
p95_latency_s: 12
cost_per_task_usd: 0.08
slices: [intent, language, tool]

The release gate

The gate compares a candidate run against the pinned baseline. Two details matter more than the thresholds themselves. First, agents are stochastic, so I run each item more than once and compare distributions, not single scores. Second, a regression is only real if it is larger than the noise, so the gate uses a bootstrap interval on the difference rather than a raw comparison of means.

eval_gate.py
from dataclasses import dataclass
import random
import statistics
@dataclass
class MetricRule:
metric_id: str
threshold: float # absolute floor
max_regression: float # allowed drop vs baseline
@dataclass
class GateResult:
metric_id: str
passed: bool
candidate_mean: float
delta_low: float # lower bound of the bootstrap interval on (candidate - baseline)
reason: str
def bootstrap_delta_low(candidate, baseline, n=2000, alpha=0.05, seed=7):
"""Lower bound of a (1 - alpha) interval on mean(candidate) - mean(baseline)."""
rng = random.Random(seed)
deltas = []
for _ in range(n):
c = [rng.choice(candidate) for _ in candidate]
b = [rng.choice(baseline) for _ in baseline]
deltas.append(statistics.fmean(c) - statistics.fmean(b))
deltas.sort()
return deltas[int(alpha / 2 * n)]
def evaluate_gate(rules, candidate_scores, baseline_scores):
"""candidate_scores / baseline_scores: metric_id -> list of per-run scores in [0, 1]."""
results = []
for rule in rules:
cand = candidate_scores.get(rule.metric_id, [])
base = baseline_scores.get(rule.metric_id, [])
if not cand:
results.append(GateResult(rule.metric_id, False, 0.0, 0.0, "no scores: treat as failure"))
continue
mean = statistics.fmean(cand)
if mean < rule.threshold:
results.append(GateResult(rule.metric_id, False, mean, 0.0,
f"mean {mean:.3f} below floor {rule.threshold}"))
continue
low = bootstrap_delta_low(cand, base) if base else 0.0
if low < -rule.max_regression:
results.append(GateResult(rule.metric_id, False, mean, low,
f"regression beyond {rule.max_regression} is plausible"))
continue
results.append(GateResult(rule.metric_id, True, mean, low, "ok"))
return results
if __name__ == "__main__":
rules = [MetricRule("tool_correctness", 0.95, 0.0), MetricRule("faithfulness", 0.90, 0.02)]
# Illustrative scores, not measured results.
candidate = {"tool_correctness": [1, 1, 1, 0, 1, 1, 1, 1, 1, 1] * 5,
"faithfulness": [0.93, 0.88, 0.95, 0.91, 0.97] * 10}
baseline = {"tool_correctness": [1] * 50,
"faithfulness": [0.92, 0.90, 0.94, 0.89, 0.96] * 10}
for r in evaluate_gate(rules, candidate, baseline):
print(r)
if not all(r.passed for r in evaluate_gate(rules, candidate, baseline)):
raise SystemExit(1)

Missing scores fail the gate. An evaluation that silently did not run is the most common way a gate turns into decoration.

Continuous evaluation and feedback

  1. 01Change: model, prompt, tool or config
  2. 02Offline suites in CI
  3. 03Gate against baseline
  4. 04Canary with online sampling
  5. 05Batch scoring of live traces
  6. 06Triage and add to regression set
Figure 2. The continuous evaluation loop. Every change runs the same suites, and production failures become tomorrow’s regression cases.

Offline evaluation checks a candidate on known tasks. Online evaluation checks whether the known tasks still resemble reality. I sample live traces into the batch plane at a rate set per risk tier, with the sampler biased towards traces the real-time plane flagged (tool errors, long trajectories, guardrail hits, negative user feedback). Explicit feedback such as a thumbs-down is a weak label on its own, so it routes the trace to review rather than counting as a failure directly.

Triaged failures go back into the regression set with their trace. Over time it becomes the most valuable dataset you own, because it describes how your agent actually fails.

Validation strategy

An evaluation platform is itself a system under test. I validate it in four ways.

Judge calibration. Every LLM-as-a-judge metric is checked against a human-labelled calibration set before it can gate a release, and again whenever the judge model or rubric changes. I track agreement with human labels and, more usefully, the judge’s false-pass rate: how often it passes an answer that humans failed.

Seeded defects. I keep a small set of deliberately broken agent builds: a tool schema with a renamed argument, a prompt that skips a confirmation step, a retriever pointed at a stale index. Each must fail the relevant gate. If one passes, the metric or threshold is wrong.

Determinism checks on scorers. Deterministic scorers must return identical results on replay. Judge-based scorers are run several times on the same item to measure their own variance, which then informs how many runs per item the suite needs.

Trace completeness. A detector checks that every agent run has the expected span structure: a root invoke_agent span, a tool span for every tool call the model requested, and token usage on every inference span. Missing spans make every downstream score unreliable, so trace completeness is monitored like any other SLO.

Metrics and measurement

The platform reports two families of numbers: quality scores from the evaluation planes, and operational telemetry derived from spans. The OpenTelemetry conventions define metrics such as gen_ai.client.operation.duration, gen_ai.invoke_agent.duration and gen_ai.execute_tool.duration, and token usage is recorded on inference spans through gen_ai.usage.input_tokens and gen_ai.usage.output_tokens. Cost is then a join between token counts and a versioned price table, never a number the agent reports about itself.

Metric Definition Plane Example gate (illustrative)
Task completion Share of tasks where the goal was achieved Batch, offline at least 0.85
Tool correctness Expected tools called, judged per task Real-time and offline at least 0.95, no regression
Argument correctness Tool arguments valid and semantically right Offline at least 0.92
Faithfulness Share of answer claims supported by tool results or context Batch, offline at least 0.90
Hallucinated actions Claimed actions with no matching tool span Real-time zero tolerated in high tier
Hand-off completeness Required context present in agent-to-agent hand-off Batch at least 0.90
p95 latency per task From root span duration Real-time within budget
Cost per task Tokens times price table, summed over the trace Real-time within budget

Every metric is reported per slice as well as overall. An average that holds steady while one intent collapses is the classic way a regression ships.

Challenges and trade-offs

Judge cost and latency. Judge-based scoring is the most informative and the most expensive part of the platform. Keeping it in the batch plane, sampling by risk, and caching scores by input hash keep the cost bounded. The trade-off is that some quality regressions are seen minutes or hours after they start, not seconds.

Multiple valid trajectories. Exact trajectory matching punishes an agent for finding a better path. I use exact matching only where order genuinely matters (for example, check before write), set-based matching elsewhere, and a judge for plan quality when the golden set lists alternatives.

Privacy versus evaluability. Tool arguments and results are what make tool-use evaluation possible, and they are also where personal data lives. Redaction before export, per-tool opt-in and short retention for raw payloads are the compromise, even though some scorers are slightly weaker on redacted text.

Convention churn. The GenAI semantic conventions are still in development, and attribute names have changed before. Pinning a version and mapping it in the collector keeps scorers stable when instrumentation libraries move.

Gate fatigue. Too many tight thresholds produce a gate that fails for noise, and teams learn to override it. Fewer metrics, chosen by risk and compared statistically, keep it credible.

Tool choice. I have evaluated tracing and evaluation tooling such as Langfuse, Arize and the NVIDIA NeMo Agent Toolkit. The lesson is less about which tool and more about the boundary: keep the span schema and the scorer interface yours, and treat vendors as interchangeable back ends.

Outcome and lessons learned

The value of this architecture is that every question in the problem statement gets an answer backed by a trace. A failing release points to a metric, a slice, a dataset item and a span. A production incident becomes a regression case with the evidence attached.

Further reading

Try “evaluation”, “red-teaming”, “governance” or “agents”.