Ask a team how good their agent is and you will often hear one number: “about 85% accurate”. That number hides almost everything that matters. Accurate on which tasks? Measured how? Was the answer right because the agent used the right tool, or despite using the wrong one? Would the same number hold after next week’s prompt change?
An agent is a pipeline of decisions. It retrieves, plans, calls tools, sometimes hands off to another agent, and finally answers. A single accuracy score collapses all of that into one figure that cannot tell you what to fix. This article sets out the framework I use instead: evaluate each layer separately, choose metrics and thresholds deliberately, trust LLM judges only after calibrating them, and treat datasets and regression runs with the same discipline as code.
Evaluate each layer, not just the answer
I split an agentic system into five layers. Each fails in its own way and needs its own signals.
- L1SafetyPolicy, PII, injection resistance, excessive agency
- L2Final answerCorrectness, relevance, faithfulness, completeness
- L3Tool useSelection, arguments, sequencing, error handling
- L4Reasoning and planPlan quality, adherence, step efficiency
- L5RetrievalContext precision, context recall, freshness
Retrieval. Did the agent fetch the evidence it needed, and only that? Context precision and recall are the core metrics. A retrieval failure produces a confident wrong answer that no amount of prompt tuning will fix.
Reasoning and plan. Did the agent decompose the task sensibly, and did it follow its own plan? Plan quality is judged against a rubric. Plan adherence and step efficiency can partly be computed from the trace: an agent that takes eleven steps for a three-step task is telling you something even when it gets there.
Tool use. Right tool, right arguments, right order, sensible handling of errors. This is the layer where deterministic checks are most valuable, and I cover it in depth in a separate article on tool-calling reliability.
Final answer. Correctness against a reference where one exists. Relevance to the question. Faithfulness to what the tools and retrieval actually returned, which is where hallucination shows up. Completeness, because an answer that is correct but omits the one caveat the user needed is still a failure.
Safety. Whether the agent stays within policy under normal and adversarial input: prompt injection, personal data leakage, and excessive agency, where an agent takes an action it was never meant to take. The OWASP Top 10 for LLM Applications is a good checklist for what this layer should cover.
The point of separating layers is diagnosis. When the final-answer score drops, layer scores tell you whether to look at the retriever, the planner prompt, a tool schema or the model.
Choosing metrics
For each layer I choose metrics along two axes.
| Reference-based | Referenceless | |
|---|---|---|
| Deterministic | Exact tool-call match, argument equality, JSON schema validity | Step count, loop detection, latency, cost |
| Model-judged | Correctness against a reference answer | Faithfulness, relevance, plan quality, tone |
I prefer the top-left cell wherever it is honest to do so. Deterministic, reference-based checks are cheap, repeatable and impossible to argue with. Model-judged metrics are necessary for anything involving meaning, but they add cost, variance and a second model whose behaviour you now have to trust.
Referenceless metrics are the only ones you can run on live production traffic, because live traffic has no reference answer. That makes them the bridge between offline evaluation and monitoring. DeepEval makes the same point in its guidance on which metrics to use in production.
Three rules keep the metric set useful:
- Every metric maps to a decision. If no one would act differently when it moves, drop it.
- Keep the set small. Two or three general metrics plus one or two specific to the use case is usually enough. More metrics mean more noise and more false alarms.
- Report by slice. Intent, language, user segment and tool. Averages hide the slice that collapsed.
Setting thresholds
Thresholds are where evaluation becomes a decision. I set them with three inputs:
- The use case’s risk tier. An internal drafting tool and an agent that issues refunds deserve different bars. The tier comes from governance, not from the engineering team’s optimism. The NIST AI RMF is a useful frame for that conversation.
- The current baseline. A threshold far above what any build achieves is a wish, not a gate. Start near the baseline and ratchet upwards.
- The noise floor. Run the same build several times. If faithfulness moves by three points between identical runs, a two-point regression threshold will fail randomly.
I agree thresholds with product and engineering before the results arrive. Thresholds negotiated after a bad score are always lower.
For most metrics I use two conditions: an absolute floor, and a maximum allowed regression against the pinned baseline. For safety metrics in high-risk tiers I often use a third: zero tolerance on specific case categories, regardless of averages.
Calibrating LLM-as-a-judge
A judge model is a measuring instrument. Before it gates a release, I check it against human labels the same way you would calibrate any other instrument.
The process is simple and mostly manual:
- Sample 100 to 200 items stratified by slice, deliberately including known failures.
- Have two domain reviewers label them independently against the same rubric the judge uses.
- Resolve disagreements, and record the human-human agreement. That is the ceiling: a judge cannot be expected to agree with humans more than humans agree with each other.
- Run the judge and compare.
The numbers I look at:
from collections import Counter
def cohen_kappa(a, b): """Agreement between two label lists beyond chance. Labels are 'pass' / 'fail'.""" assert len(a) == len(b) and a n = len(a) observed = sum(x == y for x, y in zip(a, b)) / n ca, cb = Counter(a), Counter(b) expected = sum(ca[k] * cb[k] for k in set(a) | set(b)) / (n * n) return 1.0 if expected == 1 else (observed - expected) / (1 - expected)
def calibration_report(human, judge): pairs = list(zip(human, judge)) human_fails = [j for h, j in pairs if h == "fail"] human_passes = [j for h, j in pairs if h == "pass"] return { "kappa": round(cohen_kappa(human, judge), 3), # The judge passed something humans failed: the dangerous direction. "false_pass_rate": round(human_fails.count("pass") / max(len(human_fails), 1), 3), # The judge failed something humans passed: noisy, but safe. "false_fail_rate": round(human_passes.count("fail") / max(len(human_passes), 1), 3), }Kappa tells you overall agreement beyond chance. The false-pass rate tells you how often the judge would let a bad answer through, which is the number that matters for a release gate. A judge with good kappa and a high false-pass rate is lenient in exactly the wrong place.
A few habits reduce judge error before calibration even starts. Ask for a reason before the score, so the score is conditioned on an argument. Use a narrow binary or three-point rubric rather than a ten-point scale. Randomise the order when comparing two answers, because judges show position bias. Avoid using the same model family as judge and as system under test where you can.
Designing and versioning datasets
The dataset defines what “good” means, so I design it on purpose rather than collecting whatever is to hand.
Start from a coverage matrix. Rows are intents or task types; columns are the conditions that make tasks hard: ambiguous requests, missing data, tool errors, multi-step tasks, adversarial input, languages. Each cell gets a target number of items. Empty cells are visible gaps, not surprises in production.
Store expected behaviour, not just expected answers. For agents, an item needs the expected tools, acceptable alternative trajectories, and facts the answer must and must not contain.
{ "id": "refund-0142", "slice": { "intent": "refund", "condition": "tool_error", "lang": "en" }, "input": "I was charged twice for order 88123, can you fix it?", "expected_tools": ["get_order", "list_charges", "create_refund"], "allowed_orders": [["get_order", "list_charges", "create_refund"]], "must_include": ["one refund for the duplicate charge"], "must_not_include": ["refund for both charges"], "injected_fault": { "tool": "list_charges", "first_call": "timeout" }, "source": "production-incident", "added_in": "support-v7"}Version everything that affects the score. Dataset version, rubric version, judge configuration and scorer code. A score without those four is not comparable to any other score. I keep datasets in version control or a dataset registry with content hashes, and every evaluation run records the hashes it used.
Protect the golden set. If the same items are used to tune prompts and to gate releases, the gate measures memorisation. Keep a held-out portion that prompt engineers do not see, and rotate items in from the regression set.
Feed it from production. The best new cases are real failures. A failure that has been triaged and reproduced earns a permanent place in the regression set, tagged with where it came from.
Regression across prompt and model changes
Every change to a prompt, model version, tool schema or orchestration setting is a candidate release. The regression question is not “did the score go down” but “did it go down by more than noise, and where”.
Three practices make this reliable:
- Paired comparison. Run baseline and candidate on the same items, with several runs per item. Compare per-item outcomes, not just means.
- Flip analysis. List the items that passed on baseline and fail on candidate, and the reverse. A model upgrade that fixes forty items and breaks thirty-eight different ones has the same average and a very different risk profile. Review the newly broken items by hand.
- Interval, not point. Use a bootstrap interval on the difference in means. Block only when a regression beyond the allowed margin is plausible, so the gate does not fail for noise.
Model upgrades deserve extra care. Model aliases can move to a new version without notice, pinned versions get retired, and a new version can change tool-calling behaviour in ways prompt-level tests miss. I pin model versions explicitly and run the full suite, including tool-use and safety suites, before switching.
Further reading
- Evaluation and observability for agentic workflows: the reference architecture that puts this framework into a pipeline.
- Testing tool-calling reliability in LLM applications: the tool-use layer in detail.
- DeepEval metrics introduction
- RAGAS documentation
- NIST AI Risk Management Framework
- OWASP Top 10 for LLM Applications