In traditional testing, a passing test is a strong statement. Run it again with the same inputs and it passes again. Most of the measurement habits I built over a decade of QA rest on that assumption.
AI systems break it. Run the same agent on the same task five times and you may get five different tool sequences and three different answers, some correct. A test suite that runs each case once and reports “92% passed” is telling you about one draw from a distribution, not about the system.
This article is about measuring the distribution. It is marked “working” because the practice is still settling, but the parts below are the ones I rely on today.
Capability is not reliability
There are two different questions you can ask of an agent:
- Can it do this task? Capability. One success is evidence.
- Will it do this task every time? Reliability. One success is close to no evidence.
Code generation research popularised pass@k: the probability that at least one of k sampled attempts succeeds (Chen et al., 2021). That is a capability metric, and it is the right one when a human or a test harness picks the best of several attempts.
Agents in production rarely get k attempts. A customer gets one answer. The τ-bench benchmark introduced pass^k: the probability that all k independent attempts at the same task succeed (Yao et al., 2024). That is a reliability metric. It answers “if this task arrives k times, how often do we get it right every time?”
The two diverge quickly. If a task succeeds 80% of the time and attempts are independent, pass@k climbs towards 100% as k grows while pass^k falls as 0.8 raised to the power k.
pass@k versus pass^k for a task with 80% per-attempt success
| pass@8 | 100% | |
|---|---|---|
| pass^1 | 80% | |
| pass^2 | 64% | |
| pass^4 | 41% | |
| pass^8 | 17% |
The same agent looks excellent on one metric and unacceptable on the other. Neither is wrong. They answer different questions, and the reliability question is usually the one the business is asking.
Measure from repeated runs
To estimate either metric you run each task n times and count c successes. The unbiased estimators are straightforward:
from math import comb
def pass_at_k(n: int, c: int, k: int) -> float: """Probability at least one of k attempts succeeds (Chen et al., 2021).""" if n - c < k: return 1.0 return 1.0 - comb(n - c, k) / comb(n, k)
def pass_hat_k(n: int, c: int, k: int) -> float: """Probability all k attempts succeed (tau-bench style pass^k).""" return comb(c, k) / comb(n, k)
# Per-task results from 8 runs each: (task_id, successes)results = [("refund-01", 8), ("refund-02", 6), ("lookup-07", 8), ("escalate-03", 4)]for task, c in results: print(task, round(pass_at_k(8, c, 1), 2), round(pass_hat_k(8, c, 4), 2))Compute per task, then average across the suite. The per-task view matters: a suite average of 85% can hide a handful of tasks that succeed half the time, and those are the ones that generate incidents.
How many runs? For release gating I start with five to ten runs per task on the regression suite. It is expensive, so I prioritise: every run for high-risk tasks (anything with side effects or regulatory weight), fewer for low-risk read-only tasks.
Variance is a first-class result
A success rate without an interval is half a result. With 50 trials and 45 successes, the 95% Wilson interval for the true success rate is roughly 79% to 96%. With 1,000 trials at the same rate it narrows to roughly 88% to 92%. A release that “dropped from 90% to 86%” on 50 runs may not have changed at all.
Three kinds of variance are worth tracking separately:
| Variance | What it tells you | How I measure it |
|---|---|---|
| Outcome variance | Does the task succeed consistently? | Per-task success rate and pass^k over repeated runs |
| Path variance | Does the agent take the same route? | Distinct tool sequences per task, step count spread |
| Output variance | Is the answer stable when it succeeds? | Semantic similarity or judge agreement across runs |
Path variance is an early signal. An agent that solves a task through four different tool sequences is fragile even when all four succeed, because small changes to a prompt or a model shift which path it takes.
SLOs for non-deterministic systems
Pre-release metrics tell you whether to ship. In production, reliability is managed with service level objectives. The Google SRE book chapter on service level objectives sets out the vocabulary: an SLI is a measured quantity, an SLO is a target for it over a window, and the SLO should reflect what users care about rather than what is easy to measure.
For agents, the SLIs I use most:
- Task success rate. Tasks that completed and passed validation and review, over all tasks. Success has to be defined per task type, and sometimes judged after the fact by sampling with an evaluator or a human.
- p95 task latency. End to end, including retries and fallbacks.
- Cost per task. Tokens and tool costs per completed task. Retry storms and repair loops show up here first.
- Escalation rate. Tasks handed to a human. Not a failure, but a rising rate is an early warning.
- Fallback rate. Tasks that needed any fallback. A rising rate means the primary path is degrading even if success holds.
Non-determinism changes two things about SLOs. First, success often cannot be measured synchronously. A task can return a confident, well-formed, wrong answer. I combine cheap online signals (validation failures, reviewer rejections, user corrections) with sampled offline evaluation and treat the offline result as the authoritative SLI, accepting the delay. Second, windows need enough volume. A daily SLO on a workflow that runs 40 times a day will flap on noise; a 28-day window is usually steadier.
Error budgets
An error budget is the complement of the SLO. A 95% task success SLO over 28 days allows 5% of tasks to fail in that window. The budget turns reliability into a shared, spendable resource: while there is budget left, the team ships; when it runs out, reliability work takes priority. The SRE workbook’s error budget policy is a good template to adapt.
For AI systems I add one rule to the policy. A model upgrade, prompt change or new tool counts as a change that can spend budget, and needs the same release gate as code. Many AI reliability regressions come from configuration that never went through a pipeline.
Alerting works best on burn rate rather than raw failure counts, as described in the SRE workbook chapter on alerting on SLOs. A fast burn (budget being consumed many times faster than sustainable) pages someone; a slow burn opens a ticket.
How to set thresholds
Thresholds are where most measurement programmes stall. My approach:
- Measure the baseline first. Run the current system on the regression suite with repeated runs and record success,
pass^k, latency and cost with their intervals. - Set targets from risk, not from the baseline. A read-only summarisation task and a task that issues refunds deserve different targets. Tie targets to the use case risk tier.
- Gate on regression with a tolerance. For release decisions, compare the candidate to the baseline and block if the drop exceeds a tolerance that is larger than measurement noise. If the tolerance is smaller than the confidence interval, the gate is a coin toss.
- Agree thresholds before the results arrive. A threshold set after seeing the numbers is a rationalisation.
- Review quarterly. As volume grows, intervals tighten and targets can rise.
As an example, and only as an example: a team might gate releases on suite-level pass^4 not dropping more than 5 points from baseline, high-risk tasks individually holding at least 90% success over ten runs, and p95 latency and cost per task staying within 20% of baseline.
Dashboards that answer questions
A reliability dashboard should answer, in order: are we within SLO, how fast are we spending the budget, and where is it going. The panels I put on it:
- SLO status and remaining error budget per workflow, with burn rate.
- Task success, escalation and fallback rates over time, annotated with releases and model changes.
- Failures broken down by failure mode (timeout, malformed output, tool failure, loop) and by agent.
- p95 latency and cost per task, with retry counts overlaid.
- Per-task
pass^kfrom the latest regression run, sorted worst first.
The telemetry underneath is ordinary tracing. Instrumenting model and tool calls with the OpenTelemetry GenAI semantic conventions and tagging each span with its failure mode makes every one of these panels a query rather than a project. The handling patterns that produce those tags are in Agent failure handling patterns.