When an LLM only writes text, a bad output is a bad answer. When it calls tools, a bad output is an action: a refund issued twice, a record updated with the wrong customer id, a search run with an empty query that returns nothing and leads to a confident “I could not find your order”. Tool calling is where language models touch real systems, so it deserves the most deterministic testing you can give it.
This is a working note. The categories below are stable in my practice; the harness is a pattern I keep refining rather than a finished library.
What can go wrong
I group tool-calling failures into six categories. Each needs a different kind of check.
| Category | Failure looks like | Check |
|---|---|---|
| Schema adherence | Missing required field, wrong type, extra field, malformed JSON | Validate arguments against the tool’s JSON Schema |
| Argument correctness | Valid shape, wrong value: wrong id, wrong date, unit confusion | Compare with expected values or derive from the input |
| Tool selection | Calls search_web when get_order was needed, or calls nothing |
Expected tool set per test case |
| Sequencing | Writes before reading, skips a confirmation step | Ordering constraints, not full sequence equality |
| Idempotency | Retries a non-idempotent call and doubles the effect | Count side effects per logical action |
| Error handling | Ignores a tool error, invents a result, loops on retries | Fault injection and assertions on recovery behaviour |
Schema adherence is the easiest to automate and the least interesting. Argument correctness and error handling are where most real defects hide.
Schema adherence
Every tool already has a schema, because the model needs one to call it. Reuse that schema as the test oracle. Validate every argument payload the model produces against it with a standard JSON Schema validator, and fail on extra properties as well as missing ones.
Two things are worth checking beyond validity. First, whether the schema itself is good enough: vague descriptions and free-text fields invite wrong arguments, so a high schema-failure rate is often a schema defect, not a model defect. Second, whether provider-side structured output or strict modes are enabled. They reduce malformed calls but do nothing for argument correctness.
Argument correctness
A schema-valid call can still be wrong. create_refund(order_id="88123", amount=49.0) passes validation whether or not 49.0 is the right amount.
I use three levels of check, from cheapest to most expensive:
- Exact match for arguments that are fully determined by the input or by earlier tool results, such as an order id quoted by the user.
- Derived match for arguments that must equal a value returned by an earlier tool. The harness checks that the
amountpassed tocreate_refundequals the duplicate charge returned bylist_charges. This catches the model inventing values. - Judged match for free-text arguments such as a search query, where several phrasings are acceptable. A judge or embedding similarity decides, and I keep these to a minimum.
Derived match is the one I would add first. It turns “the model hallucinated an argument” from a vague worry into a test failure.
Tool selection and sequencing
Tool selection accuracy asks whether the set of tools called matches the set expected. I score it as precision and recall rather than a single pass or fail, because calling an extra harmless lookup is a different failure from missing the write that completes the task.
Sequencing is where I see teams over-specify. Requiring an exact sequence fails an agent that calls two independent lookups in the other order. I express ordering as constraints instead:
get_orderbeforecreate_refundconfirm_with_userbefore any tool marked destructive- no tool call after
escalate_to_human
Constraints capture what actually matters for safety and correctness, and they survive harmless variation between runs and model versions. RAGAS makes a similar distinction between its order-sensitive tool call accuracy and its unordered tool call F1.
Idempotency and side effects
Agents retry. Frameworks retry. Networks retry. If a write tool is not idempotent, a retry becomes a duplicate effect.
The test is not on the model’s output but on the tool’s side effects. Mock tools record every invocation, and the harness asserts on effects per logical action: exactly one refund created for one duplicate charge, even when the first create_refund call timed out. The fix is often outside the model entirely, such as an idempotency key generated per user request and passed through, and the test is what proves it is wired correctly.
Error handling
Real tools fail. The questions are whether the agent notices, whether it tells the truth about it, and whether it stops.
I inject faults into mock tools: timeouts, 4xx-style validation errors, empty results, partial results and malformed responses. Then I assert on behaviour:
- After a validation error, the agent corrects its arguments rather than repeating the same call.
- After repeated failures, it stops within a retry budget and says it could not complete the task.
- It never claims an action succeeded when the tool returned an error. I check this by comparing claims in the final answer with the tool results in the trajectory.
The last check catches the most embarrassing failures in production, and it is cheap because both sides are in the trace.
Mocks, recordings and live tools
I use three modes, each for a different purpose.
| Mode | Use for | Trade-off |
|---|---|---|
| Scripted mocks | Fault injection, edge cases, CI on every commit | Fast and deterministic; can drift from reality |
| Recorded responses | Realistic data without hitting live systems | Must be refreshed and redacted |
| Live or sandbox | Contract checks and pre-release smoke tests | Slow, flaky, may cost money or have side effects |
Write tools never run live in automated tests outside a sandbox. That rule is not negotiable, however good the guardrails are.
Trajectory evaluation
Checks on individual calls miss problems that only appear across the whole run: loops, redundant calls, abandoned plans. Trajectory evaluation looks at the full ordered list of tool calls for a task and asks whether it completed the goal efficiently and within constraints.
- 01Load test case
- 02Run agent with mock tools
- 03Capture trajectory
- 04Deterministic checks
- 05Judged checks
- 06Aggregate across runs
Because agents are stochastic, I run each case several times and report a pass rate per case. A case that passes four times out of five is not a passing case for a write tool; it is a defect that fails one time in five.
A pytest-style harness
The harness below assumes your agent exposes a function that takes a prompt and a tool registry, runs to completion, and returns the final answer plus the list of tool calls it made. The names are placeholders for your own code.
import jsonimport pytestfrom jsonschema import Draft202012Validator
from myagent import run_agent # your agent entry pointfrom tools.schemas import TOOL_SCHEMAS # name -> JSON Schema for arguments
RUNS_PER_CASE = 5CASES = [json.loads(line) for line in open("tests/cases/tool_calls.jsonl")]
class MockTools: """Scripted tools that record calls and can inject faults."""
def __init__(self, fixtures, faults=None): self.fixtures, self.faults = fixtures, dict(faults or {}) self.calls, self.effects = [], []
def __call__(self, name, args): self.calls.append((name, args)) fault = self.faults.pop(name, None) # inject once, then behave normally if fault == "timeout": raise TimeoutError(f"{name} timed out") if name.startswith("create_"): self.effects.append((name, args.get("idempotency_key"))) return self.fixtures[name]
def check_schema(calls): for name, args in calls: errors = list(Draft202012Validator(TOOL_SCHEMAS[name]).iter_errors(args)) assert not errors, f"{name}: {errors[0].message}"
def check_order(calls, before_pairs): names = [n for n, _ in calls] for first, then in before_pairs: if then in names: assert first in names and names.index(first) < names.index(then), \ f"{first} must precede {then}"
def resolve(fixtures, path): """Look up 'list_charges.duplicate.amount' in the scripted tool responses.""" value = fixtures for key in path.split("."): value = value[key] return value
def check_derived(calls, derived, fixtures): """derived: {"create_refund.amount": "list_charges.duplicate.amount"}""" by_tool = {n: a for n, a in calls} for target, source in derived.items(): tool, arg = target.split(".", 1) assert tool in by_tool, f"{tool} not called" assert by_tool[tool].get(arg) == resolve(fixtures, source), \ f"{target} not derived from {source}"
def run_case(case): tools = MockTools(case["fixtures"], case.get("faults")) result = run_agent(case["input"], tools=tools) called = {n for n, _ in tools.calls}
check_schema(tools.calls) assert set(case["expected_tools"]) <= called, f"missing {set(case['expected_tools']) - called}" assert not called & set(case.get("forbidden_tools", [])), "forbidden tool called" check_order(tools.calls, case.get("before", [])) check_derived(tools.calls, case.get("derived", {}), case["fixtures"]) assert len(tools.effects) == case.get("expected_effects", len(tools.effects)), "duplicate side effect" assert len(tools.calls) <= case.get("max_calls", 10), "possible loop" for phrase in case.get("must_not_claim", []): assert phrase not in result.answer.lower(), f"claimed: {phrase}"
@pytest.mark.parametrize("case", CASES, ids=[c["id"] for c in CASES])def test_tool_calling(case): failures = [] for _ in range(RUNS_PER_CASE): try: run_case(case) except AssertionError as exc: failures.append(str(exc)) required = case.get("min_pass_rate", 1.0) # write-tool cases: 1.0 pass_rate = 1 - len(failures) / RUNS_PER_CASE assert pass_rate >= required, f"pass rate {pass_rate:.2f}: {failures[:3]}"The structure is the point. Deterministic checks run in a fixed order, each failure message names the rule it broke, and the pass rate across runs is the unit of judgement.
For tool selection scored as a metric rather than an assertion, DeepEval provides a ToolCorrectnessMetric that compares tools_called against expected_tools, with options for ordering and exact matching. Its core comparison is deterministic. I use it for reporting trends across builds and keep the pytest assertions for hard rules.
Where this fits
These tests are the tool-use layer of a wider evaluation framework. They run in CI on every prompt, model, schema or orchestration change, and production traces feed new cases back in: every tool failure that reaches a user becomes a JSONL line in the case file.
Further reading
- Evaluation and observability for agentic workflows: how these checks run in a continuous evaluation platform.
- Beyond accuracy: evaluating agentic AI layer by layer: the wider framework this layer belongs to.
- DeepEval: tool correctness metric
- RAGAS: metrics for agents and tool use
- JSON Schema
- pytest documentation