Skip to content
Satya Prakash Solanki

When an LLM only writes text, a bad output is a bad answer. When it calls tools, a bad output is an action: a refund issued twice, a record updated with the wrong customer id, a search run with an empty query that returns nothing and leads to a confident “I could not find your order”. Tool calling is where language models touch real systems, so it deserves the most deterministic testing you can give it.

This is a working note. The categories below are stable in my practice; the harness is a pattern I keep refining rather than a finished library.

What can go wrong

I group tool-calling failures into six categories. Each needs a different kind of check.

Category Failure looks like Check
Schema adherence Missing required field, wrong type, extra field, malformed JSON Validate arguments against the tool’s JSON Schema
Argument correctness Valid shape, wrong value: wrong id, wrong date, unit confusion Compare with expected values or derive from the input
Tool selection Calls search_web when get_order was needed, or calls nothing Expected tool set per test case
Sequencing Writes before reading, skips a confirmation step Ordering constraints, not full sequence equality
Idempotency Retries a non-idempotent call and doubles the effect Count side effects per logical action
Error handling Ignores a tool error, invents a result, loops on retries Fault injection and assertions on recovery behaviour

Schema adherence is the easiest to automate and the least interesting. Argument correctness and error handling are where most real defects hide.

Schema adherence

Every tool already has a schema, because the model needs one to call it. Reuse that schema as the test oracle. Validate every argument payload the model produces against it with a standard JSON Schema validator, and fail on extra properties as well as missing ones.

Two things are worth checking beyond validity. First, whether the schema itself is good enough: vague descriptions and free-text fields invite wrong arguments, so a high schema-failure rate is often a schema defect, not a model defect. Second, whether provider-side structured output or strict modes are enabled. They reduce malformed calls but do nothing for argument correctness.

Argument correctness

A schema-valid call can still be wrong. create_refund(order_id="88123", amount=49.0) passes validation whether or not 49.0 is the right amount.

I use three levels of check, from cheapest to most expensive:

  1. Exact match for arguments that are fully determined by the input or by earlier tool results, such as an order id quoted by the user.
  2. Derived match for arguments that must equal a value returned by an earlier tool. The harness checks that the amount passed to create_refund equals the duplicate charge returned by list_charges. This catches the model inventing values.
  3. Judged match for free-text arguments such as a search query, where several phrasings are acceptable. A judge or embedding similarity decides, and I keep these to a minimum.

Derived match is the one I would add first. It turns “the model hallucinated an argument” from a vague worry into a test failure.

Tool selection and sequencing

Tool selection accuracy asks whether the set of tools called matches the set expected. I score it as precision and recall rather than a single pass or fail, because calling an extra harmless lookup is a different failure from missing the write that completes the task.

Sequencing is where I see teams over-specify. Requiring an exact sequence fails an agent that calls two independent lookups in the other order. I express ordering as constraints instead:

  • get_order before create_refund
  • confirm_with_user before any tool marked destructive
  • no tool call after escalate_to_human

Constraints capture what actually matters for safety and correctness, and they survive harmless variation between runs and model versions. RAGAS makes a similar distinction between its order-sensitive tool call accuracy and its unordered tool call F1.

Idempotency and side effects

Agents retry. Frameworks retry. Networks retry. If a write tool is not idempotent, a retry becomes a duplicate effect.

The test is not on the model’s output but on the tool’s side effects. Mock tools record every invocation, and the harness asserts on effects per logical action: exactly one refund created for one duplicate charge, even when the first create_refund call timed out. The fix is often outside the model entirely, such as an idempotency key generated per user request and passed through, and the test is what proves it is wired correctly.

Error handling

Real tools fail. The questions are whether the agent notices, whether it tells the truth about it, and whether it stops.

I inject faults into mock tools: timeouts, 4xx-style validation errors, empty results, partial results and malformed responses. Then I assert on behaviour:

  • After a validation error, the agent corrects its arguments rather than repeating the same call.
  • After repeated failures, it stops within a retry budget and says it could not complete the task.
  • It never claims an action succeeded when the tool returned an error. I check this by comparing claims in the final answer with the tool results in the trajectory.

The last check catches the most embarrassing failures in production, and it is cheap because both sides are in the trace.

Mocks, recordings and live tools

I use three modes, each for a different purpose.

Mode Use for Trade-off
Scripted mocks Fault injection, edge cases, CI on every commit Fast and deterministic; can drift from reality
Recorded responses Realistic data without hitting live systems Must be refreshed and redacted
Live or sandbox Contract checks and pre-release smoke tests Slow, flaky, may cost money or have side effects

Write tools never run live in automated tests outside a sandbox. That rule is not negotiable, however good the guardrails are.

Trajectory evaluation

Checks on individual calls miss problems that only appear across the whole run: loops, redundant calls, abandoned plans. Trajectory evaluation looks at the full ordered list of tool calls for a task and asks whether it completed the goal efficiently and within constraints.

  1. 01Load test case
  2. 02Run agent with mock tools
  3. 03Capture trajectory
  4. 04Deterministic checks
  5. 05Judged checks
  6. 06Aggregate across runs
Figure 1. Harness flow. Deterministic checks run first and fail fast; judged checks run only on trajectories that pass them.

Because agents are stochastic, I run each case several times and report a pass rate per case. A case that passes four times out of five is not a passing case for a write tool; it is a defect that fails one time in five.

A pytest-style harness

The harness below assumes your agent exposes a function that takes a prompt and a tool registry, runs to completion, and returns the final answer plus the list of tool calls it made. The names are placeholders for your own code.

tests/test_tool_calling.py
import json
import pytest
from jsonschema import Draft202012Validator
from myagent import run_agent # your agent entry point
from tools.schemas import TOOL_SCHEMAS # name -> JSON Schema for arguments
RUNS_PER_CASE = 5
CASES = [json.loads(line) for line in open("tests/cases/tool_calls.jsonl")]
class MockTools:
"""Scripted tools that record calls and can inject faults."""
def __init__(self, fixtures, faults=None):
self.fixtures, self.faults = fixtures, dict(faults or {})
self.calls, self.effects = [], []
def __call__(self, name, args):
self.calls.append((name, args))
fault = self.faults.pop(name, None) # inject once, then behave normally
if fault == "timeout":
raise TimeoutError(f"{name} timed out")
if name.startswith("create_"):
self.effects.append((name, args.get("idempotency_key")))
return self.fixtures[name]
def check_schema(calls):
for name, args in calls:
errors = list(Draft202012Validator(TOOL_SCHEMAS[name]).iter_errors(args))
assert not errors, f"{name}: {errors[0].message}"
def check_order(calls, before_pairs):
names = [n for n, _ in calls]
for first, then in before_pairs:
if then in names:
assert first in names and names.index(first) < names.index(then), \
f"{first} must precede {then}"
def resolve(fixtures, path):
"""Look up 'list_charges.duplicate.amount' in the scripted tool responses."""
value = fixtures
for key in path.split("."):
value = value[key]
return value
def check_derived(calls, derived, fixtures):
"""derived: {"create_refund.amount": "list_charges.duplicate.amount"}"""
by_tool = {n: a for n, a in calls}
for target, source in derived.items():
tool, arg = target.split(".", 1)
assert tool in by_tool, f"{tool} not called"
assert by_tool[tool].get(arg) == resolve(fixtures, source), \
f"{target} not derived from {source}"
def run_case(case):
tools = MockTools(case["fixtures"], case.get("faults"))
result = run_agent(case["input"], tools=tools)
called = {n for n, _ in tools.calls}
check_schema(tools.calls)
assert set(case["expected_tools"]) <= called, f"missing {set(case['expected_tools']) - called}"
assert not called & set(case.get("forbidden_tools", [])), "forbidden tool called"
check_order(tools.calls, case.get("before", []))
check_derived(tools.calls, case.get("derived", {}), case["fixtures"])
assert len(tools.effects) == case.get("expected_effects", len(tools.effects)), "duplicate side effect"
assert len(tools.calls) <= case.get("max_calls", 10), "possible loop"
for phrase in case.get("must_not_claim", []):
assert phrase not in result.answer.lower(), f"claimed: {phrase}"
@pytest.mark.parametrize("case", CASES, ids=[c["id"] for c in CASES])
def test_tool_calling(case):
failures = []
for _ in range(RUNS_PER_CASE):
try:
run_case(case)
except AssertionError as exc:
failures.append(str(exc))
required = case.get("min_pass_rate", 1.0) # write-tool cases: 1.0
pass_rate = 1 - len(failures) / RUNS_PER_CASE
assert pass_rate >= required, f"pass rate {pass_rate:.2f}: {failures[:3]}"

The structure is the point. Deterministic checks run in a fixed order, each failure message names the rule it broke, and the pass rate across runs is the unit of judgement.

For tool selection scored as a metric rather than an assertion, DeepEval provides a ToolCorrectnessMetric that compares tools_called against expected_tools, with options for ordering and exact matching. Its core comparison is deterministic. I use it for reporting trends across builds and keep the pytest assertions for hard rules.

Where this fits

These tests are the tool-use layer of a wider evaluation framework. They run in CI on every prompt, model, schema or orchestration change, and production traces feed new cases back in: every tool failure that reaches a user becomes a JSONL line in the case file.

Further reading

AI evaluation & observability

Try “evaluation”, “red-teaming”, “governance” or “agents”.