Skip to content
Satya Prakash Solanki

Agents that test software

ChangeTest agentData agentBug agentQuality board
Abstract architecture pattern. Internal systems and product details are confidential.

Overview

Quality engineering is full of repeatable judgement: running suites, checking data, writing up defects and telling the right people. These are exactly the tasks agentic AI is starting to handle well, if the agents are designed and observed carefully.

I build and deploy agentic AI systems for QA: a testing agent, a data validation agent, a bug logging agent and a notification agent, orchestrated with CrewAI, LangGraph and Agno, using LangChain components and MCP for tool access, traced with LangSmith and surfaced in an internal quality dashboard. The implementation is internal, so this page covers the design and how such a system is held to account.

Problem statement

Single-prompt automation is brittle. Real QA work spans several steps and tools, needs context from earlier steps, and must be trustworthy enough that engineers act on the output.

Three things make QA a demanding domain for agents:

  • Multi-step work. Running a suite, interpreting failures, checking whether data is the cause, writing a defect and notifying an owner is a chain. An error early in the chain propagates.
  • Real tools with real effects. Agents call test runners, query data stores and create tickets. A wrong action creates noise, or worse, hides a real problem.
  • Trust. If engineers cannot see why an agent filed a defect, they will re-check everything by hand and the automation adds no value.

Engineering objectives

  1. Narrow roles. Each agent does one job with only the tools that job needs.
  2. Explicit hand-offs. Context passes between agents in a defined structure, not as loose conversation.
  3. Observable by default. Every run can be inspected step by step: which agent, which tool, which input, which output.
  4. Held to the same assurance standards as any agentic system: evaluated, reliability-tested and bounded in what it can do.
  5. Useful to humans. Results land in a dashboard the team already uses, with enough evidence to act on.

Solution architecture

01 Agents

  • Testing agent
  • Data validation agent
  • Bug logging agent
  • Notification agent

02 Orchestration

  • CrewAI
  • LangGraph
  • Agno
  • LangChain components

03 Tools

  • MCP tool access
  • Test runners
  • Data checks
  • Issue tracker and messaging

04 Visibility

  • LangSmith traces
  • Internal quality dashboard
Figure 1. Specialised agents, orchestration, tools and visibility. Simplified; implementation details are internal.

Each agent has a narrow role:

  • The testing agent runs and interprets test suites.
  • The data validation agent checks whether data issues explain a failure or are failures in their own right.
  • The bug logging agent turns confirmed problems into well-formed defect reports.
  • The notification agent tells the right people, in the right place.

Different orchestration frameworks suit different shapes of workflow. Graph-based orchestration such as LangGraph makes control flow, branching and state explicit. Role-based frameworks such as CrewAI make it quick to define collaborating agents. MCP gives agents a consistent way to reach tools, which keeps tool integrations separate from agent logic.

Technical approach

Narrow tools, narrow permissions

An agent should only be able to do what its role requires. The notification agent does not need to run tests; the testing agent does not need to create tickets. Scoping tools per agent limits the damage a confused or manipulated agent can do, and makes behaviour easier to reason about.

Structured hand-offs

Hand-offs work best as structured data rather than free text, so the next agent receives exactly the fields it needs. An illustrative shape for a hand-off from the testing agent to the bug logging agent:

{
"run_id": "illustrative-001",
"test": "checkout_applies_discount",
"status": "failed",
"evidence": ["assertion: expected 90.00, got 100.00", "trace link"],
"data_check": "passed",
"suspected_area": "pricing service"
}

With a schema, a malformed hand-off fails loudly at the boundary instead of producing a vague defect report two steps later.

Observability from day one

Agent runs are traced with LangSmith, so each step of a workflow can be inspected. Results surface in an internal quality dashboard, so humans can see what the agents did and why. Traces are also the raw material for evaluating the agents themselves.

Human checkpoints

Some actions benefit from a person in the loop, at least until the agent has earned trust. A common pattern is to let agents draft defect reports and notifications automatically, while keeping a review step for actions that are costly to undo.

Validation strategy

These agents are held to the evaluation and reliability practices applied to any agentic system:

  • Task-level evaluation. Known scenarios, such as a seeded failing test or a deliberately corrupted dataset, check whether the workflow reaches the right conclusion end to end.
  • Tool-call correctness. Checking that each agent calls the right tool with valid arguments, not just that the final output looks plausible.
  • Defect report quality. Generated reports are reviewed for accuracy, reproducibility and duplication against existing tickets.
  • Consistency. Running the same scenario repeatedly shows how much agent behaviour varies, and whether the variation changes the outcome.
  • Failure handling. Tool timeouts, unavailable services and malformed hand-offs are simulated to confirm the workflow fails safely and visibly.

Metrics and measurement

Metrics tracked for an agentic QA workflow typically include:

  • End-to-end task success rate on known scenarios.
  • Tool-call error rate and invalid-argument rate per agent.
  • Precision of filed defects: how many are real, reproducible and not duplicates.
  • Steps, tokens, cost and latency per workflow run.
  • Human override rate at checkpoints, which shows where agents are not yet trusted.

Challenges and trade-offs

  • Autonomy versus control. More autonomy saves more time but increases the cost of mistakes. Checkpoints are a dial, not a switch.
  • Framework choice. No single orchestration framework fits every workflow. Using more than one adds integration effort, but forcing every workflow into one shape adds complexity of a different kind.
  • Noise. An agent that files many low-quality defects costs more time than it saves. Precision matters more than volume.
  • Non-determinism. The same failure can be interpreted differently on different runs. Structured hand-offs and evaluation on fixed scenarios keep this visible.
  • Tool security. Agents with tool access are an attack surface. Least privilege and input handling at tool boundaries apply here as much as in any production agent.

Outcome and lessons learned

Routine QA steps now run as an agentic workflow, with full visibility for the team. Building these systems first-hand also keeps my assurance work practical: I design controls for agents knowing how they are actually built.

The lessons:

  • Narrow agents with narrow tools are easier to build, test and trust than one general agent.
  • Structured hand-offs remove a whole class of silent failures.
  • Observability is not optional. If the team cannot see why an agent acted, they will not rely on it.

Further reading

Try “evaluation”, “red-teaming”, “governance” or “agents”.