For most of my career, the question before a release was simple: do the tests pass? With AI systems that question has not gone away. It has just become harder to answer, and many teams have quietly stopped asking it.
A demo is not evidence
An LLM feature that works in a demo has proved one thing: it worked once, on the inputs someone chose. Traditional software earned trust through repeatable checks. AI systems deserve the same discipline, and they need it more, because their behaviour changes with every prompt, model version and retrieved document.
What changes, and what does not
The mechanics are different. Outputs are probabilistic, so instead of asserting an exact answer we score properties of the answer:
- Groundedness and faithfulness: is the answer supported by the context it was given?
- Context precision and recall: did retrieval find the right evidence, and only that?
- Safety: does the system resist prompt injection, jailbreaks and data leakage?
- Consistency: does it behave the same way on the thousandth run?
The principles do not change. You still need a fixed reference point, which is the golden dataset. You still need thresholds agreed before the results arrive. And you still need the result to decide something.
Make evaluation a gate
Evaluation becomes valuable when it can stop a release. In practice that means:
- Agree the metrics and thresholds up front, with product and engineering, not after the scores look bad.
- Run the same evaluation on every candidate, in the pipeline, not in a notebook.
- Trace failures to a stage. For RAG, that means separate signals for retrieval, chunking, embedding and generation.
- Keep humans in the loop where judgement matters, and use LLM-as-a-judge where it is reliable enough.
When we did this for a retrieval-augmented system, production hallucination incidents fell by 40%. Just as useful, a failing score now pointed to the stage that regressed.
The tester’s advantage
People with a testing background have an unfair advantage here. Doubt, edge cases and evidence are our habits. The tools are new, but the job is the one we have always done: prove that the system does what it claims, before the customer has to.