Proving a RAG system is grounded
Overview
Retrieval-augmented generation is often presented as the cure for hallucination. In practice, a RAG system can fail at every stage: the wrong chunks are retrieved, the right chunks are split badly, embeddings miss the meaning, or the model ignores good context and answers from memory.
I designed end-to-end evaluation pipelines for LLM and RAG products, with results enforced as quantitative release gates and made visible through a custom RAG evaluation dashboard. Production hallucination incidents fell by 40%. The products involved are confidential, so this page focuses on the method.
Problem statement
Judged on demos and spot checks, a RAG system gives no way to tell which stage failed when an answer is wrong, and no objective bar a release must clear.
That creates three practical problems:
- No diagnosis. A wrong answer could come from retrieval, chunking, embedding or generation. Without stage-level measurement, teams guess, and fixes often target the wrong stage.
- No baseline. Without a fixed reference set, “better” and “worse” depend on whoever happened to try the system that week.
- No gate. If quality is not quantified, it cannot block a release. Regressions ship and are found by users.
Engineering objectives
- Evaluate every stage, not just the final answer, so a failure can be traced to where it started.
- Use a stable reference in the form of golden datasets, so every run is compared against the same evidence.
- Turn quality into numbers that can be compared across versions and enforced as thresholds.
- Combine automated and human judgement, using LLM-as-a-judge where it is reliable and people where it is not.
- Make results visible to engineering and product, not only to the assurance team.
Solution architecture
The pipeline follows the life of a question through the system, measuring each stage before the final answer is scored.
- 01Golden dataset
- 02Retrieval metrics
- 03Chunking and embedding checks
- 04Generation metrics
- 05Release gate
- 06Dashboard
The components are deliberately simple:
- Golden datasets: curated questions paired with the evidence that should support the answer, and where useful, a reference answer.
- Evaluators: DeepEval, RAGAS and Promptfoo, each used where it is strongest, behind one pipeline so results land in a common format.
- Gate: a step in the release process that compares results against agreed thresholds and fails the build if they are not met.
- Dashboard: a custom RAG evaluation view that shows scores per stage, per version and per question category.
Technical approach
Measure what matters at each stage
| Stage | What can go wrong | Typical measures |
|---|---|---|
| Retrieval | Relevant evidence not found, or buried under noise | Context precision, context recall |
| Chunking | Evidence split across chunks or diluted by unrelated text | Recall per chunking strategy, inspection of failing cases |
| Embedding | Semantically similar questions retrieve different evidence | Retrieval comparison across embedding models |
| Generation | Answer ignores or contradicts context, or adds unsupported claims | Faithfulness, groundedness, answer relevance, hallucination detection |
The value comes from reading these together. High context recall with low faithfulness points at generation. Low context recall with high faithfulness usually means the model is faithfully answering from the wrong evidence.
Golden datasets
Golden datasets are the reference point for every run. Good ones cover the question categories users actually ask, include questions the system should decline, and include cases where the evidence is ambiguous. They are versioned alongside the code, because a change to the dataset changes what the scores mean.
LLM-as-a-judge with human review
Metrics such as faithfulness rely on a judge model. Judges are fast and consistent but not infallible, so the approach combines LLM-as-a-judge scoring with human-in-the-loop review for borderline cases and for periodic calibration of the judge itself.
Thresholds as configuration
Gates are easiest to maintain when thresholds are declared, not buried in code. An illustrative example:
# Example values only, not recommended defaults.gate: dataset: golden/v12 thresholds: context_recall: 0.80 context_precision: 0.70 faithfulness: 0.90 answer_relevance: 0.80 fail_on: any_below_thresholdValidation strategy
An evaluation pipeline is itself software, and it can be wrong. Validation covers:
- Known-bad cases. Deliberately ungrounded answers are included to confirm that the evaluators catch them.
- Judge calibration. Judge scores are compared with human labels on a sample, so drift in the judge is spotted before it moves the gate.
- Stability. The same version is evaluated more than once to understand how much scores vary between runs, which informs how strict a threshold can sensibly be.
- Production feedback. Real incidents are traced back to the stage that failed and, where useful, added to the golden dataset so the same failure is caught next time.
Metrics and measurement
Two kinds of measurement matter:
- Pre-release: stage-level scores per version, pass or fail against thresholds, and the specific questions that regressed.
- Production: hallucination incidents reported or detected after release. This is the number that tells you whether the pre-release gate is measuring the right things.
The custom RAG evaluation dashboard, built on this stage-level evaluation, reduced production hallucination incidents by 40%.
Challenges and trade-offs
- Cost of evaluation. Judge-based metrics use model calls. A typical trade-off is running the full suite on release candidates and a smaller subset on every change.
- Threshold strictness. Too strict and the gate blocks good releases because of scoring noise. Too loose and it lets regressions through. Thresholds are revisited as the dataset and judge mature.
- Metric disagreement. Different tools define similar-sounding metrics differently. Using one tool per metric, and documenting which, avoids comparing numbers that are not comparable.
- Dataset maintenance. Golden datasets go stale as content and user questions change. They need owners and a review cadence.
Outcome and lessons learned
Production hallucination incidents fell by 40%. Just as important, quality conversations changed: instead of “it looked fine in the demo”, teams could point to which stage regressed and by how much.
The lessons from this work:
- Measure the stages, not just the answer. Diagnosis is what makes improvement fast.
- A release gate changes behaviour more than a report does.
- Visibility matters. A dashboard that product owners actually open does more for quality than a perfect metric nobody sees.