How to Evaluate RAG Pipeline Quality: Metrics and Test Harness A developer outlines a RAG evaluation harness that separates retrieval failures from generation failures, using a dataset of question/answer/context triplets and four core metrics. The approach uses an LLM-as-judge to score faithfulness — with scores below 0.8 typically indicating the retriever is fetching irrelevant chunks that the generator then hallucinates from — and a zero-cost token-overlap heuristic for context recall that is fast enough to run in CI. Most teams building RAG systems spend 90% of their time on retrieval and generation, and 10% on evaluation. That ratio is backwards. Without a rigorous test harness, you're shipping a black box — and you'll only find out it's broken when a user does. A RAG pipeline has two distinct failure modes: retrieval failures the right context isn't fetched and generation failures the model hallucinates or ignores the retrieved context . Your evaluation framework needs to catch both independently. The four metrics that matter in practice: You can measure these with a reference dataset ground-truth question/answer/context triplets or, more scalably, with a language model acting as a judge. Both approaches are worth understanding. The foundation of any test harness is a dataset of question/answer/context triplets. For each item you need: the question, the correct answer, and the document chunks that should be retrieved. python import json from dataclasses import dataclass from typing import List, Optional @dataclass class RAGSample: question: str ground truth: str ground truth contexts: List str generated answer: Optional str = None retrieved contexts: Optional List str = None def load test set path: str - List RAGSample : with open path as f: data = json.load f return RAGSample item for item in data Build this dataset by exporting real user queries paired with your best-known answers and the source chunks. Even 50 well-curated samples will catch most regressions. Seed it with edge cases: short questions, ambiguous phrasings, questions that span multiple documents. Faithfulness is hard to measure with string matching — you need semantic understanding. The standard approach is using a language model as a judge. The judge receives the retrieved context and the generated answer, then returns a score from 0 to 1. python import json from openai import OpenAI client = OpenAI FAITHFULNESS PROMPT = "You are evaluating whether an answer is faithful to the given context.\n" "Faithful means: every factual claim in the answer is supported by the context.\n" "Score from 0.0 completely unfaithful to 1.0 fully faithful .\n\n" "Context:\n{context}\n\nAnswer:\n{answer}\n\n" 'Respond with JSON: {"score": float, "reason": string}' def score faithfulness context: str, answer: str - float: prompt = FAITHFULNESS PROMPT.format context=context, answer=answer resp = client.chat.completions.create model="gpt-4o-mini", messages= {"role": "user", "content": prompt} , response format={"type": "json object"}, temperature=0, result = json.loads resp.choices 0 .message.content return float result "score" A faithfulness score below 0.8 typically means the retriever is fetching irrelevant chunks and the generator is hallucinating from them — a retrieval bug that looks like a model bug until you measure it properly. For context recall — "did we retrieve what we needed?" — a simple token overlap heuristic works well as a fast baseline before reaching for a semantic similarity model: python from collections import Counter import re from typing import List def tokenize text: str - Counter: tokens = re.findall r'\b\w+\b', text.lower return Counter tokens def token recall retrieved chunks: List str , ground truth chunks: List str - float: retrieved text = " ".join retrieved chunks retrieved tokens = tokenize retrieved text total, found = 0, 0 for chunk in ground truth chunks: for token, count in tokenize chunk .items : total += count found += min count, retrieved tokens token return found / total if total 0 else 0.0 This is fast, zero-cost, and good enough to catch retrieval regressions in CI. For production monitoring, replace it with a semantic similarity model such as a bi-encoder from sentence-transformers — token overlap misses synonyms and paraphrases. python import statistics from typing import Callable, Tuple, List def run rag evaluation test set: List RAGSample , rag pipeline: Callable str , Tuple str, List str , - dict: faithfulness scores = recall scores = for sample in test set: answer, retrieved = rag pipeline sample.question sample.generated answer = answer sample.retrieved contexts = retrieved f score = score faithfulness " ".join retrieved , answer r score = token recall retrieved, sample.ground truth contexts faithfulness scores.append f score recall scores.append r score print f"Q: {sample.question :60 }..." print f" Faithfulness: {f score:.2f} Recall: {r score:.2f}" sorted f = sorted faithfulness scores results = { "faithfulness mean": statistics.mean faithfulness scores , "faithfulness p10": sorted f max 0, len sorted f // 10 , "recall mean": statistics.mean recall scores , "n samples": len test set , } return results The p10 faithfulness score — the 10th percentile — matters more than the mean. Your pipeline's worst cases are what end up in user complaints. Evaluation should run automatically on every PR that touches the retriever or the prompt template. Gate on a hard threshold: - name: RAG evaluation run: | python evaluate rag.py \ --test-set tests/rag test set.json \ --faithfulness-threshold 0.85 \ --recall-threshold 0.80 env: OPENAI API KEY: ${{ secrets.OPENAI API KEY }} If faithfulness drops below 0.85, the PR fails. This forces anyone changing retrieval logic to justify the tradeoff explicitly. The friction is intentional: silently degrading quality is worse than a blocked PR. For security-sensitive RAG applications — customer support bots with access to internal docs, compliance tools, incident response assistants — you also need a factual grounding check to prevent the model from leaking information outside the retrieved context. This type of control is part of what we cover in our security hardening checklists https://ayinedjimi-consultants.fr/checklists . Once your harness works on 50 samples, three problems appear at scale: Evaluation drift : the language model judge you use today may score differently after a provider model update. Pin the judge model version and store raw scores as artifacts, not just pass/fail flags. Dataset staleness : as your knowledge base evolves, ground-truth contexts become outdated. Tag each sample with the document version it was created from, and rebuild affected samples when source documents change. Cost : at 1000 samples x 2 LLM calls faithfulness + answer relevance , each evaluation run costs roughly $0.50-$1.00 with gpt-4o-mini. That is acceptable for weekly regression runs but expensive for every commit. Use token overlap metrics in fast CI, LLM judges in nightly runs. RAG evaluation is not optional — it is what separates a demo from a system you would put in production. Start with a small, curated test set, compute faithfulness and recall on every change, and gate your CI pipeline on minimum thresholds. The typical failure mode in production RAG is not hallucination at the generation step. It is over-retrieval at the search step: fetching 10 chunks when 2 are relevant actively degrades faithfulness, because the model gets confused by noise. The metrics above will surface this clearly. Measure first, then tune. I run AYI NEDJIMI Consultants https://ayinedjimi-consultants.fr , a cybersecurity consulting firm. We publish free security hardening checklists https://ayinedjimi-consultants.fr/checklists — PDF and Excel.