{"slug": "how-to-evaluate-rag-pipeline-quality-metrics-and-test-harness", "title": "How to Evaluate RAG Pipeline Quality: Metrics and Test Harness", "summary": "A developer outlines a RAG evaluation harness that separates retrieval failures from generation failures, using a dataset of question/answer/context triplets and four core metrics. The approach uses an LLM-as-judge to score faithfulness — with scores below 0.8 typically indicating the retriever is fetching irrelevant chunks that the generator then hallucinates from — and a zero-cost token-overlap heuristic for context recall that is fast enough to run in CI.", "body_md": "Most teams building RAG systems spend 90% of their time on retrieval and generation, and 10% on evaluation. That ratio is backwards. Without a rigorous test harness, you're shipping a black box — and you'll only find out it's broken when a user does.\n\nA RAG pipeline has two distinct failure modes: retrieval failures (the right context isn't fetched) and generation failures (the model hallucinates or ignores the retrieved context). Your evaluation framework needs to catch both independently.\n\nThe four metrics that matter in practice:\n\nYou can measure these with a reference dataset (ground-truth question/answer/context triplets) or, more scalably, with a language model acting as a judge. Both approaches are worth understanding.\n\nThe foundation of any test harness is a dataset of question/answer/context triplets. For each item you need: the question, the correct answer, and the document chunks that should be retrieved.\n\n``` python\nimport json\nfrom dataclasses import dataclass\nfrom typing import List, Optional\n\n@dataclass\nclass RAGSample:\n    question: str\n    ground_truth: str\n    ground_truth_contexts: List[str]\n    generated_answer: Optional[str] = None\n    retrieved_contexts: Optional[List[str]] = None\n\ndef load_test_set(path: str) -> List[RAGSample]:\n    with open(path) as f:\n        data = json.load(f)\n    return [RAGSample(**item) for item in data]\n```\n\nBuild this dataset by exporting real user queries paired with your best-known answers and the source chunks. Even 50 well-curated samples will catch most regressions. Seed it with edge cases: short questions, ambiguous phrasings, questions that span multiple documents.\n\nFaithfulness is hard to measure with string matching — you need semantic understanding. The standard approach is using a language model as a judge. The judge receives the retrieved context and the generated answer, then returns a score from 0 to 1.\n\n``` python\nimport json\nfrom openai import OpenAI\n\nclient = OpenAI()\n\nFAITHFULNESS_PROMPT = (\n    \"You are evaluating whether an answer is faithful to the given context.\\n\"\n    \"Faithful means: every factual claim in the answer is supported by the context.\\n\"\n    \"Score from 0.0 (completely unfaithful) to 1.0 (fully faithful).\\n\\n\"\n    \"Context:\\n{context}\\n\\nAnswer:\\n{answer}\\n\\n\"\n    'Respond with JSON: {\"score\": float, \"reason\": string}'\n)\n\ndef score_faithfulness(context: str, answer: str) -> float:\n    prompt = FAITHFULNESS_PROMPT.format(context=context, answer=answer)\n    resp = client.chat.completions.create(\n        model=\"gpt-4o-mini\",\n        messages=[{\"role\": \"user\", \"content\": prompt}],\n        response_format={\"type\": \"json_object\"},\n        temperature=0,\n    )\n    result = json.loads(resp.choices[0].message.content)\n    return float(result[\"score\"])\n```\n\nA faithfulness score below 0.8 typically means the retriever is fetching irrelevant chunks and the generator is hallucinating from them — a retrieval bug that looks like a model bug until you measure it properly.\n\nFor context recall — \"did we retrieve what we needed?\" — a simple token overlap heuristic works well as a fast baseline before reaching for a semantic similarity model:\n\n``` python\nfrom collections import Counter\nimport re\nfrom typing import List\n\ndef tokenize(text: str) -> Counter:\n    tokens = re.findall(r'\\b\\w+\\b', text.lower())\n    return Counter(tokens)\n\ndef token_recall(retrieved_chunks: List[str], ground_truth_chunks: List[str]) -> float:\n    retrieved_text = \" \".join(retrieved_chunks)\n    retrieved_tokens = tokenize(retrieved_text)\n\n    total, found = 0, 0\n    for chunk in ground_truth_chunks:\n        for token, count in tokenize(chunk).items():\n            total += count\n            found += min(count, retrieved_tokens[token])\n\n    return found / total if total > 0 else 0.0\n```\n\nThis is fast, zero-cost, and good enough to catch retrieval regressions in CI. For production monitoring, replace it with a semantic similarity model such as a bi-encoder from sentence-transformers — token overlap misses synonyms and paraphrases.\n\n``` python\nimport statistics\nfrom typing import Callable, Tuple, List\n\ndef run_rag_evaluation(\n    test_set: List[RAGSample],\n    rag_pipeline: Callable[[str], Tuple[str, List[str]]],\n) -> dict:\n    faithfulness_scores = []\n    recall_scores = []\n\n    for sample in test_set:\n        answer, retrieved = rag_pipeline(sample.question)\n        sample.generated_answer = answer\n        sample.retrieved_contexts = retrieved\n\n        f_score = score_faithfulness(\" \".join(retrieved), answer)\n        r_score = token_recall(retrieved, sample.ground_truth_contexts)\n\n        faithfulness_scores.append(f_score)\n        recall_scores.append(r_score)\n\n        print(f\"Q: {sample.question[:60]}...\")\n        print(f\"  Faithfulness: {f_score:.2f}  Recall: {r_score:.2f}\")\n\n    sorted_f = sorted(faithfulness_scores)\n    results = {\n        \"faithfulness_mean\": statistics.mean(faithfulness_scores),\n        \"faithfulness_p10\": sorted_f[max(0, len(sorted_f) // 10)],\n        \"recall_mean\": statistics.mean(recall_scores),\n        \"n_samples\": len(test_set),\n    }\n    return results\n```\n\nThe p10 faithfulness score — the 10th percentile — matters more than the mean. Your pipeline's worst cases are what end up in user complaints.\n\nEvaluation should run automatically on every PR that touches the retriever or the prompt template. Gate on a hard threshold:\n\n```\n- name: RAG evaluation\n  run: |\n    python evaluate_rag.py \\\n      --test-set tests/rag_test_set.json \\\n      --faithfulness-threshold 0.85 \\\n      --recall-threshold 0.80\n  env:\n    OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}\n```\n\nIf faithfulness drops below 0.85, the PR fails. This forces anyone changing retrieval logic to justify the tradeoff explicitly. The friction is intentional: silently degrading quality is worse than a blocked PR.\n\nFor security-sensitive RAG applications — customer support bots with access to internal docs, compliance tools, incident response assistants — you also need a factual grounding check to prevent the model from leaking information outside the retrieved context. This type of control is part of what we cover in our [security hardening checklists](https://ayinedjimi-consultants.fr/checklists).\n\nOnce your harness works on 50 samples, three problems appear at scale:\n\n**Evaluation drift**: the language model judge you use today may score differently after a provider model update. Pin the judge model version and store raw scores as artifacts, not just pass/fail flags.\n\n**Dataset staleness**: as your knowledge base evolves, ground-truth contexts become outdated. Tag each sample with the document version it was created from, and rebuild affected samples when source documents change.\n\n**Cost**: at 1000 samples x 2 LLM calls (faithfulness + answer relevance), each evaluation run costs roughly $0.50-$1.00 with gpt-4o-mini. That is acceptable for weekly regression runs but expensive for every commit. Use token overlap metrics in fast CI, LLM judges in nightly runs.\n\nRAG evaluation is not optional — it is what separates a demo from a system you would put in production. Start with a small, curated test set, compute faithfulness and recall on every change, and gate your CI pipeline on minimum thresholds.\n\nThe typical failure mode in production RAG is not hallucination at the generation step. It is over-retrieval at the search step: fetching 10 chunks when 2 are relevant actively degrades faithfulness, because the model gets confused by noise. The metrics above will surface this clearly. Measure first, then tune.\n\n*I run [AYI NEDJIMI Consultants](https://ayinedjimi-consultants.fr), a cybersecurity consulting firm. We publish [free security hardening checklists](https://ayinedjimi-consultants.fr/checklists) — PDF and Excel.*", "url": "https://wpnews.pro/news/how-to-evaluate-rag-pipeline-quality-metrics-and-test-harness", "canonical_source": "https://dev.to/ayinedjimi-consultants/how-to-evaluate-rag-pipeline-quality-metrics-and-test-harness-4n24", "published_at": "2026-09-29 10:04:25+00:00", "updated_at": "2026-09-29 10:17:04.763975+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "mlops", "ai-agents"], "entities": ["OpenAI", "GPT-4o-mini", "sentence-transformers"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-to-evaluate-rag-pipeline-quality-metrics-and-test-harness", "markdown": "https://wpnews.pro/news/how-to-evaluate-rag-pipeline-quality-metrics-and-test-harness.md", "text": "https://wpnews.pro/news/how-to-evaluate-rag-pipeline-quality-metrics-and-test-harness.txt", "jsonld": "https://wpnews.pro/news/how-to-evaluate-rag-pipeline-quality-metrics-and-test-harness.jsonld"}}