Most teams building RAG systems spend 90% of their time on retrieval and generation, and 10% on evaluation. That ratio is backwards. Without a rigorous test harness, you're shipping a black box — and you'll only find out it's broken when a user does.
A RAG pipeline has two distinct failure modes: retrieval failures (the right context isn't fetched) and generation failures (the model hallucinates or ignores the retrieved context). Your evaluation framework needs to catch both independently.
The four metrics that matter in practice:
You can measure these with a reference dataset (ground-truth question/answer/context triplets) or, more scalably, with a language model acting as a judge. Both approaches are worth understanding.
The foundation of any test harness is a dataset of question/answer/context triplets. For each item you need: the question, the correct answer, and the document chunks that should be retrieved.
import json
from dataclasses import dataclass
from typing import List, Optional
@dataclass
class RAGSample:
question: str
ground_truth: str
ground_truth_contexts: List[str]
generated_answer: Optional[str] = None
retrieved_contexts: Optional[List[str]] = None
def load_test_set(path: str) -> List[RAGSample]:
with open(path) as f:
data = json.load(f)
return [RAGSample(**item) for item in data]
Build this dataset by exporting real user queries paired with your best-known answers and the source chunks. Even 50 well-curated samples will catch most regressions. Seed it with edge cases: short questions, ambiguous phrasings, questions that span multiple documents.
Faithfulness is hard to measure with string matching — you need semantic understanding. The standard approach is using a language model as a judge. The judge receives the retrieved context and the generated answer, then returns a score from 0 to 1.
import json
from openai import OpenAI
client = OpenAI()
FAITHFULNESS_PROMPT = (
"You are evaluating whether an answer is faithful to the given context.\n"
"Faithful means: every factual claim in the answer is supported by the context.\n"
"Score from 0.0 (completely unfaithful) to 1.0 (fully faithful).\n\n"
"Context:\n{context}\n\nAnswer:\n{answer}\n\n"
'Respond with JSON: {"score": float, "reason": string}'
)
def score_faithfulness(context: str, answer: str) -> float:
prompt = FAITHFULNESS_PROMPT.format(context=context, answer=answer)
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
response_format={"type": "json_object"},
temperature=0,
)
result = json.loads(resp.choices[0].message.content)
return float(result["score"])
A faithfulness score below 0.8 typically means the retriever is fetching irrelevant chunks and the generator is hallucinating from them — a retrieval bug that looks like a model bug until you measure it properly.
For context recall — "did we retrieve what we needed?" — a simple token overlap heuristic works well as a fast baseline before reaching for a semantic similarity model:
from collections import Counter
import re
from typing import List
def tokenize(text: str) -> Counter:
tokens = re.findall(r'\b\w+\b', text.lower())
return Counter(tokens)
def token_recall(retrieved_chunks: List[str], ground_truth_chunks: List[str]) -> float:
retrieved_text = " ".join(retrieved_chunks)
retrieved_tokens = tokenize(retrieved_text)
total, found = 0, 0
for chunk in ground_truth_chunks:
for token, count in tokenize(chunk).items():
total += count
found += min(count, retrieved_tokens[token])
return found / total if total > 0 else 0.0
This is fast, zero-cost, and good enough to catch retrieval regressions in CI. For production monitoring, replace it with a semantic similarity model such as a bi-encoder from sentence-transformers — token overlap misses synonyms and paraphrases.
import statistics
from typing import Callable, Tuple, List
def run_rag_evaluation(
test_set: List[RAGSample],
rag_pipeline: Callable[[str], Tuple[str, List[str]]],
) -> dict:
faithfulness_scores = []
recall_scores = []
for sample in test_set:
answer, retrieved = rag_pipeline(sample.question)
sample.generated_answer = answer
sample.retrieved_contexts = retrieved
f_score = score_faithfulness(" ".join(retrieved), answer)
r_score = token_recall(retrieved, sample.ground_truth_contexts)
faithfulness_scores.append(f_score)
recall_scores.append(r_score)
print(f"Q: {sample.question[:60]}...")
print(f" Faithfulness: {f_score:.2f} Recall: {r_score:.2f}")
sorted_f = sorted(faithfulness_scores)
results = {
"faithfulness_mean": statistics.mean(faithfulness_scores),
"faithfulness_p10": sorted_f[max(0, len(sorted_f) // 10)],
"recall_mean": statistics.mean(recall_scores),
"n_samples": len(test_set),
}
return results
The p10 faithfulness score — the 10th percentile — matters more than the mean. Your pipeline's worst cases are what end up in user complaints.
Evaluation should run automatically on every PR that touches the retriever or the prompt template. Gate on a hard threshold:
- name: RAG evaluation
run: |
python evaluate_rag.py \
--test-set tests/rag_test_set.json \
--faithfulness-threshold 0.85 \
--recall-threshold 0.80
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
If faithfulness drops below 0.85, the PR fails. This forces anyone changing retrieval logic to justify the tradeoff explicitly. The friction is intentional: silently degrading quality is worse than a blocked PR.
For security-sensitive RAG applications — customer support bots with access to internal docs, compliance tools, incident response assistants — you also need a factual grounding check to prevent the model from leaking information outside the retrieved context. This type of control is part of what we cover in our security hardening checklists.
Once your harness works on 50 samples, three problems appear at scale:
Evaluation drift: the language model judge you use today may score differently after a provider model update. Pin the judge model version and store raw scores as artifacts, not just pass/fail flags.
Dataset staleness: as your knowledge base evolves, ground-truth contexts become outdated. Tag each sample with the document version it was created from, and rebuild affected samples when source documents change.
Cost: at 1000 samples x 2 LLM calls (faithfulness + answer relevance), each evaluation run costs roughly $0.50-$1.00 with gpt-4o-mini. That is acceptable for weekly regression runs but expensive for every commit. Use token overlap metrics in fast CI, LLM judges in nightly runs.
RAG evaluation is not optional — it is what separates a demo from a system you would put in production. Start with a small, curated test set, compute faithfulness and recall on every change, and gate your CI pipeline on minimum thresholds.
The typical failure mode in production RAG is not hallucination at the generation step. It is over-retrieval at the search step: fetching 10 chunks when 2 are relevant actively degrades faithfulness, because the model gets confused by noise. The metrics above will surface this clearly. Measure first, then tune.
I run AYI NEDJIMI Consultants, a cybersecurity consulting firm. We publish free security hardening checklists — PDF and Excel.