cd /news/ai-research/my-eval-said-rag-made-things-up-my-e… · home › topics › ai-research › article
[ARTICLE · art-147231] src=dev.to ↗ pub= topic=ai-research verified=true sentiment=· neutral

My Eval Said RAG Made Things Up. My Eval Was Wrong.

A developer maintaining django-explain-errors, a Django middleware that uses an LLM to explain unhandled exceptions, found that their RAG evaluation harness falsely flagged retrieval-augmented explanations as fabricated because the judge model had less context than the generator. The judge was penalizing specificity it couldn't verify—such as a post_id parameter that appeared in retrieved source code but not in the traceback—so the developer restructured the eval to feed the judge targeted source excerpts and derive a no_fabrication score from per-claim verdicts rather than a direct yes/no question.

by read5 min views5 publishedOct 8, 2026

I maintain django-explain-errors, a Django middleware that catches unhandled exceptions in development and asks an LLM to explain them. It works in two modes:

The whole argument for RAG-on is grounding. If the model can read your actual code, it should stop guessing at function names and parameters. So I built an eval harness to test that claim, and the first results said the opposite: RAG-on appeared to fabricate more than RAG-off.

It took three versions of a single eval question to find out why. The short version: every layer in an eval pipeline can be working from less information than the thing it judges. When that happens, your metric measures the judge's blind spot, not the model.

The harness runs against a small, deliberately breakable Django blog app with 15 fixtures. Each fixture is a URL that triggers one realistic failure: a bad queryset lookup, a missing null check, a typo'd template name, a recursive __str__. Ten of them have their cause in the app's own code. Five have a cause the traceback already names.

For each fixture, the harness gets two explanations from gpt-4o-mini, one with RAG and one without. A separate judge model (Claude Sonnet via OpenRouter) compares them pairwise. It doesn't know which explanation used RAG, and which one it sees as "A" is randomized every time. The judge answers a fixed set of questions: did it find the cause, point to the fix location, propose a working fix, and explain it for a learner? Notice what's missing from that list.

I didn't ask about fabrication at first. But reading the judge's reasoning, I kept seeing it penalize explanations for "inventing" details, usually when the other questions came out tied. The judge was deciding fabrication on its own, as an unstated tiebreaker.

That meant fabrication was affecting the scores without showing up anywhere I could measure or audit. So I added it as an explicit question: no_fabrication.

The first no_fabrication judge saw the traceback and a short list of known facts about each fixture: the real cause and the real fix location. Its question was whether the explanation stated anything not supported by those.

RAG-on lost. It consistently got flagged for inventing details.

Then I looked at what was flagged. In the unexpected_kwarg fixture, RAG-on's explanation named a post_id parameter, and the judge called it fabricated. It wasn't. post_id was right there in the view's source, which RAG-on had retrieved and the judge had never seen.

The judge wasn't detecting fabrication. It was detecting specificity it couldn't verify. Any concrete detail that came from source looked invented, because the judge's only reference was the traceback. So the metric punished RAG-on for exactly the thing RAG is supposed to do well.

I reworded the criterion from "does not appear in" to "does not contradict," which helped at the edges. But the core problem remained: a judge that knows less than the generator will score the generator's extra knowledge as error.

The obvious fix was to give the judge the source. So I gave it the whole source modules involved in each failure.

It ignored them. The judge's answers barely changed, and its reasoning showed why: faced with 300 mostly irrelevant lines, marking a detail "unverified" is easier than finding the one line that settles it. The source was in the context window. It just wasn't being used.

What finally worked was changing the shape of the task, not the amount of input.

The judge now gets a small, relevant source excerpt: the failing function itself, extracted by line number, plus any sibling methods it reaches directly. Instead of a yes/no on fabrication, it must list every checkable claim in each explanation and give each one a verdict:

no_fabrication is then derived from those verdicts in the parser, not asked as a direct question. The judge can't hand-wave anymore. Every "contradicted" points at a specific claim, which I can check.

With this in place, the result flipped. On the 10 fixtures whose cause lives in app code, RAG-on avoided fabrication in 26 comparisons against RAG-off's 17. On the other 5, it was 14 against 5. RAG-on hadn't been fabricating more. The judge had been unable to check its work.

That second number surprised me in a different way. I designed those 5 fixtures as a control: the traceback already names the cause, so RAG shouldn't matter. It mattered more there than anywhere. The eval was testing my assumptions, not just the model.

What RAG-off fabricates is revealing. Without source, gpt-4o-mini tends to invent a plausible function signature rather than say it doesn't know. In one run, RAG-off confidently wrote def post_preview(request): when the real signature is def post_preview(request, post_id). In another, it said a template needed {% load humanize %} added when line 1 already had it.

While chasing fabrication, I found a second problem affecting a different metric. RAG-on was winning points_to_fix_location by a wide margin. But the package truncated long tracebacks by keeping only the tail. For a Django ORM error, the tail is mostly library internals, so the frame naming the failing app function routinely got cut. RAG-off often never saw where the error happened.

I didn't find this by looking at scores. I found it by looking at what RAG-off was actually sent.

After fixing truncation to preserve app frames, RAG-off's fix-location score on app-code fixtures rose from 6 of 30 to 13. RAG-on stayed roughly where it was. About half of the original gap was a truncation bug. The other half is real. And because the truncation lives in the middleware, the fix didn't just correct the eval. It made the package better for everyone using it, with or without RAG.

The final numbers across three runs: RAG-on won 35 of 43 scored comparisons, RAG-off 5, with 3 ties. A full three-run pass costs about $1.37, almost all of it the judge writing out claim lists. But there are limits I can't wave away:

missing_post_key fixture, RAG-on anchored on an adjacent retrieved template and pointed the fix there instead of at the view. Retrieval reduces fabrication. It doesn't eliminate wrong answers. The harness, fixtures, judge prompt, and full results are in the repo under evals/, with the history of these changes in evals/METHODOLOGY.md.

── more in #ai-research 4 stories · sorted by recency
── more on @django-explain-errors 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/my-eval-said-rag-mad…] indexed:0 read:5min 2026-10-08 · —