# My Eval Said RAG Made Things Up. My Eval Was Wrong.

> Source: <https://dev.to/topunix/my-eval-said-rag-made-things-up-my-eval-was-wrong-47e>
> Published: 2026-10-08 00:11:55+00:00

I maintain [django-explain-errors](https://github.com/topunix/django-explain-errors), a Django middleware that catches unhandled exceptions in development and asks an LLM to explain them. It works in two modes:

The whole argument for RAG-on is grounding. If the model can read your actual code, it should stop guessing at function names and parameters. So I built an eval harness to test that claim, and the first results said the opposite: RAG-on appeared to fabricate *more* than RAG-off.

It took three versions of a single eval question to find out why. The short version: every layer in an eval pipeline can be working from less information than the thing it judges. When that happens, your metric measures the judge's blind spot, not the model.

The harness runs against a small, deliberately breakable Django blog app with 15 fixtures. Each fixture is a URL that triggers one realistic failure: a bad queryset lookup, a missing null check, a typo'd template name, a recursive `__str__`. Ten of them have their cause in the app's own code. Five have a cause the traceback already names.

For each fixture, the harness gets two explanations from `gpt-4o-mini`, one with RAG and one without. A separate judge model (Claude Sonnet via OpenRouter) compares them pairwise. It doesn't know which explanation used RAG, and which one it sees as "A" is randomized every time. The judge answers a fixed set of questions: did it find the cause, point to the fix location, propose a working fix, and explain it for a learner?

Notice what's missing from that list.

I didn't ask about fabrication at first. But reading the judge's reasoning, I kept seeing it penalize explanations for "inventing" details, usually when the other questions came out tied. The judge was deciding fabrication on its own, as an unstated tiebreaker.

That meant fabrication was affecting the scores without showing up anywhere I could measure or audit. So I added it as an explicit question: `no_fabrication`.

The first `no_fabrication` judge saw the traceback and a short list of known facts about each fixture: the real cause and the real fix location. Its question was whether the explanation stated anything not supported by those.

RAG-on lost. It consistently got flagged for inventing details.

Then I looked at what was flagged. In the `unexpected_kwarg` fixture, RAG-on's explanation named a `post_id` parameter, and the judge called it fabricated. It wasn't. `post_id` was right there in the view's source, which RAG-on had retrieved and the judge had never seen.

The judge wasn't detecting fabrication. It was detecting *specificity it couldn't verify*. Any concrete detail that came from source looked invented, because the judge's only reference was the traceback. So the metric punished RAG-on for exactly the thing RAG is supposed to do well.

I reworded the criterion from "does not appear in" to "does not contradict," which helped at the edges. But the core problem remained: a judge that knows less than the generator will score the generator's extra knowledge as error.

The obvious fix was to give the judge the source. So I gave it the whole source modules involved in each failure.

It ignored them. The judge's answers barely changed, and its reasoning showed why: faced with 300 mostly irrelevant lines, marking a detail "unverified" is easier than finding the one line that settles it. The source was in the context window. It just wasn't being used.

What finally worked was changing the shape of the task, not the amount of input.

The judge now gets a small, relevant source excerpt: the failing function itself, extracted by line number, plus any sibling methods it reaches directly. Instead of a yes/no on fabrication, it must list every checkable claim in each explanation and give each one a verdict:

`no_fabrication` is then *derived* from those verdicts in the parser, not asked as a direct question. The judge can't hand-wave anymore. Every "contradicted" points at a specific claim, which I can check.

With this in place, the result flipped. On the 10 fixtures whose cause lives in app code, RAG-on avoided fabrication in 26 comparisons against RAG-off's 17. On the other 5, it was 14 against 5. RAG-on hadn't been fabricating more. The judge had been unable to check its work.

That second number surprised me in a different way. I designed those 5 fixtures as a control: the traceback already names the cause, so RAG shouldn't matter. It mattered more there than anywhere. The eval was testing my assumptions, not just the model.

What RAG-off fabricates is revealing. Without source, `gpt-4o-mini` tends to invent a plausible function signature rather than say it doesn't know. In one run, RAG-off confidently wrote `def post_preview(request):` when the real signature is `def post_preview(request, post_id)`. In another, it said a template needed `{% load humanize %}` added when line 1 already had it.

While chasing fabrication, I found a second problem affecting a different metric.

RAG-on was winning `points_to_fix_location` by a wide margin. But the package truncated long tracebacks by keeping only the tail. For a Django ORM error, the tail is mostly library internals, so the frame naming the failing app function routinely got cut. RAG-off often never saw *where* the error happened.

I didn't find this by looking at scores. I found it by looking at what RAG-off was actually sent.

After fixing truncation to preserve app frames, RAG-off's fix-location score on app-code fixtures rose from 6 of 30 to 13. RAG-on stayed roughly where it was. About half of the original gap was a truncation bug. The other half is real. And because the truncation lives in the middleware, the fix didn't just correct the eval. It made the package better for everyone using it, with or without RAG.

The final numbers across three runs: RAG-on won 35 of 43 scored comparisons, RAG-off 5, with 3 ties. A full three-run pass costs about $1.37, almost all of it the judge writing out claim lists. But there are limits I can't wave away:

`missing_post_key` fixture, RAG-on anchored on an adjacent retrieved template and pointed the fix there instead of at the view. Retrieval reduces fabrication. It doesn't eliminate wrong answers.
The harness, fixtures, judge prompt, and full results are in the repo under [`evals/`](https://github.com/topunix/django-explain-errors/tree/main/evals), with the history of these changes in `evals/METHODOLOGY.md`.
