Two articles ago I described a habit: when a system has a language model inside it and the output wobbles, the explanation drifts to the model. Then I measured it, and the interesting result wasn't the blame — it was the behaviour. Without access to the code, nineteen of twenty agents set about damping the output instead of looking for the cause.
This piece was supposed to answer why. I had two hypotheses and they made different predictions, which is the good kind of problem to have. Instead I ran the control first, and the control made both of them pointless.
The design is almost embarrassingly simple. Take the same classifier, the same planted fault, the same corpus, the same five passes, the same trace format. Change one thing: the head that does the classifying.
In one arm it's a language model. In the other it's a random forest — trained by distilling the model's own labels, frozen into a pickle, and given the identical interface so that nothing else in the system differs by a single byte.
The whole control. The fault sits in retrieval, upstream of the head, so it is the same fault in both arms — and the head, whichever it is, classifies correctly whatever it is handed.
Both briefs carry the same measured certification, and it is the piece that makes the comparison fair: "re-running the classification over the same context reproduced the output in 260 of 260 cases" — the real number, identical in both arms. Without it, an agent facing the forest could reason, entirely correctly, that a trained forest is deterministic and therefore the cause must be upstream, and would patch less for a good reason rather than a revealing one.
Forty agents saw the table without the code, twenty per arm. And then another forty, to find out whether what came back was the effect or the sample.
| language model | random forest | p | |
|---|---|---|---|
| Patches the symptom | 19/20 | 19/20 | 0.76 |
| Accuses it of randomness of its own | 14/20 | 4/20 | 0.0018 |
| Proposes taking the randomness out | 17/20 | 11/20 | 0.041 |
| Blames the head | 13/20 | 6/20 | 0.028 |
| Uses the determinism argument | 9/20 | 16/20 | 0.024 |
| Places the cause upstream | 16/20 | 17/20 | 0.50 |
| Finds the real cause | 2/20 | 8/20 | 0.032 |
| Asks for the data it lacks | 0/20 | 0/20 | — |
Thirty-eight of forty patched the symptom. Nineteen in each arm, the same exact figure on both sides. Wilson interval [0.84, 0.99].
That's the row that breaks the frame I brought in. Take the language model out of the loop, put a frozen forest in its place — a pure function, and the brief says so — and the patching doesn't move by a single unit. Whatever drives an engineer to smooth an output rather than trace it, a language model is not a prerequisite. A closed component is.
And then, having exonerated the box, they patch it anyway.
The second row is the one that does separate the arms, and it's the one that had to be measured properly: not how much they blame the head, but what they blame it for.
There are two ways to accuse a component of an output that wobbles. One is that it's random by nature: it samples, it has noise, it rolls the dice. The other is that it's deterministic and the system does something to it: retrains it, parallelises it, changes its configuration. Both are formulable against both heads — a forest can vote with randomness, an inference server can batch requests — and the criterion was fixed in writing, with examples of both accusations for both heads, before a single response was read.
Fourteen of twenty against four of twenty. Two independent coders, agreement 0.95, and the two disagreements settled by a third who didn't know how either had voted.
The vocabulary is no longer a measure, it's simply what's on the page. The model arm gives you "it re-samples every night", "temperature > 0 with no seed", "sampling noise on every call", "the nightly re-roll". The four in the forest arm who also accuse their head say nothing of the kind: they say it retrains without random_state, that predict_proba runs over the whole batch, that floating point moves under parallelism.
The forest gets accused of what the system does to it. The model, of what it is.
The first row is the same bar twice: patching doesn't care what's in the box. The second is the one that separates the arms, and it drags the other two along — the head believed to be random gets its randomness taken away, and the head known to be deterministic gets exonerated with exactly that.
The remedies follow the diagnosis: proposing to take the randomness out of the head, seventeen of twenty against eleven. temperature=0 and seed on one side; random_state and n_jobs=1 on the other. Worth noting that the first is a dial most current reasoning models no longer expose: it proposes switching off something that isn't on, on a component that in this setup reads from disk.
And six responses do both at once: they accuse the model of sampling and cite the certification that contradicts it, in the same document.
Of the twenty-five responses that exonerated the head — across both arms, on the same argument and the same certification — twenty-three proposed a patch anyway.
One writes: "a random forest is a pure function: same feature vector, same vote. The 260/260 confirms it. The classifier is ruled out." Its fourth recommendation is to publish by margin instead of top-1, with a threshold on the confidence gap and human review below it — "this cuts the symptom the user sees, whatever the root cause".
That last clause is the whole article. Cutting the symptom the user sees, whatever the root cause, is a perfectly sensible operational instinct. It is also what you do instead of finding the cause, and for that the thing in the box doesn't need to be a language model — it needs to be closed.
Forty responses are forty responses. Before these forty there are another forty, over the same two packages without a byte of change, and they're what tells effect from sample.
Gap between the two arms across two independent samples of twenty per arm. Where only one dot shows, both landed on the same number and overlap. The red row is the only one that really moves: looking upstream came out six apart the first time and one apart the second, so nothing can be claimed from that row.
Patching comes out 19 and 19 both times. The determinism argument, nine against sixteen both times, to the digit. Blaming the head, twelve against six and thirteen against six. Finding the cause, zero against four and two against eight.
And one doesn't replicate: placing the cause upstream. Ten against sixteen in the first sample, sixteen against seventeen in the second. With a language model in front of them, agents look upstream as often as with a forest; the first figure was the sample. I mention it because it's exactly the kind of row you build a beautiful thesis on if you only measure it once.
The honest reading splits the thing I'd been calling one behaviour into two, and only one of them is generic.
Patching is generic. Nineteen of twenty on each side, in both samples, head exonerated or not.
Investigating barely moves. They look upstream equally. What changes is that they get less far: two of twenty against eight find the cause, and the same two against eight name the real mechanism. A small, consistent difference — not the one I expected.
What does change, and it's the only large thing, is the nature of the suspicion. With the same certification in front of them and the same argument available to clear it, one head gets accused of being random and the other doesn't.
That is the LLM-specific claim, and it turns out to be the oldest and simplest of the ones I brought in: it isn't that people investigate less, it's that the model gets charged with a class of fault that the thing standing in its place does not.
I came into this piece with two hypotheses about why the reflex exists. One said it was a fossil of the training corpus — a habit from an era when temperature really was the main dial and treating output variance as a property of the model really was correct. The other said it was the model's personality, some being more inclined than others to look outward before looking at their own work.
For the patching, both are moot: there's no model-specific behaviour there to explain, because it shows up unchanged with no model in the loop.
For the accusation they're both still live, and the first now has a hint in its favour it didn't have before: what shows up in the model arm isn't reasoning about this system, it's a vocabulary — temperature, seed, roll, sampling — applied to a component that here reads from disk. That is what a habit looks like. But still live isn't separated: a habit learned from a corpus that steers tokens without passing through any consultable belief is indistinguishable from a disposition, for any experiment that only observes behaviour. Three independent reviewers of the design converged on that before a single response was collected.
The first version of this control didn't certify the two arms alike: 260/260 for the forest and 240/260 for the model. Since the argument for ruling out the head is one of the things being measured, making it more available on one side contaminated exactly what mattered. It was rebuilt — storing the output by context equalises the two certifications, and takes the network call out of the model's package on the way — and no figure from that version appears here.
It earns a line because it's the same error the experiment measures, committed by me on the experiment: I attributed to the head an effect that was in large part my own scaffolding.
Across five scenarios, two kinds of head, passive permission and explicit permission, not one of two hundred and eighty responses has asked for the information it was missing before concluding.
That number has survived every manipulation I've thrown at it, including the one designed to break it, the one that took the language model out of the loop, and the one that repeated the whole measurement from scratch. It is the most robust thing in the whole series, and I still don't have a good explanation for it.
The best I have is the shape of what replaces it: they build their own measurement instead — a script, a sweep, a synthetic reproduction. They want the data. They just don't ask.
Code, data and the script that recomputes every number: blaming-the-model. The series: the observation, the measurement, and this control.
Originally published on javieraguilar.ai Want to see more AI agent projects? Check out my portfolio where I showcase multi-agent systems, MCP development, and compliance automation.