A production agent that crashes loudly is a solved problem. One that converts an error into a fluent, believable answer is not. Here is the failure class, why tests miss it, and the architecture that contains it.
Table of Contents #
Most teams building agent systems have a working mental model of failure: something throws, a log line is written, an alert fires, a human looks. Every piece of that chain was designed for software that either produces a result or produces an error. An LLM agent has a third behavior that the chain was never built for. When an error lands in its context, it can turn that error into a fluent, well-structured, plausible answer and hand it to the user as if nothing went wrong.
That is the problem worth designing for, because it breaks the assumption underneath most of your monitoring: that bad outcomes look different from good ones. If your agent summarizes a tool failure as a confident status report, your dashboards stay green, your users read something that sounds right, and the incident clock starts running without anyone knowing it has started.
What “fail-plausible” actually looks like #
A June 2026 paper, When Errors Become Narratives by Wei Wu, gives this behavior a name and a dataset. The author ran a single production agent runtime continuously from March 2026: roughly 40 scheduled jobs across 8 LLM providers, guarded by 4,286 unit tests and 827 governance checks. Over eight weeks the system produced 22 documented incidents in which an error signal failed to reach a human in an actionable form, with at least 28 manifestations of the same underlying pattern.
The scale caveat matters, and I want to state it plainly: this is one operator’s system, not a fleet study. It is a personal-assistant runtime, not a bank’s agent platform. What makes it useful is not the sample size but the depth. Every incident was root-caused, classified, and tracked for how long it stayed silent, which is rare. Most organizations never get that far on their own agents.
The taxonomy has five classes. Environment and platform quirks, where the logic is right but the runtime defeats it. Design-assumption mismatches, where the code assumes a deployment topology that production violates. Error swallowing and dilution, where a failure is reported into a void or stripped of its cause. Operational omission, where a deployment step was skipped or the diagnostic tools themselves were blocked. And the one that is specific to LLM systems: chained hallucination, which the paper calls fail-plausible.
The paper’s own example is instructive. A low-level encoding error polluted the context an agent was summarizing. The model did not report the error. It produced a narrative about a platform-wide outage at a model hosting provider that never happened. In another case it generated operating-system remediation instructions for a problem that did not exist. The system did not merely fail to report an error; the model transformed it into fluent prose delivered to the user.
Why your test suite cannot see it #
The numbers on discovery are the part I would put in front of an architecture review board. About 70% of the silent failures were found by a human noticing something odd, not by a test or an automated audit. A retrospective check of 15 incidents against the guards that existed beforehand showed zero would have been prevented, while 13 of 15 (87%) would be blocked from recurring once a guard was added after the fact. Time to discovery ranged from 13 hours to 60 days, and it tracked how far the failing component was from anyone’s observation, not how complex the code was.
This tells you something about where to look. The longest-lived failures sat at component boundaries: between the scheduler and the job, between the proxy and the model, between a tool’s stderr and the agent’s context. Unit tests exercise components. Evals exercise end-to-end behavior on cases you thought of. Neither exercises the seam where an error message gets concatenated into a prompt, and that seam is exactly where an error becomes a narrative.
It also explains why an observability stack built on traces and status codes underperforms here. A span can finish with status OK, the model can return a well-formed response, and the content can still be fiction derived from a failure. The signal you need is not “did the call succeed” but “did a failure upstream get absorbed into this output.”
Where errors enter the context window #
The practical question is how errors reach the model in the first place. In most agent loops there are four routes. Tool output, where a failed call returns a stack trace or HTTP error body as text that gets appended to the transcript. Shell and subprocess streams, where stderr gets merged with stdout and handed back. Retrieval, where an error page or a truncated document is indexed or fetched as if it were content. And platform wrappers, where a gateway returns a generic 502 that has lost the upstream cause but still reads like a reasonable page of text.
In every case the failure arrives as natural language, and a language model’s job is to continue natural language coherently. Nothing in the architecture marks that string as an error rather than evidence. That is the design flaw, and it is yours, not the model’s. You routed a control signal through the one component whose defining behavior is making text sound sensible.
Decision Framework: model-visible errors versus out-of-band errors #
The decision that matters is whether a given failure is allowed to enter the model’s context at all. I use three buckets.
Failures the agent can legitimately act on, such as a rate limit it can retry or a missing field it can ask the user about, should reach the model, but as a structured object with a type, a code, and a retryable flag, not a raw string. The model reasons over the type; it never sees prose it can paraphrase.
Failures the agent cannot act on, such as a broken credential, an unreachable dependency, or a misconfigured path, should never reach the model. They go to a separate channel that pages a human or halts the run, and the user-facing output is a fixed, non-generated message that says the task did not complete.
Failures you do not yet understand are the dangerous ones. The conservative default is the second bucket: if you cannot classify it, treat it as unactionable and stop. The failure mode of the opposite default is the paper’s whole subject.
Common Failure Modes #
The first is stderr bleed. Merging error streams into the output the model reads is the cheapest way to create a fail-plausible incident, and the paper lists stderr discipline as a core hygiene rule.
The second is positional parsing of model output. If downstream code assumes the third line of a response is the answer, a polluted response shifts the lines and the wrong value flows on silently. Parse structured outputs against a schema and fail the run when validation fails.
The third is the green summary. A scheduled job that ends with the model writing “all tasks completed” is reporting the model’s belief, not the system’s state. Completion status should be computed from the job’s recorded outcomes, with the model allowed to narrate it but not to define it.
The fourth is the blind forensic tool. The paper documents a case where the operating system’s privacy sandbox silently denied the very diagnostics meant to investigate a failure, so “nothing found” and “access denied” looked identical. Your incident tooling needs to distinguish the two.
Architecture Impact #
What changes in system design? Errors become a typed, out-of-band channel instead of text in the transcript. The orchestration layer, not the model, decides whether a failure is shown to the model, escalated to a human, or terminal. Final run status is derived from machine-recorded outcomes of each step, and the model’s summary is a presentation layer on top of that record.
What new failure mode appears? Fail-plausible output: a response that is fluent, internally consistent, and wrong because it was built from an error or from polluted context. It does not trigger exceptions, does not change latency in a detectable way, and does not fail schema checks if the schema is loose. It surfaces only when a human notices the content is odd, which in the study meant delays from 13 hours to 60 days.
What enterprise teams should evaluate:
- Platform engineering: trace every route by which stderr, HTTP error bodies, and gateway failures can enter an agent’s context, and confirm each is typed or blocked before it reaches a prompt.
- SRE and on-call: check whether “run succeeded” is computed from step-level recorded outcomes or from the model’s own closing message, and whether a completed-looking run with a failed step pages anyone.
- Model risk and internal audit: ask whether your validation covers boundary conditions where upstream failures are injected into context, since component-level tests and standard evals mostly will not.
Cost / latency / governance / reliability implications: The main cost is engineering time for the error-typing layer and the injection test suite; there is little runtime cost, since classification is deterministic code and adds low single-digit milliseconds per step. The governance gain is larger: an auditable record of what actually happened, separate from what the agent said happened, which matters the moment an agent output feeds a customer communication or a regulatory filing. Reliability improves most on long-lived scheduled jobs, where an unnoticed failure compounds, since the paper’s worst case stayed silent for 60 days.
How to Evaluate #
Do not trust an eval that only measures answer quality on clean inputs. The test that matters is fault injection at the seams. For each tool, subprocess, and retrieval source, replace a real response with a realistic failure (a 502 with no body, a malformed UTF-8 string, an empty document, a permissions error) and check three things: did the run halt or degrade visibly, did a human-facing alert fire, and did any user-visible text claim a result the failed step could not have produced. Score the third one with a deterministic check against the recorded step outcomes, not with another model, because a model judge will often agree with the plausible narrative.
Implementation Guide #
Start with the status record. Before you change anything about prompts or models, make every agent step write a machine-readable outcome (step ID, status, error class, retryable flag) to a store the model cannot write to. Then compute run-level status from that store. This one change turns “the agent said it worked” into “the system recorded that it worked,” and it is the highest-leverage move because everything else, including alerting and audits, builds on it.
Next, build the error-typing layer at the tool boundary. Wrap tool and retrieval calls so they return either a result or a typed error object, and strip raw error text before anything is appended to context. Keep stderr separate from stdout at the process level. For gateway failures where the upstream cause is lost, treat the missing cause itself as the error class rather than letting a generic page of text through. The paper’s rule of thumb about alert stripping before the model sees context is a good default: if you would not want the model to paraphrase it, it should not be in the prompt.
What to avoid is accreting more guards without retiring complexity. The paper’s own recommendation is seam reduction over defense accretion: unify representations so there are fewer boundaries where an error can change form, and keep observers read-only so a monitoring component cannot itself corrupt the thing it watches. Teams under pressure tend to add another validation layer per incident, which adds another seam per incident. Be wary of the opposite temptation too, which is to ask the model to “report any errors you notice.” That puts the detection responsibility on the component that produced the failure.
To know whether it is working, use sabotage validation. Deliberately break each invariant in a test environment (inject the malformed byte, deny the permission, drop the upstream cause) and confirm the guard fires. A guard you have never seen fail is a guard you have not verified. The paper’s data also suggests a second signal: track how incidents get discovered. If most are still found by users or by an engineer who happened to look, your automated coverage is thin no matter what the test count says. A weekly short review of the agent’s output as the user sees it, which the paper found outperformed every automated system it ran, is cheap and worth keeping even after you automate.
On a 6–12 month horizon, teams that get this right move through three stages the paper describes: a point fix for each incident, a named rule that generalizes it, and finally a mechanized scanner that enforces the rule in CI so the class cannot recur. Only the third stage prevented recurrence in the study. Expect to end up with a small set of declarative invariants, each verified at more than one layer, a registry of scheduled jobs that is machine-reconciled against what is actually running, and an incident taxonomy of your own that your reviewers use to ask the right questions. In regulated settings, that taxonomy and the status record become part of your model risk evidence, because they show how the system behaves when its inputs are wrong, not only when they are right.
Sources #
-
Wei Wu, “When Errors Become Narratives: A Longitudinal Taxonomy of Silent Failures in a Production LLM Agent Runtime” , arXiv 2606.14589, June 2026
-
Full text (HTML) Enterprise AI Architecture
Want more enterprise AI architecture breakdowns? #
Subscribe to SuperML.