Structured Output From an LLM Can Be Perfect JSON and Completely Wrong A first-person account from Mudassir Khan reports that a risk scoring pipeline ran for three weeks returning every assessment marked low because the team measured success only by whether the LLM output parsed as JSON, not whether it was semantically correct. The AI in Production 2026 Benchmark Report traced 43 percent of AI workflow failures to malformed model outputs, and prompting LLMs for JSON structure still fails at roughly 5–15% in production, while native structured output enforces a supplied JSON Schema at generation time rather than checking it afterward. The account distinguishes JSON mode, which only guarantees the bytes parse, from native structured output, which narrows token choices so the result matches the schema's constraints. Schema compliance is the easy half of the problem. The half that quietly costs money is semantic validity, and almost nothing in a normal stack is checking it. I once lost three weeks to a pipeline that never threw a single error. It was a risk scoring workflow. A model read a document, produced structured output, and a downstream service acted on the result. Every response was valid JSON. Every field name matched. Every type was right. The parser never complained, the retry counter stayed at zero, and the dashboard was a flat green line. Every risk assessment came back marked low. Not most of them. Every one. The large language model LLM output was structurally perfect, yet semantically worthless. It ran that way for weeks because we had built our entire idea of “did it work” around whether the parse step threw an exception. The language models didn’t produce broken JSON. However, it never solved the actual problem that we were trying to address. When using LLMs to output quality JSON, there are three tiers. You write “respond only with JSON” at the end of your prompt and parse whatever comes back. However, the LLM can hallucinate in this tier. Each of those creates bad data downstream in any pipeline that lacks monitoring across the whole path. Prompting LLMs for efficient JSON structure still fails at a rate of roughly 5–15% in production. The provider guarantees the output will parse. That is a real improvement from tier one, and a lot of production systems currently sit at this tier. You hand the model a JSON Schema, and the schema is enforced at generation time rather than checked afterward. Schema compliance stops being probabilistic, and works 100 percent of the time for the guarantee it actually makes. The mistake is assuming tier two and tier three are the same product with different marketing. JSON mode and native structured output are not the same thing JSON mode promises exactly one thing: the bytes you get back will parse. It says nothing about whether the object matches your schema. Native structured output narrows the token choices available at each step of the LLM so that the final output matches the constraints of the schema. The AI in Production 2026 Benchmark Report traced 43 percent of AI workflow failures back to malformed model outputs. Not model reasoning. Not retrieval. Output handling. Most teams I talk to spend their debugging energy on prompts and their infrastructure energy on retrieval. Output handling rarely gets treated with the attention it deserves, for teams discard it as monotonous as plumbing. And just like plumbing, output handling can fail unexpectedly and without warning if not designed correctly. My risk scoring pipeline is proof of this. Valid JSON, correct field names, correct types, wrong answers, running for weeks. There was no exception to catch, no error rate to alert on, no red line on any graph. The system was confidently and consistently incorrect inside a shape that every automated check we had was designed to approve. That is what makes silent failure expensive. A crash costs you an afternoon. Silent bad data costs you whatever decisions got made on top of it, plus the archaeology to find out how far back the problem goes. If you are building anything where model output feeds an action rather than a human, this is the failure mode worth instrumenting for, and it belongs in your LLM observability in production https://mudassirkhan.me/blog/llm-observability-production setup rather than in your parser. Let’s revisit Tier 2: JSON mode. If your schema wants an integer, the model could return one of the following to make sure the JSON object parses semantically. Every one of those parses cleanly. Your JSON validity metric shows 100 percent success, yet doesn’t return the values that your application is expecting. The application either fails and throws a type error, or silently coerces the value to continue. LLM JSON output reliability is a misleading metric to track on its own. In Tier 3, native structured output makes sure that the values the JSON outputs are the values your application is expecting. As of now, OpenAI, Anthropic, and Google Gemini all support native structured output. Constrained generation helps enforce native structured output on a categorical instead of a statistical level. It doesn’t solely rely on checking the shape of the JSON to validate model output. It also asks if this JSON output is a plausible answer to the question you asked. This is real AI workflow output validation. Some of the separable checks can include None of that is expressible in a JSON schema. Schema validation is free and automatic once you turn it on. Semantic validation requires you to think about your domain, write real assertions, and accept that some of them will be judgment calls. JSON structural correctness is about the container. Semantic validity is about the contents. While constrained decoding gives you the container for free, it can also give you the contents with the help of a validation library. A library my team uses is Pydantic LLM. Its validation goes beyond simple type coercion. Pydantic validators or Zod refinements for TypeScript developers , let you attach domain rules to the same object the schema already guarantees the shape of: python from pydantic import BaseModel, field validator, model validatorclass RiskAssessment BaseModel : score: int 0..100 severity: str low | medium | high evidence: list str @field validator "score" @classmethod def score in range cls, v: int - int: if not @model validator mode="after" def severity matches score self : a "low" severity paired with a high score is incoherent output even though both fields are individually schema-valid if self.severity == "low" and self.score 60: raise ValueError "severity/score contradiction" if not self.evidence: raise ValueError "assessment with no supporting evidence" return self It is twenty lines of ordinary domain logic sitting in the one place every model response has to pass through. Although simple to write, these rules convert a silent failure into a loud one to easily notice and address. A malformed tool call usually errors out safely. A well-formed tool call with wrong arguments executes. If you are wiring models to tools, the argument validation matters more than the schema does. I recommend reading up on how the LLM function calling guide https://mudassirkhan.me/blog/llm-function-calling-guide handles that boundary before you ship your application. While constrained decoding is effective for my pipeline, I should point out some caveats on when it's not always the answer. First, a very restrictive schema can push the model toward filling fields it should have left empty. If your schema marks eight fields as required and the document genuinely only supports three, constrained generation will not let the model decline. It will produce something. The workaround is to identify which fields are optional and give the model an explicit “insufficient information” path. Second, constrained decoding is not a reasoning upgrade. It guarantees the answer fits the output schema and addresses the question. It says nothing about whether the answer is right. I still recommend using constrained decoding whenever the output feeds a machine. Any time a downstream service, database write, or tool invocation consumes the response, you want the structural guarantee and enforcement at generation. For anything related to exploratory work, genuine prose, and early prototyping, I’d hold off on using constrained decoding to fit a rigid schema on the response. Locking a schema before you understand the shape of the answer is its own kind of mistake. For JSON model output, track two metrics, not one. Schema compliance tells you the container is right, and with native structured output it should sit at 100 percent permanently. Once it does, that metric stops being informative and you can stop watching it. Semantic validity is the metric nobody has instrumented, and it is where the 43 percent of failures actually live. Write the distribution check. Write the contradiction check. Assert that the answer references something that existed in the input. The container has been solved for a while now. The contents are still on you. What is structured output in LLMs? Structured output is any mechanism that makes a language model return data in a predefined machine-readable format, typically JSON conforming to a schema you supply, rather than free prose. It spans a range from prompt instructions at the weak end to constrained decoding against a JSON Schema at the strong end, where the format is enforced during generation. Why does JSON mode still fail in production? Because JSON mode only guarantees that output parses, not that it matches your schema or that the values are correct. You can get syntactically perfect JSON with missing required fields, wrong types, extra invented keys, or values that are complete nonsense. Naive JSON prompting fails outright at roughly 5 to 15 percent, and JSON mode fixes that specific problem while leaving schema compliance and semantic validity entirely unaddressed. What is the difference between JSON mode and native structured output? JSON mode guarantees parseable JSON and nothing more. Native structured output takes your JSON Schema and enforces it at generation time through constrained decoding, so the response is schema-valid by construction, 100 percent of the time. JSON mode is a formatting hint. Native structured output is a hard constraint on the decoder. Syed Ali, Founder and CEO, Echonos Syed Ali is cofounder of Echonos, an audio-aware AI music video pipeline for indie artists, managers, and small labels. He was previously COO at Tabler and a data science consultant at Deloitte, Booz Allen Hamilton, and Accenture. He writes about the economics of music release at the intersection of streaming, AI tooling, and indie artist strategy. Structured Output From an LLM Can Be Perfect JSON and Completely Wrong https://pub.towardsai.net/structured-output-from-an-llm-can-be-perfect-json-and-completely-wrong-cbd630be2306 was originally published in Towards AI https://pub.towardsai.net on Medium, where people are continuing the conversation by highlighting and responding to this story.