{"slug": "structured-output-from-an-llm-can-be-perfect-json-and-completely-wrong", "title": "Structured Output From an LLM Can Be Perfect JSON and Completely Wrong", "summary": "A first-person account from Mudassir Khan reports that a risk scoring pipeline ran for three weeks returning every assessment marked low because the team measured success only by whether the LLM output parsed as JSON, not whether it was semantically correct. The AI in Production 2026 Benchmark Report traced 43 percent of AI workflow failures to malformed model outputs, and prompting LLMs for JSON structure still fails at roughly 5–15% in production, while native structured output enforces a supplied JSON Schema at generation time rather than checking it afterward. The account distinguishes JSON mode, which only guarantees the bytes parse, from native structured output, which narrows token choices so the result matches the schema's constraints.", "body_md": "*Schema compliance is the easy half of the problem. The half that quietly costs money is semantic validity, and almost nothing in a normal stack is checking it.*\n\nI once lost three weeks to a pipeline that **never threw a single error.**\n\nIt was a risk scoring workflow. A model read a document, produced structured output, and a downstream service acted on the result. Every response was valid JSON. Every field name matched. Every type was right. The parser never complained, the retry counter stayed at zero, and the dashboard was a flat green line.\n\nEvery risk assessment came back marked low. Not most of them. Every one.\n\nThe large language model (LLM) output was structurally perfect, yet semantically worthless. It ran that way for weeks because we had built our entire idea of “did it work” around whether the parse step threw an exception.\n\nThe language models didn’t produce broken JSON. However, it never solved the actual problem that we were trying to address.\n\nWhen using LLMs to output quality JSON, there are three tiers.\n\nYou write “respond only with JSON” at the end of your prompt and parse whatever comes back.\n\nHowever, the LLM can hallucinate in this tier.\n\nEach of those creates bad data downstream in any pipeline that lacks monitoring across the whole path.\n\nPrompting LLMs for efficient JSON structure still fails at a rate of roughly 5–15% in production.\n\nThe provider guarantees the output will parse.\n\nThat is a real improvement from tier one, and a lot of production systems currently sit at this tier.\n\nYou hand the model a JSON Schema, and the schema is enforced at generation time rather than checked afterward.\n\nSchema compliance stops being probabilistic, and works 100 percent of the time for the guarantee it actually makes.\n\nThe mistake is assuming tier two and tier three are the same product with different marketing.\n\nJSON mode and native structured output are not the same thing\n\nJSON mode promises exactly one thing: the bytes you get back will parse. It says nothing about whether the object matches your schema.\n\nNative structured output narrows the token choices available at each step of the LLM so that the final output matches the constraints of the schema.\n\nThe AI in Production 2026 Benchmark Report traced 43 percent of AI workflow failures back to malformed model outputs.\n\nNot model reasoning. Not retrieval. Output handling.\n\nMost teams I talk to spend their debugging energy on prompts and their infrastructure energy on retrieval. Output handling rarely gets treated with the attention it deserves, for teams discard it as monotonous as plumbing. And just like plumbing, output handling can fail unexpectedly and without warning if not designed correctly.\n\nMy risk scoring pipeline is proof of this. Valid JSON, correct field names, correct types, wrong answers, running for weeks. There was no exception to catch, no error rate to alert on, no red line on any graph. The system was confidently and consistently incorrect inside a shape that every automated check we had was designed to approve.\n\nThat is what makes silent failure expensive. A crash costs you an afternoon. Silent bad data costs you whatever decisions got made on top of it, plus the archaeology to find out how far back the problem goes.\n\nIf you are building anything where model output feeds an action rather than a human, this is the failure mode worth instrumenting for, and it belongs in your [LLM observability in production](https://mudassirkhan.me/blog/llm-observability-production) setup rather than in your parser.\n\nLet’s revisit Tier 2: JSON mode.\n\nIf your schema wants an integer, the model could return one of the following to make sure the JSON object parses semantically.\n\nEvery one of those parses cleanly. Your JSON validity metric shows 100 percent success, yet doesn’t return the values that your application is expecting. The application either fails and throws a type error, or silently coerces the value to continue.\n\nLLM JSON output reliability is a misleading metric to track on its own.\n\nIn Tier 3, native structured output makes sure that the values the JSON outputs are the values your application is expecting. As of now, OpenAI, Anthropic, and Google Gemini all support native structured output.\n\nConstrained generation helps enforce native structured output on a categorical instead of a statistical level. It doesn’t solely rely on checking the shape of the JSON to validate model output. It also asks if this JSON output is a plausible answer to the question you asked.\n\nThis is real AI workflow output validation. Some of the separable checks can include\n\nNone of that is expressible in a JSON schema. Schema validation is free and automatic once you turn it on. Semantic validation requires you to think about your domain, write real assertions, and accept that some of them will be judgment calls.\n\nJSON structural correctness is about the container. Semantic validity is about the contents.\n\nWhile constrained decoding gives you the container for free, it can also give you the contents with the help of a validation library.\n\nA library my team uses is Pydantic LLM. Its validation goes beyond simple type coercion. Pydantic validators (or Zod refinements for TypeScript developers), let you attach domain rules to the same object the schema already guarantees the shape of:\n\n``` python\nfrom pydantic import BaseModel, field_validator, model_validatorclass RiskAssessment(BaseModel):   score: int # 0..100  severity: str # low | medium | high  evidence: list[str]@field_validator(\"score\") @classmethod   def score_in_range(cls, v: int) -> int: if not @model_validator(mode=\"after\") def severity_matches_score(self):  # a \"low\" severity paired with a high score is incoherent output  # even though both fields are individually schema-valid  if self.severity == \"low\" and self.score > 60:    raise ValueError(\"severity/score contradiction\")  if not self.evidence:    raise ValueError(\"assessment with no supporting evidence\")  return self\n```\n\nIt is twenty lines of ordinary domain logic sitting in the one place every model response has to pass through. Although simple to write, these rules convert a silent failure into a loud one to easily notice and address.\n\nA malformed tool call usually errors out safely. A well-formed tool call with wrong arguments executes.\n\nIf you are wiring models to tools, the argument validation matters more than the schema does. I recommend reading up on how the [LLM function calling guide](https://mudassirkhan.me/blog/llm-function-calling-guide) handles that boundary before you ship your application.\n\nWhile constrained decoding is effective for my pipeline, I should point out some caveats on when it's not always the answer.\n\nFirst, a very restrictive schema can push the model toward filling fields it should have left empty. If your schema marks eight fields as required and the document genuinely only supports three, constrained generation will not let the model decline. It will produce something. The workaround is to identify which fields are optional and give the model an explicit “insufficient information” path.\n\nSecond, constrained decoding is not a reasoning upgrade. It guarantees the answer fits the output schema and addresses the question. It says nothing about whether the answer is right.\n\nI still recommend using constrained decoding whenever the output feeds a machine. Any time a downstream service, database write, or tool invocation consumes the response, you want the structural guarantee and enforcement at generation.\n\nFor anything related to exploratory work, genuine prose, and early prototyping, I’d hold off on using constrained decoding to fit a rigid schema on the response. Locking a schema before you understand the shape of the answer is its own kind of mistake.\n\nFor JSON model output, track two metrics, not one.\n\nSchema compliance tells you the container is right, and with native structured output it should sit at 100 percent permanently. Once it does, that metric stops being informative and you can stop watching it.\n\nSemantic validity is the metric nobody has instrumented, and it is where the 43 percent of failures actually live. Write the distribution check. Write the contradiction check. Assert that the answer references something that existed in the input.\n\nThe container has been solved for a while now. The contents are still on you.\n\n**What is structured output in LLMs?**\n\nStructured output is any mechanism that makes a language model return data in a predefined machine-readable format, typically JSON conforming to a schema you supply, rather than free prose. It spans a range from prompt instructions at the weak end to constrained decoding against a JSON Schema at the strong end, where the format is enforced during generation.\n\n**Why does JSON mode still fail in production?**\n\nBecause JSON mode only guarantees that output parses, not that it matches your schema or that the values are correct. You can get syntactically perfect JSON with missing required fields, wrong types, extra invented keys, or values that are complete nonsense. Naive JSON prompting fails outright at roughly 5 to 15 percent, and JSON mode fixes that specific problem while leaving schema compliance and semantic validity entirely unaddressed.\n\n**What is the difference between JSON mode and native structured output?**\n\nJSON mode guarantees parseable JSON and nothing more. Native structured output takes your JSON Schema and enforces it at generation time through constrained decoding, so the response is schema-valid by construction, 100 percent of the time. JSON mode is a formatting hint. Native structured output is a hard constraint on the decoder.\n\nSyed Ali, Founder and CEO, Echonos\n\n*Syed Ali is cofounder of Echonos, an audio-aware AI music video pipeline for indie artists, managers, and small labels. He was previously COO at Tabler and a data science consultant at Deloitte, Booz Allen Hamilton, and Accenture. He writes about the economics of music release at the intersection of streaming, AI tooling, and indie artist strategy.*\n\n[Structured Output From an LLM Can Be Perfect JSON and Completely Wrong](https://pub.towardsai.net/structured-output-from-an-llm-can-be-perfect-json-and-completely-wrong-cbd630be2306) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/structured-output-from-an-llm-can-be-perfect-json-and-completely-wrong", "canonical_source": "https://pub.towardsai.net/structured-output-from-an-llm-can-be-perfect-json-and-completely-wrong-cbd630be2306?source=rss----98111c9905da---4", "published_at": "2026-09-30 13:31:02+00:00", "updated_at": "2026-09-30 13:46:44.868698+00:00", "lang": "en", "topics": ["large-language-models", "structured-data", "ai-safety", "mlops", "ai-tools"], "entities": ["Mudassir Khan", "AI in Production 2026 Benchmark Report"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/structured-output-from-an-llm-can-be-perfect-json-and-completely-wrong", "markdown": "https://wpnews.pro/news/structured-output-from-an-llm-can-be-perfect-json-and-completely-wrong.md", "text": "https://wpnews.pro/news/structured-output-from-an-llm-can-be-perfect-json-and-completely-wrong.txt", "jsonld": "https://wpnews.pro/news/structured-output-from-an-llm-can-be-perfect-json-and-completely-wrong.jsonld"}}