{"slug": "your-llm-returned-almost-valid-json-stop-hand-patching-it", "title": "Your LLM returned almost-valid JSON. Stop hand-patching it.", "summary": "A developer built coerce-json, a zero-dependency library that takes almost-valid LLM JSON output and coerces it to fit an existing Zod schema while reporting every change it makes. The library strips markdown fences, extracts JSON from prose, casts strings to numbers and booleans, applies defaults, and handles enum casing, returning an ordered audit log of fixes. On a hand-built corpus of 37 representative LLM output mistakes, the developer reports validation rising from 13.5% under plain Zod to 86.5% with coerce and 100% with fuzzy matching enabled.", "body_md": "You asked the model for `{ id: number, active: boolean }`. Here's what came\n\nback:\n\n```\n{\"id\": \"42\", \"active\": \"true\", \"role\": \"admin\"}\n```\n\nEvery field is *almost* right. `id` is a string. `active` is a string. There's\n\na `role` you never asked for. And the whole thing is wrapped in a markdown code\n\nfence. So this:\n\n``` js\nconst user = User.parse(JSON.parse(raw)); // 💥\n```\n\nthrows twice — once on the fence, and again on the types. So you start writing\n\nthe patches by hand:\n\n``` js\nconst unfenced = raw.replace(/```\n{% endraw %}\njson\\n?|\\n?\n{% raw %}\n```/g, \"\");\nconst obj = JSON.parse(unfenced);\nobj.id = Number(obj.id);\nobj.active = obj.active === \"true\"; // and \"yes\"? and \"on\"? and \"1\"?\n// ...and now do it again for the next schema\n```\n\nThis is the part of every LLM integration nobody writes a blog post about. Let's\n\nfix it properly.\n\n**1. `z.coerce`.** Zod can coerce primitives — `z.coerce.number()`,\n\n`z.coerce.boolean()`. But you opt in field by field, it can't touch the fence or\n\nthe prose, and `z.coerce.boolean()` just calls `Boolean(x)` — so `\"false\"`\n\nbecomes `true`. Worse: when it does coerce, it does it **silently**. You can't\n\nlog what changed.\n\n**2. `jsonrepair` / `json-repair`.** Great at fixing broken *syntax* — unclosed\n\nbraces, missing quotes, trailing commas. But they're **schema-blind**. Feed one\n\n`{\"id\":\"42\"}` and it returns `{\"id\":\"42\"}`, because that's already valid JSON.\n\nThe string-vs-number mismatch is exactly the problem, and it can't see it.\n\n**3. `instructor-js` and re-prompting.** Validate, and on failure send the error\n\nback to the model to try again. It works — but you've turned a `\"42\"` → `42`\n\ncast, which is a few microseconds of local work, into another network round\n\ntrip, more tokens, and a non-deterministic retry.\n\nThe common thread: each tool owns one slice. Cast primitives, *or* fix syntax,\n\n*or* re-prompt. None of them read the schema you already have, coerce the value\n\nto fit it, and **tell you what they changed**.\n\n[`coerce-json`](https://www.npmjs.com/package/coerce-json) is a\n\nzero-dependency library that takes almost-valid model output and makes it fit\n\nyour schema — reporting every fix:\n\n``` js\nimport { coerce } from \"coerce-json\";\nimport { z } from \"zod\";\n\nconst User = z.object({ id: z.number(), active: z.boolean(), role: z.string().default(\"user\") });\n\nconst { value, ok, changes } = coerce('```\n\njson\\n{\"id\":\"42\",\"active\":\"true\"}\\n\n\n```', User);\n// value   → { id: 42, active: true, role: \"user\" }\n// ok      → true\n// changes → strip-fence, string->number @ id, string->boolean @ active, fill-default @ role\n```\n\nThree things make that line pull its weight:\n\n`\"42\"` becomes a\nnumber because `id` wants a number, the missing `role` gets its documented\ndefault.`changes` is an ordered, auditable log. A coercion\nis never a silent black box — you can log it, warn on it, or gate on it.\nHow much does that actually buy you? I built a corpus of 37 representative LLM\n\nmistakes — numbers as strings, `\"yes\"` booleans, fenced and prose-wrapped\n\nobjects, enum casing, missing defaults. Under plain Zod, **13.5%** of them\n\nvalidate. Run them through `coerce` with defaults, and **86.5%** validate. Flip\n\non `fuzzy: true` and it's **100%**. (It's a hand-built, illustrative corpus, not\n\na scientific sample — the harness and the full breakdown are in\n\n[BENCHMARKS.md](https://github.com/H1manshu01/coerce-json/blob/main/BENCHMARKS.md),\n\nso re-run it before quoting the numbers.)\n\n```\n// prose-wrapped output\ncoerce('Sure! Here it is: {\"a\":\"1\"}', z.object({ a: z.number() }));\n// → { a: 1 }   (extract-json, then string->number)\n\n// booleans the way models spell them\ncoerce('{\"active\":\"yes\"}', z.object({ active: z.boolean() })); // → { active: true }\n\n// enum casing, always safe\ncoerce('{\"status\":\"ACTIVE\"}', z.object({ status: z.enum([\"active\",\"inactive\"]) }));\n// → { status: \"active\" }   (case-insensitive, on by default)\n```\n\nLossier guesses — enum *near-misses* (`\"activ\"` → `\"active\"`) and key re-casing\n\n(`first_name` → `firstName`) — are **opt-in** behind `{ fuzzy: true }`, because\n\nthey can change meaning. And an ambiguous fuzzy match is refused, not guessed.\n\nNot a Zod shop? Same API, with an optional Ajv hook for authoritative\n\nvalidation:\n\n``` js\nimport { coerce, coerceWithAjv } from \"coerce-json/json-schema\";\n\ncoerce('{\"id\":\"5\",\"active\":\"yes\"}', {\n  type: \"object\",\n  properties: { id: { type: \"integer\" }, active: { type: \"boolean\" }, role: { type: \"string\", default: \"user\" } },\n  required: [\"id\", \"active\"],\n  additionalProperties: false,\n});\n// → { id: 5, active: true, role: \"user\" }, ok: true\n```\n\nIf a library silently rewrites your data, you can't trust it in a pipeline. So\n\n`coerce-json` holds four invariants, checked by property tests:\n\n`__proto__` / `constructor` / `prototype` keys\nare dropped and logged.\nThis is a companion to [`trickle-json`](https://www.npmjs.com/package/trickle-json),\n\nmy incremental partial-JSON parser for LLM streams. Together they're the\n\nbackbone of a streaming structured-output flow:\n\n```\nfetch → SSE → parse partial JSON (trickle-json) → repair/coerce to schema (coerce-json) → validate\n```\n\n`trickle-json` gives you the best value available on every streamed chunk\n\nwithout throwing; `coerce-json` makes that value fit your schema and hands you\n\nthe receipts:\n\n``` js\nimport { StreamingJsonParser } from \"trickle-json\";\nimport { coerce } from \"coerce-json/zod\";\n\nconst parser = new StreamingJsonParser();\nparser.on(\"snapshot\", renderPreview);\nfor await (const chunk of res.body) parser.write(chunk);\n\nconst { value, ok, changes } = coerce(parser.end(), Answer);\nif (ok) save(value);\nelse console.warn(\"could not fully repair:\", changes);\n```\n\n\"But provider structured outputs already fix this!\" — they help, a lot. But they\n\ndon't cover local and open models, older endpoints, streamed partials, or\n\nanything wrapped in prose or a fence. That's exactly `coerce-json`'s niche — and\n\nwhen the output *is* already clean, it's a near no-op, so it's safe to leave in.\n\n```\nnpm install coerce-json\n```\n\n`zod` and `ajv` are optional peers — bring them only if you use them.\nIf it mangles some input it shouldn't, open an issue with the string and the\n\nschema — the change log and the \"never fabricate\" guarantee are the whole point,\n\nso I want to know. ⭐ appreciated if it saves you a pile of hand-written casts.", "url": "https://wpnews.pro/news/your-llm-returned-almost-valid-json-stop-hand-patching-it", "canonical_source": "https://dev.to/h1manshu01/your-llm-returned-almost-valid-json-stop-hand-patching-it-2fma", "published_at": "2026-10-08 15:09:22+00:00", "updated_at": "2026-10-08 15:20:36.497968+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "developer-tools", "structured-data"], "entities": ["coerce-json", "Zod", "jsonrepair", "instructor-js", "H1manshu01"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/your-llm-returned-almost-valid-json-stop-hand-patching-it", "markdown": "https://wpnews.pro/news/your-llm-returned-almost-valid-json-stop-hand-patching-it.md", "text": "https://wpnews.pro/news/your-llm-returned-almost-valid-json-stop-hand-patching-it.txt", "jsonld": "https://wpnews.pro/news/your-llm-returned-almost-valid-json-stop-hand-patching-it.jsonld"}}