cd /news/ai-tools/your-llm-returned-almost-valid-json-… Β· home β€Ί topics β€Ί ai-tools β€Ί article
[ARTICLE Β· art-147650] src=dev.to β†— pub= topic=ai-tools verified=true sentiment=↑ positive

Your LLM returned almost-valid JSON. Stop hand-patching it.

A developer built coerce-json, a zero-dependency library that takes almost-valid LLM JSON output and coerces it to fit an existing Zod schema while reporting every change it makes. The library strips markdown fences, extracts JSON from prose, casts strings to numbers and booleans, applies defaults, and handles enum casing, returning an ordered audit log of fixes. On a hand-built corpus of 37 representative LLM output mistakes, the developer reports validation rising from 13.5% under plain Zod to 86.5% with coerce and 100% with fuzzy matching enabled.

by read5 min views1 publishedOct 8, 2026

You asked the model for { id: number, active: boolean }. Here's what came

back:

{"id": "42", "active": "true", "role": "admin"}

Every field is almost right. id is a string. active is a string. There's

a role you never asked for. And the whole thing is wrapped in a markdown code

fence. So this:

const user = User.parse(JSON.parse(raw)); // πŸ’₯

throws twice β€” once on the fence, and again on the types. So you start writing

the patches by hand:

const unfenced = raw.replace(/```
{% endraw %}
json\n?|\n?
{% raw %}
```/g, "");
const obj = JSON.parse(unfenced);
obj.id = Number(obj.id);
obj.active = obj.active === "true"; // and "yes"? and "on"? and "1"?
// ...and now do it again for the next schema

This is the part of every LLM integration nobody writes a blog post about. Let's

fix it properly.

1. z.coerce. Zod can coerce primitives β€” z.coerce.number(),

z.coerce.boolean(). But you opt in field by field, it can't touch the fence or

the prose, and z.coerce.boolean() just calls Boolean(x) β€” so "false"

becomes true. Worse: when it does coerce, it does it silently. You can't

log what changed.

2. jsonrepair / json-repair. Great at fixing broken syntax β€” unclosed

braces, missing quotes, trailing commas. But they're schema-blind. Feed one

{"id":"42"} and it returns {"id":"42"}, because that's already valid JSON.

The string-vs-number mismatch is exactly the problem, and it can't see it.

3. instructor-js and re-prompting. Validate, and on failure send the error

back to the model to try again. It works β€” but you've turned a "42" β†’ 42

cast, which is a few microseconds of local work, into another network round

trip, more tokens, and a non-deterministic retry.

The common thread: each tool owns one slice. Cast primitives, or fix syntax,

or re-prompt. None of them read the schema you already have, coerce the value

to fit it, and tell you what they changed.

coerce-json is a

zero-dependency library that takes almost-valid model output and makes it fit

your schema β€” reporting every fix:

import { coerce } from "coerce-json";
import { z } from "zod";

const User = z.object({ id: z.number(), active: z.boolean(), role: z.string().default("user") });

const { value, ok, changes } = coerce('```

json\n{"id":"42","active":"true"}\n

```', User);
// value   β†’ { id: 42, active: true, role: "user" }
// ok      β†’ true
// changes β†’ strip-fence, string->number @ id, string->boolean @ active, fill-default @ role

Three things make that line pull its weight:

"42" becomes a number because id wants a number, the missing role gets its documented default.changes is an ordered, auditable log. A coercion is never a silent black box β€” you can log it, warn on it, or gate on it. How much does that actually buy you? I built a corpus of 37 representative LLM

mistakes β€” numbers as strings, "yes" booleans, fenced and prose-wrapped

objects, enum casing, missing defaults. Under plain Zod, 13.5% of them

validate. Run them through coerce with defaults, and 86.5% validate. Flip

on fuzzy: true and it's 100%. (It's a hand-built, illustrative corpus, not

a scientific sample β€” the harness and the full breakdown are in

BENCHMARKS.md,

so re-run it before quoting the numbers.)

// prose-wrapped output
coerce('Sure! Here it is: {"a":"1"}', z.object({ a: z.number() }));
// β†’ { a: 1 }   (extract-json, then string->number)

// booleans the way models spell them
coerce('{"active":"yes"}', z.object({ active: z.boolean() })); // β†’ { active: true }

// enum casing, always safe
coerce('{"status":"ACTIVE"}', z.object({ status: z.enum(["active","inactive"]) }));
// β†’ { status: "active" }   (case-insensitive, on by default)

Lossier guesses β€” enum near-misses ("activ" β†’ "active") and key re-casing

(first_name β†’ firstName) β€” are opt-in behind { fuzzy: true }, because

they can change meaning. And an ambiguous fuzzy match is refused, not guessed.

Not a Zod shop? Same API, with an optional Ajv hook for authoritative

validation:

import { coerce, coerceWithAjv } from "coerce-json/json-schema";

coerce('{"id":"5","active":"yes"}', {
  type: "object",
  properties: { id: { type: "integer" }, active: { type: "boolean" }, role: { type: "string", default: "user" } },
  required: ["id", "active"],
  additionalProperties: false,
});
// β†’ { id: 5, active: true, role: "user" }, ok: true

If a library silently rewrites your data, you can't trust it in a pipeline. So

coerce-json holds four invariants, checked by property tests:

__proto__ / constructor / prototype keys are dropped and logged. This is a companion to trickle-json,

my incremental partial-JSON parser for LLM streams. Together they're the

backbone of a streaming structured-output flow:

fetch β†’ SSE β†’ parse partial JSON (trickle-json) β†’ repair/coerce to schema (coerce-json) β†’ validate

trickle-json gives you the best value available on every streamed chunk

without throwing; coerce-json makes that value fit your schema and hands you

the receipts:

import { StreamingJsonParser } from "trickle-json";
import { coerce } from "coerce-json/zod";

const parser = new StreamingJsonParser();
parser.on("snapshot", renderPreview);
for await (const chunk of res.body) parser.write(chunk);

const { value, ok, changes } = coerce(parser.end(), Answer);
if (ok) save(value);
else console.warn("could not fully repair:", changes);

"But provider structured outputs already fix this!" β€” they help, a lot. But they

don't cover local and open models, older endpoints, streamed partials, or

anything wrapped in prose or a fence. That's exactly coerce-json's niche β€” and

when the output is already clean, it's a near no-op, so it's safe to leave in.

npm install coerce-json

zod and ajv are optional peers β€” bring them only if you use them. If it mangles some input it shouldn't, open an issue with the string and the

schema β€” the change log and the "never fabricate" guarantee are the whole point,

so I want to know. ⭐ appreciated if it saves you a pile of hand-written casts.

── more in #ai-tools 4 stories Β· sorted by recency
── more on @coerce-json 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/your-llm-returned-al…] indexed:0 read:5min 2026-10-08 Β· β€”