You asked the model for { id: number, active: boolean }. Here's what came
back:
{"id": "42", "active": "true", "role": "admin"}
Every field is almost right. id is a string. active is a string. There's
a role you never asked for. And the whole thing is wrapped in a markdown code
fence. So this:
const user = User.parse(JSON.parse(raw)); // π₯
throws twice β once on the fence, and again on the types. So you start writing
the patches by hand:
const unfenced = raw.replace(/```
{% endraw %}
json\n?|\n?
{% raw %}
```/g, "");
const obj = JSON.parse(unfenced);
obj.id = Number(obj.id);
obj.active = obj.active === "true"; // and "yes"? and "on"? and "1"?
// ...and now do it again for the next schema
This is the part of every LLM integration nobody writes a blog post about. Let's
fix it properly.
1. z.coerce. Zod can coerce primitives β z.coerce.number(),
z.coerce.boolean(). But you opt in field by field, it can't touch the fence or
the prose, and z.coerce.boolean() just calls Boolean(x) β so "false"
becomes true. Worse: when it does coerce, it does it silently. You can't
log what changed.
2. jsonrepair / json-repair. Great at fixing broken syntax β unclosed
braces, missing quotes, trailing commas. But they're schema-blind. Feed one
{"id":"42"} and it returns {"id":"42"}, because that's already valid JSON.
The string-vs-number mismatch is exactly the problem, and it can't see it.
3. instructor-js and re-prompting. Validate, and on failure send the error
back to the model to try again. It works β but you've turned a "42" β 42
cast, which is a few microseconds of local work, into another network round
trip, more tokens, and a non-deterministic retry.
The common thread: each tool owns one slice. Cast primitives, or fix syntax,
or re-prompt. None of them read the schema you already have, coerce the value
to fit it, and tell you what they changed.
coerce-json is a
zero-dependency library that takes almost-valid model output and makes it fit
your schema β reporting every fix:
import { coerce } from "coerce-json";
import { z } from "zod";
const User = z.object({ id: z.number(), active: z.boolean(), role: z.string().default("user") });
const { value, ok, changes } = coerce('```
json\n{"id":"42","active":"true"}\n
```', User);
// value β { id: 42, active: true, role: "user" }
// ok β true
// changes β strip-fence, string->number @ id, string->boolean @ active, fill-default @ role
Three things make that line pull its weight:
"42" becomes a
number because id wants a number, the missing role gets its documented
default.changes is an ordered, auditable log. A coercion
is never a silent black box β you can log it, warn on it, or gate on it.
How much does that actually buy you? I built a corpus of 37 representative LLM
mistakes β numbers as strings, "yes" booleans, fenced and prose-wrapped
objects, enum casing, missing defaults. Under plain Zod, 13.5% of them
validate. Run them through coerce with defaults, and 86.5% validate. Flip
on fuzzy: true and it's 100%. (It's a hand-built, illustrative corpus, not
a scientific sample β the harness and the full breakdown are in
so re-run it before quoting the numbers.)
// prose-wrapped output
coerce('Sure! Here it is: {"a":"1"}', z.object({ a: z.number() }));
// β { a: 1 } (extract-json, then string->number)
// booleans the way models spell them
coerce('{"active":"yes"}', z.object({ active: z.boolean() })); // β { active: true }
// enum casing, always safe
coerce('{"status":"ACTIVE"}', z.object({ status: z.enum(["active","inactive"]) }));
// β { status: "active" } (case-insensitive, on by default)
Lossier guesses β enum near-misses ("activ" β "active") and key re-casing
(first_name β firstName) β are opt-in behind { fuzzy: true }, because
they can change meaning. And an ambiguous fuzzy match is refused, not guessed.
Not a Zod shop? Same API, with an optional Ajv hook for authoritative
validation:
import { coerce, coerceWithAjv } from "coerce-json/json-schema";
coerce('{"id":"5","active":"yes"}', {
type: "object",
properties: { id: { type: "integer" }, active: { type: "boolean" }, role: { type: "string", default: "user" } },
required: ["id", "active"],
additionalProperties: false,
});
// β { id: 5, active: true, role: "user" }, ok: true
If a library silently rewrites your data, you can't trust it in a pipeline. So
coerce-json holds four invariants, checked by property tests:
__proto__ / constructor / prototype keys
are dropped and logged.
This is a companion to trickle-json,
my incremental partial-JSON parser for LLM streams. Together they're the
backbone of a streaming structured-output flow:
fetch β SSE β parse partial JSON (trickle-json) β repair/coerce to schema (coerce-json) β validate
trickle-json gives you the best value available on every streamed chunk
without throwing; coerce-json makes that value fit your schema and hands you
the receipts:
import { StreamingJsonParser } from "trickle-json";
import { coerce } from "coerce-json/zod";
const parser = new StreamingJsonParser();
parser.on("snapshot", renderPreview);
for await (const chunk of res.body) parser.write(chunk);
const { value, ok, changes } = coerce(parser.end(), Answer);
if (ok) save(value);
else console.warn("could not fully repair:", changes);
"But provider structured outputs already fix this!" β they help, a lot. But they
don't cover local and open models, older endpoints, streamed partials, or
anything wrapped in prose or a fence. That's exactly coerce-json's niche β and
when the output is already clean, it's a near no-op, so it's safe to leave in.
npm install coerce-json
zod and ajv are optional peers β bring them only if you use them.
If it mangles some input it shouldn't, open an issue with the string and the
schema β the change log and the "never fabricate" guarantee are the whole point,
so I want to know. β appreciated if it saves you a pile of hand-written casts.