cd /news/large-language-models/stop-asking-ai-to-pull-out-the-key-d… · home › topics › large-language-models › article
[ARTICLE · art-148838] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Stop Asking AI to 'Pull Out the Key Details': Use a Schema-First Extraction Contract

A developer argues that most LLM extraction failures are contract failures rather than intelligence failures, and proposes a schema-first prompt that splits the task into four labelled inputs: unstructured_text, extraction_targets, output_format_schema, and missing_data_policy. The approach forbids conversational filler and fabrication, requires explicit evidence for every field, and gives the model three honest moves for absent data — null values, YYYY-MM-DD date normalization, and a confidence: "unconfirmed" flag — to prevent invented names, dates, metrics and parameters from passing validation.

by read6 min views1 publishedOct 10, 2026

Paste a meeting transcript into ChatGPT, type "pull out the key details", and you get a friendly paragraph. It reads well. It is also useless to your code, because a script cannot parse "Sarah from Acme seems keen to move around 420 seats by mid-November."

So you add "return JSON". Now you get a code block, but it opens with "Sure! Here's the extracted data:", the dates are in three different formats, and a budget figure has appeared that nobody mentioned on the call. Most extraction failures are not intelligence failures. They are contract failures.

A request like "pull out the key details" is a summarization task wearing an extraction costume. Summaries are allowed to paraphrase, skip, and smooth over gaps. Extraction is not. It is a mapping from evidence in the source to fields in a schema, and every field either has support in the text or it does not.

A model asked to summarize will do what summaries do. It mixes facts with commentary, drops attributes that feel minor, and fills awkward gaps with something plausible. If a field is missing from the source, the most natural continuation is a value that looks right.

Three failure modes show up again and again:

JSON.parse(). The third one is the expensive one. A broken parser fails loudly. A fabricated value passes validation and lands in your database.

Author's Comment: I stopped judging extraction prompts by how the output looks and started judging them by what happens when a field is absent. If the prompt has no answer for that case, the model will invent one.

The fix is to stop describing what you want and start specifying what a valid answer is. In the Structured Data Extraction free prompt, that contract is split into four separate inputs, each doing one job:

unstructured_text: the raw source, kept apart from every instruction. extraction_targets: the exact attributes to find. output_format_schema: the structure the answer must take. missing_data_policy: what to do when evidence is absent or ambiguous. Separating these matters more than it looks. When the source text, the field list and the rules sit in one paragraph, the model has to guess which sentences are data and which are instructions. Four labelled slots remove that guess.

The prompt body then runs a short, fixed procedure: read the whole source, extract each target from explicit evidence only, normalize formats, apply the missing-data rule, and serialize to the schema. The constraints at the bottom are blunt on purpose.

- Zero Conversational Filler: Output ONLY the structured payload.
- Zero Fabrication: Never invent, round, or infer missing names,
  dates, metrics, or parameters not explicitly in the source.
- Schema & Type Integrity: Every record contains the exact keys
  defined, with consistent data types across all entries.

None of this is clever. That is the point. Each line closes one of the failure modes above.

If I could keep one part of the prompt, it would be missing_data_policy. It is the same idea I wrote about in why hard constraints beat better prompts: a rule needs an exit. Tell a model to "extract the PO owner" and give it no honest way to say the owner is unknown, and you have created pressure to fabricate.

The default policy in the prompt gives the model three honest moves. Unstated fields become null. Dates normalize to YYYY-MM-DD. Tentative items are flagged with confidence: "unconfirmed".

Here is what that does with a realistic input. Take this fragment from a customer call note:

Budget approved at $84,000 ARR. Next sync Oct 11 at 10:00 AM EST;
Marcus to deliver the hotfix before then. Q1 expansion could add
~150 seats, though timeline and PO owner are not yet confirmed.

A naive prompt returns something like "Q1 expansion: 150 seats, owner: Procurement". The schema-first version returns the number, marks it unconfirmed, and leaves the rest empty:

{
  "approved_arr_amount": 84000,
  "currency": "USD",
  "projected_q1_expansion_seats": 150,
  "q1_expansion_timeline": null,
  "q1_expansion_po_owner": null,
  "q1_expansion_confidence": "unconfirmed"
}

Two nulls and a flag. Less impressive to read, far safer to ingest. Your downstream code can now treat null as a real signal instead of discovering a hallucination in a quarterly report.

There is a fair objection here. Forcing a model into a rigid format is not free. A study titled Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models found a significant decline in reasoning ability when models were restricted to structured formats such as JSON, and stricter constraints generally produced greater degradation on reasoning tasks.

That finding should shape how you use this pattern. Extraction is mostly a lookup and normalization task, not a multi-step reasoning task, so the cost is small in practice. But if you need the model to analyze, compare, or decide, do not force that thinking through a JSON schema. Run the reasoning in free text first, then extract from the result in a second pass.

The same reasoning explains why the output format is a variable rather than a hard-coded choice. The prompt ships with options such as a strict JSON object, a flat JSON array, a Markdown table, a dual output (a human-readable table followed by a JSON block), and RFC 4180 CSV. If you choose JSON, "valid RFC 8259" is a testable target; the RFC 8259 specification defines exactly what a parser will accept, which is a better standard than "looks like JSON".

Practical Pitfall Avoidance Guide: Always validate the output with a real parser before it touches a database. A prompt contract reduces bad output; it does not eliminate it. Treat the model as an untrusted data source and keep a schema validator in the pipeline.

The prompt's Pro Tip is the most underrated part. Add a source_quote (or evidence_span) field next to each extracted value, and require the model to copy the exact 5 to 10 word substring it relied on.

This does two things. It forces the model to find actual evidence before emitting a normalized value, which cuts extraction drift. And it gives you a cheap audit: a script can check that every source_quote really appears in the original text. A quote that does not match is a flagged fabrication, no human review needed.

The prompt offers this as the "Audit Trail Mode" option for missing_data_policy, so you can switch it on without rewriting anything.

Good fits are the ones where the source is messy and the target is strict:

Poor fits are worth naming too. Do not use this for creative rewriting or open-ended editorial work, where prose flow matters more than field boundaries. And if your source is already clean CSV, SQL or valid JSON, use a deterministic tool like jq or pandas. A parser is faster, cheaper, and never hallucinates.

Pick a transcript or log you already have. Write down the five fields you actually need, decide what should happen when each one is missing, and choose the output format before you touch the model.

Then load the prompt template, which includes preset options for every variable and a full sample output, and save it to Prompt Vault so your schema and null policy stay the same from run to run. Check the first result against the source line by line. The fields that come back as null will tell you more about your data than the ones that come back filled in.

── more in #large-language-models 4 stories · sorted by recency
── more on @chatgpt 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/stop-asking-ai-to-pu…] indexed:0 read:6min 2026-10-10 · —