cd /news/large-language-models/your-llm-keeps-making-the-same-extra… Β· home β€Ί topics β€Ί large-language-models β€Ί article
[ARTICLE Β· art-144231] src=dev.to β†— pub= topic=large-language-models verified=true sentiment=↑ positive

Your LLM Keeps Making the Same Extraction Mistake. Here's How to Make It Learn.

An AWS engineer open-sourced a sample pipeline, sample-prompt-correction-memory, that captures human QA corrections to LLM document extraction and reuses them so recurring field errors are not repeated. The system routes fields through three tiers β€” deterministic rules synthesized from repeated corrections, Amazon Bedrock Claude Haiku 4.5 with confidence scoring, and Claude Sonnet 4.5 few-shot re-extraction using retrieved past corrections β€” shifting traffic toward the zero-cost rule tier over time without retraining or redeployment.

by read4 min views2 publishedOct 3, 2026

> TL;DR β€” I open-sourced an AWS sample that makes LLM document extraction learn from human corrections. Each QA fix is captured once and reused, so the pipeline gets more accurate and cheaper over time β€” with no retraining and no redeployment. Repo: github.com/aws-samples/sample-prompt-correction-memory. You can run the whole self-healing loop locally in 30 seconds, no AWS account needed.

If you've built an LLM extraction pipeline, you know this pain.

Your QA analyst opens today's invoice. The model extracted payment_terms as "quarterly." That's wrong β€” "quarterly" is the billing frequency; the actual payment term is Net 30. The analyst corrects it. Done.

Tomorrow, an identical invoice from the same vendor arrives. The model extracts... "quarterly" again. Same mistake. The analyst corrects it again.

The correction evaporated. It fixed one document and taught the system nothing.

Multiply that across thousands of documents and dozens of recurring error patterns, and you get the two options most teams settle for:

There's a third way that needs neither.

What if every human correction became a permanent, reusable asset?

That's the pattern behind sample-prompt-correction-memory. Instead of throwing corrections away, it stores them and feeds them back into future extractions β€” so the same mistake is never made twice.

It works as three tiers, and the magic is that traffic gradually shifts from the expensive tier to the free one:

When the same correction pattern recurs enough times, the system synthesizes a deterministic rule (a regex, a lookup, a normalization) that handles that field with zero LLM cost and sub-millisecond latency. "Net 30" β†’ 30 becomes a rule; it never needs a model again.

Fields with no matching rule go to Amazon Bedrock (Claude Haiku 4.5). Every field comes back with a confidence score. If confidence clears the threshold, accept it and move on.

If confidence is below threshold, the system retrieves the most relevant past corrections, formats them as few-shot examples, and re-extracts with a stronger model (** Claude Sonnet 4.5**). The model learns from the specific mistakes it made before β€” in-context, no training.

Every QA correction flows into a correction log (DynamoDB) that powers both the self-healing retrieval and the rule graduation. So the more the system is used, the smarter and cheaper it gets.

The clearest fit is high-repetition, back-office document processing β€” accounts payable, contracts, compliance filings, supply-chain docs β€” where the same field errors recur and analyst time is the real expense.

Fully serverless, deployed with AWS SAM:

Document ──► S3 ──► EventBridge ──► SQS ──► Lambda (extract)
                                              β”‚
                          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                          β–Ό                    β–Ό                   β–Ό
                   Tier 1: Rules       Tier 2: Bedrock      Tier 3: Self-heal
                   (deterministic)     (Claude Haiku 4.5)   (Claude Sonnet 4.5)
                          β–²                                        β”‚
                          β”‚                                        β–Ό
                          └──────── Correction Log (DynamoDB) β—„β”€β”€β”€β”€β”˜
                                    (grows from QA feedback)

QA corrections are uploaded as JSON to an S3 corrections/ prefix; a second Lambda validates and ingests them into the correction log. Everything is encrypted with a customer-managed KMS key, and IAM is scoped to the specific Bedrock model ARNs the sample uses.

The quickstart runs the entire self-healing loop with a mocked Bedrock client, so you can watch the mechanism without deploying anything:

git clone https://github.com/aws-samples/sample-prompt-correction-memory
cd sample-prompt-correction-memory
pip install -e ".[dev]"
python examples/quickstart.py

You'll see output like this:

Step 1: Initial extraction (no correction memory)
  Field:      effective_date
  Value:      March 2024        Confidence: 0.55   ❌ NO

Step 2: Self-healing triggers (0.55 < threshold 0.70)
  Retrieving corrections for 'effective_date'...
  Found: "March 2024" β†’ "2024-03-01"

Step 3: Re-extraction with correction memory
  Value:       2024-03-01       Confidence: 0.95   βœ“ YES
  Self-Healed: True

Low-confidence extraction β†’ retrieve past correction β†’ re-extract β†’ correct answer. No retraining. No redeployment.

With the AWS CLI, SAM CLI, and Bedrock model access for Claude Haiku 4.5 + Sonnet 4.5:

make deploy    # S3, DynamoDB, Lambda, EventBridge, SQS, KMS
make seed      # load sample corrections into the correction log
make trigger   # upload a sample document β†’ triggers real extraction
make verify    # read back results (fields, confidence, self_healed)
make destroy   # empty buckets + delete the stack

make verify prints each extracted field with its confidence and whether self-healing kicked in β€” your proof it works end-to-end.

This sample provides semantic correction memory (what was extracted wrong, and why). Its companion, sample-textract-field-memory, provides spatial memory (where fields appear on a document layout). Together they form a dual memory for document pipelines: use the cheap spatial lookup when you're confident where a field is, and fall back to self-healing extraction when you're not.

Because it's more useful when it's used well:

The repo is open source under aws-samples, MIT-0 licensed, with a benchmark suite, an interactive dashboard, and an offline test harness.

πŸ‘‰ github.com/aws-samples/sample-prompt-correction-memory

What repetitive extraction error is your team still fixing by hand? I'd love to hear which patterns you'd want a system like this to learn.

This is an open-source AWS sample intended for demonstration and non-production use.

── more in #large-language-models 4 stories Β· sorted by recency
── more on @aws 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/your-llm-keeps-makin…] indexed:0 read:4min 2026-10-03 Β· β€”