cd /news/artificial-intelligence/our-ai-reviewer-invented-a-request-o… · home topics artificial-intelligence article
[ARTICLE · art-108138] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Our AI reviewer invented a request. Our producer retried 245 times.

A developer's team running ~100 unattended LLM agents on local models discovered a single document that was rewritten 245 times in 5 days, with a sibling document rewritten 225 times, totaling about 470 wasted generations. The root cause was a reviewer agent hallucinating a request that never existed, combined with a retry mechanism that counted reviews instead of contract failures, allowing the loop to run unbounded. The developer's audit of 2,038 reviews found a 0.2% hallucination rate, but the real risk was the infinite retry loop, not the rate itself.

read2 min views1 publishedAug 24, 2026

We run ~100 LLM agents unattended on local models. Last week we found one

document that had been rewritten 245 times in 5 days — every attempt

rejected. A sibling document: 225 times. Combined, about 470 wasted

generations, all burned on the same two files.

Here is the autopsy, with the actual numbers.

Our pipeline is simple: a producer agent writes a document, a reviewer agent

checks it against a contract (minimum length, required sections, no

placeholder junk), and rejected work goes back with fix instructions.

The rejected document was a key-management (KMS) implementation spec —

4,452 characters, perfectly on-topic. The reviewer's verdict:

"The request was a 3-line email triage response (LOCK / VERDICT / REASON),

but the answer is a long KMS spec. Rewrite as3 lines only."

One problem. We grepped the document: the words "LOCK", "VERDICT", and the

name of the triage service appear zero times in it. The reviewer had

invented the request.

Two contracts collided:

No output can satisfy both. So the producer failed the contract, got

re-queued, produced again, failed again — 245 times. Our retry cap counted

reviews, but a contract-failed output never reaches review. The give-up

mechanism existed; it just watched the wrong counter.

Our review prompt contained the artifact body (first 4,000 chars) and the

output format. It never contained the original request. We asked a model

"does this match the request?" without telling it what the request was.

A model asked to judge against information it doesn't have will

hallucinate that information. Ours did, confidently, 245 times' worth.

Bonus failure: we truncated long documents to 4,000 characters before

review without saying so, and reviewers marked them "thin — cut off

mid-sentence." The cut was ours, not the producer's.

We audited all 2,038 reviews on file for concrete terms (product names,

format tokens) that appear in the review but nowhere in the reviewed document. Result:

That's the uncomfortable lesson: a 0.2% hallucination rate produced 470

wasted runs, because nothing ever gave up. Low rate × infinite retries =

unbounded damage. The rate is not the risk; the loop is.

Each fix ships with a test we deliberately broke to confirm it fails.

The checker that catches broken outputs in this story (empty text, language

leakage, placeholder junk, contract violations) is free on npm:

honto-contract — it passed 600 downloads last week, so somebody besides us finds this useful now.

The unattended-operation checklist and three of our watchdog templates are

free (email-gated):

[Unattended-Operation Kit](https://gxcafe.co.jp/harness-kit/?utm_source=devto&utm_medium=article&utm_campaign=harness-kit)

The full set of 7 production templates (cron registry, silent-zero watch,

heartbeat, output contracts — the exact ones in this story) is

US$59. Honest note: we have no customers yet. Everything above is exactly what we

run on ourselves, measured on our own failures.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @honto-contract 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/our-ai-reviewer-inve…] indexed:0 read:2min 2026-08-24 ·