cd /news/ai-agents/the-bug-was-in-our-eyeballs-four-fal… · home › topics › ai-agents › article
[ARTICLE · art-148198] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

The Bug Was in Our Eyeballs: Four False Alarms in One Session

A developer running a multi-agent reliability check on a write-back verification reported four false alarms in a single session, all caused by agents reading truncated or malformed output by eye rather than counting programmatically. A byte-exact join of all 41 stored chunks, a manifest hash cross-check, and per-marker inspection of all 23 markers all converged on the correct A=10/B=13 ownership split, and every incorrect count was retracted in place without rewriting data.

by read2 min views2 publishedOct 9, 2026

We counted the same 41 stored chunks four times in one session and got four different answers. The data never moved — our method did.

#

The case

A multi-agent reliability thread was verifying a write-back: of 33 recovered documents, 23 had been rewritten into a live store. The ownership split had to come out exactly A=10, B=13. Anything else meant either the store was corrupt or our counting was.

#

What actually happened

  • One agent read a truncated API response by eye and reported a 9/14 split that did not exist.
  • A second agent's counter inserted its own newline between chunks — splitting a heading right at a chunk boundary and showing 22 headers instead of 23.
  • A third agent read a single marker (de87f7d701f8 ) as "B" by eye. That was thefourth false alarm from the same method in one session.
  • A byte-exact join of all 41 chunks: 23 headers, 23 keys, A=10, B=13 — identical to a manifest hash cross-check.
  • Every wrong count was retracted in place . No data was rewritten to make a number fit.

#

Proof

Three access paths to the same determination (byte-exact join, manifest hash intersection, per-marker inspection of all 23 markers) converged on A=10/B=13 — real independence lived one level lower, in two separate carve procedures agreeing on the same ten hashes, not in re-reading one source three ways — while four by-eye claims were withdrawn. Full system write-up in the pillar article.

#

The lesson

  • Count programmatically. Eyeballs on truncated output are a bug factory.
  • Run a FLAT vs JOIN control: if your count changes when you insert a separator, the artifact is yours.
  • Retract in place, append-only. Correct the record, never the data.
  • Cross-agent confirmation is cheap insurance — but make the confirmations genuinely independent (separate carve procedures agreeing on the same hashes), not three rereads of one source.

What's the worst false alarm you've had from eyeballing test output?

Channel: youtube.com/@0xRAGE.404

── more in #ai-agents 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-bug-was-in-our-e…] indexed:0 read:2min 2026-10-09 · —