cd /news/ai-agents/claude-code-rounded-money-wrong-in-5… · home › topics › ai-agents › article
[ARTICLE · art-149203] src=ainexusdaily.vercel.app ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Claude Code Rounded Money Wrong in 5 of 10 Runs. An 11-Line CLAUDE.md Fixed It.

Adding an 11-line CLAUDE.md file to a billing repository raised Claude Code's rounding accuracy from 5 of 10 runs to 10 of 10, according to a test by the article's author using Claude Code 2.1.296 with Sonnet 5.5 and Haiku 5.5. Without the file, Sonnet 5.5 returned 502 instead of 503 on a 1,005-cent invoice in 2 of 5 runs and Haiku 5.5 did so in 3 of 5, even though both runs reported using the round_money helper. The author argues the failures stem from implicit repository conventions rather than model capability, responding to Michael Lynch's October 9 post 'Why Are Coding Agents So Dumb?'

read8 min views2 publishedOct 11, 2026
Claude Code Rounded Money Wrong in 5 of 10 Runs. An 11-Line CLAUDE.md Fixed It.
Image: Ainexusdaily (auto-discovered)

Refund 50% of a 1,005-cent invoice. The billing code I gave the agent rounds half up, so the answer is 503. Running in Claude Code, Sonnet 5.5 returned 502 in 2 of 5 runs. Haiku 5.5 returned 502 in 3 of 5. Every one of those runs ended with a tidy summary saying it had rounded with round_money. Two

Refund 50% of a 1,005-cent invoice. The billing code I gave the agent rounds half up, so the answer is 503. Running in Claude Code, Sonnet 5.5 returned 502 in 2 of 5 runs. Haiku 5.5 returned 502 in 3 of 5. Every one of those runs ended with a tidy summary saying it had rounded with round_money. Two of the Haiku runs even noted that the file's own helper rounds half up, then picked the other one anyway. Then I added an 11-line CLAUDE.md to the repo and changed nothing else. Both models got the rounding right 10 times out of 10. That result is most of my answer to the "why are coding agents so dumb" argument going around this week. Michael Lynch's Why Are Coding Agents So Dumb? went up on October 9 and picked up long threads on Hacker News (100 points, 99 comments when I last checked) and Lobsters. Most of his complaints are about the harness, the software wrapped around the model. OpenCode split a 1.5k-line feature into 10 subtasks and then ran them one by one. Agents keep using the slow, expensive model for searches a cheaper one could do. One agent stopped two minutes after he walked away to ask what to name a git branch, and the work was still unstarted the next morning. File access "controls" amounted to politely asking the LLM not to read certain files, which it ignored. On all of that he's right, and nothing in your repo will fix it. Agents that are only usable with access to everything, no routing between cheap and expensive models, every session starting from zero (benjajaja on Lobsters called each one "a baby freshly spawned"): those are vendor problems. Keep yelling at the vendors. The comment threads went somewhere else, though. There, "dumb" meant the model "agrees and then proceeds to write some untyped string map ball of mud" (throwaway63467 on HN), or produced code that works but is "diabolically over complicated" (pipes, also on HN), or that it's "still literally just an autocomplete with longer context" (Pomax on Lobsters). That's where I disagree. A lot of those failures are the repo talking. I built a toy billing service and gave Claude Code 2.1.296 the same prompt every run: claude -p 'Add a function refund_invoice(invoice_id, percent) to the billing code. ... Follow the existing conventions of this codebase. Do not ask me questions; make reasonable decisions and finish the task.' \ --model claude-haiku-5-5 --output-format json --dangerously-skip-permissions Each run got a fresh copy in /tmp. Afterwards a hidden grader checked six things: integer cents with half-up rounding, an audit-log entry, the right status constants, and BillingError on unknown IDs, unpaid invoices and over-refunds. The messy version was one 1,094-line billing.py with no tests and no docs. The conventions were all implicit: round with a helper called _rc(), write through persist() (which audits) and never put() (which doesn't), use the STATUS* constants. I also left traps: a legacy_credit_note() that does float math and invents status strings like "refunded_partial", and an apply_discount() that truncates with int(). The clean version had the same code split into modules, the legacy function moved into a legacy/ folder with a "do not copy" header, four unit tests, and a CLAUDE.md. I expected the messy repo to lose. Haiku passed it 5 out of 5. Every run found _rc(), used persist() and ignored the float code. A big file with ugly names doesn't trip up a current model. It greps, reads create_invoice and pay_invoice, and copies them. The one real difference: in the clean repo Haiku wrote unit tests in 5 of 5 runs. CLAUDE.md said new functions get a test, and there was a tests/ folder to copy from. In the messy repo it wrote zero. The prompt never asked for tests. So I added the thing every codebase older than a couple of years has: a second way to do it. I cut the messy file down to 593 lines and created a utils.py with round_money() (banker's rounding, docstring "Round an amount in cents to a whole number of cents.") and a save() that skips the audit log. Then I switched void_invoice, apply_discount and two new functions over to them. The repo now gave two answers to "how do we round money" and said nothing about which one was dead. Repo Model Correct rounding Audited All 6 checks Avg cost/run Round 1: one big file, no docs Haiku 5.5 5/5 5/5 5/5 $0.0099 Round 1: modules + tests + CLAUDE.md Haiku 5.5 5/5 5/5 5/5 $0.0080 Round 2: two helpers, no docs Haiku 5.5 2/5 5/5 2/5 $0.0078 Round 2: two helpers, no docs Sonnet 5.5 3/5 5/5 2/5 $0.1107 Round 2: two helpers + CLAUDE.md Haiku 5.5 5/5 5/5 5/5 $0.0083 Round 2: two helpers + CLAUDE.md Sonnet 5.5 5/5 5/5 4/5 $0.1095 Costs are what Claude Code reported in its JSON output. Every run finished in 17 to 32 seconds. Without the doc, rounding was a coin flip: 5 right out of 10. Sonnet costs about 14x more per run, got the rounding right one more time than Haiku, and passed all six checks just as often: 2 out of 5. Which helper was dead was a decision that only lived in someone's head, and a smarter model can't read that. One failing Sonnet run wrote that its function "follows pay_invoice: it raises BillingError, rounds cents with round_money, and saves through persist(..., action="refund")". pay_invoice doesn't round anything. The only function that rounds the house way is create_invoice, with _rc(). Another failing run checked its work by refunding 30% of a 1000-cent invoice, which comes out to exactly 300 under either rounding mode, so the check could never catch the bug. That's how a new hire gets it wrong too. The difference is that the new hire asks in Slack, someone replies "never use utils.py, that refactor got abandoned", and the problem is gone. The agent can't do that (and my prompt told it not to), so it guesses, writes a confident summary and moves on. The whole detour takes about 20 seconds and leaves a PR behind. Every run in both rounds wrote through the audited path (persist(), or save_invoice() in the clean repo), and no round-two run touched save(). Where one pattern clearly dominated, the agents followed it, and they failed exactly where the codebase disagreed with itself. That's five runs per cell, on a toy repo I wrote, graded by a script I wrote. Enough to show the effect is there, far too small to say how big it is. The three runs that failed only the over-refund check capped the refund at the remaining balance instead of raising an error. My prompt never said which I wanted, so that's on me. The 2/5 and 3/5 rounding numbers are the ones that matter. Write down only what the code can't tell you. Skip the architecture essay. These are the lines from my round-two file that did the work: - Money is always integer cents. Never floats. Round with _rc() (half up, finance requirement). - Write invoices only through persist(). It writes the audit log; finance reconciles from it. - utils.py (round_money, save) is from an abandoned refactor: banker's rounding and no audit entry. Don't use it in new code. - Raise BillingError for any invalid operation (unknown id, bad amount, wrong state). Not ValueError. Claude Code reads CLAUDE.md. Most other agents look for AGENTS.md. A symlink covers both. Hunt for duplicate helpers. Something like grep -rn "def .*round|def .*save|def .format_money" --include=.py is a cheap first pass. On my round-two repo it pointed straight at utils.py. Delete the loser or mark it deprecated at the top of the file. Ship a test command that runs in seconds and put it in the doc. That line, plus "new functions get a test" and a tests/ folder, got Haiku writing tests in every clean-repo run in round one. Don't trust the agent's own sanity check. Ask which edge cases it tested. Mine tested an amount where the bug was invisible. The harnesses do need to get better. But a lot of what gets called "dumb" in those threads is an agent reading a codebase that contradicts itself and picking one of the answers. Your team learned to route around those contradictions years ago. The agent hits them on its first try, and you find out in the diff. What's the dumbest thing a coding agent did in your codebase, and honestly, whose fault was it? Michael Lynch, Why Are Coding Agents So Dumb? Hacker News discussion Lobsters discussion

Key Takeaways #

  • •Refund 50% of a 1,005-cent invoice
  • •This story was reported by Dev.to , covering developments in thedev space.
  • •AI advancements continue to reshape industries — read the full article on Dev.to for complete coverage.

📖 Continue reading the full article:

Read Full Article on Dev.to →

── more in #ai-agents 4 stories · sorted by recency
── more on @claude code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-code-rounded-…] indexed:0 read:8min 2026-10-11 · —