cd /news/large-language-models/same-quantity-different-answer-numer… · home topics large-language-models article
[ARTICLE · art-137793] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Same Quantity, Different Answer: Numerical Representation Invariance in Language Models

A study of 3,600 exact-rational word problems and 8,600 prompts across five open-weight language models found canonical accuracy of 0.969-0.996 but orbit correctness of only 0.848-0.981 and orbit invariance of 0.851-0.981 when numerically equivalent quantities were rewritten as decimals, fractions, percentages, number words, scientific notation, or converted units. The authors report that most strict-parser failures stem from multiplication-form scientific notation falling outside the implemented number grammar, while Mistral Small 4 scored 0.699 on unit-converted inputs and produced 265 errors differing from the label by exact powers of ten. In a separate 9,000-call experiment with equal calls per arm, representation consensus did not outperform paraphrase consensus on a low-error subset and generated substantially more false alarms.

by read1 min views1 publishedSep 23, 2026

arXiv:2609.25009v1 Announce Type: new Abstract: Numerically equivalent word problems should yield the same canonical answer whether a quantity is written as a decimal, fraction, percentage, number word, scientific notation, or an exactly converted unit. We generate 3,600 exact-rational problems and 8,600 prompts spanning five identity-preserving transformation families, and evaluate five open-weight systems. After a fixed syntax audit that normalizes common answer forms without an LLM judge, canonical accuracy is 0.969-0.996, but orbit correctness falls to 0.848-0.981 and orbit invariance to 0.851-0.981; invariant-but-wrong orbits account for at most 0.003. Most of the broad strict-parser collapse arises because multiplication-form scientific notation lies outside the implemented number grammar, illustrating how evaluator interfaces can masquerade as reasoning failures. A distinct semantic pathology remains: Mistral Small 4 scores 0.699 on unit-converted inputs and produces 265 errors differing from the label by exact powers of ten. In a separate 9,000-call experiment that allocates equal calls to the compared arms, representation consensus does not outperform paraphrase consensus on a low-error subset and produces substantially more false alarms. The accompanying ancillary archive contains the frozen benchmark, evaluation and audit records, consensus raw responses, manifests, analysis code, and a one-command paper build.

── more in #large-language-models 4 stories · sorted by recency
── more on @mistral small 4 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/same-quantity-differ…] indexed:0 read:1min 2026-09-23 ·