{"slug": "same-quantity-different-answer-numerical-representation-invariance-in-language", "title": "Same Quantity, Different Answer: Numerical Representation Invariance in Language Models", "summary": "A study of 3,600 exact-rational word problems and 8,600 prompts across five open-weight language models found canonical accuracy of 0.969-0.996 but orbit correctness of only 0.848-0.981 and orbit invariance of 0.851-0.981 when numerically equivalent quantities were rewritten as decimals, fractions, percentages, number words, scientific notation, or converted units. The authors report that most strict-parser failures stem from multiplication-form scientific notation falling outside the implemented number grammar, while Mistral Small 4 scored 0.699 on unit-converted inputs and produced 265 errors differing from the label by exact powers of ten. In a separate 9,000-call experiment with equal calls per arm, representation consensus did not outperform paraphrase consensus on a low-error subset and generated substantially more false alarms.", "body_md": "arXiv:2609.25009v1 Announce Type: new \nAbstract: Numerically equivalent word problems should yield the same canonical answer whether a quantity is written as a decimal, fraction, percentage, number word, scientific notation, or an exactly converted unit. We generate 3,600 exact-rational problems and 8,600 prompts spanning five identity-preserving transformation families, and evaluate five open-weight systems. After a fixed syntax audit that normalizes common answer forms without an LLM judge, canonical accuracy is 0.969-0.996, but orbit correctness falls to 0.848-0.981 and orbit invariance to 0.851-0.981; invariant-but-wrong orbits account for at most 0.003. Most of the broad strict-parser collapse arises because multiplication-form scientific notation lies outside the implemented number grammar, illustrating how evaluator interfaces can masquerade as reasoning failures. A distinct semantic pathology remains: Mistral Small 4 scores 0.699 on unit-converted inputs and produces 265 errors differing from the label by exact powers of ten. In a separate 9,000-call experiment that allocates equal calls to the compared arms, representation consensus does not outperform paraphrase consensus on a low-error subset and produces substantially more false alarms. The accompanying ancillary archive contains the frozen benchmark, evaluation and audit records, consensus raw responses, manifests, analysis code, and a one-command paper build.", "url": "https://wpnews.pro/news/same-quantity-different-answer-numerical-representation-invariance-in-language", "canonical_source": "https://arxiv.org/abs/2609.25009", "published_at": "2026-09-23 04:00:00+00:00", "updated_at": "2026-09-23 04:25:17.568488+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "natural-language-processing", "machine-learning"], "entities": ["Mistral Small 4", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/same-quantity-different-answer-numerical-representation-invariance-in-language", "markdown": "https://wpnews.pro/news/same-quantity-different-answer-numerical-representation-invariance-in-language.md", "text": "https://wpnews.pro/news/same-quantity-different-answer-numerical-representation-invariance-in-language.txt", "jsonld": "https://wpnews.pro/news/same-quantity-different-answer-numerical-representation-invariance-in-language.jsonld"}}