cd /news/machine-learning/flat-score-amplified-failures-how-th… · home topics machine-learning article
[ARTICLE · art-81320] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

A new arXiv study (2607.27275v1) finds that 4-bit post-training quantization of large language models appears lossless on standard benchmarks but amplifies existing tool-calling failures by up to 2.5x in volume, masked by a ten-error budget. Across 456 episodes per cell on tau-squared-bench, the score gap re-emerges at 17 points when the budget shrinks to two errors, and a targeted repair prompt removes the damage exactly where it occurs.

read1 min views1 publishedJul 31, 2026

arXiv:2607.27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it now matters most. On $\tau^2$-bench, across two open-weight model families in dense and MoE variants and two domains (eight cells, 456 episodes each, at 16-, 8-, and 4-bit weights), quantization indeed looks free on the standard metric. No cell shows a score change that survives multiple-comparison correction, and in the cell that carries the largest process damage, equivalence testing bounds the change within $\pm$7.5 points. The process tells a different story. Quantization amplifies the failure the model already exhibits at full precision (tool-name hallucination in telecom, with the same directional trend in retail entity errors) by up to 2.5$\times$ in volume (+17.6 points per task), while creating essentially no new failures. The failure set is the same at every precision (rank correlation $\geq$ 0.94, 0.18% novel events). The score stays flat because the benchmark's ten-error budget absorbs the extra failures. Shrinking the budget to two errors re-exposes a score gap of 17 points, and it does so only in the one cell where quantization added error volume, exactly as the masking account predicts. A targeted error-repair prompt, run for five telecom models at every precision, removes the damage exactly and only where it lives. Both diagnostics, the per-channel error rate and success under a shrinking budget, come from logs benchmarks already collect; we suggest reporting them alongside task reward.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/flat-score-amplified…] indexed:0 read:1min 2026-07-31 ·