{"slug": "compress-and-forget-bitsandbytes-quantization-amplifies-proactive-interference", "title": "Compress and Forget: bitsandbytes Quantization Amplifies Proactive Interference in LLMs", "summary": "A new study from researchers including Shayan Shahrabi finds that bitsandbytes 4-bit (INT4/NF4) post-training quantization significantly amplifies proactive interference in large language models, reducing retrieval accuracy under high interference from 81.0% to 68.3% for Qwen2.5-7B-Instruct, with paired McNemar's tests confirming significance (p ≤ 2.6 × 10^-6). The effect, also present but smaller at INT8 in two of three models, is mechanistically linked to increased same-key intrusion errors (from 21.5% to 24.6%, p = 4.8 × 10^-7) and originates in the quantized transformer backbone, suggesting a hidden cost for applications with long, updatable contexts.", "body_md": "arXiv:2608.18578v1 Announce Type: new\nAbstract: Proactive interference (PI) is a documented failure mode in large language models in which retrieval of a repeatedly overwritten value degrades as prior overwrites accumulate, mirroring a classical phenomenon in human working memory. Post-training quantization (PTQ) is now the default deployment path for open-weight models, yet its effect on this failure mode has not been tested. We evaluate three precision levels (FP16, INT8, INT4/NF4, via bitsandbytes) across three architecturally distinct instruction-tuned models (Qwen2.5-7B-Instruct, Mistral-7B-Instruct-v0.3, Phi-3.5-mini-instruct), holding the retrieval task fixed. INT4 quantization significantly reduces accuracy under high interference in every model (e.g., from 81.0% to 68.3% for Qwen), confirmed by paired McNemar's tests ($p \\le 2.6 \\times 10^{-6}$) and a mixed-effects regression spanning all interference levels; INT8, often assumed safe, also carries a smaller but real penalty in two of three models. The effect is specific to semantically similar (word-type) distractors and reverses sign under a numeric control condition, and is mechanistically linked to a rise in same-key intrusion errors under INT4 (from 21.5% to 24.6% of trials, $p = 4.8 \\times 10^{-7}$). A follow-up ablation shows the effect originates in the quantized transformer backbone rather than the output projection layer. These results suggest that bitsandbytes 4-bit quantization can impose an additional cost on applications relying on long, updatable, semantically dense contexts, even when aggregate benchmark accuracy appears largely unaffected. We release our code and tokenizer-verified vocabulary construction method at https://github.com/ShayanShahrabi/compress-and-forget", "url": "https://wpnews.pro/news/compress-and-forget-bitsandbytes-quantization-amplifies-proactive-interference", "canonical_source": "https://www.machinebrief.com/news/compress-and-forget-bitsandbytes-quantization-amplifies-proa-tx6y", "published_at": "2026-08-20 04:00:00+00:00", "updated_at": "2026-08-20 05:14:45.105147+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-infrastructure"], "entities": ["bitsandbytes", "Qwen2.5-7B-Instruct", "Mistral-7B-Instruct-v0.3", "Phi-3.5-mini-instruct", "Shayan Shahrabi", "arXiv", "GitHub"], "alternates": {"html": "https://wpnews.pro/news/compress-and-forget-bitsandbytes-quantization-amplifies-proactive-interference", "markdown": "https://wpnews.pro/news/compress-and-forget-bitsandbytes-quantization-amplifies-proactive-interference.md", "text": "https://wpnews.pro/news/compress-and-forget-bitsandbytes-quantization-amplifies-proactive-interference.txt", "jsonld": "https://wpnews.pro/news/compress-and-forget-bitsandbytes-quantization-amplifies-proactive-interference.jsonld"}}