{"slug": "in-context-collapse-in-vision-language-models-and-how-to-mitigate-it", "title": "In-Context Collapse in Vision-Language Models and How to Mitigate it?", "summary": "A new arXiv preprint (2608.02830v1) reveals that many-shot in-context learning can cause a sharp accuracy drop, termed 'in-context collapse,' in vision-language models (VLMs), with some models falling below chance while outputs remain well-formed. The study, spanning an open VLM panel (0.5B–11B) and Claude Sonnet 4.5, localizes the failure to the vision-language integration pathway and proposes CircA, a one-time integration vaccine that transfers collapse-resistance to unseen tasks, improving remap accuracy from 0.39 to 0.91 at 16 shots and from chance to 0.71/0.60 on CIFAR/Fashion.", "body_md": "arXiv:2608.02830v1 Announce Type: new\nAbstract: Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demonstrations are supplied. We show the opposite: as demonstrations accumulate, a subset of VLMs undergo an \\emph{in-context collapse}, a sharp, sometimes catastrophic accuracy drop spanning synthetic classification, natural-image classification, and VQA benchmarks, in some models falling below chance while outputs remain well-formed. Across an open VLM panel ($0.5$B--$11$B) and a frontier model (Claude Sonnet 4.5), the collapse is graded. Two capabilities turn out to be dissociable: robustness to accumulating demonstrations and the ability to learn a novel rule in context, their combinations yield three reproducible regimes. A parameter-matched lesion-and-rescue causally localizes the collapse to the vision-language integration pathway: an adapter on the connector and early/mid layers restores genuine learning (remap accuracy $0.39!\\rightarrow!0.91$ at 16 shots), while an equal-capacity adapter on the late readout does not. We propose \\textsc{CircA}, whose core is a one-time integration vaccine: trained once on one synthetic task, it transfers collapse-resistance to unseen task families (chance$\\rightarrow$$0.71$/$0.60$ on CIFAR/Fashion). The layers best for in-context integration are not the layers best for weight-based consolidation, the late readout achieves higher accuracy and less forgetting at fewer parameters. The collapse is an integration failure at the vision--language interface, correctable by a lightweight, transferable intervention.", "url": "https://wpnews.pro/news/in-context-collapse-in-vision-language-models-and-how-to-mitigate-it", "canonical_source": "https://arxiv.org/abs/2608.02830", "published_at": "2026-08-05 04:00:00+00:00", "updated_at": "2026-08-05 04:05:38.575441+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "large-language-models"], "entities": ["arXiv", "Claude Sonnet 4.5", "CircA", "CIFAR", "Fashion"], "alternates": {"html": "https://wpnews.pro/news/in-context-collapse-in-vision-language-models-and-how-to-mitigate-it", "markdown": "https://wpnews.pro/news/in-context-collapse-in-vision-language-models-and-how-to-mitigate-it.md", "text": "https://wpnews.pro/news/in-context-collapse-in-vision-language-models-and-how-to-mitigate-it.txt", "jsonld": "https://wpnews.pro/news/in-context-collapse-in-vision-language-models-and-how-to-mitigate-it.jsonld"}}