{"slug": "understanding-in-context-multimodal-jailbreaks-via-posterior-reweighting", "title": "Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting", "summary": "A new arXiv paper (2609.10613v1) proposes a posterior reweighting framework that models safety-aligned multimodal large language models (MLLMs) as implicitly operating over competing behavioral modes, explaining in-context learning jailbreaks as inference-time evidence that shifts the model's posterior preference between safe and harmful behaviors. The framework yields predictive scaling laws for demonstration count, harmful ratio, adversarial strength, and semantic diversity, and the authors introduce a posterior-aware inference-time defense that injects benign counter-evidence based on estimated risk. The defense reportedly achieves a significantly improved robustness-utility trade-off versus existing in-context defenses under a fixed intervention budget.", "body_md": "arXiv:2609.10613v1 Announce Type: cross \nAbstract: In-context learning (ICL) jailbreaks reveal a critical vulnerability in multimodal large language models (MLLMs): harmful demonstrations in the prompt can induce unsafe outputs without modifying model parameters. Despite extensive empirical evidence, existing work lacks a principled understanding of why such jailbreaks reliably succeed or how their effectiveness scales with context composition. We propose a posterior reweighting framework that models a safety-aligned MLLM as implicitly operating over competing behavioral modes, and interprets in-context demonstrations as inference-time evidence that dynamically shifts the model's posterior preference between safe and harmful behaviors. This view formalizes jailbreak as a process of evidence accumulation, yielding predictive scaling laws with respect to demonstration count, harmful ratio, adversarial strength, and semantic diversity. Guided by this framework, we introduce a posterior-aware inference-time defense that adaptively injects benign counter-evidence based on estimated risk, effectively suppressing harmful posterior drift while preserving model utility. Compared to existing in-context defenses, our method achieves a significantly improved robustness-utility trade-off under a fixed intervention budget. Together, our results establish posterior reweighting as a unifying and predictive framework for understanding and mitigating ICL jailbreak in MLLMs.", "url": "https://wpnews.pro/news/understanding-in-context-multimodal-jailbreaks-via-posterior-reweighting", "canonical_source": "https://www.machinebrief.com/news/understanding-in-context-multimodal-jailbreaks-via-posterior-igjb", "published_at": "2026-09-11 04:00:00+00:00", "updated_at": "2026-09-11 05:27:05.386110+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "ai-research", "artificial-intelligence"], "entities": ["arXiv", "multimodal large language models", "MLLMs"], "alternates": {"html": "https://wpnews.pro/news/understanding-in-context-multimodal-jailbreaks-via-posterior-reweighting", "markdown": "https://wpnews.pro/news/understanding-in-context-multimodal-jailbreaks-via-posterior-reweighting.md", "text": "https://wpnews.pro/news/understanding-in-context-multimodal-jailbreaks-via-posterior-reweighting.txt", "jsonld": "https://wpnews.pro/news/understanding-in-context-multimodal-jailbreaks-via-posterior-reweighting.jsonld"}}