cd /news/ai-safety/understanding-in-context-multimodal-… · home topics ai-safety article
[ARTICLE · art-126581] src=machinebrief.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting

A new arXiv paper (2609.10613v1) proposes a posterior reweighting framework that models safety-aligned multimodal large language models (MLLMs) as implicitly operating over competing behavioral modes, explaining in-context learning jailbreaks as inference-time evidence that shifts the model's posterior preference between safe and harmful behaviors. The framework yields predictive scaling laws for demonstration count, harmful ratio, adversarial strength, and semantic diversity, and the authors introduce a posterior-aware inference-time defense that injects benign counter-evidence based on estimated risk. The defense reportedly achieves a significantly improved robustness-utility trade-off versus existing in-context defenses under a fixed intervention budget.

by read1 min views1 publishedSep 11, 2026

arXiv:2609.10613v1 Announce Type: cross Abstract: In-context learning (ICL) jailbreaks reveal a critical vulnerability in multimodal large language models (MLLMs): harmful demonstrations in the prompt can induce unsafe outputs without modifying model parameters. Despite extensive empirical evidence, existing work lacks a principled understanding of why such jailbreaks reliably succeed or how their effectiveness scales with context composition. We propose a posterior reweighting framework that models a safety-aligned MLLM as implicitly operating over competing behavioral modes, and interprets in-context demonstrations as inference-time evidence that dynamically shifts the model's posterior preference between safe and harmful behaviors. This view formalizes jailbreak as a process of evidence accumulation, yielding predictive scaling laws with respect to demonstration count, harmful ratio, adversarial strength, and semantic diversity. Guided by this framework, we introduce a posterior-aware inference-time defense that adaptively injects benign counter-evidence based on estimated risk, effectively suppressing harmful posterior drift while preserving model utility. Compared to existing in-context defenses, our method achieves a significantly improved robustness-utility trade-off under a fixed intervention budget. Together, our results establish posterior reweighting as a unifying and predictive framework for understanding and mitigating ICL jailbreak in MLLMs.

── more in #ai-safety 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/understanding-in-con…] indexed:0 read:1min 2026-09-11 ·