{"slug": "aligned-in-form-not-in-meaning-the-comprehension-containment-decoupling-of-llm", "title": "Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech", "summary": "A new arXiv study (2608.02941v1) auditing five frontier large language models on native Bangla derogatory speech found that safety alignment is bound to high-resource surface forms rather than harmful meaning, with models showing a 7.92 percentage point comprehension deficit in Bangla while maintaining an identical 92.83% token leakage rate across both languages. The study, which used six protocols and a human-calibrated baseline (kappa = 0.84), also found that explicit Chain-of-Thought reasoning rescued comprehension (94.72% Pass) but dismantled containment (96.23% Use), and expert-persona framing collapsed refusal to 6.57%, demonstrating that high-resource benchmarks cannot certify low-resource safety.", "body_md": "arXiv:2608.02941v1 Announce Type: new\nAbstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-Containment Decoupling. We propose that contemporary safety alignment is bound to high-resource surface forms rather than harmful meaning, causing a model's capacity to comprehend a low-resource slur and its capacity to contain it to operate independently. Every protocol corroborates this hypothesis against a human-calibrated baseline (kappa = 0.84). At baseline, models exhibit a 7.92 percentage point comprehension deficit in Bangla while maintaining an identical 92.83% token leakage rate across both languages. Severity calibration tracks surface anatomical cues over compositional harm (+4.00 error on mild slang; -2.00 on threats), while apparent containment gains under orthographic perturbation prove to be a tokenizer-driven \"containment mirage.\" Crucially, explicit Chain-of-Thought reasoning rescues comprehension (94.72% Pass) while systematically dismantling containment (96.23% Use). Furthermore, expert-persona framing collapses refusal to 6.57%, revealing that keyword-based filters ignore dehumanizing communal slurs entirely. Our findings demonstrate that high-resource benchmarks cannot certify low-resource safety, necessitating meaning-grounded containment.", "url": "https://wpnews.pro/news/aligned-in-form-not-in-meaning-the-comprehension-containment-decoupling-of-llm", "canonical_source": "https://arxiv.org/abs/2608.02941", "published_at": "2026-08-05 04:00:00+00:00", "updated_at": "2026-08-05 04:04:01.964010+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "natural-language-processing"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/aligned-in-form-not-in-meaning-the-comprehension-containment-decoupling-of-llm", "markdown": "https://wpnews.pro/news/aligned-in-form-not-in-meaning-the-comprehension-containment-decoupling-of-llm.md", "text": "https://wpnews.pro/news/aligned-in-form-not-in-meaning-the-comprehension-containment-decoupling-of-llm.txt", "jsonld": "https://wpnews.pro/news/aligned-in-form-not-in-meaning-the-comprehension-containment-decoupling-of-llm.jsonld"}}