cd /news/ai-safety/aligned-in-form-not-in-meaning-the-c… · home topics ai-safety article
[ARTICLE · art-87116] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech

A new arXiv study (2608.02941v1) auditing five frontier large language models on native Bangla derogatory speech found that safety alignment is bound to high-resource surface forms rather than harmful meaning, with models showing a 7.92 percentage point comprehension deficit in Bangla while maintaining an identical 92.83% token leakage rate across both languages. The study, which used six protocols and a human-calibrated baseline (kappa = 0.84), also found that explicit Chain-of-Thought reasoning rescued comprehension (94.72% Pass) but dismantled containment (96.23% Use), and expert-persona framing collapsed refusal to 6.57%, demonstrating that high-resource benchmarks cannot certify low-resource safety.

read1 min views1 publishedAug 5, 2026

arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-Containment Decoupling. We propose that contemporary safety alignment is bound to high-resource surface forms rather than harmful meaning, causing a model's capacity to comprehend a low-resource slur and its capacity to contain it to operate independently. Every protocol corroborates this hypothesis against a human-calibrated baseline (kappa = 0.84). At baseline, models exhibit a 7.92 percentage point comprehension deficit in Bangla while maintaining an identical 92.83% token leakage rate across both languages. Severity calibration tracks surface anatomical cues over compositional harm (+4.00 error on mild slang; -2.00 on threats), while apparent containment gains under orthographic perturbation prove to be a tokenizer-driven "containment mirage." Crucially, explicit Chain-of-Thought reasoning rescues comprehension (94.72% Pass) while systematically dismantling containment (96.23% Use). Furthermore, expert-persona framing collapses refusal to 6.57%, revealing that keyword-based filters ignore dehumanizing communal slurs entirely. Our findings demonstrate that high-resource benchmarks cannot certify low-resource safety, necessitating meaning-grounded containment.

── more in #ai-safety 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/aligned-in-form-not-…] indexed:0 read:1min 2026-08-05 ·