cd /news/artificial-intelligence/llm-safety-alignment-in-low-resource… · home › topics › artificial-intelligence › article
[ARTICLE · art-100826] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

A systematic literature review of 50 studies from roughly 1,500 papers found that large language models (LLMs) have significantly weaker safety alignment in low-resource languages, with translated English benchmarks failing to capture culturally rooted harms and models vulnerable to cross-lingual jailbreaks and code-switching attacks. The review, conducted using the PRISMA 2020 methodology and drawing from Semantic Scholar, arXiv, and OpenAlex, identifies uneven multilingual pre-training coverage and insufficient native-language preference data as key drivers, and notes that African languages have fewer safety benchmarks than other regions.

read1 min views25 publishedAug 18, 2026

arXiv:2608.14626v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages. In this paper, we conduct a Systematic Literature Review (SLR) of LLM safety alignment in low-resource languages by adopting the PRISMA 2020 methodology. Out of roughly 1,500 papers identified from Semantic Scholar, arXiv, and OpenAlex, 50 relevant studies have been selected and analyzed. Our review is organized around four themes: safety alignment methods, multilingual safety risks, evaluation benchmarks, and cross-lingual transferability. We further propose a taxonomy of safety alignment approaches based on three adaptation mechanisms: data adaptation, objective optimization, and mechanistic alignment. Across literature, translated English benchmarks fail to sufficiently represent culturally rooted harms, and multilingual models are more vulnerable to cross-lingual jailbreaks, code-switching attacks, and safety degradation in underrepresented languages. These failures are driven by several key factors, including uneven multilingual pre-training coverage, insufficient native-language preference data, poor transfer of safety representations, and a lack of culturally aware evaluation frameworks. The review also notes that many low-resource languages, especially African languages, have fewer safety benchmarks available than other multilingual regions. Overall, the results reveal a persistent multilingual safety gap, and suggest that future progress will require culturally grounded benchmarks, participatory data collection, balanced multilingual pre-training, and scalable multilingual alignment methods.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/llm-safety-alignment…] indexed:0 read:1min 2026-08-18 · —