cd /news/artificial-intelligence/redteaming-leading-arabic-llms-with-… · home topics artificial-intelligence article
[ARTICLE · art-109702] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Redteaming Leading Arabic LLMs with ASAS

Researchers introduced the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming large language models, containing 801 prompts across 8 safety categories and 8 attack strategies. In evaluations of seven leading models including GPT-4o, Claude 3.7 Sonnet, ALLaM, and FANAR, most models failed to defend against 50% of unsafe prompts, with major gaps in high-harm categories such as weapons and illicit substances. The findings highlight that language alignment does not transfer across languages and that automated safety judges like GPT-4o underperform compared to human annotators.

read1 min views1 publishedAug 25, 2026

arXiv:2608.21985v1 Announce Type: new Abstract: As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safety and cultural alignment is increasingly critical. However, Arabic LLM safety remains underexplored, especially in adversarial evaluation settings. We introduce the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming LLMs. ASAS contains 801 prompts spanning 8 safety categories and 8 attack strategies, with ideal responses in Modern Standard Arabic (MSA). We conduct a redteaming evaluation across seven leading models with Arabic capabilities, including GPT-4o, Claude 3.7 Sonnet, and regional models such as ALLaM and FANAR. Human annotators rate responses using a structured 4-point safety scale, revealing that most models fail to defend against 50% of unsafe prompts. Our findings highlight major safety gaps in high-harm categories such as weapons and illicit substances, with direct and obfuscation-based attacks proving most effective. The results also show that language alignment does not readily transfer across languages, and that automated safety judges (e.g., GPT-4o) perform poorly compared to human annotators. ASAS provides a culturally grounded benchmark and redteaming protocol to drive progress in Arabic LLM safety.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @asas 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/redteaming-leading-a…] indexed:0 read:1min 2026-08-25 ·