{"slug": "redteaming-leading-arabic-llms-with-asas", "title": "Redteaming Leading Arabic LLMs with ASAS", "summary": "Researchers introduced the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming large language models, containing 801 prompts across 8 safety categories and 8 attack strategies. In evaluations of seven leading models including GPT-4o, Claude 3.7 Sonnet, ALLaM, and FANAR, most models failed to defend against 50% of unsafe prompts, with major gaps in high-harm categories such as weapons and illicit substances. The findings highlight that language alignment does not transfer across languages and that automated safety judges like GPT-4o underperform compared to human annotators.", "body_md": "arXiv:2608.21985v1 Announce Type: new\nAbstract: As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safety and cultural alignment is increasingly critical. However, Arabic LLM safety remains underexplored, especially in adversarial evaluation settings. We introduce the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming LLMs. ASAS contains 801 prompts spanning 8 safety categories and 8 attack strategies, with ideal responses in Modern Standard Arabic (MSA). We conduct a redteaming evaluation across seven leading models with Arabic capabilities, including GPT-4o, Claude 3.7 Sonnet, and regional models such as ALLaM and FANAR. Human annotators rate responses using a structured 4-point safety scale, revealing that most models fail to defend against 50% of unsafe prompts. Our findings highlight major safety gaps in high-harm categories such as weapons and illicit substances, with direct and obfuscation-based attacks proving most effective. The results also show that language alignment does not readily transfer across languages, and that automated safety judges (e.g., GPT-4o) perform poorly compared to human annotators. ASAS provides a culturally grounded benchmark and redteaming protocol to drive progress in Arabic LLM safety.", "url": "https://wpnews.pro/news/redteaming-leading-arabic-llms-with-asas", "canonical_source": "https://www.machinebrief.com/news/redteaming-leading-arabic-llms-with-asas-2136", "published_at": "2026-08-25 04:00:00+00:00", "updated_at": "2026-08-25 05:14:29.158145+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-research"], "entities": ["ASAS", "GPT-4o", "Claude 3.7 Sonnet", "ALLaM", "FANAR"], "alternates": {"html": "https://wpnews.pro/news/redteaming-leading-arabic-llms-with-asas", "markdown": "https://wpnews.pro/news/redteaming-leading-arabic-llms-with-asas.md", "text": "https://wpnews.pro/news/redteaming-leading-arabic-llms-with-asas.txt", "jsonld": "https://wpnews.pro/news/redteaming-leading-arabic-llms-with-asas.jsonld"}}