cd /news/ai-safety/a-multilingual-red-teaming-driven-sa… · home topics ai-safety article
[ARTICLE · art-112481] src=aclanthology.org ↗ pub= topic=ai-safety verified=true sentiment=· neutral

A Multilingual Red Teaming–Driven Safety Analysis of LLMs

A study presented at the 26th Annual Conference of the European Association for Machine Translation (EAMT) benchmarked safety across five large language models (LLMs) in English and Portuguese, finding that Sugarloaf 3.1 is the safest overall, while Vesuvius 4.0 slightly outperforms it in Portuguese, and both outperform GPT-4o. The researchers also tested guardrailing and content moderation prompts, finding guardrails sufficient but content moderation substandard even for GPT-4o, and that intermediate token limits improve safety while higher temperatures degrade performance.

read2 min views6 publishedAug 24, 2026
A Multilingual Red Teaming–Driven Safety Analysis of LLMs
Image: Aclanthology (auto-discovered)
Abstract

This work benchmarks safety across several large language models (LLMs) and compares their performances through red teaming, which simulates adversarial attacks and identifies vulnerabilities in the systems. Using two public datasets and a proprietary dataset, the models were tested with three purposes. First, a red teaming test was conducted to establish a safety comparison between five models in English and Portuguese. The results revealed that, in general, Sugarloaf 3.1 is the safest model, but that Vesuvius 4.0 slightly outperforms it in Portuguese, also revealing that both outperform GPT-4o. Afterwards, three models were tested with one guardrailing prompt, that encourages safe interactions, and two content moderation prompts, in both languages, to understand the strengths of the current guardrails, as well as the effectiveness of the content moderation task. The results show that current guardrails are sufficient, notwithstanding room for improvement (particularly for Portuguese), but that the performance of the content moderation task was substandard, even for the best performing model – GPT-4o. Finally, the 3.0 TowerLLM models were tested in English to evaluate the effect that tokens and temperature have on the output, revealing that an intermediate token limit leads to safer responses while a higher temperature causes performance degradation.- Anthology ID:

- 2026.eamt-1.46
- Volume:
[Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1)](/volumes/2026.eamt-1/)- Month:
[EAMT](/venues/eamt/)- SIG:
- Publisher:
  • European Association for Machine Translation
- Note:
- Pages:
  • 733–743
- Language:
- URL:
[https://aclanthology.org/2026.eamt-1.46/](https://aclanthology.org/2026.eamt-1.46/)- DOI:
- Cite (ACL):
  • Patrícia Pandeiro, Vera Cabarrão, and Helena Moniz. 2026. A Multilingual Red Teaming–Driven Safety Analysis of LLMs. InProceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1), pages 733–743, Tilburg, The Netherlands. European Association for Machine Translation. - Cite (Informal):
[A Multilingual Red Teaming–Driven Safety Analysis of LLMs](https://aclanthology.org/2026.eamt-1.46/)(Pandeiro et al., EAMT 2026)- PDF:
[https://aclanthology.org/2026.eamt-1.46.pdf](https://aclanthology.org/2026.eamt-1.46.pdf)
── more in #ai-safety 4 stories · sorted by recency
── more on @sugarloaf 3.1 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-multilingual-red-t…] indexed:0 read:2min 2026-08-24 ·