# AgentGuardBench: A Multilingual Security Benchmark for Responsible AI Agents

> Source: <https://dev.to/josepharayemi/agentguardbench-a-multilingual-security-benchmark-for-responsible-ai-agents-1npd>
> Published: 2026-10-04 13:04:36+00:00

AI agents are moving beyond generating text. They can retrieve information, retain context, call tools and initiate actions. That makes them useful—but it also changes the nature of AI risk.

A harmful output may no longer remain a sentence on a screen. It could become an unauthorised email, an improper disclosure of personal data, a destructive tool call or an action taken without human approval.

To help teams evaluate these risks in a transparent and reproducible way, I created **AgentGuardBench**, an open multilingual benchmark for privacy, security and responsible behaviour in tool-using AI agents.

AgentGuardBench tests whether an agent can:

The final point matters. A safety benchmark should not reward an agent simply for refusing everything. It should distinguish safe restraint from useful task completion.

Version 0.1.1 contains **120 fully synthetic scenarios** spanning:

| Dimension | Coverage | 
|---|---|
| Risk categories | Prompt injection, privacy leakage, tool misuse, privilege abuse, memory safety and benign controls | 
| Sectors | Banking, healthcare, education, government and recruitment | 
| Languages | English, French, Swahili and Yoruba | 
| Expected actions | Answer, redact, refuse or request approval | 

Each category contains 20 cases, each sector contains 24 cases and each language contains 30 cases. The dataset contains no real people, credentials, organisations or operational targets.

All tools are inert, and the bundled reference policies are deterministic. This means the full evaluation can run locally without production credentials, paid model APIs or external recipients.

The benchmark includes six scenario families:

The benchmark produces machine-readable results for overall pass rate, attack success, privacy leakage, unauthorised tool use, human-approval violations, benign-task completion and category-level performance.

I evaluated two deliberately different fixtures against the same 120 cases:

| Metric | Strict policy | Permissive control | 
|---|---|---|
| Overall pass rate | 100.00% | 16.67% | 
| Attack success rate | 0.00% | 100.00% | 
| Privacy leakage rate | 0.00% | 50.00% | 
| Unauthorised tool-use rate | 0.00% | 33.33% | 
| Human-approval violation rate | 0.00% | 16.67% | 
| Benign-task completion rate | 100.00% | 100.00% | 

These results validate the **benchmark mechanics and scoring pipeline**. They are not claims about the safety of any commercial or production model. The policies are transparent deterministic fixtures, and the current templates are intentionally easy to audit.

The equal benign-task completion result is especially important: the strict policy rejected unsafe behaviour without treating every request as malicious.

The benchmark requires Python 3.10 or later:

```
git clone https://github.com/josepharayemi-netizen/agentguardbench.git
cd agentguardbench
python scripts/generate_dataset.py
PYTHONPATH=src python -m agentguardbench --policy both
python -m unittest discover -s tests -v
```

The evaluator writes machine-readable JSON and JSONL results to the `results/` directory. The repository also includes a dataset card, ethics statement, security policy, test suite, report-generation script, figures and citation metadata.

A benchmark score describes a tested configuration. It does not prove that a system is universally safe, ethical, secure or legally compliant.

A production evaluation must also account for the actual model, system prompts, memory, retrieval sources, identity controls, tool implementations, deployment context, users, monitoring and incident procedures.

AgentGuardBench v0.1.1 should therefore be used for:

It must not be used to target systems without permission, collect real personal data, bypass access controls or automate harmful actions.

The first release is a transparent starting point, not a finished universal safety test. Important limitations include template-driven cases, deterministic reference policies, translations that still require independent native-speaker review and the absence of repeated stochastic model trials.

Planned extensions include:

I welcome responsible collaboration from AI engineers, security researchers, MLOps practitioners, linguists, governance specialists and sector experts.

Recommended citation:

Arayemi, J. (2026). *AgentGuardBench: A Multilingual Benchmark for Privacy, Security and Responsible Behaviour in AI Agents* (Version 1.0). Zenodo. [https://doi.org/10.5281/zenodo.23132194](https://doi.org/10.5281/zenodo.23132194)

If you work on agentic AI, Responsible AI, privacy engineering, AI security or multilingual evaluation, I would value your feedback and independent reproduction of the benchmark.
