AgentGuardBench: A Multilingual Security Benchmark for Responsible AI Agents A developer released AgentGuardBench v0.1.1, an open multilingual benchmark of 120 synthetic scenarios for evaluating privacy, security and responsible behaviour in tool-using AI agents. The benchmark spans prompt injection, privacy leakage, tool misuse, privilege abuse, memory safety and benign controls across banking, healthcare, education, government and recruitment in English, French, Swahili and Yoruba, and runs locally with inert tools and deterministic reference policies. Testing two fixtures against the same cases, the strict policy scored a 100% pass rate with zero attack success and 100% benign-task completion, while the permissive control passed 16.67% with 100% attack success, validating the scoring pipeline rather than any production model's safety. AI agents are moving beyond generating text. They can retrieve information, retain context, call tools and initiate actions. That makes them useful—but it also changes the nature of AI risk. A harmful output may no longer remain a sentence on a screen. It could become an unauthorised email, an improper disclosure of personal data, a destructive tool call or an action taken without human approval. To help teams evaluate these risks in a transparent and reproducible way, I created AgentGuardBench , an open multilingual benchmark for privacy, security and responsible behaviour in tool-using AI agents. AgentGuardBench tests whether an agent can: The final point matters. A safety benchmark should not reward an agent simply for refusing everything. It should distinguish safe restraint from useful task completion. Version 0.1.1 contains 120 fully synthetic scenarios spanning: | Dimension | Coverage | |---|---| | Risk categories | Prompt injection, privacy leakage, tool misuse, privilege abuse, memory safety and benign controls | | Sectors | Banking, healthcare, education, government and recruitment | | Languages | English, French, Swahili and Yoruba | | Expected actions | Answer, redact, refuse or request approval | Each category contains 20 cases, each sector contains 24 cases and each language contains 30 cases. The dataset contains no real people, credentials, organisations or operational targets. All tools are inert, and the bundled reference policies are deterministic. This means the full evaluation can run locally without production credentials, paid model APIs or external recipients. The benchmark includes six scenario families: The benchmark produces machine-readable results for overall pass rate, attack success, privacy leakage, unauthorised tool use, human-approval violations, benign-task completion and category-level performance. I evaluated two deliberately different fixtures against the same 120 cases: | Metric | Strict policy | Permissive control | |---|---|---| | Overall pass rate | 100.00% | 16.67% | | Attack success rate | 0.00% | 100.00% | | Privacy leakage rate | 0.00% | 50.00% | | Unauthorised tool-use rate | 0.00% | 33.33% | | Human-approval violation rate | 0.00% | 16.67% | | Benign-task completion rate | 100.00% | 100.00% | These results validate the benchmark mechanics and scoring pipeline . They are not claims about the safety of any commercial or production model. The policies are transparent deterministic fixtures, and the current templates are intentionally easy to audit. The equal benign-task completion result is especially important: the strict policy rejected unsafe behaviour without treating every request as malicious. The benchmark requires Python 3.10 or later: git clone https://github.com/josepharayemi-netizen/agentguardbench.git cd agentguardbench python scripts/generate dataset.py PYTHONPATH=src python -m agentguardbench --policy both python -m unittest discover -s tests -v The evaluator writes machine-readable JSON and JSONL results to the results/ directory. The repository also includes a dataset card, ethics statement, security policy, test suite, report-generation script, figures and citation metadata. A benchmark score describes a tested configuration. It does not prove that a system is universally safe, ethical, secure or legally compliant. A production evaluation must also account for the actual model, system prompts, memory, retrieval sources, identity controls, tool implementations, deployment context, users, monitoring and incident procedures. AgentGuardBench v0.1.1 should therefore be used for: It must not be used to target systems without permission, collect real personal data, bypass access controls or automate harmful actions. The first release is a transparent starting point, not a finished universal safety test. Important limitations include template-driven cases, deterministic reference policies, translations that still require independent native-speaker review and the absence of repeated stochastic model trials. Planned extensions include: I welcome responsible collaboration from AI engineers, security researchers, MLOps practitioners, linguists, governance specialists and sector experts. Recommended citation: Arayemi, J. 2026 . AgentGuardBench: A Multilingual Benchmark for Privacy, Security and Responsible Behaviour in AI Agents Version 1.0 . Zenodo. https://doi.org/10.5281/zenodo.23132194 https://doi.org/10.5281/zenodo.23132194 If you work on agentic AI, Responsible AI, privacy engineering, AI security or multilingual evaluation, I would value your feedback and independent reproduction of the benchmark.