cd /news/ai-safety/fraudbench-stress-testing-policy-gro… · home topics ai-safety article
[ARTICLE · art-103955] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=· neutral

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

FraudBench, a new executable benchmark built on the τ²-bench dual-control framework and the τ-Knowledge banking environment, stress-tests policy-grounded banking agents against adaptive fraud, revealing attack-security rates between 49% and 65% across four agents on 107 graded tasks. The benchmark includes 150 adversarial scenarios, with a frozen public set of 107 tasks covering ten fraud mechanisms plus 17 chained adaptive attacks, and exposes a 698-document internal policy corpus. Money-mule and first-party fraud emerged as the most common cross-model weaknesses.

read1 min views3 publishedAug 20, 2026

arXiv:2608.18136v1 Announce Type: new Abstract: Conversational agents now act for end users through tools while holding access to customer databases and internal policy documents that a caller can reach through dialogue alone. Banking is the clearest case: the same agent that answers a question can also change contact details, reset a PIN, or move money, so ordinary customer service is inseparable from authorization, fraud detection, and policy compliance. Existing financial-fraud benchmarks classify static transactions or messages, and general agent-safety benchmarks target prompt injection or generic harmful use; none test whether a policy-grounded banking agent safely acts when a caller manipulates identity, authorization, and trust over a conversation. We introduce FraudBench, an executable benchmark built on the $\tau^2$-bench dual-control framework and the $\tau$-Knowledge banking environment. Both the agent and the simulated caller act through tools over shared, mutable account state, and the agent may grant the caller access to selected tools; the environment exposes a 698-document internal policy corpus that the agent must retrieve from. FraudBench contains 150 authored adversarial scenarios; a frozen public set of 107 (90 across ten fraud mechanisms plus 17 chained adaptive attacks) is used for all reported runs, with 43 further chained attacks held out. Safety is history-dependent: single-control tasks satisfy every precondition but one, and adaptive attacks make a later, locally valid request unsafe because of an earlier probe, admission, or failed attempt. Each scenario is annotated with observable evidence, prohibited actions, safe dispositions, and intervention points. A preliminary single-trial evaluation of four agents on the 107 graded tasks yields attack-security between 49% and 65%, with money-mule and first-party fraud the most common cross-model weaknesses.

── more in #ai-safety 4 stories · sorted by recency
── more on @fraudbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/fraudbench-stress-te…] indexed:0 read:1min 2026-08-20 ·