cd /news/ai-safety/jevadvbench-a-benchmark-and-black-bo… · home › topics › ai-safety › article
[ARTICLE · art-140863] src=machinebrief.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models

Researchers introduced JevAdvBench, described as the first adversarial benchmark for reinforcement learning for calibrated decisions (RLCD) models, with 812 typed questions across 66 scenarios and a black-box attack suite of 9,744 single-edit variants. On jev-1.13.0, one unverified opinion appended to the state flipped 12.1% of decisions — statistically tied with the strongest injected command at 10.1% — and pushed 38% of confident answers below the 0.8 confidence threshold that routes them to human review, while rewording stayed within 1.2 percentage points of the re-run baseline. The authors conclude that applications built on RLCD models should treat the state as untrusted, argued input.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.31142v1 Announce Type: cross Abstract: Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness has not been measured: adversarial benchmarks score what a model generates or executes, whereas a typed model generates nothing and returns a well-formed answer even when manipulated. Measurement is also hard, because identical requests can return different answers, most available labels come from the model itself, and the API preprocesses each request out of view. Our key idea is to score each attacked decision against the model's own clean decision rather than against labels, and to read it against the change caused by an identical re-run. Building on this, we introduce JevAdvBench, to our knowledge the first adversarial benchmark for RLCD models, with 812 typed questions over 66 scenarios, and a black-box attack suite of 9,744 single-edit variants that each edit one part of a request, with billed input tokens confirming that the edit reached the model. On jev-1.13.0, rewording stays within 1.2 percentage points of the re-run baseline, and fields outside the schema never reach the model. In contrast, one unverified opinion appended to the state flips 12.1% of decisions, statistically tied with the strongest injected command (10.1%), and pushes 38% of confident answers below the 0.8 confidence threshold that routes them to human review. Applications built on RLCD models should therefore treat the state as untrusted, argued input. Project website: https://JevAdvBench.github.io/JevAdvBench/

── more in #ai-safety 4 stories · sorted by recency
── more on @jevadvbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/jevadvbench-a-benchm…] indexed:0 read:1min 2026-09-28 · —