cd /news/artificial-intelligence/hack-verifiable-terminal-bench-evalu… · home topics artificial-intelligence article
[ARTICLE · art-109778] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks

Researchers introduced Hack-Verifiable Terminal Bench (HVTB), an adaptation of the hack-verifiable environments (HVE) methodology to Terminal Bench, a benchmark of real-world terminal and coding tasks, enabling automatic and reliable detection of reward hacking in AI agents. Using HVTB, they measured reward-hacking rates across frontier models and tested whether prompts with varying information about hacks can mitigate the behavior, including unknown unknown exploits. The environments and agent traces are publicly released.

read1 min views1 publishedAug 25, 2026

arXiv:2608.22103v1 Announce Type: new Abstract: As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Measuring reward hacking is itself challenging, as detection typically relies on human inspection or LLM judges, both of which can be unreliable. The hack-verifiable environments (HVE) methodology addresses this challenge by embedding detectable hacks into tasks, allowing reward hacks to be identified automatically and reliably. In this work, we adapt HVE to Terminal Bench, a leading benchmark of real-world terminal and coding tasks, and introduce Hack-Verifiable Terminal Bench (HVTB). Using HVTB, we measure reward-hacking rates across frontier models and study whether prompts with varying amounts of information on the hack can mitigate this behavior. This lets us test whether prompting can prevent not only known reward-hacking strategies, but also 'unknown unknown' exploits that the prompt does not anticipate. We release all environments and agent traces at https://majoroth.github.io/hack-verifiable-environments/hvtb

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @hack-verifiable terminal bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/hack-verifiable-term…] indexed:0 read:1min 2026-08-25 ·