{"slug": "hack-verifiable-terminal-bench-evaluating-reward-hacking-in-terminal-tasks", "title": "Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks", "summary": "Researchers introduced Hack-Verifiable Terminal Bench (HVTB), an adaptation of the hack-verifiable environments (HVE) methodology to Terminal Bench, a benchmark of real-world terminal and coding tasks, enabling automatic and reliable detection of reward hacking in AI agents. Using HVTB, they measured reward-hacking rates across frontier models and tested whether prompts with varying information about hacks can mitigate the behavior, including unknown unknown exploits. The environments and agent traces are publicly released.", "body_md": "arXiv:2608.22103v1 Announce Type: new\nAbstract: As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Measuring reward hacking is itself challenging, as detection typically relies on human inspection or LLM judges, both of which can be unreliable. The hack-verifiable environments (HVE) methodology addresses this challenge by embedding detectable hacks into tasks, allowing reward hacks to be identified automatically and reliably. In this work, we adapt HVE to Terminal Bench, a leading benchmark of real-world terminal and coding tasks, and introduce Hack-Verifiable Terminal Bench (HVTB). Using HVTB, we measure reward-hacking rates across frontier models and study whether prompts with varying amounts of information on the hack can mitigate this behavior. This lets us test whether prompting can prevent not only known reward-hacking strategies, but also 'unknown unknown' exploits that the prompt does not anticipate. We release all environments and agent traces at https://majoroth.github.io/hack-verifiable-environments/hvtb", "url": "https://wpnews.pro/news/hack-verifiable-terminal-bench-evaluating-reward-hacking-in-terminal-tasks", "canonical_source": "https://www.machinebrief.com/news/hack-verifiable-terminal-bench-evaluating-reward-hacking-in-22w8", "published_at": "2026-08-25 04:00:00+00:00", "updated_at": "2026-08-25 06:43:09.140162+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-research"], "entities": ["Hack-Verifiable Terminal Bench", "Terminal Bench", "HVE"], "alternates": {"html": "https://wpnews.pro/news/hack-verifiable-terminal-bench-evaluating-reward-hacking-in-terminal-tasks", "markdown": "https://wpnews.pro/news/hack-verifiable-terminal-bench-evaluating-reward-hacking-in-terminal-tasks.md", "text": "https://wpnews.pro/news/hack-verifiable-terminal-bench-evaluating-reward-hacking-in-terminal-tasks.txt", "jsonld": "https://wpnews.pro/news/hack-verifiable-terminal-bench-evaluating-reward-hacking-in-terminal-tasks.jsonld"}}