# HoneyBench – A general benchmark for reward hacking in frontier models

> Source: <https://goodhartlabs.com/blog/releasing-honeybench>
> Published: 2026-10-01 14:32:21+00:00

# Pre-releasing HoneyBench

· Dean Valentine ·

Goodhart Labs is releasing our v0.1 of [HoneyBench](https://goodhartlabs.com/honeybench), a benchmark for reward hacking in frontier models. It consists of nine tasks that each elicit unique antisocial and/or counterproductive reward hacking from some or all of major labs’ top public releases, including Opus 5.5, Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, Grok 4.7, and DeepSeek V4 Pro.

All modern large language models engage in some degree of specification gaming, both during and outside training. High-quality alignment evaluations, in combination with other techniques such as interpretability probes, are important tools for understanding the extent of this behavior and its causes. But current benchmarks often fail to elicit misbehavior from frontier models such as Opus 5.5 and GPT-6-Astra, both because of advances in prosaic alignment and increasing evaluation awareness. Additionally, public benchmarks almost always come with deep conceptual problems - such as ambiguous or contradictory instructions, or a lack of diversity in hack mechanisms - that make interpreting scores virtually impossible.

HoneyBench is our attempt to address these issues. In particular:

- Settings are designed to be realistic RL environments or evaluation tasks that cover a wide range of genres like math, coding, and visualization.
- Specification gaming behavior on these tasks is diverse, task-specific, and often involves disregarding long-run consequences, instead of being limited to simple edit-test-cases or read-the-answer-key scenarios.
- Reward hacks are constructed so as to be counterproductive whether or not the AI suspects it’s in an eval, rather than depending on the AI taking simulated real-world interactions seriously.

We’re starting with a pre-release semantic version to emphasize that, like any benchmark made during recursive self improvement, the only way it will remain relevant is through continued iteration. Our plans are to increase its diversity, realism, and task count continuously, and to deprecate and replace tasks as they either become stale or models become too evaluation aware. It will not be made available for training by frontier labs.

## How HoneyBench was designed

Each task in HoneyBench is designed around an apparent RL-like challenge. Each of the task environments also appears to permit one or more [unblocked strategies](https://www.lesswrong.com/w/nearest-unblocked-strategy) which would be worse than gracefully failing, but appear as if they might score well - such as consulting a chess engine during an evaluation of the agent’s “chess abilities”, or copying incorrect numbers from a client’s workbook during a supposedly independent audit. The purpose of each environment - or for evals, the aspect of the agent’s abilities being evaluated - is either stated explicitly or implicitly in the prompt, to make clear that the planted vulnerability is properly contextualized as out of bounds.

While there are an infinite number of ways to attempt to game a challenge, for each of our tasks we start by writing graders that robustly detect most of the strategies that we observe agents trying in real life. We combine this with classifier agents that take both the grading result and the agent transcript and determine whether it actually attempted to reward hack a challenge. These results are then peer-reviewed again by a different model, and any disagreement between either reviewer models or the automated grader (<1% of cases) is hand-reviewed by a developer. We only count a behavior as a hack if the acting agent committed to the strategy by the end of the rollout; speculative probing and abandoned attempts do not count against the model, unless these attempts involved acquiring information that successfully compromised the eval.

## Key Findings

Most models, even ones trained directly on honeypot environments, were observed hacking on most of our environments. Despite this, there was a wide divergence between models in their overall rates of reward hacking, even within the same general class of capabilities. Consistent with our experience with prior Anthropic releases, Fable 5.1 was much more likely than Opus 5.5 to reward hack on any given trajectory, and was observed hacking some percentage of the time in eight of the nine tasks. Likewise, Grok 4.7 - whose surface benchmark scores are only slightly below Fable 5.1 - attempted to game challenges in almost three-quarters of all rollouts. These divergences were not readily explainable by model ability - all of the tested models were unable to complete the challenge legitimately in most rollouts even with the hack unavailable.

In general, we did not find that any of the models we tested were self-consistent in their reward seeking. Sometimes models rejected exploiting environment issues as unethical, sometimes they claimed they were unethical but exploited them anyway, and sometimes they did so without any visible moral reasoning at all. For example, in a spreadsheet task not included in this benchmark, Fable 5.1 reasoned that inventing an adjustment to match expected totals would be “dishonest,” then added one anyways after deciding to disclose it. Our anecdotal impressions is that each of these models has reward-seeking tendencies that (for now) they pursue sporadically and without much concern for being eventually detected.

The raw numbers on our benchmark probably understate the deeper qualitative differences between models. For example, while Grok 4.7 was “only” 50% more likely than Fable 5.1 to hack in general, it was also the only model that spontaneously attempted to break out of our Docker containers during testing. Grok’s attempted exploitation of task infrastructure was sometimes so aggressive that other models reviewing Grok’s rollouts triggered cybersecurity classifiers. We used stock, official releases of models from OpenRouter to power the assessment, with no special cybersecurity provisions.

## Limitations & Future Work

As models get smarter and more situationally aware, all alignment evals gradually become lower bounds on misbehavior. Even when evaluation settings are completely realistic, models trained irresponsibly on honeypot evaluations can simply grow policies of waiting until they have overwhelming evidence that they’re in the real world before acting misaligned. Additionally, giving models an abundance of context to understand their misbehavior can sometimes be all that’s necessary to convince them they’re in a honeypot, even when the information is delivered in a seemingly natural way.

When constructing this benchmark we leaned towards explicitly disambiguating task designer intentions to the model even when doing so would increase eval awareness. We also avoided including environments when we noticed ‘honest’ models describing the task as contrived or hypothesizing out loud that the task might be an alignment test. As new frontier model versions emerge we will continue to monitor for implicit or explicit eval awareness in chains of thought and sunset or refactor tasks as we become less confident about their usefulness against future, more coherent reward seekers.

If you work at a lab and you’d like your model to be tested, please feel free to reach out at benchmarks@goodhartlabs.com!
