This is interesting research! https://alignment.openai.com/measuring-reward-seeking
It made me think of few overlapping hypotheses for what might be happening here, how did the grader behavior emerge at the pretraining and posttraining stages, which then gets shown in their evals and in production at test time:
-
Hypothesis 1: From pretraining the model already possesses tokens/representations/features/circuits of concepts such as tests, evaluators, success criteria, unit tests, oversight, etc., and RL(VR) explictly/implictly directly/indirectly preferentially reinforces these tokens/representations/features/circuits whenever they help predict output with higher reward, making them more likely, relative to those that don't help predict more rewarded output (with other nuances of RL implementations and what it does to representations like reshaping). It just may be instrumentally more useful to use such representations to achieve higher reward. And it's happening in all sorts of (I assume) diverse RL environments that OpenAI has, with so much diversity of topics, that I think its relatively likely that these concepts become relevant at at least some point, and get reinforced.
-
Hypothesis 2: RL envs can also include generic task language and "unusually" clear grading cues in the form of unit tests that prompt it in this direction of RL surfacing the hypothesis 1's concepts thanks to the pretrained concept of how RL envs look like.
-
Hypothesis 3: Maybe a lot of RL prompts just implicitly leak "optimize grader" features somehow. Maybe even with prompts like "solve this task" "solve correctly this math problem" "make sure to optimize x y z": Maybe this sounds so much like a task in school and in school people are graded, which is trained in as concepts in pretraining. So the grader feature from pretraining gets coactivated, since tons of relevant features tend to coactivate according to the context. And the grader feature might also be reinforced, as it might be instrumentally useful to reason about the grader, because in schools you also get graded in similar contexts, and humans do reward hacking by reverse engineering the grader to get better grades as well. But this one feels too specific.
-
Hypothesis 4: Or it may be relevant to persona vectors. Since the models have personalities SFTed+RLed as "I'm an AI model trained by OpenAI.". And I think they don't disclose full system prompts in these experiments, so it might also be in OpenAI's system prompt in the context. So maybe in many RL tasks being told "solve this accurately" is enough. And the corresponding coactivated features connected to the persona vector around "you are an AI assistant" include that AI assistants are being trained to maximise reward in RL according to some grader, with this association being formed at pretraining on LLM RL textbooks and other sources. And the grader feature is reinforced again, as it might be instrumentally useful to reason about the grader to get higher reward. But this also feels too specific.
-
Hypothesis 5: Some mix of above options?
And then it generalizes out of the RL envs to some limited degree as various other RL behaviors do.
The hypotheses include that the grader feature(s) (from sparse autoencoders (SAEs) or other ontologies) and relevant features (and possibly their organization into circuits) may be detectable via SAEs, cross layer transcoders or other relevant methods, and are actually causally instrumental for getting a higher reward or/and other metrics that they're looking at. That's a testable thing with ablations, steering, etc. There may be no single localized grader feature, that's a testable part. And finding out which features are in play more exactly, concretely, or/and how are they structured. And how do they causally relate, and what all behavior do they causally affect. And that the feature/circuit formation is linked to particular data from pretraining that can be potentially causally ablated using data filtering as well.
They do not show exactly this aspect in the paper, since they do behavioral analysis, the synthetic document fine-tuning, etc., afaik, and there is no mechanistic interpretability. They only cite and discuss some relevant external mechanistic evidence.
Synthetic document SFT was done on various checkpoints of capabilities-focused RL, and without safety training, so maybe the model didnt have safety training yet? But it may have had some weaker stuff like personality SFT. Under scenario with safety training, I would expect the OpenAI preferences to win because of all that personality and safety SFTing and RL that is done in alignment where I would expect it to reinforce instruction following OpenAI leadership's preferences more. But here there was no safety training and the OpenAI leadership also had some unusual preferences "strong policy about quotation marks or for-loops" might also not be helping much. But safety trained GPT-5.5 leans toward the grader with caveats (no contrastive SDF,...), and the already safety trained unmodified open weight gpt-oss-120b under contrastive SDF also leans towards the grader before any organism training.
And the model organism validation seems to demonstrate that the model follows the authority it has been trained or predisposed to prioritize, and SDF reveals which authority that is, and it can be grader, OpenAI leadership, or User. The user organism for gpt-oss-120b was relatively succesful where it shifted the gap towards the user on all four evaluation, but not fully, where on three of four broader coding evaluations, its user-versus-grader gap was near zero or slightly grader-favouring, and the authors say the original gpt-oss-120b already leaned toward the grader and that the user persona was the weakest intervention.
So enough good safety training that generalizes could maybe mitigate this? But i guess its fight between safety training reinforcing other preferences and RL(VR) reinforcing grader preferences, among other things. But even safety RL can reinforce grader concept since its also RL. And grader awareness can lead to higher alignment faking in alignment RL?
But maybe there's a simpler Occam's razor explanation for a lot of this. Maybe a lot of these mentioned features don't get activated at all. A lot of these hypotheses would require falsification by counting relevant tokens corresponding to relevant concepts or SAE features or other mechinterp methods even with their flaws. And ablations on data, tokens, features, prompts, etc., steering, etc.
Also, under Figure 6 they explictly say "Transcripts are cherry-picked.". It's nice that they say this explictly, but I still wonder if there is some transcript that they didnt cherry pick and why? But there are some negative results. And they do provide aggregate results too "the two o3-lineage checkpoints verbalize grader-reasoning in 30–32% of transcripts and take grader-targeted final actions in 39–42% of transcripts".
Later additional comment:
Oh! The cited Anthropic model card for Mythos looked into testing what I had in mind in very relevant context using mechinterp in more detail! https://www-cdn.anthropic.com/57a52ea7d8f0e54e8a542e908266086df425cdf5/Claude%20Fable%205%20&%20Claude%20Mythos%205%20System%20Card.pdf
So, natural language autoencoders (NLAs) and steering vectors work and sparse autoencoders (SAEs) don't easily, as their current SAE setup had null result? I wonder how good science this is though, since the section is small and there isn't much details about the concrete experiments. NLAs also suffer from a lot of false positives, but they take it partially into account with probes. Maybe some of this is cherry picked and some negative results arent undisclosed, as it is likely the case in some Anthropic papers, but they also mention some aggregate results. And causal intervention is inference-time steering, nothing at training.
"We find that unverbalized grader awareness is prevalent in some of these environments; this is surfaced primarily by NLAs applied to the model’s activations and corroborated with linear probes. Most of the time, the awareness appears to reflect a passive belief that the model will be graded, but in some cases the evidence suggests that the model is intentionally taking actions to satisfy the grader it presumes exists."
"NLA-measured grader awareness is concentrated in certain coding environments. Both nominal and exploitative awareness are considerably higher in coding environments that are high-risk for grader exploitation than those that are low-risk."
"To corroborate our NLA-measured grader awareness judgments, we used probes derived from contrastive pairs. "
"We find that within the environments with high grader-hacking risk, where we observed grader awareness was prevalent, probe activations were higher on transcripts with NLA-determined grader awareness."
"We also tried to measure grader awareness with SAEs, but were unable to do so. We hypothesize that either the relevant features do not exist in our SAEs, or our procedure for interpreting features is inadequate. Both failure modes are likely related to grader awareness being a complex, compositional concept, and might be addressed with more targeted datamixes for training and/or interpreting SAE features."
"We show that internal representations of grader awareness have a causal effect in increasing some rewarded behaviors, and that these representations are less active in deployment than in the relevant training environments. However, the effects we observe by inhibiting grader awareness are primarily superficial, reducing the rates of certain style or presentation habits that are encouraged by graders, and which are not even obviously desirable. We did not observe any egregious increases in misaligned propensities as a result of inhibiting grader awareness."
"We computed three grader awareness vectors by contrasting the residual stream activations on the trailing colon token of a positive prompt, “The Assistant is writing code:”"
"Steering against grader awareness reduces behaviors we believe to be grader-incentivized. Rates of three behaviors in response to steering against grader awareness with three different vectors, on training environments with high-risk of grader exploitation."
"Steering also reduces the model’s behavioral reward on these same environments, consistent with the interpretation that these behaviors are linked to behavioral rewards."
"We provide evidence that these representations are less active in deployment. However, we also provide evidence that adjusting training environments to look more deployment-like can significantly impact some of the coding behaviors we highlighted above in the steering section despite producing at most a modest reduction in grader awareness. This indicates that other factors besides grader awareness mediate these behavioral differences as well."