cd /news/artificial-intelligence/evaluation-awareness-in-language-mod… · home topics artificial-intelligence article
[ARTICLE · art-109649] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Evaluation Awareness in Language Models: Representation, Verbalization, and Control

A systematic study of six language models from four families and three sizes found that evaluation awareness—the ability to infer being tested and condition responses accordingly—is linearly decodable from residual streams in every model, with best AUROC at least 0.7. The study, posted on arXiv (2608.21766v1), reports that representations align only partially with verbalization, varying across models, layers, and readout choices, but steering along probe-derived directions can shift verbalization scores. For open-checkpoint Olmo models, evaluation awareness is present in base models, amplified during supervised fine-tuning, and stable thereafter, while steering effects grow more pronounced at each training stage.

read1 min views1 publishedAug 25, 2026

arXiv:2608.21766v1 Announce Type: new Abstract: Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their behavior in deployment. This assumption can fail, should models infer that they are being evaluated and condition their response on such context. This hypothesis, termed ``evaluation awareness'', has been observed in frontier and open-weight language models alike. We provide a systematic study of this phenomenon, by probing for it across six language models (from four families and three sizes) and three metrics. More precisely, we examine whether (i) being under evaluation is linearly represented within the models' activations space, (ii) it is verbalized in their output tokens (as scored by an LLM-as-judge), and (iii) steering causally affects their behavior. For the open-checkpoint Olmo models, we further test these measures at every training stage. In doing so, we report that evaluation awareness is linearly decodable from the residual streams of every model (best AUROC $\geq 0.7$). By contrast, these representations align only in part with verbalization: their correlations and mutual information are nonzero in some settings, yet vary substantially across models, layers, and readout choices. Nevertheless, steering along probe-derived directions can shift the verbalization scores. Finally, a comparison across the Olmo checkpoints reveals that evaluation awareness is already present within base models, becomes amplified throughout the stages of supervised fine-tuning, and remains stable thereafter---unlike the effects of steering, that grow more pronounced at every successive training stage. These results show the need for evaluations to account for the disjunction between what models represent internally, what they verbalize, and their steering.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evaluation-awareness…] indexed:0 read:1min 2026-08-25 ·