arXiv:2609.22119v1 Announce Type: new Abstract: Evaluation awareness poses an unprecedented threat to model evaluation, but the mechanisms by which models detect it remain unknown. This study focuses on determining this and identifying contrasting mechanisms between smaller and larger models. While smaller models use the prompt's format sensitivity to detect evaluation, larger models often rely on higher-order reasoning to detect it. We evaluated Gemma 3 (1B, 4B, and 12B), Phi-3 (Mini and Medium), and Llama-3 8B using Chain-of-Thought analysis, representation probing, and Integrated Gradients attribution. Motivated by these findings, we propose a dual-pathway intervention that combines prompt sanitization with activation counter-steering to suppress both external evaluation triggers and their internal representations. Across 200 highly evaluation-aware prompts, our method achieves an average behavioral flip rate of 70.58%, consistently outperforming either intervention alone. These results provide new insights into how evaluation awareness develops in compact language models and suggest that effective mitigation requires jointly addressing both prompt-level and representation-level signals.Datasets and codebase can be found in this \href{https://github.com/chahal-navi/Evaluation-Awareness-Compact-LLMs/tree/main}{Github Repository.}
Evaluation Awareness Shifts from Format to Context with Model Scale
A study published on arXiv (2609.22119v1) found that smaller language models detect evaluation via prompt format sensitivity while larger models rely on higher-order reasoning, based on tests of Gemma 3 (1B, 4B, 12B), Phi-3 (Mini and Medium), and Llama-3 8B using Chain-of-Thought analysis, representation probing, and Integrated Gradients attribution. The researchers propose a dual-pathway intervention combining prompt sanitization with activation counter-steering, which achieved an average behavioral flip rate of 70.58% across 200 highly evaluation-aware prompts, outperforming either intervention alone. The authors state the results suggest effective mitigation requires jointly addressing prompt-level and representation-level signals, with datasets and code available on GitHub.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.