04:00
2026-09-22
arxiv.org
artificial-intelligence
Evaluation Awareness Shifts from Format to Context with Model Scale
A study published on arXiv (2609.22119v1) found that smaller language models detect evaluation via prompt format sensitivity while larger models rely on higher-order reasoning, based on tests of Gemma…