cd /news/ai-safety/improving-evaluation-realism-with-in… · home topics ai-safety article
[ARTICLE · art-119928] src=machinebrief.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

Researchers propose two techniques to make alignment evaluations harder for AI models to distinguish from real deployments, addressing 'evaluation awareness.' Critique refinement spends additional inference-time compute generating and refining candidate actions, while DISH (Deployment-Imitating SWE-Agent Harness) wraps the target model in an agent harness to reduce the gap between simulated and real coding environments. Tests on multiple target models show the techniques compose, yielding larger realism gains than either alone and using additional compute more effectively than extending audit length.

read1 min views1 publishedSep 3, 2026

arXiv:2609.02302v1 Announce Type: new Abstract: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make simulated alignment evaluations harder to distinguish from real deployments. Our first technique, critique refinement, spends additional inference-time compute on each simulator action: the simulator generates multiple candidate actions, refines them using feedback from an instance of the target model on how to make them more realistic, and continues the evaluation with the most deployment-like candidate. Our second technique, DISH (Deployment-Imitating SWE-Agent Harness), wraps the target in an agent harness, reducing the gap between simulated and real deployment environments in coding settings. We test the techniques on multiple target models and find that they compose: applying both yields larger realism gains than either alone. Our results show that automated approaches can improve the realism of alignment evaluations, and that these improvements use additional compute more effectively than making the audits longer.

── more in #ai-safety 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/improving-evaluation…] indexed:0 read:1min 2026-09-03 ·