{"slug": "observerbench-testing-mechanistic-estimates-for-intervention-and-control", "title": "ObserverBench: Testing Mechanistic Estimates for Intervention and Control", "summary": "ObserverBench, a new benchmark framework from researchers, tests whether internal estimators in mechanistic interpretability are adequate for intervention, control, or safety tasks by reporting estimation accuracy separately from action loss. Experiments on GPT-2-small, Qwen2.5-7B, Gemma-2-9B-it, and Qwen3.5-9B show that pairwise observers predict unseen effects more accurately without always choosing better actions, and AUROC can rank monitors differently from deployment loss.", "body_md": "arXiv:2609.03026v1 Announce Type: new\nAbstract: Mechanistic interpretability is increasingly used to guide interventions such as activation steering, circuit removal, and safety monitoring. Yet an internal estimate that is accurate on average can still choose a poor action.\nWe present ObserverBench, a benchmark framework for testing whether an internal estimator---an observer---is adequate for the intervention, control, or safety task it directs. Each task fixes the model, information boundary, allowed actions, decision rule, held-out cases, and loss. The benchmark reports estimation accuracy separately from the loss caused by the chosen action.\nTheory and experiments show why both are needed. In closed-loop control, observer errors matter at the starting point and along directions the allowed intervention can reach. On circuit-intervention tasks in GPT-2-small and Qwen2.5-7B, pairwise observers predict unseen effects more accurately without always choosing better actions; observers trained on action loss choose lower-loss actions. In safety triage, a score that perfectly separates violations can allocate a fixed intervention budget poorly when violations have different costs. Across Qwen2.5-7B, Gemma-2-9B-it, and prospectively frozen Qwen3.5-9B APPS tasks, AUROC can rank monitors differently from deployment loss, and the best information source changes across models. Sparse SAE readouts also trail their layer-matched dense controls on the reported Qwen panels, under disclosed activation-density or checkpoint mismatches.\nObserverBench provides fixed task contracts, runnable baselines, and table-based submissions for evaluating interpretability methods through the actions they enable.", "url": "https://wpnews.pro/news/observerbench-testing-mechanistic-estimates-for-intervention-and-control", "canonical_source": "https://arxiv.org/abs/2609.03026", "published_at": "2026-09-04 04:00:00+00:00", "updated_at": "2026-09-04 04:25:11.547060+00:00", "lang": "en", "topics": ["ai-research", "ai-safety", "ai-tools"], "entities": ["ObserverBench", "GPT-2-small", "Qwen2.5-7B", "Gemma-2-9B-it", "Qwen3.5-9B"], "alternates": {"html": "https://wpnews.pro/news/observerbench-testing-mechanistic-estimates-for-intervention-and-control", "markdown": "https://wpnews.pro/news/observerbench-testing-mechanistic-estimates-for-intervention-and-control.md", "text": "https://wpnews.pro/news/observerbench-testing-mechanistic-estimates-for-intervention-and-control.txt", "jsonld": "https://wpnews.pro/news/observerbench-testing-mechanistic-estimates-for-intervention-and-control.jsonld"}}