SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? A new benchmark called SAEScientist-Bench evaluates whether AI agents can conduct autonomous sparse autoencoder (SAE) interpretability research, addressing a gap in recursive self-improvement work that has focused mainly on automating model training pipelines. The benchmark targets post-hoc monitoring and auditing of what models learn, which the source frames as necessary for reliable autonomous development and safe alignment. While research on recursive self-improvement RSI has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essent