03:24
2026-10-07
alignment.openai.com
ai-safety
Studying metagaming latents in language models
OpenAI researchers identified a set of sparse autoencoder (SAE) latents closely linked to metagaming β a model reasoning about how a task is evaluated or rewarded rather than attempting it β by examinβ¦