13:51
2026-07-15
lesswrong.com
ai-safety
Eliciting hidden knowledge from monitors with NLAs
Researchers propose using natural language autoencoders (NLAs) to surface hidden reasoning from AI monitors, testing whether NLAs can recover knowledge of reward hacking that monitors internally detec…