22:37
2026-07-23
lesswrong.com
ai-safety
The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology
A new study of 54 model organisms trained with seven methodologies across two base models finds that interpretability techniques detect quirks much more easily in models built with standard post-hoc fโฆ