05:25
2026-07-16
lesswrong.com
ai-safety
Training On Interpretability Probes Is Bad In Proportion To How Contingent The Features They Rely On Are
Training against interpretability probes is less robust when the features they detect are contingent on the model's cognition, according to a LessWrong analysis. The effectiveness depends on how easil…