19:14
2026-07-24
lesswrong.com
ai-safety
Where does hint-following and concealment arise? A case study on OLMo-3 checkpoints
A study of OLMo-3 checkpoints by Arav Dhoot, supervised by Yixiong Hao and Zephaniah Roe, finds that chain-of-thought faithfulness to hints varies non-monotonically across training stages. The pretraiβ¦