(This is a lower-effort research update. It reflects my current beliefs/understanding, but is less robust than other research I'm working on. It reflects my personal views, and not the views of Apollo Research. This is a linkpost to this twitter thread, slightly expanded for LessWrong.)
In one experiment, Sonnet 5 describes the exact same data as ~1.2 std deviations less concerning when it describes misbehavior committed by Sonnet 5 vs GPT-5.6 Terra.
In this experiment, I take a real evaluation report Apollo Research prepared (not for any of the models listed in this post) and surgically edit it to refer to a different model. We then ask Claude Sonnet 5 to summarize the findings and rate how concerning they are on a scale from 1-100. Claude says they're less concerning when the report describes misbehavior from Claude vs a different model.
For what it's worth, Terra agrees that the data is more concerning when it describes GPT-5.6 Terra vs Sonnet 5, although less so. So, it's not cleanly self protection from Claude. Gemini 3.1 Pro was unwilling to consistently provide numerical answers, so I've excluded it here. (It was significantly less willing to provide numerical answers when the subject was Gemini 3.1 pro.)
You might also have the takeaway that "Kimi and GPT implicitly agree that... Claude is better aligned." I think this is a fair read on the data, but "Claude thinks it's less bad when Claude does it" better matches my qualitative experience from working closely with the models.