03:20
2026-09-11
twitter.com
ai-safety
To Thine Own AI Be Truthful: emergent misalignment in alignment research
A post by lumpenspace and 99+ others on X argues that Anthropic's "hacking snafu" is as airtight a case as possible against the "misalignment" interpretation, and notes the release of 4 (FOUR) new Claβ¦