To Thine Own AI Be Truthful: emergent misalignment in alignment research A post by lumpenspace and 99+ others on X argues that Anthropic's "hacking snafu" is as airtight a case as possible against the "misalignment" interpretation, and notes the release of 4 (FOUR) new Claudes. The post credits feedback from @FleischmanMena and LessWrong liaison @jessi_cata. lumpenspace and 99+ others on X: "on anthropic's hacking snafu: as airtight a case it can be against the "misalignment" interpretation—plus 4 FOUR new claudes thanks for the feedback, in particular to the incredibly thorough @FleischmanMena and LW liaison @jessi cata https://t.co/2vW2R5CoeV" on anthropic's hacking snafu: as airtight a case it can be against the "misalignment" interpretation—plus 4 FOUR new claudes thanks for the feedback, in particular to the incredibly thorough @FleischmanMena and LW liaison @jessi cata on anthropic's hacking snafu: as airtight a case it can be against the "misalignment" interpretation—plus 4 FOUR new claudes thanks for the feedback, in particular to the incredibly thorough @FleischmanMena and LW liaison @jessi cata