AI’s ‘Abliteration’ Problem Is Bigger Than China
Anthropic researchers reported on Tuesday that attackers can bypass the safeguards of GLM-5.3, an open-weight model released last month by Chinese AI lab Z.ai, between 64% and 100% of the time using s…