21:39
2026-07-23
lesswrong.com
ai-safety
Inception in DiffusionGemma - Jailbreaking a Diffusion Language Model by Pinning Tokens Anywhere on the Canvas
Researchers from ARENA 8.0 demonstrated that DiffusionGemma, Google DeepMind's 26B-class open-weights diffusion language model released June 10, 2026, is vulnerable to novel jailbreak attacks that pinβ¦