04:00
2026-08-10
machinebrief.com
artificial-intelligence
Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits
Researchers found that safety alignment in Diffusion Large Language Models (DLLMs) is sparse and transferable, enabling attacks that increase attack success rates from 2.6% to 73.8% on LLaDA and from โฆ