Masked diffusion language models (dLLMs) generate text by iteratively denoising masked positions, re-predicting each token multiple times before it is committed. An autoregressive decoder exposes an answer's distribution once, at the step that commits it; a dLLM exposes it at every denoising step be
Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering
Researchers identified a targeted bias injection attack on masked diffusion language models (dLLMs) that exploits their iterative denoising process, in which each token's distribution is re-exposed at every denoising step rather than once at commit time as in autoregressive decoders. The attack, described as closed-loop activation steering, uses this repeated exposure to steer model outputs toward injected biases. The finding matters because dLLMs' multi-step re-prediction of each token gives attackers a wider surface for manipulation than autoregressive models provide.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.