Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering Researchers identified a targeted bias injection attack on masked diffusion language models (dLLMs) that exploits their iterative denoising process, in which each token's distribution is re-exposed at every denoising step rather than once at commit time as in autoregressive decoders. The attack, described as closed-loop activation steering, uses this repeated exposure to steer model outputs toward injected biases. The finding matters because dLLMs' multi-step re-prediction of each token gives attackers a wider surface for manipulation than autoregressive models provide. Masked diffusion language models dLLMs generate text by iteratively denoising masked positions, re-predicting each token multiple times before it is committed. An autoregressive decoder exposes an answer's distribution once, at the step that commits it; a dLLM exposes it at every denoising step be