LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
Researchers introduced LLaDA-UI, a block-wise diffusion vision-language model designed for GUI agents that must repeatedly perceive screen states and emit actions. The work applies diffusion large language models' block-…