04:00
2026-08-28
arxiv.org
artificial-intelligence
Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models
Researchers propose the Visual Information-Guided Sampler (VIG-Sampler), a new decoding method for diffusion multimodal large language models that prioritizes tokens based on their attention to image …