cd /news/artificial-intelligence/visual-information-guided-parallel-d… · home topics artificial-intelligence article
[ARTICLE · art-113830] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models

Researchers propose the Visual Information-Guided Sampler (VIG-Sampler), a new decoding method for diffusion multimodal large language models that prioritizes tokens based on their attention to image tokens, outperforming the Info-Gain Sampler by an average of 19.3 CIDEr points across captioning benchmarks and surpassing it on COCO Caption while using only half as many decoding steps.

read1 min views1 publishedAug 28, 2026

arXiv:2608.26580v1 Announce Type: new Abstract: Diffusion multimodal large language models (dMLLMs) have recently emerged as a new decoding paradigm for multimodal generation. Starting from a fully masked sequence, dMLLMs progressively decode the sequence by unmasking a subset of the remaining masked positions at each step. Since the selected tokens serve as the prediction context for subsequent steps, deciding which tokens to decode is crucial to the quality of the final output. The most common strategy prioritizes tokens based on a certainty measure that tends to favor tokens frequently observed in the training data. Recent approaches instead order tokens according to their influence on subsequent predictions, but do not explicitly account for the input image. We propose the Visual Information-Guided Sampler (VIG-Sampler), which prioritizes tokens based on their attention to image tokens. We further impose a constraint that penalizes candidate tokens whose image-attention distributions are similar to those of previously selected tokens, thereby increasing the information gain of the decoded subset. Extensive experiments on 7 captioning and VQA benchmarks with 3 open-source dMLLMs demonstrate the effectiveness of VIG-Sampler, which outperforms the Info-Gain Sampler by an average of 19.3 CIDEr points across the captioning benchmarks and surpasses it on COCO Caption while using only half as many decoding steps.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @vig-sampler 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/visual-information-g…] indexed:0 read:1min 2026-08-28 ·