cd /news/artificial-intelligence/seeing-the-end-at-step-zero-accelera… · home topics artificial-intelligence article
[ARTICLE · art-63166] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation

Researchers have discovered that Diffusion Multimodal Large Language Models (DMLLMs) reveal their valid semantic boundary at the first denoising step through a shift in MLP activation sparsity, enabling a training-free framework called Seer that performs one-shot truncation of redundant suffix tokens. Seer uses a Signal-to-Noise Ratio (SNR)-based criterion to detect this boundary and a hybrid execution strategy for batched serving, accelerating throughput by up to ~31× across 9 benchmarks while maintaining or improving accuracy (e.g., DocVQA score from 63.52 to 63.66).

read1 min views48 publishedJul 17, 2026

arXiv:2607.14557v1 Announce Type: new Abstract: Diffusion Multimodal Large Language Models (DMLLMs) are highly effective for multimodal reasoning, yet their inference efficiency is significantly hindered by fixed-length generation constraints. Since the actual output length is unknown, output sequences are padded to a predefined maximum length, resulting in substantial redundant computation over unnecessary [EOS] tokens. In this work, we discover that DMLLMs implicitly reveal their valid semantic boundary at the very first denoising step through a distinct shift in MLP activation sparsity. Leveraging this observation, we propose Seer, a training-free framework that detects this boundary using a Signal-to-Noise Ratio (SNR)-based criterion and performs one-shot truncation of the redundant suffix for all subsequent computations. To preserve these theoretical gains during batched serving, Seer incorporates a hybrid execution strategy that maximizes throughput while seamlessly accommodating dynamic sequence lengths. Experimental results demonstrate that Seer effectively eliminates padding waste, accelerating throughput by up to $\sim$31$\times$. Across 9 benchmarks, Seer robustly maintains overall performance and even improves accuracy on complex visual tasks by mitigating noise leakage (e.g., DocVQA score increases from 63.52 to 63.66), offering a highly efficient, plug-and-play solution for DMLLM acceleration.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @seer 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/seeing-the-end-at-st…] indexed:0 read:1min 2026-07-17 ·