cd /news/large-language-models/videomm-adaptive-macro-micro-inferen… · home topics large-language-models article
[ARTICLE · art-131034] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

Researchers introduced VideoMM, an adaptive macro-micro inference framework for video multimodal large language models that decouples region selection from detailed reasoning, according to an arXiv paper (arXiv:2609.16722v1). VideoMM achieves a 6.13x speedup and a 7.4% accuracy gain over full-context baselines on LongVideoBench, and accelerates inference by 2.73x over current leading methods. The framework performs semantic filtering on a cost-effective Macro Proxy derived from downscaled frames, projecting selected regions onto high-fidelity Micro Tokens only when necessary, with code available at https://github.com/adfh917k/VideoMM.

by read1 min views1 publishedSep 16, 2026

arXiv:2609.16722v1 Announce Type: new Abstract: Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly rely on auxiliary models for token reduction but face a fundamental dilemma: lightweight encoder-driven approaches often overlook critical semantic information, whereas heavyweight MLLM-driven reduction negates the efficiency gains. {In this work, we identify a more fundamental inefficiency underlying this dilemma: while fine-grained visual details are essential for detailed understanding, they are largely redundant for the preliminary task of selecting semantically relevant regions. } Motivated by this, we introduce \textbf{VideoMM}, which marks a paradigm shift from model-centric downsizing to adaptive perceptual granularity. Specifically, our framework {decouples selection from reasoning} by executing semantic filtering on a cost-effective \textit{Macro Proxy} (derived from downscaled frames), and projecting the selected regions onto high-fidelity \textit{Micro Tokens} for detailed understanding only when necessary. Extensive evaluations show that VideoMM significantly outperforms existing solutions. It achieves a 6.13$\times$ speedup and a 7.4% accuracy gain over full-context baselines on LongVideoBench, and further accelerates inference by 2.73$\times$ over current leading methods, establishing a highly scalable paradigm for long-video understanding. Our code is available at: https://github.com/adfh917k/VideoMM.

── more in #large-language-models 4 stories · sorted by recency
── more on @videomm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/videomm-adaptive-mac…] indexed:0 read:1min 2026-09-16 ·