cd /news/computer-vision/efficient-unified-multimodal-underst… · home topics computer-vision article
[ARTICLE · art-133348] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=↑ positive

Efficient Unified Multimodal Understanding (EUMU): Winning Solution for the MUMU Track at the 8th LSVOS Challenge

Efficient Unified Multimodal Understanding (EUMU) won the MUMU Track of the 8th LSVOS Challenge with a final score of 17.3409, using a single shared pretrained multimodal model to jointly perform multi-concept image tagging, open-vocabulary object detection, and image captioning. EUMU contains 239.169M parameters, requires 23.947 GFLOPs, and uses 4.5 GB of peak inference memory, with task-aware inference refinement that reuses detection, caption, and tagging outputs as cross-task cues. Code and models are available at https://github.com/Dayoung-Kil/EUMU.

by read1 min views1 publishedSep 18, 2026

arXiv:2609.19451v1 Announce Type: new Abstract: The Mobile Unified Multimodal Understanding (MUMU) Challenge requires a single efficient model to jointly perform multi-concept image tagging, open-vocabulary object detection, and image captioning. We present Efficient Unified Multimodal Understanding (EUMU), the winning solution for the MUMU Track of the 8th LSVOS Challenge. EUMU builds on a shared pretrained multimodal model, using its prompt-based capabilities for detection and captioning and training lightweight heads on shared visual features to predict quality, scene, and event tags. Rather than treating the three tasks independently, EUMU applies task-aware inference refinement by reusing task outputs as cross-task cues. For detection, caption cues help recover objects missed by the initial detection. For captioning, detection cues help refine the caption to better reflect the detected objects. For tagging, image statistics refine quality predictions, while caption and detection cues refine scene and event predictions. This design unifies all three tasks within a single model while satisfying the challenge's resource constraints. EUMU contains 239.169M parameters, requires 23.947 GFLOPs, uses 4.5 GB of peak inference memory, and achieves a final challenge score of 17.3409. Code and models are available at https://github.com/Dayoung-Kil/EUMU.

── more in #computer-vision 4 stories · sorted by recency
── more on @efficient unified multimodal understanding 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/efficient-unified-mu…] indexed:0 read:1min 2026-09-18 ·