{"slug": "mage-vl-an-efficient-codec-native-streaming-multimodal-foundation-model", "title": "Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model", "summary": "Researchers introduce Mage-VL, an efficient codec-native streaming multimodal foundation model that reduces visual token consumption by over 75% while matching or outperforming flagship encoders. The model achieves up to a 3.5x wall-clock inference speedup and surpasses the 15B Phi-4-reasoning-vision baseline on static and video tasks.", "body_md": "arXiv:2607.24904v1 Announce Type: new\nAbstract: Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.", "url": "https://wpnews.pro/news/mage-vl-an-efficient-codec-native-streaming-multimodal-foundation-model", "canonical_source": "https://arxiv.org/abs/2607.24904", "published_at": "2026-07-29 04:00:00+00:00", "updated_at": "2026-07-29 04:21:48.838090+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "computer-vision", "ai-research", "ai-infrastructure"], "entities": ["Mage-VL", "Mage-ViT", "Qwen3-VL-4B", "Phi-4-reasoning-vision"], "alternates": {"html": "https://wpnews.pro/news/mage-vl-an-efficient-codec-native-streaming-multimodal-foundation-model", "markdown": "https://wpnews.pro/news/mage-vl-an-efficient-codec-native-streaming-multimodal-foundation-model.md", "text": "https://wpnews.pro/news/mage-vl-an-efficient-codec-native-streaming-multimodal-foundation-model.txt", "jsonld": "https://wpnews.pro/news/mage-vl-an-efficient-codec-native-streaming-multimodal-foundation-model.jsonld"}}