Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
Researchers introduce Mage-VL, an efficient codec-native streaming multimodal foundation model that reduces visual token consumption by over 75% while matching or outperforming flagship encoders. The model achieves up to…