Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings Alibaba's Ovis team introduced Ovis-Embedding, an omni-modal embedding family that encodes text, image, video, and audio through a single shared multimodal backbone rather than separate modality towers. The model is positioned as a state-of-the-art universal embedding system built on native integration across all four modalities. In this report, we introduce Ovis-Embedding, a state-of-the-art omni-modal embedding family built on native integration of text, image, video, and audio. Instead of assembling separate modality towers, Ovis-Embedding uses a shared multimodal backbone to encode different modalities in a common repres