{"slug": "fusion-embedding-a-unified-embedding-space-for-text-image-video-and-audio", "title": "Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio", "summary": "Researchers have introduced Fusion Embedding, a family of models that unify text, image, video, and audio into a single embedding space, enabling cross-modal retrieval without paired audio-visual training data. The first generation (fusion-embedding-1) trains a 16.4M-parameter connector between a frozen audio tower and a frozen vision-language base, while the second generation (fusion-embedding-2) adds 44.2M-parameter modality-gated deep adapters that do not alter outputs for non-audio inputs. Both models train in hours on a single GPU and are openly released with weights, code, and evaluation harness.", "body_md": "arXiv:2607.18666v1 Announce Type: new\nAbstract: A single embedding space that covers text, images, video, and audio lets one index serve every query a user can pose. Embedding models built on vision-language backbones now lead text/image/video retrieval benchmarks but lack audio entirely, while audio-text retrieval is led by specialist systems that serve no other modality. We present the Fusion Embedding family, which adds audio to a frozen vision-language embedding base whose parameters are never updated: generation 1 (fusion-embedding-1) trains only a 16.4M-parameter connector between a frozen audio tower and the frozen base, and generation 2 (fusion-embedding-2) adds modality-gated deep adapters (44.2M parameters) whose branch never executes on text, image, or video inputs: their outputs are bit-for-bit those of the released base, verified after every training run. Because the base already binds text, images, and video, aligning audio to text alone makes audio-image retrieval emerge, with zero paired audio-visual training data. Alongside the recipe we map its design space with controlled negative results (rewriting training captions with an LLM, substituting a leaderboard-stronger audio tower, and widening the connector each reduce retrieval) and with training-protocol findings that we expect to transfer to any frozen decoder-LM embedding backbone. Both generations train in hours on a single GPU. Weights, code, and the evaluation harness are openly released.", "url": "https://wpnews.pro/news/fusion-embedding-a-unified-embedding-space-for-text-image-video-and-audio", "canonical_source": "https://arxiv.org/abs/2607.18666", "published_at": "2026-07-22 04:00:00+00:00", "updated_at": "2026-07-22 04:16:22.192606+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "generative-ai", "ai-research", "ai-tools"], "entities": ["Fusion Embedding", "arXiv", "fusion-embedding-1", "fusion-embedding-2"], "alternates": {"html": "https://wpnews.pro/news/fusion-embedding-a-unified-embedding-space-for-text-image-video-and-audio", "markdown": "https://wpnews.pro/news/fusion-embedding-a-unified-embedding-space-for-text-image-video-and-audio.md", "text": "https://wpnews.pro/news/fusion-embedding-a-unified-embedding-space-for-text-image-video-and-audio.txt", "jsonld": "https://wpnews.pro/news/fusion-embedding-a-unified-embedding-space-for-text-image-video-and-audio.jsonld"}}