cd /news/artificial-intelligence/fusion-embedding-a-unified-embedding… · home topics artificial-intelligence article
[ARTICLE · art-68014] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio

Researchers have introduced Fusion Embedding, a family of models that unify text, image, video, and audio into a single embedding space, enabling cross-modal retrieval without paired audio-visual training data. The first generation (fusion-embedding-1) trains a 16.4M-parameter connector between a frozen audio tower and a frozen vision-language base, while the second generation (fusion-embedding-2) adds 44.2M-parameter modality-gated deep adapters that do not alter outputs for non-audio inputs. Both models train in hours on a single GPU and are openly released with weights, code, and evaluation harness.

read1 min views1 publishedJul 22, 2026

arXiv:2607.18666v1 Announce Type: new Abstract: A single embedding space that covers text, images, video, and audio lets one index serve every query a user can pose. Embedding models built on vision-language backbones now lead text/image/video retrieval benchmarks but lack audio entirely, while audio-text retrieval is led by specialist systems that serve no other modality. We present the Fusion Embedding family, which adds audio to a frozen vision-language embedding base whose parameters are never updated: generation 1 (fusion-embedding-1) trains only a 16.4M-parameter connector between a frozen audio tower and the frozen base, and generation 2 (fusion-embedding-2) adds modality-gated deep adapters (44.2M parameters) whose branch never executes on text, image, or video inputs: their outputs are bit-for-bit those of the released base, verified after every training run. Because the base already binds text, images, and video, aligning audio to text alone makes audio-image retrieval emerge, with zero paired audio-visual training data. Alongside the recipe we map its design space with controlled negative results (rewriting training captions with an LLM, substituting a leaderboard-stronger audio tower, and widening the connector each reduce retrieval) and with training-protocol findings that we expect to transfer to any frozen decoder-LM embedding backbone. Both generations train in hours on a single GPU. Weights, code, and the evaluation harness are openly released.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @fusion embedding 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/fusion-embedding-a-u…] indexed:0 read:1min 2026-07-22 ·