# EmbeddingGemma 2: One Model for Text, Code, Images, and Audio

> Source: <https://byteiota.com/embeddinggemma-2-multimodal-on-device-rag/>
> Published: 2026-10-08 12:09:12+00:00

Google DeepMind shipped EmbeddingGemma 2 on October 6 — a 740M parameter open-weight model that encodes text, code, images, audio, and video into the same vector space. The full model runs in 567MB of RAM. That combination of cross-modal retrieval, a sub-1GB footprint, and an Apache 2.0 license is genuinely new. If you are building RAG pipelines, local search, or agent memory systems, this is worth your attention. One Vector Space for Everything The headline capability is cross-modal retrieval: query with text, get back a matching image. Query with an image, retrieve a related audio clip. All from one model, […]

The post
