EmbeddingGemma 2: One Model for Text, Code, Images, and Audio Google DeepMind shipped EmbeddingGemma 2 on October 6, a 740M-parameter open-weight model that encodes text, code, images, audio, and video into a single vector space and runs in 567MB of RAM under an Apache 2.0 license. The cross-modal retrieval capability, sub-1GB footprint, and permissive license make the model relevant to developers building RAG pipelines, local search, and agent memory systems. Google DeepMind shipped EmbeddingGemma 2 on October 6 — a 740M parameter open-weight model that encodes text, code, images, audio, and video into the same vector space. The full model runs in 567MB of RAM. That combination of cross-modal retrieval, a sub-1GB footprint, and an Apache 2.0 license is genuinely new. If you are building RAG pipelines, local search, or agent memory systems, this is worth your attention. One Vector Space for Everything The headline capability is cross-modal retrieval: query with text, get back a matching image. Query with an image, retrieve a related audio clip. All from one model, … The post