EmbeddingGemma 2 Runs 5-Modality Search in 567MB RAM Google DeepMind released EmbeddingGemma 2, a 740-million-parameter open embedding model that maps text, code, images, video, and audio into a single 768-dimensional vector space and runs in 567MB of RAM quantized on-device. The model scores 69.67 on MTEB English v2 versus OpenAI text-embedding-3-large's 64.6, improves code retrieval 14% over EmbeddingGemma 1 (68.76 to 78.68 on MTEB Code), and expands its context window to 8,192 tokens, with video retrieval at 50.67 and audio at 49.39. It is Apache 2.0 licensed, built on Gemma 4 with modular encoders ranging from 270M parameters (text and code only, ~191MB quantized) to the full 740M multimodal configuration, and ships with sentence-transformers support. Google DeepMind shipped EmbeddingGemma 2 yesterday. It is a 740-million-parameter open model that maps text, code, images, video, and audio into a single vector space — and does it in 567MB of RAM. Not a cloud endpoint. A model you run locally, on your phone, with no data leaving the device. The first version had 20 million downloads. This one is better in every dimension. Five Modalities, One Vector Space Most embedding models do one thing well. CLIP does vision. text-embedding-3 does text. Stitching them into a unified search index requires maintaining multiple models, multiple indexes, and cross-modal retrieval hacks. EmbeddingGemma 2 skips all of that. The model maps text, code, images, video, and audio into a shared 768-dimensional space. One query vector — regardless of its source modality — can retrieve results across all five. A text query finds matching video frames. A voice memo retrieves related documents. That is not a prompting trick. That is how the model was trained. It is built on Gemma 4, with a modular architecture that lets you load only the encoders you need from a single checkpoint: - Text and code only: 270M parameters ~191MB quantized - Text and vision: 440M parameters - Text and audio: 570M parameters - Full multimodal all five : 740M parameters ~567MB quantized Every configuration shares the same vector space. You can build a text-only index today and add vision retrieval later without re-embedding your existing data. The Benchmarks Are Not Polite About OpenAI EmbeddingGemma 2 scores 69.67 on MTEB English v2. OpenAI’s text-embedding-3-large scores 64.6. The open-source model outperforms the paid API on the standard benchmark — while being self-hosted, Apache 2.0 licensed, and free to run at scale. Code retrieval improved 14% over EmbeddingGemma 1, jumping from 68.76 to 78.68 on MTEB Code https://huggingface.co/spaces/mteb/leaderboard . For developers who want to index their own codebase and search it with natural language, that is a meaningful number. The context window grew to 8,192 tokens — four times the original — which is enough to embed 5.5 minutes of audio, 29 images, or 58 video frames in a single pass. The newer modalities are honest about where they stand. Video retrieval scores 50.67 and audio scores 49.39. Those are solid baselines for a first-generation multimodal embedding model at this parameter count. They will improve. The text and code numbers are already production-ready. Getting Started in Five Lines The model ships with first-class support for sentence-transformers https://developers.googleblog.com/embeddinggemma-2-the-developer-guide/ . Installation pulls the audio, image, and video dependencies automatically: pip install -U "sentence-transformers image,audio,video " transformers python from sentence transformers import SentenceTransformer model = SentenceTransformer "google/embeddinggemma-2" query emb = model.encode "database connection timeout", prompt name="SearchQuery" doc emb = model.encode "ConnectionPool exhausted after 30s retry", prompt name="Document" print model.similarity query emb, doc emb The prompt name parameter distinguishes queries from documents at index time — the same distinction most production RAG pipelines need anyway. The Real Case for This Model: Offline RAG The combination nobody is talking about yet: EmbeddingGemma 2 for retrieval and Gemma 4 for generation. Both open-weight, both Apache 2.0, both designed to run on consumer hardware. Together they form a complete offline RAG stack. No API keys. No rate limits. No telemetry. No cost per query at scale. For enterprise use cases — legal documents, medical records, proprietary code — that privacy guarantee matters more than a few percentage points on a benchmark. GDPR enforcement in 2026 has made “data stays on-device” a procurement argument, not just a preference. Google’s edge deployment guide https://developers.googleblog.com/google-ai-edge-with-embeddinggemma-2/ frames it this way explicitly. The 567MB full-model footprint changes the constraint. The old version of the conversation was “you can have privacy or you can have good embeddings.” EmbeddingGemma 2 removes that tradeoff. Where to Get It The weights are live on Hugging Face https://huggingface.co/google/embeddinggemma-2 and Kaggle today. Unsloth’s GGUF quantized version is already available for CPU-only inference. Google’s Model Garden integration is listed as “coming soon,” which likely means Vertex AI and Android AI tooling will follow. The official announcement https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/ covers the full roadmap. If you are still paying per-token for embeddings on text that never leaves your company, this is worth evaluating this week.