# EmbeddingGemma 2 Runs 5-Modality Search in 567MB RAM

> Source: <https://byteiota.com/embeddinggemma-2-multimodal-on-device-embeddings/>
> Published: 2026-10-07 00:05:50+00:00

Google DeepMind shipped EmbeddingGemma 2 yesterday. It is a 740-million-parameter open model that maps text, code, images, video, and audio into a single vector space — and does it in 567MB of RAM. Not a cloud endpoint. A model you run locally, on your phone, with no data leaving the device. The first version had 20 million downloads. This one is better in every dimension.

## Five Modalities, One Vector Space

Most embedding models do one thing well. CLIP does vision. text-embedding-3 does text. Stitching them into a unified search index requires maintaining multiple models, multiple indexes, and cross-modal retrieval hacks. EmbeddingGemma 2 skips all of that.

The model maps text, code, images, video, and audio into a shared 768-dimensional space. One query vector — regardless of its source modality — can retrieve results across all five. A text query finds matching video frames. A voice memo retrieves related documents. That is not a prompting trick. That is how the model was trained.

It is built on Gemma 4, with a modular architecture that lets you load only the encoders you need from a single checkpoint:

- Text and code only: 270M parameters (~191MB quantized)
- Text and vision: 440M parameters
- Text and audio: 570M parameters
- Full multimodal (all five): 740M parameters (~567MB quantized)

Every configuration shares the same vector space. You can build a text-only index today and add vision retrieval later without re-embedding your existing data.

## The Benchmarks Are Not Polite About OpenAI

EmbeddingGemma 2 scores 69.67 on MTEB English v2. OpenAI’s text-embedding-3-large scores 64.6. The open-source model outperforms the paid API on the standard benchmark — while being self-hosted, Apache 2.0 licensed, and free to run at scale.

Code retrieval improved 14% over EmbeddingGemma 1, jumping from 68.76 to 78.68 on [MTEB Code](https://huggingface.co/spaces/mteb/leaderboard). For developers who want to index their own codebase and search it with natural language, that is a meaningful number. The context window grew to 8,192 tokens — four times the original — which is enough to embed 5.5 minutes of audio, 29 images, or 58 video frames in a single pass.

The newer modalities are honest about where they stand. Video retrieval scores 50.67 and audio scores 49.39. Those are solid baselines for a first-generation multimodal embedding model at this parameter count. They will improve. The text and code numbers are already production-ready.

## Getting Started in Five Lines

The model ships with [first-class support for sentence-transformers](https://developers.googleblog.com/embeddinggemma-2-the-developer-guide/). Installation pulls the audio, image, and video dependencies automatically:

```
pip install -U "sentence-transformers[image,audio,video]" transformers
python
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("google/embeddinggemma-2")

query_emb = model.encode("database connection timeout", prompt_name="SearchQuery")
doc_emb   = model.encode("ConnectionPool exhausted after 30s retry", prompt_name="Document")

print(model.similarity(query_emb, doc_emb))
```

The `prompt_name` parameter distinguishes queries from documents at index time — the same distinction most production RAG pipelines need anyway.

## The Real Case for This Model: Offline RAG

The combination nobody is talking about yet: EmbeddingGemma 2 for retrieval and Gemma 4 for generation. Both open-weight, both Apache 2.0, both designed to run on consumer hardware. Together they form a complete offline RAG stack. No API keys. No rate limits. No telemetry. No cost per query at scale.

For enterprise use cases — legal documents, medical records, proprietary code — that privacy guarantee matters more than a few percentage points on a benchmark. GDPR enforcement in 2026 has made “data stays on-device” a procurement argument, not just a preference. [Google’s edge deployment guide](https://developers.googleblog.com/google-ai-edge-with-embeddinggemma-2/) frames it this way explicitly.

The 567MB full-model footprint changes the constraint. The old version of the conversation was “you can have privacy or you can have good embeddings.” EmbeddingGemma 2 removes that tradeoff.

## Where to Get It

The weights are live on [Hugging Face](https://huggingface.co/google/embeddinggemma-2) and Kaggle today. Unsloth’s GGUF quantized version is already available for CPU-only inference. Google’s Model Garden integration is listed as “coming soon,” which likely means Vertex AI and Android AI tooling will follow. The [official announcement](https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/) covers the full roadmap.

If you are still paying per-token for embeddings on text that never leaves your company, this is worth evaluating this week.
