What if an AI search application could retrieve information from internal documentation, source code, screenshots, and recordings without sending those files to an external embedding API?
On October 6, 2026, Google DeepMind introduced EmbeddingGemma 2, an open-weight multimodal embedding model released under the Apache 2.0 license. It brings text, code, images, video, and audio into a shared 768-dimensional vector space.
For developers building Retrieval-Augmented Generation (RAG) systems, the most interesting question is architectural: which parts of retrieval can run directly on the user's device?
A typical RAG pipeline prepares documents, creates embeddings, retrieves relevant passages, and optionally sends those passages to a generative language model. Depending on its design, the pipeline may transmit content to third-party services.
Local retrieval offers a different trade-off:
This does not make cloud APIs obsolete. For public data or centrally managed services, a remote API can still be the right engineering choice.
The model uses modular encoders that share a compatible embedding space:
| Configuration | Parameters | Typical content |
|---|---|---|
| Text and code | 270M | Documentation, codebases, knowledge bases |
| Text + vision | 440M | Images, diagrams, video frames |
| Text + audio | 570M | Recordings and speech |
| Full multimodal | 740M | Text, code, images, video, audio |
The shared representation makes it possible to compare a text query with an image or audio segment without requiring a separate caption or transcript in every workflow.
Important: EmbeddingGemma 2 is an embedding model, not a generative LLM. It helps find relevant content. If you want a written answer, you need a separate generation step. For fully on-device RAG, that generative model must also run locally.
A service provider computes the vectors; the application stores and queries an index locally or on a server. This can simplify deployment, but introduces considerations around network access, data handling, and recurring costs.
A compact text encoder and retrieval index live on the user's device. This is a practical starting point for notes, internal documentation, and code search. Document parsing, chunking, and access control still matter.
Additional encoders make screenshots, diagrams, sound, and video searchable within a common representation. This can improve retrieval when essential information is not purely textual, at the cost of more processing and memory.
Google documents EmbeddingGemma 2 with sentence-transformers 6.1.0 or newer. Install the dependencies in your Python environment:
pip install -U "sentence-transformers[image,audio,video]" transformers
For a text-only prototype, omit the encoders you do not need:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={
"vision_config": None,
"audio_config": None,
},
)
documents = [
"The application enforces document access permissions.",
"The meeting notes describe a local retrieval architecture.",
]
question = "Where can I find information about offline search?"
document_vectors = model.encode(
documents,
prompt_name="Document",
normalize_embeddings=True,
)
query_vector = model.encode(
question,
prompt_name="SearchQuery",
normalize_embeddings=True,
)
scores = model.similarity(query_vector, document_vectors)[0]
best_match = int(scores.argmax())
print(documents[best_match])
This returns a retrieved passage, not an AI-generated answer. The first run requires down model weights. Subsequent encoding can work offline when all necessary resources have been cached.
A production search system should also store stable document identifiers and source locations, handle updates incrementally, and enforce document-level permissions.
With the full model, media can be embedded by passing a dictionary keyed by modality:
model = SentenceTransformer("google/embeddinggemma-2")
image_vector = model.encode({"image": "architecture.png"})
audio_vector = model.encode({"audio": "meeting.wav"})
query_vector = model.encode(
"discussion about application architecture",
prompt_name="SearchQuery",
)
print("Image:", model.similarity(query_vector, image_vector))
print("Audio:", model.similarity(query_vector, audio_vector))
The sample files must exist and be compatible with the runtime. For longer recordings, index meaningful segments with timestamps: a vector for an entire video cannot reliably locate a particular moment. PDFs with diagrams may require image-based retrieval alongside extracted text and, for scanned pages, OCR.
EmbeddingGemma 2 supports Matryoshka Representation Learning, so embeddings can be shortened from 768 dimensions to 512, 256, or 128. For example, one million raw float32 vectors require approximately:
These figures exclude metadata and index overhead. Queries and indexed content need matching dimensions, and smaller vectors should be evaluated for retrieval quality on your actual dataset.
Google has also published memory figures for optimized configurations on specific devices. Those are not universal memory guarantees for a Python application or every browser.
Not automatically. A privacy-oriented deployment also needs to answer:
Local embeddings reduce certain data transfers; they do not replace a complete security design.
A useful evaluation can compare remote embeddings, local text retrieval, and local multimodal retrieval using the same authorized corpus, questions, and target hardware.
This is a proposed evaluation method, not a report of independent benchmarks or a description of a deployed product.
EmbeddingGemma 2 provides an intriguing building block for multimodal search over sensitive, heterogeneous content on local devices. It does not remove the need for good segmentation, source metadata, evaluation, permissions, or a separate generative model.
The key engineering decision is where each stage of the retrieval pipeline should run, given the application's performance, privacy, and usability constraints.
For a more detailed discussion of the architecture, examples, and practical limitations, see the original English article on SDX Development.
Discussion: Which stage would you move on-device first: embeddings, retrieval, or generation?