cd /news/artificial-intelligence/how-to-build-multimodal-rag-with-goo… · home › topics › artificial-intelligence › article
[ARTICLE · art-147468] src=mindstudio.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

How to Build Multimodal RAG with Google's EmbeddingGemma 2

Google's open-weight EmbeddingGemma 2 model embeds text, code, images, video, and audio into a single shared vector space using a 270 million parameter Gemma-based text backbone plus optional 170M vision and 300M audio encoders, all under a billion parameters total. In hands-on testing on 1,000 Flickr photos, text-to-photo retrieval returned the correct image first about 86% of the time and in the top five 97% of the time, while raw everyday sound effects were retrieved correctly only about 24% of the time versus 86% for photos. The full multimodal RAG pipeline runs in a free Colab notebook on a T4 GPU in roughly 15 minutes, and Matryoshka representation learning lets the 768-dimension output be truncated to 512, 256, or 128 dimensions to shrink the index.

by read9 min views1 publishedOct 8, 2026
How to Build Multimodal RAG with Google's EmbeddingGemma 2
Image: Mindstudio (auto-discovered)

A practical guide to building cross-modal photo, video, audio, and document search with Google's EmbeddingGemma 2 in a free Colab notebook.

What is EmbeddingGemma 2 and why does it matter for RAG? #

EmbeddingGemma 2 is Google’s open-weight embedding model that places text, code, images, video, and audio into one shared vector space, using under a billion parameters total. That matters because most retrieval augmented generation systems today only search text, which means photos, screen recordings, voice memos, and videos stay invisible to search unless someone writes captions or transcripts first. EmbeddingGemma 2 skips that step entirely by embedding the raw image, audio, or video alongside text, letting you search across modalities with a single model small enough to run on a phone.

TL;DR #

  • EmbeddingGemma 2 embeds text, code, images, video, and audio into the same vector space using a shared backbone under a billion parameters.
  • The model uses a 270 million parameter Gemma-based text backbone , plus optional vision (170M) and audio (300M) encoders that only load when you need them.
  • Contrastive training pulls matching pairs (a photo and its caption, a clip and its transcript) together in vector space while pushing unrelated pairs apart, which is how a flamingo photo ends up near the word “flamingo.”
  • Matryoshka representation learning lets you truncate the 768-dimension output down to 512, 256, or 128 dimensions, trading some accuracy for a much smaller index.
  • In hands-on testing on 1,000 Flickr photos, text-to-photo retrieval hit the correct image first about 86% of the time , and landed in the top five 97% of the time.
  • Audio search works far better for speech than for sound effects , with raw everyday sounds (dog barks, rain) returning the correct clip only about 24% of the time versus 86% for photos.
  • The whole pipeline runs in a free Colab notebook on a T4 GPU in roughly 15 minutes , making it accessible to anyone without dedicated hardware.

Remy is new. The platform isn't. #

Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.

How did multimodal retrieval work before this? #

Before single-backbone models like this one, developers had two main workarounds. The first was converting everything to text: run a captioning model on every image, a speech-to-text model on every audio clip, then embed the resulting text. It works, but you end up chaining three separate models together, and your search is only as good as what the captioner or transcriber wrote down. If a captioning model describes a video as “a dog in the park” and never mentions the frisbee the dog is catching, a text search for “frisbee” will never surface that clip. Google’s own benchmarking on audio retrieval found embedding raw audio scored higher (73.99) than embedding transcripts of that audio (70.40), confirming that the conversion-to-text step loses information.

The second approach came from CLIP in 2021, which trained separate encoders for images and text and aligned them so matching pairs land close together. This made text-to-image search genuinely useful, but it only covers one pair of modalities and was trained mostly on short captions. Meta’s ImageBind extended this idea in 2023 to six modalities by aligning everything to images. Document-focused tools like ColPali took a different angle, embedding screenshots of entire pages to skip OCR, though that approach stores many vectors per page and makes indexes grow quickly.

EmbeddingGemma 2 represents the newer pattern: one model, one shared backbone, reading every modality directly. Google’s larger Gemini Embedding API model works the same way, and EmbeddingGemma 2 brings that same architecture down to open weights under a billion parameters.

How does EmbeddingGemma 2 work under the hood? #

An embedding is a list of numbers that captures the meaning of an input, so that inputs with similar meaning get similar numbers. The challenge is that a text model can’t read pixels or sound waves directly, so EmbeddingGemma 2 adds two small encoders in front of the shared backbone. The vision encoder (170 million parameters) converts an image or video frame into tokens. The audio encoder (300 million parameters) does the same for sound. Those tokens then join the text tokens in one sequence, so from the model’s perspective a photo is just a few hundred extra tokens of input.

By default, an image costs 280 tokens, a video frame costs 140 tokens, and a second of audio costs about 25 tokens. With an 8,192-token context window, a single input can hold roughly 29 images, 58 video frames, or around five and a half minutes of audio, meaning you can mix a product description, a photo, and a short video clip into one embedding.

Other agents start typing. Remy starts asking. #

Scoping, trade-offs, edge cases — the real work. Before a line of code.

All of this flows through a shared 270 million parameter text backbone based on Gemma. The model averages every token’s output into a single vector (mean pooling) and projects it into 768 dimensions. Training relies on contrastive learning: the model sees pairs that belong together (a photo with its caption, a clip with its transcript) and learns to pull those pairs closer while pushing everything else in the batch apart. Repeat that across a large, mixed dataset and every modality converges into one shared map.

What are Matryoshka embeddings and why do they matter? #

Matryoshka representation learning means the model is trained so that a truncated slice of its output vector, say the first 256 or 128 numbers out of 768, still functions as a usable embedding on its own. This matters for production because smaller vectors mean smaller, cheaper indexes. Cutting from 768 to 128 dimensions stores roughly six times less data.

The tradeoff shows up clearly in testing. On a 1,000-photo retrieval test, the correct image came back first 86% of the time at full dimensions. Dropping to 256 dimensions barely hurt accuracy (84%) while cutting storage to a third. But dropping further to 128 dimensions caused a bigger drop, down to 75%. Google’s own published numbers show a similar pattern: text retrieval barely changes at 128 dimensions, but multimodal scores drop more sharply (from around 59 to about 46 in one reported benchmark). The lesson is to test truncation on your own data rather than assuming a fixed cutoff works for every modality.

How do you build the Colab notebook pipeline step by step? #

The reference notebook runs on a free T4 GPU in about 15 minutes and uses sentence-transformers with image and video extras installed. Two practical gotchas come up during setup: the video extra doesn’t install torchcodec by default, so video decoding fails unless you add it manually, and the audio extra pulls in an older NumPy version that pip compiles from source on Colab, adding several minutes of install time (audio embedding works fine without that extra).

Precision matters too. Running the model in float16 causes activation overflow, producing NaN or garbage embeddings instead of a clean error. Float32 is the safe default; on newer GPUs, bfloat16 works correctly.

A typical test setup uses a few different datasets to cover each modality: a set of photos with human-written captions, a library of short everyday sound clips, a few minutes of video footage, and a PDF of a research paper rendered as page images. The pipeline is the same regardless of modality: embed everything once to build an index (the slow part, taking several minutes on a T4 for a thousand images), then embed each incoming query and compare it against the index.

Is EmbeddingGemma 2 good at every modality, or does it have weak spots? #

It’s strongest on photo and document search, noticeably weaker on raw sound effects, and multilingual on the text side. Photo search in testing retrieved the correct image first about 86% of the time and placed it in the top five 97% of the time, numbers that hold up well for a sub-billion-parameter model. Because the text backbone is based on Gemma, queries work across more than 100 languages: a German query for “dog running in the snow” and a Spanish query for “woman playing guitar” both returned accurate matches.

One coffee. One working app. #

You bring the idea. Remy manages the project.

Document search worked by embedding full page images with no OCR or text extraction at all, and queries for specific diagrams or tables correctly located the right page in a multi-page PDF. Video search worked similarly well, pulling the correct five-second clip out of several minutes of footage based on a plain text description of the scene, with no captions involved.

Audio is the weaker spot. Voice-based cross-modal search, speaking a query aloud instead of typing it, worked well and returned the same results as the typed version. But retrieval of generic sound effects (a dog barking, rain falling) only found the correct clip first about 24% of the time, compared to 86% for photos. Some sounds, like a toilet flushing, still mapped correctly, but others got confused (a barking dog matched to a crow sound). The audio encoder appears tuned more for speech than for environmental sound effects, so anyone building sound-effect search should test thoroughly on their own clips before relying on it.

Frequently Asked Questions #

What is EmbeddingGemma 2 used for?

It’s used for building search and retrieval systems that work across text, images, video, and audio using one embedding model, rather than chaining together separate captioning, transcription, and embedding models.

Can EmbeddingGemma 2 run on a phone?

Yes. Because the encoders are modular, you can load only the parts you need. The text-only portion is 270 million parameters, and Google has reported quantized RAM usage low enough for on-device use on phones like a Pixel.

Does EmbeddingGemma 2 need transcription or captions to search images and audio?

No. It embeds raw images, video frames, and audio directly into the same vector space as text, which avoids the information loss that happens when you rely on a captioning or transcription model as an intermediate step.

How accurate is multimodal search with EmbeddingGemma 2?

It varies by modality. Photo and document search performed strongly in testing, with the correct result returned in the top five most of the time. Raw sound effect search performed much weaker than speech or photo search.

What hardware do you need to try it?

A free Colab notebook with a T4 GPU is enough to run the full pipeline, including building an index and testing search across photos, audio, video, and documents, in roughly 15 minutes.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-to-build-multimo…] indexed:0 read:9min 2026-10-08 · —