Google / Wikimedia Commons (Public domain)
The multimodal embedding model handles text, code, images, video and audio, and is built to run on phones and laptops
Google has released EmbeddingGemma 2, an open-weight multimodal embedding model with 740 million parameters. It ships under an Apache 2.0 license, which means developers can use it commercially.
Google DeepMind launched the model on October 6, 2026. It’s designed to run on the hardware already in your pocket or backpack rather than in a distant data center.
What EmbeddingGemma 2 actually does #
An embedding model converts content into a list of numbers, called a vector, so software can measure how related two pieces of content are.
EmbeddingGemma 2 places everything into a shared 768-dimensional vector space. It accepts text (including code), images, video, and audio.
Because all of those formats share a single map, the model can match a text query to an image, or link any one modality to another.
The model’s context window is now 8K tokens. Google says that is four times larger than the window in the original EmbeddingGemma.
The numbers under the hood #
At its core sits a 270M-parameter backbone dedicated to text and code. Optional vision and audio encoders bolt on top, bringing the full package up to 740 million parameters.
The model builds on the Gemma 4 architecture. On Pixel devices, the text-only configuration has an active RAM footprint of roughly 191MB.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
Code performance saw the biggest jump. EmbeddingGemma 2 posted a 9.92-point gain on the MTEB Code benchmark, climbing from 68.76 to 78.68.
The model also maintains multilingual text handling across more than 100 languages.
Nesting dolls and storage bills #
One of the more practical features is Matryoshka Representation Learning, or MRL. With MRL, the most important information is packed into the front of each vector, allowing you to trim the vector down and still keep a smaller, usable version inside.
According to the research findings, MRL can cut embedding dimensions by up to six times. Smaller vectors mean less storage and faster lookups.
Open weights and where to get them #
Google has made the model weights available immediately on Hugging Face and Kaggle. The Apache 2.0 license permits commercial use.
The model works with sentence-transformers, LiteRT/MediaPipe, transformers, MLX, and Ollama. It also supports fine-tuning, so teams can adapt it to their own data.
The original EmbeddingGemma surpassed 20 million downloads.
What this means for developers and device makers #
The clearest beneficiaries are developers building privacy-sensitive apps. When embeddings are generated on the device, personal photos, voice notes, and documents don’t have to be uploaded to a server just to become searchable.
The modular design also lowers the barrier to entry. A note-taking app can run the 270M text backbone alone, while a photo or media app can add the vision and audio encoders only when needed.
Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our