# Google expands EmbeddingGemma beyond text to images, audio and video

> Source: <https://siliconangle.com/2026/10/06/google-expands-embeddinggemma-beyond-text-to-images-audio-and-video/>
> Published: 2026-10-06 22:18:04+00:00

### Google expands EmbeddingGemma beyond text to images, audio and video

Google LLC today [released](https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/) EmbeddingGemma 2, an open multimodal embedding model small enough to run on a smartphone.

The release takes the EmbeddingGemma line beyond text, which was all the first version handled when Google introduced it in  [September 2025](https://developers.googleblog.com/en/introducing-embeddinggemma/). Images, audio and video now share one embedding space with text. An app built on the new model could take a voice memo, for example, and find the matching moment in a video without the data ever leaving the phone.

The response to the original “blew past our expectations,” Google DeepMind research engineers Sahil Dua and Henrique Schechter Vera wrote in the announcement. By their count, developers have downloaded it more than 20 million times.

Built on the Gemma 4 architecture Google released [in April](https://siliconangle.com/2026/04/02/googles-new-gemma-4-models-bring-complex-reasoning-skills-low-power-devices/), the new version is more than twice the size of the original at 740 million parameters. Most of the growth is in the vision and audio encoders, which apps that only work with text can leave off. On its own, the 270 million-parameter text core used about 191 megabytes of memory when Google tested a quantized build on a Google Pixel 11 Pro.

An app also has to store what the model produces. Each embedding is a list of 768 numbers, and every photo, clip or document it indexes adds one more to a local vector database. A training technique called Matryoshka Representation Learning lets developers cut those lists to as few as 128 numbers, which reduces the space they take up by as much as six times. At 256, Google’s [developer guide says](https://developers.googleblog.com/en/embeddinggemma-2-the-developer-guide/), image, video and speech retrieval keep about 95% of their full quality.

Code showed the biggest benchmark gain. EmbeddingGemma 2 scored 78.68 on the code section of the Massive Text Embedding Benchmark, almost 10 points above the first version. Google is pitching the result at developers who build retrieval for coding agents. Multilingual text scores barely moved.

The company also claims leading results among multimodal embedding models under 1 billion parameters, and it said the model outperforms some specialist models more than twice its size on image, video and audio tasks.

Because EmbeddingGemma 2 shares a text tokenizer and an audio encoder with Gemma 4, an on-device retrieval-augmented generation setup running both needs less memory than two unrelated models would. Google’s AI Edge Foresight meeting app for Mac already runs the pair together. In Google’s AI Edge Gallery demo app, a Video Moments Finder feature locates a scene inside a video from a typed or spoken query.

Model weights are available now from Hugging Face Inc. and Google’s Kaggle under an Apache 2.0 license that permits commercial use. Google said the model will reach the Model Garden in Gemini Enterprise Agent Platform soon, and the weights already work with open-source serving tools such as vLLM, llama.cpp and Ollama.

##### Image: Google

# A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. **Join theCUBE’s Alumni Trust Network**, where technology leaders connect, share intelligence and create opportunities.

- **15M+ viewers of theCUBE videos** , powering conversations across AI, cloud, cybersecurity and more
- **11.4k+ theCUBE alumni** — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

### Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: [https://siliconangle.com/aws-marketplace/](https://siliconangle.com/aws-marketplace/)

##### **About SiliconANGLE Media**

[SiliconANGLE](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fsiliconangle.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=SiliconANGLE&index=9&md5=646b1b564e2259100a2b8638aab0a552),

[theCUBE Network](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecube.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Network&index=10&md5=7de2a85f95ab4a4a495cede20b8cb1da),

[theCUBE Research](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fthecuberesearch.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Research&index=11&md5=7bb33676722925eb57d588ec343e4f6f),

[CUBE365](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.cube365.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=CUBE365&index=12&md5=d310fb35919714e66ad8d42c9c0c1bc6),

[theCUBE AI](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecubeai.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+AI&index=13&md5=b8b98472f8071b23ebb10ab9a8dd0683)and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.
