cd /news/artificial-intelligence/how-to-run-embeddinggemma-2-locally-… · home › topics › artificial-intelligence › article
[ARTICLE · art-146887] src=mindstudio.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

How to Run EmbeddingGemma 2 Locally With Sentence Transformers

Google DeepMind released EmbeddingGemma 2, an Apache 2.0 open embedding model on Hugging Face that maps text, code, images, video, and audio into a single shared 768-dimension vector space, with a 270 million parameter text backbone and optional 170M vision and 300M audio encoders for 740M total. A hands-on guide using the sentence-transformers and transformers libraries measured under 2GB (1.7GB) of VRAM for text embedding on a 48GB GPU, and Matryoshka truncation cut the 768-dimension vector to 256 or 128 dimensions for up to 6x storage savings with minimal quality loss above 256. In testing, text-to-image search matched all six queries to their target images, and Google's published benchmark put EmbeddingGemma 2 at around 78 on code search quality, ahead of a 1.5 billion parameter retriever.

by read8 min views1 publishedOct 7, 2026
How to Run EmbeddingGemma 2 Locally With Sentence Transformers
Image: Mindstudio (auto-discovered)

Step-by-step guide to installing Google's EmbeddingGemma 2 locally, with measured VRAM use and text, image, and Matryoshka embedding tests.

What is EmbeddingGemma 2 and why run it locally? #

EmbeddingGemma 2 is an open embedding model from Google DeepMind, released under Apache 2.0 with weights on Hugging Face. It converts text, code, images, video, and audio into a single shared 768-dimension vector space, so a text query can retrieve a matching image or audio clip without separate models for each modality. It’s small enough to run on a single consumer GPU, or even a CPU for text-only use, which makes it practical for local search and retrieval projects without sending data to an external API.

TL;DR #

  • EmbeddingGemma 2 maps text, code, images, video, and audio into one 768-dimension vector space, letting you search across modalities with a single model.
  • The model is modular : the text backbone is 270 million parameters, with optional vision (170M) and audio (300M) encoders loaded only when needed, for a 740M total.
  • Installation only requires the sentence-transformers and transformers libraries , and the model downloads directly through a sentence-transformers call.
  • Measured VRAM use was under 2GB (1.7GB) for text embedding tasks on a 48GB GPU, confirming the model is lightweight for its capability.
  • Matryoshka truncation lets you cut the 768-number vector down to 256 or 128 dimensions, shrinking storage up to 6x with minimal quality loss above 256.
  • In hands-on testing, text-to-image search correctly matched all six queries to their target images, and results held up even after truncating vectors to 256 dimensions.
  • On Google’s published benchmark, EmbeddingGemma 2 scored around 78 on code search quality , ahead of larger models like a 1.5 billion parameter retriever.

Built like a system. Not vibe-coded.

Remy manages the project — every layer architected, not stitched together at the last second.

What makes EmbeddingGemma 2 different from other embedding models? #

Most embedding models are built for one modality: text embedders handle text, image embedders handle images, and so on. EmbeddingGemma 2 routes everything into the same 768-dimension space. Text goes through a tokenizer, images and video frames go through a vision encoder, and audio goes through an audio encoder, but all three outputs land in the same shared text backbone. That shared space is what lets a plain text query find a matching photo, video frame, or audio clip using one model instead of three.

The modularity matters for deployment. If your application only needs text search, you load the 270 million parameter text component and skip the vision and audio encoders entirely. If you need image or audio search too, you add the relevant encoder (170 million for vision, 300 million for audio), bringing the full multimodal model to roughly 740 million parameters total. You’re not forced into carrying the full model weight for a text-only use case.

The model also supports an 8K token context window and Matryoshka Representation Learning, a training technique (named after Russian nesting dolls) that front-loads the most important information into the first numbers of the embedding vector. That means you can truncate a 768-number embedding down to 512, 256, or 128 numbers and discard the rest, shrinking storage requirements without retraining or re-embedding anything.

How do you install and run EmbeddingGemma 2 locally? #

Getting EmbeddingGemma 2 running locally requires two Python libraries: transformers, which handles and running models generally, and sentence-transformers, a framework built on top of it that adds pooling and normalization so you get ready-to-use embeddings through a simple API rather than raw model outputs.

The basic setup flow:

  1. Create a Python environment.
  2. Install the prerequisites: pip install transformers sentence-transformers (an upgraded transformers build with image support may be flagged as needed).
  3. Load the model directly through sentence-transformers, which handles the download automatically.
  4. Pass sentences, images, or other supported inputs through the model to get 768-dimension embeddings back.

No special model server or heavy framework is required. The combination of sentence-transformers’ pooling layer and the model’s small footprint means a basic text embedding script can be running within a few commands.

How much VRAM does EmbeddingGemma 2 actually use? #

In a hands-on test on a system with 48GB of GPU VRAM, running EmbeddingGemma 2 for text embedding tasks consumed just 1.7GB of VRAM, monitored live with nvtop. That’s a small footprint for a model offering multimodal embedding capability, and it means EmbeddingGemma 2 doesn’t require a high-end GPU to run. Because the text-only component is just 270 million parameters, it’s also light enough to run on CPU for text use cases where GPU access isn’t available.

This VRAM number reflects text-only embedding. the vision and audio encoders alongside the text backbone for full multimodal use will increase memory use, though the overall model (740 million parameters across all modalities) remains far smaller than large general-purpose language models.

#

Plans first. Then code.

Remy writes the spec, manages the build, and ships the app.

Does the embedding quality actually hold up in testing? #

A text similarity test comparing three sentence pairs (near-identical meaning, loosely related, and unrelated) produced similarity scores of 0.9558 for the same-meaning pair, 0.8632 for the loosely related pair, and 0.7331 for the unrelated pair. The ranking order was correct: closest meaning scored highest, unrelated scored lowest. The unrelated pair still scored 0.7331 rather than near zero, which Google’s documentation notes is expected behavior for this model. What matters for retrieval tasks is the relative ranking between candidates, not the absolute similarity score.

A cross-modal test pushed this further: six AI-generated images were embedded alongside six plain text queries (using a search-specific prompt format), with no text prefix added to the images. Every query correctly matched its intended image at the top of the ranking, including queries like “something delicious” correctly identifying a curry image without the word “curry” ever appearing in the query, and “a person sleeping in a car” correctly matching an image of a seatbelt. Similarity scores for text-to-image matches ran lower than text-to-text matches (roughly 65 to 80), which is normal when comparing across modalities, since what matters is the gap between the top result and the runner-up.

When the same test was repeated with vectors truncated from 768 to 256 dimensions (a 3x reduction in storage), all six image matches stayed identical. This is the practical payoff of Matryoshka truncation: you can cut storage costs substantially without changing retrieval outcomes, at least on smaller test sets.

Is EmbeddingGemma 2 worth using over larger embedding models? #

On Google’s published code search benchmark, EmbeddingGemma 2 scored around 78, placing it near the top of the quality-versus-size chart and ahead of models several times its parameter count, including a 1.5 billion parameter retriever model. For teams building local retrieval-augmented generation pipelines, semantic search tools, or cross-modal search features, this combination of small size, low VRAM footprint, and competitive benchmark performance makes it a reasonable default rather than reaching for a much larger model.

The tradeoffs to weigh: quality stays strong down to 256-dimension truncation per Google’s own published table, but drops off at 128 dimensions, especially for image, video, and audio embeddings. If your application needs maximum compression and primarily handles text, 128 dimensions might be acceptable. If you’re doing multimodal retrieval, staying at 256 dimensions or higher is the safer choice.

Frequently Asked Questions #

What is Matryoshka truncation in embedding models?

It’s a training technique where the model learns to put the most important information at the front of its embedding vector. This lets you truncate the vector (for example, from 768 numbers down to 256 or 128) and discard the rest while retaining most of the original meaning, reducing storage needs without re-running the model.

Can EmbeddingGemma 2 run on a CPU instead of a GPU?

Yes, at least for text-only embedding. The text component is only 270 million parameters and lightweight enough for CPU use. A GPU helps for speed and is more relevant when the vision or audio encoders for multimodal tasks.

Remy doesn't build the plumbing. It inherits it. #

Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.

Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.

How much VRAM do I need to run EmbeddingGemma 2?

In testing, text embedding tasks used just 1.7GB of VRAM on a 48GB GPU. The model doesn’t require a high-end GPU for text use cases, though full multimodal use (text, image, and audio encoders loaded together) will use more memory.

What libraries do I need to run EmbeddingGemma 2 locally?

The transformers library for and running the model, and sentence-transformers, which adds pooling and normalization on top so you get usable embeddings through a simple API rather than raw model outputs.

Does truncating the embedding vector hurt search accuracy?

In testing, cutting vectors from 768 to 256 dimensions didn’t change any of the top search results in a text-to-image matching test. Google’s own documentation states quality holds up well down to 256 dimensions, but degrades at 128, particularly for image, video, and audio embeddings.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google deepmind 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-to-run-embedding…] indexed:0 read:8min 2026-10-07 · —