{"slug": "embeddinggemma-2-the-developer-guide", "title": "EmbeddingGemma 2: The Developer Guide", "summary": "Google released EmbeddingGemma 2, an Apache 2.0-licensed open multimodal embedding model built on Gemma 4 that maps text, code, images, video, and audio into a unified 768-dimensional space, scaling from 270M parameters for text and code up to 740M parameters for all modalities. The model runs through the sentence-transformers library (v6.1.0 or later), and developers can omit unused modality encoders at load time by setting vision_config or audio_config to None, cutting memory to 270M parameters for text-only, 440M for text and images, and 570M for text and audio. EmbeddingGemma 2 uses short task instructions via prompt_name in encode(), such as SearchQuery and Document, to steer representations for retrieval and cross-modal search.", "body_md": "Modern search and retrieval augmented generation (RAG) applications increasingly need to work across diverse content types, from technical documentation and source code to images, video clips, and audio recordings. The challenge is finding models that deliver strong retrieval accuracy while maintaining low latency across all these formats without requiring massive compute infrastructure to run and index.\n\n[**EmbeddingGemma 2**](https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2) is designed to provide a single, compact open model released under the Apache 2.0 license, delivering exceptional multimodal performance for its size. Based on Gemma 4, this sub-1B model maps **text, code, images, video, and audio** into a unified 768-dimensional space. Its modular architecture lets you load only what you need, scaling from 270M parameters for text and code up to 740M parameters for all modalities.\n\nKey capabilities include:\n\nEmbeddingGemma 2 replaces chained models with modular encoders that project into a shared 768-dimensional space:\n\nEven though each modality is processed by a specialized encoder, all inputs are processed through the shared backbone and their resulting embeddings occupy the same dimensional space:\n\nYou can run EmbeddingGemma 2 across text, code, images, video, and audio with the [sentence-transformers](https://github.com/huggingface/sentence-transformers) library (v6.1.0 or later):\n\n```\npip install -U sentence-transformers[image,audio,video] transformers\n```\n\nThe full model embeds text, code, images, video, and audio:\n\n``` python\nfrom sentence_transformers import SentenceTransformer\n\n# Full model: all modalities (740M parameters)\nMODEL_ID = \"google/embeddinggemma-2\"\nmodel = SentenceTransformer(MODEL_ID)\n```\n\nTo minimize memory usage, you can omit unused modality encoders at load time by setting `vision_config` or `audio_config` to `None` in `config_kwargs`. Disabled encoders are never loaded into memory, so the savings apply to both the weights and peak allocation:\n\n```\n# Text only (270M parameters)\ntext_only_model = SentenceTransformer(\n    MODEL_ID,\n    config_kwargs={\"vision_config\": None, \"audio_config\": None},\n)\n\n# Text, images, and video (440M parameters)\ntext_image_model = SentenceTransformer(\n    MODEL_ID,\n    config_kwargs={\"audio_config\": None},\n)\n\n# Text and audio (570M parameters)\ntext_audio_model = SentenceTransformer(\n    MODEL_ID,\n    config_kwargs={\"vision_config\": None},\n)\n```\n\nEmbeddingGemma 2 is trained with short task instructions to steer representations for specific tasks. Set prompt_name in encode() to add it for you:\n\nFor retrieval, encode queries and documents with different prompts:\n\n```\nquery = \"What causes the northern lights?\"\ndocument = \"The northern lights are caused by charged particles from the sun..\" # truncated\n\n# Embed using `prompt_name`\nquery_emb = model.encode(query, prompt_name=\"SearchQuery\")\ndoc_emb = model.encode(document, prompt_name=\"Document\")\n\nprint(model.similarity(query_emb, doc_emb))\n```\n\nPass media as a dictionary keyed by modality without a prompt. To embed text and media together, mark where each item goes with `<|image|>`, `<|video|>`, or `<|audio|>`:\n\n```\n# Cross-modal search: one text query against a photo and a sound recording\nimage_emb = model.encode({\"image\": \"sunset_beach.jpg\"})\naudio_emb = model.encode({\"audio\": \"ocean_waves.wav\"})\nquery_emb = model.encode(\"ocean waves at sunset\", prompt_name=\"SearchQuery\")\n\nprint(model.similarity(query_emb, image_emb))\nprint(model.similarity(query_emb, audio_emb))\n\n# Interleaved: one embedding for a product listing with text, photo, and video\nlisting_emb = model.encode({\n    \"text\": \"Waterproof trail shoe. <|image|> Grip test on wet rock: <|video|>\",\n    \"image\": \"trail_shoe.jpg\",\n    \"video\": \"grip_test.mp4\",\n})\nquery_emb = model.encode(\"waterproof trail shoes\", prompt_name=\"SearchQuery\")\n\nprint(model.similarity(query_emb, listing_emb))\n```\n\nDespite coming from different modalities, the embeddings generated by EmbeddingGemma 2 occupy the same dimensional space and can be compared on their semantic similarity.\n\nPass `truncate_dim` (`512`, `256`, or `128`) with `normalize_embeddings=True` to get shorter, unit-length vectors. Queries and documents must use the same dimension:\n\n```\n# Truncate the query\nquery_emb = model.encode(\n    query,\n    prompt_name=\"SearchQuery\",\n    truncate_dim=256,\n    normalize_embeddings=True,\n)\n```\n\nTo use one dimension for every call, set it at load time instead: SentenceTransformer(MODEL_ID, truncate_dim=256). For instance, in bfloat16 precision, storing a million 768-dimensional vectors takes roughly 1.5 GB of memory, while truncating them to 128 dimensions requires just 250 MB. That 6x reduction allows you to store six times as many embeddings in the same memory budget, making it much easier to fit large indexes in memory or on-device.\n\nIn sentence-transformers, all four encoder setups load from the same checkpoint, so they share one vector space: a query embedded with the 270M text-only setup can be matched directly against documents embedded with the full model.\n\nAs a general guideline, load only the encoders your data needs, and truncate dimensions only when storage or search speed require it.\n\nIf you start with a text-only index and later add image or audio embeddings, simply reload the model with the additional encoder enabled. Embeddings you have already computed do not need to be re-computed.\n\nAll modalities share the 8,192-token context window, at fixed rates:\n\nThe maximums assume a single modality with no text. Pass media as file paths (MP4 for video), URLs (images and audio), or in-memory PIL images, arrays, and tensors. Video is sampled at 1 frame per second by default, and audio should be 16 kHz mono.\n\nEmbeddingGemma 2 scores 14% higher than EmbeddingGemma 1 on MTEB (Code), adds image, video, and audio retrieval while retaining the accuracy on multilingual text of EmbeddingGemma 1\n\nReady to explore multimodal embeddings? Take a look at the following resources to find out more:", "url": "https://wpnews.pro/news/embeddinggemma-2-the-developer-guide", "canonical_source": "https://developers.googleblog.com/embeddinggemma-2-the-developer-guide/", "published_at": "2026-10-06 16:17:10.868479+00:00", "updated_at": "2026-10-06 16:17:13.519554+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-products", "ai-tools", "developer-tools"], "entities": ["Google", "EmbeddingGemma 2", "Gemma 4", "sentence-transformers", "Hugging Face"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/embeddinggemma-2-the-developer-guide", "markdown": "https://wpnews.pro/news/embeddinggemma-2-the-developer-guide.md", "text": "https://wpnews.pro/news/embeddinggemma-2-the-developer-guide.txt", "jsonld": "https://wpnews.pro/news/embeddinggemma-2-the-developer-guide.jsonld"}}