How to Fine-Tune EmbeddingGemma on Your Own Data With LoRA Fine-tuning EmbeddingGemma with LoRA adapters can raise retrieval accuracy on custom data in minutes on a single GPU, according to a practical guide from Sentence Transformers, which reports accuracy on unseen everyday-sounds clips rising from about 25% to roughly two in three after a few minutes of training. The guide says LoRA adjusts under 2% of total parameters and relies on multiple negative ranking loss, where every other item in a batch serves as a free negative example, though duplicate passages in a batch can punish correct answers unless a no-duplicate sampler is used. It warns that because EmbeddingGemma shares one backbone across text, image, audio, and video, fine-tuning on one modality can cost a few points of general text search accuracy, and that reloading an audio-tuned adapter can silently break if the processor config is swapped out. How to Fine-Tune EmbeddingGemma on Your Own Data With LoRA A practical guide to fine-tuning EmbeddingGemma with LoRA for custom retrieval, covering training pairs, negative sampling, and pitfalls. What does it mean to fine-tune an embedding model? Fine-tuning an embedding model means adjusting it so it places your own questions and documents closer together in vector space, without rewriting the model’s entire set of weights. Unlike fine-tuning a large language model, there’s no exact output to match. You only tell the model which pairs of inputs a question and the passage that answers it, for example should end up near each other. Using LoRA adapters, this can be done in minutes on a single GPU, often improving retrieval accuracy on your own data noticeably while leaving most of the model untouched. TL;DR - EmbeddingGemma can be adapted to your own documents, product names, or sounds using LoRA adapters that train in minutes rather than hours. - Training data for embedding models is just matched pairs a question and its answer passage, or a label and an audio clip , with no need to hand-label wrong answers. - Sentence Transformers’ multiple negative ranking loss turns every other item in a training batch into a free negative example, but duplicate passages in the same batch can accidentally punish correct answers unless you use a no-duplicate sampler. - In one test on an everyday-sounds dataset, accuracy on unseen clips went from about 25% to roughly two in three after a few minutes of LoRA training. - The fine-tuned model learned the specific labels it was trained on , not a general improvement in hearing or reading, so anything you care about must appear in training data. - Because EmbeddingGemma shares one backbone across text, image, audio, and video, fine-tuning on one modality can quietly shift performance on the others, in one case costing a few points of general text search accuracy . - Saving and reloading an audio-tuned adapter can silently break things if the processor config gets swapped out, so retesting after reload is worth doing. Plans first. Then code. Remy writes the spec, manages the build, and ships the app. Why fine-tune an embedding model instead of using it off the shelf? A pretrained embedding model like EmbeddingGemma knows general language and, in its multimodal form, general sounds and images. It has no built-in knowledge of your product names, internal documentation, or the specific audio events your application cares about. Off the shelf, it will embed those things using whatever general patterns it learned during pretraining, which often works reasonably well but misses the specifics that matter to your use case. Fine-tuning closes that gap by nudging the model’s embedding space so that your own queries land near your own correct answers. The appeal of doing this with LoRA is that it’s cheap. Instead of updating every weight in the network, LoRA inserts small trainable matrices alongside existing layers and freezes the original weights. In practice, this means adjusting a small fraction of the total parameters under 2% in one demonstrated run while still getting a meaningful boost in retrieval accuracy on unseen examples. How does LoRA fine-tuning for embeddings actually work? The core idea is a contrastive loss function. During training, the model processes a batch of paired examples, say 32 question-passage pairs. For each question, its own passage is the correct match, and every other passage in that batch becomes an automatic negative example. The loss function pulls each question toward its correct passage and pushes it away from all the others in the batch. This is the mechanism Sentence Transformers calls multiple negative ranking loss, the same general approach CLIP used to align images and text in one space. The practical benefit is that you never have to manually collect or label wrong answers. Every batch generates its own negatives automatically, just by containing multiple pairs. There’s a catch worth knowing before you start. If two different questions share the same correct passage, and both pairs land in the same batch, the loss will treat the second occurrence of that passage as wrong for the first question, even though it’s actually correct. The practical fix is a batch sampler that guarantees no duplicate text appears twice in a single batch what Sentence Transformers calls a “no duplicate” sampler . Skipping this step means your model quietly learns incorrect signals whenever duplicate passages happen to co-occur. Prompt formatting matters too. EmbeddingGemma expects a specific prompt structure, a short task prefix in front of search queries, and a title-and-text format for documents. Whatever format is used during training has to match exactly what’s used at inference time, the same way prompt templates need to stay consistent for large language models. What does the training data look like in practice? Unlike LLM fine-tuning, there’s no “correct output” to write down. You only need pairs of things that belong together: - A customer support question paired with the help-center passage that answers it. - A short audio clip paired with a sentence describing the sound turning a classification task into a search task . - A viewer-style question paired with the transcript passage it was generated from. Remy is new. The platform isn't. Remy is the latest expression of years of platform work. Not a hastily wrapped LLM. In one demonstrated example, a dataset of everyday sound clips 50 categories, including things like dog barks, rain, and chainsaws was converted into pairs by turning each category label into a sentence, like “the sound of a chainsaw.” Classification effectively became retrieval: the model embeds the clip, embeds the label sentence, and the correct label should come back as the closest match. For a text-based example, transcripts from a batch of videos were cut into short passages. Since no questions existed for those passages, Gemini was used to generate one viewer-style question per passage, producing the question-passage pairs needed for training. How much does fine-tuning actually improve retrieval accuracy? In the audio example, the untrained model picked the correct sound label first only about one time in four, roughly 25% accuracy. Notably, using the model’s built-in classification prompt performed worse than just treating the label as a search query, meaning prompt engineering alone wasn’t enough to fix the gap. After roughly 7.5 minutes of LoRA training on an A100 GPU, accuracy on sounds the model had never heard during training rose to about two in three correct. Some categories, like the sound of brushing teeth, went from frequently wrong to nearly always correct. For the text example, built from 120 videos’ worth of transcripts, the correct passage came back first about two times in three before training. After less than 3 minutes of training, that rose to roughly three in four. Does the model generalize or just memorize labels? This is the critical pitfall to understand before fine-tuning for a real application. A follow-up test hid 10 sound categories entirely from training and then tested the model on those unseen categories. Performance on the trained categories improved as expected, but performance on the 10 hidden categories did not improve at all. The conclusion: the model learns the specific labels and examples it’s given, not a general improvement in “hearing” or “understanding.” For anyone fine-tuning their own application, this means every label, category, or type of question that matters needs to be explicitly represented in the training data. The model won’t extrapolate to categories it has never seen, even if they’re conceptually similar to ones it has. What does fine-tuning cost elsewhere in the model? Because EmbeddingGemma routes text, images, video, and audio through a shared backbone, training on one modality can affect the others. After fine-tuning on audio, three checks were run: photo search, voice search, and general text search. Photo and voice search held steady, but general text search accuracy dropped by about four points. The same roughly four-point drop in general text search showed up again after the separate text fine-tuning run. This means fine-tuning isn’t free even when it only targets one modality. If an application relies on multiple types of search, each fine-tuning pass should be followed by a check across all the modalities the application actually uses, not just the one being improved. - ✕a coding agent - ✕no-code - ✕vibe coding - ✕a faster Cursor The one that tells the coding agents what to build. A second pitfall involves saving and reloading adapters. After reloading a fine-tuned audio adapter, voice search accuracy initially collapsed to almost zero, looking like the model had completely forgotten speech. The adapter weights themselves were fine. The actual bug was in the processor saved alongside the adapter, which was truncating audio clips down to a few milliseconds, making every sound effectively indistinguishable. The fix was reusing the processor from the base model instead of the one bundled with the adapter. Anyone saving and reloading an audio-tuned adapter should retest after reload rather than assuming it behaves identically to the in-memory version. Frequently Asked Questions What is LoRA and why use it for embedding models? LoRA low-rank adaptation adds small trainable matrices alongside a model’s existing layers while keeping the original weights frozen. For embedding models, this means a fine-tuning run can touch a small fraction of total parameters, train in minutes rather than hours, and produce a lightweight adapter file instead of a full copy of the model. What format does training data need to be in? Training data for embedding models is pairs: a question paired with the passage that answers it, or a label paired with the audio/image/text it describes. No explicit negative examples or wrong-answer labels are needed, since other items in the same training batch automatically serve as negatives. Why do duplicate passages cause problems during training? If the same passage is the correct answer for two different questions and both pairs land in the same training batch, the loss function will treat the repeated passage as a wrong answer for one of the questions. This actively trains the model to treat a correct answer as incorrect, which is why a sampler that avoids duplicate text within a batch is important. Does fine-tuning make the model generally better, or just better at specific labels? Based on testing with hidden categories, fine-tuning improves performance specifically on the labels, questions, or categories included in training. It does not generalize to unseen categories, even conceptually similar ones. Every label or type of query that matters for an application needs to be represented in the training set. Can fine-tuning for one modality hurt performance on others? Yes. Since EmbeddingGemma shares one backbone across text, image, audio, and video, fine-tuning on audio data, for instance, can reduce accuracy on general text search even though the text-handling layers weren’t the explicit training target. Testing all relevant modalities after fine-tuning is the only reliable way to catch this.