{"slug": "embedding-gemma2-use-cases", "title": "Embedding Gemma2 Use Cases", "summary": "Med Karim Bchini ported four jeffhub.ai use cases — ground, support-intents, spam, and triage — onto Google's EmbeddingGemma-2 embedding model, replacing jeffhub's 0.8B \"System 1\" decider with cosine-similarity scoring over labelled exemplars. The project runs EmbeddingGemma-2 at int8 (onnx-community/embeddinggemma-2-ONNX, ~314 MB) via Node.js and @huggingface/transformers v4 on CPU with no Python or GPU, measuring p50 embedding latency of 168 ms and p95 of 177 ms at batch size 1, plus 24 docs/s batched throughput (128 docs in 5.36 s, 41.8 ms/doc).", "body_md": "**Author:** Med Karim Bchini @karimtn\n\nPorts of **[jeffhub.ai](https://jeffhub.ai)** use cases onto Google's\n**EmbeddingGemma-2** embedding model, with a runnable demo and a\n**latency + quality** performance benchmark.\n\njeffhub ships a 0.8B \"System 1\" decider model plus one small adapter per task;\neach adapter takes an input and a list of options and returns a probability for\nevery option. This project keeps that request shape (\"pick one of your options,\nwith a probability\") but swaps the 0.8B decider for **EmbeddingGemma-2**\nembeddings: each option is scored by cosine similarity to a few labelled\nexemplars (classification) or by dense passage ranking (retrieval), and a\nsoftmax turns the scores into probabilities.\n\n| jeffhub adapter | Category | Ported here as | Quality metric | \n|---|---|---|---|\n| `ground` | Retrieval | Dense passage retrieval / re-ranking | Recall@1, Recall@5, MRR | \n| `support-intents` | Support | Nearest-exemplar intent classification | accuracy, macro-F1 | \n| `spam` | Safety | Legitimate / spam / phishing classification | accuracy, macro-F1 | \n| `triage` | Support | Ticket routing to a team | accuracy, macro-F1 | \n\nThe runtime here has **no Python, no GPU**, and a small disk budget, so the\nexamples use:\n\n- **Node.js + [`@huggingface/transformers`](https://huggingface.co/docs/transformers.js) v4** (ONNX Runtime, CPU).\n- **`onnx-community/embeddinggemma-2-ONNX`** at**int8** (` dtype: \"q8\"` →`model_quantized.onnx` , ~314 MB).\n\nEmbeddingGemma-2 is a 740M-parameter model (270M text backbone) that maps text\ninto a unified 768-d space. It is *task-steered*: a short instruction prefix\ntells it what representation to produce. The prefixes used here come straight\nfrom the model card:\n\n| Purpose | Prefix | \n|---|---|\n| Retrieval query | `task: search result \\| query:`  | \n| Document | `title: none \\| text:`  | \n| Classification | `task: classification \\| query:`  | \n\n```\nnpm install          # installs transformers.js + onnxruntime-node\nnpm run demo         # runs every use case on sample inputs\nnpm run bench        # latency + quality benchmark (writes bench/results.json)\nnpm run bench:quick  # smaller timing loops\n```\n\nThe first run downloads the ONNX weights (~850 MB, all three encoders) into\n`./models/`, which is git-ignored. Later runs load from disk in ~1–2 s\n(warm page cache).\n\n```\nsrc/model.js       EmbeddingGemma-2 loader, task prefixes, mean pooling, MRL truncation\nsrc/knn.js         nearest-exemplar classifier -> per-option probabilities (softmax)\nsrc/similarity.js  cosine / dot / softmax / ranking helpers\nsrc/metrics.js     accuracy, macro-F1, Recall@k, MRR, percentile\nsrc/usecases.js    the four use cases (ground, support-intents, spam, triage)\nsrc/index.js       public re-exports\ndata/*.json        small labelled datasets (train exemplars + held-out test)\nexamples/demo.js   runnable end-to-end example\nbench/perf.js      latency + throughput + quality harness\njeffhub use case: ground  (Retrieval — pick the passage that answers)\nQ: Which planet is known as the Red Planet?\n  1. [0.8068] d1  Mars, known for its reddish appearance, is often referred to as ...\n  2. [0.7056] d3  Jupiter is the largest planet in the solar system, with a ...\n  3. [0.6650] d2  Venus is often called Earth's twin because of its similar size ...\n\njeffhub use case: spam\nInput: Your account has been suspended, verify your password immediately at the link below.\n  -> phishing  (confidence 99.7%)\n     options: phishing=99.7%  spam=0.2%  legitimate=0.1%\n```\n\nMeasured on this machine (CPU-only, int8, single process). Numbers vary a little run-to-run with machine load.\n\n```\nEmbedding latency (batch = 1)\n  p50 168 ms | p95 177 ms | mean 168 ms   (n=30)\n\nEmbedding throughput (batched)\n  24 docs/s   (128 docs in 5.36 s, 41.8 ms/doc)\n\nPer-use-case quality + end-to-end latency\n  use case          | n  | quality                                   | p50 ms | p95 ms\n  ------------------+----+-------------------------------------------+--------+-------\n  ground            | 17 | Recall@1 100.0% | Recall@5 100.0% | MRR 1.000 |  166.9 |  171.5\n  support-intents   | 13 | accuracy 100.0% | macro-F1 100.0%         |  184.0 |  233.2\n  spam              | 13 | accuracy 100.0% | macro-F1 100.0%         |  170.7 |  185.2\n  triage            | 12 | accuracy 100.0% | macro-F1 100.0%         |  163.6 |  169.2\n```\n\n**How to read this honestly:**\n\n- **Latency/throughput** are the meaningful performance numbers here: ~170 ms\nper single-text embedding on CPU, ~24 docs/s batched. That is the price of a\n740M model on CPU with no GPU.\n- **Quality is at ceiling (100%)** because the datasets are small and curated,\nand the held-out rows are paraphrases of the exemplars. This shows the ports\nwork end-to-end; it is**not** a competitive accuracy claim. A rigorous claim\nneeds a real labelled benchmark (e.g. an MTEB retrieval subset or genuine\nsupport tickets).\n\n- **Pooling / normalization:** mean pooling, L2-normalized, so dot product ==\ncosine.\n- **Classification:** each option is scored by its nearest labelled exemplar\n(cosine), then a softmax (temperature 0.05) yields the per-option\nprobabilities jeffhub returns.\n- **Retrieval:** documents are embedded with the document prefix, queries with\nthe retrieval-query prefix, then ranked by cosine.\n- **Matryoshka (MRL):**`embed(texts, { dim: 256 })` truncates to 128/256/512-d\nand re-normalizes, trading a little quality for a 3–6× smaller vector store.\n\n- No GPU here, so throughput is CPU-bound; the same code runs on WebGPU by\npassing `device: \"webgpu\"` to`loadModel` .\n- The ONNX build loads all three encoders (text + vision + audio). A text-only\ndeployment could ship just `onnx/model_quantized.*` to save space.\n- Datasets are tiny by design (they live in `data/` ), meant to demonstrate the\npipeline rather than to benchmark the model.", "url": "https://wpnews.pro/news/embedding-gemma2-use-cases", "canonical_source": "https://github.com/karimtn/embeddingGemma2-jeff", "published_at": "2026-10-08 09:32:16+00:00", "updated_at": "2026-10-08 09:49:34.653710+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-tools", "developer-tools", "natural-language-processing"], "entities": ["Med Karim Bchini", "jeffhub.ai", "EmbeddingGemma-2", "Google", "@huggingface/transformers", "onnx-community/embeddinggemma-2-ONNX", "Node.js", "ONNX Runtime"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/embedding-gemma2-use-cases", "markdown": "https://wpnews.pro/news/embedding-gemma2-use-cases.md", "text": "https://wpnews.pro/news/embedding-gemma2-use-cases.txt", "jsonld": "https://wpnews.pro/news/embedding-gemma2-use-cases.jsonld"}}