Embedding Gemma2 Use Cases Med Karim Bchini ported four jeffhub.ai use cases — ground, support-intents, spam, and triage — onto Google's EmbeddingGemma-2 embedding model, replacing jeffhub's 0.8B "System 1" decider with cosine-similarity scoring over labelled exemplars. The project runs EmbeddingGemma-2 at int8 (onnx-community/embeddinggemma-2-ONNX, ~314 MB) via Node.js and @huggingface/transformers v4 on CPU with no Python or GPU, measuring p50 embedding latency of 168 ms and p95 of 177 ms at batch size 1, plus 24 docs/s batched throughput (128 docs in 5.36 s, 41.8 ms/doc). Author: Med Karim Bchini @karimtn Ports of jeffhub.ai https://jeffhub.ai use cases onto Google's EmbeddingGemma-2 embedding model, with a runnable demo and a latency + quality performance benchmark. jeffhub ships a 0.8B "System 1" decider model plus one small adapter per task; each adapter takes an input and a list of options and returns a probability for every option. This project keeps that request shape "pick one of your options, with a probability" but swaps the 0.8B decider for EmbeddingGemma-2 embeddings: each option is scored by cosine similarity to a few labelled exemplars classification or by dense passage ranking retrieval , and a softmax turns the scores into probabilities. | jeffhub adapter | Category | Ported here as | Quality metric | |---|---|---|---| | ground | Retrieval | Dense passage retrieval / re-ranking | Recall@1, Recall@5, MRR | | support-intents | Support | Nearest-exemplar intent classification | accuracy, macro-F1 | | spam | Safety | Legitimate / spam / phishing classification | accuracy, macro-F1 | | triage | Support | Ticket routing to a team | accuracy, macro-F1 | The runtime here has no Python, no GPU , and a small disk budget, so the examples use: - Node.js + @huggingface/transformers https://huggingface.co/docs/transformers.js v4 ONNX Runtime, CPU . - onnx-community/embeddinggemma-2-ONNX at int8 dtype: "q8" → model quantized.onnx , ~314 MB . EmbeddingGemma-2 is a 740M-parameter model 270M text backbone that maps text into a unified 768-d space. It is task-steered : a short instruction prefix tells it what representation to produce. The prefixes used here come straight from the model card: | Purpose | Prefix | |---|---| | Retrieval query | task: search result \| query: | | Document | title: none \| text: | | Classification | task: classification \| query: | npm install installs transformers.js + onnxruntime-node npm run demo runs every use case on sample inputs npm run bench latency + quality benchmark writes bench/results.json npm run bench:quick smaller timing loops The first run downloads the ONNX weights ~850 MB, all three encoders into ./models/ , which is git-ignored. Later runs load from disk in ~1–2 s warm page cache . src/model.js EmbeddingGemma-2 loader, task prefixes, mean pooling, MRL truncation src/knn.js nearest-exemplar classifier - per-option probabilities softmax src/similarity.js cosine / dot / softmax / ranking helpers src/metrics.js accuracy, macro-F1, Recall@k, MRR, percentile src/usecases.js the four use cases ground, support-intents, spam, triage src/index.js public re-exports data/ .json small labelled datasets train exemplars + held-out test examples/demo.js runnable end-to-end example bench/perf.js latency + throughput + quality harness jeffhub use case: ground Retrieval — pick the passage that answers Q: Which planet is known as the Red Planet? 1. 0.8068 d1 Mars, known for its reddish appearance, is often referred to as ... 2. 0.7056 d3 Jupiter is the largest planet in the solar system, with a ... 3. 0.6650 d2 Venus is often called Earth's twin because of its similar size ... jeffhub use case: spam Input: Your account has been suspended, verify your password immediately at the link below. - phishing confidence 99.7% options: phishing=99.7% spam=0.2% legitimate=0.1% Measured on this machine CPU-only, int8, single process . Numbers vary a little run-to-run with machine load. Embedding latency batch = 1 p50 168 ms | p95 177 ms | mean 168 ms n=30 Embedding throughput batched 24 docs/s 128 docs in 5.36 s, 41.8 ms/doc Per-use-case quality + end-to-end latency use case | n | quality | p50 ms | p95 ms ------------------+----+-------------------------------------------+--------+------- ground | 17 | Recall@1 100.0% | Recall@5 100.0% | MRR 1.000 | 166.9 | 171.5 support-intents | 13 | accuracy 100.0% | macro-F1 100.0% | 184.0 | 233.2 spam | 13 | accuracy 100.0% | macro-F1 100.0% | 170.7 | 185.2 triage | 12 | accuracy 100.0% | macro-F1 100.0% | 163.6 | 169.2 How to read this honestly: - Latency/throughput are the meaningful performance numbers here: ~170 ms per single-text embedding on CPU, ~24 docs/s batched. That is the price of a 740M model on CPU with no GPU. - Quality is at ceiling 100% because the datasets are small and curated, and the held-out rows are paraphrases of the exemplars. This shows the ports work end-to-end; it is not a competitive accuracy claim. A rigorous claim needs a real labelled benchmark e.g. an MTEB retrieval subset or genuine support tickets . - Pooling / normalization: mean pooling, L2-normalized, so dot product == cosine. - Classification: each option is scored by its nearest labelled exemplar cosine , then a softmax temperature 0.05 yields the per-option probabilities jeffhub returns. - Retrieval: documents are embedded with the document prefix, queries with the retrieval-query prefix, then ranked by cosine. - Matryoshka MRL : embed texts, { dim: 256 } truncates to 128/256/512-d and re-normalizes, trading a little quality for a 3–6× smaller vector store. - No GPU here, so throughput is CPU-bound; the same code runs on WebGPU by passing device: "webgpu" to loadModel . - The ONNX build loads all three encoders text + vision + audio . A text-only deployment could ship just onnx/model quantized. to save space. - Datasets are tiny by design they live in data/ , meant to demonstrate the pipeline rather than to benchmark the model.