cd /news/artificial-intelligence/embedding-gemma2-use-cases · home › topics › artificial-intelligence › article
[ARTICLE · art-147463] src=github.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Embedding Gemma2 Use Cases

Med Karim Bchini ported four jeffhub.ai use cases — ground, support-intents, spam, and triage — onto Google's EmbeddingGemma-2 embedding model, replacing jeffhub's 0.8B "System 1" decider with cosine-similarity scoring over labelled exemplars. The project runs EmbeddingGemma-2 at int8 (onnx-community/embeddinggemma-2-ONNX, ~314 MB) via Node.js and @huggingface/transformers v4 on CPU with no Python or GPU, measuring p50 embedding latency of 168 ms and p95 of 177 ms at batch size 1, plus 24 docs/s batched throughput (128 docs in 5.36 s, 41.8 ms/doc).

read4 min views1 publishedOct 8, 2026
Embedding Gemma2 Use Cases
Image: Michielbdejong (auto-discovered)

Author: Med Karim Bchini @karimtn

Ports of jeffhub.ai use cases onto Google's EmbeddingGemma-2 embedding model, with a runnable demo and a latency + quality performance benchmark.

jeffhub ships a 0.8B "System 1" decider model plus one small adapter per task; each adapter takes an input and a list of options and returns a probability for every option. This project keeps that request shape ("pick one of your options, with a probability") but swaps the 0.8B decider for EmbeddingGemma-2 embeddings: each option is scored by cosine similarity to a few labelled exemplars (classification) or by dense passage ranking (retrieval), and a softmax turns the scores into probabilities.

jeffhub adapter Category Ported here as Quality metric
ground Retrieval Dense passage retrieval / re-ranking Recall@1, Recall@5, MRR
support-intents Support Nearest-exemplar intent classification accuracy, macro-F1
spam Safety Legitimate / spam / phishing classification accuracy, macro-F1
triage Support Ticket routing to a team accuracy, macro-F1

The runtime here has no Python, no GPU, and a small disk budget, so the examples use:

  • Node.js + @huggingface/transformers v4 (ONNX Runtime, CPU).
  • onnx-community/embeddinggemma-2-ONNX atint8 ( dtype: "q8" →model_quantized.onnx , ~314 MB).

EmbeddingGemma-2 is a 740M-parameter model (270M text backbone) that maps text into a unified 768-d space. It is task-steered: a short instruction prefix tells it what representation to produce. The prefixes used here come straight from the model card:

Purpose Prefix
Retrieval query task: search result | query:
Document title: none | text:
Classification task: classification | query:
npm install          # installs transformers.js + onnxruntime-node
npm run demo         # runs every use case on sample inputs
npm run bench        # latency + quality benchmark (writes bench/results.json)
npm run bench:quick  # smaller timing loops

The first run downloads the ONNX weights (~850 MB, all three encoders) into ./models/, which is git-ignored. Later runs load from disk in ~1–2 s (warm page cache).

src/model.js       EmbeddingGemma-2 , task prefixes, mean pooling, MRL truncation
src/knn.js         nearest-exemplar classifier -> per-option probabilities (softmax)
src/similarity.js  cosine / dot / softmax / ranking helpers
src/metrics.js     accuracy, macro-F1, Recall@k, MRR, percentile
src/usecases.js    the four use cases (ground, support-intents, spam, triage)
src/index.js       public re-exports
data/*.json        small labelled datasets (train exemplars + held-out test)
examples/demo.js   runnable end-to-end example
bench/perf.js      latency + throughput + quality harness
jeffhub use case: ground  (Retrieval — pick the passage that answers)
Q: Which planet is known as the Red Planet?
  1. [0.8068] d1  Mars, known for its reddish appearance, is often referred to as ...
  2. [0.7056] d3  Jupiter is the largest planet in the solar system, with a ...
  3. [0.6650] d2  Venus is often called Earth's twin because of its similar size ...

jeffhub use case: spam
Input: Your account has been suspended, verify your password immediately at the link below.
  -> phishing  (confidence 99.7%)
     options: phishing=99.7%  spam=0.2%  legitimate=0.1%

Measured on this machine (CPU-only, int8, single process). Numbers vary a little run-to-run with machine load.

Embedding latency (batch = 1)
  p50 168 ms | p95 177 ms | mean 168 ms   (n=30)

Embedding throughput (batched)
  24 docs/s   (128 docs in 5.36 s, 41.8 ms/doc)

Per-use-case quality + end-to-end latency
  use case          | n  | quality                                   | p50 ms | p95 ms
  ------------------+----+-------------------------------------------+--------+-------
  ground            | 17 | Recall@1 100.0% | Recall@5 100.0% | MRR 1.000 |  166.9 |  171.5
  support-intents   | 13 | accuracy 100.0% | macro-F1 100.0%         |  184.0 |  233.2
  spam              | 13 | accuracy 100.0% | macro-F1 100.0%         |  170.7 |  185.2
  triage            | 12 | accuracy 100.0% | macro-F1 100.0%         |  163.6 |  169.2

How to read this honestly:

  • Latency/throughput are the meaningful performance numbers here: ~170 ms per single-text embedding on CPU, ~24 docs/s batched. That is the price of a 740M model on CPU with no GPU.

  • Quality is at ceiling (100%) because the datasets are small and curated, and the held-out rows are paraphrases of the exemplars. This shows the ports work end-to-end; it isnot a competitive accuracy claim. A rigorous claim needs a real labelled benchmark (e.g. an MTEB retrieval subset or genuine support tickets).

  • Pooling / normalization: mean pooling, L2-normalized, so dot product == cosine.

  • Classification: each option is scored by its nearest labelled exemplar (cosine), then a softmax (temperature 0.05) yields the per-option probabilities jeffhub returns.

  • Retrieval: documents are embedded with the document prefix, queries with the retrieval-query prefix, then ranked by cosine.

  • Matryoshka (MRL):embed(texts, { dim: 256 }) truncates to 128/256/512-d and re-normalizes, trading a little quality for a 3–6× smaller vector store.

  • No GPU here, so throughput is CPU-bound; the same code runs on WebGPU by passing device: "webgpu" toloadModel .

  • The ONNX build loads all three encoders (text + vision + audio). A text-only deployment could ship just onnx/model_quantized.* to save space.

  • Datasets are tiny by design (they live in data/ ), meant to demonstrate the pipeline rather than to benchmark the model.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @med karim bchini 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/embedding-gemma2-use…] indexed:0 read:4min 2026-10-08 · —