# Embedding Gemma2 Use Cases

> Source: <https://github.com/karimtn/embeddingGemma2-jeff>
> Published: 2026-10-08 09:32:16+00:00

**Author:** Med Karim Bchini @karimtn

Ports of **[jeffhub.ai](https://jeffhub.ai)** use cases onto Google's
**EmbeddingGemma-2** embedding model, with a runnable demo and a
**latency + quality** performance benchmark.

jeffhub ships a 0.8B "System 1" decider model plus one small adapter per task;
each adapter takes an input and a list of options and returns a probability for
every option. This project keeps that request shape ("pick one of your options,
with a probability") but swaps the 0.8B decider for **EmbeddingGemma-2**
embeddings: each option is scored by cosine similarity to a few labelled
exemplars (classification) or by dense passage ranking (retrieval), and a
softmax turns the scores into probabilities.

| jeffhub adapter | Category | Ported here as | Quality metric | 
|---|---|---|---|
| `ground` | Retrieval | Dense passage retrieval / re-ranking | Recall@1, Recall@5, MRR | 
| `support-intents` | Support | Nearest-exemplar intent classification | accuracy, macro-F1 | 
| `spam` | Safety | Legitimate / spam / phishing classification | accuracy, macro-F1 | 
| `triage` | Support | Ticket routing to a team | accuracy, macro-F1 | 

The runtime here has **no Python, no GPU**, and a small disk budget, so the
examples use:

- **Node.js + [`@huggingface/transformers`](https://huggingface.co/docs/transformers.js) v4** (ONNX Runtime, CPU).
- **`onnx-community/embeddinggemma-2-ONNX`** at**int8** (` dtype: "q8"` →`model_quantized.onnx` , ~314 MB).

EmbeddingGemma-2 is a 740M-parameter model (270M text backbone) that maps text
into a unified 768-d space. It is *task-steered*: a short instruction prefix
tells it what representation to produce. The prefixes used here come straight
from the model card:

| Purpose | Prefix | 
|---|---|
| Retrieval query | `task: search result \| query:`  | 
| Document | `title: none \| text:`  | 
| Classification | `task: classification \| query:`  | 

```
npm install          # installs transformers.js + onnxruntime-node
npm run demo         # runs every use case on sample inputs
npm run bench        # latency + quality benchmark (writes bench/results.json)
npm run bench:quick  # smaller timing loops
```

The first run downloads the ONNX weights (~850 MB, all three encoders) into
`./models/`, which is git-ignored. Later runs load from disk in ~1–2 s
(warm page cache).

```
src/model.js       EmbeddingGemma-2 loader, task prefixes, mean pooling, MRL truncation
src/knn.js         nearest-exemplar classifier -> per-option probabilities (softmax)
src/similarity.js  cosine / dot / softmax / ranking helpers
src/metrics.js     accuracy, macro-F1, Recall@k, MRR, percentile
src/usecases.js    the four use cases (ground, support-intents, spam, triage)
src/index.js       public re-exports
data/*.json        small labelled datasets (train exemplars + held-out test)
examples/demo.js   runnable end-to-end example
bench/perf.js      latency + throughput + quality harness
jeffhub use case: ground  (Retrieval — pick the passage that answers)
Q: Which planet is known as the Red Planet?
  1. [0.8068] d1  Mars, known for its reddish appearance, is often referred to as ...
  2. [0.7056] d3  Jupiter is the largest planet in the solar system, with a ...
  3. [0.6650] d2  Venus is often called Earth's twin because of its similar size ...

jeffhub use case: spam
Input: Your account has been suspended, verify your password immediately at the link below.
  -> phishing  (confidence 99.7%)
     options: phishing=99.7%  spam=0.2%  legitimate=0.1%
```

Measured on this machine (CPU-only, int8, single process). Numbers vary a little run-to-run with machine load.

```
Embedding latency (batch = 1)
  p50 168 ms | p95 177 ms | mean 168 ms   (n=30)

Embedding throughput (batched)
  24 docs/s   (128 docs in 5.36 s, 41.8 ms/doc)

Per-use-case quality + end-to-end latency
  use case          | n  | quality                                   | p50 ms | p95 ms
  ------------------+----+-------------------------------------------+--------+-------
  ground            | 17 | Recall@1 100.0% | Recall@5 100.0% | MRR 1.000 |  166.9 |  171.5
  support-intents   | 13 | accuracy 100.0% | macro-F1 100.0%         |  184.0 |  233.2
  spam              | 13 | accuracy 100.0% | macro-F1 100.0%         |  170.7 |  185.2
  triage            | 12 | accuracy 100.0% | macro-F1 100.0%         |  163.6 |  169.2
```

**How to read this honestly:**

- **Latency/throughput** are the meaningful performance numbers here: ~170 ms
per single-text embedding on CPU, ~24 docs/s batched. That is the price of a
740M model on CPU with no GPU.
- **Quality is at ceiling (100%)** because the datasets are small and curated,
and the held-out rows are paraphrases of the exemplars. This shows the ports
work end-to-end; it is**not** a competitive accuracy claim. A rigorous claim
needs a real labelled benchmark (e.g. an MTEB retrieval subset or genuine
support tickets).

- **Pooling / normalization:** mean pooling, L2-normalized, so dot product ==
cosine.
- **Classification:** each option is scored by its nearest labelled exemplar
(cosine), then a softmax (temperature 0.05) yields the per-option
probabilities jeffhub returns.
- **Retrieval:** documents are embedded with the document prefix, queries with
the retrieval-query prefix, then ranked by cosine.
- **Matryoshka (MRL):**`embed(texts, { dim: 256 })` truncates to 128/256/512-d
and re-normalizes, trading a little quality for a 3–6× smaller vector store.

- No GPU here, so throughput is CPU-bound; the same code runs on WebGPU by
passing `device: "webgpu"` to`loadModel` .
- The ONNX build loads all three encoders (text + vision + audio). A text-only
deployment could ship just `onnx/model_quantized.*` to save space.
- Datasets are tiny by design (they live in `data/` ), meant to demonstrate the
pipeline rather than to benchmark the model.
