Author: Med Karim Bchini @karimtn
Ports of jeffhub.ai use cases onto Google's EmbeddingGemma-2 embedding model, with a runnable demo and a latency + quality performance benchmark.
jeffhub ships a 0.8B "System 1" decider model plus one small adapter per task; each adapter takes an input and a list of options and returns a probability for every option. This project keeps that request shape ("pick one of your options, with a probability") but swaps the 0.8B decider for EmbeddingGemma-2 embeddings: each option is scored by cosine similarity to a few labelled exemplars (classification) or by dense passage ranking (retrieval), and a softmax turns the scores into probabilities.
| jeffhub adapter | Category | Ported here as | Quality metric |
|---|---|---|---|
ground |
Retrieval | Dense passage retrieval / re-ranking | Recall@1, Recall@5, MRR |
support-intents |
Support | Nearest-exemplar intent classification | accuracy, macro-F1 |
spam |
Safety | Legitimate / spam / phishing classification | accuracy, macro-F1 |
triage |
Support | Ticket routing to a team | accuracy, macro-F1 |
The runtime here has no Python, no GPU, and a small disk budget, so the examples use:
- Node.js +
@huggingface/transformersv4 (ONNX Runtime, CPU). onnx-community/embeddinggemma-2-ONNXatint8 (dtype: "q8"→model_quantized.onnx, ~314 MB).
EmbeddingGemma-2 is a 740M-parameter model (270M text backbone) that maps text into a unified 768-d space. It is task-steered: a short instruction prefix tells it what representation to produce. The prefixes used here come straight from the model card:
| Purpose | Prefix |
|---|---|
| Retrieval query | task: search result | query: |
| Document | title: none | text: |
| Classification | task: classification | query: |
npm install # installs transformers.js + onnxruntime-node
npm run demo # runs every use case on sample inputs
npm run bench # latency + quality benchmark (writes bench/results.json)
npm run bench:quick # smaller timing loops
The first run downloads the ONNX weights (~850 MB, all three encoders) into
./models/, which is git-ignored. Later runs load from disk in ~1–2 s
(warm page cache).
src/model.js EmbeddingGemma-2 , task prefixes, mean pooling, MRL truncation
src/knn.js nearest-exemplar classifier -> per-option probabilities (softmax)
src/similarity.js cosine / dot / softmax / ranking helpers
src/metrics.js accuracy, macro-F1, Recall@k, MRR, percentile
src/usecases.js the four use cases (ground, support-intents, spam, triage)
src/index.js public re-exports
data/*.json small labelled datasets (train exemplars + held-out test)
examples/demo.js runnable end-to-end example
bench/perf.js latency + throughput + quality harness
jeffhub use case: ground (Retrieval — pick the passage that answers)
Q: Which planet is known as the Red Planet?
1. [0.8068] d1 Mars, known for its reddish appearance, is often referred to as ...
2. [0.7056] d3 Jupiter is the largest planet in the solar system, with a ...
3. [0.6650] d2 Venus is often called Earth's twin because of its similar size ...
jeffhub use case: spam
Input: Your account has been suspended, verify your password immediately at the link below.
-> phishing (confidence 99.7%)
options: phishing=99.7% spam=0.2% legitimate=0.1%
Measured on this machine (CPU-only, int8, single process). Numbers vary a little run-to-run with machine load.
Embedding latency (batch = 1)
p50 168 ms | p95 177 ms | mean 168 ms (n=30)
Embedding throughput (batched)
24 docs/s (128 docs in 5.36 s, 41.8 ms/doc)
Per-use-case quality + end-to-end latency
use case | n | quality | p50 ms | p95 ms
------------------+----+-------------------------------------------+--------+-------
ground | 17 | Recall@1 100.0% | Recall@5 100.0% | MRR 1.000 | 166.9 | 171.5
support-intents | 13 | accuracy 100.0% | macro-F1 100.0% | 184.0 | 233.2
spam | 13 | accuracy 100.0% | macro-F1 100.0% | 170.7 | 185.2
triage | 12 | accuracy 100.0% | macro-F1 100.0% | 163.6 | 169.2
How to read this honestly:
-
Latency/throughput are the meaningful performance numbers here: ~170 ms per single-text embedding on CPU, ~24 docs/s batched. That is the price of a 740M model on CPU with no GPU.
-
Quality is at ceiling (100%) because the datasets are small and curated, and the held-out rows are paraphrases of the exemplars. This shows the ports work end-to-end; it isnot a competitive accuracy claim. A rigorous claim needs a real labelled benchmark (e.g. an MTEB retrieval subset or genuine support tickets).
-
Pooling / normalization: mean pooling, L2-normalized, so dot product == cosine.
-
Classification: each option is scored by its nearest labelled exemplar (cosine), then a softmax (temperature 0.05) yields the per-option probabilities jeffhub returns.
-
Retrieval: documents are embedded with the document prefix, queries with the retrieval-query prefix, then ranked by cosine.
-
Matryoshka (MRL):
embed(texts, { dim: 256 })truncates to 128/256/512-d and re-normalizes, trading a little quality for a 3–6× smaller vector store. -
No GPU here, so throughput is CPU-bound; the same code runs on WebGPU by passing
device: "webgpu"toloadModel. -
The ONNX build loads all three encoders (text + vision + audio). A text-only deployment could ship just
onnx/model_quantized.*to save space. -
Datasets are tiny by design (they live in
data/), meant to demonstrate the pipeline rather than to benchmark the model.