Setting up vector search for github starred repos A developer built an on-device semantic search feature for GitHub starred repositories using Google's EmbeddingGemma 300M model, running through ONNX Runtime in a desktop app's Deno process so no query leaves the machine. The implementation embeds queries with a dedicated query prompt and ranks stored vectors in PGlite with pgvector via cosine distance, keeping a single shared model instance for both indexing and search. Ask the starred-repo corpus questions in natural language: embed the query on-device, rank the stored vectors, and show the matching repos in the TanStack Start UI. This is where the earlier chapters pay off. Auth chapter 2 got us a GitHub token, the worker chapter 3 turned stars into vectors, the bus chapter 4 kept the UI honest while it ran, and now a search box turns a sentence into a ranked list. "graph databases in Rust" | | 600 ms debounce, ?q= in the URL v useQuery - Eden: GET /api/elysia/enrich/starred/search?q=... | v Elysia route TypeBox: 1..2000 chars | v embedQuery text -- EmbeddingGemma 300M, query prompt | ONNX Runtime onnxruntime-node, CPU , Q4 weights v Float32Array 768 | v PGlite + pgvector: ORDER BY embedding <= $query LIMIT 50 | v rows no vector blob - list of repo links Everything below the Elysia route runs inside the desktop app's Deno process. No request leaves the machine. EmbeddingGemma https://ai.google.dev/gemma/docs/embeddinggemma is Google's 308M-parameter embedding model built on Gemma 3, designed for phones and laptops. The properties that matter here: We run the onnx-community/embeddinggemma-300m-ONNX https://huggingface.co/onnx-community/embeddinggemma-300m-ONNX export through @kessler/gemma-embedding https://github.com/kessler/gemma-embedding , a small wrapper over Transformers.js that uses native onnxruntime-node in Node-compatible runtimes. Our own package, packages/gemma-embedding https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/packages/gemma-embedding/src/node.ts , adds a shared instance, quantization switching, download progress, and cache inspection on top. EmbeddingGemma is asymmetric : the text is wrapped in a different prompt depending on whether it's something to find or something to search with. From the wrapper's source: js const prefixed = mode === "query" ? task: search result | query: ${text} : title: none | text: ${text} ; So the worker embeds repos with embedDocument and search embeds the user's sentence with embedQuery . Mixing the two up still returns results, just noticeably worse ones. That's why both are exported as separate named functions rather than a mode flag callers could forget instance.ts https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/packages/gemma-embedding/src/node-runtime/instance.ts : js export async function embedDocument text: string { const embedding = await getServerGemmaEmbedding ; return embedding.embed text, "document" ; } export async function embedQuery text: string { const embedding = await getServerGemmaEmbedding ; return embedding.embed text, "query" ; } Loading the model takes a few seconds and a few hundred MB, so there's exactly one instance per process. It's created on first use, and concurrent callers await the same load promise: export async function getServerGemmaEmbedding options?: GemmaEmbeddingOptions { if options?.dtype && options.dtype == getActiveGemmaDtype { await disposeInstance ; // switching quantization: drop the old one setActiveDtypeInternal options.dtype ; } let instance = getEmbeddingInstance ; if instance?.isLoaded return instance; if instance { const { GemmaEmbedding } = await import "@kessler/gemma-embedding" ; instance = new GemmaEmbedding resolveServerGemmaOptions { ...options, dtype: getActiveGemmaDtype } ; setEmbeddingInstance instance ; } if getEmbeddingLoadPromise setEmbeddingLoadPromise instance.load / + progress + error handling / ; await getEmbeddingLoadPromise ; return getEmbeddingInstance ; } The worker and the search route share this instance, so indexing and searching at the same time don't load the model twice. The ONNX export ships several weight files. Settings lets the user pick one catalog.ts https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/packages/gemma-embedding/src/catalog.ts : | Variant | Download | Notes | |---|---|---| | q4 | ~197 MB | Default. Smallest, and fine for repo search | | q8 | ~309 MB | Balanced | | fp16 | ~618 MB | Higher quality | | fp32 | ~1.2 GB | Full precision | All variants run on the CPU device: "cpu" . GEMMA DTYPE and GEMMA MODEL PATH override the choice from the environment, which is handy for pointing at a pre-downloaded model. onnxruntime-node is a native addon, a .node binary per OS and architecture. Bundling every platform's copy into the desktop binary would bloat it for no benefit, so the packaging step excludes it: "desktop:build": "... deno desktop ... --exclude-unused-npm --exclude ./.output/server/node modules/onnxruntime-node --compress ..." At runtime, ort-runtime.ts https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/apps/desktop/src/lib/embedding-gemmma/ort-runtime.ts resolves ORT in two steps. In dev, the normal node modules copy just works. In a packaged build, it uses a copy downloaded on first run into the config dir: export async function ensureOrtReady : Promise