{"slug": "setting-up-vector-search-for-github-starred-repos", "title": "Setting up vector search for github starred repos", "summary": "A developer built an on-device semantic search feature for GitHub starred repositories using Google's EmbeddingGemma 300M model, running through ONNX Runtime in a desktop app's Deno process so no query leaves the machine. The implementation embeds queries with a dedicated query prompt and ranks stored vectors in PGlite with pgvector via cosine distance, keeping a single shared model instance for both indexing and search.", "body_md": "Ask the starred-repo corpus questions in natural language: embed the query on-device, rank the stored vectors, and show the matching repos in the TanStack Start UI.\n\nThis is where the earlier chapters pay off. Auth (chapter 2) got us a GitHub token, the worker (chapter 3) turned stars into vectors, the bus (chapter 4) kept the UI honest while it ran, and now a search box turns a sentence into a ranked list.\n\n```\n \"graph databases in Rust\"\n        |\n        |  600 ms debounce, ?q= in the URL\n        v\n useQuery -> Eden: GET /api/elysia/enrich/starred/search?q=...\n        |\n        v\n Elysia route (TypeBox: 1..2000 chars)\n        |\n        v\n embedQuery(text)  -- EmbeddingGemma 300M, query prompt\n        |             ONNX Runtime (onnxruntime-node, CPU), Q4 weights\n        v\n Float32Array(768)\n        |\n        v\n PGlite + pgvector:  ORDER BY embedding <=> $query  LIMIT 50\n        |\n        v\n rows (no vector blob) -> list of repo links\n```\n\nEverything below the Elysia route runs inside the desktop app's Deno process. No request leaves the machine.\n\n[EmbeddingGemma](https://ai.google.dev/gemma/docs/embeddinggemma) is Google's 308M-parameter embedding model built on Gemma 3, designed for phones and laptops. The properties that matter here:\n\nWe run the [`onnx-community/embeddinggemma-300m-ONNX`](https://huggingface.co/onnx-community/embeddinggemma-300m-ONNX) export through [`@kessler/gemma-embedding`](https://github.com/kessler/gemma-embedding), a small wrapper over Transformers.js that uses native `onnxruntime-node` in Node-compatible runtimes. Our own package, [`packages/gemma-embedding`](https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/packages/gemma-embedding/src/node.ts), adds a shared instance, quantization switching, download progress, and cache inspection on top.\n\nEmbeddingGemma is **asymmetric**: the text is wrapped in a different prompt depending on whether it's something to find or something to search with. From the wrapper's source:\n\n``` js\nconst prefixed =\n  mode === \"query\"\n    ? `task: search result | query: ${text}`\n    : `title: none | text: ${text}`;\n```\n\nSo the worker embeds repos with `embedDocument()` and search embeds the user's sentence with `embedQuery()`. Mixing the two up still returns results, just noticeably worse ones. That's why both are exported as separate named functions rather than a `mode` flag callers could forget ([`instance.ts`](https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/packages/gemma-embedding/src/node-runtime/instance.ts)):\n\n``` js\nexport async function embedDocument(text: string) {\n  const embedding = await getServerGemmaEmbedding();\n  return embedding.embed(text, \"document\");\n}\n\nexport async function embedQuery(text: string) {\n  const embedding = await getServerGemmaEmbedding();\n  return embedding.embed(text, \"query\");\n}\n```\n\nLoading the model takes a few seconds and a few hundred MB, so there's exactly one instance per process. It's created on first use, and concurrent callers await the same load promise:\n\n```\nexport async function getServerGemmaEmbedding(options?: GemmaEmbeddingOptions) {\n  if (options?.dtype && options.dtype !== getActiveGemmaDtype()) {\n    await disposeInstance(); // switching quantization: drop the old one\n    setActiveDtypeInternal(options.dtype);\n  }\n\n  let instance = getEmbeddingInstance();\n  if (instance?.isLoaded()) return instance;\n\n  if (!instance) {\n    const { GemmaEmbedding } = await import(\"@kessler/gemma-embedding\");\n    instance = new GemmaEmbedding(resolveServerGemmaOptions({ ...options, dtype: getActiveGemmaDtype() }));\n    setEmbeddingInstance(instance);\n  }\n\n  if (!getEmbeddingLoadPromise()) setEmbeddingLoadPromise(instance.load() /* + progress + error handling */);\n  await getEmbeddingLoadPromise();\n  return getEmbeddingInstance()!;\n}\n```\n\nThe worker and the search route share this instance, so indexing and searching at the same time don't load the model twice.\n\nThe ONNX export ships several weight files. Settings lets the user pick one ([`catalog.ts`](https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/packages/gemma-embedding/src/catalog.ts)):\n\n| Variant | Download | Notes | \n|---|---|---|\n| `q4` | ~197 MB | Default. Smallest, and fine for repo search | \n| `q8` | ~309 MB | Balanced | \n| `fp16` | ~618 MB | Higher quality | \n| `fp32` | ~1.2 GB | Full precision | \n\nAll variants run on the CPU (`device: \"cpu\"`). `GEMMA_DTYPE` and `GEMMA_MODEL_PATH` override the choice from the environment, which is handy for pointing at a pre-downloaded model.\n\n`onnxruntime-node` is a native addon, a `.node` binary per OS and architecture. Bundling every platform's copy into the desktop binary would bloat it for no benefit, so the packaging step excludes it:\n\n```\n\"desktop:build\": \"... deno desktop ... --exclude-unused-npm --exclude ./.output/server/node_modules/onnxruntime-node --compress ...\"\n```\n\nAt runtime, [`ort-runtime.ts`](https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/apps/desktop/src/lib/embedding-gemmma/ort-runtime.ts) resolves ORT in two steps. In dev, the normal `node_modules` copy just works. In a packaged build, it uses a copy downloaded on first run into the config dir:\n\n```\nexport async function ensureOrtReady(): Promise<void> {\n  if (await canImportOrt()) return; // bundled / dev node_modules\n  if (!downloadedOrtReady()) return; // nothing downloaded yet\n  ensureOrtModulePath(); // prepend ~/.config/tangerine-desktop/native/node_modules to NODE_PATH\n}\n```\n\nThe download pulls the pinned `onnxruntime-node` tarball (the version must match what Transformers.js expects) straight from the npm registry. It extracts the tarball, deletes the binaries for other operating systems, and adds `onnxruntime-common` next to it:\n\n``` js\nconst tarballUrl = `https://registry.npmjs.org/onnxruntime-node/-/onnxruntime-node-${ORT_NPM_VERSION}.tgz`;\n// fetch with progress -> tar -xzf -> keep bin/napi-v6/<this OS>/ -> fetch onnxruntime-common\n```\n\nEvery embed call goes through `ensureOrtReady()` before importing the model package, which is why that line appears in both `embedRepo()` and the search helper.\n\nA fresh install has neither ORT nor model weights. The first time the app opens, it fetches both in the background with a progress toast. It's modelled on an IDE's first-run downloads and is cancellable from the toast.\n\nThe server side ([`embedding-bootstrap.ts`](https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/apps/desktop/src/lib/embedding-gemmma/embedding-bootstrap.ts)) runs the two downloads in order:\n\n``` js\nvoid (async () => {\n  await ensureOrtReady();\n  await beginOrtRuntimeDownload(); // ~40 MB, skipped if ORT already imports\n  await awaitOrtRuntimeDownload();\n\n  const q4 = inspectGemmaCache().variants.find((v) => v.id === \"q4\");\n  if (q4?.ready) return;\n  beginServerGemmaDtypeSwitch(\"q4\"); // ~197 MB of weights, then load\n})();\n```\n\nProgress is streamed with the chapter 4 SSE pattern. This source is a status snapshot rather than an event bus, so the generator polls once a second, only sends when something changed, and closes itself when the downloads finish ([`models/bootstrap.ts`](https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/apps/desktop/src/elysia/routes/models/bootstrap.ts)):\n\n``` js\n.get(\"/events\", async function* ({ request }) {\n  let previous: string | null = null;\n  while (!request.signal.aborted) {\n    const status = await getEmbeddingBootstrapStatus();\n    const serialized = JSON.stringify(status);\n    if (serialized !== previous) {\n      previous = serialized;\n      yield sse({ data: status });\n    }\n    if (!isLive(status)) break;\n    await sleep(1000, request.signal);\n  }\n})\n```\n\nOn the client, [`EmbeddingBootstrapHost`](https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/apps/desktop/src/components/embeddings/EmbeddingBootstrapHost.tsx) is mounted once. It calls `POST /embedding/bootstrap/start` when the server says `shouldAutoStart`, and renders the toast with a progress bar and a Cancel button.\n\nVectors live next to the repo rows in embedded Postgres ([PGlite](https://pglite.dev/docs/) with the [pgvector](https://github.com/pgvector/pgvector) extension). The column width comes from the same constant the model package exports, so the two can't drift apart ([`project-enrichment-outputs.ts`](https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/apps/desktop/src/pglite/schema/project-enrichment-outputs.ts)):\n\n``` js\n// packages/gemma-embedding/src/constants.ts\nexport const EMBEDDING_MODEL_ID = \"embeddinggemma-300m\";\nexport const EMBEDDING_DIMENSIONS = 768;\njs\nexport const projectEnrichmentOutputs = pgTable(\"project_enrichment_outputs\", {\n  id: text(\"id\").primaryKey().default(sql`gen_random_uuid()`),\n  owner: text(\"owner\").notNull(),\n  name: text(\"name\").notNull(),\n  type: text(\"type\").$type<\"starred\" | \"repos\" | \"other\">(),\n  description: text(\"description\"),\n  summary: text(\"summary\"),\n  payload: jsonb(\"payload\").notNull(), // { text } = the exact document that was embedded\n  modelId: text(\"model_id\"), // which model produced `embedding`\n  embedding: vector(\"embedding\", { dimensions: EMBEDDING_DIMENSIONS }),\n  embeddedAt: timestamp(\"embedded_at\", { withTimezone: true }),\n  // ...\n});\n```\n\nStoring `payload.text` and `modelId` makes the index easy to reason about later. You can see exactly what was embedded, and if the model ever changes you know which rows to re-embed.\n\nThe whole retrieval step is one function ([`search.ts`](https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/apps/desktop/src/elysia/routes/enrich/starred/helpers/search.ts)): embed the query, then let Postgres sort by cosine distance using Drizzle's `cosineDistance` (pgvector's `<=>` operator):\n\n```\nexport async function searchStarredByQuery(q: string): Promise<StarredSearchHit[]> {\n  const text = q.trim().slice(0, EMBED_TEXT_MAX_CHARS);\n  if (!text) return [];\n\n  await ensureOrtReady();\n  const { embedQuery, getServerGemmaEmbedding } = await import(\"@repo/gemma-embedding/node\");\n  await getServerGemmaEmbedding({ dtype: readGemmaPrefs().dtype });\n  const vector = Array.from(await embedQuery(text));\n\n  const distance = sql<number>`${cosineDistance(projectEnrichmentOutputs.embedding, vector)}`;\n\n  return db\n    .select({ id: projectEnrichmentOutputs.id, owner: projectEnrichmentOutputs.owner, name: projectEnrichmentOutputs.name, description: projectEnrichmentOutputs.description, /* … */ distance })\n    .from(projectEnrichmentOutputs)\n    .where(and(eq(projectEnrichmentOutputs.type, \"starred\"), isNotNull(projectEnrichmentOutputs.embedding)))\n    .orderBy(distance)\n    .limit(50);\n}\n```\n\nThe select lists columns explicitly and leaves out `embedding`, so 768 floats per row never cross into the UI.\n\nRight now this is an **exact scan**: Postgres computes the distance to every starred row. For a personal star list (hundreds to a few thousand rows) that's a small amount of work and always returns the true nearest neighbours. If the corpus grows much larger, an approximate HNSW index is one custom migration. drizzle-kit can't generate `USING hnsw`, so it would be hand-written:\n\n```\nCREATE INDEX project_enrichment_outputs_embedding_hnsw\n  ON project_enrichment_outputs USING hnsw (embedding vector_cosine_ops);\n```\n\nThe route exposes the search with validation and OpenAPI docs ([`enrich/starred/index.ts`](https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/apps/desktop/src/elysia/routes/enrich/starred/index.ts)):\n\n``` js\n.get(\"/search\", ({ query }) => searchStarredByQuery(query.q), {\n  query: t.Object({\n    q: t.String({ minLength: 1, maxLength: 2000, description: \"Natural-language query; embedded then ranked by cosine distance.\" }),\n  }),\n})\n```\n\nThe starred page ([`EnrichedStarred.tsx`](https://github.com/tigawanna/tangerine/blob/cac981aca4b365f2ee1293fbf5248da5548c3b5a/apps/desktop/src/routes/_dashboard/$user/enriched/-components/starred/EnrichedStarred.tsx)) reuses the app's normal list scaffold. The search box writes `?q=` to the URL after a 600 ms debounce, and a query hook calls the typed Eden client:\n\n``` js\nconst q = (routeApi.useSearch().q ?? \"\").trim();\n\nconst semantic = useQuery({\n  queryKey: [\"enriched-starred-search\", q],\n  enabled: q.length > 0,\n  placeholderData: (previous) => previous, // keep old hits on screen while the next query embeds\n  queryFn: async () => {\n    const { data, error } = await getElysiaTreaty().enrich.starred.search.get({ query: { q } });\n    if (error) throw new Error(treatyErrorMessage(error));\n    return data ?? [];\n  },\n});\n\nconst rows = q ? (semantic.data ?? []) : allRows;\n```\n\nA few small choices make it feel local rather than remote:\n\n`placeholderData`` q` shows every embedded star, paginated, and that list keeps growing live during indexing (chapter 4).`/$user/repos/$repo`).\nThe finished loop: sign in once, click **Embed starred**, watch repos stream into the list, and search them by meaning. It all runs on your machine with a ~200 MB model.\n\nIt's deliberately \"retrieval only\": the answer is a ranked list of your own repos, not generated text. That keeps it fast, obviously correct (every result is a real star), and free of a second, much larger model.\n\nNatural next steps, roughly in order of payoff:\n\n`WHERE` next to the vector sort.\nI tried hard to keep the shipped binary small:\n\n`backend: \"webview\"` uses the OS webview instead of bundling Chromium, which saves roughly 100 MB (chapter 1).\nEven so, the Deno runtime alone is about **70 MB**, and that's a floor this approach can't go below. For a tool whose job is \"search my stars\", that feels like a lot.\n\nI suspect a Go shell with [Wails](https://wails.io/) (a single Go binary using the OS webview) could come in well under that. So I've started learning Go, and I'm looking forward to rebuilding a slice of this and writing up what I find.", "url": "https://wpnews.pro/news/setting-up-vector-search-for-github-starred-repos", "canonical_source": "https://dev.to/tigawanna/setting-up-vector-search-for-github-starred-repos-6f5", "published_at": "2026-09-30 17:37:38+00:00", "updated_at": "2026-09-30 17:47:09.469902+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "large-language-models", "ai-infrastructure"], "entities": ["GitHub", "Google", "EmbeddingGemma", "ONNX Runtime", "PGlite", "pgvector", "TanStack Start", "Elysia"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/setting-up-vector-search-for-github-starred-repos", "markdown": "https://wpnews.pro/news/setting-up-vector-search-for-github-starred-repos.md", "text": "https://wpnews.pro/news/setting-up-vector-search-for-github-starred-repos.txt", "jsonld": "https://wpnews.pro/news/setting-up-vector-search-for-github-starred-repos.jsonld"}}