Ask the starred-repo corpus questions in natural language: embed the query on-device, rank the stored vectors, and show the matching repos in the TanStack Start UI.
This is where the earlier chapters pay off. Auth (chapter 2) got us a GitHub token, the worker (chapter 3) turned stars into vectors, the bus (chapter 4) kept the UI honest while it ran, and now a search box turns a sentence into a ranked list.
1 "graph databases in Rust"
2 |
3 | 600 ms debounce, ?q= in the URL
4 v
5 useQuery -> Eden: GET /api/elysia/enrich/starred/search?q=...
6 |
7 v
8 Elysia route (TypeBox: 1..2000 chars)
9 |
10 v
11 embedQuery(text) -- EmbeddingGemma 300M, query prompt
12 | ONNX Runtime (onnxruntime-node, CPU), Q4 weights
13 v
14 Float32Array(768)
15 |
16 v
17 PGlite + pgvector: ORDER BY embedding <=> $query LIMIT 50
18 |
19 v
20 rows (no vector blob) -> list of repo links
Everything below the Elysia route runs inside the desktop app's Deno process. No request leaves the machine.
EmbeddingGemma is Google's 308M-parameter embedding model built on Gemma 3, designed for phones and laptops. The properties that matter here:
- 768-dimensional vectors. The size can be cut down to 512, 256, or 128 through Matryoshka representation learning. We keep 768.
- 2K-token context, which is roomy for a repo's metadata plus the top of its README (chapter 3).
- 100+ languages, so non-English READMEs and queries still land.
- Small when quantized. Google quotes under 200 MB of RAM with quantization.
- Fully offline once the weights are on disk.
We run the onnx-community/embeddinggemma-300m-ONNX export through @kessler/gemma-embedding, a small wrapper over Transformers.js that uses native onnxruntime-node in Node-compatible runtimes. Our own package, packages/gemma-embedding, adds a shared instance, quantization switching, download progress, and cache inspection on top.
EmbeddingGemma is asymmetric: the text is wrapped in a different prompt depending on whether it's something to find or something to search with. From the wrapper's source:
1const prefixed =
2 mode === "query"
3 ? `task: search result | query: ${text}`
4 : `title: none | text: ${text}`;
So the worker embeds repos with embedDocument() and search embeds the user's sentence with embedQuery(). Mixing the two up still returns results, just noticeably worse ones. That's why both are exported as separate named functions rather than a mode flag callers could forget (instance.ts):
1export async function embedDocument(text: string) {
2 const embedding = await getServerGemmaEmbedding();
3 return embedding.embed(text, "document");
4}
5
6export async function embedQuery(text: string) {
7 const embedding = await getServerGemmaEmbedding();
8 return embedding.embed(text, "query");
9}
Loading the model takes a few seconds and a few hundred MB, so there's exactly one instance per process. It's created on first use, and concurrent callers await the same load promise:
1export async function getServerGemmaEmbedding(options?: GemmaEmbeddingOptions) {
2 if (options?.dtype && options.dtype !== getActiveGemmaDtype()) {
3 await disposeInstance();
4 setActiveDtypeInternal(options.dtype);
5 }
6
7 let instance = getEmbeddingInstance();
8 if (instance?.isLoaded()) return instance;
9
10 if (!instance) {
11 const { GemmaEmbedding } = await import("@kessler/gemma-embedding");
12 instance = new GemmaEmbedding(resolveServerGemmaOptions({ ...options, dtype: getActiveGemmaDtype() }));
13 setEmbeddingInstance(instance);
14 }
15
16 if (!getEmbeddingLoadPromise()) setEmbeddingLoadPromise(instance.load() );
17 await getEmbeddingLoadPromise();
18 return getEmbeddingInstance()!;
19}
The worker and the search route share this instance, so indexing and searching at the same time don't load the model twice.
The ONNX export ships several weight files. Settings lets the user pick one (catalog.ts):
All variants run on the CPU (device: "cpu"). GEMMA_DTYPE and GEMMA_MODEL_PATH override the choice from the environment, which is handy for pointing at a pre-downloaded model.
onnxruntime-node is a native addon, a .node binary per OS and architecture. Bundling every platform's copy into the desktop binary would bloat it for no benefit, so the packaging step excludes it:
1"desktop:build": "... deno desktop ... --exclude-unused-npm --exclude ./.output/server/node_modules/onnxruntime-node --compress ..."
At runtime, ort-runtime.ts resolves ORT in two steps. In dev, the normal node_modules copy just works. In a packaged build, it uses a copy downloaded on first run into the config dir:
1export async function ensureOrtReady(): Promise<void> {
2 if (await canImportOrt()) return;
3 if (!downloadedOrtReady()) return;
4 ensureOrtModulePath();
5}
The download pulls the pinned onnxruntime-node tarball (the version must match what Transformers.js expects) straight from the npm registry. It extracts the tarball, deletes the binaries for other operating systems, and adds onnxruntime-common next to it:
1const tarballUrl = `https://registry.npmjs.org/onnxruntime-node/-/onnxruntime-node-${ORT_NPM_VERSION}.tgz`;
2
Every embed call goes through ensureOrtReady() before importing the model package, which is why that line appears in both embedRepo() and the search helper.
A fresh install has neither ORT nor model weights. The first time the app opens, it fetches both in the background with a progress toast. It's modelled on an IDE's first-run downloads and is cancellable from the toast.
The server side (embedding-bootstrap.ts) runs the two downloads in order:
1void (async () => {
2 await ensureOrtReady();
3 await beginOrtRuntimeDownload();
4 await awaitOrtRuntimeDownload();
5
6 const q4 = inspectGemmaCache().variants.find((v) => v.id === "q4");
7 if (q4?.ready) return;
8 beginServerGemmaDtypeSwitch("q4");
9})();
Progress is streamed with the chapter 4 SSE pattern. This source is a status snapshot rather than an event bus, so the generator polls once a second, only sends when something changed, and closes itself when the downloads finish (models/bootstrap.ts):
1.get("/events", async function* ({ request }) {
2 let previous: string | null = null;
3 while (!request.signal.aborted) {
4 const status = await getEmbeddingBootstrapStatus();
5 const serialized = JSON.stringify(status);
6 if (serialized !== previous) {
7 previous = serialized;
8 yield sse({ data: status });
9 }
10 if (!isLive(status)) break;
11 await sleep(1000, request.signal);
12 }
13})
On the client, EmbeddingBootstrapHost is mounted once. It calls POST /embedding/bootstrap/start when the server says shouldAutoStart, and renders the toast with a progress bar and a Cancel button.
Vectors live next to the repo rows in embedded Postgres (PGlite with the pgvector extension). The column width comes from the same constant the model package exports, so the two can't drift apart (project-enrichment-outputs.ts):
1
2export const EMBEDDING_MODEL_ID = "embeddinggemma-300m";
3export const EMBEDDING_DIMENSIONS = 768;
1export const projectEnrichmentOutputs = pgTable("project_enrichment_outputs", {
2 id: text("id").primaryKey().default(sql`gen_random_uuid()`),
3 owner: text("owner").notNull(),
4 name: text("name").notNull(),
5 type: text("type").$type<"starred" | "repos" | "other">(),
6 description: text("description"),
7 summary: text("summary"),
8 payload: jsonb("payload").notNull(),
9 modelId: text("model_id"),
10 embedding: vector("embedding", { dimensions: EMBEDDING_DIMENSIONS }),
11 embeddedAt: timestamp("embedded_at", { withTimezone: true }),
12
13});
Storing payload.text and modelId makes the index easy to reason about later. You can see exactly what was embedded, and if the model ever changes you know which rows to re-embed.
The whole retrieval step is one function (search.ts): embed the query, then let Postgres sort by cosine distance using Drizzle's cosineDistance (pgvector's <=> operator):
1export async function searchStarredByQuery(q: string): Promise<StarredSearchHit[]> {
2 const text = q.trim().slice(0, EMBED_TEXT_MAX_CHARS);
3 if (!text) return [];
4
5 await ensureOrtReady();
6 const { embedQuery, getServerGemmaEmbedding } = await import("@repo/gemma-embedding/node");
7 await getServerGemmaEmbedding({ dtype: readGemmaPrefs().dtype });
8 const vector = Array.from(await embedQuery(text));
9
10 const distance = sql<number>`${cosineDistance(projectEnrichmentOutputs.embedding, vector)}`;
11
12 return db
13 .select({ id: projectEnrichmentOutputs.id, owner: projectEnrichmentOutputs.owner, name: projectEnrichmentOutputs.name, description: projectEnrichmentOutputs.description, distance })
14 .from(projectEnrichmentOutputs)
15 .where(and(eq(projectEnrichmentOutputs.type, "starred"), isNotNull(projectEnrichmentOutputs.embedding)))
16 .orderBy(distance)
17 .limit(50);
18}
The select lists columns explicitly and leaves out embedding, so 768 floats per row never cross into the UI.
Right now this is an exact scan: Postgres computes the distance to every starred row. For a personal star list (hundreds to a few thousand rows) that's a small amount of work and always returns the true nearest neighbours. If the corpus grows much larger, an approximate HNSW index is one custom migration. drizzle-kit can't generate USING hnsw, so it would be hand-written:
1CREATE INDEX project_enrichment_outputs_embedding_hnsw
2 ON project_enrichment_outputs USING hnsw (embedding vector_cosine_ops);
The route exposes the search with validation and OpenAPI docs (enrich/starred/index.ts):
1.get("/search", ({ query }) => searchStarredByQuery(query.q), {
2 query: t.Object({
3 q: t.String({ minLength: 1, maxLength: 2000, description: "Natural-language query; embedded then ranked by cosine distance." }),
4 }),
5})
The starred page (EnrichedStarred.tsx) reuses the app's normal list scaffold. The search box writes ?q= to the URL after a 600 ms debounce, and a query hook calls the typed Eden client:
1const q = (routeApi.useSearch().q ?? "").trim();
2
3const semantic = useQuery({
4 queryKey: ["enriched-starred-search", q],
5 enabled: q.length > 0,
6 placeholderData: (previous) => previous,
7 queryFn: async () => {
8 const { data, error } = await getElysiaTreaty().enrich.starred.search.get({ query: { q } });
9 if (error) throw new Error(treatyErrorMessage(error));
10 return data ?? [];
11 },
12});
13
14const rows = q ? (semantic.data ?? []) : allRows;
A few small choices make it feel local rather than remote:
- The placeholder hints at meaning, not keywords: "Try natural language — e.g. graph databases in Rust…".
- The loading state says "Embedding query…" instead of a generic spinner, which also explains the short wait on the very first search while the model loads.
placeholderData keeps the previous results visible while you refine the query, so the list doesn't flash empty on every keystroke.- No query means the full corpus. An empty
q shows every embedded star, paginated, and that list keeps growing live during indexing (chapter 4). - Each hit is a link to the in-app repo page (
/$user/repos/$repo).
The finished loop: sign in once, click Embed starred, watch repos stream into the list, and search them by meaning. It all runs on your machine with a ~200 MB model.
It's deliberately "retrieval only": the answer is a ranked list of your own repos, not generated text. That keeps it fast, obviously correct (every result is a real star), and free of a second, much larger model.
Natural next steps, roughly in order of payoff:
- Filters on language, topic, or owner, as a
WHERE next to the vector sort. - Hybrid ranking: blend in Postgres full-text search, so exact names ("tokio") always rank first.
- Chunked READMEs for deep-content matches (see the note in chapter 3).
- An HNSW index once the corpus outgrows an exact scan.
- Optional local generation: pass the top hits to a small local LLM for a one-paragraph answer that cites them.
I tried hard to keep the shipped binary small:
- Nothing heavy in the bundle. ONNX Runtime and the EmbeddingGemma weights are fetched at runtime (the first-run bootstrap above), not packaged.
- Lazy imports. The model package is only imported the first time something embeds.
- No CEF.
backend: "webview" uses the OS webview instead of bundling Chromium, which saves roughly 100 MB (chapter 1).
Even so, the Deno runtime alone is about 70 MB, and that's a floor this approach can't go below. For a tool whose job is "search my stars", that feels like a lot.
I suspect a Go shell with Wails (a single Go binary using the OS webview) could come in well under that. So I've started learning Go, and I'm looking forward to rebuilding a slice of this and writing up what I find.