All projects

Corpus index
Running; the code is private
Python, DuckDB, ONNX Runtime, nomic-embed-text v1.5, nomic-embed-vision v1.5, NumPy

Search by meaning across 18 years of my own records

An index over 763,186 dated records from 53 sources, searched by meaning with a local model. A new embedder has to reproduce the stored vectors before I trust it.

Draft

Data in or outA model proposesCode checks and decidesA person acts or approves

  1. InDated records from 53 sources, left where they are
  2. CodeIndex stores each record with a link to its original
  3. ModelLocal embedding model turns passages into vectors
  4. CodeA new embedder must reproduce stored vectors first
  5. InA question typed in plain words
  6. OutRanked passages, each pointing to its source

The problem

I have about 18 years of my own records spread across 53 sources and nine storage technologies, from database files to plain exports. Answering one question about one week used to take seven different ways of getting at the data. I wanted one place to ask, without putting any original at risk.

What it does

It is an index over 763,186 dated records from those 53 sources, from 2008 to 2026. Every row carries a reference back to where the record actually lives. On top of it, search by meaning covers about 891,000 passages, using an embedding model that runs on the server’s CPU with no network calls. The text and image models share one 768-dimensional space, so a typed description can rank photographs as well.

How it’s built

The index is DuckDB: one file, no server. The questions I ask it are analytical (counts by source, coverage by month), and there is one writer. Builds are atomic. A run writes to a temporary file and replaces the index only on success, so a crash leaves the previous index intact. The sources use six different timestamp conventions, and all of them go through one shared function that converts to a local day, because mixing them silently moves overnight activity by about five hours.

The embedder is nomic-embed v1.5 exported to ONNX, behind one small interface that everything calls. Stored vectors are 8-bit integers with a scale per row, which kept the top 20 results identical to full-precision vectors over the whole set, at a quarter of the memory.

Decisions

How it broke, and what changed

In September the old embedding service was replaced with the local ONNX model. Before trusting it, I had it re-embed about 40 texts whose stored vectors already existed and compared them. The bar is a median cosine of 0.999 or better, with the lowest score explained. It took four variants to match the main index. The model is asymmetric: stored text needs a “search_document:” prefix and questions a “search_query:” prefix. A note written five days earlier said no prefix on either side, and the sample showed the note was wrong.

The same comparison on my search-history sessions turned up something worse. The median looked perfect and the damage was all in the tail. The model’s vocabulary is uncased, so text has to be lowercased before it is split into tokens, and the old service’s build of the same model skipped that step. Any text with a capital letter had been embedded from partly wrong tokens. One word scored a cosine of 0.36 against its own lower-case spelling. That was about 16% of 18,853 sessions, for months, and nothing errored. Search just got quietly worse.

All 18,853 were re-embedded in the main index’s configuration, with the old vectors kept in a side column so the change could be undone. Checking the lowest scores, as well as the median, is now part of accepting any embedder.

What’s still rough

How this was made: I wrote the specification, made the design decisions and tested the result. AI coding agents (Claude Code) wrote the code. More on how I work.