The problem
I have about 18 years of my own records spread across 53 sources and nine storage technologies, from database files to plain exports. Answering one question about one week used to take seven different ways of getting at the data. I wanted one place to ask, without putting any original at risk.
What it does
It is an index over 763,186 dated records from those 53 sources, from 2008 to 2026. Every row carries a reference back to where the record actually lives. On top of it, search by meaning covers about 891,000 passages, using an embedding model that runs on the server’s CPU with no network calls. The text and image models share one 768-dimensional space, so a typed description can rank photographs as well.
How it’s built
The index is DuckDB: one file, no server. The questions I ask it are analytical (counts by source, coverage by month), and there is one writer. Builds are atomic. A run writes to a temporary file and replaces the index only on success, so a crash leaves the previous index intact. The sources use six different timestamp conventions, and all of them go through one shared function that converts to a local day, because mixing them silently moves overnight activity by about five hours.
The embedder is nomic-embed v1.5 exported to ONNX, behind one small interface that everything calls. Stored vectors are 8-bit integers with a scale per row, which kept the top 20 results identical to full-precision vectors over the whole set, at a quarter of the memory.
Decisions
- Index everything, move nothing. A migration makes the new copy authoritative, so losing it loses data, while an index is derived and deleting it costs a rebuild.
- The configuration is part of the index. Model version, task prefix, tokenizer behaviour, pooling and normalisation all have to match for two vectors to be comparable, and a mismatch fails silently.
- Full-precision weights, compressed storage. The quantised version of the model drifted to a median cosine of 0.95, so I kept the full model and compressed the stored vectors instead, after checking the rankings held.
- Photos rank separately from text. Text-to-image cosines run around 0.05 to 0.10 while text-to-text runs 0.6 to 0.8, so one merged list on raw cosine would bury every photo.
How it broke, and what changed
In September the old embedding service was replaced with the local ONNX model. Before trusting it, I had it re-embed about 40 texts whose stored vectors already existed and compared them. The bar is a median cosine of 0.999 or better, with the lowest score explained. It took four variants to match the main index. The model is asymmetric: stored text needs a “search_document:” prefix and questions a “search_query:” prefix. A note written five days earlier said no prefix on either side, and the sample showed the note was wrong.
The same comparison on my search-history sessions turned up something worse. The median looked perfect and the damage was all in the tail. The model’s vocabulary is uncased, so text has to be lowercased before it is split into tokens, and the old service’s build of the same model skipped that step. Any text with a capital letter had been embedded from partly wrong tokens. One word scored a cosine of 0.36 against its own lower-case spelling. That was about 16% of 18,853 sessions, for months, and nothing errored. Search just got quietly worse.
All 18,853 were re-embedded in the main index’s configuration, with the old vectors kept in a side column so the change could be undone. Checking the lowest scores, as well as the median, is now part of accepting any embedder.
What’s still rough
- It runs on my personal records, so there is no public demo. One would need a synthetic corpus.
- A 3D map of my search history (clusters found with HDBSCAN, projected with UMAP) was last built in June, before the re-embedding, and hasn’t been rebuilt from the corrected vectors.
- The record and passage counts can only be re-derived on my server. A weekly check does that (see the verification case file), but a reader has to take its word.