All projects

Sovereign
Running; the code is private
Python, FastAPI, TOTP two-factor, Fernet, HMAC-SHA256, FHIR, Synthea, pytest

De-identifying clinical notes: the model points, code removes

Strips patient identifiers from clinical text on my own server. Rules go first, a model only points at what they missed, and code makes every removal.

Draft

Data in or outA model proposesCode checks and decidesA person acts or approves

  1. InA clinical note and the identifiers already on file
  2. CodeRules replace every known identifier with a tag
  3. CodeOutbound gate blocks anything identifier-shaped
  4. ModelLists leftover spans, from 7 allowed categories
  5. CodeDrops bad spans, redacts the rest locally
  6. CodeStricter release gate refuses if any remain
  7. OutA redacted note, or a refusal

The problem

Clinical notes are useful to an AI model and dangerous to send to one. Rules handle structured fields well: a name, a birth date, a health-card number. Free text defeats them. “She still sees Dr. Halloran, and her sister Mei drives her” holds two identifiers that no field list or regular expression will catch, because what makes “Mei” an identifier is the words around it. Finding them is a language problem, so a language model is the right tool for the search. It is the wrong tool to trust with the result.

What it does

Sovereign removes identifiers on my own server so an outside model can work with a note without seeing who it is about.

The summary path is reversible: known identifiers become tags such as [PATIENT_NAME_ab12cd], the tagged text goes out for a summary, and the summary is re-identified locally from an encrypted map.

The release path is for text that leaves for good, so nothing is kept that could reverse it. Rules replace everything they know. The model reads the tagged text and returns only a list of substrings it thinks are leftover identifiers, each with a category. It never writes the document. Code applies the redactions, and a final gate refuses the release if anything identifier-shaped survives.

Underneath is a batch engine for structured patient records, built around the Ontario privacy commissioner’s 2025 de-identification guidelines: identifiers removed or hashed with a secret salt, dates shifted per patient, age and region generalised until every group of look-alike records is big enough. On 1,473 synthetic records the smallest group held 94 and the estimated re-identification risk was 0.053%.

How it’s built

The web app is FastAPI with mandatory two-factor login, server-side roles, encrypted re-identification maps and an audit log that records categories, never matched text. Everything bound for the model passes an outbound gate that looks for the document’s own known identifiers and for identifier-shaped patterns. A hit stops the call.

The batch engine is standard-library Python under two Unix accounts. My coding agent works in one and never sees the other, which owns the sealed area for real records and the hashing salt. The only bridge is one approved admin command that takes no arguments, so there is nothing to inject.

Everything has run on synthetic data only.

Decisions

How it broke, and what changed

It broke twice, both times where an identifier turned up in a form the code didn’t expect.

In June the released output was called leak-free after a check for a fixed list of fields: name, birth date, postal code, city, coordinates. In July an audit ran the pipeline on fresh synthetic records and found each record’s mother’s maiden name and birthplace passing straight through, in a part of the record (FHIR extensions) the code never read. The fix inverted the rule: every extension is dropped unless it is on an allowlist, and the allowlist is empty. It was checked by running the pipeline over 30 records and inspecting the output.

In August a realistic assessment note, written as a test document, leaked on its first run. The stored identifier was the fictional patient’s full name, replacement matched whole strings, and the note used the first name on its own throughout. No gate pattern matches a bare first name, so it reached the outside model in the clear. Name fields now add their first-name and surname parts to the map, with guards: parts of four letters or more, whole words only, so a surname like Fisher can’t rewrite “Fisherman”.

What’s still rough

How this was made: I wrote the specification, made the design decisions and tested the result. AI coding agents (Claude Code) wrote the code. More on how I work.