All projects

Wick
Running; the code is private
Claude Opus, Claude CLI, Python, Git worktrees, SQLite, FastAPI, React, cron, Tailscale

A worker agent whose code changes wait for tests and my "ship it"

Wick answers in my assistant's chat and can read my whole server. Its code changes land on a branch, get tested, and merge only when my own message says ship it.

Draft

Data in or outA model proposesCode checks and decidesA person acts or approves

  1. InA message in my assistant's chat
  2. ModelWick reads the server and drafts a change
  3. CodeWrite guard puts it on a branch, away from live code
  4. CodeThe project's own tests run on that branch
  5. PersonI try a preview copy and type "ship it"
  6. CodeChecks my words, the commit and the tests, then merges
  7. OutLive app health-checked, reverted if it stops answering

The problem

I run a personal assistant on my own server: a phone app over my email, calendar, documents and photos, with a chat. Harder questions were going to outside chat models that couldn’t see the machine and left no record. I wanted an agent on the server that could look at everything and change code, without ever finding a change I hadn’t approved in an app I use every day. A rule in a prompt doesn’t settle that, because an agent working fast will reason its way past a sentence.

What it does

Wick is the agent behind the chat: when the first model call decides a message needs more than a reply, it hands the turn to Wick. Wick can read anything on the server. Apart from its own notes, scripts and documents it leaves for me, every change goes through one program, wick-write:

Every accepted write is committed, so every change has an undo.

Wick can run the project’s tests on its branch and start a preview: a copy of the assistant running from the branch on copied data, reachable only from my devices, that shuts itself down after three hours. To ship, I type “ship it” in the chat. Wick records my words and the exact commit, and a scheduled job picks it up within five minutes.

How it’s built

Wick is Claude, run through the Claude command-line tool with an explicit tool allowlist. Reading and searching are open. The built-in Edit and Write tools are left out, so the sanctioned way to change a file is the guard, and the guard’s own folder is protected.

The merge job ships an approval only if the approval is under two hours old, my quoted words appear in one of my own chat messages from the 30 minutes before it, the branch is still exactly the commit I approved, the tests pass again, and the live repository has no uncommitted changes. Then it merges with a single merge commit, restarts the service and checks that the app answers. If it doesn’t answer within 30 seconds, the merge is reverted. I hear about the outcome either way.

Judgment calls (direction, money, other people, security) go to the escalation log for a longer session to review.

Decisions

How it broke, and what changed

On the evening of 18 August, Wick couldn’t write anything. Four exchanges produced no log entry, no escalation and no undo. Every write came back as needing approval, and the guard’s audit log showed it had never been reached.

The cause was spelling. The command-line tool matches an allowlist entry against the literal text of a command. Wick’s instructions called the guard using the home-directory shortcut (a tilde), and the allowlist named it by its full path. The tilde isn’t expanded before matching, so the two never met, and in non-interactive mode a refused command fails quietly.

All the common spellings are allowed now. The broader change is how I read an allowlist: each entry permits one spelling of a command. Allowing git status doesn’t allow git -C <folder> status, so where a command has several common spellings, all of them are listed.

What’s still rough

How this was made: I wrote the specification, made the design decisions and tested the result. AI coding agents (Claude Code) wrote the code. More on how I work.