The problem
My 2009 Infiniti G37 has an intermittent knock and a check-engine light. An OBD-II reader shows the numbers but can’t say which ones matter. A good mechanic forms a few explanations, then runs the one test that tells them apart. I wanted an agent that works that way and can’t do anything to the car a careful mechanic wouldn’t.
What it does
Watch mode is the default. A tablet in the car logs engine readings for the whole drive and stays quiet apart from three safety alerts. One button, MARK, records the moment the knock happens. After the drive, code computes the findings (fuel trims per bank, every timing pull of 4 degrees or more under load, the seconds around each MARK) and the model writes a report from those numbers only.
In diagnose mode the agent names its competing explanations and picks the test that separates them. If both banks run lean, a vacuum leak shrinks as rpm rises and a weak fuel pump gets worse under load, so it asks the driver to hold 2500 rpm. The tablet shows a gauge and records the trims while the needle is in the band.
How it’s built
The server half, in Python, runs the loop: the model proposes an action, a guard checks it, the tablet runs it, the raw reply is decoded. Claude Opus does the reasoning through the Claude command-line tool. The tablet half is a Kotlin Android app that talks Bluetooth to the reader and a WebSocket to the server over my private network.
Two Claude Code instances built the halves on two machines and never shared code, only a written protocol whose every change I approve.
To evaluate it, I had a simulated G37 built that answers in the raw bytes a real ELM327 adapter sends, with six hidden faults. Four of them produce the same two codes and only separate under the right test. The model never sees which fault is loaded. The first run found 5 of 6, calling an under-reading airflow sensor a fuel-supply problem. A second run with road tests off also got 5 of 6, and there the miss was a fault you can’t see without driving: it said “inconclusive” and named the road test that would settle it.
Decisions
- Safety limits live in code, in two places. The server’s guard caps rpm and refuses wide-open throttle, stall tests and clearing codes, and the tablet separately refuses anything outside a read-only whitelist, so a bug on one side can’t reach the car alone.
- I didn’t tune the prompt to fix the missed fault. The simulator is mine, so tuning to it would only teach the agent my simulator.
- Code computes the findings and the model explains them. The numbers are reproducible, and if the model fails the report still goes out with them.
- The tablet keeps every reading until the server confirms it’s on disk. Resending is harmless, so a drive with no signal still arrives whole, later.
How it broke, and what changed
On 2 October the after-drive step saved a faster watch plan that read rpm and speed only inside a combined request. The tablet refuses that, because its “engine running” and “car moving” checks read exactly the rpm and speed commands. The server’s guard had approved it. The guard asked whether rpm and speed were read somewhere in the plan; the tablet asked whether those exact commands were in it. Both rules came from the same sentence in the protocol, and they still disagreed.
The server treated the refusal as a failed session and hung up. The tablet reconnected, was offered the same plan and refused again: 1,824 connections in a row over more than two hours. The tablet kept logging on its old plan, so no data was lost, but no faster plan ever ran.
The guard now asks the tablet’s exact question, and a refused plan no longer ends the session: the server logs the reason and falls back to the default. The default had drifted too, with three copies in the code asking for two readings this car never answers, so a test now reads it straight out of the protocol document and compares.
What’s still rough
- The knock that started this hasn’t been located in a log. The first one fell in a 208-second gap when the original version read nothing between requests, which is why watch mode exists.
- Per-cylinder misfire counters and freeze-frame data aren’t built.
- The server half is started by hand.