All projects

obd-agent
Running; the code is private
Python, Claude Opus, Claude Code, Kotlin, Android, WebSocket, Tailscale, ELM327, pytest

An AI mechanic that picks the next test, inside limits it can't change

An agent reads my car's engine data over OBD-II and decides what to test next. Code checks every request twice, on the server and the tablet, before it reaches the car.

Draft

Data in or outA model proposesCode checks and decidesA person acts or approves

  1. InThe complaint, trouble codes and live engine readings
  2. ModelNames competing causes, proposes the test that splits them
  3. CodeServer guard checks rpm caps and read-only commands
  4. CodeTablet checks the command again against its whitelist
  5. PersonDriver holds the rpm the tablet asks for
  6. OutDiagnosis, the evidence, and the cheapest next step

The problem

My 2009 Infiniti G37 has an intermittent knock and a check-engine light. An OBD-II reader shows the numbers but can’t say which ones matter. A good mechanic forms a few explanations, then runs the one test that tells them apart. I wanted an agent that works that way and can’t do anything to the car a careful mechanic wouldn’t.

What it does

Watch mode is the default. A tablet in the car logs engine readings for the whole drive and stays quiet apart from three safety alerts. One button, MARK, records the moment the knock happens. After the drive, code computes the findings (fuel trims per bank, every timing pull of 4 degrees or more under load, the seconds around each MARK) and the model writes a report from those numbers only.

In diagnose mode the agent names its competing explanations and picks the test that separates them. If both banks run lean, a vacuum leak shrinks as rpm rises and a weak fuel pump gets worse under load, so it asks the driver to hold 2500 rpm. The tablet shows a gauge and records the trims while the needle is in the band.

How it’s built

The server half, in Python, runs the loop: the model proposes an action, a guard checks it, the tablet runs it, the raw reply is decoded. Claude Opus does the reasoning through the Claude command-line tool. The tablet half is a Kotlin Android app that talks Bluetooth to the reader and a WebSocket to the server over my private network.

Two Claude Code instances built the halves on two machines and never shared code, only a written protocol whose every change I approve.

To evaluate it, I had a simulated G37 built that answers in the raw bytes a real ELM327 adapter sends, with six hidden faults. Four of them produce the same two codes and only separate under the right test. The model never sees which fault is loaded. The first run found 5 of 6, calling an under-reading airflow sensor a fuel-supply problem. A second run with road tests off also got 5 of 6, and there the miss was a fault you can’t see without driving: it said “inconclusive” and named the road test that would settle it.

Decisions

How it broke, and what changed

On 2 October the after-drive step saved a faster watch plan that read rpm and speed only inside a combined request. The tablet refuses that, because its “engine running” and “car moving” checks read exactly the rpm and speed commands. The server’s guard had approved it. The guard asked whether rpm and speed were read somewhere in the plan; the tablet asked whether those exact commands were in it. Both rules came from the same sentence in the protocol, and they still disagreed.

The server treated the refusal as a failed session and hung up. The tablet reconnected, was offered the same plan and refused again: 1,824 connections in a row over more than two hours. The tablet kept logging on its old plan, so no data was lost, but no faster plan ever ran.

The guard now asks the tablet’s exact question, and a refused plan no longer ends the session: the server logs the reason and falls back to the default. The default had drifted too, with three copies in the code asking for two readings this car never answers, so a test now reads it straight out of the protocol document and compares.

What’s still rough

How this was made: I wrote the specification, made the design decisions and tested the result. AI coding agents (Claude Code) wrote the code. More on how I work.