Testing AI tooling for a living, writing about it in public.
I run structured evaluations of AI coding agents and data pipelines — the kind of work where the interesting part is never the demo, but what breaks on the third run. This site collects the method: task books, acceptance scripts, failure taxonomies, and the numbers that come out.
Latest
- What goes on this site, and what does not 2026-10-07
The scope of Field Notes, how often it updates, and why the numbers here have to come from a script.
Newsletter
One long read, roughly every two weeks. Signup is not wired up yet — the form lands here once the mailing list is set up.