Field Notes

Testing AI tooling for a living, writing about it in public.

I run structured evaluations of AI coding agents and data pipelines — the kind of work where the interesting part is never the demo, but what breaks on the third run. This site collects the method: task books, acceptance scripts, failure taxonomies, and the numbers that come out.

Latest