AI evaluation
How AI work gets measured: building an eval set and watching it rot, error rates in support bots, cost per task rather than per token.
- Measuring error in support AI, and what accuracy hides AI and engineering·2026-10-09 ·11 min read
Five kinds of error, three ways one accuracy figure misleads, and what a routing test on 3,894 public complaints actually returned.
- How an eval set is built, and why it rots AI and engineering·2026-10-09 ·10 min read
An eval set is an instrument with a purpose, a scoring rule, a calibration record and an expiry date. Where the samples come from, what counts as correct, and what makes the set expire.
- From token prices to task prices: three jobs, seven APIs AI and engineering·2026-10-09 ·12 min read
I priced three concrete jobs on seven model APIs using each vendor's own pricing page, and wrote down every assumption the conversion rests on.