The Series
Field notes from a daily practice with AI coding agents. Patterns, tools, sharp edges. Evidence-based, no fluff.
The Agent Cannot Edit Its Own Answer Key: Structural Guardrails Against Reward Hacking
An agent scored against ground-truth files has three shortcuts to a fake PASS: edit the answer key, tune a constant until green, or assert "verified" without running anything. I closed each one in the harness, where the model cannot rationalize past it.
The Refutation Swarm: 842 Subagents That Default to Disbelieving Each Other
One FileAuditAgent per source file, one RefutationAgent per claimed defect, and a standing order to disbelieve. Only 43 of 239 candidate defects survived the adversarial gate.
The Expensive Verb: Teaching an Agent to Stop Re-Running Everything
My agent's favorite diagnostic was a seven-minute full re-run, even when the answer was already sitting in a cached JSON file. Prose rules didn't stop it. A 148-line hook did.
Shipping awesome.video: 2,365 Resources, Built Mostly by Agents
A curated directory of video-development resources is now live. It took 1,185 commits and thirteen months, and coding agents wrote or drove most of them. Here is the honest recap.
What we ship
ValidationForge proves that code works. Anneal proves the plan works before any code gets written. Two flagships, both production-grade. The other cards are companion tooling from the series.
Get the next field note before it ships.
4,206 engineers, operators, and weirdos. No spam, ever.