Measured findings, including the ones that said no. Nulls published next to wins — the method is the product.
Null result 2026-07
We built skill checklists for our agents, believed in them, and measured them anyway. Across 88 arm-randomised scored executions, quality moved +0.002 — inside the noise. The null, the method that caught it, and the boundary on what it does not say.
Register 2026-07
A running record of independent practitioners, an academic study, and frontier labs arriving at the conclusions we built on — in their own words, with dates, sources and caveats. Updated as the record grows.
Essay 2026-06
We audited 36 mechanical quality checks and found about 20 had caught nothing in a month of production. The scaffolding you build for a weak model becomes a cage for a strong one.
Essay 2026-06
A GDPR-certified consent tool got a client fined anyway — because blocking a tracker doesn't make you invisible, it makes you a specific, reconstructable kind of missing data. On MNAR statistics, Bayesian imputation, and what enforcement actually catches.
Essay 2026-03
Per-job quality scores in multi-agent systems almost always land around 85%. That number is telling you something different than you think it is. 4,600+ scored executions from our own production, and the pattern we found.
Essay 2026-04
We ran a controlled A/B across 88 scored agent executions, with and without skill checklists in the prompt. Quality moved +0.002 — inside the noise. Why front-loaded instructions don't land, and what does.