Engineering note 2026-07
Our platform ran a night of maintenance on itself and produced five defects, zero of them model failures. The worker declared the job unsatisfiable three times before failing. Every warning was written down correctly, on time, and read by nothing.
Null result 2026-07
We built skill checklists for our agents, believed in them, and measured them anyway. Across 88 arm-randomised scored executions, quality moved +0.002 — inside the noise. The null, the method that caught it, and the boundary on what it does not say.
Register 2026-07
A running record of independent practitioners, an academic study, and frontier labs arriving at the conclusions we built on — in their own words, with dates, sources and caveats. Updated as the record grows.
Essay 2026-06
We bet the system around the model matters more than the model itself. Over the last few months, five independent practitioners and a frontier lab arrived at the same conclusion — none of them ours.
Essay 2026-06
We audited 36 mechanical quality checks and found about 20 had caught nothing in a month of production. The scaffolding you build for a weak model becomes a cage for a strong one.
Essay 2026-06
A GDPR-certified consent tool got a client fined anyway — because blocking a tracker doesn't make you invisible, it makes you a specific, reconstructable kind of missing data. On MNAR statistics, Bayesian imputation, and what enforcement actually catches.
Library · MIT 2026-04
Lazy skill loading for agent systems. Claude Skills are useful when loaded on demand, expensive when loaded up front. context-steward is the npm package that solves the second problem without breaking the first.
Essay 2026-03
Per-job quality scores in multi-agent systems almost always land around 85%. That number is telling you something different than you think it is. 4,600+ scored executions from our own production, and the pattern we found.
Essay 2026-04
We ran a controlled A/B across 88 scored agent executions, with and without skill checklists in the prompt. Quality moved +0.002 — inside the noise. Why front-loaded instructions don't land, and what does.
Essay 2026-05
Multi-agent delivery doesn't break on the model. It breaks on file conflicts, bad sequencing, and per-job metrics that miss the deployment. The unglamorous machinery that decides whether a sprint ships.