Architecture & Method · What we run
BPOS: how we build software with agents.
And what decides what ships.
For a year we wrote about what we learned, not about the machine that taught us. This is the machine: how a goal becomes shipped code, what decides what ships, what it has built, and what it does not do.
BPOS is the platform Bouletteproof builds software with. A goal becomes a blueprint, the blueprint becomes sprints of small jobs, agents carry out each job in a developer's own container, every job is scored against its acceptance criteria, and the result either commits on its own or goes to a person.
Why we were quiet
Claims are cheap in this field and records are not. We wanted the record first: every execution scored, every correction dated, and what went wrong published next to what worked. That takes time to accumulate.
It has now. Since our records start in February 2026, BPOS has run more than 2,000 sprints and scored more than 19,000 agent executions, across 30 model identifiers (counts as of 9 October 2026; an identifier can be a version of the same model).
The machine
A goal is a sentence a person can check. BPOS turns it into a blueprint: the sprints needed, the jobs inside each sprint, and for each job a contract that states the objective, the constraints and the acceptance criteria.
Generating the blueprint creates its sprints. A person starts a sprint, and agents work its jobs inside a container that holds that developer's own copy of the code. When a sprint is one step of a longer program, finishing it can start the next step without anyone in between.
What decides what ships
Each job is scored by EQS (Execution Quality Scorer): the share of the job's acceptance criteria that a judge model finds met, from 0 to 1. Code that does not compile scores 0. Then the rules are mechanical:
- Tests failed their last run: the commit is blocked and the job goes to a person, with the failure in front of them.
- The graders disagree strongly: a person decides.
- The score is below 0.60: a person decides.
- Otherwise the job commits on its own. At the end of the sprint only approved jobs' files are committed, and every commit, pull and push is recorded.
A sprint can also ask for a person on every job, whatever the score. In the 30 days to 9 October 2026, 86 of 268 sprints did. The rest committed what cleared the bar without anyone looking. That is a choice made per sprint, and it is on the record.
What it has built
HikrLink, our link and click-attribution service, was built this way: 122 scored executions across 29 sprints and four repositories, between 11 March and 12 July 2026. The receipt, with the query that produces it and the corrections we made to it, is on its own page. Parts of this website went through BPOS too, including some of the articles on it.
What others found
In March 2026 AprilNEA, a 19-year-old engineer who founded ArcBox Labs, reverse-engineered the binary behind Claude Code Web. According to that teardown, it contains a platform Anthropic had not announced: a Go orchestration engine, authentication through a GitHub app, skills, sandboxed containers per session, a deploy step with quality-gate hooks, and a loop that learns from usage. Mandar Karhade's write-up in AI Advances took it to a wider audience.
That is the shape we had been building. Anthropic does not know we exist; this is independent convergence, not an endorsement.
There are two differences, one each way. By the teardown's account, their gate checks build-level problems before a session stops; ours also judges the acceptance criteria and blocks a commit when tests fail. Their containers are sandboxed per session; ours are per developer, so one developer's jobs share a workspace.
What it does not do
The score does not see the tests. A failed test run blocks the commit, but it does not lower the score. On a small 14-day sample in October, executions whose test run failed averaged 0.975 and those whose run passed averaged 0.980. Those scores feed what the platform learns, so a failure can be learned as a success. We have measured this. We have not fixed it yet.
It does not make the model irrelevant. On the evidence we have checked, the model explains most of the variation in results. The system around it is the second effect, and the one you can keep and correct. We changed our own claim when the evidence changed, in The Pigeon Was Not Superstitious.
It is not autonomous, and it is not supervised everywhere either. Two in three sprints commit unattended above the threshold, and programs move from one sprint to the next on their own. Where a person decides is a rule you can read, not a promise.
What comes next
This is the first of a series on what BPOS runs, one mechanism at a time, from the scale of a goal down to a single step. Each piece carries the outside work that arrived at the same place, says whether the mechanism is built, measured or still designed, and publishes the part that did not work.
- Plan. A goal has necessary stages: how a program carries them, and how BPOS Lite was built through one. Every job is a contract: from a blueprint to the frame each job receives.
- Execute. The mesh: how a job is worked by many small steps. What a step chooses to read before it acts. When to stop retrying, and why the answer is convergence, not a fixed number of attempts.
- Judge. What the gate checks, and what it does not. Where a person decides.
- Learn. A skillset is not a SKILL.md: tools, model, procedure and output schema declared in pages, not in code. Procedure outside the weights. Predicting before running.
The series closes on the control room: what a person sees, and decides, when the machine has learned something.
Common questions
What is BPOS?
The platform Bouletteproof builds software with. A goal becomes a blueprint of sprints and jobs, agents carry out the jobs, each job is scored against its acceptance criteria, and the result either commits on its own or goes to a person.
Does a person review every change?
No. A person decides when a job's tests failed, when the graders disagree, when the score is below 0.60, or when the sprint asks for a person on every job. In the 30 days to 9 October 2026, 86 of 268 sprints asked. Other work at or above the threshold commits on its own.
What is EQS?
EQS stands for Execution Quality Scorer. It scores a job from 0 to 1 as the share of its acceptance criteria that a judge model finds met. Code that does not compile scores 0. The score does not include the test result: a failed test run blocks the commit separately.
Sources
- AprilNEA (ArcBox Labs), reverse-engineering Claude Code Web and Antspace, March 2026.
- Mandar Karhade, "Anthropic's Antspace: The Secret PaaS Nobody Was Supposed to Find", AI Advances, March 2026.
- Our own records: sprint, execution and model counts as of 9 October 2026; the share of sprints with a person on every job, 30 days to 9 October 2026; the test-run sample, 14 days to 4 October 2026.
Related writing
- HikrLink, built with BPOS. The receipt for one product, with its corrections.
- The Pigeon Was Not Superstitious. The Reader Was.. How the platform learns, and the evidence against us.
- A Green Build Is Not Knowledge. Why a passing build is not the same as a job done.
- Independent Convergence. The running record of outside work that arrived where we are.