Architecture & Method · Independent convergence
Clarify before you build.
Then let the work correct it.
Google's new SDLC paper says the work is moving to specification and verification. A second Google paper says a hand-written specification can make an agent worse. Both are right, and the gap between them is where we work.
Spec-first agentic engineering is a way of building software where the specification, the tests and the quality checks are written before an AI agent writes the code, so the first implementation is close to correct and every change is verified before it ships. It only pays if the specification is corrected by what happens when the agent uses it.
What Google says
In May 2026 Google published The New SDLC With Vibe Coding (Addy Osmani, Shubham Saboo, Sokratis Kartakis; 51 pages). It was the Day 1 reading for Google's 5-Day AI Agents Intensive on Kaggle, 15–19 June 2026.
The paper puts AI-assisted development on a spectrum: vibe coding (accept what the model gives you) → structured AI-assisted work → agentic engineering (specs, architecture docs, tests, CI gates, evals, guardrails). Where you sit depends on the stakes. A weekend prototype and a payment flow do not get the same process.
Four of its points matter here:
- Agent = Model + Harness. The model is one component. The instructions, tools, sandbox, gates and observability around it decide whether it produces useful work.
- Tests and evals are a contract with the AI. Write them before implementation. They define what "correct" means.
- Output evaluation is not trajectory evaluation. Checking the final artifact is not the same as checking whether the agent took sensible steps to get there.
- The factory model. The developer's output is no longer code. It is the system that produces code, and the quality bar that system has to clear.
The paper's own summary of the shift: How much can we clarify before implementation so that the first implementation is already close to correct?
What a second Google paper says
Four months later, researchers at Google, Georgia Tech and Peking University published Procedural Graphs (Lu, Chen, Wu, Arık, September 2026). On one benchmark an agent with no guidance scored 87.50. Given an expert's hand-written map of the task, it scored 58.93. The score recovered, to 92.86, only when the map could change during use, each change kept only if a held-out test did not fall.
We wrote about that paper in A Hand-Written Map Can Be Worse Than No Map. The point stands on its own: the value was never in the map. It was in the loop that corrects it.
Read together, the two papers say one thing. Clarify first, and then let the work correct the clarification. A specification nobody corrects is a liability that sounds authoritative.
What we run
The same shape, in production, since 2025.
- Every job starts as a contract. Objective, constraints, acceptance checks. The agent's task is to satisfy the contract. The review question is "did it do what was asked", not "does this look fine".
- Pass or fail against the request, before anything ships. A green build is a compilation event. Verification is a check against what was actually asked, with a reason you can point to. A Green Build Is Not Knowledge is the long version.
- After a pass, we ask why. Each completed run gets an explanation step: what worked, the evidence, what to reuse. A run can pass for the wrong reason, and a loop that only studies failures never finds out.
- Context is pulled per job, not dumped. Relevance-gated, with a token budget. Google calls this static versus dynamic context. The Procedural Graphs paper measured it: showing the agent two steps around its position beat showing the whole map by 27 points at 70% fewer tokens.
- Nobody types the weights. What the platform prefers moves only from what happened to it. Operator judgement enters as an observation, counted rather than imposed. The rule is written out here.
Where Google is ahead of us
Scale and publication. We are a small studio; our evidence is our own runs. Google's paper ships with a course, a vocabulary, and a company that builds most of its software this way.
Where we go further, and where we don't
We study the green runs. Their factory feeds failures back. We also record why a successful run succeeded, because a pass with no reason is a coin that landed heads.
Our correction loop is still open. The part of our system that proposes corrections to our own knowledge pages runs at volume. The part that measures whether a correction helped has measured six. Until that closes, our hand-written pages are a hypothesis, not an asset. We said so in September and it is still true.
We do not claim the harness beats the model. On the evidence we have checked, the model explains most of the variation in results. Structure is the consistent second effect, and the one you can keep, inspect and correct. Google's "Agent = Model + Harness" is a statement about components, not proportions, and we read it that way.
The part that did not work
Clarifying up front is a cost, and sometimes it exceeds the saving.
METR's July 2025 study put 16 experienced open-source developers on 246 tasks in their own mature codebases with early-2025 tools. With AI they were 19% slower, while believing they were 20% faster. On a mature system with an expert in the seat, writing the spec and checking the agent can cost more than typing the change.
Our experience matches. Contracts pay on migrations, test coverage, integrations and repeated change types. They do not pay on one-off exploratory work, and we don't use them there.
The other failure is quieter: if the contract is wrong, the agent builds the wrong thing faster. Once implementation is cheap, the specification is where the risk lives. That is the Procedural Graphs result in one sentence.
The same question, at the scale of a company
Google's question is about code. It applies to the whole business. What do we actually know? Which rules apply? Who decided, and is that decision still current? When the answers live in people's heads, every agent and every new hire starts from zero.
A wiki is a hand-written map. The Knowledge Layer is the corrected one: you record a claim (a fact, a rule or a decision), look up what is current, and record who decided and what it replaced. Claims get superseded, not overwritten. It is what the contract is to the agent, at the scale of the company, with the correction loop built in.
Common questions
Is vibe coding bad?
Not for prototypes, throwaway tools and experiments. Google's paper says the same. It becomes a problem when code nobody read goes into a system people rely on.
Should we write a detailed specification before letting an agent build?
Write the acceptance check first: what "done" looks like in a form a test can verify. Keep the map short and let the results correct it. In the Procedural Graphs paper, an uncorrected expert map scored below no map at all.
Does a better model fix bad agent output?
Partly. The model explains most of the variation on the evidence we have checked. The system around it is the second effect, and the one you can keep and correct.
Sources
- Osmani, Saboo, Kartakis, The New SDLC With Vibe Coding, Google, 2026 (51 pp.). Day 1 reading of the 5-Day AI Agents Intensive on Kaggle, 15–19 June 2026.
- Lu, Chen, Wu, Arık, Procedural Graphs: Self-Evolving Execution Structures for LLM Agents, arXiv:2609.09153, 8 September 2026.
- Rahul Gite, "What Google's New SDLC Paper Got Right About AI-Assisted Software Development", Level Up Coding, 28 September 2026.
- METR, July 2025 developer productivity study, and the February 2026 design update.
Related writing
- A Hand-Written Map Can Be Worse Than No Map. The correction loop matters more than the map.
- A Green Build Is Not Knowledge. Verification is a check against what was asked, with a reason.
- The Pigeon Was Not Superstitious. The Reader Was.. The learning rule, and the evidence against us.
- A Decision Is Not a Fact. Why a contract is a decision, and what that changes.
- Independent Convergence. The running record this entry joins.