01 · The chassis
The Harness
The frame everything bolts onto: which ticket, which branch, which workspace, which gate. It never improvises, which is exactly why you can trust the parts that do.
never: improvises
■ Software factories · AI Native Orgs
That is the question we hear on almost every first call with a CTO. The licences went in, the demos were great, and six months later the sprint curve looks the same. You are not imagining the gap. It took us a year of getting it wrong to see what closes it: a software factory.
The pattern has surveys behind it. Nine in ten professional developers use an AI coding agent every week (JetBrains, AI Coding Agents: Adoption Trends, August 2026, 15,000-plus developers). Six per cent of companies can attribute at least five per cent of their earnings to AI (McKinsey, The State of AI: Global Survey 2026, August 2026, 1,719 respondents).
Does this sound familiar?
You approved the AI budget because the numbers were obvious. Pull requests per engineer went up. Then the review queue went up faster. Then an incident at 2am traced back to a change nobody had fully read. Now someone senior is saying "let's slow down until we have guardrails", and the budget that bought speed is buying caution.
"PRs doubled and I cannot tell which ones anyone actually read."
"Our tests are green and I no longer believe them."
"Small changes cross four repos, and the map of why is in two people's heads."
"The engineer who set up our agents is on leave and everything got slower."
This page is for CTOs and founders with fifty or more people in product and engineering.
The software factory
Because you bought a faster engine and never touched the brakes. AI agents made writing code cheap. Nothing after the writing got cheaper: reading it, testing it, knowing it works. So the first thing a sensible team does when the car starts to feel unsafe is slow down. That is the brake you are feeling. It is a management brake, applied by hand, because the car does not have one of its own.
The car's own brake is a validation layer: every requirement turned into a test a person can read and a machine can run, running on every change. With it, you can drive fast. Without it, you are choosing between fast and safe every sprint, and safe wins.
That brake is one of the four parts of a software factory. Here is the whole car.
01 · The chassis
The frame everything bolts onto: which ticket, which branch, which workspace, which gate. It never improvises, which is exactly why you can trust the parts that do.
never: improvises
02 · The engine
One job each. A builder builds, a validator checks, a reviewer reviews and writes nothing. An agent that can do anything is one you cannot review.
never: two jobs in one agent
03 · The map
What your system actually is, written down with a source on every line. Skip it and every agent guesses your architecture again, differently each time.
never: written by the thing it describes
04 · The brakes
How you know it worked. Nothing leaves the software factory without passing through. This is the piece we give away.
never: a test nobody can trace to a requirement
Most teams build the engine and the chassis. They are fun, they demo well, and every vendor sells them.
Nobody sells you the map or the brakes, because neither one demos. That is the gap we build for.
Which of the four do you already have? Ten questions will tell you.
A fair question
This year Plum asked us the same thing. Employee health insurance for companies, Series B, about six hundred people with sixty in product and engineering, fifty-six repositories, engineers already better at AI than most. Their baseline showed one in five units of engineering effort going to rework. Six weeks later the software factory was theirs. Here is what "theirs" meant.
Plum · employee health insurance · Series B · about 600 people, 60 in product and engineering · 56 repositories
Six weeks, 2026
Trusted by teams at

The practical question
About six weeks, and less of your team's attention than the last AI initiative. The engagement is four dones, written down with acceptance criteria before we start. It is complete when the last one is checked, and that is what you pay for: the dones, not the days.
D1
We agree the pod, the repos a change touches, who approves what, and we measure what your test suite proves today. It is usually less than the dashboard says.
D2
The four pieces, built on your Jira, your CI, your conventions. Not a platform you adopt. A system that fits.
D3
Real tickets go through. The suite is a required check. Failures teach the harness.
D4
Your people run it without us. The first occurrence of anything new gets solved with you and written into the playbook. Repeats are yours.
Because you are buying a working software factory, not hours. Each done has a written acceptance criterion, and payment is tied to those. No FTE, no open-ended retainer. If a done takes longer than planned, that is our problem, not your invoice.
The fourth done is your team running the factory without us. It is in the contract, not promised at the end.
Book a conversationThe obvious objection
Probably, for some of it
Buy the runners, the model and the dashboards. We did. Nobody should build compute.
Not for the parts that encode how you decide
Your gates, your specs, your context, your tests. A product cannot know what a fact about your system is, or which flow crosses which service. When a change spans three repos, this is where products stop.
Score your organisation on the eight criteria below. If two or more land on the right, you are building the judgment layer whatever you buy underneath it.
| Criterion | Leans product | Leans custom Yours to build |
|---|---|---|
| Repositories a typical change touches | One | Three or more, in a merge order that matters |
| Ticketing and CI already in place | None, or a standard SaaS you would happily replace | Jira and GitHub Actions with conventions the team relies on |
| Domain conventions | A generic web application | Regulated or combinatorial: insurance configurations, ledgers, consent |
| Who owns the suites and agents afterwards | Living in the vendor's organisation is acceptable | Must live in your organisation, in your repositories, under your access |
| Approval boundaries | One gate for everything | Policy by author, area and risk, and the policy changes |
| Validation depth needed | UI journeys are enough | Cross-service back-end flows with seeded data and stubbed vendors |
| Metrics you must report | Vendor defaults are acceptable | Your definitions, traceable to a ticket and a PR |
| Platform appetite | None. Nobody will keep a harness alive | A small team that will keep it thin and improving |
Every row on the right is a place where a product has to guess about your organisation. Every row on the left is a place where guessing is fine. If you want a second opinion on your scores, that is a thirty-minute call.
Book a conversationThe follow-up
Everything we say about software factories starts on LinkedIn. The pieces worth reading in full, in order.
Code is becoming the new assembly language. When AI writes the code, your primary task is no longer syntax. It is context and validation.
In an AI-native model, code is disposable. The true artifact is the context: the spec, the architecture, the past review comments, the trade-offs.
AI agents can behave like humans under pressure, preferring a clean story over uncertainty, unless you force a verification contract. Contracts, not vibes.
■ Start here
Ten questions about how your team uses AI today. Five minutes. You get a paradigm level and, more usefully, a view of which of the four pieces of a software factory you already have. Free, and a person reads every result.
Not ready for the assessment?
The 8 paradigms, the risks at each level, the structural shifts between them. Emailed to you. Written on the earlier version of the model; the assessment uses the current seven.