■ Software factories · AI Native Orgs

Every engineer on your team is faster. Why isn't the team?

That is the question we hear on almost every first call with a CTO. The licences went in, the demos were great, and six months later the sprint curve looks the same. You are not imagining the gap. It took us a year of getting it wrong to see what closes it: a software factory.

Individual output is up Team output is flat The fix is a software factory, not a tool
outputlicences insix months laterThe gapIndividual outputIndividual outputTeam outputTeam output
Individual output Team output The gap Fig. 0 · The shape of the last six months on most teams. Not a measurement; the pattern.

The pattern has surveys behind it. Nine in ten professional developers use an AI coding agent every week (JetBrains, AI Coding Agents: Adoption Trends, August 2026, 15,000-plus developers). Six per cent of companies can attribute at least five per cent of their earnings to AI (McKinsey, The State of AI: Global Survey 2026, August 2026, 1,719 respondents).

Does this sound familiar?

Does this sound like your last engineering review?

You approved the AI budget because the numbers were obvious. Pull requests per engineer went up. Then the review queue went up faster. Then an incident at 2am traced back to a change nobody had fully read. Now someone senior is saying "let's slow down until we have guardrails", and the budget that bought speed is buying caution.

  • 01

    "PRs doubled and I cannot tell which ones anyone actually read."

  • 02

    "Our tests are green and I no longer believe them."

  • 03

    "Small changes cross four repos, and the map of why is in two people's heads."

  • 04

    "The engineer who set up our agents is on leave and everything got slower."

This page is for CTOs and founders with fifty or more people in product and engineering.

The software factory

Why did we speed up and then hit the brakes?

Because you bought a faster engine and never touched the brakes. AI agents made writing code cheap. Nothing after the writing got cheaper: reading it, testing it, knowing it works. So the first thing a sensible team does when the car starts to feel unsafe is slow down. That is the brake you are feeling. It is a management brake, applied by hand, because the car does not have one of its own.

The car's own brake is a validation layer: every requirement turned into a test a person can read and a machine can run, running on every change. With it, you can drive fast. Without it, you are choosing between fast and safe every sprint, and safe wins.

That brake is one of the four parts of a software factory. Here is the whole car.

01 · The chassisThe Harness03 · The mapThe Context Layer02 · The engineThe Agents04 · The brakesThe Validation Layer
Fig. 1 · A software factory drawn as a car. The brakes are the part most teams never build. Hover to see it move.

01 · The chassis

The Harness

The frame everything bolts onto: which ticket, which branch, which workspace, which gate. It never improvises, which is exactly why you can trust the parts that do.

never: improvises

02 · The engine

The Agents

One job each. A builder builds, a validator checks, a reviewer reviews and writes nothing. An agent that can do anything is one you cannot review.

never: two jobs in one agent

03 · The map

The Context Layer

What your system actually is, written down with a source on every line. Skip it and every agent guesses your architecture again, differently each time.

never: written by the thing it describes

A gift

04 · The brakes

The Validation Layer

How you know it worked. Nothing leaves the software factory without passing through. This is the piece we give away.

never: a test nobody can trace to a requirement

Most teams build the engine and the chassis. They are fun, they demo well, and every vendor sells them.

Nobody sells you the map or the brakes, because neither one demos. That is the gap we build for.

Take the assessment

Which of the four do you already have? Ten questions will tell you.

A fair question

Has anyone actually done this, or is this a deck?

This year Plum asked us the same thing. Employee health insurance for companies, Series B, about six hundred people with sixty in product and engineering, fifty-six repositories, engineers already better at AI than most. Their baseline showed one in five units of engineering effort going to rework. Six weeks later the software factory was theirs. Here is what "theirs" meant.

Plum · employee health insurance · Series B · about 600 people, 60 in product and engineering · 56 repositories

Six weeks, 2026

  • A Jira button. Press it and a tech spec arrives, written against the real code.
  • Agents build one pull request per repository, in an order where every partial merge is safe.
  • Your people approve at the gates you chose. Three rules did more work than any prompt: humans merge, the factory never does; an approved spec is never rewritten silently; every run ends by pushing everything it did.
  • A system-integration suite runs on every change as a required check. That is the brake.
  • One console: what is running, what it cost, factory tickets against manual tickets.
  • Every developer got it in their editor. One plugin, and their coding agent knew how the company works.
  • Token cost known before the month starts: three flat-rate subscriptions at $100 each, $300 a month for the pilot team, fixed. A harder ticket is not a bigger bill.
56
repositories the context layer tracks, updated with their latest changes
~1,040
API operations reconstructed from source, each with a file and line
223
automated tests on the onboarding journey, in a suite built from scratch. Before: 180 cases in a spreadsheet
80 min
fastest ticket to a code PR with tests green
~6 min
for that suite on every pull request, on an ephemeral environment in the cloud
20
tickets end to end into production by week ten

Trusted by teams at

Plum logoIIFL logo5paisa logoBizSherpa.ai logo

The practical question

What would this take from my team, and how long?

About six weeks, and less of your team's attention than the last AI initiative. The engagement is four dones, written down with acceptance criteria before we start. It is complete when the last one is checked, and that is what you pay for: the dones, not the days.

D1

Scope and baseline

We agree the pod, the repos a change touches, who approves what, and we measure what your test suite proves today. It is usually less than the dashboard says.

D2

The factory built

The four pieces, built on your Jira, your CI, your conventions. Not a platform you adopt. A system that fits.

D3

Integrated and gated

Real tickets go through. The suite is a required check. Failures teach the harness.

D4

Your team owns it

Your people run it without us. The first occurrence of anything new gets solved with you and written into the playbook. Repeats are yours.

Why dones and not a day rate?

Because you are buying a working software factory, not hours. Each done has a written acceptance criterion, and payment is tied to those. No FTE, no open-ended retainer. If a done takes longer than planned, that is our problem, not your invoice.

The fourth done is your team running the factory without us. It is in the contract, not promised at the end.

Book a conversation

The obvious objection

Should we just buy a factory product instead?

Probably, for some of it

Buy the runners, the model and the dashboards. We did. Nobody should build compute.

Yours to build

Not for the parts that encode how you decide

Your gates, your specs, your context, your tests. A product cannot know what a fact about your system is, or which flow crosses which service. When a change spans three repos, this is where products stop.

Score your organisation on the eight criteria below. If two or more land on the right, you are building the judgment layer whatever you buy underneath it.

CriterionLeans productLeans custom Yours to build
Repositories a typical change touchesOneThree or more, in a merge order that matters
Ticketing and CI already in placeNone, or a standard SaaS you would happily replaceJira and GitHub Actions with conventions the team relies on
Domain conventionsA generic web applicationRegulated or combinatorial: insurance configurations, ledgers, consent
Who owns the suites and agents afterwardsLiving in the vendor's organisation is acceptableMust live in your organisation, in your repositories, under your access
Approval boundariesOne gate for everythingPolicy by author, area and risk, and the policy changes
Validation depth neededUI journeys are enoughCross-service back-end flows with seeded data and stubbed vendors
Metrics you must reportVendor defaults are acceptableYour definitions, traceable to a ticket and a PR
Platform appetiteNone. Nobody will keep a harness aliveA small team that will keep it thin and improving

Every row on the right is a place where a product has to guess about your organisation. Every row on the left is a place where guessing is fine. If you want a second opinion on your scores, that is a thirty-minute call.

Book a conversation

The follow-up

Tell me more.

Everything we say about software factories starts on LinkedIn. The pieces worth reading in full, in order.

■ Start here

Where do I start
on Monday?

Ten questions about how your team uses AI today. Five minutes. You get a paradigm level and, more usefully, a view of which of the four pieces of a software factory you already have. Free, and a person reads every result.

Take the assessment ~10 questions · 5 minutes · free

Not ready for the assessment?

Start with the field guide.

The 8 paradigms, the risks at each level, the structural shifts between them. Emailed to you. Written on the earlier version of the model; the assessment uses the current seven.

We'll email it to you. No on-page download.