Skip to content
Founders, CTOs and engineering leads

More code is arriving. Fewer people can vouch for it.

You gave the team AI tooling and output went up straight away. Then the review queue got longer. Then a bug came back from a file nobody had touched. Someone said the word “refactor” and everyone looked at their shoes.

Nobody on your team is doing anything wrong, and the model is not broken. You are running into something structural, and it has a shape.

Responsible for a system you did not build? That has its own page, and so does software an AI tool built.

You are not imagining it, and it is not only you.

Every one of these was published by someone else in 2026. We have linked each so you can check it rather than take our word.

1.7×

more issues in pull requests containing AI-assisted code than in human-written code

arXiv preprint 2603.28592, March 2026
30–41%

rise in technical debt within six months of widespread AI tool adoption

Innovative Group, 2026
45%

of AI-generated code introduced a security vulnerability

Innovative Group, 2026
4×

maintenance costs by year two, for teams that do not actively manage the debt

Innovative Group, 2026

Published figures, directional rather than audited. Three of the four come from the same 2026 analysis; the first is a separate preprint. Every one is linked so you can weigh it yourself.

Maintenance cost compounds to roughly four times baseline by year two when AI-generated code is not actively managed.adopt6 monthsyear 1year 2managed: roughly flatunmanaged: ~4× by year twolooks fine here
Published 2026 industry figure, directional rather than audited. The first months look fine — which is the problem: the curve is flattest exactly when you are deciding whether to worry.
Why it happens

Two ceilings, and neither is about the model.

We wrote these in July 2026, from running a factory past 165,000 lines of our own product. The figures above are other people’s — check them rather than take our word. We are not claiming we got there first: most of those sources carry a year and no month, and one is a March 2026 preprint, so the ordering is not something you could verify if we asserted it.

Your reviewer is your throughput

AI can 10× how much code gets written. It cannot 10× how much a human can safely review. If every line still passes through one person's eyes, that person's reading speed is your real ceiling — no matter how fast the model is.

It degrades exactly where you are going

AI is excellent on day one: a blank file and an afternoon get you a week of senior work. Real products cross into six figures of lines, where parts depend on each other in ways nobody wrote down. The model does not get worse — the blast radius of confidently wrong grows with the codebase, and nothing tells you unless something is built to.
Generation capacity rises steeply while review capacity stays flat; the widening gap becomes unreviewed code.AI tooling adoptedmonths latervolumecode generatedcode safely reviewedthe gap shipsanyway
The gap is not idle work. It is code that ships having been read less carefully than the code before it. We think that is what the 1.7× figure is measuring, though the study reports the association and not the cause.
Why the usual answers fail

Each one addresses generation. Generation was never the constraint.

More AI tooling

Raises output and hands the review queue straight back to the same people. You are buying more of the input to a step that is already saturated.

Hire senior engineers

The correct answer, and the slow expensive one. The scarcity of those people is why you reached for AI in the first place.

Bring in an agency

Ask the last agency that pitched you for the check results for a specific change they shipped, and see what arrives. You inherit their output and their assumptions either way.
What we do instead

Multiply the reviewers, not the reviewer.

If one person has to read everything, their reading speed is your throughput and no model changes that. The way through is a ladder: automated checks that block on failure, and narrow human judgement only where judgement is actually required.

Six checks run before anything reaches your product. The last one asks whether a real person can finish the job on the running system, because that is the failure a green test suite will not catch — we know, because ours did not.

And the bar rises as the product grows. The coverage that was enough on the first ten files is not enough at 165,000 lines — a fixed quality bar is a falling one.

What it produced

On a live product, over 165,000 lines.

98%
Work completed
1,355 of 1,381 items closed
182/184
Features delivered
99% of the planned programme
2.64
Defects per 1,000 lines
every line shipped — within the 1–5 published for a disciplined team
17 days
To reach this scope
against a 9–18 month norm — see /results for the caveat

Position as at 2 August 2026, recomputed from the work record. The same board says 43% of work was later redone and that queueing time rose as more work ran in parallel — we publish those too. All of it →

Find out what your codebase actually looks like.

Give us read access to one system. We map it and come back with what we found, what we would build and what it would take — before any commitment. The assessment is yours whether or not we go further.

Talk to us