More code is arriving. Fewer people can vouch for it.
You gave the team AI tooling and output went up straight away. Then the review queue got longer. Then a bug came back from a file nobody had touched. Someone said the word “refactor” and everyone looked at their shoes.
Nobody on your team is doing anything wrong, and the model is not broken. You are running into something structural, and it has a shape.
Responsible for a system you did not build? That has its own page, and so does software an AI tool built.
You are not imagining it, and it is not only you.
Every one of these was published by someone else in 2026. We have linked each so you can check it rather than take our word.
more issues in pull requests containing AI-assisted code than in human-written code
arXiv preprint 2603.28592, March 2026maintenance costs by year two, for teams that do not actively manage the debt
Innovative Group, 2026Published figures, directional rather than audited. Three of the four come from the same 2026 analysis; the first is a separate preprint. Every one is linked so you can weigh it yourself.
Two ceilings, and neither is about the model.
We wrote these in July 2026, from running a factory past 165,000 lines of our own product. The figures above are other people’s — check them rather than take our word. We are not claiming we got there first: most of those sources carry a year and no month, and one is a March 2026 preprint, so the ordering is not something you could verify if we asserted it.
Your reviewer is your throughput
It degrades exactly where you are going
Each one addresses generation. Generation was never the constraint.
More AI tooling
Hire senior engineers
Bring in an agency
Multiply the reviewers, not the reviewer.
If one person has to read everything, their reading speed is your throughput and no model changes that. The way through is a ladder: automated checks that block on failure, and narrow human judgement only where judgement is actually required.
Six checks run before anything reaches your product. The last one asks whether a real person can finish the job on the running system, because that is the failure a green test suite will not catch — we know, because ours did not.
And the bar rises as the product grows. The coverage that was enough on the first ten files is not enough at 165,000 lines — a fixed quality bar is a falling one.
On a live product, over 165,000 lines.
Position as at 2 August 2026, recomputed from the work record. The same board says 43% of work was later redone and that queueing time rose as more work ran in parallel — we publish those too. All of it →
Find out what your codebase actually looks like.
Give us read access to one system. We map it and come back with what we found, what we would build and what it would take — before any commitment. The assessment is yours whether or not we go further.
Talk to us