Four options. We are the right answer for one of them.
You are not choosing between vendors, you are choosing between shapes: rent capacity offshore, retain a US agency, hand it to an AI-native studio, or arm the team you already have. Each is genuinely better than the others at something.
So each option below carries the row a comparison page normally leaves out — where it beats us. On rate, we lose outright, and we would tell you that on the call rather than let you find out in month three.
An offshore development shop
A contracted team, usually in another timezone, billed by the person.
Raw capacity at a rate no US supplier can match, and depth of bench when you need five people next month.
Timezone lag on every decision, senior scarcity, quality that varies by who was assigned, and usually no evidence of what was checked before it shipped.
If the constraint is genuinely headcount at low cost, this wins and we cannot compete on rate. We would say so.
A US agency or dev studio
A local team on a monthly retainer, with a named account lead.
Trust, proximity, references you can call, and someone who will sit in your office.
The most expensive option by a distance, margin bounded by headcount, and AI usually bolted onto the same process rather than changing it.
If you need people in the room and relationships matter more than throughput, this is the right shape.
An AI-native studio
A small team leaning heavily on generation, often at a flat monthly fee.
Speed and price. Genuinely fast on a greenfield build.
Usually thin process and no delivery evidence — which is the exact profile the 2026 defect data has taught buyers to be careful about.
We are priced in the same band. The difference is what happens after generation, and you should ask both of us to prove it.
AI coding tools for your own team
Copilot, Cursor and their kin, per seat.
Cheap, instant, and your team keeps full control.
It raises generation and hands the review queue back to the same people. The bottleneck moves; it does not lift.
You should be doing this anyway. It is complementary, not an alternative — the playbook is written to help whether or not you hire anyone.
The last row is not a rhetorical device. Two of these four are the correct choice for problems we are asked about most weeks, and the fourth is something you should be doing regardless of who you hire — the playbook is written for that case.
There is no price row here, and you should be suspicious of one.
A comparison table that prices four categories against each other has to invent three of the four numbers. We are not going to quote somebody else’s rate card to make ours look better, and a single figure for “an offshore shop” describes nothing you would actually be quoted.
Engagements are monthly and sized to scope, and sized to land below the fully-loaded cost of one senior US engineer.
What moves that figure is written out on the engagements page — size of the estate, how much of it nobody can explain any more, what it has to comply with, and how much you want to own yourself.
Everyone on this page uses AI. Almost nobody can hand you the evidence.
The models are the commodity part. Any of the four options above can call the same ones you can, at the same price, this afternoon. That is not where suppliers differ any more and a page claiming it as a differentiator is a page written before 2026.
What is hard to copy is the record. 6 checks run before anything reaches your product, each one recorded against the change that triggered it, so a year later you can ask what was verified before a particular thing shipped and get an answer instead of a recollection. Building that is a year of unglamorous work on the process, not a model upgrade.
We run it on our own product first — 165,000 lines of it — which is why the numbers on this site are ours rather than a case study we were handed.
The evidence
The live board
The decisions, written down
43% of our work was later redone, as at 2 August 2026, and queueing time rose as more work ran in parallel. Keeping the record is what makes those visible; it does not make them go away. We publish them because a supplier whose own board shows nothing wrong is a supplier who is not measuring. The full table, caveats included →
"Typical" means commonly published software-engineering ranges for work of comparable scope. The effort and calendar rows are an order-of-magnitude estimate, not a precise measurement — shown with that caveat rather than without it.
Five questions that separate the four options faster than any table.
Ask us the same five. If our answers are worse than someone else’s, you have learned that in a call rather than a quarter.
- 01
Show me the check results for one change you shipped last month.
Not a policy, not a test badge — the record for one specific change. Most suppliers can describe their process and cannot produce its output.
- 02
What share of your work gets redone?
We publish ours: 43%, as at 2 August 2026. A supplier who has never measured it will tell you it is low.
- 03
Who owns the repository, and what do I lose if I stop?
Ask specifically about proprietary runtimes and anything hosted you would have to migrate off.
- 04
Where does our code go, and which models see it?
This is the question your first enterprise customer will ask you, six months from now, in writing.
- 05
What would you refuse to take on?
A supplier with no answer has never turned work down, which tells you how the next scoping call will go.
Four situations where we would point you at one of the other three.
These are the ones we say on the call. It is cheaper for everyone to find them now than in the second month.
The scope is genuinely unknown
Nobody inside wants to own the gates
A simpler tool would do
You need people in the room
Compare us on output, not on this page.
Give us read access to one system. We map it and come back with what we found, what we would build and what it would take — before any commitment. Put that write-up next to what the other options give you at the same stage, which is usually a proposal.
Talk to us