Skip to content
Quality

Every test was green. The product was broken.

·4 min read

The test suite passed. Coverage was above eighty per cent. Every gate reported success, and the product was broken.

Two screens failed to load. Of nine journeys a customer might actually want to complete, exactly one could be completed end to end. None of that showed up anywhere, because none of it was what anyone had thought to test.

Tests check what someone thought of

That is not a criticism of testing. It is the definition of it. A test encodes an expectation, and an expectation is something a person had. The failures that hurt are the ones nobody expected — which is precisely the set no test covers.

You can push the boundary outward with more tests, and you should. But the boundary moves; it never closes. At some point the only remaining signal is whether a person can finish the job on the running system.

So that became the last gate

Six checks run before anything reaches a product we deliver. The first five are the ones you would expect: code health, tests, integration, a full rebuild with a vulnerability scan, and a sampled score of what the AI actually produced.

The sixth asks a different kind of question. Not "did the code behave as specified" but "can a real person get through this". It is the one check that cannot be satisfied by writing more code to satisfy it.

Tests can only check what someone thought to test. The one signal that cannot be faked is whether a person can finish the job.

The bar has to rise with size

The other thing that incident taught us: what was good enough at five thousand lines is not good enough at a hundred and sixty-five thousand. A fixed quality bar is a falling quality bar, because the number of ways a system can be broken grows faster than the system does.

So the thresholds move as the product grows. That is unpopular with anyone who wants a stable definition of done, and it is the only version that stays honest.

Also worth reading

Start with one piece of work.

Give us one piece of work and read access to one system. We map it and come back with what we found, what we would build and what it would take — before any commitment. You decide using our output, not our pitch.

Talk to us