Skip to content
Case study

165,000 lines of cybersecurity platform, one operator, twelve specialist roles.

Every figure on this site comes from one product. This is that product: what it does, the constraint that decided its architecture, how it was delivered, and the parts of the record we would rather not print.

Start with the thing you were about to think anyway. This is our own product, not a client engagement — we wrote the requirements as well as met them. What that does and does not prove is set out at the foot of the page, and it is worth reading before you weigh anything above it.

What was built

A scanner tells you a finding exists. This answers what it means here.

The product is a knowledge graph of a security estate. Assets, the vulnerabilities published against them, the detections that would catch an attempt to use one, and the countermeasures already deployed are held as one connected structure rather than four disconnected tools, so a question about any of them can be answered in terms of the others.

That changes what a security team can ask. Not “how many criticals do we have” — a number nobody can act on — but whether a new advisory reaches machines that actually exist here, which of the controls already running would see it, and what is left once you subtract those. A finding stops being an item in a queue and becomes a statement about one real estate.

We do not publish how large the graph is. The node and edge counts size the asset closely enough to tell a competitor what rebuilding it would cost them, so they stay out of public material — including this page.

It links four things that normally sit apart

Assets, vulnerabilities, detections and countermeasures, connected rather than exported into one another. The connections are the product; the lists were already available everywhere.

It answers in terms of your estate

The same advisory means something different on two estates. The answer names the machines, the controls in place, and the part that is genuinely still open.

It runs on models we host

No request leaves the customer’s environment on its way to an answer. That was a requirement before it was a feature — see below.
The constraint that shaped it

No external AI API. That rules out most of the obvious architecture.

A security estate is the one dataset an organisation will not post to a third party: it is a map of where they are weakest. So the product had to run inside the customer’s own infrastructure, with no call to a hosted model and no managed service in the answer path.

That single line deletes the architecture most teams would reach for. No frontier model doing the reasoning, no hosted vector service, no provider whose terms you have to read. What remains is smaller models running locally — and a set of consequences you inherit whether you planned for them or not.

The strongest models were not available to us

Work has to be cut into tasks a locally hosted model can finish reliably, with the structure around each call doing more and the model itself doing less.

The graph carries what the model cannot

Relationships are traversed, not inferred. The model turns a result into language; it is not asked to remember what connects to what, because that is the part it would get confidently wrong.

Offline is a build problem, not only a runtime one

Weights, datasets and dependencies all have to install on a machine with no route out. That cost lands in packaging and reproducibility, and it is invisible on any project that can reach the internet.

The honest reading: this is the slower way to build. With a hosted frontier model in the loop, several parts would have taken less time and some answers would probably be better today. The constraint bought deployability in environments that would otherwise be closed, and it was paid for in engineering.

How it was delivered

Four steps on a loop, and six checks a change has to clear.

This is the same process every client engagement runs on — the product was built on it, which is why we are willing to describe it in this much detail.

01

Understand

We read your problem — or your existing system — into a map of what it is, what depends on what, and where the gaps are. You get a plan you can challenge before anything is built.

02

Build

The work is split into small, tracked items with agreed acceptance criteria and assigned to the right specialist. Every change is isolated and reviewed.

03

Prove

Six checks run before anything reaches your product — ending with whether a real person can finish the job on the running system.

04

Report

Progress, quality and cost are recorded as the work happens. One live board, always current. Blocked work shows the day it blocks.

Every item was assigned to a role with its own scope and its own definition of done, never to a general-purpose prompt. Twelve of them worked on this product, under one accountable person:

TLTechnical LeadPOProduct OwnerBABusiness AnalystBEBackend EngineerFEFrontend EngineerDEData EngineerMLEML EngineerQAQA / DevOps EngineerSESecurity EngineerCMCompliance EngineerIEIntegration EngineerUXUI/UX Designer
A change passes through six gates; any one can stop it, and only a change clearing all six reaches production.changeyour product1Code health2Tests3Integration4Whole system5AI output6Real journeystopped here → back to the author
Six is the one that cannot be satisfied by writing more code to satisfy it: it asks whether a person can finish the job on the running system.
#CheckWhat it established on this buildRuns
01Security and code healthInsecure code, leaked credentials and unreadable work stopped at the door.every change
02Automated testsEvery change proved against the test suite — and the bar rises as your product grows.every change
03IntegrationConfirms the parts still work together, not just on their own.on integration
04Whole systemRebuilds everything from scratch and scans the full stack for vulnerabilities.on integration
05AI output qualityWhat the AI produced is sampled and scored. Nothing is assumed.before release
06Real user journeyCan a real person finish the job, end to end, on the running system.before release
The quality board part-way through a run on this product: gates listed with their individual sub-checks, each showing a pass, warn or fail state.
This product's own quality board, part-way through a run. Each check is named and holds its state while it is live, so a change that leaves one failing is visible on the day it fails rather than in a summary written afterwards. The same board is handed over on an engagement; there is no internal edition.
What the record shows

Read out of the work record, including the rows that do not flatter us.

Position as at 2 August 2026, recomputed from the record rather than typed into a slide. These are a dated capture, not a live feed. The board they come from keeps running, and the figures move with it — up and down.

98%
Work completed
1,355 of 1,381 items closed
182/184
Features delivered
99% of the planned programme
2.64
Defects per 1,000 lines
every line shipped — within the 1–5 published for a disciplined team
17 days
To reach this scope
against a 9–18 month norm — see /results for the caveat
80.6
Quality score
coverage, checks passed and defects, combined
10.9 h
Active work per item
time actually spent, not elapsed
87%
Steps with no human
of the process runs unattended
165,000
Lines of working software
the product these numbers come from
MeasureThis buildTypical for comparable workRead
Work delivered as committed94–96%60–85%Better
Time actually spent working28%10–20%Better
Test coverage80%70–80% targetIn line
Defects per 1,000 lines2.641–5 for a disciplined teamIn line
Effort to reach this scope≈ 1 person-month~200 person-monthsFar lower
Calendar time to reach it17 days9–18 monthsFar lower
Work later redone43%no published normWatch item

"Typical" means commonly published software-engineering ranges for work of comparable scope. The effort and calendar rows are an order-of-magnitude estimate, not a precise measurement — shown with that caveat rather than without it.

The two rows to look at first

43% of the work on this product was later redone. That is the highest number on the board and it has no published norm to hide behind, so we cannot tell you whether it is good. We can tell you it is measured, on its own line, rather than absorbed into a total where nobody would ever find it.

Time lost to queueing rose sharply as more work ran in parallel. Running more items at once made the average item slower to finish — which is the opposite of what the parallelism was for, and it is still on the board today.

The full record →
What went wrong

Every test was green. The product was broken.

Part-way through this build the suite passed, coverage was above eighty per cent, and every check reported success. Two screens were failing to load, and of nine journeys a customer might want to complete, one could be completed end to end.

Nothing was faulty about the tests. A test encodes an expectation, and an expectation is something a person had; the failures that hurt are the ones nobody thought of, which is precisely the set no test covers. We had been reading a green board as evidence of a working product, and those are different claims.

The sixth check exists because of that week. It asks whether a real person can finish the job on the running system — the one question that cannot be satisfied by writing more code to satisfy it.

The full account →
What we would do differently

Four changes we made from this build, and would make from day one on the next.

Put rework on the board in week one

We found the 43% figure by computing it, late, rather than by watching it. A rate that high is a signal about how work is being specified, and it arrived months after it would have been useful.

Cap how much runs at once

Queueing time rose with parallelism until adding work made everything slower. A limit on items in flight would have cost throughput on paper and bought it back in finished work.

Run a thin journey check at integration

The real-person check ran before release, which is where it caught the broken product. A cut-down version at integration would have turned a bad week into a bad afternoon.

Plan the offline packaging with the first architecture

Installing on a machine with no route out was treated as a late concern and behaved like an early one. It belongs in the first plan, priced alongside the parts everyone remembers to price.
Limits of this evidence

We set the requirements as well as met them, so discount it accordingly.

This is our own product. No client changed their mind in the middle of a sprint, no third party was late, no stakeholder had to be talked round, and when a requirement proved expensive we were free to move it. A large share of what makes commercial delivery slow was simply absent, and the calendar figure above reflects that as much as it reflects the process.

The quality figures are counted by us, from our own board, against our own definition of a defect. Nobody external has audited them. The comparison table is an order-of-magnitude reading against published ranges, not a controlled trial — one product, one domain, one operator, which makes it an existence proof rather than a distribution you can plan against.

What it does show is narrower and still worth something: a gated process ran to this scope, the record survived being published with its worst rows intact, and the same board and the same checks are what an engagement is delivered on. Whether that transfers to your codebase is a different question, and the honest way to answer it is to point the process at one of your systems and see.

Point it at one of your systems.

Give us read access to a single service or repository. We map what it is, what depends on what and where the gaps are, then write up what we found — what we would build, what we would not, and what we are assuming. About a week, no charge, and the assessment is yours whether or not we go further.

Talk to us