Tool 04 · Production Readiness Ledger
Can you actually ship it?
Fourteen questions, each about an artefact that either exists or does not. No score out of 100, no maturity curve, no peer cohort. The output is GO, NOT YET or NO — with the exact missing artefacts named. If it says NOT YET, it tells you not to hire anyone, including us, until you have them.
Ninety seconds, nothing to estimate
- Questions
- Blocking
- Output
- Score
Every gate is an artefact, not an opinion, so there is nothing here you can flatter yourself into. Runs entirely in your browser: no email, no account, nothing stored.
Gartner names three causes. These are the three column headings.
Over 40% of agentic AI projects are forecast to be cancelled by the end of 2027 — on escalating cost, unclear business value and inadequate risk controls. Every gate below sits under one of those three causes, because the cheapest way to avoid being in that 40% is to check whether the artefacts exist before anyone is paid to build anything.
Nobody owns the number, the released time or the system a year from now.
Nothing to measure against, so nothing can be proven.
No access, no authority, no way back.
Answer for one workflow, not for the organisation
Pick the workflow you would actually put an agent on. Answer honestly — a don't know is counted as a no, because an artefact nobody can point to is not one you have. The verdict updates as you go and the link carries every answer.
Step zero — where the work sits
InputsEscalating cost
Nobody owns the number, the released time or the system a year from now.
Unclear business value
Nothing to measure against, so nothing can be proven.
Inadequate risk controls
No access, no authority, no way back.
NO — not buildable yet
This cannot be built by anyone, including us, until G01, G02, G03 and G04 exist. They are free. They are also the 4 things that decide whether the build survives contact with production.
Blocking artefacts missing
A written definition of what counts as a wrong answer for this workflow, agreed by the people who own the process.
Read access to the system of record in a non-production environment, available today.
A named individual with authority to let the agent act without review, and the authority to stop it.
A rollback that does not require a vendor to execute.
We would not take this build yet either. Thirty minutes on closing the blockers — no pitch. The blockers below are internal work, and they are free. That half hour is worth more to you right now than a proposal.
Rule 1 — any blocking gate missing returns NO. Rule 2 — 3 or more non-blocking gates missing returns NOT YET. Rule 3 — everything else returns GO. There is no score, no maturity curve and no peer cohort. See the full rubric →
You cannot state the error rate of the humans doing this work today. That means you will have no way to prove the agent is better, and no way to defend it the first time it is wrong.
A number, from a sample somebody actually checked, for how often the current process gets it wrong. Not an estimate from memory.
Every gate, its cause group, and where you stand
- In place
- 0
- Missing
- 14
- Don't know
- 0
- Gates
- 14
| Gate | Artefact | Status |
|---|---|---|
| Escalating cost | ||
A number the CFO already believes for what this workflow costs today. A cost figure that has survived finance once already. A number invented for the business case will be challenged exactly when you need it. | ||
A decision on what happens to the time the agent frees up, made by someone who owns that budget. A written answer to 'and then what?' from the person whose budget the released hours sit in. Without it, the saving is a slide. | ||
Somebody who will still own this system in twelve months. A named owner with a role that survives the project. An unowned agent is a liability with a monthly bill. | ||
| Unclear business value | ||
A written definition of what counts as a wrong answer for this workflow, agreed by the people who own the process. A page naming the failure modes and, for each, what the correct outcome would have been — signed off by the process owner, not written by the project team alone. | ||
A measured baseline error rate for the humans doing this work today. A number, from a sample somebody actually checked, for how often the current process gets it wrong. Not an estimate from memory. | ||
At least 50 real historical items, with their correct outcomes, that could be used as an evaluation set. Fifty real cases with the right answer attached — the minimum from which an evaluation harness can be built. | ||
A test environment where a failure costs nothing. Somewhere the agent can be wrong repeatedly without touching a customer, a payment or a patient. | ||
| Inadequate risk controls | ||
Read access to the system of record in a non-production environment, available today. Working credentials against a non-production instance with representative data — not a promise that access can be requested. | ||
A named individual with authority to let the agent act without review, and the authority to stop it. One person, named, who can both switch it on and switch it off. A committee is not an answer to this question. | ||
A rollback that does not require a vendor to execute. A documented, rehearsed way for your own team to reverse what the agent did, without opening a support ticket. | ||
A written escalation path for the case the agent gets wrong — who is told, how fast, what they can do. A named recipient, a stated time, and a stated remedy. One page. | ||
Write access to the system of record, or a documented decision that the agent will only recommend. Either the write path exists, or somebody has decided in writing that it will not — both are fine, ambiguity is not. | ||
A named owner for the data the agent will read, who can grant access without a committee. One person who owns the data and can say yes. | ||
An agreed retention and logging policy that covers model inputs and outputs. A written policy stating what is logged, where it lives and how long it is kept — covering prompts and completions, not only application logs. | ||
Every rule, printed before you answer
No competitor in this landscape shows its weights. We have none to hide: there is no arithmetic here, no coefficient and no score — only a list of gates and three rules applied in order.
The verdict rules
- Any blocking gate (G01–G04) answered No or Don't know.
- Three or more non-blocking gates answered No or Don't know.
- Everything else.
A don't know counts as a no for the verdict, and is reported separately and prominently. The threshold in rule 2 is 3 — it is not tuned, it is published, and you are welcome to argue with it.
The 4 blocking gates
These four are not weighted more heavily. They are absolute: any one of them missing returns NO on its own, because without them nobody can build this — including us.
- A written definition of what counts as a wrong answer for this workflow, agreed by the people who own the process.
- Read access to the system of record in a non-production environment, available today.
- A named individual with authority to let the agent act without review, and the authority to stop it.
- A rollback that does not require a vendor to execute.
The gate list and these rules are one pure module that both this page and the tool import, so the disclosure cannot drift from the code that produces the answer. Read the full methodology →
What this does not measure
Transparency is the only thing we can honestly substitute for scale. So here is what this instrument is not.
This is our rubric, not an industry benchmark. Every gate and every verdict rule is published above and on the methodology page.
It scores production-readiness for one workflow, not organisational AI maturity. Those are different questions and conflating them is how readiness quizzes became worthless.
We deliberately do not place you against a peer cohort, because we do not have one. Cisco can, with 7,985 double-blind respondents across 30 markets. Borrowing that language without the sample would be the fastest way to be dismissed.
Sources
Over 40% of agentic AI projects forecast cancelled by end-2027 on escalating cost, unclear business value and inadequate risk controls (poll of 3,412). The three named causes are the three column headings of this ledger.
Roughly 95% of enterprise GenAI pilots show no measurable P&L return, with barriers that are organisational rather than model-related.
Cited as the reason we do NOT claim benchmarking: their cohort is 7,985 double-blind respondents across 30 markets. Ours does not exist.
Not sure which of these you need? Start with the necessity test → Or go back to the toolkit →