Toolkit · four instruments
Decide whether to build it — before you pay anyone
Four connected tools that follow the order you actually ask the questions in: should this be an agent, what shape would it be, does it pay back, can you ship it. Every one of them can tell you not to build. No email, no score out of 100, no peer cohort we do not have.
What you get, and what it costs you
- A complete answer on screen, instantly — nothing held back behind a form
- The whole rubric printed beside the verdict, so you can argue with the weights
- A copyable link and a printable one-page model sheet, both ungated
- A verdict that can be “do not build this”, and often is
Nothing on these pages calls a language model. Every figure is arithmetic you can reproduce from the published formula.
Four questions, in the order a buyer actually asks them
Each instrument answers one question and hands its state to the next, so nothing you have already told us has to be typed twice. You can also start anywhere — every tool stands on its own.
- 01TOOL 01
Should this be an agent?
Answer for one workflow. The rubric scores it on two axes — how much autonomy the work genuinely needs, and how much control it demands — and returns one named verdict on the spectrum below. Every weight, every threshold and every override is printed underneath the answer, so you can check the scale is not tilted toward the expensive end.
Possible verdicts
T0T1T2T3T4Where this costs us the sale
T0, T1, T2 — 3 of the 5 verdicts — all conclude that this should not be an agent. T0 is deliberately costed below our own published $25,000 floor, and says so on the page.
- 02TOOL 02
What shape would it be?
Describe the workflow and its constraints and get a real architecture back: the orchestration pattern, the components we would build and the ones we ruled out with the reason for each, a model tier per role, and the envelope it runs in — tokens and cost per run, monthly spend, p50 and p95 latency, end-to-end reliability, and where it breaks.
Possible patterns
Where this costs us the sale
The first rule in the pattern tree returns NOT AN AGENT and hands you back to the necessity test. It fires on settings close to the defaults, which is the point.
- 03TOOL 03
Does it pay back?
Size one manual workflow and set every coefficient yourself — none of them is hidden, and none of them is described as a benchmark unless it is published with a source. The payback is the zero crossing of a 36-month cash flow, with our own 12-20 weeks build drawn as the trough you have to climb out of first.
Possible verdicts
Where this costs us the sale
A payback shorter than the build is structurally impossible here. The negative and slow verdicts route you sideways to the necessity test rather than into our contact form, and the model calls its own output implausible rather than printing a four-figure return.
- 04TOOL 04
Can you actually ship it?
Every question is about an artefact that either exists or does not — an evaluation set, a named owner, a measured human error rate — never a maturity opinion, so there is nothing to estimate and nothing to flatter. The output names the exact artefacts you are missing.
Possible verdicts
Where this costs us the sale
NO means we would not take this build either. It names the work to do first, and none of it is work you need to pay us — or anyone — for.
Every weight, coefficient and formula behind those four verdicts is published in one place.
Read the methodology →Four instruments and the reference behind them
All four are live, ungated and free. Each one is built to be able to tell you not to build — and to say so as clearly as it would say the opposite.
The question
Does this workflow need an agent, a single model call, or plain code?
What it returns
A named tier from T0 to T4, with the full scoring table printed beside it.
The question
What would it be made of, and what would it cost to run every month?
What it returns
A blueprint, a bill of materials including the parts we ruled out, and a cost, latency and reliability envelope.
The question
When does the build pay for itself, if it ever does?
What it returns
A 36-month cash flow and a payback band — conservative, expected and aggressive, never one number.
The question
Do you have the artefacts a production deployment needs?
What it returns
GO, NOT YET or NO, with the missing artefacts named one by one.
The question
Where does every number in these tools come from?
What it returns
Every rubric weight, formula, coefficient and citation, rendered from the same modules the tools run on.
No email. No gate. No cohort.
A tool that trades your answer for your address is a lead form wearing a calculator’s clothes. These are not that, and the constraint is written into how they are built rather than into a paragraph of reassurance.
No email address, at any point
Results are instant and complete, and there is no “email me this report” field anywhere. The durable artefact is a print sheet you produce yourself with Cmd-P and a link you copy to your clipboard — both ungated, both forwardable to your CFO without going through us.
Every weight published
The methodology page renders from the same modules the tools calculate with, so what we publish is mechanically incapable of drifting from what we compute. Nothing that moves the answer is hidden, and nothing that is our assumption is labelled as a benchmark.
Deterministic, with no model in the loop
No tool here calls an LLM. The arithmetic is pure and client-side, which means a colleague opening your link sees exactly the verdict you saw — and means we can publish the rubric in full, which a model in the loop would make impossible.
What this toolkit will not do
Six things we decided against, with the reason for each. A tool that will not name its own limits is asking you to take the rest of it on faith.
- Benchmark you against peers
- We have no respondent cohort, so we will not pretend to place you in one. No percentiles, no quartiles, no “compared to similar organisations”. What we publish instead is the whole rubric, so you can argue with it.
- Give you a score out of 100
- A score invites a curve, and a curve invites a cohort we do not have. You get a named verdict and the reasoning that produced it, or a band — never a single number carrying more precision than the inputs deserve.
- Price the run cost as a share of the build
- Inference scales with volume multiplied by steps, not with what the build was quoted at. The configurator prices it bottom-up from tokens and the orchestration pattern, and the payback model consumes that figure rather than guessing a percentage.
- Convert currencies
- Every figure is US dollars. We have no cited exchange-rate source, and a selector that changes the symbol without converting the number is worse than not offering one.
- Read your free text
- Where a tool has a text field, whatever you type is carried to the report header and to your enquiry as a label. It is never an input to a verdict — a verdict you cannot reproduce is a verdict you cannot check.
- Hide the bad answer
- Negatives render as negatives. Nothing is ever replaced with an em dash, a blur or a “contact us to see this”, and the answer that costs us the sale gets the same design weight as the one that wins it.
Where to start
Start at the top of the chain
The necessity test takes about two minutes and can end the conversation before it costs you anything. If it says an agent is the wrong instrument, that is a finished answer — take it, and keep your budget.
Rather talk it through? Request a strategy call →