Skip to main content

Toolkit · four instruments

Decide whether to build it — before you pay anyone

Four connected tools that follow the order you actually ask the questions in: should this be an agent, what shape would it be, does it pay back, can you ship it. Every one of them can tell you not to build. No email, no score out of 100, no peer cohort we do not have.

What you get, and what it costs you

  • A complete answer on screen, instantly — nothing held back behind a form
  • The whole rubric printed beside the verdict, so you can argue with the weights
  • A copyable link and a printable one-page model sheet, both ungated
  • A verdict that can be “do not build this”, and often is

Nothing on these pages calls a language model. Every figure is arithmetic you can reproduce from the published formula.

01The chain

Four questions, in the order a buyer actually asks them

Each instrument answers one question and hands its state to the next, so nothing you have already told us has to be typed twice. You can also start anywhere — every tool stands on its own.

  1. 01
    TOOL 0110 questions · about two minutes

    Should this be an agent?

    Answer for one workflow. The rubric scores it on two axes — how much autonomy the work genuinely needs, and how much control it demands — and returns one named verdict on the spectrum below. Every weight, every threshold and every override is printed underneath the answer, so you can check the scale is not tilted toward the expensive end.

    Possible verdicts

    T0Rules or a scriptT1Deterministic workflow / RPAT2One LLM call inside a deterministic pipelineT3Retrieval-grounded single agentT4Supervised multi-agent with human-in-the-loop

    Where this costs us the sale

    T0, T1, T2 — 3 of the 5 verdicts — all conclude that this should not be an agent. T0 is deliberately costed below our own published $25,000 floor, and says so on the page.

    Agent Necessity TestCarries forward to TOOL 02tier · steps · volume · industry
  2. 02
    TOOL 02about ninety seconds

    What shape would it be?

    Describe the workflow and its constraints and get a real architecture back: the orchestration pattern, the components we would build and the ones we ruled out with the reason for each, a model tier per role, and the envelope it runs in — tokens and cost per run, monthly spend, p50 and p95 latency, end-to-end reliability, and where it breaks.

    Possible patterns

    Not an agentSingle agentSequential pipelineRouter + specialistsSupervisor + specialistsParallel fan-out + evaluator

    Where this costs us the sale

    The first rule in the pattern tree returns NOT AN AGENT and hands you back to the necessity test. It fires on settings close to the defaults, which is the point.

    Agent Architecture & Run-Cost ConfiguratorCarries forward to TOOL 03cost per run · volume · industry
  3. 03
    TOOL 0336-month cash flow · conservative, expected, aggressive

    Does it pay back?

    Size one manual workflow and set every coefficient yourself — none of them is hidden, and none of them is described as a benchmark unless it is published with a source. The payback is the zero crossing of a 36-month cash flow, with our own 12-20 weeks build drawn as the trough you have to climb out of first.

    Possible verdicts

    ClearsSlowNegativeImplausible

    Where this costs us the sale

    A payback shorter than the build is structurally impossible here. The negative and slow verdicts route you sideways to the necessity test rather than into our contact form, and the model calls its own output implausible rather than printing a four-figure return.

    Agent Payback ModelCarries forward to TOOL 04industry · volume
  4. 04
    TOOL 0414 gates · about ninety seconds

    Can you actually ship it?

    Every question is about an artefact that either exists or does not — an evaluation set, a named owner, a measured human error rate — never a maturity opinion, so there is nothing to estimate and nothing to flatter. The output names the exact artefacts you are missing.

    Possible verdicts

    GONOT YETNO

    Where this costs us the sale

    NO means we would not take this build either. It names the work to do first, and none of it is work you need to pay us — or anyone — for.

Every weight, coefficient and formula behind those four verdicts is published in one place.

Read the methodology →
02Tool index

Four instruments and the reference behind them

All four are live, ungated and free. Each one is built to be able to tell you not to build — and to say so as clearly as it would say the opposite.

TOOL 01Agent Necessity Test

The question

Does this workflow need an agent, a single model call, or plain code?

What it returns

A named tier from T0 to T4, with the full scoring table printed beside it.

TOOL 02Agent Architecture & Run-Cost Configurator

The question

What would it be made of, and what would it cost to run every month?

What it returns

A blueprint, a bill of materials including the parts we ruled out, and a cost, latency and reliability envelope.

TOOL 03Agent Payback Model

The question

When does the build pay for itself, if it ever does?

What it returns

A 36-month cash flow and a payback band — conservative, expected and aggressive, never one number.

TOOL 04Production Readiness Ledger

The question

Do you have the artefacts a production deployment needs?

What it returns

GO, NOT YET or NO, with the missing artefacts named one by one.

REFERENCEMethodology

The question

Where does every number in these tools come from?

What it returns

Every rubric weight, formula, coefficient and citation, rendered from the same modules the tools run on.

03The promise

No email. No gate. No cohort.

A tool that trades your answer for your address is a lead form wearing a calculator’s clothes. These are not that, and the constraint is written into how they are built rather than into a paragraph of reassurance.

No email address, at any point

Results are instant and complete, and there is no “email me this report” field anywhere. The durable artefact is a print sheet you produce yourself with Cmd-P and a link you copy to your clipboard — both ungated, both forwardable to your CFO without going through us.

Every weight published

The methodology page renders from the same modules the tools calculate with, so what we publish is mechanically incapable of drifting from what we compute. Nothing that moves the answer is hidden, and nothing that is our assumption is labelled as a benchmark.

Deterministic, with no model in the loop

No tool here calls an LLM. The arithmetic is pure and client-side, which means a colleague opening your link sees exactly the verdict you saw — and means we can publish the rubric in full, which a model in the loop would make impossible.

04Limits

What this toolkit will not do

Six things we decided against, with the reason for each. A tool that will not name its own limits is asking you to take the rest of it on faith.

Benchmark you against peers
We have no respondent cohort, so we will not pretend to place you in one. No percentiles, no quartiles, no “compared to similar organisations”. What we publish instead is the whole rubric, so you can argue with it.
Give you a score out of 100
A score invites a curve, and a curve invites a cohort we do not have. You get a named verdict and the reasoning that produced it, or a band — never a single number carrying more precision than the inputs deserve.
Price the run cost as a share of the build
Inference scales with volume multiplied by steps, not with what the build was quoted at. The configurator prices it bottom-up from tokens and the orchestration pattern, and the payback model consumes that figure rather than guessing a percentage.
Convert currencies
Every figure is US dollars. We have no cited exchange-rate source, and a selector that changes the symbol without converting the number is worse than not offering one.
Read your free text
Where a tool has a text field, whatever you type is carried to the report header and to your enquiry as a label. It is never an input to a verdict — a verdict you cannot reproduce is a verdict you cannot check.
Hide the bad answer
Negatives render as negatives. Nothing is ever replaced with an em dash, a blur or a “contact us to see this”, and the answer that costs us the sale gets the same design weight as the one that wins it.

Where to start

Start at the top of the chain

The necessity test takes about two minutes and can end the conversation before it costs you anything. If it says an agent is the wrong instrument, that is a finished answer — take it, and keep your budget.

Start with the necessity test

We reply within one business day.

Rather talk it through? Request a strategy call →