Skip to main content
03

Tool 03 · Agent Payback Model

Does an agent build actually pay back?

Size one manual workflow, set every assumption yourself, and read a 36-month cash flow whose zero crossing is the payback. Our own published 12-20 weeks build is drawn as the trough you have to climb out of — so a payback shorter than the build is impossible here.

What makes this different

  • Conservative, expected and aggressive — never one to-the-dollar figure from seven guesses.
  • Build cost decomposed into the delivery tracks published on our own price list, so you can audit it line by line.
  • Run cost modelled bottom-up from your volume, not as a percentage of the build price.
  • It can conclude that this does not pay back, and send you somewhere else.

No email required, no gate, no “contact us to see this”. Every coefficient is visible and adjustable, and each one is labelled either as published with a source and a date, or as an assumption that is yours to set.

01The model

Agent Payback Model

A planning model, not a forecast. Pick the segment first — it pre-fills every downstream default — then move anything you disagree with. Results recalculate as you drag; nothing is hidden behind a Calculate button.

Your current process
01 · Industry / workflow domain

Pick the segment first. It pre-fills headcount, hours, rate, volume and the automatable share, so you never face an empty slider — and changing it resets those five back to that segment's defaults.

Assumptionyours to set(segment pre-fills, all adjustable below)

people

Only the people inside this one workflow. A first production build covers one workflow, not a whole function.

Assumptionyours to set(yours — counted, not modelled)

h / week

Assumptionyours to set(yours — counted, not modelled)

/ hour

Salary, employer taxes, benefits and overhead — not the base wage. This rate also prices the human review of exceptions further down.

Assumptionyours to set(yours — counted, not modelled)

items / mo

Load-bearing: run cost is modelled bottom-up from this figure, not as a percentage of the build price. The scale is logarithmic because it spans four orders of magnitude.

Assumptionyours to set(yours — counted, not modelled)

≈ 60,000 a year
06 · Systems the workflow must touch

Drives the enterprise integration line in the build bill of materials. One system means no integration track at all.

PublishedVelocityMind published pricing, /pricing, 26 Jul 2026

07 · Distinct workflow variants in scope

Two or more variants pulls in the multi-agent architecture design track as well as development.

PublishedVelocityMind published pricing, /pricing, 26 Jul 2026

Cross-check · average handling time

That works out to 11.2 minutes per item. Does that match?

(people × hours × 52 × 60) ÷ (volume × 12) · believable range 0.5240 minutes

Assumptions you can change9 coefficients

Every coefficient that materially moves this answer is here, visible and adjustable. Nothing is hidden behind the result.

%

The largest single lever in this model, which is exactly why it is a slider. McKinsey Global Institute puts 60–70% of work hours in scope as technically automatable with generative AI. We use that as a ceiling on this slider, never as the number itself. McKinsey Global Institute, 2023, 26 Jul 2026.

Assumptionyours to set(resets when you change the segment)

%

Released hours are capacity, not a guaranteed headcount reduction. This is your call on how much of that capacity turns into budget. Our 60% default has no published source — Forrester's TEI methodology recaptures 50% of hours saved in its own composite model, so ours is on the optimistic side of that.

Assumptionyours to set

/ item
≈ $7,200 a year of inference
%

For invoice processing, Ardent Partners' ePayables 2024 reports 9% best-in-class against 22% typical. That is context, not our default — and we could not link a public copy of the primary document, so treat it as a reference point you should check. Our 15% sits between the two and is yours to move.

Assumptionyours to set

min

Priced at the same fully-loaded hourly cost you set above. On most default scenarios this line costs several times more than the inference does — the review loop is the reason pilots do not reach production.

Assumptionyours to set

≈ $39,000 a year of human review
/ year

Gateway, vector store, tracing, evaluation runs and on-call. Fixed rather than per-item.

Assumptionyours to set

A7 · Benefit ramp after go-live

Nothing reaches full value on day one. The ramp starts the month after the build ends, and both benefit and run cost ramp together.

Assumptionyours to set

%

The de facto convention, matching Forrester's TEI 10% and three-year NPV treatment. Compounded monthly across the 36-month horizon.

Assumptionyours to set

A9 · Regulated evidence pack required

Validation documentation, change control and a reproducible audit trace. Applies a ×1.25 multiplier to the build and pulls in the architecture design track.

Assumptionyours to set

Modelled results

Baseline · manual cost
$730K12 × 18 h × 52 × $65 — your inputs, not a model
Net annual benefit
$142K–$204KAfter inference, platform and human review. Conservative to aggressive.
Payback
9 mo–11 moFrom kickoff, including the published 12-20 weeks build.
3-year NPV
$250K–$383KAt 10% a year, compounded monthly over 36 months.

Ranges, never a point estimate. Negative figures render as negatives — nothing here is suppressed when the answer is bad.

36-month cumulative cash flow

The build is spent across the first 4 months — the mid-point of our published 12-20 weeks delivery window — and nothing returns until it ends. Where the curve crosses zero is the payback.

Expected scenario. A payback shorter than the build is structurally impossible in this model.

Conservative · expected · aggressive

Forrester's TEI methodology risk-adjusts benefits down 10–15% and costs up 5–10%. The conservative column uses the mid-points of those two ranges. A calculator returning one to-the-dollar figure from seven guesses is not being more precise than this; it is being less honest.

MeasureConservativeExpectedAggressive
Net / yr$141,796$176,726$204,029
Payback11 mo10 mo9 mo
3-yr NPV$249,744$324,693$382,934
Yr-1 return230%342%>400%

PublishedForrester Total Economic Impact methodology, 26 Jul 2026

The bridge — baseline to net annual benefit

Full figures, no compact rounding, so the arithmetic is checkable end to end.

StepAmountRunning
Baseline — annual cost of the manual work$730,080
Automatable share — 55%-$328,536$401,544
Realization — 60% converted to budget-$160,618$240,926
Less inference-$7,200$233,726
Less platform and observability-$18,000$215,726
Less human review-$39,000$176,726
Net annual benefit$176,726

Human review $39,000 against inference $7,200. The review loop is usually the larger number, and it is the one that decides whether a pilot reaches production.

Build — bill of materials

Decomposed into the delivery tracks published on our price list, so you can audit it line by line instead of trusting a black box.

TrackFormulaPublishedModelled
Agent development and deployment25,000 + 5,000 × (1 variant − 1)$25K - $75K$25,000
Enterprise integration15,000 + 8,000 × (2 systems − 2)$15K - $50K$15,000
Domain multiplierDocument Intelligenceassumption×1.00
Modelled build — returns are computed off this figure$40,000

Published implementation band $25,000 – $150,000 over 12–20 weeks from kickoff to production. See how the tracks are packaged and priced →

PublishedVelocityMind published pricing, /pricing, 26 Jul 2026

What moves this answer most

Each assumption perturbed ±10% and the payback re-solved. These two are worth arguing about; the rest are noise at your settings.

  1. automatable shareA 10% move against you on automatable share pushes payback from 10 months to 10 months.

  2. realization rateA 10% move against you on realization rate pushes payback from 10 months to 10 months.

Verdict

Clears the bar

This clears

Net $176,726 a year against a modelled $40,000 build, crossing zero at month 10 from kickoff — including the 12-20 weeks build itself.

Your inputs, the assumptions you chose and the figures they produced all travel with this link, so the first reply can argue with the model rather than ask you to repeat it.

Pressure-test these numbers with us

We reply within one business day.

This is a planning model, not a forecast, and it cannot see your data quality, your approval chain or your change-control burden — the three things that actually decide whether a build ships. The readiness ledger covers those →
02How this is modelled

Every coefficient, and where it comes from

Published in full so you can argue with it. This is our rubric, not an industry benchmark — we have no cohort to compare you against, so we disclose the whole model instead.

The arithmetic, in full

build        = (development + architecture + integration)
               × domainMultiplier × (regulated ? 1.25 : 1)
development  = 25,000 + 5,000 × (variants − 1)
architecture = (variants ≥ 2 OR regulated) ? 10,000 + 2,500 × (variants − 1) : 0
integration  = (systems ≥ 2) ? 15,000 + 8,000 × (systems − 2) : 0

baseline     = people × hours × 52 × rate
released     = baseline × A1
realized     = released × A2

inference    = volume × 12 × A3
platform     = A6
review       = volume × 12 × A4 × (A5 ÷ 60) × rate
runCost      = inference + platform + review
netAnnual    = realized − runCost

cashflow[m]  = m ≤ 4 ? −(build ÷ 4)
               : (netAnnual ÷ 12) × ramp(m − 4)
payback      = first m where Σ cashflow[1..m] ≥ 0
npv          = Σ cashflow[t] ÷ (1 + A8 ÷ 12)^t   for t = 0..35

aht          = (people × hours × 52 × 60) ÷ (volume × 12)

Build spend is spread evenly across the first 4 months — the mid-point of our published 12-20 weeks delivery window — and no benefit is allowed before it ends, which is what makes a payback shorter than the build structurally impossible rather than merely unlikely. Returns are computed off the modelled build, never off a figure clamped to the published band; the band $25,000 – $150,000 is used for engagement framing only.

PublishedVelocityMind published pricing, /pricing, 26 Jul 2026

The nine coefficients you can change

A1

Automatable share of in-scope hours

Default 55% · range 10% – 80%

The largest single lever in this model, which is exactly why it is a slider. Defaults come from the industry profile and reset when you change it. Bounded above by the published 60–70% technically-automatable ceiling, which is a ceiling and not our number.

Assumptionyours to set(The largest single lever in this model, which is exactly why it is a slider. Defaults come from the industry profile and reset when you change it. Bounded above by the published 60–70% technically-automatable ceiling, which is a ceiling and not our number.)

A2

Share of released hours converted to budget

Default 60% · range 20% – 100%

Released hours are capacity, not a guaranteed headcount reduction. This is your call on how much of that capacity turns into budget. Our 60% default is a modelling assumption with no published source — Forrester's TEI methodology recaptures 50% of hours saved in its own composite model, so ours is on the optimistic side of that.

Assumptionyours to set(Released hours are capacity, not a guaranteed headcount reduction. This is your call on how much of that capacity turns into budget. Our 60% default is a modelling assumption with no published source — Forrester's TEI methodology recaptures 50% of hours saved in its own composite model, so ours is on the optimistic side of that.)

A3

Cost per item processed by the agent

Default $0.12 · range $0.01 – $2.00

An assumption unless you bring it from the Architecture Configurator, which derives it bottom-up from tokens and orchestration pattern.

Assumptionyours to set(An assumption unless you bring it from the Architecture Configurator, which derives it bottom-up from tokens and orchestration pattern.)

A4

Exception rate

Default 15% · range 2% – 50%

For invoice processing, best-in-class is 9% against 22% typical (Ardent Partners, ePayables 2024). That is shown as context. Our 15% default is an assumption sitting between the two, and it is yours to set.

Assumptionyours to set(For invoice processing, best-in-class is 9% against 22% typical (Ardent Partners, ePayables 2024). That is shown as context. Our 15% default is an assumption sitting between the two, and it is yours to set.)

A5

Review minutes per exception

Default 4 min · range 1 min – 20 min

A modelling assumption. If you time it, use your number.

Assumptionyours to set(A modelling assumption. If you time it, use your number.)

A6

Platform and observability, annual

Default $18,000 · range $6,000 – $60,000

Gateway, vector store, tracing, evaluation runs and on-call. A modelling assumption.

Assumptionyours to set(Gateway, vector store, tracing, evaluation runs and on-call. A modelling assumption.)

A8

Discount rate for NPV

Default 10% · range 0% – 20%

The de facto convention, matching Forrester's TEI 10% and three-year NPV treatment.

Assumptionyours to set(The de facto convention, matching Forrester's TEI 10% and three-year NPV treatment.)

A9

Regulated evidence pack, build multiplier

Default ×1.25 · range fixed — not adjustable

A modelling assumption for validation documentation, change control and the reproducible audit trace a regulated deployment requires.

Assumptionyours to set(A modelling assumption for validation documentation, change control and the reproducible audit trace a regulated deployment requires.)

A7

Benefit ramp after go-live

Fast (3 months to full) — 50%, 80%, 100% · Standard (7 months to full) — 40%, 40%, 40%, 70%, 70%, 70%, 100% · Slow (12 months to full) — 25%, 25%, 25%, 25%, 50%, 50%, 50%, 50%, 75%, 75%, 75%, 75%, 100%

The last value in each schedule repeats for the rest of the horizon. Nothing reaches full value on day one, and pretending otherwise is how a model produces a payback that never happens.

Assumptionyours to set

Thresholds and cross-checks

  • Negative verdict — net annual benefit at or below zero. The figures still render, as negatives.
  • Slow verdict — payback beyond 24 months, or no crossing inside 36 months. The crossover headcount is solved live for your settings rather than asserted.
  • Implausible verdict — first-year return above 400%. We print the threshold instead of a four-figure percentage and tell you to re-check the scope.
  • Handling-time cross-check — the model divides your headcount and hours by your volume and warns when the result falls outside 0.5240 minutes per item, because that usually means the two inputs describe different processes.
  • Risk adjustment Forrester's TEI methodology risk-adjusts benefits down 10–15% and costs up 5–10%. The conservative column uses the mid-points of those two ranges. A calculator returning one to-the-dollar figure from seven guesses is not being more precise than this; it is being less honest.
  • The automatable ceiling McKinsey Global Institute puts 60–70% of work hours in scope as technically automatable with generative AI. We use that as a ceiling on this slider, never as the number itself.

Sources

  1. Forrester Total Economic Impact methodology

    Risk-adjusting benefits down 10–15% and costs up 5–10%; 50% recapture of hours saved; the 10% discount rate and three-year NPV convention. Cited by name as the basis for the conservative band.

  2. McKinsey Global Institute, 2023

    60–70% of work hours technically automatable with generative AI. Used strictly as a CEILING on the automatable-share assumption, never as the assumption itself.

  3. Ardent Partners, ePayables 2024

    9% best-in-class versus 22% typical invoice exception rate, shown as context beside the exception-rate slider and explicitly not as our default.

    No public link to the primary document — treat this as a reference point to verify with the publisher, not as our figure.

  4. VelocityMind published pricing

    Agent strategy $5K–$15K, multi-agent architecture design $10K–$30K, agent development and deployment $25K–$75K, enterprise integration $15K–$50K, and the published $25,000–$150,000 / 12–20 week implementation band.

Every weight and formula on this page is also published on the methodology page, rendered from the same module this calculator runs on, so the disclosure cannot drift from the code. Prices come from our published price list.

What this model cannot tell you

It values labour hours. For predictive maintenance the real driver is avoided downtime, for semiconductor work it is yield and test cost, and in healthcare it is throughput — none of which this model prices. It also cannot see your data quality, your approval chain or your change-control burden. If the answer here matters to you, those are the parts worth a conversation.

Should this be an agent at all? →Price the run cost bottom-up →Can you actually ship it? →