Practical guides
How to calculate ROI for an AI initiative
Most AI business cases are built on hours saved multiplied by a loaded hourly rate. That number is almost always wrong, and it is wrong in a predictable direction.
Every AI business case eventually reduces to one arithmetic sentence: this much value, against this much cost, arriving by this date. The difficulty is that both halves are routinely mis-stated, and the error compounds.
The value side is usually inflated by counting time saved as money earned. The cost side is usually understated by counting only the build. What follows is a way to do the calculation so that the answer survives a finance review.
Start with a baseline you already have
An initiative without a baseline cannot produce a return, only an opinion. Before estimating anything, establish what the workflow costs today using data that already exists in your systems.
Volume
How many times a month this decision or task happens. Take it from a system of record, not from a manager's estimate.
Unit effort
Time per instance, and by whom. Include the rework, the chasing, and the second pair of eyes.
Unit error
How often it goes wrong today and what a wrong one costs. This is the line most business cases omit entirely.
Cycle time
Elapsed time from request to resolution, including waiting. Effort and cycle time are different numbers and they pay differently.
If none of those four can be produced within a week, that is the finding. The initiative is not ready to be sized, and the first piece of work is instrumentation rather than AI.
Four ways value actually arrives
Be explicit about which of these you are claiming, because they convert to money at very different rates.
Hours saved is an input. Only capacity redeployed, cycle time that changes an outcome, errors avoided, or margin moved are outputs.
The UK government's AI adoption research found businesses reporting improved workforce productivity alongside no change in revenue. That is this distinction, observed at national scale.
The cost lines that get left out
Build cost is the line everyone estimates and the smallest of the four over three years.
Inference at real volume
Consumption pricing is cheap at pilot scale. Model cost per decision at expected volume and at three times it, because usage expands as people find new uses.
Evaluation and monitoring
Building and maintaining a held-out evaluation set of real cases, plus the dashboards that detect quality drift. Ongoing, not one-off.
Human review
The escalation path for cases the system should not decide alone. Budget it as a permanent percentage of volume, not a transitional cost.
Process change
Retraining, documentation, policy updates, and the productivity dip while people adjust. This is where most overruns live, and it is not a technology cost at all.
Two structural notes. Run the calculation over three years, because the recurring lines dominate. And model a version where the quality target is missed by ten per cent, because that scenario decides whether the initiative is fragile.
Put it in one line
Annual return equals volume, times the value per instance you are actually claiming, times the share of volume the system can handle unaided, minus annual inference, evaluation, review, and process cost. Divide the build cost by that figure to get payback in years.
Two parameters carry most of the sensitivity. The share of volume the system can handle unaided is usually assumed to be higher than it is, and the value per instance is usually the loaded hourly rate when it should be the marginal one. Vary both first.
The result is not meant to be precise. It is meant to be honest enough that a threshold can be agreed before launch: the figure that would justify scaling, and the figure below which you stop. Both should be written down while nobody is defending anything.
Four tests a number has to pass
Before a business case leaves the room, run it against these.
We have written about the organisational half of this at more length in AI ROI is an organisational problem, and about the specific places money escapes between pilot and return in the value gap. The measurement discipline itself is part of how we scope work.
A business case that passes these four is usually smaller than the one that walked in. That is the point of doing it.
Bring us the workflow. We will help you size it.
Discovery produces the baseline, the threshold, and the stopping condition. If the number does not clear, we would rather tell you that than build.