Measuring AI ROI without lying to yourself
Here is how to measure AI ROI in one sentence: measure the process before the AI arrives, measure the same things after, count every cost including the human ones, and let the subtraction speak. Nobody disputes this method. Almost nobody follows it, which is why most AI ROI numbers in circulation are somewhere between soft and fictional.
The benchmark data makes the self-deception visible. In Wharton’s 2025 AI adoption study, roughly three in four enterprise leaders reported positive returns on generative AI. Over the same period, MIT’s Project NANDA examined actual deployments and found about 95 percent of pilots with no measurable P&L impact. Both can be true at once only one way: most of the reported ROI was never measured, it was felt. Executives are not lying to the surveys. They are reporting vibes, because their organizations never built the apparatus that would let them report numbers.
Baseline first, or the debate never ends
The single most common mistake is deploying first and measuring later. Once the AI is in the workflow, the before state is gone. You cannot reconstruct how long invoice handling took last spring; you can only argue about it, and every budget conversation becomes an exchange of anecdotes between believers and skeptics.
The baseline is a week of unglamorous work: hours spent on the process, error or rework rate, throughput, backlog age, cost per unit handled. Most of it already sits in your ticketing system, your ERP timestamps, or a simple time sample. This week of work is what makes every later claim checkable, and its absence is one of the recurring killers we describe in why most AI projects die. If a use case cannot be baselined at all, that is evidence you picked the wrong one; measurability is a selection criterion in choosing your first AI use case.
What to count as return
Time saved is the workhorse metric, with two honesty rules attached. Count net time: minutes saved per task, minus the minutes a human now spends reviewing the AI’s output. Review time is real work and it belongs in the equation. And distinguish capacity from cash: 2,000 hours freed is a valid return only when you can say what those hours now produce, whether that is absorbed growth without hiring, reduced overtime and backlog, or reassigned work. "Efficiency" with no destination is where soft numbers are born.
Beyond time, the returns that survive scrutiny are error rates (each error priced at its downstream cost of rework, refunds, or write-offs), throughput and latency where speed has business value (quotes answered in an hour instead of two days), and occasionally revenue effects, which are the hardest to attribute and should be claimed last, not first.
The costs everyone forgets
The visible costs are the build and the model bills. The forgotten ones routinely double the denominator. Human review time, again, because a system that saves ten minutes of drafting and adds six of checking is a fifteen-minute story someone will tell as thirty. Maintenance: models get deprecated, prompts and evaluation sets need upkeep, sources change; systems degrade silently without this spend, so it is not optional. Monitoring and evaluation infrastructure. Integration upkeep as the systems around the AI evolve. And adoption: training, workflow redesign, the champion’s time.
Model usage itself deserves one caution: it is usually modest at pilot scale and it grows with success. Price the ROI at target volume, not pilot volume, an exercise we detail in your AI bill at ten times the volume.
What a defensible AI ROI calculation looks like
The shape we hold our own projects to: a baseline measured before deployment; returns counted as net time valued at loaded cost, plus priced error reduction, plus throughput effects where they carry business value; costs counted in full, build, run, review, maintenance, adoption; the comparison run over a realistic period; and a named owner who signs the number the way a CFO signs a forecast.
A useful sanity test: could a skeptical board member recompute your ROI from the underlying data? If the answer is no, you do not have an ROI, you have a narrative with a percent sign.
None of this is exotic. It is a week of baseline work, a few honest subtractions, and the discipline to keep counting after the launch applause fades. We bake that discipline into every custom AI system we build: the baseline is captured before the prototype starts, so by the time production is on the table, the ROI case is arithmetic rather than argument.