How to measure ROI on an AI pilot: metrics that survive board scrutiny
Most organisations can tell you they ran an AI pilot. Far fewer can tell you what it actually returned. McKinsey’s most recent global survey found 88% of organisations now use AI in at least one business function, but only 39% can point to any measurable EBIT impact from it - and just 6% qualify as high performers attributing 5% or more of EBIT to AI. The gap isn’t that the pilots don’t work. It’s that most were never set up to produce a number a board would accept. If you’re working out how to measure ROI on an AI pilot before you ask for scale-up budget, the measurement plan has to exist before the pilot starts, not get reconstructed afterwards from whatever data happens to be lying around.
Why most ROI claims don’t survive the first hard question
The pattern shows up the same way in firm after firm: a pilot runs, the team reports it went well, and the first question back from finance or the board is “compared to what?” There’s rarely a good answer, because nobody measured the workflow before the pilot touched it. One recent industry study of global capability centres found over 70% of leaders piloting or scaling AI had no structured framework to measure the return at all. Without a baseline, “it feels faster” is the only claim left, and that claim doesn’t survive scrutiny.
The fix isn’t a more sophisticated dashboard after the fact. It’s capturing the baseline before anyone switches a tool on, and deciding in advance which three or four numbers will carry the entire argument for scaling.
The metrics that actually hold up
A pilot that survives board scrutiny reports on a small set of numbers, chosen before launch, not a long list assembled afterwards to make the result look better:
- A financial number, not a productivity proxy. Hours saved and tasks automated are useful internally, but a board wants the dollar version: cost avoided, revenue protected, or net benefit after the pilot’s own run costs (licensing, integration, oversight time). If you can’t convert the pilot’s effect into a dollar figure, that conversion - not the pilot - is the actual gap to close first.
- A baseline taken before the pilot started. Turnaround time, error rate, or cost per matter on the exact workflow the pilot touches, measured for at least a few weeks beforehand. Without it, every post-pilot number is a guess dressed up as a result.
- A payback period. How many months until the net benefit covers what the pilot cost to build and run. Boards fund things with a payback horizon attached; “it’s going well” doesn’t have one.
- A risk and quality indicator. Error rate, rework rate, or audit exceptions on the pilot’s output, tracked alongside the efficiency number. A pilot that moves faster but quietly raises the error rate hasn’t actually improved anything - and a board will ask this before they ask about speed.
Four numbers, agreed before the pilot starts, beats twelve numbers argued over afterwards.
Set the threshold before you see the result
The number itself isn’t enough - you also need the bar it’s being measured against, written down before the pilot runs. That means agreeing, in advance, what result justifies scaling (for example, a 15% improvement on the baseline with no new audit exceptions), what result justifies stopping, and who signs off on that call. Deciding the threshold after you’ve already seen a promising-looking result is how organisations end up scaling pilots that never earned it - and how a bad result gets quietly reframed as “learnings” instead of a stop decision.
This is also where the timeframe matters. A 60- to 90-day pilot in one team or workflow gives you a clean before-and-after window. A pilot that drifts on for six months without a review date rarely produces a number anyone trusts, because too much else changed in the meantime to isolate the pilot’s actual effect.
Report it the way a board actually reads it
A board doesn’t need the methodology - it needs four lines: what the baseline was, what the result was, what it’s worth in dollars, and what you’re asking for next (scale, adjust, or stop). Build the report around those four lines from day one, and the measurement work you did during the pilot becomes the board pack itself, not a translation exercise you do under deadline pressure the week before the meeting.
Where to start
Almost none of this is possible if nobody captured the baseline - and most firms haven’t, because nobody looked closely enough at the data behind the workflow before the pilot started. Our complimentary AI-Readiness Data Check is built to surface exactly that: twelve questions, about ten minutes, and a plain picture of what you can actually measure before you commit a pilot budget to it. Treat it as the step that makes the ROI conversation possible at all, rather than something you reach for after the board has already asked the question you can’t answer.