How to measure the real ROI of AI

In the boardroom someone puts up a slide that says we saved time, and nobody asks the question that matters: compared to what. Without that answer, the number measures nothing.

Carlos Andrés Ramírez ·

In the boardroom someone puts up a slide that says we saved time, and nobody asks the question that matters: compared to what. Without that answer, the number measures nothing. It's a figure shaped like an argument.

I have sat in more than one board meeting where, six months after approving an AI budget, someone comes back with an adoption report. Active seats, prompts per day, a satisfaction score with a high number. All of that measures usage. None of it answers the question the CFO actually asked when the budget got approved, which was different: did this move anything on the P&L.

And that is where it falls apart, because nobody in the room can honestly answer that. Not for lack of data. There is plenty of data. What is missing is the method to separate what the AI did from what the business would have done anyway, and in most cases the two are impossible to tell apart by eye.

The symptom

How do you measure the actual value of an AI initiative?

The answer circulating today is a list of vanity metrics dressed up as rigor: adoption, satisfaction, number of use cases in production. They work for a quarterly update and fail the moment someone has to defend a second year of budget, because none of them isolates the AI's effect from everything else happening in the business at the same time.

  • The report compares the quarter with AI against the quarter without it, while headcount, demand or the process itself also changed in between.
  • Nobody set a baseline before starting, so the 'before' gets reconstructed from memory six months later.
  • The time the tool frees up gets reported as savings even though nobody checked what that hour actually went toward.
  • ROI gets calculated on the pilot, with the most motivated team and the easiest case, then projected as if it would repeat identically across the rest of the operation.
  • The number that climbs on the dashboard is usage, not outcome, because usage is the one number that exists without asking finance for anything.

The problem underneath

Freed-up time isn't a return until it lands somewhere.

An hour AI hands back to an analyst is worth nothing on its own. It is worth something if it turns into fewer hours worked, more output from the same team, or a different task that used to go undone. If that hour refills itself with more of the same queue, which is what happens most of the time, there is no return, no matter what the usage report says.

And without a baseline or a comparison window, any improvement you observe could come from the AI or from the quarter simply being better for another reason: a new hire, a simplified process, lower demand. Crediting the whole gain to AI because it happened at the same time is the most common way to buy smoke with real data.

An adoption dashboard is not a business case. It's proof that someone used the tool, not that the company gained anything from it.

BECOME

The method

Four variables to fix before measuring, not after.

Baseline
The state of the process before AI, measured with the same indicator you will use afterward: time per case, cost per transaction, error rate. Set before you start, not reconstructed from memory at the end.
Attribution window
The period over which you compare before and after, tracking what else changed in that stretch: volume, headcount, process. Without that list, the improvement gets credited to AI by default.
Where the time went
Fewer heads, more output from the same team, or a new task of different value. If nobody can name the destination, that hour doesn't count as return.
Comparison group
An equivalent team or process that didn't use the AI over the same period. No lab experiment required: a similar business unit works as a reference for how much would have changed anyway.

None of the four requires more budget. They require deciding on them before the initiative gets approved, not when someone asks for the number in committee. The team that fixes them upfront shows up with a business case. The one that doesn't shows up with a usage dashboard and hopes nobody asks compared to what.

Take the AI initiative that has been running longest and ask, for every result credited to it, what that same number would have been without the tool. If nobody in the room can answer with a number, there is no business case. There is a hunch with a chart attached.

Frequently asked questions

How do you measure the actual value of an AI initiative?

By fixing a baseline before you start, using the same indicator you will compare later, an attribution window that tracks what else changed in the business, and a defined destination for the time the tool frees up. Without those three decisions made in advance, any number presented afterward is an estimate dressed up as measurement.

Why isn't adoption the same as AI return?

Adoption measures how many people use the tool and how often, which is a behavioural number. Return measures whether that changed cost, time or revenue, which is a financial one. An initiative can show high adoption and zero return if the freed-up time dissolves into more of the same queue.

How do you separate AI's effect from other business changes?

With an attribution window that records what else moved in the same period: headcount, volume, process changes. Without that record, any improvement observed gets credited to AI by default even when the real cause was something else, and that easy attribution never survives a CFO's first serious question.

What happens to AI-freed time if nobody tracks where it goes?

It refills itself with more of the same queue, so the person's day still looks just as full even though the tool is working. That time only counts as return once someone explicitly decides its destination, fewer hours, more output or a different task, and checks afterward that it actually went there.

Let's build your business case

From the idea to the operation

Adoption is not communicated: it is designed with the teams who will operate the capability, and measured against a baseline.

About the author

Carlos Andrés Ramírez — Transformation Director

Specialist in business transformation and reinvention. Director of Specialised Programmes and lecturer in Artificial Intelligence at UPC's Graduate School.

LinkedIn