Four numbers, three of them proxies
A slide from a credible source puts four AI outcomes side by side. Three of them measure time saved, which is the input. The fourth is the only one a CFO can bank.
ICONIQ published a slide this year on how AI is landing inside G and A functions at their portfolio companies. It is a good slide. It is specific, it names sequences rather than tools, and its footer is more careful than most: illustrative examples observed at select portfolio companies, and not intended to represent benchmarks or survey findings. I would quote it in a room.
Across the top it carries four outcomes. Annual cost savings of two to four million dollars. Forty to sixty per cent faster time-to-hire. Forty-five to eighty per cent time saved in knowledge retrieval and analytics. Thirty to eighty per cent time saved on the finance close and reporting.
Read those again as measurements rather than as results. One is money. The other three are time.
Faster time-to-hire does not say whether the people hired were better, or stayed. Time saved in retrieval does not say whether anyone made a different decision with what they found. Time saved on the close does not say whether the close was more accurate, or whether the quarter it reported was read correctly.
Time saved is what the work costs. It is not what the work returned.
Time saved is what the work costs. It is not what the work returned.
This is the ordinary shape of a proxy metric. A proxy is a number that moves with the thing you care about, chosen because it is easier to observe. It earns its place right up until the moment it comes loose, and then it keeps reporting while the thing underneath stops moving. The failure is quiet, because the number still goes up.
Three of these four are the same proxy in different departments. That is worth noticing, because it tells you what the market currently finds easy to measure, and what it has not started measuring at all.
The one that is different is the two to four million. That is a number a finance team can put in a plan and check against an actual. It has a denominator and a period and an owner. It is the only one of the four I would build a case on, and I notice it is the one the slide marks as most impactful.
There is a second number further down the same page that I think is more useful than any of the four, and it is buried in a quote rather than set in the outcome row. A CFO at a five-hundred-million-dollar ARR company says their token spend went from near zero to roughly five to ten per cent of payroll, and that they pulled it out of the future headcount plan.
That sentence does something the percentages cannot. It gives AI spend a denominator that already exists in every company, and a budget line it can be taken from. Most AI ROI numbers in circulation are numerators with nothing underneath them, which is why a finance team cannot act on them even when they are true.
The same CFO opens with a line worth holding onto: the hard part is finding the value, not controlling the cost. Once a use case is valuable and expensive, the cost work is known. Mid-tier model defaults, caching, tiering by what the task is worth. Their spend came down while usage climbed. None of that is the hard part.
The hard part is knowing which use case was worth having in the first place, and then whether it did what you thought it did after it shipped. Neither of those is a cost question, and neither is answered by a percentage of time saved.
I am not saying the slide is wrong. The practices on it are the right practices, in what looks to me like the right order: standardise the manual workflow, put the person who owns the work in the builder seat, then share what they built as governed infrastructure. I would run an engagement in that order.
I am saying that if you lift the outcome row into your own board deck, you will be reporting three numbers that measure effort and calling them results. Somebody at that table will eventually ask what changed for a customer, and the row will not have an answer.
The fix is not more measurement. It is picking the one number per bet that would look bad if the bet failed, agreeing it before the work starts, and reading it against the proxy afterwards. When the proxy climbs and that number does not, you have learned something. That gap is the finding.