Your AI Metrics Measure Expense, Not Value
Every consumption dashboard is a cost report wearing a performance costume.
Consumption belongs on the expense line until an outcome is attached to it. A usage chart tells you what an effort cost. Presented in a performance review it silently claims to say what the effort produced. The two live in different columns of the same business, and moving a number from one column to the other without evidence is the whole mechanism of AI theater.
Published
Key Takeaways
- A consumption figure is cost telemetry. It reports how much purchased capacity was drawn down, which is a fact about spending and contains no information about result.
- The error is not falsification, it is filing. On the expense line the number says what we spent. Moved to the performance line, unchanged, it says what we accomplished. Nothing about the measurement changed. Everything about the claim did.
- The correction is cheap. Keep the dashboard, relabel it as cost, send it to the function that owns cost, and leave the performance slot empty until an outcome measure can fill it honestly.
Adoption is not value. Consumption is not even adoption. It is what the effort cost, presented as what the effort produced.
From Builders Build, Chapter 9
That sentence is an accounting statement before it is a management one. It says that a number generated by the act of spending has been placed where a number describing production is supposed to sit. Everything that follows in this article is an attempt to prove it, because organizations that see the substitution clearly tend to fix it in a single reporting cycle, and organizations that do not can spend years defending a chart that was never evidence of anything.
What a consumption figure actually reports
Take the most impressive AI number in your current reporting pack and describe it without adjectives. It records how much of a purchased capacity was drawn down, by which teams, over which period, compared with the period before. That is the complete content of the measurement. It is generated as a byproduct of the transaction, which is why it is available so early and in such fine detail.
Notice what the measurement never touches. It does not observe the work the capacity was spent on. It does not observe whether that work was needed. It does not observe whether the output was used, reviewed, discarded, or paid for a second time when someone had to correct it. The instrument sits entirely on the purchasing side of the transaction and reports what was purchased.
This is why the same figure can be produced by two organizations with nothing in common. One redesigned a claims process and now resolves cases in a third of the time. The other regenerated the same board memo nine times because editing did not feel like progress. The dashboards are indistinguishable. Any measure that cannot separate those two organizations is not measuring performance, and no amount of visual polish will make it start.
None of this makes the number bad. A meter that reports electricity drawn is an excellent meter. It is simply not a report on what the factory built.
The column it belongs in
Every organization already operates a discipline for keeping this straight. It is called accounting, and it is unglamorous precisely because it refuses to let a fact about spending drift into a claim about production.
A consumption figure has a home, and the home is the expense line. It supports capacity planning, vendor negotiation, budget forecasting, unit cost modeling, and the detection of runaway processes that are quietly burning money at three in the morning. A finance function that failed to track it would be failing at its job. On the expense line the number is not only defensible, it is necessary.
The performance line answers a different question. It reports what the organization produced for the people it serves. Cases resolved. Changes shipped. Cycle time. Defects escaped. Volume absorbed without adding headcount. Those measures are harder to obtain, slower to arrive, and frequently shared across functions, which is exactly why they are the ones worth reporting.
The failure I keep encountering is not that leaders confuse these two ideas in the abstract. Ask any executive whether spending is the same as producing and the answer is immediate and correct. The failure is procedural. The consumption chart is built by the team that owns the AI program, and the AI program is asked to report progress, so the chart goes into the progress slide. Nobody made an argument. A number crossed a column boundary because that was the slide it was pasted into, and the claim changed underneath it.
On the expense line the chart says: this is what we spent. On the performance line, unchanged in every particular, it says: this is what we accomplished. Same number. Different assertion. Only one of them has evidence behind it.
What has to be true before it moves columns
A consumption figure can legitimately appear in a performance discussion. It just has to earn the move, and the conditions are specific enough to check in a meeting.
The outcome has to be named before the period starts. An outcome selected afterward, from whatever happened to move, is not a finding. It is the reporting equivalent of drawing the target around the arrow, and it is the most common way honest people produce a dishonest slide. Naming it in advance costs nothing and removes the entire failure mode.
The outcome has to be measured in a system the organization already ran. If the only instrument for the result is one the AI program built for itself, then the program is grading its own work with a tool it designed, and the baseline will be reconstructed rather than recorded. Ticketing systems, financial systems, delivery pipelines and case management systems were all running before any of this started. They have history, and history is what makes a comparison mean something.
The consumption and the outcome have to cover the same period and the same population of work. Consumption reported across the enterprise, set beside an improvement observed in one team, is not a relationship. It is two facts sharing a slide.
And someone has to be willing to state, in advance, what result would count as the effort not working. This is the condition that gets dropped first, because it is the only one with a cost attached. A program that has defined no disconfirming result cannot produce evidence, only illustration. If nothing could have come back negative, nothing that came back positive is informative.
Meet those four and the consumption figure has a place on the performance line, as a denominator. Miss any of them and it is still an expense number, and presenting it as performance is a claim the organization cannot support.
Why finance already knows this and the AI program does not
Finance has been burned into this discipline over a century. It has standards, an audit function, and a professional culture that treats the confusion of cost with value as a category error rather than an optimistic interpretation. Nobody in a finance function would present the utility bill as evidence of manufacturing performance, and if someone tried, the correction would arrive before they finished the sentence.
The AI program has none of that. It is usually new, usually reporting into a leadership team still learning the domain, and usually under pressure to demonstrate return on a substantial investment long before any outcome could plausibly have matured. Its available instrumentation reports consumption by default. Its unavailable instrumentation would report results. Under a deadline, the available number wins, and it wins without anyone deciding to be misleading.
This is not a story about weak people. It is a story about a function that has not yet been given the reporting standards that every other spending function in the organization already has. Robert Solow observed in 1987 that the computer age was visible everywhere except in the productivity statistics, and the gap took years to close. Erik Brynjolfsson and Lorin Hitt examined what eventually closed it and found that the returns to information technology were not the returns to the equipment. They came from the organizational changes built around it, changes that cost real money, took real time, and were largely invisible in conventional measurement. Brynjolfsson, Rock and Syverson revisited the same question for artificial intelligence in 2017 and reached a compatible conclusion: the lag between a general purpose technology arriving and its effect showing up in the numbers is normal, and it is caused by the time required to build the complementary assets that let the technology pay off.
Read that literature in one sentence and it says something uncomfortable and useful. The purchase is not the thing that pays. The reorganization is. A metric that tracks the purchase will therefore look strongest during exactly the period when the least value is being created, which is the period before the reorganization has happened.
There is a second reason to correct this early. Ridgway pointed out in 1956 that quantitative measures reliably generate behavior aimed at the measure rather than at the purpose behind it. A consumption number that people know they are judged by does not stay descriptive. It becomes a target, and the organization starts spending real money to move a figure whose only honest job was to describe spending. The measure degrades as it climbs.
Cost per outcome versus cost per interaction
Once an organization accepts that the raw number belongs to finance, the next question is what the performance line should say instead. Most teams reach first for cost per interaction, and it is worth understanding why that is still the wrong column.
Cost per interaction divides spending by activity. Queries handled, sessions run, documents generated, requests served. It looks like a productivity measure because it is a ratio, and ratios feel rigorous. But both terms describe the inside of the machine. The number improves whenever activity grows faster than spending, which can happen while the organization delivers nothing it was not already delivering. A call center that generates twice as many interactions at the same cost has not necessarily served a single additional customer.
Cost per outcome divides spending by a result the organization was already trying to achieve. Cost per resolved case. Cost per shipped change. Cost per closed engagement. Cost per audit completed. Cost per unit of whatever this organization exists to produce. The denominator is chosen from the business, not from the tool, and it existed before the tool arrived.
The difference is not cosmetic. Cost per outcome moves in the right direction only when the organization gets better at producing the outcome, and it moves in the wrong direction when spending rises without effect. It is capable of delivering bad news, which is the entire reason to trust it. Cost per interaction is not, which is the entire reason it is popular.
Kaplan and Norton made the general form of this argument in 1992. Their point was that no single measure describes performance, and that measures of internal activity have to be balanced against measures of what the customer actually receives. That principle is uncontroversial in every other part of the business. It has simply not yet been applied to the AI line item, because the AI line item arrived recently enough to still be treated as its own category rather than as ordinary spending that has to justify itself in ordinary terms.
A practical note on choosing the denominator. Take it from a number that was already in the operating review before any of this began. If the outcome measure has to be invented to accommodate the AI program, that is a signal worth attending to, because it usually means the program was never attached to a result anyone was accountable for.
What to do with the dashboard you already built
Do not delete it. That instinct is the wrong correction and it will cost the organization something real.
The dashboard is accurate. It took work to build, it reports genuine cost telemetry, and finance needs it. Deleting it removes visibility into a growing expense at precisely the moment that expense is compounding, and it punishes the team that built the one instrument the program actually has. The problem was never the chart. The problem was the slide it was pasted into.
Retitle it first. If it is currently called AI Adoption, or AI Momentum, or Transformation Progress, rename it to what it measures. AI Spend by Team. Capacity Consumed. Cost Telemetry. The rename is not a cosmetic exercise. A title is a claim, and every person who reads the chart reads the title first. Renaming the chart to describe its contents ends the substitution instantly and costs nothing but the awkwardness of one meeting.
Then route it to the function that owns cost. It belongs in the finance review, next to the other things the organization buys, where it will be used for forecasting and negotiation rather than for reassurance. Consumption presented to finance is information. The same consumption presented to a board as progress is a costume.
Then leave the performance slot empty until an outcome measure is ready to fill it. This is the part that requires nerve, and it is the part that separates organizations that recover from this quickly from the ones that do not. A leadership team that can look at a blank space on the AI section of the reporting pack, and say plainly that the outcome measure is not ready yet, has just done something more valuable than any chart it could have shown. It has told the truth about the state of the program, which means the next decision about that program will be made on real ground.
The uncomfortable version of the exercise is that the empty slot frequently stays empty for a while. That is not a failure of the exercise. It is the finding. If an organization cannot name an outcome, obtain a baseline for it, and attach spending to it, then it has not yet built the conditions that would let the technology pay, and no dashboard was going to reveal that as clearly as the blank space just did.
Consumption is not the enemy here, and neither is the technology. What these systems can do is remarkable and the trajectory is difficult to overstate. The mistake is smaller and more ordinary than any argument about capability. A cost number was placed in a value column, nobody objected because nobody had to, and an entire program was evaluated on the strength of its own invoice.
Sources
- Flynn, Dan. Builders Build: The Four A’s of Organizational Readiness. Mission Intelligence Systems LLC. Chapter 12, “What I Kept Finding,” on the pattern this article describes, and Chapter 9, the source of the quotation that adoption is not value and consumption is not even adoption.
- Ridgway, V. F. “Dysfunctional Consequences of Performance Measurements.” Administrative Science Quarterly, vol. 1, no. 2, 1956, pp. 240–247. doi.org/10.2307/2390989.
- Kaplan, Robert S., and David P. Norton. “The Balanced Scorecard: Measures That Drive Performance.” Harvard Business Review, vol. 70, no. 1, 1992, pp. 71–79.
- Brynjolfsson, Erik, and Lorin M. Hitt. “Beyond Computation: Information Technology, Organizational Transformation and Business Performance.” Journal of Economic Perspectives, vol. 14, no. 4, 2000, pp. 23–48. doi.org/10.1257/jep.14.4.23.
- Brynjolfsson, Erik, Daniel Rock, and Chad Syverson. “Artificial Intelligence and the Modern Productivity Paradox: A Clash of Expectations and Statistics.” NBER Working Paper 24001, National Bureau of Economic Research, November 2017. doi.org/10.3386/w24001.
- Solow, Robert M. “We’d Better Watch Out.” New York Times Book Review, 12 July 1987, p. 36. The origin of the productivity paradox observation.
About the Author
Dan Flynn
Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build
Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.
His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.
