Organizational Diagnostics · Alignment

Measure the Outcome, Not the Interaction

The replacement metrics, and what each one costs to collect.

Complaining about a measure is cheap. Anyone can point out that token counts and active-user percentages describe expense rather than performance. The harder and more useful thing is to name what goes in their place, and to be honest about what each replacement costs to collect, because a measure presented as free is a measure nobody will maintain past the second quarter.

Published

Key Takeaways

Why the replacement has to be specific

Consumption metrics do not survive because nobody has noticed they are weak. Most executives reporting them already suspect it. They survive because the person holding the reporting obligation has a date on the calendar and needs a number before that date arrives. Remove the number and offer nothing in its place, and the old number returns within a quarter, usually with a new name on it.

So the replacement has to be concrete enough to be built by a specific person on a specific system, and it has to arrive with its price tag attached. Every measure below is stated three ways: what it is, what it costs to collect, and what it is vulnerable to. The third part matters as much as the first. A measure whose weaknesses are known can be defended. A measure sold as clean and free gets adopted enthusiastically, degrades quietly, and takes the credibility of the whole effort down with it.

One more constraint before the list. None of these is a single number that settles the question. Kaplan and Norton made that argument in 1992 for the same reason it applies here: a scorecard is a small set of measures chosen to show the same strategy from more than one angle, precisely because any one of them promoted to a verdict starts distorting the behavior underneath it. Deming had said something adjacent and harsher in 1986, warning that running an organization on visible figures alone ignores the figures that are unknown and unknowable and often matter most. Both cautions apply to everything that follows.

The five measures

1. Cycle time to a decision that sticks

Measure the elapsed time from the moment a decision is first framed to the moment it is made and stays made. The second half of that sentence is where the value is. A decision sticks if it is not reopened within a window you define in advance, typically sixty or ninety days, absent genuinely new external information. Reopening because a competitor moved or a regulator ruled is legitimate and should not count against the measure. Reopening because the decision never had real authority behind it is exactly what you are trying to detect.

What it costs to collect. Moderate. It requires a decision log with two timestamps and a named owner who reviews the log at the end of each window. Realistically that is a few minutes per decision at entry and an hour a month for the review. The cost is not the tooling, which is a spreadsheet. The cost is that somebody has to own it and keep owning it after the novelty wears off.

What it is vulnerable to. Three things. Definition drift, where the boundary of what counts as a decision quietly widens to include easy calls. Suppression, where people avoid formally reopening a bad decision because reopening it would damage the number, which is the worst possible outcome because it converts a measurement into a reason to leave a mistake standing. And late framing, where the clock does not start until the decision is nearly settled, which produces excellent cycle times and tells you nothing.

2. Rework rate

Measure the share of AI-generated output that required material correction or was discarded before use. Material has to be defined before you start collecting, not after you see the first result. A useful threshold is whether a human had to change the substance rather than the surface. Fixing formatting is not rework. Rewriting the analysis because the premise was wrong is.

What it costs to collect. This is the most expensive of the five, because it requires a human judgment at the point of use rather than a count from a system. Do not attempt a census. Sample instead: pull a random set of outputs per period, have someone other than the author classify them, and accept the wider error bars that come with sampling. A weekly sample of twenty items reviewed by a rotating pair is sustainable. Tagging every artifact is not, and organizations that try it produce three weeks of excellent data and then nothing.

What it is vulnerable to. Self-review above all. If the person who generated the output classifies it, the rate will be low and meaningless. It is also vulnerable to being read in one direction only. A rework rate near zero is not a triumph, it is a signal that the organization is either using AI for trivial work or accepting output without examining it. Both are findings. Neither is success.

3. Decision conversion

Divide the number of decisions actually made in a period by the number of decisions generated in the same period. Analysis that surfaces a genuine choice generates a decision. Someone with the authority to commit resources making that call converts it. The ratio between the two is the single most revealing number in this article.

What it costs to collect. Low, which is unusual for a measure this informative. It is two counts against a list you probably already keep in some form, whether that is a steering committee agenda, a decision log, or the standing set of open items in a leadership meeting. There is no instrumentation to build. The cost is definitional discipline in the first month while people agree on what counts.

What it is vulnerable to. The denominator. When the ratio starts being watched, the easiest way to improve it is to stop logging generated decisions, which shrinks the denominator without changing anything real. It is also vulnerable to inflation of the numerator, where a decision is recorded as made when what actually happened was a meeting agreeing to discuss it again. Guard both by defining a made decision as one with a named owner and a committed resource, and by keeping the generation count owned by someone other than the person the ratio reflects on.

4. Customer-visible change

State one specific claim that someone outside the organization would notice, in language they would recognize. Not improved efficiency. Something closer to: the median time from a customer request to a delivered answer moved from nine days to four, or we now answer this category of question on the first contact rather than the third. If nobody outside can perceive it, it is an internal rearrangement, which may still be worthwhile but is not evidence of transformation.

What it costs to collect. Low to collect and high to earn. The underlying data usually exists already in service levels, cycle times, defect rates or support volumes, so the extraction cost is small. The real cost is the discipline required to state a claim narrow enough to be falsifiable, which is a political cost rather than an operational one.

What it is vulnerable to. Attribution, and there is no clean way around it. Many things change at once, and the AI program is rarely the only variable. Do not paper over this with statistical machinery you cannot defend. State the claim, state the confounders in the same breath, and let the reader weigh them. A modest claim with its confounders named survives scrutiny. A large claim presented as clean attribution does not, and when it collapses it takes the other four measures with it.

5. Cost per outcome

Put a completed outcome in the denominator rather than an interaction. Cost per resolved case, per shipped change, per closed finding, per delivered analysis that someone acted on. The numerator has to be fully loaded, which means it includes the human review time and the rework, not just the platform invoice.

What it costs to collect. Moderate, and it requires finance to participate rather than merely approve. The hard part is not arithmetic, it is agreeing on the outcome unit, because that argument surfaces genuine disagreement about what the function is for. Budget several weeks for that conversation and treat it as valuable rather than as delay.

What it is vulnerable to. Unit shopping, where the outcome unit is quietly chosen to be whichever one makes the number look best, and changed again when it stops flattering. It is also vulnerable to excluding review time from the numerator, which is the single most common error, because review time is precisely where the cost migrated when generation got cheap. A cost-per-outcome figure that omits human review is a cost-per-interaction figure with a better title.

If you can only build one

Build decision conversion.

It wins on three grounds. It is the cheapest, requiring two counts rather than new instrumentation or a sampling protocol. It is a ratio, which means it cannot be improved by simply doing more of the thing, and the entire failure mode of consumption metrics is that doing more improves them. And it detects the specific problem that AI introduces, which is that generation capacity rises quickly while the capacity to convert options into commitments does not move at all. Authority does not scale by license.

The practical build is small enough to start this month. Pick one leadership body. For one quarter, keep two counts: how many decisions were put in front of it, and how many left with an owner and a committed resource. Do nothing else. If the ratio is falling while AI usage rises, you have found the bottleneck, and it is not the technology. It is the number of people permitted to say yes.

What to keep from the old dashboard

Do not delete the consumption panel. The numbers on it are accurate, they were always accurate, and they answer a question that genuinely matters: what does this program cost, and are we paying for capacity nobody is using? Seat utilization is a real procurement signal. Token spend is a real budget line. Neither was ever dishonest.

The error was placement, not arithmetic. So change the placement. Retitle the panel from adoption to spend and utilization, move it next to the other cost lines where it belongs, and stop citing it in any sentence about performance. That single relabeling does most of the work, because it removes the unexamined inference that rising consumption implies rising capability without discarding anything useful. It also spares you the argument with whoever built the dashboard, since you are not telling them their work was wrong. You are telling them it was filed under the wrong heading.

Choosing the measure is choosing the attention

Chapter 13 of Builders Build draws a distinction that governs this whole subject. Awareness is passive: things arrive in front of you and you register them. Attention is chosen: out of everything you are aware of, some small set gets your organization’s time, and that choice is made whether or not anyone admits to making it. The diagnostic question that follows is whether an organization’s attention actually reflects what matters, or whether it has drifted toward what is merely visible, urgent and loud.

A consumption dashboard is visible, urgent and loud by construction. It updates daily, it always has a number, and the number usually moves in a direction someone can present. The five measures above are none of those things. They resolve slowly, several require judgment, and two of them will occasionally deliver news nobody wanted.

Which is the point. Whatever appears on the dashboard is what gets discussed, and whatever gets discussed is what the organization is actually paying attention to, regardless of what the strategy document says. Choosing the measure is choosing the attention. It is an attention allocation decision arriving in the costume of a reporting decision, and treating it as merely administrative is how organizations end up with their most senior people spending an hour a month on a chart that describes their own spending.

These will be gamed too

Nothing here is exempt from the pattern that broke the old measures. Ridgway documented in 1956 that quantitative performance measures reliably produce behavior optimized for the measure rather than the purpose. Campbell restated it in 1979 from the evaluation side. Strathern gave it the phrasing everyone remembers, that a measure which becomes a target ceases to be a good measure. Those findings apply to decision conversion and rework rate exactly as they applied to lines of code and hours at a desk. Assume they will be gamed, because they will be.

The defense is not a cleverer metric. It is structure, and it has two parts. First, pair each measure with one that moves the wrong way when the first is being manipulated. Decision conversion pairs with cycle time to a decision that sticks, because converting faster by rubber-stamping shows up as reopened decisions. Rework rate pairs with customer-visible change, because suppressing rework counts shows up as nothing improving outside. Cost per outcome pairs with rework rate, because shrinking the numerator by cutting review inflates the rework the following quarter. A pair is harder to game than a number, since it requires manipulating two things in opposite directions at once.

Second, put the measures themselves on a review cycle. Once a year, someone should be accountable for asking a different question than how are we doing on these. The question is what behavior is each of these now producing, and is any of it behavior we did not intend. Answering it honestly will eventually retire a measure that has stopped working, which is a sign the system is functioning rather than a sign it failed. The organizations that get this wrong are not the ones whose metrics decayed. Every set of metrics decays. They are the ones who never scheduled the conversation that would have noticed.

Sources

  1. Flynn, Dan. Builders Build: The Four A’s of Organizational Readiness. Mission Intelligence Systems LLC. Chapter 13, “The First A: Attention,” on the distinction between passive awareness and chosen attention, and the diagnostic question of whether an organization’s attention reflects what matters or what is visible, urgent and loud.
  2. Ridgway, V. F. “Dysfunctional Consequences of Performance Measurements.” Administrative Science Quarterly, vol. 1, no. 2, 1956, pp. 240–247. doi.org/10.2307/2390989.
  3. Campbell, Donald T. “Assessing the Impact of Planned Social Change.” Evaluation and Program Planning, vol. 2, no. 1, 1979, pp. 67–90. doi.org/10.1016/0149-7189(79)90048-X.
  4. Strathern, Marilyn. “‘Improving Ratings’: Audit in the British University System.” European Review, vol. 5, no. 3, 1997, pp. 305–321. Source of the widely quoted formulation that when a measure becomes a target, it ceases to be a good measure.
  5. Kaplan, Robert S., and David P. Norton. “The Balanced Scorecard: Measures That Drive Performance.” Harvard Business Review, vol. 70, no. 1, 1992, pp. 71–79. The argument for a small paired set of measures rather than a single number promoted to a verdict.
  6. Deming, W. Edwards. Out of the Crisis. MIT Press, 1986. On the hazard of managing by visible figures alone, when the figures that matter most are often unknown and unknowable.
DF

About the Author

Dan Flynn

Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build

Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.

His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.