Why Measuring Effort Feels Safer Than Measuring Outcomes
Input metrics are legible, immediate, and belong to somebody. Outcomes are none of those things.
It is easy to look at a dashboard full of activity counts and conclude that whoever built it does not understand measurement. That conclusion is comfortable, and it is almost always wrong. The people building those dashboards can usually explain, unprompted, why activity is a poor proxy for value. They built it anyway, because of what they were asked for and when they were asked for it.
Published
Key Takeaways
- Input metrics win because they are defensible. They exist today, resolve to one number, and attach to one team. Outcome metrics are lagging, shared across functions, and contestable, which makes them dangerous to carry into a review.
- This is not a failure of intelligence. Asked to prove an investment on a reporting cycle that arrives before any outcome could mature, a rational person reaches for the number that exists. Calling that stupidity guarantees it continues.
- The fix sits upstream of the dashboard. Change what the organization asks for and when, name the maturation window before the work starts, and remove the penalty for reporting an honest flat line.
The defensibility problem
Start with the actual decision a manager faces, because most commentary on metrics skips it.
A leader has three weeks until a review. An investment has been made, in AI or in anything else, and the review will ask what came of it. The leader now has to select a number to put on a slide. That number has to survive being questioned by people who were not involved in the work, who have competing claims on the same budget, and who will remember the figure longer than the caveats attached to it.
Judged against that requirement, an input metric is an excellent instrument. It has three properties that matter enormously and that almost nobody names out loud. It is legible: it exists right now, in a system that already collects it, with no definitional work required. It is singular: it resolves to one figure that means the same thing to everyone looking at it. It is attributable: it belongs unambiguously to one team, which means someone can be credited for it and, just as importantly, someone can be held to it.
An outcome metric fails all three tests on the day the slide is due. It is not available yet. It does not resolve cleanly, because it usually requires a definition somebody has to agree to. And it does not belong to anyone in particular, because outcomes are produced jointly.
So the choice is not between a good metric and a bad one. In that room, on that day, it is between a number that can be presented and a number that cannot. Framed that way, the widespread preference for consumption charts stops looking like a failure of understanding and starts looking like what it is, which is a competent response to a badly formed question.
Why lagging and shared metrics feel dangerous to report
Take the three failures in turn, because each produces a distinct kind of exposure and the exposure is what actually drives the behavior.
Lagging means the reporter has to stand in front of a flat line. Real outcomes mature on the organization’s clock, not the reporting calendar’s. Cycle time does not fall the month a tool is deployed. Error rates do not improve until the process around them changes. Cost to serve moves after the work is reorganized, not after the license is bought. Every one of those is a normal, expected shape. But a flat line presented in a review is not read as a normal shape. It is read as an absence of progress, and the person presenting it is the person standing next to the absence.
Shared means the credit is diffuse and the blame is not. An outcome like time from customer request to resolution is produced by sales, operations, engineering, and support acting together. If it improves, four functions have a plausible claim to the improvement. If it does not, the function that put the number on the slide is the one the number is now associated with. That asymmetry is obvious to anyone who has sat through a quarterly review, and it is sufficient on its own to keep outcome measures off the page.
Contestable means the conversation moves. An input number ends discussion, because there is nothing to argue with. An outcome number invites a peer to propose a different baseline, a different definition, a different attribution window, or a different denominator. None of those objections have to be made in bad faith to be damaging. Once the meeting is discussing methodology, it has stopped discussing the result, and the reporter has lost the room regardless of what the underlying work achieved.
Kahneman and Tversky documented the general shape of this in 1979: people weigh potential losses considerably more heavily than equivalent gains. A manager choosing between metrics is not weighing accuracy against inaccuracy. They are weighing a modest upside, being seen as rigorous, against a concrete downside, standing in front of a flat and contestable number in a room full of peers. The asymmetry in that weighing is not irrationality. It is the well documented default behavior of human beings under evaluation.
Chris Argyris described the organizational consequence in 1977. Organizations get very good at single loop learning, which is adjusting behavior to better hit the existing target, and remain very bad at double loop learning, which is questioning whether the target is the right one. Asking whether the metric itself is the wrong metric is a double loop question. It requires someone to challenge the frame in a meeting that was convened to report inside the frame, and the people most able to see the problem are usually the people most exposed by naming it.
The moth pattern
There is a version of this that operates below the reporting layer, in what individuals learn to be rewarded for. I wrote about it in Builders Build as the moth effect.
A pattern of running hard toward every fire and away from the patient, invisible work of building something that would not keep producing fires. A pattern of substituting velocity for
From Builders Build, Chapter 8
Substituting velocity for direction. That is the completion, and the reason the substitution keeps happening is that velocity is visible and direction is not. The person who runs at the fire is seen running. Their effort is observed in real time, by an audience, and it is recognized within days. The person who spent the same period removing the conditions that produce fires has nothing to show at the end of it, because the deliverable is the absence of an event. Nobody schedules a review to discuss the incident that did not occur.
Consumption metrics are the visible fire. They are the organizational equivalent of being seen running: immediate, observable, and reliably recognized. Outcome metrics are the patient invisible work: real, more valuable, and structurally disadvantaged in every forum where recognition is distributed.
This is why the problem does not yield to better information. A manager can fully understand that the consumption chart is a poor measure and still, correctly, observe that the organization has never once rewarded anyone for producing the other kind of evidence. The moth is not confused about what fire is. It is responding to the only signal strong enough to navigate by.
What the reporting cycle actually demands
Look closely at the request itself, because that is where the defect lives.
The request is usually some version of: show me the return on this investment at the next review. Embedded in it is an assumption nobody states, which is that a return exists at the next review in a form that can be measured. For most substantial changes, it does not. The organization has asked for evidence on a timeline that precedes the evidence, and then treated the resulting substitution as a reporting choice rather than as a consequence of its own question.
Herbert Simon explained the underlying dynamic in 1971, in an argument about information that has aged into something close to prophecy. A wealth of information creates a poverty of attention, and the scarce resource in an information rich environment is the attention of the people who have to consume it. Consumption metrics are abundant and cheap. They arrive automatically, in fine granularity, without anyone having to decide anything. Outcome metrics are scarce and expensive, because each one requires a definition, an agreement across functions, and instrumentation that somebody has to fund and maintain.
An attention constrained review defaults to the abundant input. Not because anyone chose abundance over relevance, but because the abundant number was already on the page and the relevant one would have required work that nobody was tasked with and nobody had time to do.
The pattern then hardens. Once a consumption figure has been reported successfully, it becomes the expected figure. The next review asks how it moved. At that point the metric has become a target, and Ridgway in 1956, Campbell in 1979, and Strathern in 1997 all describe what happens next. The number continues to rise and its usefulness as evidence continues to fall, because more and more of its movement is explained by the fact that it is being watched.
The cost of being right too late
There is a defense of input metrics that deserves an honest hearing, because in my experience it is what people actually believe even when they do not say it.
The defense is this: an outcome measure that arrives eighteen months from now is worthless to a decision that has to be made this quarter. A leader who insists on measuring only what matters, and therefore reports nothing for four consecutive quarters, does not get to be vindicated in the fifth. The program gets cancelled in the third, the budget goes to a peer who produced charts, and the eventual proof that the approach was correct lands on a program that no longer exists.
That is not a hypothetical. It is the most common way I see rigorous measurement get punished. The person who refused to report activity was right about measurement and wrong about the organization, and being right about measurement is not a survivable position on its own.
This is what makes moralizing about metrics so ineffective. Telling a manager that consumption is a poor measure supplies information they already have, and supplies nothing at all about the problem they actually face, which is what to put on the slide in three weeks without losing the program. Any recommendation that does not address the three week problem will be nodded at and then ignored, correctly.
So the useful move is not to demand outcome purity. It is to make an honest interim report survivable. That means naming the outcome and its expected maturation window before the work begins, so that a flat line at week six is a reading the organization predicted rather than a failure the reporter has to absorb. It means accepting leading indicators explicitly as leading, on the record, with the outcome they stand in for named beside them, so the substitution is visible instead of silent. A leading indicator labeled as one is an honest instrument. The same number presented as a result is the whole problem.
Changing what gets asked for
Every fix worth having sits upstream of the dashboard, in the request rather than the response.
Define the outcome before the investment, not after. The moment to agree on what success will look like is while the money is being approved, when there is appetite for the conversation and no one is yet exposed by the answer. An outcome defined at that moment is a shared commitment. The same outcome proposed nine months later, in a review, is an argument about whether the program worked, and it will be litigated as one.
Publish the maturation window with the outcome. If cycle time is the measure and cycle time will not move for two quarters, then two quarters of flat readings are the expected result and should be written down as such in advance. This single move removes most of the personal exposure that drives the substitution, because the reporter is no longer standing next to an unexplained absence. They are reporting a result the organization already agreed to expect.
Assign the shared metric to a person anyway. Outcomes are produced jointly, which is why they end up owned by nobody and reported by nobody. Someone has to be accountable for reporting the number without being solely accountable for producing it. Those are different assignments and organizations routinely collapse them, which is precisely why cross functional outcomes go unmeasured.
Make the absence of an outcome measure a reportable finding. If a program cannot name an outcome it expects to move, that is information worth having and it belongs in the report as a line item rather than being papered over with a consumption chart. The most valuable sentence in many progress reviews is a version of: we have no outcome measure for this yet, here is when we will, and here is the leading indicator we are using until then. That sentence is currently unsayable in most organizations. Making it sayable is most of the work.
And remove the penalty for the honest flat line. As long as reporting no movement is more costly than reporting activity, capable people will keep reporting activity, and they will be right to. Measurement education does not survive contact with a review that punishes candor. The condition has to change first.
None of this requires anyone to become smarter about metrics. It requires the organization to stop asking a question that can only be answered dishonestly, and to start asking one that a careful person can answer on the record without exposure. Managers reaching for the number that exists are behaving rationally inside the conditions they were given. Change the conditions and the behavior changes with them, which is the whole argument for treating this as a design problem rather than a character problem.
Sources
- Flynn, Dan. Builders Build: The Four A’s of Organizational Readiness. Mission Intelligence Systems LLC. Chapter 8, “The Moth Effect,” on running toward visible fires and away from the patient, invisible work that would prevent them.
- Ridgway, V. F. “Dysfunctional Consequences of Performance Measurements.” Administrative Science Quarterly, vol. 1, no. 2, 1956, pp. 240–247. doi.org/10.2307/2390989.
- Campbell, Donald T. “Assessing the Impact of Planned Social Change.” Evaluation and Program Planning, vol. 2, no. 1, 1979, pp. 67–90. doi.org/10.1016/0149-7189(79)90048-X.
- Strathern, Marilyn. “‘Improving Ratings’: Audit in the British University System.” European Review, vol. 5, no. 3, 1997, pp. 305–321. Source of the widely quoted formulation that when a measure becomes a target, it ceases to be a good measure.
- Simon, Herbert A. “Designing Organizations for an Information-Rich World.” In Computers, Communications and the Public Interest, Johns Hopkins Press, 1971, pp. 37–72. On a wealth of information creating a poverty of attention.
- Argyris, Chris. “Double Loop Learning in Organizations.” Harvard Business Review, vol. 55, no. 5, 1977, pp. 115–125. On the difficulty organizations have questioning the frame rather than optimizing inside it.
- Kahneman, Daniel, and Amos Tversky. “Prospect Theory: An Analysis of Decision Under Risk.” Econometrica, vol. 47, no. 2, 1979, pp. 263–291. doi.org/10.2307/1914185.
About the Author
Dan Flynn
Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build
Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.
His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.
