Organizational Diagnostics · Attention

The Token Is the New Timesheet

We already learned that counting the input tells you nothing about the output.

Every organization I work with has already survived a version of this. They stopped counting hours at a desk because it produced longer days and no more work. They stopped counting lines of code because it produced more code and worse software. Then a new unit arrived, attached to a technology impressive enough to suspend judgment, and the lesson did not travel.

Published

Key Takeaways

The lesson we already paid for

The timesheet was never a measure of contribution. It was a measure of attendance that organizations agreed to treat as a measure of contribution, because attendance was easy to observe and contribution was not. The consequence was predictable and it took decades to unwind. People stayed later. Presence became a performance. The person who left at a reasonable hour having finished the work looked worse than the person who stayed until eight and finished less. Nobody lied about anything. The hours were real. What was false was the inference laid on top of them.

Software repeated the error in its own dialect. Lines of code were countable, and countable things become metrics whether or not they mean anything. Frederick Brooks warned about this in The Mythical Man-Month in 1975, in the same book that explained why adding people to a late project makes it later. A line of code is a cost of production, not a unit of output. Measuring a programmer by lines written rewards verbosity and penalizes the deletion of code, which is frequently the most valuable thing an engineer does all week. The industry eventually accepted the point, and it accepted it expensively.

Call centers ran the same experiment on a different axis. Average handle time is a clean number, easy to collect and easy to compare, and organizations that made it the primary measure got exactly what they asked for: shorter calls. They also got customers who called back three times about the same problem, because the fastest way to end a call is to end it before the problem is solved. Handle time went down and total cost went up.

None of this was a surprise even at the time. V. F. Ridgway published a short paper in Administrative Science Quarterly in 1956 titled “Dysfunctional Consequences of Performance Measurements,” which argued that quantitative measures reliably produce behavior optimized for the measure rather than for the purpose behind it. Donald Campbell reached the same conclusion from a different direction in 1979. Marilyn Strathern supplied the phrasing most people can quote: when a measure becomes a target, it ceases to be a good measure. W. Edwards Deming spent an entire career arguing that management by visible figures alone is one of the deadly diseases of Western management, precisely because the figures that matter most are frequently the ones nobody has instrumented.

So this is not new knowledge. It is knowledge the organization already bought and already paid for.

Why the new unit slips past the filter

Here is the part worth sitting with. The executives now approving token dashboards are, in many cases, the same executives who dismantled the timesheet culture and who would laugh at a proposal to rank engineers by keystrokes. Their judgment did not deteriorate. The pattern simply arrived wearing clothes they did not recognize.

A token is an unfamiliar unit. It sounds technical. It belongs to a domain most senior leaders are still learning, and the natural posture in an unfamiliar domain is deference. When someone presents a chart of monthly token consumption by team, the executive is not evaluating a measurement proposal. They believe they are being briefed on an emerging technology. The frame is “here is how AI adoption is progressing,” not “here is a proposal to rank your organization by an input.”

Say it in the old unit and the room reacts instantly. Propose ranking the engineering organization by keystrokes per developer per week and you will be interrupted before you finish the sentence. Propose ranking the same organization by tokens per developer per week and the conversation is about which chart to use. The two proposals are the same proposal. The second one just does not trip the alarm, because the subject appears to be artificial intelligence rather than counting.

There is a second reason the filter does not fire, and it is structural rather than psychological. Consumption data arrives for free. Every AI platform reports usage by default, in fine granularity, attributable to teams and individuals, updated daily. Outcome data requires somebody to define an outcome, agree on it across functions, instrument it, and wait. Faced with a board meeting in three weeks and a substantial invoice already paid, the available number wins. That is not weakness. It is a rational response to a reporting cycle that demands evidence before evidence exists.

What the number actually is

A token count is a cost-per-unit-of-consumption figure. That is its complete and honest description. It tells you how much of a purchased capacity was drawn down, by whom, over what period. On a utility bill this would be uncontroversial. Nobody presents electricity consumption to a board as evidence of manufacturing performance, and nobody would accept it if they did, because everyone understands intuitively that a factory can burn a great deal of power producing scrap.

Consumption belongs on the expense line. It is genuinely useful there. It supports capacity planning, vendor negotiation, budget forecasting, and the detection of runaway processes. A finance function that does not track it is not doing its job. None of that is in dispute.

The failure is a filing error with consequences. The number gets moved from the expense line to the performance line, and in the move it silently changes what it claims. On the expense line it says: this is what we spent. On the performance line it says: this is what we accomplished. Nothing about the number changed. Everything about the claim did.

A team consuming enormous capacity might be transforming the business. It might also be generating analyses nobody has the authority to act on, drafts nobody has the capacity to review, and code that will be maintained by whoever inherits it. In my field engagements the most consistent finding is not that high consumption teams are unproductive. It is that consumption and contribution turn out to be almost entirely uncorrelated once you go looking, and that nobody had checked before the dashboard went up.

What happens once people know they are measured by it

Everything above describes the metric as a passive observation. It stops being passive the moment people learn it exists.

This is the mechanism Ridgway described and Strathern named, and it is not a story about dishonest employees. It is a story about rational ones. When an organization publishes what it counts, it has told people what it values, and capable people respond to that signal quickly and without malice. A manager whose team shows low consumption now has something to explain in a review. The cheapest way to stop having something to explain is to consume more.

So queries get run that would not otherwise have been run. Work that a competent person could complete directly in ten minutes gets routed through a tool so that it registers. Drafts get regenerated rather than edited, because regeneration produces a number and editing does not. A team that quietly figured out how to solve a class of problem without AI at all, which is a genuine and valuable finding, now has an incentive to keep that finding to itself.

None of that is fraud. Every individual action is defensible on its own terms. The aggregate is an organization spending real money to move a number that was only ever meant to describe spending. And the metric degrades exactly as it climbs: the higher it goes, the less it tells you, because more of its movement is now explained by the fact that it is being watched.

The timesheet did this. Lines of code did this. Handle time did this. There is no version of this pattern where the new unit is the exception.

The calendar test applied to AI

There is a test for all of this, and it does not require new instrumentation.

In Builders Build, Chapter 11 is called “The Calendar Never Lies.” The argument is that an organization’s stated priorities and its actual priorities are two different documents, and only one of them is auditable. Strategy decks describe intent. Calendars describe what happened. When the two disagree, the calendar is correct, because a calendar cannot be aspirational after the fact.

Because the calendar is not a scheduling document. The calendar is a confession.

From Builders Build, Chapter 11

This lineage is not an invented parallel. Chapter 9 of the same book names AI theater directly as priority theater’s descendant, and defines priority theater as what happens when the document says one thing and the calendar says another. The pattern was already identified. What follows is simply the same test pointed at a newer dashboard.

The token dashboard is the official story. It is the strategy deck of the AI program: produced deliberately, presented upward, describing intent and effort. The outcome record is the confession. It is what actually happened, and it is available in systems the organization already runs. Cycle times. Defect and rework rates. Cases resolved. Time from request to decision. Cost to serve. Volume absorbed without adding headcount.

Put the two side by side over the same period. If consumption climbed steeply and the outcome record is flat, you have your answer, and you did not need a consultant to reach it. If consumption climbed and cycle time fell, you have something real, and now you know where and can go find out why. Either way you have replaced an assertion with evidence.

The uncomfortable version of this test is that most organizations have never run it. The consumption chart is presented, the outcome record sits untouched in another system, and the two are never placed on the same page. That gap is not an accident of tooling. It is the gap the theater lives in.

What to put on the dashboard instead

Keep the consumption number. Move it.

On the expense line it stays exactly as it is, tracked by finance, used for forecasting and negotiation. On the performance line it appears only as a denominator, attached to something the organization was already trying to achieve. Cost per resolved case. Cost per shipped change. Cost per closed engagement. Cost per unit of whatever this organization exists to produce. A consumption figure divided by an outcome figure is a productivity measure and it improves when the organization gets better. A consumption figure alone improves when the organization spends more.

Kaplan and Norton made the general version of this argument in 1992 when they introduced the balanced scorecard, and their core observation applies directly: no single measure can describe performance, and measures that describe internal activity have to be balanced against measures that describe what customers actually receive. That is not a new idea for anyone running a business. It just has not yet been applied to the AI line item.

Choosing the right outcome metric, and defending it against the pressure to report something sooner, is its own problem with its own failure modes. That is the next article. For now the useful move is smaller and available immediately: find the consumption chart in your reporting pack, and ask what it would be divided by if it were honest.

Sources

  1. Flynn, Dan. Builders Build: The Four A’s of Organizational Readiness. Mission Intelligence Systems LLC. Chapter 11, “The Calendar Never Lies,” and Chapter 9, where AI theater is named as priority theater’s descendant.
  2. Ridgway, V. F. “Dysfunctional Consequences of Performance Measurements.” Administrative Science Quarterly, vol. 1, no. 2, 1956, pp. 240–247. doi.org/10.2307/2390989.
  3. Campbell, Donald T. “Assessing the Impact of Planned Social Change.” Evaluation and Program Planning, vol. 2, no. 1, 1979, pp. 67–90. doi.org/10.1016/0149-7189(79)90048-X.
  4. Strathern, Marilyn. “‘Improving Ratings’: Audit in the British University System.” European Review, vol. 5, no. 3, 1997, pp. 305–321. Source of the widely quoted formulation that when a measure becomes a target, it ceases to be a good measure. doi.org/10.1002/(SICI)1234-981X.
  5. Brooks, Frederick P. The Mythical Man-Month: Essays on Software Engineering. Addison-Wesley, 1975. Source of the argument that lines of code measure activity rather than progress.
  6. Deming, W. Edwards. Out of the Crisis. MIT Press, 1986. On management by visible figures alone as a deadly disease of Western management.
  7. Kaplan, Robert S., and David P. Norton. “The Balanced Scorecard: Measures That Drive Performance.” Harvard Business Review, vol. 70, no. 1, 1992, pp. 71–79.
DF

About the Author

Dan Flynn

Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build

Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.

His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.