Leadership · Attention

Twelve Hours of Tokens and Nothing Shipped

The person maximizing consumption is responding rationally to the scoreboard you handed them.

Someone in your organization sits down before seven in the morning and stands up after seven at night, and in between they drive as much AI capacity as one person can drive. Their usage numbers are the best on the team. If you asked them what actually reached a customer this week, the honest answer would be uncomfortable for both of you. This article is not about that person. It is about the scoreboard they were handed, and about whoever handed it to them.

Published

Key Takeaways

The day itself

It starts early, because the quiet hours are the productive ones. Four sessions open by half past seven. A refactor running in one, a market analysis in another, a set of draft policy documents in a third, and something exploratory in the fourth that might turn into a proposal if it holds together. Everything is moving. Everything is producing. The work is real work, done attentively, by someone who cares about doing it well.

By eleven there is more output than there was capacity to review, so review gets deferred. That is not laziness. It is arithmetic. One person directing four parallel streams of generation will always produce faster than one person can read, and the person knows this, and keeps going anyway, because stopping to read would show up as a gap in the only number anyone has asked about.

The afternoon is more of the same at a higher intensity. Around six there is a stretch that feels genuinely good, where a hard problem comes apart cleanly and something valuable exists that did not exist that morning. That hour is not theater. It is the reason the person is in this job.

The day ends somewhere after seven. The consumption chart is excellent. The refactor is waiting on a review that nobody has time to do. The market analysis produced eleven recommendations and no owner for any of them. The policy drafts are in a folder. The exploratory work is promising and unread. Nothing shipped, and the person closing the laptop is aware of that in a way the dashboard is not.

I want to be exact about the tone here, because the reflex in most executive conversations is contempt, and contempt is both cruel and analytically wrong. Nothing in that day was slacking. It was twelve hours of hard, directed, cognitively expensive work performed by someone taking the instruction seriously.

What the scoreboard asked for, and what it got

The organization did something specific. It published a measure of AI consumption, put it in a dashboard, and attached recognition to movement in that measure. Perhaps not formally. Formality is not required. It is enough that consumption is the number named first in the weekly review, or the number that appeared in the slide about who is leading on adoption.

Once that happens, the measure stops being an observation and becomes an instruction. The instruction was legible, it was quantified, and it was the only quantified thing in the room. The person who moved it furthest was not exploiting a loophole. They were the most obedient person on the team.

This is the ordinary behavior of measurement, and it has been documented for seventy years. V. F. Ridgway wrote in 1956 that quantitative measures, however well intended, reliably produce behavior optimized for the measure rather than for the purpose behind it. Donald Campbell arrived at the same conclusion from social program evaluation in 1979. Marilyn Strathern supplied the phrasing most people know, that when a measure becomes a target it ceases to be a good measure.

None of those authors were writing about dishonest people. That is the part executives consistently skip. The dysfunction Ridgway described does not require anyone to cheat. It only requires people to pay attention to what they are judged by, which is the behavior every organization also claims to want.

So the scoreboard asked for consumption and got consumption. It worked exactly as designed. The design was the problem.

We have run this experiment before

There is an older version of this trap, and most organizations reading this already survived it once.

For decades, the standard proxy for contribution was presence. Hours at a desk, arrival before the manager, departure after. Everyone eventually agreed this was a poor measure, and the agreement usually came with an insult attached, as though presenteeism were caused by employees who wanted credit without doing work.

That was never what it was. Presenteeism was an organizational failure, not a personal one. It happened because contribution is genuinely hard to see, especially in knowledge work, and presence is easy to see. Faced with something important and invisible next to something trivial and visible, organizations counted the visible thing. Then people supplied the visible thing, because they are not fools.

W. Edwards Deming spent the last part of his career arguing that most of what looks like individual performance is produced by the system the individual works inside, and that management by visible numbers alone leads organizations to blame people for outcomes the system determined. His point was not that people never underperform. It was that when you see a whole population behaving a certain way, the population is not the variable. The system is.

Consumption metrics reproduce the structure precisely. Contribution from AI-assisted work is hard to see. Consumption is trivially easy to see, because the platform reports it. So the organization counts what the platform reports and calls it adoption. The unit changed from hours to tokens. The category of error did not change at all.

The twelve-hour day described above is presenteeism with better tooling. It is presence, expressed in a unit that did not exist five years ago, produced by a person who would have stayed late at a desk in 1996 for exactly the same reason.

What it costs the person

Most writing on bad metrics goes straight to the cost to the business. That ordering is a mistake, because the damage does not land there first.

It lands on the person, and it lands in three stages.

The first is exhaustion, and it is not the ordinary kind. Directing several streams of generated work at once is continuous decision-making without recovery. There is no natural pause, no waiting on a build, no walk to a colleague’s desk. Every completion arrives with another prompt available immediately. The day has no slack in it anywhere, and slack is where thinking normally happens.

The second is quieter and worse. The person knows the work is not landing. They know the refactor is unreviewed, the analysis is unowned, the documents are unread. Nobody has told them. The dashboard is congratulating them. But they can see their own output ending nowhere, and there is no event to point at, no failed launch, no missed date, nothing that would let them raise it as a problem without it sounding like an excuse for a number that already looks good.

Raising it requires believing you can say something uncomfortable without cost, which is precisely the condition Amy Edmondson identified in 1999 as the difference between teams that learn and teams that repeat. In its absence, the person says nothing and works harder, because working harder is the only move available that does not require anyone else to agree with them.

The third stage is the one that ends careers quietly. Given enough months, the person stops expecting the work to land. They stop treating the gap between output and outcome as a problem to solve and start treating it as the nature of the job. The high consumption continues. The care behind it does not. That is how an organization converts a builder into someone who produces volume, and it happens without a single conversation in which anyone says anything untrue.

How capable people get written off

Here is where the damage becomes permanent, and where I would ask any executive reading this to be honest about their own past judgments.

I wrote about a specific person in Builders Build, in the chapter about what people actually want from work.

Her manager had told me, without apparent irony, that she was capable but not leadership material. That she did not take initiative. That she was fine where she was.

From Builders Build, Chapter 5

Every one of those three judgments was an observation about conditions, delivered as a verdict about a person. She did not take initiative because initiative in that environment produced friction and no result. She was fine where she was because the alternative had never been made available to her. She was not leadership material in the sense that nothing about the way she was being measured would ever have produced evidence of leadership.

The measurement regime is part of the conditions. That is the connection worth holding onto. When the only number attached to a person describes what their effort cost, every judgment formed about that person is formed from an incomplete picture, and the missing part is the part that matters. Consumption cannot show initiative, judgment, restraint, or the decision not to generate something because it would not have helped. Under a consumption measure, the person who thinks for two hours and then writes forty deliberate lines looks like the weakest performer on the team.

So the twelve-hour builder ends up in the same position as the woman in Chapter 5. High effort, visible compliance, no evidence of the qualities anyone claims to promote for. And in eighteen months a manager will describe them as reliable but not strategic, and will believe it, and will be describing a measurement system rather than a human being.

What a manager changes on Monday

None of this requires a transformation program. It requires a small number of specific changes, and they are available immediately.

Change the first question. Whatever gets asked about first in the next one-on-one is the real scoreboard, and everything said afterward is commentary. If the opening line is about usage, the measure survives no matter how the rest of the conversation goes. Open with what reached someone outside the team, and what is currently stuck waiting on review or on a decision.

Name one outcome for the week, and name it out loud. Not a theme, not a focus area. One thing that will be visibly different by Friday and that someone outside the team would notice. A person given a real target will regulate their own consumption without being asked, because generation becomes a means again rather than the point.

Say the quiet part explicitly. Tell the person that a lower consumption number attached to a shipped result is the better week, and that you will defend that trade if it is questioned upward. This has to be said, not implied. They have months of evidence pointing the other way and one sentence from you pointing at this one.

Protect review capacity as deliberately as generation capacity. The bottleneck in that twelve-hour day was never production. It was the absence of anyone with time to read, decide, and accept. An organization that funds generation without funding judgment has built a machine for producing unreviewed work, and then it blames the person operating it.

And if you are the executive who set the consumption target, retract it in the same forum where you set it. Not in a memo, and not by quietly adding better metrics beside it. The measure has to visibly stop being the measure, publicly, from the person who made it one. Everyone downstream is watching what you count, not what you say, and until the counting changes they will assume the rest is decoration.

The person working twelve hours a day did nothing wrong. They read the instruction and followed it further than anyone else would. Give them a better instruction and you will find out what they can actually do.

Sources

  1. Flynn, Dan. Builders Build: The Four A’s of Organizational Readiness. Mission Intelligence Systems LLC. Chapter 5, “People Want to Build.”
  2. Ridgway, V. F. “Dysfunctional Consequences of Performance Measurements.” Administrative Science Quarterly, vol. 1, no. 2, 1956, pp. 240–247. doi.org/10.2307/2390989.
  3. Campbell, Donald T. “Assessing the Impact of Planned Social Change.” Evaluation and Program Planning, vol. 2, no. 1, 1979, pp. 67–90. doi.org/10.1016/0149-7189(79)90048-X.
  4. Strathern, Marilyn. “‘Improving Ratings’: Audit in the British University System.” European Review, vol. 5, no. 3, 1997, pp. 305–321. Source of the widely quoted formulation that when a measure becomes a target, it ceases to be a good measure.
  5. Deming, W. Edwards. Out of the Crisis. MIT Press, 1986. The argument that performance is produced by the system rather than by the individual, and that management by visible numbers alone leads organizations to blame people for outcomes the system determined.
  6. Edmondson, Amy C. “Psychological Safety and Learning Behavior in Work Teams.” Administrative Science Quarterly, vol. 44, no. 2, 1999, pp. 350–383. doi.org/10.2307/2666999.
DF

About the Author

Dan Flynn

Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build

Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.

His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.