AI Theater
The activity resembles transformation closely enough to be reported as transformation.
There is a kind of organization that is hard to see clearly, because from a distance it looks like a success story and up close it behaves like a failure. The tools are everywhere. Usage is climbing. The adoption dashboard would impress any board. And nothing the organization produces has actually changed.
Published
Key Takeaways
- AI theater is activity that resembles transformation closely enough to be reported as transformation. The consumption numbers say transformation while the outcomes say expense.
- It survives because every number in it is real. Usage is genuinely up, the charts genuinely move, and nobody is asking what changed because of any of it. Nothing has to be falsified for the reporting to be wrong.
- The exit is not a better dashboard. It is a different question, asked of the outcome rather than the activity: what would a customer notice, and what did the organization stop doing?
I named this pattern in Builders Build while writing about the thing I kept finding in organization after organization. It deserves its own treatment, because it is spreading faster than anything else I encounter.
Priority theater is what happens when the document says one thing and the calendar says another. AI theater is its descendant. The consumption numbers say transformation while the outcomes say expense. The activity resembles transformation closely enough to be reported as transformation, and in most organizations it will be, because the numbers are real, the charts move in the right direction, and nobody is asking what changed because of any of it. Adoption is not value. Consumption is not even adoption. It is what the effort cost, presented as what the effort produced.
From Builders Build, Chapter 9
That last sentence is the whole diagnosis. It is what the effort cost, presented as what the effort produced. Every other symptom follows from it.
Nothing has to be falsified
The reason AI theater is durable is that it requires no dishonesty. Compare it to the failures executives are trained to look for. Fraud requires someone to misstate a number. Overpromising requires someone to make a claim that later proves false. Both leave evidence, and both have a person attached to them.
AI theater leaves no such trail. The usage figures are accurate. The adoption percentages are correctly calculated. The month-over-month growth is real growth. Every individual element of the report survives scrutiny, because every individual element is true. What fails is the inference: that because consumption rose, capability rose with it.
No one has to argue for that inference. It arrives unexamined, in the gap between a chart that is going up and a question nobody asked.
The oldest measurement error, in a new unit
This is not a new failure. It is the most durable failure in the history of organizational measurement, and it has been documented for seventy years.
In 1956, V. F. Ridgway published a short paper in Administrative Science Quarterly with a title that has aged remarkably well: “Dysfunctional Consequences of Performance Measurements.” His argument was that quantitative measures, however well intended, reliably produce behavior optimized for the measure rather than for the purpose the measure was meant to serve. Donald Campbell reached the same conclusion from a different direction in 1979. Marilyn Strathern gave it the phrasing most people know, writing that when a measure becomes a target, it ceases to be a good measure.
Every organization I work with knows this. Most of them can name the version of it they already survived. They stopped counting lines of code because it produced more code and worse software. They stopped counting hours at a desk because it produced longer days and no more output. They learned, expensively, that an input measure describes consumption rather than contribution.
Then a new input arrived with a new unit attached, and the lesson did not travel.
Token consumption is hours at a desk. It is lines of code. It is the same category of number wearing unfamiliar clothing, and it is unfamiliar enough that the pattern recognition does not fire. An executive who would immediately reject a proposal to rank engineers by keystrokes will accept a dashboard ranking teams by tokens, because the second one sounds like it is about AI rather than about counting.
What high consumption can actually mean
A team consuming enormous quantities of AI capacity might be transforming the business. It might also be doing something considerably worse than nothing, and the consumption figure cannot distinguish between the two.
Consider what twelve hours of maximum consumption is capable of producing in an organization whose conditions are weak. Code that no one has the capacity to review, and that will therefore be maintained by whoever is unlucky enough to inherit it. Analyses that generate decision points faster than the organization can assign anyone to decide them. Documents that exist to demonstrate effort rather than to be read. Options where there was previously a plan.
All of that registers as high adoption. Some of it registers as high productivity. None of it reaches a customer, and a meaningful share of it will be paid for twice, once when it was generated and again when someone has to undo it.
The consumption number is not wrong about what happened. It is silent about whether what happened was worth doing, and silence is easily mistaken for endorsement when the chart is going up.
Why capable executives keep choosing it
It would be convenient if this were a failure of intelligence or of care. It is neither, and treating it that way guarantees you will not fix it.
Consumption metrics get chosen because they are available. An executive asked to demonstrate return on a substantial AI investment, on a reporting cycle that arrives long before any outcome could plausibly mature, has a narrow set of options. Outcome measures are lagging, shared across functions, and contestable. Consumption measures exist today, resolve to a single number, and belong unambiguously to one team.
Faced with a board meeting in three weeks, most people reach for the number that exists. That is not weakness. It is a rational response to a reporting structure that demands evidence before evidence is available.
Which means the fix is not to inform executives that consumption is a poor measure. Most of them already suspect it. The fix is to change what the organization asks for, and when.
The paradox has a precedent
There is a longer pattern here worth holding onto, because it argues for patience rather than panic.
Robert Solow observed in 1987 that the computer age was visible everywhere except in the productivity statistics. The gap was real, and it persisted for years. It closed not when better computers arrived but when organizations reorganized the work around them. Erik Brynjolfsson, Daniel Rock and Chad Syverson revisited exactly this question for AI in 2017, arguing that the lag between a general purpose technology’s arrival and its measured productivity effect is normal, and is caused by the time required to build the complementary organizational assets that let the technology pay off.
That is the argument of this entire body of work, stated by economists. The technology is not the constraint. The conditions around it are. The organizations that got value from computers were not the ones that bought the most computers.
AI theater is what the intervening years look like when an organization measures the purchase instead of the reorganization.
Two questions that end the performance
Diagnosis without a test is just commentary. There are two questions that reliably separate theater from transformation, and neither requires new instrumentation.
What would a customer notice? Real transformation surfaces outside the organization. Something resolves faster, fails less often, costs materially less, or becomes possible that was not possible before. If every piece of evidence for an AI program is internal, and none of it would be visible to the people the organization exists to serve, the program has produced activity. Internal evidence is not disqualifying on its own. Internal evidence with nothing behind it is.
What did we stop doing? Transformation removes work. If AI has genuinely changed how an organization operates, something that used to consume effort no longer does, and someone can name it. Theater only ever adds: new tools, new reviews, new dashboards, new coordination, and the original work still sitting underneath all of it. An organization that cannot name a single thing it stopped has not transformed anything. It has bought an addition.
Ask both in the same meeting. The silence after the second one is usually more informative than any answer to the first.
What this is not
This is not an argument that AI is overhyped, and it should not be read as one.
What these systems can already do is remarkable, and the trajectory is genuinely difficult to overstate. I am not describing a technology that failed to deliver. I am describing organizations that deployed something powerful into conditions that could not convert it, and then measured the deployment instead of the conversion.
AI does not create organizational problems. It inherits them. Scattered attention produces AI output that pulls in a dozen directions. Weak alignment produces results that are locally optimized and organizationally incoherent. Concentrated authority turns the technology into one more instrument of centralization. Brittle adaptability leaves an organization unable to learn at the speed its environment now demands.
None of that is the technology’s doing. All of it was true before the tools arrived. What changed is that the tools made it visible, expensive, and fast.
The organizations getting compounding value from AI are, without exception in my experience, the organizations that had done the underlying work first, consciously or not. The Four A’s are not a response to AI. They are the organizational readiness that AI requires.
Sources
- Flynn, Dan. Builders Build: The Four A’s of Organizational Readiness. Mission Intelligence Systems LLC. Chapter 9, “The Pattern I Couldn’t Ignore,” where the term AI theater is defined.
- Ridgway, V. F. “Dysfunctional Consequences of Performance Measurements.” Administrative Science Quarterly, vol. 1, no. 2, 1956, pp. 240–247. doi.org/10.2307/2390989.
- Campbell, Donald T. “Assessing the Impact of Planned Social Change.” Evaluation and Program Planning, vol. 2, no. 1, 1979, pp. 67–90. doi.org/10.1016/0149-7189(79)90048-X.
- Strathern, Marilyn. “‘Improving Ratings’: Audit in the British University System.” European Review, vol. 5, no. 3, 1997, pp. 305–321. Source of the widely quoted formulation that when a measure becomes a target, it ceases to be a good measure. doi.org/10.1002/(SICI)1234-981X.
- Solow, Robert M. “We’d Better Watch Out.” New York Times Book Review, 12 July 1987, p. 36. The origin of the productivity paradox observation.
- Brynjolfsson, Erik, Daniel Rock, and Chad Syverson. “Artificial Intelligence and the Modern Productivity Paradox: A Clash of Expectations and Statistics.” NBER Working Paper 24001, National Bureau of Economic Research, November 2017. doi.org/10.3386/w24001.
About the Author
Dan Flynn
Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build
Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.
His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.
