AI Theater Is a Business Risk
Four exposures that accumulate while the adoption chart goes up.
AI theater is usually discussed as waste, which makes it a budget conversation and therefore a low priority one. That framing is wrong, and it is expensive. Activity reported as transformation creates specific exposures with names, owners and early indicators. This article is written for the people who hold that register.
Published
Key Takeaways
- AI theater creates four nameable exposures: cost with no outcome ceiling, debt accumulating faster than it is retired, false confidence traveling upward and then being acted on, and a claim record that cannot survive audit.
- Each exposure is invisible to budget review, because nothing was overspent and nothing was misstated. Each is visible to risk review, because each attaches uncertainty directly to a stated objective.
- The mechanism is the same one that hides every other organizational risk. Problems do not travel upward at the size they started, and the residue is where the exposure accumulates.
Why this belongs in the risk register and not just the budget review
A budget review asks three questions. Was the spending approved, did it stay within the approved envelope, and does it fit the plan for the year. An AI program producing activity without outcome answers all three cleanly. The spending was approved. It stayed inside the envelope, or it grew through a process that was itself approved. It fits the plan, because the plan said adopt AI and the organization adopted AI. Nothing in that review has any mechanism for noticing the problem.
A risk review asks something else. What could prevent the organization from achieving what it committed to achieving, who owns that possibility, what would tell us early, and what would we do about it. ISO 31000:2018 defines risk as the effect of uncertainty on objectives, and it is explicit that risk management should be integrated into governance and decision making rather than run alongside them as a separate exercise. That definition is the whole argument. A program consuming capital and management attention while producing no measurable change in what the organization delivers is uncertainty attached directly to objectives. It belongs on the register by the standard’s own definition, and the fact that it rarely appears there is itself a finding about how the register is built.
What follows are four exposures. They are stated separately because they have different owners, different indicators and different remedies. They compound, which is the reason to name them before they do.
Exposure one: cost with no ceiling
Every disciplined investment has a stopping rule. It may be a hurdle rate, a milestone gate, a pilot threshold or a stated condition under which the organization walks away. The rule exists so that spending is bounded by something other than the enthusiasm of the people spending it.
A program measured by consumption has no such rule available to it. If the evidence of success is usage, then more usage is more success, and there is no level of usage that would ever trigger a review. The metric that justifies continuation is the same metric that increases with continuation. That is not a budget problem, it is a control failure, and it is the kind that a finance function will not catch because every individual expenditure is legitimate and correctly coded.
The literature on cost behavior in large programs is unambiguous about where this leads. Bent Flyvbjerg, writing in the Project Management Journal in 2006, documented that cost underestimation in major projects is systematic rather than occasional, persistent across decades and geographies, and driven by a combination of optimism bias and strategic misrepresentation. His proposed remedy, reference class forecasting, works by forcing an outside view: instead of asking what this project will cost, you ask what projects of this class actually cost. The relevance to AI programs is direct. Internal forecasts for AI investment are being produced by the people most invested in the answer, using an inside view, against a class of projects the organization has never completed before.
The exposure is not that AI spending is large. It is that the spending has no defined outcome threshold below which it would be reduced, and no independent basis for the forecast that justified it. A risk owner should be able to read the stopping rule. In most organizations, there is no document to read.
Exposure two: debt outpacing retirement
The second exposure is quieter and takes longer to surface. AI programs generate artifacts at a rate the organization has never handled before. Code that must be reviewed and maintained. Documents that enter the record. Analyses that create decision points. Tools that require administration, access control and vendor management. Review steps added to catch the failure modes of the new tools. Every one of those is a liability with a carrying cost.
The question that determines whether this is healthy or dangerous is simple and almost never asked. What is the retirement rate. Which systems were decommissioned, which review steps were removed, which documents left the record, which manual process was actually shut down rather than left running in parallel as a fallback. Transformation removes work. When nothing is removed, the organization has not replaced a process, it has added a second one and kept paying for the first.
There is a specific version of this that deserves board attention, and it was described forty years before the current tools existed. Lisanne Bainbridge, in Ironies of Automation in 1983, observed that automating the routine parts of a task leaves the human operator with exactly the residual work the designer could not automate, which is the hardest work, and simultaneously removes the routine practice through which the operator maintained the skill needed to do it. The application is uncomfortable and precise. The people reviewing generated output need the deepest expertise in the organization, and they are the people whose daily practice of that expertise is being reduced fastest. Reviewing at volume degrades over time, and it degrades invisibly, because the reviews are still happening and the sign off is still being recorded.
The exposure is a growing stock of assets nobody has capacity to maintain, guarded by a review function whose capability is quietly eroding. That is a maintainability and operational resilience risk, and it belongs to whoever owns those.
Exposure three: false confidence traveling upward
The third exposure is the one that turns a measurement problem into a capital allocation problem.
Activity metrics are reported upward because they are what exists. They are accurate. Usage did rise, deployment did happen, the percentages are calculated correctly. But an accurate input measure presented on the performance line is read as evidence of capability, and once it is read that way it becomes an input to decisions. Headcount plans are adjusted on the assumption that the capability is present. Delivery commitments are made on the assumption that throughput has changed. Further investment is approved on the strength of a trend line describing consumption. The original report was not false. The decisions built on it are exposed anyway.
Senior executives are not naturally protected against this, and there is direct evidence on the point. Itzhak Ben-David, John Graham and Campbell Harvey, publishing in The Quarterly Journal of Economics in 2013, examined a long series of forecasts made by senior financial executives and found them severely miscalibrated. The confidence intervals those executives supplied were far too narrow, meaning outcomes fell outside their stated ranges far more often than their own stated confidence implied. The finding is about capable, experienced, financially sophisticated people. It is not a comment on competence. It is a comment on how confidence behaves when it is not measured.
Now combine that with a reporting stream that supplies confirming activity data and nothing that could disconfirm it. The organization is not merely uncertain about its AI program. It is confident, it is wrong, and it has no instrument that would tell it which. That is the definition of an unmanaged risk.
Exposure four: an unsubstantiable claim record
The fourth exposure is the one that arrives last and hurts most, because it arrives in the form of a question from someone the organization cannot brush off.
At some point a benefit will be claimed. It will appear in an annual report, an investor call, a regulatory filing, a customer contract or a board minute. Efficiency gains. Productivity improvement. Cost reduction attributed to AI. And at some later point somebody will ask the ordinary follow up question that auditors, regulators, litigators and serious board members all ask. Compared to what, measured how, over what period, and how do you know it was this.
Substantiating that claim requires four things that must have existed beforehand: a baseline captured before the change, a defined measure, a defined period, and a documented basis for attribution. An activity trail supplies none of them. Seat counts, usage growth and deployment milestones describe what the organization did. They cannot evidence what changed because of it. The record is voluminous and it is not evidence, which is the worst combination available, because volume creates the appearance of documentation and invites reliance on it.
The measurement failure underneath this is old. V. F. Ridgway, writing in Administrative Science Quarterly in 1956, set out how quantitative performance measures reliably produce behavior optimized toward the measure rather than the purpose the measure was meant to serve. The organizational record follows the measure. If the measure is consumption, then consumption is what the record preserves, and the record will be assembled around it for years. When the substantiation question finally arrives, the organization does not get to go back and capture a baseline. That window closed at deployment.
This exposure has a name in any mature risk taxonomy. It is a disclosure and assurance risk, and it does not require anyone to have lied.
The travel problem
All four exposures share a single mechanism, and it is not specific to AI.
Chapter 7 of Builders Build is about what does not travel. The principle it sets out is that problems almost never reach leadership at the size they started. Each time an issue passes upward it is compressed, softened, aggregated and given context, usually by people acting reasonably and without any intent to mislead. What arrives at the top is a rounded version of something that was sharp several levels down. The real risk in an organization is rarely in what leadership was told. It accumulates in the part that failed to make the trip.
Every exposure described above lives in that residue. The engineer who knows the generated code is unreviewable says so in a standup, not in a board paper. The analyst who knows the baseline was never captured mentions it once, to a manager, who has no field on the reporting template to put it in. The team lead who knows the manual process is still running in parallel does not raise it, because raising it sounds like resistance to a program the organization has publicly committed to. None of these people are hiding anything. The reporting path simply has no channel for the information, and by the time the summary reaches the risk committee it contains adoption percentages and nothing else.
This is why the exposures compound rather than self correct. A risk that travels gets managed. A risk that does not travel gets larger, and it stays out of sight until it surfaces as something else: a maintenance crisis, a missed commitment, a restated claim, a program that has to be written off at a size nobody expected.
What a risk owner should actually ask for
The remedy is not a better dashboard, and it is not a new committee. It is five artifacts, each of which is a document rather than an opinion, and each of which either exists or does not.
Ask for the baseline. A record of the specific measure the AI program is expected to move, captured before deployment, with a date and a named owner. If it does not exist, that is the first finding, and it should be logged as one rather than treated as an oversight to be corrected quietly. A missing baseline is not a documentation gap. It is a permanent loss of the ability to evaluate the investment.
Ask for the stopping rule. A statement written in advance describing the outcome threshold below which the program is reduced, restructured or ended, and the date at which that threshold will be tested. This is the ceiling that consumption metrics cannot supply. It also converts an open ended commitment into a bounded one, which is the entire purpose of a gate.
Ask for the retirement ledger. A list of what the organization has stopped doing since the program began: systems decommissioned, steps removed, parallel processes shut down. An empty ledger after a year of investment tells you the program has added rather than transformed, and it quantifies the carrying cost that has been taken on.
Ask for recorded confidence. When a forecast is made about what an AI investment will deliver, require that the range be written down at the time it is made, along with who made it. This costs nothing and it makes calibration checkable later, which is the only known defense against confidence that has never been scored.
Ask for one substantiable claim. Not all of them. One. Take a single benefit the organization has asserted and assemble it to the standard an external auditor would apply: baseline, measure, period, attribution. If one claim can be built to that standard, the program is real and the rest can follow the same pattern. If it cannot, the organization has learned something important while the cost of learning it is still small.
None of this is hostile to AI investment, and it should not be presented that way in the room. An organization that can produce these five artifacts is in a stronger position to invest more, because it can distinguish the parts of the program that are working from the parts that are only running. The exposures described here are what accumulate in the absence of that instrumentation. They are ordinary risks, they respond to ordinary risk discipline, and the only unusual thing about them is how rarely anyone writes them down.
Sources
- Flynn, Dan. Builders Build: The Four A’s of Organizational Readiness. Mission Intelligence Systems LLC. Chapter 7, “What Doesn’t Travel,” on why problems rarely reach leadership at the size they started.
- Flyvbjerg, Bent. “From Nobel Prize to Project Management: Getting Risks Right.” Project Management Journal, vol. 37, no. 3, 2006, pp. 5–15. doi.org/10.1177/875697280603700302.
- Ben-David, Itzhak, John R. Graham, and Campbell R. Harvey. “Managerial Miscalibration.” The Quarterly Journal of Economics, vol. 128, no. 4, 2013, pp. 1547–1584. doi.org/10.1093/qje/qjt023.
- Bainbridge, Lisanne. “Ironies of Automation.” Automatica, vol. 19, no. 6, 1983, pp. 775–779. doi.org/10.1016/0005-1098(83)90046-8.
- International Organization for Standardization. Risk Management: Guidelines. ISO 31000:2018. Source of the definition of risk as the effect of uncertainty on objectives, and of the requirement that risk management be integrated into governance and decision making.
- Ridgway, V. F. “Dysfunctional Consequences of Performance Measurements.” Administrative Science Quarterly, vol. 1, no. 2, 1956, pp. 240–247. doi.org/10.2307/2390989.
About the Author
Dan Flynn
Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build
Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.
His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.
