AI Transformation · Authority
A Time Saving Is Not a Return
The arithmetic between a faster task and a number that moves.
The pilot worked. The task really is faster, and the trials that say so are good ones. The profit and loss statement still did not move, and that is not a contradiction. It is arithmetic nobody did.
slower with AI, in a trial where the same developers estimated they had been 20 percent faster
METR, Early-2025 AI and Experienced Developer Productivity
A reported saving and a measured saving are different numbers, and they can differ in sign.
Diagnose your constraint →Research Foundation
- Randomized and field-experimental AI productivity studies
- Cost behavior and management accounting research
- Internal capital and resource reallocation research
- Productivity economics of general purpose technologies
- Enterprise transformation field experience
- The Builders Build Framework
Key Takeaways
- The task-level gains are real and replicated. Randomized trials report writing tasks completed 37 percent faster, support issues resolved 14 percent faster, and consulting tasks finished 25 percent faster at higher quality. Arguing with the trials is the wrong move.
- A saving converts to money through exactly three exits. Capacity leaves the cost base, capacity earns revenue, or a planned hire is canceled. Absent one of those decisions the hours are absorbed, and cost behavior research shows absorption is the default because costs fall more slowly than they rise.
- The two numbers most business cases omit are the honest baseline and the cost of the build itself. Both are Authority problems, because converting a saving requires somebody empowered to change a cost base or move a person, and that decision is made outside the AI program.
The task-level gains are real, and that is what makes this hard
It would be easier if the technology did not work. It does, and the evidence is better than most management claims ever get. Noy and Zhang randomized 444 college-educated professionals across occupation-specific writing tasks and found completion times fell sharply while graded quality rose. Brynjolfsson, Li and Raymond studied the staggered rollout of a generative assistant across 5,172 customer support agents and measured a 14 percent average increase in issues resolved per hour, concentrated among novices at roughly 34 percent. Dell’Acqua and colleagues ran a pre-registered field experiment with 758 Boston Consulting Group consultants and found that on tasks inside the model’s competence, consultants finished 12.2 percent more tasks, 25.1 percent faster, at meaningfully higher quality.
Three independent designs, three consistent findings. A leader who responds to a disappointing return by doubting whether the tool saves time is arguing against good evidence and will lose. The interesting question sits one step later.
The same trials say the gain is conditional
Read those papers past the headline and each one carries a boundary. Dell’Acqua and colleagues named it the jagged technological frontier. On one task deliberately chosen to sit outside the model’s competence, consultants using AI were 19 percentage points less likely to reach a correct answer than consultants using nothing at all. The support study found the large gains among newer agents and minimal effect for experienced, high-skill ones. The gain is not a property of the tool. It is a property of the pairing between a tool, a task, and a person.
This matters for a business case because a pilot is almost never a random sample of the work. Pilots run on the tasks somebody expected to go well, with the people who volunteered. Extrapolating that result across a function assumes the frontier is flat, and the research says it is jagged.
A reported saving is not a measured saving
In 2025, METR ran a randomized controlled trial with 16 experienced open-source developers across 246 real tasks in mature repositories they already knew well. Before starting, the developers forecast that AI assistance would cut completion time by 24 percent. Afterward, having done the work, they estimated it had cut completion time by 20 percent. The clock said it had increased completion time by 19 percent.
The size of that error is the finding worth carrying into a boardroom. Belief and measurement did not merely diverge, they pointed in opposite directions, and the people holding the belief were experienced practitioners reporting on their own recent work. Most enterprise AI business cases are built on exactly this kind of self-report, gathered through a survey asking people how much time the tool saves them. That instrument has now been calibrated against a clock, and it failed.
If the baseline did not exist before the pilot started, the pilot did not measure a saving. It collected an impression.
A saving has exactly three exits
Grant the best case. The frontier was flat, the baseline was real, and a team of forty genuinely recovered four hours each per week. That is 160 hours a week, and it is not money yet. It becomes money through three exits and no others. Capacity leaves the cost base, through reduced contract labor, attrition that is not backfilled, or a smaller team. Capacity is redeployed onto work that produces revenue or avoids a loss, and somebody can name that work. Or a hire that was in the plan is canceled, which is the cleanest of the three because the counterfactual is written down.
Anything else is absorption. The hours return to the working day and raise the quality of what was already being done, lengthen the tasks that expand to fill available time, or disappear into the coordination load the new tool created. Absorption is not a scandal and it is sometimes the right choice, because a team running at the edge may genuinely need slack. It is simply not a return, and a business case that promised one has not been met.
The cost base does not fall on its own
The assumption hiding inside most AI business cases is that freed capacity converts itself, that once the work takes less effort the cost of the work follows it down. Management accounting settled this question two decades ago and the answer is no. Anderson, Banker and Janakiraman examined 7,629 firms over twenty years and found selling, general and administrative costs rise about 0.55 percent for each 1 percent increase in sales, while falling only about 0.35 percent for each 1 percent decrease. Costs are sticky. They go up more readily than they come down, because reducing committed resources takes a deliberate decision by someone willing to make it.
That asymmetry is the whole leak, and it is not a failure of will. It is the ordinary behavior of organizations, documented across decades and industries, and it predicts precisely what AI programs report. Activity falls. Cost does not. The difference shows up as slack rather than savings, and the finance function books no improvement because none arrived.
The build has a cost that nobody books
There is a second number missing, and it sits on the other side of the ledger. The AI initiative was staffed by people, and those people came off something. Rarely is the AI program funded by hiring; it is funded from a roster. The engineers, analysts and product managers assigned to it were previously assigned elsewhere, and whatever they were doing slowed, stopped, or shipped later. That is a real cost of the initiative and it is almost never charged against the initiative’s return.
Strategy research says organizations are poor at this in a specific and measurable way. Bardolet, Fox and Lovallo found that corporate capital allocation follows a cognitive tendency toward spreading resources evenly across units rather than concentrating them where value is. Lovallo, Brown, Teece and Bardolet then examined several thousand firms over eighteen years and found that the willingness to reallocate resources between business units correlates positively with firm performance, with inertia sustained by sunk cost reasoning, status quo bias and internal politics. Reallocation is where returns come from and it is the thing firms are worst at, which means the displaced work is both expensive and invisible.
The complement is the actual investment
None of this is new to the economics of technology. Brynjolfsson and Hitt showed that the value of information technology investment depends far more on complementary organizational investment, in redesigned processes, new skills and changed decision rights, than on the technology spend itself, and that those complements are intangible and hard to measure. Brynjolfsson, Rock and Syverson formalized the consequence as the productivity J-curve. While an organization is building the intangible complements a general purpose technology requires, measured productivity understates real progress, and it overstates progress later when those complements start paying off.
The J-curve is genuinely good news, and it is also the most easily abused idea in this article. It explains why an honest eighteen-month number can look poor while real capability is accumulating. It becomes an excuse the moment it is used without a named complement and a date. The test is simple. If a leader can say which process was redesigned, which decision right moved, and when the curve is expected to turn, the J-curve is a forecast. If not, it is a way of not answering.
What the aggregate implies about your case
There is a useful check available at the macro level. Acemoglu builds a task-based model of AI’s effects and estimates total factor productivity gains on the order of 0.7 percent over ten years, an order of magnitude below the more enthusiastic projections. Reasonable economists disagree with that estimate, and it is a modeling result rather than a measurement. Its use here is narrower than a prediction.
If credible estimates of the economy-wide effect are modest, then an organization forecasting a transformative internal return is claiming to be an outlier. That may well be true, and the 5 percent of firms capturing most of the value are somewhere. But an outlier claim requires outlier evidence, naming the specific process, the specific exit, and the specific number. It cannot rest on the general proposition that AI is important.
The Four A’s reading
The research establishes the phenomenon. The Four A’s of Organizational Readiness give an executive a way to see which condition is missing before the next approval, rather than after the write-off.
Attention decides whether a baseline exists at all, because the number a saving is measured against has to be tracked before the pilot rather than reconstructed afterward. Alignment decides whether the organization agreed which of the three exits it was taking, since a program aiming at cost reduction and a program aiming at redeployment are different programs that look identical in a slide. Authority decides whether anyone can actually take the exit, because removing a cost or moving a person is a decision made outside the AI program by somebody who owns the roster or the budget. Adaptability decides whether the change survives the quarter in which attention moves elsewhere. This article is filed under Authority because that is the condition the evidence keeps pointing at. Cost stickiness and reallocation inertia are both, at bottom, descriptions of decisions nobody was empowered or willing to make.
Evidence matrix
Each claim below separates what the research establishes, what field experience observes, and what the Four A’s interpret. The framework is a synthesis built on established science, not a replacement for it.
| Claim | Research | Field evidence | Four A’s |
|---|---|---|---|
| Task-level AI speedups are real and replicated | Noy & Zhang (2023); Brynjolfsson, Li & Raymond (2025); Dell’Acqua et al. (2025) | Successful function-level pilots | Alignment |
| The gain is conditional on task and experience | Dell’Acqua et al. jagged frontier; support-agent skill heterogeneity | Pilot results that do not survive rollout | Attention |
| Self-reported savings can be wrong in sign | METR (2025) | Survey-based benefit cases | Attention |
| Freed capacity does not leave the cost base by itself | Anderson, Banker & Janakiraman (2003) | Efficiency programs with unchanged run rate | Authority |
| The build’s opportunity cost goes unbooked | Bardolet, Fox & Lovallo (2011); Lovallo et al. (2020) | Roadmaps quietly slipped to staff an AI build | Authority |
| Returns lag while intangible complements are built | Brynjolfsson & Hitt (2000); Brynjolfsson, Rock & Syverson (2021) | Multi-year modernization programs | Adaptability |
What to settle before the next approval
Four decisions, taken before the money is committed rather than after the pilot reports. Name the number that will move and confirm it is already tracked, because a baseline reconstructed after the fact is an argument rather than a measurement. Name which of the three exits this case takes, and say it out loud, since cost reduction and redeployment are different promises made to different people. Name who has the authority to take that exit, and check it is not the program sponsor, because it rarely is. And book the build against the work it displaced, so the return is compared with the roadmap that was given up to earn it.
None of this argues for less AI. It argues that the technology has stopped being the hard part. The trials have answered whether these tools save time. Whether that time becomes money is a different question, decided by people who mostly do not attend the AI steering committee, and it is answered in the organization rather than in the model.
