The Customer Cannot Tell You Adopted AI
If nothing outside the organization changed, nothing was transformed.
Most tests of an AI program require data the organization does not have, on a timeline that outruns the reporting cycle. This one requires nothing. It can be run in a meeting, from the material already on the table, by anyone willing to ask about the outcome instead of the activity.
Published
Key Takeaways
- Real transformation surfaces outside the building. Something resolves faster, fails less, costs materially less, or becomes possible that was not possible before.
- Internal evidence is not disqualifying by itself. Internal evidence with nothing behind it is, and it accumulates faster than any other kind because it is the easiest to produce.
- Pair the customer test with the subtraction test: name what the organization stopped doing. Transformation removes work. Theater only adds.
The test, stated plainly
Would a customer be able to tell that you adopted AI?
Not from an announcement. Not from a banner on the product, a note in the release log, or a line in the annual letter. From their own experience of dealing with you. If they never read a word about your AI program, would anything in their contact with the organization be different than it was eighteen months ago?
That is the whole test. Real transformation surfaces outside the building. Something resolves faster than it used to. Something fails less often than it used to. Something costs materially less, and the difference is visible in what is charged or what is delivered. Something is now possible that was not possible before, and the customer can have it.
If every piece of evidence for an AI program is internal, and none of it would be visible to the people the organization exists to serve, the program has produced activity. That conclusion is unpleasant and it is also cheap to reach, which is the point. It costs one question and it does not wait on a measurement system.
One clarification before the rest of this, because the test is easy to overapply. Internal evidence is not disqualifying by itself. Programs have stages. Platform work, training, evaluation harnesses and internal pilots are all real work, and in the early months the evidence for them is legitimately internal. What is disqualifying is internal evidence with nothing behind it. A program a year in, with a full page of internal milestones and no external consequence either delivered or credibly scheduled, has not been slowed by sequencing. It has stopped at the part that was easy to report.
Why internal evidence accumulates so easily
Internal evidence is not chosen because people are trying to obscure anything. It is chosen because it is the only evidence available on the schedule that evidence is demanded.
An external outcome is slow, shared and contestable. Resolution time belongs to operations and to the product and to the customers themselves. Error rates take a quarter to move and another quarter to trust. Cost per unit is entangled with volume, mix and a dozen decisions that had nothing to do with AI. Any external number worth reporting has other people’s fingerprints on it, and someone in the room can always argue it moved for a different reason.
Internal evidence has none of those problems. It resolves this week. It belongs to one team. Nobody disputes it, because it describes something the team did rather than something the world did in response. Seats provisioned, workflows built, models evaluated, policies published, training completed. Each of those is a genuine accomplishment and each of them is fully within the control of the person reporting it.
So internal evidence compounds. Every reporting cycle produces more of it, and the accumulated weight of it starts to feel like proof. This is Kaplan and Norton’s argument from 1992 running in reverse. They built the balanced scorecard because financial measures alone told an incomplete story, and they insisted the customer perspective sit alongside the internal process perspective precisely because organizations that watch only their own processes optimize their own processes. An AI program reported entirely through internal milestones is a scorecard with one quadrant filled in.
Ridgway saw the mechanism earlier still. Writing in 1956, he observed that quantitative measures reliably produce behavior optimized for the measure rather than for the purpose the measure was meant to serve. Once internal milestones become the thing a program is judged by, the program will produce internal milestones. It will produce them efficiently, in volume, and on time. What it will not necessarily produce is anything a customer encounters.
Deming made the same observation about quality thirty years into his career and it holds here without modification. You cannot inspect quality into a product at the end, and you cannot report an outcome into existence at the end. The outcome is either produced by how the work is organized or it is not produced at all.
Three outcomes a customer would notice
Abstraction is what lets this test be passed with a category instead of an instance. It helps to say concretely what an external outcome looks like, so that a vague answer is recognizable as a vague answer.
Something that took days now takes hours
A claim, an approval, a quote, a permit, a discharge summary, a title search. Whatever the organization makes people wait for, they now wait less. The customer does not know why. They know they submitted something on Tuesday and heard back on Tuesday, and that last year the same request took until the following week. This outcome is the easiest of the three to verify because the organization almost certainly already records the timestamps. Nobody needs a new dashboard. Somebody needs to pull the distribution for the same request type across two comparable periods and look at it.
The weak version of this answer sounds like: our analysts can draft the response much faster now. That is an internal claim about one step. It becomes external only when the elapsed time the customer experiences moves, and it often does not, because the saved drafting time is absorbed by a queue further downstream. When the elapsed time does not move, that is information, and it is more useful than the drafting statistic. It tells you the constraint was never drafting.
Something that used to go wrong stopped going wrong
The rework rate falls. The rate of orders that ship incorrectly, applications rejected for a fixable defect, invoices disputed, escalations reopened after being marked resolved. The customer experiences this as an absence, which makes it the least glamorous of the three outcomes and often the most valuable, because errors consume capacity twice: once when the work is done and again when it is undone.
This is the outcome most likely to be genuinely produced by AI and least likely to be reported, because a program measured on adoption has no field for it. Somebody has to go and get the number. It exists, usually in the same system that tracks the complaint or the return or the resubmission.
Something is now offered that was not offered before
The organization does work it previously declined. It serves a segment it previously could not serve economically, reviews every case instead of a sample, answers in a language it did not support, or provides a level of detail it used to reserve for the largest accounts. This is the outcome Brynjolfsson and Hitt described in 2000, when they argued that the returns to information technology came not from the technology itself but from the organizational changes built around it, and that those complementary changes were what separated firms that captured value from firms that only bought equipment. The capability is the input. The reorganization around it is what reaches the customer.
This is the hardest of the three to fake, because it cannot be produced by a pilot. Offering something new requires the organization to commit to delivering it, which means the process, the staffing and the accountability all had to change. That is why its presence is such a strong signal, and why its absence across an entire program is worth taking seriously.
The second question, and the silence after it
The customer test has a companion, and the pair is stronger than either alone. Ask what the organization stopped doing.
Transformation removes work. If AI has genuinely changed how an organization operates, something that used to consume effort no longer does, and someone can name it. A report that is no longer produced. A review step that no longer exists. A weekly meeting that was dissolved because the thing it coordinated is now handled. A vendor contract not renewed. A queue that was eliminated rather than accelerated.
Theater only adds. New tools, new governance reviews, new dashboards, new coordination roles, new training obligations, and the original work still sitting underneath all of it, done the same way it was done before by the same people who are now also attending the new reviews. Every element of that addition is defensible in isolation. Together they describe an organization that has increased its own cost and called the increase progress.
The silence after this question is the diagnosis, and it is worth understanding why the silence happens. It is not that nobody knows the answer. It is that in most organizations, stopping something requires authority that nobody has been given. Adding a tool needs a budget line. Removing a report needs someone willing to be responsible if the report turns out to have been load bearing. Those are different acts, and only one of them is safe. So the additions proceed and the subtractions do not, and a year later the organization is running the old process and the new one at the same time.
That is why the second question is really a question about adaptability rather than about AI. An organization that cannot name one thing it stopped is telling you that it has no mechanism for stopping things. The AI program did not cause that condition. It revealed it, at a cost.
What the organization is broadcasting either way
Neither of these questions extracts information from an unwilling organization. Both of them just point at information that was already being transmitted.
Chapter 6 of Builders Build makes the point that an organization is always telling you what is true about itself, and that it does so through its calendar, its workarounds and its artifacts rather than through its strategy documents. Behavior is the honest channel. The recurring meeting nobody questions shows where attention actually goes. The step everyone routes around shows where the official process failed. The spreadsheet that quietly became the system of record shows what people needed and were not given. None of this is hidden and none of it requires investigation. It requires reading behavior instead of narration.
Apply that to an AI program and the two tests stop being clever and start being obvious. The program is broadcasting continuously. If the calendar filled with AI governance reviews and enablement sessions while the operational meetings ran unchanged, the organization has said that AI is a parallel activity rather than a change to the work. If people built private workarounds to get AI into their actual jobs, the organization has said the sanctioned path did not fit the work, which is worth far more than the adoption number. If the artifacts produced are decks about AI rather than changed processes, the organization has said what the program was for.
The customer test reads the same broadcast from outside. An organization that reorganized its work around a new capability cannot help but produce external evidence, because the work is what customers touch. An organization that layered a capability on top of unchanged work cannot produce it, no matter how sincere the effort or how large the budget. In both cases the outcome is not a verdict on the technology. It is a description of the conditions the technology landed in.
Running both tests in one meeting
Take the most recent AI progress report, whatever form it arrives in. Do not ask for anything new to be prepared, because the preparation is where the answers get smoothed.
Ask the first question of the report as written: which line item here would a customer have noticed, and how would they have noticed it. Require the instance, not the category. Faster claims processing is a category. Auto claims under a certain threshold now close in one day instead of four, here is the distribution, is an instance. If the answer is a category, ask which specific thing changed and where the record of it lives. If nobody can say, that is the answer, and it is not a failure of the person answering.
Then ask the second question: what did we stop doing because of this work. Wait through the pause. The pause is normal and it is the most informative part of the meeting. If the answer eventually arrives and it is real, the program is doing something. If the answer is that nothing stopped but efficiency improved, ask where the freed capacity went, because capacity that cannot be located was probably never freed.
Two concrete answers means the program is producing outcomes and the reporting is simply pointed at the wrong things, which is a straightforward fix. A confident answer to the first and silence on the second means the organization has added capability without removing anything, and the cost line will show it before the value line does. Silence on both means the program has produced activity, and the useful next step is not a better dashboard. It is picking one external outcome, naming the work that would have to stop for it to happen, and giving someone the authority to stop it.
Neither question needs instrumentation. That is the argument for asking them now rather than after the measurement program is built, because the measurement program is itself an internal milestone, and it will report on time.
Sources
- Flynn, Dan. Builders Build: The Four A’s of Organizational Readiness. Mission Intelligence Systems LLC. Chapter 6, “The Organization Is Talking,” on reading the calendar, the workarounds and the artifacts rather than the strategy document.
- Ridgway, V. F. “Dysfunctional Consequences of Performance Measurements.” Administrative Science Quarterly, vol. 1, no. 2, 1956, pp. 240–247. doi.org/10.2307/2390989.
- Kaplan, Robert S., and David P. Norton. “The Balanced Scorecard: Measures That Drive Performance.” Harvard Business Review, vol. 70, no. 1, 1992, pp. 71–79. Source of the argument that the customer perspective must sit alongside internal process measures.
- Brynjolfsson, Erik, and Lorin M. Hitt. “Beyond Computation: Information Technology, Organizational Transformation and Business Performance.” Journal of Economic Perspectives, vol. 14, no. 4, 2000, pp. 23–48. doi.org/10.1257/jep.14.4.23.
- Deming, W. Edwards. Out of the Crisis. MIT Press, 1986. On the impossibility of inspecting quality into a product after the fact.
About the Author
Dan Flynn
Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build
Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.
His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.
