Mission Intelligence Systems
Risk · Alignment

“Likely” Is Not a Probability

Why estimative words like likely fail as real probabilities.

When two members of a risk review assign different numerical meanings to the same word, the risk register is not a shared instrument. It is a collection of individual opinions formatted to look like one.

Published Last reviewed

A comparison of the word likely as a feeling read as many different percentages versus probability as a single agreed number, illustrating why a shared word is not a shared probability.

Key Takeaways

Sit in enough risk reviews and a pattern becomes unmistakable. Someone rates a risk as “likely.” Someone else rates a different risk as “probable.” The register captures both. Leadership reviews the heat map and draws conclusions about relative risk exposure. And nobody in the room has checked whether “likely” and “probable” mean the same thing to the two people who used them, or whether either term has a shared numerical meaning at all.

They almost never do. This is not a failure of rigor. It is a predictable consequence of how verbal probability language works in practice, and it is one of the most consequential unexamined assumptions in organizational risk management.

A probability is a number. It states, on a defined scale, how likely an event is to occur. A verbal probability term is a word standing in for that number, and the substitution only holds if every person using the word converts it back to the same range. When they do not, the word carries the appearance of a measurement without the content of one. That is the entire problem, and it is why “likely” is not a probability.

How much do people disagree about what 'likely' means?

The problem has been documented in decision science for decades. In a landmark study by Sherman Kent, the CIA analyst who founded the analytic standards tradition, analysts were asked to assign numerical probabilities to a set of verbal probability terms: words like “probable,” “likely,” “possible,” “unlikely,” and “remote.” The results were striking: for the same term, respondents assigned probabilities spanning ranges of 40 percentage points or more. One analyst's “probable” was another's “possible.”

This finding has been replicated many times across different professional communities: military analysts, medical professionals, financial risk teams, program managers. The specific numbers vary by context and culture, but the structural finding is consistent: verbal probability terms are interpreted differently by different people, and the spread in interpretation is large enough to change decisions if it were made visible.

The implication for organizational risk management is direct. When a risk register uses verbal probability terms without numerical anchors, it is not capturing shared risk assessments. It is capturing individual ones: formatted to look shared because they use the same vocabulary. The heat map that emerges from this process appears to rank risks. What it actually does is aggregate the intuitions of whoever rated them, expressed in terms that are legible but not comparable.

Shared vocabulary is not shared understanding. The word “probable” appears in the same column on every row, but the number behind it varies by rater, by mood, by recent experience, and by organizational culture.

Why does vague probability language break alignment?

In the Four A's of Organizational Readiness™ framework, Alignment is the condition that determines whether leaders share not just the same words but the same operating assumptions. I have written about how this gap appears in strategy: leaders leave an offsite having agreed on a direction, only to discover three months later that they were each implementing a different interpretation of what the direction meant. Risk probability language creates exactly the same failure mode, but with numbers.

Consider what happens in a risk review when probability language is not aligned. The program manager rates a vendor dependency risk as “likely”: meaning, in her experience, somewhere around 65%. The finance lead rates a budget risk as “probable”: meaning, to him, something more like 45%. The heat map places both risks in the same probability band. Leadership allocates contingency and attention based on what the heat map shows.

What leadership is actually looking at is not a ranked list of organizational risks. It is a ranked list of individual intuitions, translated into a common vocabulary that obscures the variation in what that vocabulary means. The decisions that follow from this, about where to focus attention, where to build contingency, what to escalate, are based on a false premise of shared assessment.

This is alignment debt in its purest form: the accumulated cost of decisions made without shared understanding. In risk management, that debt pays out in surprises: risks that materialize at rates no one expected because the probability estimates that should have flagged them were not actually comparable across the people who made them.

What does a calibrated probability scale require?

The solution is not to abandon verbal probability terms. They are useful for communication and easier to work with than raw numbers in most organizational contexts. The solution is to anchor them: to define, explicitly and numerically, what each verbal term means in your risk framework, and then to train everyone who rates risks to apply those definitions consistently.

A well-designed probability scale might look like this: Remote (1–10%), Unlikely (11–30%), Possible (31–50%), Likely (51–70%), Probable (71–85%), Near-Certain (86–99%). These are not universal: different organizations and different domains require different calibrations. But whatever the specific ranges, the critical requirement is that they be defined in advance, documented in the risk framework, and applied consistently by everyone who uses them.

Definition alone is insufficient. Calibration, the actual alignment of individual probability intuitions to the defined scale, requires practice. The method used in intelligence analysis and in high-stakes decision-making contexts is calibration training: raters are presented with reference cases where the actual outcome is known, asked to assign probabilities, and then shown how their assignments compare to the actual rates. Over time, raters develop the habit of anchoring their intuitions to evidence rather than to unexamined confidence.

Reference-class forecasting provides a second anchor. When rating the probability of a risk, the most reliable starting point is not introspection about this particular situation: it is the historical base rate for similar situations. What percentage of programs with this vendor profile have experienced delivery delays? What is the historical rate of regulatory change in this domain over a two-year implementation horizon? These base rates are not always available, but when they are, they provide the empirical grounding that prevents probability estimates from drifting toward optimism or anchoring to the most recent anecdote.

What conditions let calibration survive contact with the organization?

Calibrated probability assessment is not just a technical practice. It is an organizational one, and it depends on conditions that many organizations have not built.

The first condition is psychological safety. Calibration requires raters to say “I don't know” and “I might be wrong” in a professional context where confidence is often rewarded and uncertainty is sometimes read as weakness. In organizational cultures where expressing uncertainty about a project outcome is treated as a signal of insufficient commitment, raters will systematically understate probability on negative risks and overstate confidence in their ability to manage them. No calibration methodology can fix this. It is a conditions problem, not a training problem.

The second condition is learning infrastructure. Calibration improves over time only if the organization tracks what it predicted and what actually happened. This requires closing the loop on risk outcomes: recording not just that a risk materialized but the probability that was assigned to it at the time, so that the organization can see whether its risk assessments were systematically optimistic, pessimistic, or well-calibrated. Most organizations do not do this. They record risks while they are live and close them when they resolve, without capturing what their initial estimate implied about expected frequency.

The third condition is shared standards that are actually enforced. A probability scale that lives in the risk framework document but is never checked in practice is not a standard: it is a decoration. Making calibration real requires someone in the risk review to ask the question: when you said “likely,” what probability range did you have in mind? Is that consistent with our scale? That question is uncomfortable the first time it is asked. It becomes routine when the organizational culture treats it as a sign of rigor rather than a challenge to the rater's judgment.

You do not have a shared risk picture until you have shared probability language. And you do not have shared probability language until you have defined what the words mean and built the conditions for people to apply them honestly.

What changes once the probability language is shared?

Organizations that invest in probability language alignment get a specific and measurable return: their risk registers become comparable. A risk rated “likely” by one team means the same thing as a risk rated “likely” by another. Portfolio-level risk aggregation becomes possible. Leadership can look at a risk report that spans programs and functions and make resource allocation decisions with some confidence that the ratings it is looking at reflect a common standard, not a collection of individual intuitions.

This comparability is the foundation of organizational risk judgment: the capacity to see the aggregate risk picture across the enterprise, not just the individual risk picture of each program in isolation. Without calibrated probability language, enterprise risk management is a category error: you cannot aggregate what was never measured on a common scale.

The investment required to build this is not large. It is primarily a matter of organizational will: the decision to define the standard, train the people who apply it, and enforce it consistently in risk reviews. What makes it rare is not technical complexity but the same thing that makes most organizational conditions rare: it requires deliberate construction, sustained attention, and leadership behavior that reinforces the standard when it would be easier to accept a vague assessment and move on.

The risk register that uses the word “likely” without a defined numerical anchor is not providing risk information. It is providing the appearance of risk information: formatted as data, interpreted as judgment, and acting as a basis for decisions that deserve better.

When is a calibrated word still not precise enough?

Calibrating the vocabulary is the floor, not the ceiling. Agreeing that “likely” means 60 to 79 percent makes two registers comparable, which is a real gain. It does not tell a board how much money to hold against the risks in those registers, and it cannot, because a rating describes one risk at a time while the decision the board has to make is about all of them at once.

That decision requires a distribution rather than a rating. Combine the individual uncertainties and you get a range of possible outcomes with a confidence level attached to each point on it. This is where percentile language comes from: a P50 estimate is the value you have a 50 percent chance of coming in under, a P80 is the value you have an 80 percent chance of coming in under, and the distance between your target and your P80 is the contingency the arithmetic says you need. The word “likely” cannot produce that number. A calibrated probability, combined across a program, can.

This is not an academic preference. It is what oversight bodies increasingly require in writing. The Federal Transit Administration has held capital project cost contingency to the 65th percentile since 2018 and asks for a risk-informed assessment of cost and schedule ranges. The GAO cost and schedule assessment guides treat uncertainty analysis as a practice that programs are audited against. AACE International maintains recommended practices for determining contingency from a simulated range rather than from a percentage applied by habit. In each case the requirement is a number carrying a stated confidence level, and an organization that has never defined what its own probability words mean cannot produce one honestly.

So the two problems are the same problem at different resolutions. An organization that cannot agree what “likely” means will not produce a defensible P80, because the inputs to the model are the same judgments that populate the register. Calibrating the language is the prerequisite. Quantifying the aggregate is what the language was for.

References

  1. Kent, Sherman. “Words of Estimative Probability.” Studies in Intelligence, vol. 8, no. 4, 1964, pp. 49–65. Studies in Intelligence archive (PDF no longer hosted on CIA.gov). The foundational study demonstrating that verbal probability terms carry widely divergent numerical interpretations even among trained professional analysts.
  2. Beyth-Marom, Ruth. “How Probable Is Probable? A Numerical Translation of Verbal Probability Expressions.” Journal of Forecasting, vol. 1, no. 3, 1982, pp. 257–269. doi.org/10.1002/for.3980010305. Documented 40-plus percentage point spreads in numerical interpretation of the same verbal probability term across respondents.
  3. Tetlock, Philip E., and Dan Gardner. Superforecasting: The Art and Science of Prediction. Crown Publishers, 2015. Chapters 3–4 address calibration training, reference-class forecasting, and the practices that distinguish accurate probability estimators from intuitive ones.
  4. Kahneman, Daniel. Thinking, Fast and Slow. Farrar, Straus and Giroux, 2011. Chapter 11 (Anchors) and Chapter 24 (The Engine of Capitalism) address how intuitive probability estimates are anchored to vivid recent experiences rather than base rates.
  5. International Organization for Standardization. ISO 31000:2018 Risk Management - Guidelines. ISO, 2018. iso.org/standard/65694. Provides the international standard framework for risk assessment, including the requirement for consistent probability criteria across an organization.
  6. Friedman, Jeffrey A. and Richard Zeckhauser. “Assessing Uncertainty in Intelligence.” Intelligence and National Security, vol. 27, no. 6, 2012, pp. 824–847. doi.org/10.1080/02684527.2012.708275. Extends the Kent research to demonstrate the persistent calibration problem in professional analytical contexts.
DF

About the Author

Dan Flynn

Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build

Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.

His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.

Related Articles

Diagnose Your Risk Alignment

The Risk Capability Diagnostic measures sixteen capabilities, among them estimating and uncertainty, and quantitative analysis and aggregation, that determine whether risk management produces real protection or risk theater.