Mission Intelligence Systems
Risk · Adaptability

Can You Put a Number on a Risk That Never Happened?

The objection that ends most quantification conversations, and why it does not hold.

Someone proposes quantifying the exposure and someone else says there is no data. The conversation ends there, the risk stays qualitative, and the organization's single largest exposure remains a colored square on a grid for another five years. The objection sounds rigorous. It rests on a confusion between having a frequency history and having a defensible statement of uncertainty.

Published

Key Takeaways

Research foundation

Structured elicitation rests on Cooke (1991) and the applications recorded in Cooke and Goossens (2008) in Reliability Engineering & System Safety. The overconfidence finding is Lichtenstein, Fischhoff and Phillips (1982); the trainability finding is Mellers and colleagues (2014) in Psychological Science. Decomposition is illustrated with The Open Group's FAIR taxonomy, and reference class with Flyvbjerg (2006). Technique selection is anchored in ISO/IEC 31010:2019 and GAO-20-195G. The Four A's are the executive lens applied on top.

Start by separating two claims that get said in the same breath. We have never experienced this is a fact about the organization. Therefore we cannot say anything quantitative about it is a much stronger claim, and it is false.

Insurers price events that have not happened to a particular policyholder. Engineers specify against loads a structure has never seen. Both do so with numbers, and neither waits for the event.

Three routes to a number

Decompose until you reach something you know

Most unmeasurable risks are unmeasurable only at the level they are stated. The statement a major cyber incident has no natural frequency. Its components do.

This is the logic of the FAIR taxonomy, published through The Open Group, which decomposes risk into loss event frequency and loss magnitude, and decomposes each of those in turn: contact frequency, probability of action, threat capability against control strength, and on the magnitude side the separable categories of loss the organization would actually incur. At the bottom of that tree sit quantities the organization does have information about, including how many records it holds, what its notification obligations cost per record, how long recovery took the last time a system was down for any reason, and what its own logs show about attempt volume.

The same move works outside cyber. The systems integration fails resists estimation. How many interfaces there are, how many have been tested, what proportion of interfaces historically require rework, and what rework has cost per interface do not.

Decomposition does not manufacture information. It relocates the estimate to a level where the organization already has some, which is usually two or three levels below where the risk was written.

Widen the class until it has members

You have no history of your own facility experiencing a total loss. There is extensive history of facilities like yours experiencing total losses.

This is reference class reasoning applied to a discrete event rather than to a cost forecast, and it carries the same discipline: the class must be defined before the estimate is wanted, on characteristics that plausibly drive the outcome, and it must be wide enough to contain enough members to say something. The temptation, always, is to narrow the class until only favorable cases remain, which is how an outside view gets quietly converted back into an inside view.

Elicit under a protocol, not in a workshop

The third route is expert judgment, and it deserves care because it is the one most often done badly.

The core finding in the calibration literature, consolidated by Lichtenstein, Fischhoff and Phillips and replicated extensively since, is that people asked for a range give one that is too narrow. Ask for a ninety percent interval and the true value falls outside it far more than ten percent of the time. This is not a failure of intelligence or of subject knowledge, and experts are not exempt.

Cooke's response, the classical model, is to stop treating expert opinion as a single undifferentiated input. Experts are asked seed questions to which the answers are known but which they have not memorized, their responses are scored for calibration and informativeness, and their judgments on the questions of interest are weighted accordingly. The TU Delft database that Cooke and Goossens assembled records a substantial body of applications of this method across domains, which is what makes it something other than a proposal.

Structured elicitation does not assume experts are well calibrated. It measures whether they are, and weights them by the answer. That is the entire difference between elicitation and a show of hands.

Calibration also improves with training. In a large geopolitical forecasting tournament, Mellers and colleagues found that a short training intervention in probabilistic reasoning produced measurably better forecasts. An organization that runs a calibration exercise with its estimators once a year is doing something with published support behind it, and it takes an afternoon.

What this does not solve

Three limits, stated plainly, because the case for quantification is weakened rather than strengthened by overselling it.

Tail extrapolation assumes stability. Statistical methods for extremes work by fitting a tail shape to the observed portion of a distribution and extending it. That extension is only valid if the mechanism producing extreme outcomes is the same one that produced the observed range. For a genuinely novel event, or one arising from a system that has changed structurally, the assumption is unsupported and the resulting tail estimate is a modelling artifact.

Elicitation inherits its own quality. A distribution obtained from three people in a hurry, anchored on the first number spoken, is not evidence merely because it emerged from a process labelled elicitation. The protocol is what does the work, and a protocol that is skipped produces a number with more authority than the workshop it came from.

Precision is not accuracy. A P80 stated to five significant figures is a presentational choice, and it invites more confidence than the underlying inputs support. Where an estimate is built primarily on elicited judgment, that fact belongs next to the number rather than in an appendix, and the range should be presented at a resolution that reflects what is actually known.

When the number genuinely will not come

There is a reframe that resolves more of these arguments than any estimation technique, and it is underused.

Stop asking what the probability is and ask whether the answer would change the decision.

Run the analysis across the whole plausible range of the disputed value rather than at a point. If the organization would fund the same contingency, choose the same procurement route and approve the same schedule whether the annual likelihood is one percent or ten, the number was never needed and the argument about it was a way of avoiding a decision that was already determined.

If the decision does flip somewhere in that range, the analysis has located the threshold. Now the question is bounded and answerable: what would it cost to find out which side of that threshold we are on, and is that less than the difference between the two decisions. That is a tractable question about the value of information, and it is a far better use of a meeting than another round of arguing about whether the risk is a four or a five.

Evidence matrix

ClaimEvidence tierSource
People asked for a range give one that is systematically too narrowPeer reviewed, foundationalLichtenstein, Fischhoff & Phillips (1982)
Experts can be scored against known answers and weighted by calibrationPeer reviewed, foundationalCooke (1991), Experts in Uncertainty
The classical model has a substantial recorded application base across domainsPeer reviewedCooke & Goossens (2008), RESS 93(5)
Training in probabilistic reasoning improves forecast accuracyPeer reviewedMellers et al. (2014), Psychological Science 25(5)
Loss can be decomposed into frequency and magnitude components that are separately estimableIndustry standardThe Open Group, FAIR risk taxonomy and analysis standards
Outcomes are better predicted from a class of comparable cases than from the case itselfPeer reviewedFlyvbjerg (2006), PMJ 37(3)
Tail behavior of a genuinely unprecedented event can be estimated from dataNot establishedExtrapolation assumes the generating mechanism is unchanged
Most disputed probabilities do not change the decision they are being argued aboutNamed field experienceCapital program practice, Mission Intelligence Systems

What to do with this

Find the largest risk currently recorded qualitatively because there is no data, and do two things to it.

Decompose it into components until you reach a level where the organization can say something from its own records. Then run the decision at both ends of the plausible range for whatever remains unknown. One of two outcomes follows, and both are useful. Either the decision holds across the range, in which case the risk can be closed as quantified enough and the debate ends. Or it does not, in which case you now know exactly what is worth finding out and roughly what it is worth paying to find it.

References

  1. Cooke, Roger M. Experts in Uncertainty: Opinion and Subjective Probability in Science. Oxford University Press, 1991. The classical model for performance-weighted expert judgment.
  2. Cooke, Roger M., and Louis L. H. J. Goossens. “TU Delft Expert Judgment Data Base.” Reliability Engineering & System Safety, vol. 93, no. 5, 2008, pp. 657–674. sciencedirect.com.
  3. Lichtenstein, Sarah, Baruch Fischhoff, and Lawrence D. Phillips. “Calibration of Probabilities: The State of the Art to 1980.” In Daniel Kahneman, Paul Slovic and Amos Tversky, eds., Judgment Under Uncertainty: Heuristics and Biases. Cambridge University Press, 1982, pp. 306–334.
  4. Mellers, Barbara, Lyle Ungar, Jonathan Baron, Jaime Ramos, Burcu Gurcay, Katrina Fincher, Sydney E. Scott, Don Moore, Pavel Atanasov, Samuel A. Swift, Terry Murray, Eric Stone, and Philip E. Tetlock. “Psychological Strategies for Winning a Geopolitical Forecasting Tournament.” Psychological Science, vol. 25, no. 5, 2014, pp. 1106–1115. doi.org/10.1177/0956797614524255.
  5. The Open Group. Risk Analysis (O-RA) Standard and Risk Taxonomy (O-RT) Standard. The Open Group. opengroup.org. The FAIR decomposition of loss event frequency and loss magnitude.
  6. Flyvbjerg, Bent. “From Nobel Prize to Project Management: Getting Risks Right.” Project Management Journal, vol. 37, no. 3, 2006, pp. 5–15. doi.org/10.1177/875697280603700302.
  7. International Organization for Standardization and International Electrotechnical Commission. Risk Management: Risk Assessment Techniques. IEC 31010:2019. iso.org/standard/72140.html. Technique selection including expert elicitation and Bayesian methods.
  8. U.S. Government Accountability Office. Cost Estimating and Assessment Guide: Best Practices for Developing and Managing Program Costs. GAO-20-195G, March 2020. gao.gov/products/gao-20-195g.
  9. Hubbard, Douglas W., and Dylan Evans. “Problems with Scoring Methods and Ordinal Scales in Risk Assessment.” IBM Journal of Research and Development, vol. 54, no. 3, 2010, pp. 2:1–2:10. doi.org/10.1147/JRD.2010.2042914.
DF

About the Author

Dan Flynn

Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build

Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.

His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.