How to Present a Probabilistic Cost Range to a Council
The single number is not more certain. It is less honest about what it does not know.
The fear is always the same: bring a range to a public meeting and someone will say you do not know what the project costs. It is a real risk and it is manageable, but not by explaining the mathematics harder. It is managed by presenting a decision rather than a method.
Published
Key Takeaways
- Lead with the decision, not the analysis. A governing body has to choose how much downside to fund and who releases the reserve. It does not need to understand simulation to make that choice, and the presentations that fail are the ones that explain the method first.
- Two numbers and one question is the whole structure: the value the program is expected to land near, the higher value that covers most of the plausible downside, and which of the two to appropriate. Say the confidence level plainly once and then stop talking about statistics.
- Anticipate four questions, because you will always get them: which number is real, why the range is so wide, whether the contingency will simply be spent, and what happens if you are wrong. Each has an honest answer, and having it ready is the difference between rigor and retreat.
Research foundation
Three bodies of evidence sit behind this. Work on communicating uncertainty establishes that presentation format changes what an audience concludes: Spiegelhalter, Pearson and Short reviewed how uncertainty about the future can be visualized, and Budescu, Broomell and Por showed experimentally that readers systematically misinterpret verbal probability terms even when numerical ranges are supplied alongside them. Work on estimating establishes why the range matters: Flyvbjerg and colleagues documented systematic cost underestimation in public works, and Ben-David, Graham and Harvey found expert confidence intervals to be far too narrow in practice. Public sector percentile policy is set out in FTA and WSDOT documents. The Four A's are the executive lens applied to that evidence.
I have watched a good analysis lose a room in under a minute. The slide went up with a histogram on it, the presenter said the words Monte Carlo simulation, and a council member asked, reasonably, whether anyone could just tell them what the bridge was going to cost. Everything after that was recovery.
The analysis was not the problem. The order was.
Why does a single number feel more credible than a range?
Because a single number sounds like knowledge and a range sounds like hedging. That instinct is understandable and it is wrong, and it is worth being able to say why in one sentence: the single number was always a range with the uncertainty deleted.
Deleting it does not remove it. It relocates it. The uncertainty that was not in the estimate turns up later as a change order, a supplemental appropriation, or a schedule slip, at which point it is a governance failure rather than a planning input. Flyvbjerg, Skamris Holm and Buhl documented cost underestimation across hundreds of transport infrastructure projects at a scale and consistency that ordinary estimating error does not explain, and the mechanism they describe is exactly this: single-point estimates presented with a confidence they were never entitled to.
So the honest framing, and it holds up under challenge, is that you are not less certain than the person who brings one number. You are telling the council what the one number was hiding.
What does a council actually have to decide?
This is the question that reorders the whole presentation, and most technical presenters never ask it.
A governing body is not being asked to evaluate a model. It is being asked to answer two things. How much money do we set aside, and who is allowed to release the part we hope not to spend. That is it. Everything in the analysis either serves those two decisions or belongs in an appendix.
Once you accept that, the structure writes itself. You need the value the program is expected to land near, a higher value that covers most of the plausible downside, the difference between them expressed as a reserve, and a rule for who releases it and on what evidence. Four items. The percentile is one sentence inside that, not the subject of the presentation.
How do you present a percentile without teaching statistics?
Say it once, in outcome language rather than distribution language, and then move on.
What works is a sentence of the form: funding the higher figure covers roughly four out of five plausible outcomes for this program. What does not work is explaining what an eightieth percentile is, because the moment you begin defining terms you have signalled that the audience needs to understand your method in order to trust your recommendation, and they will conclude, correctly, that they do not understand your method.
Two further presentation choices matter more than people expect. Give the range a small number of named drivers rather than a distribution shape: three items that account for most of the spread, in plain language, so the width has a cause a non-specialist can hold. And show the two numbers with the gap between them labelled as what it is, a reserve with a purpose, rather than showing a probability curve. Spiegelhalter, Pearson and Short's review of visualizing uncertainty makes the general point that format is not neutral, and the specific lesson for a public meeting is that the audience should be looking at a decision, not at a density function.
One trap to avoid, and it is well evidenced. Do not substitute words for the numbers in the belief that it will land more softly. Budescu, Broomell and Por showed that readers reinterpret verbal probability terms according to their own priors even when numerical ranges are printed alongside them, which means saying a cost overrun is unlikely will be heard as several different probabilities by several different council members. If you have a number, use the number.
What will you be asked, and what are the honest answers?
Four questions, near enough every time.
Which number is the real one?
Neither, and that is the point. Both are outputs from the same analysis at different confidence levels. The lower one is what the program is likely to land near if things go about as expected; the higher one is what it takes to cover most of what could go differently. The recommendation is to appropriate the higher figure and manage to the lower one.
Why is the range so wide?
Name the drivers. A range is wide for reasons, usually two or three of them, and saying so converts an apparent weakness into a work plan. It is also the moment to say what will narrow it and when: after the geotechnical work, after the permit decision, after the procurement closes. A range with a schedule for narrowing is a plan. A range without one is a shrug.
If we approve the contingency, will it just get spent?
This is the sharpest question and it deserves the most concrete answer, because it is fundamentally about authority rather than about money. The answer is a release rule: who may authorize a draw, against what evidence, and reported to whom. Washington State DOT's policy is a useful reference model here, because it makes the structure explicit: a legislative budget value at one percentile and an operational working value at a lower one, with the gap between them held as the project risk reserve rather than distributed into the estimate. The reserve is visible, and releasing it is an event with a name. Without a rule of that kind, the concern behind the question is entirely justified, which is the argument made in Risk Ownership Without Authority.
What if you are wrong?
Answer it directly, because the alternative is worse. The range can be wrong in two ways: the ranges themselves may have been too narrow, which is the well documented tendency of expert judgment, or the program may change scope, which no confidence level covers. Ben-David, Graham and Harvey found that realized outcomes fell inside senior financial executives' own 80 percent confidence intervals only about a third of the time, which is worth knowing and worth saying, because it explains why the analysis is re-run rather than filed. A presenter who volunteers the limits of the method is markedly harder to attack than one who defends it as complete.
What happens the second time you come back?
The second appearance is what makes the first one credible, and it is the part most organizations skip.
If the analysis is re-run as the program progresses, the range narrows as uncertainty is resolved, and you can return with a smaller reserve request and the evidence for why. That is the strongest possible demonstration that the first number was analysis rather than padding. If instead the range stays identical for two years, the council will reasonably conclude it was a negotiating position.
This is also where the reserve stops being a political liability and becomes an instrument. Releasing contingency back because a risk was retired is a governance win that can be reported. ISO 31000:2018 puts the underlying principle plainly in stating that risk management is dynamic, anticipating and responding to change as the organization's context changes; a number that never moves is not managing anything.
Why is this an Alignment problem?
Because the presentation is where a technical output either becomes a shared understanding or fails to. The model can be sound, the percentile correctly chosen, the reserve properly sized, and none of it governs anything if the room leaves with a different understanding of what was approved than the one you intended.
The specific alignment failure to watch for is that the council believes it approved a cap while the program believes it approved a target, or the reverse. That divergence is invisible in the minutes and expensive eighteen months later. The cheapest protection is to state the two numbers, the reserve, and the release rule in the resolution language itself rather than in the slide deck, so that what was agreed survives the meeting. Agreement in the room is not alignment, which is the argument in Risk Appetite Is Not Alignment, and a capital appropriation is an unusually costly place to learn the difference.
Evidence matrix
| Claim | Evidence tier | Source |
|---|---|---|
| Single-point public works estimates are systematically low | Peer reviewed | Flyvbjerg, Skamris Holm & Buhl (2002); Flyvbjerg (2006) |
| Expert ranges are too narrow, so the method has stated limits | Peer reviewed | Ben-David, Graham & Harvey (2013) |
| Verbal probability terms are reinterpreted by each reader | Peer reviewed, experimental | Budescu, Broomell & Por (2009); Beyth-Marom (1982) |
| Presentation format changes what an audience concludes | Peer reviewed review | Spiegelhalter, Pearson & Short (2011) |
| A two-tier budget with an explicit reserve is established public practice | Government policy | WSDOT Policy Statement P 2047.00; FTA Oversight Procedure 40 |
| Approval without a shared understanding of the numbers is agreement, not alignment | Four A's interpretation | Builders Build, Alignment |
What to do before the next meeting
Write the resolution language first, before the deck. Two numbers, the reserve, the release rule, and the date of the next re-forecast. If you cannot state those five things in four sentences, the analysis is not ready to present, and no amount of explaining the simulation will compensate.
References
- Spiegelhalter, David, Mike Pearson, and Ian Short. “Visualizing Uncertainty About the Future.” Science, vol. 333, no. 6048, 2011, pp. 1393–1400. doi.org/10.1126/science.1191181.
- Budescu, David V., Stephen Broomell, and Han-Hui Por. “Improving Communication of Uncertainty in the Reports of the Intergovernmental Panel on Climate Change.” Psychological Science, vol. 20, no. 3, 2009, pp. 299–308. doi.org/10.1111/j.1467-9280.2009.02284.x. Readers reinterpreted verbal probability terms according to their own priors even when numerical ranges were supplied.
- Beyth-Marom, Ruth. “How Probable Is Probable? A Numerical Translation of Verbal Probability Expressions.” Journal of Forecasting, vol. 1, no. 3, 1982, pp. 257–269. doi.org/10.1002/for.3980010305.
- Ben-David, Itzhak, John R. Graham, and Campbell R. Harvey. “Managerial Miscalibration.” The Quarterly Journal of Economics, vol. 128, no. 4, 2013, pp. 1547–1584. doi.org/10.1093/qje/qjt023.
- Flyvbjerg, Bent, Mette K. Skamris Holm, and Søren L. Buhl. “Underestimating Costs in Public Works Projects: Error or Lie?” Journal of the American Planning Association, vol. 68, no. 3, 2002, pp. 279–295. doi.org/10.1080/01944360208976273.
- Flyvbjerg, Bent. “From Nobel Prize to Project Management: Getting Risks Right.” Project Management Journal, vol. 37, no. 3, 2006, pp. 5–15. doi.org/10.1177/875697280603700302.
- Klein, Gary. “Performing a Project Premortem.” Harvard Business Review, vol. 85, no. 9, September 2007, pp. 18–19. hbr.org. A practitioner source, not peer reviewed; cited for the prospective hindsight technique used to surface the drivers behind a range.
- Washington State Department of Transportation. Policy Statement P 2047.00: Estimating Project Budget and Uncertainty. Effective 20 March 2017. wsdot.wa.gov. Establishes paired legislative and operational budget values, with the difference held as the project risk reserve.
- Federal Transit Administration, Office of Capital Project Management. OP 40: Risk and Contingency Review. U.S. Department of Transportation, October 2023. transit.dot.gov. States that since July 2018 FTA has adopted the P65 confidence level to determine cost contingency for the projects it funds.
- U.S. Government Accountability Office. Cost Estimating and Assessment Guide: Best Practices for Developing and Managing Program Costs. GAO-20-195G, March 2020. gao.gov/products/gao-20-195g.
- International Organization for Standardization. Risk Management: Guidelines. ISO 31000:2018. iso.org/standard/65694.html. Clause 4(e) establishes that risk management is dynamic and responds to change in context.
About the Author
Dan Flynn
Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build
Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.
His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.
