Mission Intelligence Systems
Risk · Alignment

What P80 Means, and What It Does Not Promise

The percentile is a statement about a model, not a safety rating.

A P80 is the value a program has an 80 percent modeled chance of not exceeding. It is an output of a simulation rather than a number anyone selects, and it inherits every judgment that went into the model. Executives who understand that read the number correctly. Executives who do not tend to hear a guarantee.

Published

Key Takeaways

Research foundation

Percentile language is established practice in cost engineering and is required in writing by several public funding bodies, including the Federal Transit Administration and the guidance in the GAO cost and schedule assessment guides. The limits of the practice are equally well documented: Flyvbjerg and colleagues established the scale of systematic cost underestimation in public works, and Ben-David, Graham and Harvey found that realized market outcomes fell inside senior financial executives' own 80 percent confidence intervals only about a third of the time.

The first time a board is shown a cost range instead of a number, someone in the room asks the obvious question: which one is the real number? It is a fair question and it deserves a direct answer. There is no real number. There is a distribution of possible outcomes, and a percentile is a way of naming one point on it.

What does P80 mean?

A P80 is the value a program has an 80 percent modeled probability of not exceeding. If a simulation produces a P80 cost of 240 million dollars, the model is saying that in 80 percent of simulated outcomes the program finishes at or below 240 million, and in the remaining 20 percent it does not. A P80 date carries the same meaning in time rather than money.

Two things follow immediately, and both are routinely missed. First, the percentile is not an amount. It is a confidence level. The money is the gap between your baseline target and the percentile you have chosen to fund, so a program with a 200 million baseline and a 240 million P80 is carrying 40 million of contingency, which is 20 percent of baseline and not 80 percent of anything. Second, the percentile is not chosen. It is read off a curve that the model produced. You choose which point on the curve to fund. You do not choose where the curve sits.

Where does the P80 number come from?

It comes from combining uncertainty, not from adding a factor. Each significant cost item or schedule activity is given a range rather than a single value, discrete risk events from the register are attached with their own probabilities and impacts, correlations between items are set where they exist, and the whole structure is simulated thousands of times. Each iteration draws a plausible value for every input and produces one possible total. Run enough iterations and you have a distribution, from which any percentile can be read.

AACE International maintains recommended practices for exactly this, distinguishing methods that model inherent uncertainty by ranging the estimate from methods that model discrete risk events by expected value, and describing how the two are combined. The important point for an executive is not the method name. It is that the output is a curve, and the curve is a summary of the judgments that went in.

What does a P80 estimate not protect you against?

This is the part that gets skipped, and it is the part that matters when a program overruns a number that was supposed to be conservative.

A percentile does not cover scope you have not decided yet. If the program adds a station, a floor, or a module, the baseline moves and the old distribution is describing a different project. Nothing about the arithmetic protects you from a decision.

It does not cover what the model was not told. A simulation is exhaustive about the risks in the register and silent about everything else. Flyvbjerg, Skamris Holm and Buhl found systematic cost underestimation across hundreds of transport infrastructure projects, at a scale and consistency that is difficult to explain as ordinary estimating error. A model built by the same people, with the same optimism, will reproduce that optimism with more decimal places.

And it does not cover the possibility that the ranges themselves are too narrow. Ben-David, Graham and Harvey analyzed roughly thirteen thousand forecast distributions provided by senior financial executives and found that realized outcomes fell inside those executives' own 80 percent confidence intervals only about 36 percent of the time. These were not careless people. Overconfidence in ranges is a property of expert judgment, not a failure of diligence, and it is the single most common reason a P80 turns out to have been a P50.

The honest framing is narrow and useful: a P80 tells you that, if the model is right about the shape of the uncertainty it was given, you have covered four fifths of the modeled downside. Everything outside that clause is still yours.

Who decides which percentile a program is held to?

Usually not the program. If public money is involved, the level is set in writing by the body providing it, and the levels are not uniform.

The Federal Transit Administration states in its Oversight Procedure 40 that since July 2018 it has adopted the P65 confidence level, the 65th percentile, to determine cost contingency for the projects it funds, and that sponsors are required to provide cash funding at that level. For schedule it applies a different rule, recommending the larger of 125 percent of the stripped and adjusted base schedule remaining critical path duration or the 65th percentile, and it directs its oversight contractor to report the P40, P50, P65 and P80 confidence levels so reviewers can see the whole curve rather than one point on it.

Washington State DOT does something more interesting still. Its policy sets a pair of values rather than one: for projects under 100 million dollars the budgeted legislative value is the 60th percentile and the operational value is the 40th, and for projects over 100 million the pair is the 50th and the 30th. The difference between the two is held explicitly as the project risk reserve. Note the direction, which surprises people: the larger the project, the lower the percentile the legislature is asked to fund. That is a deliberate governance choice about where reserve is held and who releases it, not a statement that big projects are safer.

The GAO cost estimating and schedule assessment guides treat uncertainty analysis as a practice programs are assessed against rather than prescribing a single percentile, which is why federal programs are more often criticized for not having done the analysis than for having chosen the wrong confidence level.

The practical consequence is that P80 is a common industry default rather than a universal requirement, and a program that assumes it without checking can fund the wrong number in either direction.

Why does the P80 move?

Because it is a statement about remaining uncertainty, and remaining uncertainty changes. This is a feature, and it is where the number stops being a reporting artifact and starts being a management instrument.

When a high variance item is resolved, its spread leaves the model. The distribution narrows and the P80 pulls in toward the baseline. That movement is the mathematical justification for releasing contingency, and it is a far better basis for the decision than the calendar. When execution reveals that an activity is running past its most likely duration, the distribution widens and the P80 pushes out, which gives leadership a quantified warning while there is still time to act on it.

A P80 that has not moved in eighteen months on a live program is not evidence of stability. It is evidence that nobody has re-run the model.

Worked example: why you cannot add P80s

This is the most common arithmetic error in contingency setting, and it is easiest to see with numbers small enough to check by hand. Three activities, each estimated with a minimum, a most likely and a maximum, in units of days.

ActivityMinMost likelyMaxMeanIts own P80
Site preparation80100150110.0123.5
Utility relocation607512085.096.8
Structural works40509060.070.0
Total255.0290.3

The tempting move is to read 290.3 as the program's P80, since it is the sum of three P80s. Simulate the total instead, drawing each activity independently from its own distribution, and the answer is different.

So the real P80 is 274, not 290.3. Adding the individual P80s overstates the requirement by 16.3 days, about six percent of the total. Put the other way round, and this is the sentence worth carrying into a meeting: the figure obtained by adding P80s sits at roughly the 93.5th percentile of the actual distribution. It is not a P80 at all. It is close to a P95 that has been labeled as something else.

The mechanism is straightforward once stated. For all three activities to simultaneously land at or above their individual 80th percentiles requires a coincidence. Some overrun while others come in under, and the offsetting is what pulls the total in. This is the same portfolio effect that makes a diversified position less volatile than its holdings.

The exception, which matters

Run the same three activities again, but assume they are perfectly correlated, so that whatever drives one to its upper bound drives all of them there together. Now the P80 of the total is 290.3 days, exactly the sum of the individual P80s, and the naive addition turns out to be correct.

That is not a curiosity. It is the boundary condition. Adding P80s is only right when everything moves together perfectly, and nothing does, but plenty of things move together partly: a single subcontractor across three packages, one weather season, one labor market, one design authority. The gap between 274 and 290 is entirely a statement about how much the activities share, and an organization that has never modeled correlation has implicitly assumed one of these two extremes without choosing which.

The practical instruction is short. Never aggregate percentiles by addition. Aggregate the distributions and then read the percentile off the total. And when someone objects that the simulated number looks too low, the honest answer is that it looks low because the previous method was quietly reserving to P93.

Why is this an Alignment problem?

Because a percentile is a shared instrument or it is nothing. The number is produced by aggregating individual judgments about probability and consequence, which means it is only comparable across a portfolio if those judgments were made on a common scale. An organization that has never defined what its own probability language means cannot produce a defensible percentile, because the inputs to the model are the same intuitions that populate the register. That argument is made at length in Why 'Likely' Is Not a Probability, and it is the prerequisite for everything on this page.

Alignment also governs how the number survives contact with a board. A percentile presented without its assumptions invites the question it cannot answer, which is whether the program is safe. A percentile presented with what it excludes, scope decisions, unmodeled risk, and the known narrowness of expert ranges, gives a board something it can actually govern with.

Evidence matrix

ClaimEvidence tierSource
Expert confidence intervals are systematically too narrowPeer reviewedBen-David, Graham & Harvey (2013), realized outcomes inside executives' 80 percent intervals about 36 percent of the time
Public works costs are systematically underestimatedPeer reviewedFlyvbjerg, Skamris Holm & Buhl (2002); Flyvbjerg (2006)
FTA funds cost contingency at the 65th percentileGovernment requirementFTA Oversight Procedure 40, Risk and Contingency Review (October 2023)
Percentile choice is a governance decision, not a technical defaultGovernment policyWSDOT Policy Statement P 2047.00, legislative and operational percentile pairs
Verbal probability terms are not shared instrumentsPeer reviewedBeyth-Marom (1982); Budescu, Broomell & Por (2009)
A percentile is only as good as the alignment behind its inputsFour A's interpretationBuilders Build, Alignment

What to do with this

Three questions will tell you whether a percentile in front of you is worth trusting. Ask what confidence level the funding body actually requires, in writing, rather than assuming P80. Ask what the model was not told, and specifically whether the ranges were set by the same people whose plan they describe. And ask when it was last re-run, because a percentile that does not move is not a forecast.

None of that requires becoming a modeler. It requires treating the number as a claim with conditions attached, which is how every other quantitative claim reaching a board is already treated.

References

  1. Ben-David, Itzhak, John R. Graham, and Campbell R. Harvey. “Managerial Miscalibration.” The Quarterly Journal of Economics, vol. 128, no. 4, 2013, pp. 1547–1584. doi.org/10.1093/qje/qjt023. Found that realized S&P 500 returns fell within senior financial executives' own 80 percent confidence intervals only about 36 percent of the time, across roughly 13,300 forecast distributions.
  2. Flyvbjerg, Bent, Mette K. Skamris Holm, and Søren L. Buhl. “Underestimating Costs in Public Works Projects: Error or Lie?” Journal of the American Planning Association, vol. 68, no. 3, 2002, pp. 279–295. doi.org/10.1080/01944360208976273.
  3. Flyvbjerg, Bent. “From Nobel Prize to Project Management: Getting Risks Right.” Project Management Journal, vol. 37, no. 3, 2006, pp. 5–15. doi.org/10.1177/875697280603700302. Introduces reference class forecasting as a corrective to inside-view estimating.
  4. Kahneman, Daniel, and Amos Tversky. Intuitive Prediction: Biases and Corrective Procedures. Technical Report PTR-1042-77-6, Decision Research, June 1977. Prepared for the Defense Advanced Research Projects Agency. DTIC accession ADA047747. apps.dtic.mil/sti/pdfs/ADA047747.pdf.
  5. Beyth-Marom, Ruth. “How Probable Is Probable? A Numerical Translation of Verbal Probability Expressions.” Journal of Forecasting, vol. 1, no. 3, 1982, pp. 257–269. doi.org/10.1002/for.3980010305.
  6. Budescu, David V., Stephen Broomell, and Han-Hui Por. “Improving Communication of Uncertainty in the Reports of the Intergovernmental Panel on Climate Change.” Psychological Science, vol. 20, no. 3, 2009, pp. 299–308. doi.org/10.1111/j.1467-9280.2009.02284.x.
  7. Spiegelhalter, David, Mike Pearson, and Ian Short. “Visualizing Uncertainty About the Future.” Science, vol. 333, no. 6048, 2011, pp. 1393–1400. doi.org/10.1126/science.1191181.
  8. Federal Transit Administration, Office of Capital Project Management. OP 40: Risk and Contingency Review. U.S. Department of Transportation, October 2023. transit.dot.gov. States that since July 2018 FTA has adopted the P65 confidence level to determine cost contingency, and directs reporting of P40, P50, P65 and P80.
  9. U.S. Government Accountability Office. Cost Estimating and Assessment Guide: Best Practices for Developing and Managing Program Costs. GAO-20-195G, March 2020. gao.gov/products/gao-20-195g.
  10. U.S. Government Accountability Office. GAO Schedule Assessment Guide: Best Practices for Project Schedules. GAO-16-89G, December 2015. gao.gov/products/gao-16-89g.
  11. Washington State Department of Transportation. Policy Statement P 2047.00: Estimating Project Budget and Uncertainty. Effective 20 March 2017. wsdot.wa.gov. Sets legislative and operational budget values at the 60th and 40th percentiles under 100 million dollars, and the 50th and 30th above it.
  12. AACE International. Recommended Practice 118R-21: Cost Risk Analysis and Contingency Determination Using Estimate Ranging for Inherent Risks with Monte Carlo Simulation, and Recommended Practice 123R-22: Integrated Cost and Schedule Risk Analysis and Contingency Determination Using Estimate Ranging and Expected Value with Monte Carlo Simulation. Cited for scope and applicability; the practices themselves are available to AACE members.
  13. International Organization for Standardization. Risk Management: Guidelines. ISO 31000:2018. Clause 4(e) establishes that risk management is dynamic, anticipating and responding to change as an organization's context changes.
DF

About the Author

Dan Flynn

Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build

Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.

His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.