The Premortem and the Reference Class
Two corrections for the same bias, working from opposite ends.
One technique asks your team to imagine the failure in detail. The other ignores your team entirely and asks what happened to projects like yours. They are usually presented as alternatives. They are closer to complements, and knowing which question each one answers is what makes either useful.
Published
Key Takeaways
- Both techniques target the planning fallacy, the tendency to estimate from the plan in front of you rather than from what happened to comparable efforts. They attack it from opposite directions.
- A premortem produces specific local failure modes and no sense of magnitude. A reference class produces a defensible magnitude and no idea what to do. Each is weak exactly where the other is strong.
- On evidence, be careful with the premortem. There is no peer-reviewed journal study showing it improves risk identification. The support is a laboratory finding on prospective hindsight plus conference-grade evidence that it reduces overconfidence, which is a narrower claim than the one usually made for it.
Research foundation
The premortem's mechanism rests on Mitchell, Russo and Pennington (1989) on prospective hindsight, with Klein's practitioner articulation of the technique. Reference class forecasting rests on Flyvbjerg (2006, 2008), building on Kahneman and Tversky's account of the planning fallacy and the outside view. This article states a limitation that most treatments omit: no peer-reviewed journal study demonstrates that premortems improve risk identification. The available evaluation is a refereed conference paper measuring plan confidence. The Four A's are the executive lens applied to that evidence.
Both techniques exist because of the same well-documented problem. People estimate from the plan in front of them, and the plan in front of them describes the intended path rather than the distribution of paths. Kahneman and Tversky named the failure and proposed the corrective in the same breath: take the outside view, and anchor to what happened in comparable cases rather than building up from the route you intend to take.
How does a premortem work?
By changing the tense of the question.
Asking a team what might go wrong produces a familiar, cautious list. Telling a team that the project has already failed, twelve months from now, and asking them to explain why, produces something different: concrete, specific, sometimes uncomfortably candid accounts. The underlying mechanism is prospective hindsight. Mitchell, Russo and Pennington found that framing a future event as though it had already occurred increased people's ability to generate explanations for it, which is the effect the technique is engineered to exploit.
The second, less discussed mechanism is social. A premortem gives license. It converts “I have a concern about the data migration” from an act of disloyalty into an assigned exercise, which matters far more in a hierarchical organization than the cognitive trick does. That is why the technique often surfaces things people already knew, which is a strange kind of success and a real one.
Does the premortem actually work?
Less conclusively than it is usually sold, and it is worth being precise because this is the kind of claim that gets repeated until it sounds settled.
There is no peer-reviewed journal study demonstrating that premortems improve risk identification on real projects. I looked for one. What exists is a refereed conference paper by Veinott, Klein and Wiggins evaluating the technique against alternatives and finding it reduced plan confidence relative to a pro and con exercise, plus a later descriptive case study and a master's thesis. Reduced overconfidence is a real and useful outcome. It is not the same claim as better risk identification, and treating the two as equivalent is exactly the sort of slippage this library is supposed to avoid.
So the defensible position is this: the premortem reliably surfaces more specific concerns than an open question does, it reduces overconfidence in the plan, and its effect on whether the risks that actually materialize get caught has not been established. That is still a good reason to run one. It is not a reason to treat it as a control.
How does reference class forecasting work?
By refusing to look at your project at all.
Rather than building an estimate up from tasks, reference class forecasting places the project in a class of completed comparable projects and predicts from the distribution of their outcomes. Flyvbjerg set out the method for project management and later documented its application in practice, including adoption by the UK Department for Transport. The operational form is an uplift applied to the naive estimate, calibrated to how far comparable projects historically overran.
The strength is that it is immune to the specific optimism of the specific team, because it never consults them. The weakness follows directly: it produces a magnitude and no diagnosis. It will tell you that projects of this type finish thirty percent over, and nothing whatsoever about what to do on Monday.
Why do you need both?
Because each is weak precisely where the other is strong, and the failure modes of using one alone are predictable.
Use only the reference class and you get a defensible number that nobody in the room believes applies to them, because every team is certain their project is the exception. The uplift gets negotiated down, quietly, in the way that unowned numbers always do.
Use only the premortem and you get a rich list of concerns with no sense of scale. Twenty failure modes are identified, all plausible, and nothing tells you whether they add up to a two percent problem or a forty percent one. Lists without magnitudes get triaged by whoever speaks loudest.
The sequence that works runs outside-in. The reference class establishes how much uncertainty a project of this type carries, which sets the size of the container. The premortem populates what that uncertainty is made of here, which makes it actionable. Then the quantified model, described in From Risk Register to Quantified Model, attaches the specific items to specific parts of the estimate so the two views can be reconciled.
Where they disagree is the most valuable output of all. If the premortem's identified risks sum to far less than the reference class implies, the difference is what the team has not thought of yet, and that gap is the most honest thing on the page.
Why is this an Adaptability problem?
Because both techniques ask an organization to take unwelcome information seriously before anything has gone wrong, and that is a learned capability rather than a procedural one.
A premortem in an organization that punishes dissent produces a polite list of risks nobody owns. A reference class uplift in an organization that treats every project as unprecedented gets argued away in a single meeting. Neither technique fails for technical reasons. They fail because the conditions that would let a team act on the output were not there, which is the same pattern as a register that is completed and never consulted, described in The Risk Register Nobody Reads.
The diagnostic is what happened last time. If your organization has run a premortem, find the list it produced and check whether any item changed the plan. If none did, running another one will not help, and the problem was never the technique.
Evidence matrix
| Claim | Evidence tier | Source |
|---|---|---|
| Prospective hindsight increases generation of concrete explanations | Peer reviewed | Mitchell, Russo & Pennington (1989), JBDM 2(1) |
| Premortems reduce plan confidence relative to a pro and con exercise | Refereed conference paper | Veinott, Klein & Wiggins (2010), ISCRAM |
| Premortems improve risk identification | Not established | No peer-reviewed journal study located |
| Outcomes are better predicted from a class of comparable projects | Peer reviewed | Flyvbjerg (2006), PMJ 37(3); Flyvbjerg (2008), European Planning Studies 16(1) |
| Inside-view planning systematically underestimates cost and duration | Peer reviewed, foundational | Kahneman & Tversky (1977) |
| Both techniques fail on conditions, not on method | Four A's interpretation | Builders Build, Adaptability |
What to do with this
Run them in order and record both numbers. Establish the reference class uplift first, before the team has anchored on its own estimate. Then run the premortem and total the identified impacts. If the two disagree materially, do not reconcile them by adjusting either one. Write the gap down and say what it represents, because it is the part of the risk nobody has named yet.
References
- Mitchell, Deborah J., J. Edward Russo, and Nancy Pennington. “Back to the Future: Temporal Perspective in the Explanation of Events.” Journal of Behavioral Decision Making, vol. 2, no. 1, 1989, pp. 25–38. doi.org/10.1002/bdm.3960020103. The prospective hindsight finding underpinning the premortem.
- Klein, Gary. “Performing a Project Premortem.” IEEE Engineering Management Review, vol. 36, no. 2, 2008, pp. 103–104. doi.org/10.1109/EMR.2008.4534313. Reprint of the September 2007 Harvard Business Review article (reprint F0709A), cited here because the journal version carries verifiable metadata.
- Veinott, Beth, Gary Klein, and Sterling Wiggins. “Evaluating the Effectiveness of the PreMortem Technique on Plan Confidence.” Proceedings of the 7th International ISCRAM Conference, Seattle, 2010. idl.iscram.org. A refereed conference paper measuring plan confidence, not risk identification quality.
- Flyvbjerg, Bent. “From Nobel Prize to Project Management: Getting Risks Right.” Project Management Journal, vol. 37, no. 3, 2006, pp. 5–15. doi.org/10.1177/875697280603700302.
- Flyvbjerg, Bent. “Curbing Optimism Bias and Strategic Misrepresentation in Planning: Reference Class Forecasting in Practice.” European Planning Studies, vol. 16, no. 1, 2008, pp. 3–21. doi.org/10.1080/09654310701747936. Reference class forecasting as applied in practice.
- Kahneman, Daniel, and Amos Tversky. Intuitive Prediction: Biases and Corrective Procedures. Technical Report PTR-1042-77-6, Decision Research, June 1977. DTIC accession ADA047747. apps.dtic.mil. The planning fallacy and the outside view.
- Flyvbjerg, Bent, Mette K. Skamris Holm, and Søren L. Buhl. “Underestimating Costs in Public Works Projects: Error or Lie?” Journal of the American Planning Association, vol. 68, no. 3, 2002, pp. 279–295. doi.org/10.1080/01944360208976273. The empirical scale of the problem both techniques address.
- Ward, Stephen, and Chris Chapman. “Transforming Project Risk Management into Project Uncertainty Management.” International Journal of Project Management, vol. 21, no. 2, 2003, pp. 97–105. doi.org/10.1016/S0263-7863(01)00080-1.
- International Organization for Standardization. Risk Management: Guidelines. ISO 31000:2018, Clause 4(f) on best available information. iso.org/standard/65694.html.
About the Author
Dan Flynn
Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build
Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.
His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.
