How Do We Know Our Risk Management Is Working?
Nothing went wrong is not an answer. Neither is the register count.
It is the fairest question an executive can ask about a function that costs money and produces documents, and it is answered almost everywhere with metrics that cannot fail. There is a measure that genuinely tests the work. It is cheap, it uses records the organization already keeps, and it requires one thing that turns out to be the real obstacle.
Published
Key Takeaways
- Activity metrics cannot fail. The register count rises when nothing is closed, review compliance measures the diary, and mitigation completion improves by choosing easy mitigations. All three can improve while exposure worsens.
- Outcome metrics fail for the opposite reason. Rare events are absent most years regardless of what anyone did, and a program that completes once every few years gives too few observations, too late, with too many confounders.
- Calibration is the measure that works. Every bid, every completed activity and every resolved risk tests a range or a probability the organization already stated. Scored at the input level rather than the program level, this yields hundreds of observations a year instead of one a decade.
Research foundation
The value evidence is Hoyt and Liebenberg (2011) and Pagach and Warr (2011) in the Journal of Risk and Insurance, with Gatzert and Martin (2015) in Risk Management and Insurance Review for the state of the literature. Scoring rests on Brier (1950) and the calibration literature consolidated by Lichtenstein, Fischhoff and Phillips (1982), with Mellers and colleagues (2014) on trainability. One limit is stated: no peer-reviewed study establishes that calibration scoring of program risk inputs improves program outcomes. The Four A's are the executive lens applied on top.
The reporting pack answers the question with three numbers. There are 147 risks in the register, up from 138. Ninety-four percent of scheduled risk reviews were held. Eighty-one percent of mitigation actions are on track.
Now consider what would have to happen for any of those to get worse. The register count falls when risks are closed, so a rising count is a sign of diligence and a falling one is a sign of diligence. Review compliance measures whether meetings occurred. Mitigation completion rises when the team selects mitigations it can finish. Every one of these can improve steadily while the organization walks into the thing that damages it.
A metric that cannot get worse when the situation gets worse is not a measure. It is a reassurance mechanism with a number attached.
Why outcome measurement does not rescue this
The obvious correction is to measure results instead of effort, and it does not work either, for reasons worth understanding before anyone builds a scorecard on it.
Rare events are absent most of the time whatever the organization does. An entity with no risk function will also report no catastrophic loss in the majority of years, so the observation fails to distinguish between competence and luck. This is the base rate problem and it is fatal to the naive version of outcome measurement.
Then there is the counting problem. A major capital program completes every few years. That is not a sample, and each of those few outcomes is driven by procurement conditions, labor markets, political change, weather and design quality alongside anything the risk process did. Attributing the result to the risk function requires isolating a small effect inside enormous noise with a handful of observations.
And there is the lag. By the time a program lands, the people who built the model have moved, the method has been revised twice, and the feedback arrives too late to change behavior. Feedback that arrives after the behavior it should correct has already been repeated is not feedback.
What the research establishes, and what it does not
This deserves a careful reading because the summaries in circulation are more confident than the papers.
Hoyt and Liebenberg examined US insurers and found a positive and statistically significant relation between enterprise risk management and firm value. That is a real finding in a peer-reviewed journal and it is the strongest single result usually cited.
Pagach and Warr examined firms that appointed chief risk officers and found the adopters were systematically different before adoption: larger, more leveraged, with more volatile earnings. On the benefits side their results offered limited support for the outcomes the framework is expected to produce. That is a materially more cautious picture than the first result alone conveys.
Gatzert and Martin reviewed the accumulated literature and concluded that findings are generally positive but depend substantially on how adoption is measured.
That last point is the one to carry away. Nearly all of this research proxies risk management by the presence of a chief risk officer, or by an external ERM rating, or by a disclosure. Those are measures of whether the organization adopted the structure. None of them measures whether the work is done well. So the literature can tell you something about the average effect of having a function, and almost nothing about whether your function is any good, which was the question.
The measure that actually tests the work
Risk management makes claims about uncertainty. Claims about uncertainty are testable. That is the whole idea, and it is unfamiliar mainly because the claims are rarely recorded in a form that permits scoring.
Interval hit rate
If a range was stated at eighty percent confidence, then across many such statements roughly eighty percent of actual outcomes should fall inside the range. Fewer means the ranges are too narrow, which is the direction the calibration literature predicts and which Lichtenstein, Fischhoff and Phillips documented as the dominant and highly replicated finding. More means the ranges are too wide and the analysis is not discriminating.
This is computable from records the organization already has, and it needs no new process. It needs the original range preserved.
A proper score on discrete risks
Where risks carried explicit probabilities, those probabilities can be scored against what happened using a proper scoring rule, the oldest and simplest of which is Brier's. A proper rule has the property that a forecaster minimizes their expected score by stating what they actually believe, so it cannot be improved by hedging toward the middle. Score by estimator and the organization learns which of its people give useful probabilities, which is information no maturity assessment will ever produce.
Score the inputs, not the programs
Here is the move that resolves the small-sample problem, and it is the most useful idea on this page.
The organization does not have to wait for programs to complete. Every bid that arrives is a test of a modelled cost range. Every activity that finishes is a test of a duration range. Every risk that occurs or is retired is a test of a stated probability. A single mid-sized program generates hundreds of these tests a year, and they are all currently discarded.
Score at that level and the feedback loop closes in months. You learn, concretely, that the organization's ranges on utility work are consistently forty percent too narrow while its ranges on structural work are well calibrated, and that is an actionable finding about a specific estimating practice rather than a verdict on a department.
The obstacle is not analytical and not financial. It is that scoring requires the original forecast to survive unrevised, and organizations habitually update the record to reflect what they now know. A range that was quietly widened after the bid came in cannot be scored, and the widening usually feels like diligence at the time.
Time from signal to decision
One non-statistical measure worth adding, because it captures something the others miss. For each significant risk that materialized, how long elapsed between the first point at which the organization possessed the information and the point at which someone made a decision about it.
This is measurable retrospectively from documents, it is uncomfortable in a productive way, and it tests the conditions rather than the analysis. A model can be excellent while nothing happens because of it, and this is the measure that catches that.
What to stop reporting
- Register size. It measures documentation effort and moves in both directions for good reasons.
- Review attendance and schedule compliance. It measures the diary.
- Mitigation actions completed. It measures activity and rewards choosing easy mitigations over material ones.
- Maturity self-assessments unaccompanied by any outcome measure. They record what the organization believes about itself.
An honest limitation before this becomes a scorecard. No peer-reviewed study establishes that calibration scoring of program risk inputs improves program outcomes. The case for it rests on the demonstrated existence of miscalibration, on the demonstrated trainability of calibration in forecasting research, and on the ordinary principle that a claim which cannot be checked will not improve. That is a reasoned argument rather than a validated intervention, and it should be presented as one.
Evidence matrix
| Claim | Evidence tier | Source |
|---|---|---|
| Enterprise risk management is positively related to firm value among US insurers | Peer reviewed | Hoyt & Liebenberg (2011), JRI 78(4) |
| Firms hiring chief risk officers differ systematically beforehand; benefit evidence is limited | Peer reviewed | Pagach & Warr (2011), JRI 78(1) |
| Findings across the literature depend substantially on how adoption is proxied | Peer reviewed review | Gatzert & Martin (2015), RMIR 18(1) |
| Stated ranges are systematically too narrow | Peer reviewed, foundational | Lichtenstein, Fischhoff & Phillips (1982) |
| A proper scoring rule cannot be improved by stating other than your true belief | Peer reviewed, foundational | Brier (1950), Monthly Weather Review 78(1) |
| Calibration improves with training | Peer reviewed | Mellers et al. (2014), Psychological Science 25(5) |
| Calibration scoring of program risk inputs improves program outcomes | Not established | No peer-reviewed study located |
| Preserving the unrevised original forecast is the binding constraint, not the arithmetic | Named field experience | Capital program practice, Mission Intelligence Systems |
What to do with this
Take the last twenty bids received and find the cost range that was modelled for each package before the bid opened. Count how many landed inside the range. If the ranges were stated at eighty percent, about sixteen should have.
Most organizations that run this find a number closer to nine or ten, and the finding is specific, defensible and immediately actionable in a way that no maturity assessment has ever been. It also takes about a day, using files that already exist.
Then make one process change: lock the forecast at the moment it is issued and keep the locked version. Everything above depends on it, and nothing above is possible without it.
References
- Hoyt, Robert E., and Andre P. Liebenberg. “The Value of Enterprise Risk Management.” Journal of Risk and Insurance, vol. 78, no. 4, 2011, pp. 795–822. doi.org/10.1111/j.1539-6975.2011.01413.x.
- Pagach, Donald, and Richard Warr. “The Characteristics of Firms That Hire Chief Risk Officers.” Journal of Risk and Insurance, vol. 78, no. 1, 2011, pp. 185–211. doi.org/10.1111/j.1539-6975.2010.01378.x.
- Gatzert, Nadine, and Michael Martin. “Determinants and Value of Enterprise Risk Management: Empirical Evidence from the Literature.” Risk Management and Insurance Review, vol. 18, no. 1, 2015, pp. 29–53. doi.org/10.1111/rmir.12028.
- Brier, Glenn W. “Verification of Forecasts Expressed in Terms of Probability.” Monthly Weather Review, vol. 78, no. 1, 1950, pp. 1–3. doi.org. The original proper scoring rule.
- Lichtenstein, Sarah, Baruch Fischhoff, and Lawrence D. Phillips. “Calibration of Probabilities: The State of the Art to 1980.” In Daniel Kahneman, Paul Slovic and Amos Tversky, eds., Judgment Under Uncertainty: Heuristics and Biases. Cambridge University Press, 1982, pp. 306–334.
- Mellers, Barbara, et al. “Psychological Strategies for Winning a Geopolitical Forecasting Tournament.” Psychological Science, vol. 25, no. 5, 2014, pp. 1106–1115. doi.org/10.1177/0956797614524255.
- U.S. Government Accountability Office. Cost Estimating and Assessment Guide: Best Practices for Developing and Managing Program Costs. GAO-20-195G, March 2020. gao.gov/products/gao-20-195g. Estimate documentation as a precondition for later comparison.
- Committee of Sponsoring Organizations of the Treadway Commission. Enterprise Risk Management: Integrating with Strategy and Performance. COSO, 2017. coso.org/guidance-erm.
- International Organization for Standardization. Risk Management: Guidelines. ISO 31000:2018, clause 6.6 on monitoring and review. iso.org/standard/65694.html.
About the Author
Dan Flynn
Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build
Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.
His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.
