Mission Intelligence Systems
Risk · Adaptability

How Do We Know Our Risk Management Is Working?

Nothing went wrong is not an answer. Neither is the register count.

It is the fairest question an executive can ask about a function that costs money and produces documents, and it is answered almost everywhere with metrics that cannot fail. There is a measure that genuinely tests the work. It is cheap, it uses records the organization already keeps, and it requires one thing that turns out to be the real obstacle.

Published

Key Takeaways

Research foundation

The value evidence is Hoyt and Liebenberg (2011) and Pagach and Warr (2011) in the Journal of Risk and Insurance, with Gatzert and Martin (2015) in Risk Management and Insurance Review for the state of the literature. Scoring rests on Brier (1950) and the calibration literature consolidated by Lichtenstein, Fischhoff and Phillips (1982), with Mellers and colleagues (2014) on trainability. One limit is stated: no peer-reviewed study establishes that calibration scoring of program risk inputs improves program outcomes.

The reporting pack answers the question with three numbers. There are 147 risks in the register, up from 138. Ninety-four percent of scheduled risk reviews were held. Eighty-one percent of mitigation actions are on track.

Now consider what would have to happen for any of those to get worse. The register count falls when risks are closed, so a rising count is a sign of diligence and a falling one is a sign of diligence. Review compliance measures whether meetings occurred. Mitigation completion rises when the team selects mitigations it can finish. Every one of these can improve steadily while the organization walks into the thing that damages it.

A metric that cannot get worse when the situation gets worse is not a measure. It is a reassurance mechanism with a number attached.

Why outcome measurement does not rescue this

The obvious correction is to measure results instead of effort, and it does not work either, for reasons worth understanding before anyone builds a scorecard on it.

Rare events are absent most of the time whatever the organization does. An entity with no risk function will also report no catastrophic loss in the majority of years, so the observation fails to distinguish between competence and luck. This is the base rate problem and it is fatal to the naive version of outcome measurement.

Then there is the counting problem. A major capital program completes every few years. That is not a sample, and each of those few outcomes is driven by procurement conditions, labor markets, political change, weather and design quality alongside anything the risk process did. Attributing the result to the risk function requires isolating a small effect inside enormous noise with a handful of observations.

And there is the lag. By the time a program lands, the people who built the model have moved, the method has been revised twice, and the feedback arrives too late to change behavior. Feedback that arrives after the behavior it should correct has already been repeated is not feedback.

What the research establishes, and what it does not

This deserves a careful reading because the summaries in circulation are more confident than the papers.

Hoyt and Liebenberg examined US insurers and found a positive and statistically significant relation between enterprise risk management and firm value. That is a real finding in a peer-reviewed journal and it is the strongest single result usually cited.

Pagach and Warr examined firms that appointed chief risk officers and found the adopters were systematically different before adoption: larger, more leveraged, with more volatile earnings. On the benefits side their results offered limited support for the outcomes the framework is expected to produce. That is a materially more cautious picture than the first result alone conveys.

Gatzert and Martin reviewed the accumulated literature and concluded that findings are generally positive but depend substantially on how adoption is measured.

That last point is the one to carry away. Nearly all of this research proxies risk management by the presence of a chief risk officer, or by an external ERM rating, or by a disclosure. Those are measures of whether the organization adopted the structure. None of them measures whether the work is done well. So the literature can tell you something about the average effect of having a function, and almost nothing about whether your function is any good, which was the question.

The measure that actually tests the work

Risk management makes claims about uncertainty. Claims about uncertainty are testable. That is the whole idea, and it is unfamiliar mainly because the claims are rarely recorded in a form that permits scoring.

Interval hit rate

If a range was stated at eighty percent confidence, then across many such statements roughly eighty percent of actual outcomes should fall inside the range. Fewer means the ranges are too narrow, which is the direction the calibration literature predicts and which Lichtenstein, Fischhoff and Phillips documented as the dominant and highly replicated finding. More means the ranges are too wide and the analysis is not discriminating.

This is computable from records the organization already has, and it needs no new process. It needs the original range preserved.

A proper score on discrete risks

Where risks carried explicit probabilities, those probabilities can be scored against what happened using a proper scoring rule, the oldest and simplest of which is Brier's. A proper rule has the property that a forecaster minimizes their expected score by stating what they actually believe, so it cannot be improved by hedging toward the middle. Score by estimator and the organization learns which of its people give useful probabilities, which is information no maturity assessment will ever produce.

Score the inputs, not the programs

Here is the move that resolves the small-sample problem, and it is the most useful idea on this page.

The organization does not have to wait for programs to complete. Every bid that arrives is a test of a modeled cost range. Every activity that finishes is a test of a duration range. Every risk that occurs or is retired is a test of a stated probability. A single mid-sized program generates hundreds of these tests a year, and they are all currently discarded.

Score at that level and the feedback loop closes in months. You learn, concretely, that the organization's ranges on utility work are consistently forty percent too narrow while its ranges on structural work are well calibrated, and that is an actionable finding about a specific estimating practice rather than a verdict on a department.

The obstacle is not analytical and not financial. It is that scoring requires the original forecast to survive unrevised, and organizations habitually update the record to reflect what they now know. A range that was quietly widened after the bid came in cannot be scored, and the widening usually feels like diligence at the time.

Worked example: scoring twenty bids

The following is a constructed illustration rather than client data, built to show the arithmetic and the shape of the typical result. Twenty procurement packages, each carrying a modeled range stated as an eighty percent interval before the bids opened, against the award actually received. Figures in thousands of dollars.

PackageModeled 80% rangeAwardResult
Earthworks, north1,180 to 1,4201,365Inside
Piling640 to 760812Above
Drainage package A930 to 1,1201,058Inside
Utility relocation1,450 to 1,7402,090Above
Structural steel3,100 to 3,7203,610Inside
Precast units860 to 1,0301,004Inside
Facade2,240 to 2,6903,155Above
MEP fit-out4,100 to 4,9205,480Above
Fire systems720 to 865858Inside
Vertical transport1,310 to 1,5701,522Inside
Roofing980 to 1,1751,310Above
Interior finishes2,650 to 3,1803,096Inside
Signage and wayfinding310 to 375402Above
Landscape540 to 650631Inside
Security systems690 to 830905Above
Commissioning support450 to 540528Inside
Temporary works820 to 9851,142Above
Drainage package B760 to 915889Inside
Substation works1,620 to 1,9452,280Above
Testing and handover380 to 455447Inside

Eleven of twenty landed inside the range. On an eighty percent interval, sixteen should have. That is the headline finding and it took one spreadsheet column.

The second column is the more useful one. Of the nine misses, all nine came in above the range and none below. If the intervals were merely too narrow but correctly centered, misses should split roughly evenly. The chance of nine out of nine falling on the same side by accident is about one in two hundred and fifty. This organization does not have one problem, it has two, and they need fixing in the right order.

Separate the bias from the spread

A range can fail in two independent ways. It can be too narrow, which is the overconfidence finding, and it can be in the wrong place, which is optimism. Diagnosing narrowness when the real fault is location leads to the worst available outcome: enormous ranges still centered on a number that is too low.

Check location first. The median award here is 1.077 times the midpoint of its stated range, and the median award sits at position 0.92 within its own interval, meaning a typical bid lands near the very top of the range that was supposed to contain it comfortably. That is a location problem, and it says the estimating is anchored on the intended outcome rather than the likely one.

Then check spread, and only after re-centering. Applying a uniform 7.7 percent uplift and keeping the relative width unchanged lifts coverage from eleven to thirteen of twenty. Better, and still short of sixteen. So the intervals were both mis-centered and too narrow, and correcting one of the two would have left the organization believing it had solved the problem.

For the remaining narrowness the arithmetic is standard. Treating the underlying as roughly normal, a genuine eighty percent central interval spans plus or minus 1.28 standard deviations while a fifty five percent interval spans plus or minus 0.76. The implied widening is 1.28 divided by 0.76, about 1.7. Applied after the uplift, not instead of it.

Two cautions on the method, both real. Twenty observations is enough to detect a fault this large and nowhere near enough to calibrate finely, since sampling variation on a coverage estimate from twenty trials is wide. And the normal approximation used for the widening factor is a convenience: cost outcomes are right-skewed, so it will understate the correction needed in the upper tail. Treat the first run as a direction, keep scoring, and revisit once there are a hundred observations.

None of that weakens the point. An organization that has never run this has no idea whether its stated confidence levels mean anything, and one afternoon with twenty existing files converts that from an open question into a number with a correction attached.

Time from signal to decision

One non-statistical measure worth adding, because it captures something the others miss. For each significant risk that materialized, how long elapsed between the first point at which the organization possessed the information and the point at which someone made a decision about it.

This is measurable retrospectively from documents, it is uncomfortable in a productive way, and it tests the conditions rather than the analysis. A model can be excellent while nothing happens because of it, and this is the measure that catches that.

What to stop reporting

An honest limitation before this becomes a scorecard. No peer-reviewed study establishes that calibration scoring of program risk inputs improves program outcomes. The case for it rests on the demonstrated existence of miscalibration, on the demonstrated trainability of calibration in forecasting research, and on the ordinary principle that a claim which cannot be checked will not improve. That is a reasoned argument rather than a validated intervention, and it should be presented as one.

Evidence matrix

ClaimEvidence tierSource
Enterprise risk management is positively related to firm value among US insurersPeer reviewedHoyt & Liebenberg (2011), JRI 78(4)
Firms hiring chief risk officers differ systematically beforehand; benefit evidence is limitedPeer reviewedPagach & Warr (2011), JRI 78(1)
Findings across the literature depend substantially on how adoption is proxiedPeer reviewed reviewGatzert & Martin (2015), RMIR 18(1)
Stated ranges are systematically too narrowPeer reviewed, foundationalLichtenstein, Fischhoff & Phillips (1982)
A proper scoring rule cannot be improved by stating other than your true beliefPeer reviewed, foundationalBrier (1950), Monthly Weather Review 78(1)
Calibration improves with trainingPeer reviewedMellers et al. (2014), Psychological Science 25(5)
Calibration scoring of program risk inputs improves program outcomesNot establishedNo peer-reviewed study located
Preserving the unrevised original forecast is the binding constraint, not the arithmeticNamed field experienceCapital program practice, Mission Intelligence Systems

What to do with this

Take the last twenty bids received and find the cost range that was modeled for each package before the bid opened. Count how many landed inside the range. If the ranges were stated at eighty percent, about sixteen should have.

Most organizations that run this find a number closer to nine or ten, and the finding is specific, defensible and immediately actionable in a way that no maturity assessment has ever been. It also takes about a day, using files that already exist.

Then make one process change: lock the forecast at the moment it is issued and keep the locked version. Everything above depends on it, and nothing above is possible without it.

Working tool

A free spreadsheet that does this, with live formulas and a filled example row: Forecast Calibration Scorer. It scores the coverage, separates the centering fault from the width fault, and returns both corrections. It reproduces the twenty row example above exactly. No sign up. The rest are on the risk tools page.

References

  1. Hoyt, Robert E., and Andre P. Liebenberg. “The Value of Enterprise Risk Management.” Journal of Risk and Insurance, vol. 78, no. 4, 2011, pp. 795–822. doi.org/10.1111/j.1539-6975.2011.01413.x.
  2. Pagach, Donald, and Richard Warr. “The Characteristics of Firms That Hire Chief Risk Officers.” Journal of Risk and Insurance, vol. 78, no. 1, 2011, pp. 185–211. doi.org/10.1111/j.1539-6975.2010.01378.x.
  3. Gatzert, Nadine, and Michael Martin. “Determinants and Value of Enterprise Risk Management: Empirical Evidence from the Literature.” Risk Management and Insurance Review, vol. 18, no. 1, 2015, pp. 29–53. doi.org/10.1111/rmir.12028.
  4. Brier, Glenn W. “Verification of Forecasts Expressed in Terms of Probability.” Monthly Weather Review, vol. 78, no. 1, 1950, pp. 1–3. doi.org. The original proper scoring rule.
  5. Lichtenstein, Sarah, Baruch Fischhoff, and Lawrence D. Phillips. “Calibration of Probabilities: The State of the Art to 1980.” In Daniel Kahneman, Paul Slovic and Amos Tversky, eds., Judgment Under Uncertainty: Heuristics and Biases. Cambridge University Press, 1982, pp. 306–334.
  6. Mellers, Barbara, et al. “Psychological Strategies for Winning a Geopolitical Forecasting Tournament.” Psychological Science, vol. 25, no. 5, 2014, pp. 1106–1115. doi.org/10.1177/0956797614524255.
  7. U.S. Government Accountability Office. Cost Estimating and Assessment Guide: Best Practices for Developing and Managing Program Costs. GAO-20-195G, March 2020. gao.gov/products/gao-20-195g. Estimate documentation as a precondition for later comparison.
  8. Committee of Sponsoring Organizations of the Treadway Commission. Enterprise Risk Management: Integrating with Strategy and Performance. COSO, 2017. coso.org/guidance-erm.
  9. International Organization for Standardization. Risk Management: Guidelines. ISO 31000:2018, clause 6.6 on monitoring and review. iso.org/standard/65694.html.
DF

About the Author

Dan Flynn

Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build

Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.

His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.