Mission Intelligence Systems

Risk and uncertainty

Forecast Calibration Scorer

If your ranges are stated at eighty percent confidence, roughly eighty percent of outcomes should land inside them. Almost nobody checks. This checks, and tells you whether the fault is where the range sits or how wide it is.

Observed coverageCentering vs widthBoth correctionsStays in your browser

Why this is the measure that works

Counting activity does not test risk management, and the absence of disasters is not evidence of anything. Calibration is a real test: every bid, every completed activity and every resolved risk checks a range or a probability that was already stated. Scored at the input level it yields hundreds of observations a year rather than one a decade.

Two faults, two corrections

Coverage below your stated confidence has two possible causes, and they need opposite fixes. If the misses land on one side, the ranges are in the wrong place and widening them will not help. If they land on both, the ranges are too narrow. Correcting one while the other stands leaves an organization believing it has solved the problem.

This scores ranges you already stated.

It cannot check that you stated them before the outcome was known, and if you did not, every number below is decoration. Preserving the original forecast unrevised is the hard part of calibration, and it is the part no scorer can do for you.

Ranges you stated with a confidence level, scored on how often the outcome landed inside them.

ItemStated lowStated highActual outcomeResult
-
-
-
-
-
-
-
-

Fill at least one row with a low, a high and an actual outcome, and the scoring appears here. Nothing is sent anywhere: this runs entirely in your browser.

Before you act on this

What this can and cannot tell you

Read this before acting on anything above. It is here to make the result more useful, not to hedge it.

  1. 1 of 4

    A reading is one vantage point, not the organization

    Every score here comes from what people could observe and were willing to report, from where they sit. A senior view stops where the information stops reaching it, and a team view is bounded by what the team has been told. That is a real signal about the organization, and it is not the same thing as ground truth. Treat a result as a claim to check against records, calendars, and the people closest to the work, not as a finding that has already been checked.

  2. 2 of 4

    A recommendation is a hypothesis, not an instruction

    What this produces is the most defensible next question given the pattern in the answers. It has no access to your funding cycle, your contracts, your regulator, your technology estate, or the person who is about to resign, and any one of those can be the real constraint while the instrument points somewhere else. The recommendation is worth acting on when your own judgment, and the evidence you can gather, agree with it. It is worth arguing with when they do not, and the disagreement is more useful than the score.

  3. 3 of 4

    Outcomes are decided in execution, by people and by systems

    Reading a diagnostic changes nothing. What changes an organization is a specific person with authority making a specific decision, and then the follow-through surviving contact with the work: competing priorities, staffing, incentives, contracts, data quality, system limits, vendors, and everything else outside these questions. Two organizations with identical results can end a year in opposite places, and the difference is what they did and what happened to them, not what they scored.

  4. 4 of 4

    Conditions move, so a result has a shelf life

    These readings describe a moment. A reorganization, a departure, a new system, or a change in demand can move a condition faster than any plan built on the old reading. Measure again rather than assuming a result still holds, and treat a number that has not moved as a question about whether anything actually changed.

The part this cannot do

Calibration only means something if the forecast was written down before the outcome was known and never quietly edited afterwards. That is a records problem, not a scoring problem, and it is why most organizations cannot measure this at all: the original numbers are gone. If you want the reasoning behind these measures, the worked example this tool reproduces is in the article below.

The other half of this

These numbers say whether the estimates held up. They say nothing about whether anything institutional produced them, and the two answers together are more useful than either. The Risk Capability Diagnostic scores sixteen practices, five of which make claims about numbers, and its result page reads your rows from this browser and sets the two side by side. Mature practice whose forecasts still miss is the pairing worth finding.