Risk and uncertainty
Forecast Calibration Scorer
If your ranges are stated at eighty percent confidence, roughly eighty percent of outcomes should land inside them. Almost nobody checks. This checks, and tells you whether the fault is where the range sits or how wide it is.
Why this is the measure that works
Counting activity does not test risk management, and the absence of disasters is not evidence of anything. Calibration is a real test: every bid, every completed activity and every resolved risk checks a range or a probability that was already stated. Scored at the input level it yields hundreds of observations a year rather than one a decade.
Two faults, two corrections
Coverage below your stated confidence has two possible causes, and they need opposite fixes. If the misses land on one side, the ranges are in the wrong place and widening them will not help. If they land on both, the ranges are too narrow. Correcting one while the other stands leaves an organization believing it has solved the problem.
This scores ranges you already stated.
It cannot check that you stated them before the outcome was known, and if you did not, every number below is decoration. Preserving the original forecast unrevised is the hard part of calibration, and it is the part no scorer can do for you.
| Item | Stated low | Stated high | Actual outcome | Result |
|---|---|---|---|---|
| - | ||||
| - | ||||
| - | ||||
| - | ||||
| - | ||||
| - | ||||
| - | ||||
| - |
Fill at least one row with a low, a high and an actual outcome, and the scoring appears here. Nothing is sent anywhere: this runs entirely in your browser.
The part this cannot do
Calibration only means something if the forecast was written down before the outcome was known and never quietly edited afterwards. That is a records problem, not a scoring problem, and it is why most organizations cannot measure this at all: the original numbers are gone. If you want the reasoning behind these measures, the worked example this tool reproduces is in the article below.
