Are Risk Matrices Valid? What Cox Actually Proved
Not useless. Not sufficient. The distinction is mathematical, not rhetorical.
There is a published mathematical critique of the risk matrix that most practitioners have heard of and few have read. It is more specific than the slogans built on top of it, and it is more useful, because it tells you exactly which decisions a matrix can carry and which it cannot.
Published
Key Takeaways
- The critique is narrow and real. A matrix has limited resolution, can invert the true ranking of two risks under specifiable conditions, and produces ordinal labels that cannot legitimately be added or averaged. It does not follow that matrices are worthless.
- Enlarging the grid does not fix it. More cells add apparent precision without adding information, because the inputs are still categories assigned by human judgment. The published improvements concern axis scaling and coloring rules, not grid size.
- The practical rule is a division of labor: use the matrix to triage a long register into rough tiers, and quantify the small number of items that actually drive cost and schedule. The error is not using a matrix. It is asking a matrix to produce a number.
Research foundation
The primary source is Cox (2008) in Risk Analysis, which establishes the limits of risk matrices formally. It has been extended to industry practice by Thomas, Bratvold and Bickel (2014) and paralleled by Hubbard and Evans (2010) on ordinal scales. It has also been answered: Ball and Watt (2013) defend the matrix as a screening device in the same journal, and Duijm (2015) sets out how to design matrices better rather than discard them. This article reports that discussion rather than resolving it, and the Four A's are the executive lens applied to it.
Ask a risk professional whether the five by five matrix is sound and you will usually get one of two answers, both unhelpful. Either it is the standard and therefore fine, or it is mathematically discredited and anyone still using one is not serious. The first ignores a body of published work. The second overstates what that work concluded.
What is a risk matrix?
A risk matrix is a grid that places each risk in a cell according to two ordered categories, usually a likelihood band and a consequence band, and assigns a rating or color to that cell. Three by three, four by four and five by five are the common sizes. The output is a tier, typically red, amber or green, or a score formed by multiplying the two category numbers.
That last step is where the trouble starts, and it is worth being precise about why. The rows and columns are ordinal: they tell you that band four is worse than band three, but not by how much. Multiplying two ordinal positions produces a number that looks like a measurement and is not one.
What did Cox actually prove?
Cox's 2008 paper in Risk Analysis is a formal treatment rather than an opinion piece, and it establishes three distinct limitations.
The first is resolution. Because a matrix compresses a continuous range into a handful of bands, two risks whose quantitative expected loss differs by orders of magnitude can land in the same cell and receive an identical rating. The matrix is not wrong about them. It simply cannot see the difference.
The second, and the one that surprises people, is ranking reversal. Cox shows that under conditions he specifies formally, a matrix can assign a higher qualitative rating to a risk that is quantitatively smaller. The condition that most often produces this in practice is negative correlation between frequency and severity, which is extremely common: the events that would hurt most are usually the ones least likely to happen. When that relationship holds, the banding can invert the true ordering, and the organization dutifully works on the wrong risk first.
The third is non-additivity. Ordinal ratings cannot be summed or averaged in any way that preserves meaning, which means a matrix cannot legitimately aggregate to a portfolio view or support an allocation of budget. This is the limitation with the largest practical consequence, because aggregating and allocating is precisely what most organizations do with their matrix output.
Hubbard and Evans made a closely related argument about ordinal scales in risk assessment two years later in the IBM Journal of Research and Development, and Thomas, Bratvold and Bickel carried the argument into petroleum industry practice in 2014, working through how matrix-driven decisions diverge from decision-analytic ones on real cases.
When does a risk matrix actively mislead?
Not always, and knowing the difference is more useful than a blanket position. A matrix is most likely to mislead in four situations.
When the risks being compared have negatively correlated likelihood and consequence, which is the ranking reversal case. When the ratings are multiplied into a score and the scores are then summed, ranked across a portfolio, or used to size a reserve. When the bands are wide enough that the items inside a single cell differ materially, which is a resolution failure and gets worse as the program gets larger. And when different people populated the matrix without a shared definition of what the bands mean, which is a problem the matrix inherits rather than causes, and which is treated at length in Why 'Likely' Is Not a Probability.
Conversely, a matrix used to sort forty candidate risks into rough tiers so that a team knows where to spend its next two weeks of attention is doing something it is entirely adequate for. That is triage, and triage does not require arithmetic.
Does making the matrix bigger help?
Very little, and it can hurt. Moving from three by three to five by five multiplies the number of cells without changing the nature of the inputs, so it produces finer-looking output from the same coarse judgments. The visible effect is more confidence rather than more information.
The published work on improvement points somewhere else entirely. Levine argued in the Journal of Risk Research for logarithmically scaled axes, so that each step up an axis represents a consistent multiple rather than an arbitrary widening, which addresses part of the resolution problem directly. Duijm's review in Safety Science is the most practical single source: it works through axis definition, the rules by which cells are colored, and the case for a continuous probability and consequence diagram in place of a discrete grid.
If you are debating grid dimensions, the more productive question is what happens to the output. If anything downstream of the matrix needs a number, the size of the grid will not save you.
Why do organizations keep using them?
Because they solve a real problem cheaply, and because the alternative has a cost that the critique tends to skip over.
A matrix lets a group of people with different backgrounds reach a shared rough view of a long list in an afternoon. It requires no software, no distributions, and no specialist. Ball and Watt made essentially this case in their response to Cox, arguing that the matrix should be judged against the practical alternatives available to a working practitioner rather than against an ideal. That is a fair point and it deserves to be represented honestly.
There is also a less flattering reason, and it is worth naming. A red, amber and green grid produces a defensible-looking artifact quickly. It survives a gate review. Organizations that are optimizing for the appearance of risk management rather than for changed decisions will always prefer an instrument that produces a clean deliverable, which is the same dynamic that turns a register into paperwork.
What should you do instead, given real constraints?
Divide the work according to what each instrument can support.
Keep the matrix for screening, and be explicit on the page that its output is a triage tier and not a measurement. Define each band numerically, so that a likelihood band names a probability range and a consequence band names a cost or duration range, which removes the worst of the shared-meaning problem and makes the resolution limits visible. Stop multiplying the bands to make a score, or if the score is politically unavoidable, never sum or rank across it.
Then take the small number of risks that plausibly drive the outcome, usually far fewer than the register contains, and express them as ranges rather than categories. Combine those ranges and you have a distribution, and from a distribution you can read a defensible contingency figure. That is the move described in What P80 Means, and it is what the matrix was never able to give you.
Why is this an Alignment problem?
Because the matrix is a shared instrument in appearance and frequently not one in fact. Two people can place risks in the same grid, using the same words, and mean different things by every band. The format hides that disagreement rather than surfacing it, which is worse than an obvious inconsistency, because the output looks like consensus.
The deeper alignment failure is between what the instrument produces and what the organization asks of it. A matrix answers the question of rough priority order. It is routinely asked to answer questions about money. Nothing in the mathematics of the grid can close that gap, and no amount of additional cells will either. ISO 31000:2018 is clear that risk management should be based on the best available information and should be dynamic; a matrix built once at inception, multiplied into a score, and never revisited fails both tests before the question of validity even arises.
Evidence matrix
| Claim | Evidence tier | Source |
|---|---|---|
| Matrices have limited resolution and can invert true risk rankings | Peer reviewed | Cox (2008), Risk Analysis 28(2) |
| Ordinal risk scores cannot be legitimately aggregated | Peer reviewed | Hubbard & Evans (2010), IBM Journal of Research and Development 54(3) |
| The critique holds in applied industry decision making | Peer reviewed | Thomas, Bratvold & Bickel (2014), SPE Economics & Management 6(2) |
| Abandonment is contested; matrices remain defensible as screening tools | Peer reviewed, counterposition | Ball & Watt (2013), Risk Analysis 33(11) |
| Improvements lie in axis scaling and design, not grid size | Peer reviewed | Duijm (2015), Safety Science 76; Levine (2012), Journal of Risk Research 15(2) |
| The format conceals disagreement about what the bands mean | Four A's interpretation | Builders Build, Alignment |
The summary a board can act on is short. The matrix is a reasonable way to sort a long list and an unreasonable way to produce a number, and most organizations are using it for both.
References
- Cox, Louis Anthony (Tony), Jr. “What's Wrong with Risk Matrices?” Risk Analysis, vol. 28, no. 2, 2008, pp. 497–512. doi.org/10.1111/j.1539-6924.2008.01030.x. The primary formal treatment of resolution, ranking reversal, and the limits of ordinal risk categories.
- Hubbard, Douglas W., and Dylan Evans. “Problems with Scoring Methods and Ordinal Scales in Risk Assessment.” IBM Journal of Research and Development, vol. 54, no. 3, 2010, pp. 2:1–2:10. doi.org/10.1147/JRD.2010.2042914.
- Thomas, Philip, Reidar B. Bratvold, and J. Eric Bickel. “The Risk of Using Risk Matrices.” SPE Economics & Management, vol. 6, no. 2, 2014, pp. 56–66. doi.org/10.2118/166269-PA.
- Ball, David J., and John Watt. “Further Thoughts on the Utility of Risk Matrices.” Risk Analysis, vol. 33, no. 11, 2013, pp. 2068–2078. doi.org/10.1111/risa.12057. A direct response defending the matrix as a practical screening device.
- Duijm, Nijs Jan. “Recommendations on the Use and Design of Risk Matrices.” Safety Science, vol. 76, 2015, pp. 21–31. doi.org/10.1016/j.ssci.2015.02.014. The most practical single source on axis definition, coloring rules, and continuous alternatives.
- Levine, E. S. “Improving Risk Matrices: The Advantages of Logarithmically Scaled Axes.” Journal of Risk Research, vol. 15, no. 2, 2012, pp. 209–222. doi.org/10.1080/13669877.2011.634514.
- Beyth-Marom, Ruth. “How Probable Is Probable? A Numerical Translation of Verbal Probability Expressions.” Journal of Forecasting, vol. 1, no. 3, 1982, pp. 257–269. doi.org/10.1002/for.3980010305.
- Budescu, David V., Stephen Broomell, and Han-Hui Por. “Improving Communication of Uncertainty in the Reports of the Intergovernmental Panel on Climate Change.” Psychological Science, vol. 20, no. 3, 2009, pp. 299–308. doi.org/10.1111/j.1467-9280.2009.02284.x.
- Kent, Sherman. “Words of Estimative Probability.” Studies in Intelligence, vol. 8, no. 4, 1964, pp. 49–65. cia.gov.
- International Organization for Standardization. Risk Management: Guidelines. ISO 31000:2018. iso.org/standard/65694.html. Clause 4 establishes that risk management should be based on the best available information and should be dynamic.
About the Author
Dan Flynn
Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build
Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.
His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.
