Mission Intelligence Systems
Risk · Alignment

Why the Simulation Disagrees With Your Critical Path Date

Later is expected. Earlier is a symptom. They look identical on the report.

The first time a team sees a simulated finish date that does not match the schedule's own critical path date, the instinct is to distrust the simulation. Sometimes that instinct is right. More often the simulation is telling you something true that the deterministic date was never able to express.

Published

Key Takeaways

Research foundation

This is one of the best-established results in project scheduling and one of the least known outside it. Clark (1961) supplied the mathematics of the greatest of a finite set of random variables. Fulkerson (1962) and MacCrimmon and Ryavec (1964) established the optimistic bias of PERT-style estimates, and Klingel (1966) documented it on a real network. Elmaghraby (2005) restates the point plainly as the fallacy of averages. Williams (1992) and Dodin and Elmaghraby (1985) cover why the criticality index diverges from the nominal critical path. The Four A's are the executive lens applied to that evidence.

The disagreement is not a software artifact and it is not new. It was worked out in the operations research literature more than sixty years ago, and the reason it keeps surprising people is that the result is genuinely counterintuitive.

Why does the deterministic date come out too early?

Because of what happens at merge points, and it follows from a piece of mathematics that is hard to argue with.

Consider a milestone that cannot start until four parallel workstreams have all finished. Its start is the maximum of four uncertain durations. Clark established the distribution of the greatest of a finite set of random variables in 1961, and the consequence for schedules is direct: the expected value of a maximum is greater than the maximum of the expected values. A deterministic schedule computes the second quantity. Reality produces the first.

So the deterministic date is systematically optimistic wherever paths converge, and the effect intensifies as more paths of similar length merge at the same point. Fulkerson showed in 1962 that substituting expected activity times gives a bound that is beatable, meaning the naive computation understates duration. MacCrimmon and Ryavec quantified PERT's errors two years later, including the bias introduced by ignoring parallel non-critical paths. Klingel then demonstrated it on an actual project network in 1966.

Elmaghraby put the general form of the error most memorably, calling it the fallacy of averages in project risk management: plugging single-point averages into a network and reading the answer as an average is not how the arithmetic works.

The practical consequence for an executive is one sentence. The critical path date is not a conservative number and never was. It is the answer to a hypothetical in which every activity takes exactly its assigned duration, and the probability of that hypothetical is usually far below fifty percent.

So what does it mean if the simulation comes out earlier?

It means something is wrong with the inputs, and it is worth chasing rather than celebrating. Merge bias pushes the simulated date later, so an earlier result has to be explained by something overriding it. Three causes account for nearly all of it.

The baseline durations are padded

If the schedule's single durations already contain private buffer, and the three-point estimates are then built honestly around a lower most-likely value, the distribution means sit below the baseline durations. The simulation is not being optimistic; it is running an unpadded model against a padded schedule and reporting the difference. This is precisely why FTA requires a stripped and adjusted base schedule before its risk model runs.

Constraints are holding the deterministic date out

Hard constraints and artificial dates pin the deterministic computation while the simulation, depending on configuration, may be free to move around them. The deterministic date is then a statement about the constraints, not about the logic.

The distributions were assigned around the wrong anchor

If the person assigning three-point estimates treated the existing duration as the pessimistic case rather than the most likely, every distribution sits below the baseline by construction. This is common when the estimating conversation is rushed, and it is invisible in the output.

In all three cases the honest reading is the same: the simulation has surfaced something that was already in the schedule and not previously visible. That is a finding, not a fault, and it belongs in the conversation described in Is This Schedule Good Enough to Run a Risk Analysis On.

Why does the criticality index not match the critical path?

Because they answer different questions, and the divergence is informative.

The deterministic critical path is the longest path under one set of durations. The criticality index is the proportion of iterations in which an activity actually sat on the critical path. A shorter path with high variance can be critical more often than the nominal critical path, and that is exactly the path a manager should be watching, because it is the one that will surprise them.

Williams treated criticality in stochastic networks directly, and Dodin and Elmaghraby had earlier worked on approximating criticality indices in PERT networks. The result worth carrying into a review is that a project usually has more than one candidate for what will drive the finish, and the deterministic view shows only one of them.

Why is this an Alignment problem?

Because two numbers are circulating, both correct, describing different things, and nobody has said which one governs.

The schedule publishes a date. The risk analysis publishes a distribution. If the organization has not agreed which is the commitment, which is the working target, and what the gap between them is for, then the two numbers will be used selectively: the earlier one in external communication, the later one in internal caveats. That is not dishonesty, it is the predictable result of leaving the relationship undefined, and it is the same failure as a risk appetite that is published but not operative, described in Risk Appetite Is Not Alignment.

The fix is a sentence in the governance document rather than a better model. Name the deterministic date as a computation, name the percentile as the commitment, and name the gap as contingency with an owner. Then the disagreement between them stops being an argument and becomes an instrument.

Evidence matrix

ClaimEvidence tierSource
The expected maximum exceeds the maximum of expectationsPeer reviewed, foundationalClark (1961), Operations Research 9(2)
PERT-style deterministic estimates are optimistically biasedPeer reviewedFulkerson (1962), OR 10(6); MacCrimmon & Ryavec (1964), OR 12(1)
The bias is measurable on a real project networkPeer reviewedKlingel (1966), Management Science 13(4)
Single-point averages in a network do not produce an average outcomePeer reviewedElmaghraby (2005), EJOR 165(2)
Criticality index legitimately diverges from the nominal critical pathPeer reviewedWilliams (1992), JORS 43(4); Dodin & Elmaghraby (1985), Management Science 31(2)
Two uncontested numbers get used selectivelyFour A's interpretationBuilders Build, Alignment

What to do with this

When the two dates disagree, first establish the direction. Later than deterministic is expected and the conversation is about how much contingency the gap implies. Earlier than deterministic is a model or baseline problem, and the first three things to check are padding in the durations, constraints in the network, and which anchor the three-point estimates were built around.

References

  1. Clark, Charles E. “The Greatest of a Finite Set of Random Variables.” Operations Research, vol. 9, no. 2, 1961, pp. 145–162. doi.org/10.1287/opre.9.2.145. The mathematical basis of merge event bias.
  2. Fulkerson, D. R. “Expected Critical Path Lengths in PERT Networks.” Operations Research, vol. 10, no. 6, 1962, pp. 808–817. doi.org/10.1287/opre.10.6.808.
  3. MacCrimmon, Kenneth R., and Charles A. Ryavec. “An Analytical Study of the PERT Assumptions.” Operations Research, vol. 12, no. 1, 1964, pp. 16–37. doi.org/10.1287/opre.12.1.16.
  4. Klingel, A. R., Jr. “Bias in Pert Project Completion Time Calculations for a Real Network.” Management Science, vol. 13, no. 4, 1966, pp. B-194–B-201. doi.org/10.1287/mnsc.13.4.B194.
  5. Elmaghraby, Salah E. “On the Fallacy of Averages in Project Risk Management.” European Journal of Operational Research, vol. 165, no. 2, 2005, pp. 307–313. doi.org/10.1016/j.ejor.2004.04.003.
  6. Williams, Terry M. “Criticality in Stochastic Networks.” Journal of the Operational Research Society, vol. 43, no. 4, 1992, pp. 353–357. doi.org/10.1057/jors.1992.50.
  7. Dodin, Bajis M., and Salah E. Elmaghraby. “Approximating the Criticality Indices of the Activities in PERT Networks.” Management Science, vol. 31, no. 2, 1985, pp. 207–223. doi.org/10.1287/mnsc.31.2.207.
  8. Elmaghraby, Salah E. “On Criticality and Sensitivity in Activity Networks.” European Journal of Operational Research, vol. 127, no. 2, 2000, pp. 220–238. doi.org/10.1016/S0377-2217(99)00483-X.
  9. Federal Transit Administration. Oversight Procedure 40: Risk and Contingency Review. U.S. Department of Transportation, October 2023. transit.dot.gov. Requires a stripped and adjusted base schedule, which is the control for padded durations.
  10. U.S. Government Accountability Office. GAO Schedule Assessment Guide. GAO-16-89G, December 2015. gao.gov/products/gao-16-89g. Confirming a valid critical path is one of the ten practices.
DF

About the Author

Dan Flynn

Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build

Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.

His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.