Is This Schedule Good Enough to Run a Risk Analysis On?
A simulation does not repair a schedule. It inherits it.
The industry sells schedule health checks and it sells schedule risk analysis, and almost nobody connects the two. The question that should come first, and rarely does, is whether the thing you are about to simulate can carry the weight of a simulation at all.
Published
Key Takeaways
- A simulation samples durations across the logic it is given. Where the logic is wrong, the model propagates the error and reports it with more decimal places, which is worse than not running it.
- Mechanical health checks are necessary and not sufficient. They test construction, not realism, and a schedule can pass all of them while nobody believes its critical path.
- Being straight about the evidence: no peer-reviewed study validates these metric sets or their effect on simulation validity. What is established is that network structure drives risk analysis output, which is a narrower and better-supported claim.
Research foundation
The authoritative public reference is the GAO Schedule Assessment Guide, which defines ten best practices across four characteristics of a reliable schedule. The peer-reviewed evidence bearing on this question is about network structure rather than about health-check metrics: Tavares, Ferreira and Coelho (1999) on the risk of delay in terms of network morphology, and Vanhoucke (2010) on activity sensitivity and network topology. One honest limitation is carried throughout: no peer-reviewed study validates schedule-quality metric sets or measures their effect on simulation validity. The Four A's are the executive lens applied to that material.
I have been handed schedules to analyse that could not compute a critical path, because every significant milestone was pinned by a hard constraint. The request was still for a P80 date. It is possible to produce one. It would have been a number about the constraints, not about the project.
What does a simulation actually do to a bad schedule?
It amplifies it, politely.
A Monte Carlo schedule analysis samples a duration for each activity from its assigned distribution, computes the network, and records the outcome, several thousand times. Every one of those computations obeys the logic in front of it. If two activities that genuinely depend on each other have no relationship, the simulation runs them in parallel, forever, in every iteration. If a hard constraint pins a milestone, the simulation respects the pin and reports a distribution that is narrower than reality because the model was not allowed to move.
The output looks identical either way. That is the danger. A percentile carries an implicit claim that the underlying model represents the project, and nothing in the histogram tells a reader whether it does.
What are the readiness questions, in order?
Does the network compute its own critical path?
Not whether the tool displays a critical path, but whether it emerges from the logic rather than from constraints. Hard constraints, artificial dates and open ends are the defects that most directly break a simulation, because they determine which paths can move. Confirming that the critical path is valid is one of the ten GAO practices, and it belongs before any simulation rather than after.
Is remaining work at a granularity where uncertainty can show?
A twelve-month activity with a single duration distribution cannot express that its first two months are well understood and its last six are not. It will contribute a smooth spread where the reality is lumpy. The practical rule is that activities in the near term should be short enough to fail visibly.
Are the durations anybody's real estimate?
This is the one that is not technical. Durations backfilled from a required end date are not estimates, and running a distribution around them models the uncertainty in a wish. The tell is a schedule where the sum of the optimistic cases still lands exactly on the committed date.
Has contingency been stripped?
If activity durations already contain private buffer, and the model then adds uncertainty on top, the result double counts and the percentile is not comparable to anything. This is why FTA requires a stripped and adjusted base schedule before its risk model runs, and it is good practice whether or not federal money is involved.
What about the fourteen-point check?
It is useful and it is oversold, and it is worth being precise about what it is.
The fourteen checks originate in a Defense Contract Management Agency publication, the Earned Value Management System Program Analysis Pamphlet. Two caveats matter. It is not maintained as a current, publicly hosted government standard, and the implementations shipped by scheduling tools differ from one another. So it is a widely adopted hygiene screen rather than an authoritative specification, and citing it as a government standard overstates it.
More importantly, it tests construction rather than realism. Logic density, lead and lag counts, float distributions and duration lengths are all properties of how the file was built. None of them can tell you whether the sequence reflects how the work will be done, or whether the durations are believed by the people who will do it. A schedule can pass fourteen out of fourteen and still be fiction, which is why the second question in this article's title matters more than the checks.
What does the evidence actually support?
Less than the industry implies, and it is better to say so.
I went looking for peer-reviewed validation that schedule health metrics predict analysis quality, or that poor construction degrades simulation validity in a measurable way. It does not exist. The space is occupied by vendor material and consultant white papers.
What is established is adjacent and narrower. Tavares, Ferreira and Coelho showed that the risk of delay can be characterised in terms of the morphology of the project network, and Vanhoucke demonstrated that activity sensitivity and network topology information govern how project time performance should be monitored. Both establish that the shape of the network materially determines what a risk analysis will produce. That supports the argument here, which is that the network is the instrument, without pretending there is a literature on health-check scores that there is not.
Why is this an Attention problem?
Because nobody has ever been thanked for stopping to ask whether the schedule was ready.
The analysis was requested, a date was set, and the person who could see that the network was not fit for purpose is usually junior to the person who wants the number. Raising it costs time and standing. Running it costs neither, and produces something that looks like an answer. The default therefore runs toward simulating whatever is in front of you, and the failure is invisible until the forecast is contradicted by events.
The condition that prevents this is not analytical skill. It is whether the organization has made it acceptable to say the input is not ready, which is the same dynamic as a register that is filled in and never consulted, described in The Risk Register Nobody Reads. In both cases the artifact is produced on schedule and the thinking is skipped.
Evidence matrix
| Claim | Evidence tier | Source |
|---|---|---|
| A reliable schedule has ten defined practices, including a valid critical path | Government guidance | GAO-16-89G, ten best practices across four characteristics |
| Network morphology governs the risk of delay | Peer reviewed | Tavares, Ferreira & Coelho (1999), EJOR 119(2) |
| Network topology determines how performance should be monitored | Peer reviewed | Vanhoucke (2010), Omega 38(5) |
| Baselines must be stripped of contingency before modelling | Government requirement | FTA Oversight Procedure 40 (October 2023) |
| No validated literature on schedule health metrics or their effect on simulation validity | Absence of evidence, stated as such | No peer-reviewed source located; DCMA checks not maintained as a live standard |
| Nobody is rewarded for stopping to ask whether the input is ready | Four A's interpretation | Builders Build, Attention |
What to do with this
Before the next analysis, ask for one page: the number of open ends, the number of hard constraints, the longest remaining activity, and a one-sentence answer to whether the durations were estimated or derived from the target date. If that page cannot be produced in an afternoon, the schedule is not ready, and the simulation would have told you something about the file rather than about the project.
References
- U.S. Government Accountability Office. GAO Schedule Assessment Guide: Best Practices for Project Schedules. GAO-16-89G, December 2015. gao.gov/products/gao-16-89g. Ten best practices across four characteristics: comprehensive, well-constructed, credible and controlled. Conducting a schedule risk analysis is itself the eighth practice.
- Tavares, L. V., J. A. Ferreira, and J. S. Coelho. “The Risk of Delay of a Project in Terms of the Morphology of Its Network.” European Journal of Operational Research, vol. 119, no. 2, 1999, pp. 510–537. doi.org/10.1016/S0377-2217(99)00150-2.
- Vanhoucke, Mario. “Using Activity Sensitivity and Network Topology Information to Monitor Project Time Performance.” Omega, vol. 38, no. 5, 2010, pp. 359–370. doi.org/10.1016/j.omega.2009.10.001.
- Williams, Terry M. “Criticality in Stochastic Networks.” Journal of the Operational Research Society, vol. 43, no. 4, 1992, pp. 353–357. doi.org/10.1057/jors.1992.50. On why the simulated critical path and the deterministic one diverge.
- Defense Contract Management Agency. Earned Value Management System (EVMS) Program Analysis Pamphlet, DCMA-EA PAM 200.1, section 4. Source of the fourteen-point schedule assessment. Not maintained as a current publicly hosted standard; publication date is inconsistently reported.
- Federal Transit Administration. Oversight Procedure 40: Risk and Contingency Review. U.S. Department of Transportation, October 2023. transit.dot.gov. Requires a stripped and adjusted base schedule.
- U.S. Government Accountability Office. Cost Estimating and Assessment Guide. GAO-20-195G, March 2020. gao.gov/products/gao-20-195g.
- AACE International. Recommended Practice No. 132R-23: Schedule Risk Analysis Maturity Model, Rev. 18 May 2024. Cited for scope and applicability; available to AACE members.
- Flyvbjerg, Bent. “From Nobel Prize to Project Management: Getting Risks Right.” Project Management Journal, vol. 37, no. 3, 2006, pp. 5–15. doi.org/10.1177/875697280603700302. On durations derived from targets rather than evidence.
About the Author
Dan Flynn
Creator of The Four A's of Organizational Readiness™ · Enterprise Transformation Executive · Author, Builders Build
Dan Flynn has spent thirty years inside federal, defense, and commercial organizations: diagnosing the invisible conditions that determine whether capable people produce extraordinary results. He is the creator of The Four A's of Organizational Readiness™ framework, has reached more than 11,000 professionals across corporate, civic, and national security contexts, and took a federal data platform from one release every six months to seventy-two every two weeks by changing organizational conditions: not people.
His book, Builders Build: The Four A’s of Organizational Readiness™, is forthcoming.
