Schedule risk and confidence: from pass/fail to your odds of finishing on time
The DCMA checks tell you whether a schedule is well built. A schedule risk analysis tells you how likely it is to finish on time. Here is how the Monte Carlo confidence level, risk drivers, and statistical checks work.
The DCMA 14-point assessment answers one question: is this schedule built well enough to manage to? It is a structural, pass/fail check. What it cannot answer is the question a program manager actually gets asked: how likely are we to finish on time? That is a statistical question, and it needs a statistical tool: a schedule risk analysis. GanttScore now runs one on every file, on top of the structural checks.
Schedule Confidence Level: your odds, as a date
A single planned finish date hides the uncertainty behind it. Every task in the plan is an estimate, and estimates carry a range. A schedule risk analysis samples that range across the whole network, thousands of times, and reports the distribution of finish dates that results. That gives you a Schedule Confidence Level (SCL):
- P50: the date you have a 50/50 chance of hitting. If it is later than your plan, your plan is optimistic.
- P80: the date you can be 80% confident of hitting. This is the number reviewers and sponsors ask for.
- Probability of on-time: the share of simulations that finish on or before your planned date. A low number is the honest signal that the plan needs margin.
This is a Schedule Confidence Level, a schedule-only estimate. It is not a Joint Confidence Level (JCL), which combines cost and schedule; GanttScore analyzes the schedule, not the budget. Where the 70%-confidence date is mentioned, it is a common reference level, not an endorsement by any agency.
How the simulation works
GanttScore runs a Monte Carlo simulation over your logic network. On each of a thousand iterations, every incomplete task draws a duration from a range, and the resulting slip propagates forward through the real dependencies, so a delay on one task pushes everything downstream, exactly as it would in the field. Merge points, where several chains converge, correctly accumulate risk (the reason parallel work finishes later on average than any single chain suggests).
By default, durations are sampled from a conservative range (a triangular distribution centered on the plan, skewed toward overrun, because durations overrun far more often than they come in early). That default is a disclosed starting basis, not a per-task calibration. The report says so plainly. Results are reported as slip against the schedule’s own dates, which keeps the working-calendar effects already baked into your plan intact.
What drives your risk (the tornado)
Knowing you are likely to be late is only half the answer. The report ranks the tasks whose duration uncertainty most moves the finish date (a sensitivity “tornado”) and pairs each with a criticality index (how often that task lands on the driving path across the simulations). A long task with zero float that sits on the critical path 90% of the time is where your recovery effort belongs. A task with plenty of float, however large, is not, and the analysis shows you the difference, which a static critical-path view cannot.
Two statistical checks the fixed rules miss
Alongside the simulation, GanttScore adds two checks that read your schedule’s own distribution rather than a fixed threshold:
- Duration realism: the share of task durations that land on round numbers (5, 10, 20, 30 days). Heavy clustering on round numbers is a sign durations were assigned by convention rather than estimated from the work. This is a heuristic signal, not proof of padding (some round durations are genuinely correct), but reviewers look for exactly this, and no fixed rule surfaces it.
- Relative outliers: tasks whose duration or float is far outside the norm for their own WBS area, using a robust statistic (the modified z-score) that is not thrown off by the very outliers it is hunting. These pass the fixed DCMA thresholds yet stand out against your own schedule: often a units slip, a placeholder, or a missing link.
Health and risk, together
Structural health and schedule risk answer different questions, and you need both. A well-built schedule (high DCMA score) can still carry an ugly confidence curve if the work is genuinely tight. A poorly built one (low score) makes the risk numbers unreliable, because the float and critical path they rest on are themselves suspect, which is why GanttScore fixes the structure first and reads the risk second. See how the health score works.
Frequently asked questions
Is this a NASA Joint Confidence Level (JCL)?
No. A JCL is a joint cost-and-schedule confidence level. GanttScore produces a Schedule Confidence Level (the schedule half only) because the uploaded file rarely carries trustworthy cost data. We are explicit about that distinction everywhere it appears.
Where do the duration ranges come from?
By default, a conservative triangular spread centered on each planned duration and skewed toward overrun. It is a disclosed starting basis, not a calibrated per-task estimate. It gives a defensible first read; a future option will let you supply your own ranges.
Will I get the same numbers if I run the same file twice?
Yes. The simulation is seeded from the file, so an identical upload always produces identical results. The numbers are reproducible, not random from run to run.
Does the risk analysis change my health score?
No. The statistical layer is reported separately and never affects the 0-to-100 DCMA score, so the score stays comparable across every report.