Home → Part 21
“When Will It Be Done?”: Not One Date, a Range of Probabilities
Monday 12 January, 11:05, in the corridor. The deputy general manager stopped me. “When will the new account opening flow be done?” “27 March,” I said. One date, with the day. On 27 March, 31 of 47 items were done. We went live on 5 May.
- One date is an answer with its probability hidden. When I said “27 March”, I did not say how sure I was, because I did not know.
- Dividing points by velocity makes three hidden assumptions. Velocity stays the same, scope does not grow, and all the team’s points go to this work. None of them helped.
- Past throughput is enough data. How many items finished each week in the last 10 weeks? That list and a 20-line Monte Carlo gave a more honest answer than one date.
- Say two numbers: 50% and 85%. 85% means: if we did this work seven times, six times it would be done by this date or earlier.
- Run the forecast again every week. As data comes in, the range gets narrower. If the date is slipping, you see it weeks earlier.
From the field: where did 27 March come from?
That morning in the corridor, I did the calculation in my head. The account opening flow was split into 41 items, 196 points in total. The last sprint’s velocity was 34, but it was going up. I said to myself, “40 is fine.” 196 divided by 40 is 4.9 sprints. I rounded up to five sprints. If we started on 19 January, the fifth sprint would end on Friday 27 March. I added no buffer, because velocity was rising.
This calculation had three assumptions that I did not say out loud. Later, I put each one next to what really happened:
| Assumption | Reality (19 Jan–27 Mar) |
|---|---|
| Velocity will be 40 | It was, and more: 41, 44, 47, 52, 42. But most of the points came from internal work |
| All team capacity goes to this work | In 10 weeks, 31 items from this project finished: 3.1 per week. The team total was 3.6 |
| Scope stays at 41 items | It grew to 47: 6 more items, 15% growth |
The first row hurts the most. My velocity assumption was right, but it did not help. Velocity counted all of the team’s points, not this project’s points. I explained where the points came from in the cycle time post: internal items were short and clear, and they got finished while customer work waited. I had also explained what breaks when you turn points into dates, in the story points post. I wrote that post in November, and I made the same mistake myself.
On 27 March, 31 of 47 items were done. The remaining 16 items took four more weeks. The last one was done on 24 April, and with the release train and operations acceptance we went live on 5 May. When and how I reported the risk is a separate topic. This post is about the date itself: how I produced it.
Why does one date not work?
The finish date of software work is not one point. It is a distribution. Some weeks 2 items finish, some weeks 6. Some items close in one day, some wait three weeks for acceptance. When you squeeze this uncertainty into one date, you do one of two things. You pick a very optimistic point, or you add a hidden buffer. In both cases, the other person does not know how much risk they are taking.
The most common mistake is to use the average. “We finish 3.8 items per week on average. There are 42 items, so 11 weeks.” This calculation gives you a date with about 50% probability. In other words, a coin toss. Below, I will show this with our own data.
One note: inside a sprint, this problem does not exist. In sprint planning, you make a two-week commitment, and velocity is enough for that question. The problem starts when you stretch the same arithmetic to ten weeks of work. Small differences that look harmless in two weeks add up over ten weeks. One date shows none of them.
Count items, not points: throughput
Monte Carlo needs only one piece of data: in each of the last few weeks, how many items finished? A count, not points. There are two reasons. First, a count needs no estimate. The number of items that reached “Done” in Jira is not open to debate. Second, if you split work into similar sizes, the variation in the count gets smaller. In our team, most items take two to five person-days. Whether an item is 5 or 8 points hardly changes the result of the forecast.
One distinction matters. If you ask when one item will be done, the answer is the cycle time p85: “85% of the items we start finish within 10 days.” If you ask when many items, a whole project, will be done, cycle time is not enough. Items move in parallel, not one after another. There, you need throughput and a simulation.
Monte Carlo: a 20-line calculation
In early April, the next project arrived: the card application flow, 36 items. This time I did not answer in the corridor. I said, “I will bring two numbers on Friday.” For scope growth, I used the rate from the first project: 36 × 1.15 = 41.4, rounded up to 42 items. Then I ran this calculation:
import random
history = [3, 5, 2, 4, 6, 3, 4, 5, 2, 4] # last 10 weeks: items finished each week
remaining = 42 # 36 items + 15% scope growth buffer
trials = 10000
results = []
for _ in range(trials):
done, weeks = 0, 0
while done < remaining:
done += random.choice(history) # pick a random past week
weeks += 1
results.append(weeks) # how many weeks this trial took
results.sort()
for p in (50, 85, 95):
print(p, results[int(p / 100 * trials) - 1])
# output:
# 50 11
# 85 13
# 95 13
The logic is simple. We assume each future week will look like one of the past weeks, but we do not know which one. So we pick one at random and count until the work is done. When you do this ten thousand times, you do not have one answer to “how many weeks?”. You have ten thousand answers. Then you sort them and read them. The project would start on Monday 27 April, after the last items of the account opening flow. When we turned weeks into dates, the table looked like this:
| Finish week | Date (Friday) | Chance of finishing by this date |
|---|---|---|
| 10 | 3 July | 18% |
| 11 | 10 July | 53% |
| 12 | 17 July | 83% |
| 13 | 24 July | 96% |
With the average, I would have said: 42 / 3.8 = 11.05, so 11 weeks, 10 July. In the table, that is 53%. I was about to repeat my January mistake with cleaner arithmetic.
- The future looks like the past. If someone leaves the team, or two people go on holiday, the last 10 weeks do not describe the future.
- The history includes internal work. During the card project, we reduced internal work but did not stop it. We assumed the 15% buffer would cover the difference. This was the weakest assumption of the model.
- Items are split in the same way. If some of the 36 items are still big, throughput will mislead you.
Three objections from the team
When I first showed the calculation to the team, three questions came up. All three were fair, and the answers also show the limits of the model.
“Isn’t 10 weeks too little?”
This was Mehmet’s question. It is little, but not as little as you think. In the history, the worst week is 2 and the best week is 6. Next week could fall outside this range, but it is not likely. If you use a longer history, you also bring the old team and the old process into the model. We use the last 10 weeks. Every week, we drop the oldest week and add the newest. When the team changes a lot, for example when someone leaves, we drop all data from before that week.
“Why 10,000 trials?”
The number is not magic. With 1,000 trials, the results could move by a week from one run to the next. With 10,000 trials, they did not. The calculation takes less than a second, so we did not bother with fewer trials. The number of trials is less important than the quality of the input. A million trials with wrong throughput data also give a wrong range. It just looks more precise.
“But items are not the same size?”
This was Ayşe’s question, and the most important one. They are not the same size, but the items in past weeks were not the same size either. The model assumes future items come in a similar mix to past items. This holds as long as you split work in a similar way in refinement. In the first list of the card project, two items looked longer than a week. We split them before starting. 36 is the number after splitting.
How did I explain the range to management?
My first attempt was bad. On Friday 10 April, I showed the deputy general manager the four-row table above and a histogram. Five minutes later I heard: “OK, let’s say 3 July.” The first row of the table, the date with an 18% chance, had become a promise in his eyes. The tallest bar in the histogram was week 11, and he read “the tallest bar” as “the most likely date”. Both were my mistakes: too many numbers and a chart without explanation.
One week later, I brought the same information on a single card:
- Card application flow: 10 July with 50% probability, 24 July with 85% probability.
- What 85% means: if we did this project seven times, six times it would be done by 24 July or earlier. Once, we would be late.
- What it is based on: items finished in the last 10 weeks, and 42 items (including a 15% growth buffer).
- What would change it: if scope goes above 42, or someone leaves the team, the date moves.
- Update: we run it again every Monday.
The question changed. He asked, “Which date should I give?” I said, “If you make a promise outside, use 85%. Inside, let’s plan for 50%. If we slip, there are two weeks in between.” Outside, the date was 24 July. Inside, the target became 10 July.
The real gain was the “every Monday” line. A single date is said once, and after that everyone defends it. A range is calculated again every week with new data. If the date is slipping, you see it weeks earlier, not in the last week, and there is still time to decide.
Today: 31 May
Five weeks have passed. 17 items are done, 3.4 per week. That is below the 3.8 the model assumed; internal work took more space than I expected. Scope grew from 36 to 38. 21 items remain, or 24 with the 15% buffer. This morning I ran the calculation again:
| When | 50% | 85% | Items remaining |
|---|---|---|---|
| 10 April (first forecast) | 10 July | 24 July | 42 (36 + buffer) |
| 31 May (week 5) | 17 July | 24 July | 24 (21 + buffer) |
The 50% date moved by one week. The chance of the internal target, 10 July, dropped to 41%. Tomorrow morning I will tell the deputy general manager myself, before he asks. But the 24 July date we gave outside is still safe; the chance of finishing by then is 98%. In January, I had this conversation after 27 March had passed. This time, I am having it six weeks before the target date.
How it breaks
- Use the count of finished items instead of points
- Measure scope growth from a past project and state the buffer openly
- Say only two numbers: 50% and 85%
- Explain 85% as “six out of seven”
- Run the forecast again every week
- Give dates in the corridor
- Calculate with the average and call it a forecast
- Show a histogram without explanation
- Use old throughput after the team has changed
- Say the range once and never update it
Checklist
- Do I know how likely I am to meet this date?
- Did I calculate the date from points or from past throughput?
- Did I add a buffer for scope growth, and if so, did I say it?
- How much of the team’s capacity really goes to this work?
- Does the other person know the difference between 50% and 85%?
- Is the day of the next forecast update fixed?
- Did I write down what would change the forecast?
Conclusion
When I said “27 March” in the corridor on 12 January, I did not give a wrong date. I gave the wrong kind of answer. The other person asked about a probability. I gave him a promise, and I hid its probability, even from myself.
Today my answer to the same question is two dates and a day: “50% by this one, 85% by that one, I will update it on Monday.” It is a longer answer, but it creates fewer meetings. One date gives comfort. A range earns trust.
Monte Carlo forecasting with throughput and communicating with percentiles come from Daniel Vacanti’s book When Will It Be Done?. Troy Magennis’s approach to adding scope growth to the model is also a useful reference. The data, code and forecast card come from our own team.