Home → Part 20
Cycle Time: Not Velocity, the Time Your Customer Waits
Friday 13 March, 11:00, sprint review. I put the velocity chart on the screen: 52 points, the record of the last six sprints. Ece from operations raised her hand. “We have waited 47 days for the commission line fix on the statement. Who is the record for?”
- Velocity is the team’s clock, not the customer’s. In six sprints, velocity went from 34 to 52. In the same period, the p85 lead time of customer requests grew to 58 days.
- There are three clocks, and each starts at a different point. Lead time starts at the request, cycle time at the start of work, and DORA lead time at the first commit.
- Stop using the average; use percentiles. For 10 items, the median was 5.5 days and the average 9.5 days. One 41-day item made the difference.
- Most of the time is waiting, not working. Only 28% of cycle time was active work. The rest was queues.
- Velocity went down, and the customer waited less. Three sprints later, velocity was 43 and cycle time p85 dropped from 17 to 10 days.
From the field: a record sprint and 47 days
The day before that review, I had written “Record sprint, well done” in the team channel. Over the last six sprints, velocity had been 34, 38, 41, 44, 47 and 52. I showed this chart at every review, and every time I felt a little prouder.
I could not answer Ece at that moment. After the review, I sat down with Ayşe. That quarter, Ayşe owned the team’s process measurement, and she could pull the status history of every item from Jira. We looked at Ece’s item first. The 47 days looked like this:
| Stage | Time | Active work |
|---|---|---|
| Waiting in the backlog | 29 days | 0 |
| Clarifying in refinement (between two meetings) | 3 days | 0 |
| Development (4 PRs, with review waits between them) | 6 days | 3 days |
| Waiting for end-to-end test + test | 5 days | 1 day |
| Waiting for operations acceptance + acceptance | 4 days | half a day |
| Total | 47 days | 4.5 days (10%) |
Of the 47 days, 4.5 were spent working. For the other 42.5 days, the item was waiting for someone somewhere. None of these rows appeared on the velocity chart. Velocity does not count how long an item waited. It counts how many points reached the “Done” column at the end of the sprint.
Then we looked at why velocity had grown. The answer was not nice. Over the last six sprints, a growing share of finished items were internal work: refactoring, small technical tasks, improvements the team suggested itself. They were short, clear and estimated. Customer requests were less clear and needed operations acceptance, so the team pulled them into sprints less often. Some points had also grown bigger. I explained how that happens in the story points post, so I will not repeat it here. The work that grew velocity was not the work the customer was waiting for.
Three clocks: where do they start and stop?
In the first meeting, we spent half an hour on definitions. Everyone meant something different by “lead time”. In the engineering metrics post, I described DORA lead time: from the first commit to production. For us, that number had been under three days since November. The path from code to production was fast. The problem was before that path, and between two PRs. We wrote the definitions down:
| Clock | Starts | Stops | Whose question |
|---|---|---|---|
| Lead time | Request is recorded | Usable in production | Customer: “When will I have it?” |
| Cycle time | Team starts the work (moved to “In progress”) | Meets the Definition of Done | Team: “How long do we take to finish what we start?” |
| DORA lead time | First commit | Running in production | Engineering: “How fast is the path for code?” |
The end point of cycle time matters. For us, “Done” means what our Definition of Done says: in production and accepted by operations. If “merged” counts as the end, cycle time looks good, but nothing changes for Ece. The start point matters too. The clock starts when the item moves to “In progress”. Moving the item back to “To do” does not reset it. Otherwise, moving waiting work back becomes the easiest way to make the number look better.
So do not compare the 4.5 days in the WIP limits post with the numbers in this post. That measurement was in November, when “Done” still meant checked in the test environment. In December, the Definition of Done got a production layer. From then on, the same clock also counted the weekly release train, 24 hours of monitoring and operations acceptance. The work did not get slower. The finish line of the clock moved.
Why does the average lie?
Ayşe first brought a table with the average. The 29 items finished in January and February had an average cycle time of 11.4 days. Nobody could do anything with that number. Then we looked at the last 10 items finished in March, one by one:
2 3 3 4 5 6 8 9 14 41
average = 95 / 10 = 9.5 days
median = (5 + 6) / 2 = 5.5 days # p50: half of the items are shorter
p85 = 9th of 10 = 14 days # 85% of items finish in this time or less
# one 41-day item (forgotten in operations acceptance)
# pulls the average up to 9.5. without it: 54 / 9 = 6 days.
Cycle time is skewed. Most items are short, and a few are very long. The average spreads those few items over everyone and gives two wrong messages. The typical item looks longer than it is, and the long item becomes invisible. But the 41-day item was the one we really needed to talk about.
So we use two numbers. p50 describes the typical item. p85 gives a sentence you can say to a customer: “85% of the items we start finish within 14 days.” Telling Ece “9.5 days on average” promises nothing. Telling her “14 days with 85% probability” does.
Flow efficiency: how much working, how much waiting?
Ece’s item was one example. For a general picture, we split 10 items finished in February into stages in the same way. Total cycle time was 90 days, and 25 of them were active work. Flow efficiency, the share of active time in total time, was 28%.
There is no single “right” value for this ratio, and it is hard to measure. You find active time by asking people or by estimating from the status history. So we do not measure it all the time. We measure it once a month, with a sample of 10 items. But even one measurement was enough. If three quarters of the time is waiting, writing code faster hardly changes cycle time. Shortening the queues does.
What did we change?
From 16 March, we made five changes. None of them was “work faster”.
- Waiting columns. We added “Waiting for review”, “Waiting for test” and “Waiting for acceptance” to the board. Waiting could no longer hide inside “In progress”.
- The ageing work rule. The board shows the age of every item in progress. If an item is older than the p85 (17 days at that time), it is the first item in the daily.
- An acceptance window. Operations acceptance is no longer random. Ece’s team keeps one hour on Tuesday and Thursday afternoons.
- A tighter WIP limit. I explained how we set it in the WIP limits post. The only difference: we applied the limit to the waiting columns too.
- Velocity left the review. It stays in planning, to talk about capacity. The review now shows the cycle time chart and the ageing items.
Three sprints later, on 24 April, the table looked like this:
| Measure | Before (Jan–Feb, 29 items) | After (16 Mar–24 Apr, 24 items) |
|---|---|---|
| Cycle time p50 | 9 days | 5 days |
| Cycle time p85 | 17 days | 10 days |
| Throughput (items finished / week) | 3.6 | 4.0 |
| Flow efficiency (sample of 10 items) | 28% | 46% |
| Customer request lead time p85 | 58 days | 39 days |
| Velocity | 52 (sprint ending 13 March) | 43 (average of three sprints) |
Velocity went down, and in the first week this bothered me. Then I looked at the throughput row. The number of items finished per week had not dropped; it had gone up a little. There were fewer points because customer requests were less clear, so estimates were more careful. We had also moved waiting internal work to second place. The team was delivering fewer points and more value.
There was a side effect too. Our sprint is ten working days, or two calendar weeks. When p85 was 17 days, many items pulled into a sprint did not fit into it from the start. They moved to the next sprint. Part of our missed sprint goals came from here, and we thought it was an estimation problem. With p85 at 10 days, an item pulled into the sprint will most likely finish inside it. Cycle time should be shorter than the sprint itself.
Lead time p85 is still 39 days. Most of it passes in the backlog, before anyone starts the work. That part is a prioritisation question, not a team question, and we discuss it separately with the PO. Cycle time measures the part the team controls. Lead time also shows the part the team does not control. If you do not look at both, you either blame the team unfairly or forget the customer.
What did not work for me
The table looks good, but we did not get there in a straight line. I tried three things and took them back.
1. Making cycle time a target
In the first week, I told the team: “Let’s get p85 down to 10 days.” By the end of the second week, the number really started to fall. Then Ayşe noticed something on the board. Three items in the “To do” column had actually started. Their branches were open and their first PRs were written. People moved items to “In progress” late, because the clock started then. Nobody was cheating. Everyone was optimising what was measured. We removed the target number and put one question in its place: “Where is the work waiting?” I had already learned this with DORA. I learned it again with a new number.
2. Removing velocity completely
My second reaction was to delete velocity from the board. One sprint later, we had nothing to talk about capacity with in planning. The team started to answer “what fits in this sprint?” by feeling. We brought velocity back, but only to planning. The number itself was not wrong. Showing it to the customer as an answer was.
3. Measuring flow efficiency on every item
To find active time correctly, I asked everyone to log hours on every item. They lasted three days. On the fourth day, half the entries were missing. On the fifth day, nobody logged anything. When measuring costs more than the information is worth, the measurement dies. We moved to a monthly sample of 10 items and a half-hour conversation. It is less precise, but it happens every month.
How it breaks
- Write down the start and end points
- Use p50 and p85; put the average only in a footnote
- Make waiting states separate columns
- Make ageing work the first item in the daily
- Show cycle time and lead time together
- Count “merged” as done
- Reset the clock when an item moves back
- Show cycle time per person
- Put velocity and cycle time in the same chart as competitors
- Celebrate one record sprint
The third item is the most dangerous. If you connect cycle time to a person, people stop pulling hard work, or they split work into tiny pieces. I saw how fast this happens with our per-person PR dashboard. Cycle time is a team number. An item gets done by passing through several people’s hands.
What to track
- Cycle time p50 and p85 — items finished in the last 30 days, as a trend.
- Ageing work — the age of every item in progress, with the p85 line.
- Throughput — items finished per week. A count, not points.
- Items in waiting columns — where the queue is growing.
- Customer request lead time p85 — monthly, with the PO.
- Flow efficiency — monthly, with a sample of 10 items.
Checklist
- Are the start and end points of cycle time written down?
- Is “Done” the moment the customer can use it, or the merge?
- Am I showing p50 and p85 instead of the average?
- How many items in progress are already older than p85?
- Which column has the longest wait?
- When velocity went up, did I also look at how long the customer waits?
- Is there any dashboard that connects cycle time to a person?
Conclusion
The 52 points I showed on 13 March were real. The team really had worked hard. But that chart answered my question: “How much does the team produce?” Ece’s question was different: “How long do I wait?”
Now the first chart in the review is cycle time. Velocity stayed in planning, to talk about capacity. In the last review, Ece did not ask anything. She only wanted to change the time of the acceptance window. When the customer does not wait, nobody asks about records.
Reading cycle time through percentiles and the idea of ageing work come from Daniel Vacanti’s book Actionable Agile Metrics for Predictability. The definitions, measurements and changes come from our own team.