Sertaç Yıldırım field notes

Home → Part 07

Story Points: Not Time, Uncertainty

Thursday, 12 June 2025, a roadmap meeting. The director looked at our velocity chart: “The payments team closes 81 points per sprint. You close 58. And tell me this: how many hours is one point?” I opened the calculator on my phone and answered: “About five and a half hours.” We paid for that sentence for three months.

Summary
  • Points measure uncertainty, not time. Our 5-point items took from 1.5 to 9 working days. As points grow, the range grows, not the average.
  • Points die the moment they become hours. After I said “1 point = 5.5 hours”, rounds with split votes fell from 31% to 8%. The uncertainty had not gone down. We just stopped talking about it.
  • Velocity is not a comparison tool. 81 and 58 are two different rulers. After the comparison started, we blindly re-voted ten items. Seven got more points.
  • We tried #NoEstimates for six sprints. Counting stories improved forecasts and cut planning from 95 to 55 minutes. But the uncertainty signal disappeared, and mid-sprint surprises went from 3 to 7.
  • What we kept: story counts and a range go up to management, never hours. Points stay in the team room and serve one question: what do we not know about this work?

What are points for?

I explained how planning poker works, and why it should be quick, in the Sprint Planning part. Here I look at the number on the table, not the game. What does that number measure, and what does it not measure?

The book definition is this: a story point gives the size of a piece of work, relative to other work. It considers effort, complexity and uncertainty together. The word “hour” is not in this definition, and that is not an accident. Because points are relative, they belong to the team. “This is a bit bigger than the report field we built last month” only means something to the team that built that report.

In practice, the most valuable part of a point was not the number. It was the gap between votes. When one person shows 3 and another shows 13, there is a knowledge gap at the table. Someone knows something the others do not: an old module, an undocumented integration, a surprise from last time. If you talk about that gap, the uncertainty comes up during planning. If you do not, it comes up in the middle of the sprint.

What points do not measure is as important as what they measure. Points do not measure a person’s speed. Mehmet finishes a 3-point item in one day; a new person may need three days, and both voted correctly. Points do not measure value. A 13-point item may give the customer nothing, while a 1-point fix may save the call centre an hour every day. Points do not measure time either, but the rest of this post is the story of that misunderstanding.

In planning poker, the real information is not in the average. It is in the distance between the highest and the lowest card.

The real link between points and time

After the director’s question, I did what I should have done before answering. I pulled 62 items finished in 12 sprints, from October 2024 to March 2025, out of Jira. For each one I took its points and the time it spent in the “In progress” column. I counted one working day as 8 hours.

PointsItemsMedian (working days)Shortest – longestHours per point
190.60.3 – 1.24.8
2121.10.5 – 2.54.4
3131.90.8 – 45.1
5144.21.5 – 96.7
898.53 – 198.5
135176 – 3110.5

Two things stand out. First, “how many hours is one point?” has no single answer. Small items take 4–5 hours per point; big items take more than 10. My “five and a half” came from dividing capacity hours by velocity: 5 people (the sixth was on the operations lane) × 10 days × 6.5 productive hours = 325 hours, and 325 / 58 ≈ 5.6. The math was right. The question was wrong.

The second point matters more. As points grow, the range grows, not just the time. For a 2-point item, the gap between the shortest and the longest is two days. For 8 points it is 16 days, and for 13 points it is 25 days. 8 and 13 points do not mean “big work”. They mean “we do not know how long this will take”. That is also why the T-shirt scale in the Backlog Refinement part grows exponentially. As work grows, it is not only time that grows. Uncertainty grows too.

From the field: points turned into hours

1. Split votes disappeared

After 12 June, “5.5 hours” moved to the team channel and from there to the roadmap slides. Something changed in planning poker, and I did not notice it for two sprints. People started to calculate hours in their heads before showing a card. “Two days of work, 16 hours, let’s say 3 points.” When everyone does the same math, everyone shows the same card.

I counted from our planning notes. In the six sprints before that meeting, 31% of rounds had a gap of at least two cards between the lowest and the highest vote. In the next three sprints, it was 8%. The uncertainty had not gone down. It just no longer reached the table.

2. An epic planned for 1.5 sprints took 5

The “Notification preferences” epic was 89 points. The slides did the math like this: 89 × 5.5 = 490 hours, and 325 hours per sprint, so 1.5 sprints. We started on 23 June, and the roadmap said “mid-July”. The epic finished on 29 August.

Most of the delay came from two 13-point items. One was a new consent integration with our SMS provider. In the vote, Ayşe showed 21: “We do not know how long the provider’s approval takes.” The others showed 5 and 8. Before, this gap would have meant a ten-minute talk and probably a spike. A spike is a short, time-boxed task to learn something before you commit. That day we said “let’s take the middle, 13” and moved on. In our heads, 13 points was 71.5 hours, and that looked big enough. The item took 23 working days. The provider’s approval alone took 11 days.

3. Comparison inflated the points

The sentence “payments 81, you 58” reached the team. Nobody did anything on purpose. But in the vote, it became easier to say “this is actually a bit harder”, and harder to push back with “I think this is a 3”. In September I ran a test. I took 10 items finished in the spring and asked the team to vote on them again, without showing the old points. Seven got more points than before, three stayed the same, and none went down. The average went from 3.2 to 4.6.

This is how inflation happens. No lies, no agreement. The reference just slowly moves. The unit of measure gets smaller, and the same work is worth more points. The payments team’s 81 and our 58 were already two different rulers. The comparison also broke our own ruler.

The moment points leave the team, they stop being an estimate and become a performance score.

The #NoEstimates trial: six sprints of counting stories

When we saw that points were broken, we tried dropping them completely. For six sprints, from 4 August to 24 October, the rules were these:

Trial rules (August – October 2025)
1. No points. One question in planning: "Does this story fit in 3 days?"
2. If not, split it. Each part must be a change the user can see.
3. Forecast = story count of the last 6 sprints, as a range, not one number.
4. Report to management: remaining stories / stories per sprint range.

# example: 20 stories left, 5-7 stories per sprint
# 20 / 7 = 2.9   20 / 5 = 4.0   -> "between 3 and 4 sprints"

The “change the user can see” condition in rule two was not there on day one. In the fifth sprint, the story count was about to reach 13. When I looked, some stories had been split into “backend” and “frontend”. The number had gone up, but there was no more work to show. We had escaped point inflation and walked into count inflation. We added the condition and counted that sprint as 9.

Measure (6 sprints)Points (May – July)Story count (August – October)
Planning time (average)95 min55 min
Commitment accuracy (done / taken)74%78%
Stories done / sprint6.86.0 (4–8)
Stories over 3 working days / sprint2.31.5
Items whose uncertainty appeared mid-sprint (total)37

Planning time, accuracy and the number of big stories supported the trial. The splitting rule made work smaller. Smaller work was easier to forecast, and planning got shorter. The drop in stories done did not come from the trial. In the same months, we were also testing a 14-item readiness list that let less work into the sprint. That is a story for another post. “Between 3 and 4 sprints” was much more honest than “490 hours”, and the director had no problem with a range. I saw no resistance to giving a range instead of one number. The resistance was in me, in my habits.

The last row was the cost. Everyone easily said “yes” to “does it fit in 3 days?”. With points, split votes shouted “there is something we do not know here”. A yes/no question turned that voice down. Two of the seven items were big surprises. A task that read the end-of-day file from the Central Registry Agency took 8 days. A task that touched the old commission module took 11 days. In both cases someone in the team knew the area. In both cases nobody asked them, because there was no moment that made anyone ask.

How it should work: what we kept

After the trial we brought points back, but for a different job. Points are no longer a time estimate. They are a tool that brings uncertainty to the surface. Story counts do the forecasting.

Do
  • Talk about every item with split votes; if the gap is more than two cards, run a 1–2 day spike first
  • Give management a story count and a range: “between 3 and 4 sprints”
  • Split every story into parts shorter than 3 days that the user can see
  • Every quarter, blindly re-vote 10 reference items and measure the drift
Don't
  • Answer “how many hours is one point?” with a number
  • Show two teams’ velocity on the same chart
  • Compare velocity across different sprint lengths
  • Say “let’s take the middle” when votes split

We also wrote down what happens when votes split. First, the people with the lowest and the highest card speak, two minutes each. Then we vote once more. If the gap is still more than two cards, the item does not enter this sprint. A one- or two-day spike enters instead, and the answer comes to the next planning. In the first two sprints we opened four spikes. In three of them, the person with the highest card was right. In one, the work was easier than expected, and the item came back as 2 points. The four spikes took six days in total. The SMS item in “Notification preferences” alone took 23 days, so this is cheap insurance.

The third “don’t” comes from the Sprint Length part: 20 points in a one-week sprint is not the same as 40 points in a two-week sprint. The same logic is even stronger between teams. Instead of the comparison, I suggested this to the director: “If you want to compare the two teams, let’s look at delivered work and how long requests wait. Let’s not look at points.” He agreed. How we track those numbers is a topic for another post.

The first result of re-voting the reference items was a surprise. We reset the reference list, and at the planning at the end of October we took 36 points into the sprint. With the old reference, the same work would have been about 58. Nobody is working slower. The ruler is back in place.

What to watch

  • Share of rounds with split votes. If it drops suddenly, the team is calculating hours or avoiding disagreement.
  • Drift in blind re-votes of reference items. If the average goes up, points are inflating.
  • Time range per point value. If 8- and 13-point items become more common, nobody is splitting work.
  • Items whose uncertainty appears mid-sprint. This is the real test of any estimation method.

Checklist

For the next planning
  • Do points or velocity leave the team as hours or as a comparison?
  • Did we talk about items with votes more than two cards apart, or did we round to the middle?
  • Did a 13-point or bigger item enter the sprint without being split?
  • Is the forecast I give to management one date, or a range?
  • When did we last blindly re-vote the reference items?
  • When we split stories, is each part a change the user can see?

Conclusion

On 12 June the director asked me a fair question. He had to build a roadmap, and all he had was points. The mistake was not in his question. It was in my answer. I could have said “points do not measure hours; I can give you a story count and a range”. Then “Notification preferences” would still have taken 5 sprints. But nobody would have expected 1.5, and we would have talked about Ayşe’s 21.

#NoEstimates showed us that forecasts can live without points. The same trial showed that points do a job other than forecasting: they make the knowledge gap at the table visible.

A point is not the answer to “how long will it take?”. It is the answer to “how unsure are we?”.