AI Cycling
017 FTP Estimation And Fitness Models 1,694 words · 8 min

Fitting Critical Power And W’ From Your Own Ride Files

Your W’ is probably wrong. Not slightly wrong: wrong by 5 to 8 kJ, which is the difference between a model that tells you useful things about your 30-second sprint and a model that produces numbers with no relationship to what your legs can actually do.

The reason is simple and it is almost always the same. You have never gone properly deep at the short end. Your power duration curve has a beautiful, well-sampled region between 5 and 20 minutes because that is where FTP tests live and where your local climbs sit. Below about 90 seconds it is populated by whatever happened to occur in races, group rides and the occasional Zwift sprint you did half-heartedly because you were 90 minutes into a ride and your legs were cooked. The two-parameter CP model fits a straight line through whatever you give it. Give it a cooked 30-second effort and it will hand back a confidently wrong W’.

What the model actually is

Critical power is not a rebranded FTP. The two-parameter model (Monod and Scherrer, 1965, then Moritz and Whipp and a long line of work since) says the total work you can do in time t at a constant power is:

W_total = CP × t + W'

Divide both sides by t and you get the form that matters for fitting:

P(t) = W'/t + CP

That is a straight line if you plot mean power against 1/t. Slope is W’ in joules, intercept is CP in watts. That is the whole model. You can fit it in a spreadsheet, and doing it by hand once is worth more than any amount of reading about it, because you will see exactly how much leverage your shortest effort has over the slope.

Two things follow that people skip past. First, the model only holds in a window, roughly 2 to 15 minutes for most trained cyclists, sometimes out to 20. Second, because W’ is the slope and slope is governed by the extremes, your shortest data point is doing most of the work. That is the load-bearing fact of this whole post.

Fitting it: a worked example

Here are four maximal efforts from a rider I will call an honest 4.2 W/kg amateur, 72 kg, doing a proper test block on a quiet stretch of the A-road out past Blakeney. Four separate days, fresh each time, indoor trainer for consistency where possible.

DurationTime (s)Mean power (W)1/t (s⁻¹)Work (J)
3 min1803720.00555666,960
5 min3003370.003333101,100
8 min4803160.002083151,680
12 min7203020.001389217,440

Linear regression of P against 1/t gives:

CP  = 288 W
W'  = 15,200 J
R²  = 0.997

Check it against something the model didn’t see. Predicted 20-minute power: 288 + 15200/1200 = 301 W. Predicted 2-minute power: 288 + 15200/120 = 415 W. If this rider goes out and does an actual 2-minute maximal and holds 462 W, the model just under-predicted by 47 W and W’ is too small. Refit including that point and W’ will climb toward 19 kJ while CP drifts down a few watts. Same rider, same legs, a different answer because the short end finally got sampled.

That drift is not a flaw in the maths. It is the maths telling you your input set was lopsided.

Where the corruption comes from

intervals.icu will happily fit an eCP and W’ from your ride history using its power curve, and it is the best free tool for this by a wide margin. Go to Fitness → Power, set the date range to 42 or 90 days, and look at what the fit is anchored on. Now hover the 30-second and 1-minute points and click through to the actual ride. Nine times out of ten you will find:

  • A 30-second point set during a race, in a bunch, at minute 84, with 1400 kJ already in your legs
  • A 1-minute point from a Zwift sprint segment where you were chasing a bot and eased at the line
  • A 2-minute point that is actually the first two minutes of a 5-minute effort, which is the same thing as a pacing artefact

All three are submaximal. Not by much, maybe 6 to 10 percent, but the model multiplies that error. A 1-minute effort that is 8 percent low pulls roughly 2 kJ out of your W’. Two such points and you are down 4 kJ before you have made a single modelling decision.

Golden Cheetah’s extended CP model and WKO5’s mFTP both use different functional forms (3-parameter and a whole-curve model respectively), and they will produce different numbers again. WKO5 typically reports higher FRC than intervals.icu reports W’ for the same rider, partly because FRC is defined against a different anchor. Do not compare across tools and conclude you have improved. For the relationship between CP, FTP and the various models that claim to estimate them, the FTP estimation and fitness models pillar walks through why these numbers diverge and when it matters.

Calculate W’ from Strava: what you can and can’t do

Strava does not fit a CP model. It gives you a power curve on premium, and it stores your files, which is what you actually need. The workable route:

  1. Export the FIT files you care about. Individual activity → three dots → Export Original. For bulk, request your archive from Settings → My Account → Download or Delete Your Account → Request Archive. It arrives as a zip of FIT files.
  2. Get them into a tool that fits the model. intervals.icu syncs from Strava directly and is free for this. If you’d rather own the pipeline, fitparse in Python reads FIT records in about ten lines, and scipy.stats.linregress does the fit.
  3. Critically: choose your input efforts yourself. Do not let the tool pick maxima from 90 days of riding.

That third step is the whole game. A rough Python sketch, because the honest version of this is short:

from scipy.stats import linregress

# durations in seconds, mean power in watts - YOUR chosen efforts
t = [120, 180, 300, 480, 720]
p = [462, 372, 337, 316, 302]

inv_t = [1/x for x in t]
fit = linregress(inv_t, p)

print(f"CP  = {fit.intercept:.0f} W")
print(f"W'  = {fit.slope:.0f} J")
print(f"R2  = {fit.rvalue**2:.4f}")

Running that on the five-point set above:

CP  = 291 W
W'  = 20,600 J
R2  = 0.9932

CP moved 3 watts. W’ moved 5.4 kJ, a 35 percent change, from adding one honest 2-minute effort. If you take one thing from this: your CP is fairly robust to sloppy inputs, and your W’ is not.

Testing the short end properly

You need at least one genuinely maximal effort between 90 seconds and 3 minutes, done fresh, with a real warm-up and a pacing strategy that doesn’t fade. This is the most unpleasant test in cycling and that is exactly why nobody has a good version of it in their file history.

The protocol that works: 20 minutes easy, three 30-second builds at around CP with two minutes between, five minutes easy, then a single 2-minute effort where you start at roughly 130 percent of your assumed CP and hold or build. Not a sprint start. If your first 15 seconds averages 550 W and your last 15 averages 380, you have produced a fade curve, not a maximal effort, and the mean is worthless for fitting. Look at the file afterwards: you want the last 30 seconds within about 5 percent of the middle 60.

Do it on a separate day from your 12-minute effort. Not the same session, not after. A 3-minute all-out plus a 12-minute all-out in the same ride shares W’ between them and you will get a fit that looks tidy and means nothing.

Verify against a held-out effort. This step gets skipped constantly and it is the only thing separating a model from a guess:

DurationActualModel predictsError
4 min358 W377 W+19 W
10 min305 W325 W+20 W
20 min299 W308 W+9 W

A systematic over-prediction of 19 to 20 W across the middle means your CP is too high, which usually means your longest input effort was itself submaximal. Errors under about 8 W with no consistent sign are a fit you can use.

Which tools to actually run

intervals.icu for the day-to-day: the W’bal chart during a hard race file is the single most useful output of having fit the model at all, and you can override CP and W’ manually under Settings → Sport Settings, which you should do once you have your own numbers. Golden Cheetah if you want to see the 3-parameter fit and how much the extra parameter buys you (usually not much below 3 minutes, more than you’d expect above 20). WKO5 if you are already paying for it and want the full curve model, though its FRC is not your W’ and treating them as interchangeable will confuse you.

For the fitting itself, do it in Python or a spreadsheet the first time. I know that sounds like busywork when intervals.icu does it for free. It isn’t. Once you’ve watched your W’ jump 5 kJ because you swapped one input point, you will never again accept an auto-fitted W’ from 90 days of ride history without checking what it was anchored on.

The riders who get real value out of W’bal are the ones who tested their short end on a specific Tuesday in February, wrote down the number, and put it in manually. Everyone else has a decorative chart.