AI Cycling
011 FTP Estimation And Fitness Models 1,860 words · 8 min

Ramp Test, 20-Minute Test Or Modelled FTP: Which Should You Actually Run?

Two riders turn up to the same Wednesday club run. Both have a true critical power of 300 W. Both weigh roughly the same. One is a 25-mile TT specialist who hasn’t sprinted since 2019; the other is a cyclocross rider who wins from a two-lap move and dies on any climb longer than four minutes. They both open Zwift, both run the standard ramp test, and they walk away with FTPs 24 watts apart.

Neither test malfunctioned. The ramp protocol is behaving exactly as designed. The problem is that its design embeds an assumption about your anaerobic capacity, and if you sit away from that assumption, the number it hands you is wrong in a direction you can predict in advance. That is the whole argument of this piece: ftp test protocols compared on paper look like a trade-off between pain and accuracy, but in practice the deciding variable is your phenotype, not your pain tolerance.

What each protocol is actually measuring

Strip the branding off and there are three families.

Ramp tests (Zwift’s Ramp Test, TrainerRoad’s Ramp Test, most lab MAP protocols) push power up in steps until you fail, take your best one-minute power, and multiply by 0.75. What they measure directly is maximal aerobic power. FTP is inferred.

Constant-effort tests (the classic 20-minute test after a 5-minute blowout, the 8-minute pair, the 3-minute all-out test) measure a power you genuinely produced for a duration close to the thing you care about. The 20 gets multiplied by 0.95, the 8s by 0.90, and the 3-minute all-out test takes your final 30 seconds as a direct read on CP.

Modelled FTP (intervals.icu eFTP, WKO5 mFTP, Xert Threshold Power, Strava’s estimate) fits a curve to efforts you already did and reads threshold off the fit. No test day. The quality of the answer is entirely the quality of the efforts feeding it.

The ramp test’s hidden term

Run the critical power model through a linear ramp and the algebra is short. You fail when you have spent your W’, so:

peak power at failure = CP + sqrt(W' × r / 30)

  W' in joules, r = ramp rate in W/min
  best 1-min ≈ peak − (half a step)

Now put the two riders in. Diesel has W’ = 15.0 kJ. Cross rider has W’ = 26.0 kJ. Zwift’s standard ramp is 20 W/min.

Diesel (W’ 15.0 kJ)Cross rider (W’ 26.0 kJ)
True CP300 W300 W
Predicted 3-min383 W444 W
Predicted 20-min313 W322 W
Ramp best 1-min~390 W~422 W
Ramp FTP (×0.75)293 W317 W
20-min FTP (×0.95)297 W306 W
Ramp error−7 W+17 W
20-min error−3 W+6 W

Both multipliers are calibrated, it turns out, for almost exactly the same rider: one whose W’ is about 60 joules per watt of CP. Set W’/CP = 63 and the 0.95 rule returns your CP exactly. Set it near 60 and the ramp does too. That is not the interesting part. The interesting part is what happens as you move away from that calibration point, because the ramp carries W’ inside a square root multiplied by 0.75, while the 20-minute test carries it as W’/1200 multiplied by 0.95. Differentiate both and the ramp is roughly three times as sensitive to your anaerobic capacity. Same drift in phenotype, triple the error.

Worse, the ramp has a second free parameter the 20 doesn’t have: the ramp rate. Take the cross rider and change nothing except the step size.

W' = 26.0 kJ, CP = 300 W

  15 W/min  →  best 1-min ~407 W  →  FTP 305 W
  20 W/min  →  best 1-min ~422 W  →  FTP 317 W
  25 W/min  →  best 1-min ~435 W  →  FTP 326 W

Twenty-one watts of spread from the protocol’s own configuration. TrainerRoad partly sidesteps this by scaling steps to about 6% of your current FTP per minute, which normalises the rate across rider sizes, but that also means a wrong starting FTP gives you a wrong ramp rate which gives you a wrong new FTP. The error is self-reinforcing across test cycles. Zwift’s fixed 20 W/min (and the smaller-step Ramp Test Lite) avoids the feedback loop and instead makes the rate relatively harsher for smaller riders.

Modelled FTP: better mechanism, worse inputs

Modelled estimates fix the phenotype problem in principle, because a two- or three-parameter fit separates CP from W’ instead of conflating them. In practice they fail on data hygiene.

intervals.icu’s eFTP reads off your power curve from efforts of roughly three minutes and up, and takes the best estimate it can find. If you spent the last six weeks doing Zwift racing, your curve is a wall of 3-to-5 minute maximals and a desert past 10 minutes. The model is being asked to extrapolate a threshold from the exact region where a high-W’ rider’s curve is most inflated. Riders in that situation routinely see eFTP sitting 10-15 W above anything they can hold for 40 minutes, and the platform is not wrong to report it: you did produce those numbers.

WKO5 is the most useful of the group because it tells you the thing you need. Alongside mFTP it reports FRC in kilojoules (its W’ analogue) and TTE, the modelled time you can hold mFTP. Use TTE as the lie detector. An mFTP of 320 W with a TTE of 27 minutes is not a threshold, it’s a 30-minute power with a threshold label on it. Real FTP typically models out at 40 to 70 minutes of TTE. If yours is under 35, the number is high and your FRC is probably why.

Xert is the outlier worth naming: it refuses to hand you a single number, fitting Threshold Power, High Intensity Energy and Peak Power simultaneously from every ride. A signature like TP 298 W / HIE 25.4 kJ / PP 1140 W cannot flatter you the way a ramp can, because the punch has somewhere to live. The cost is that Xert needs weeks of varied, genuinely hard riding before the signature stabilises, and it will drift on a diet of endurance rides.

Work out your phenotype first, in about four minutes of arithmetic

You need two maximal efforts, fresh, from the last six weeks, on similar terrain: a best 3-minute and a best 12-minute. Then:

W' = (P3 − P12) / (1/180 − 1/720)      =  (P3 − P12) / 0.0041667
CP = P12 − W' / 720

A gravel and CX rider I’ll call Rachel has a best 3-minute of 395 W and a best 12-minute of 318 W.

W' = (395 − 318) / 0.0041667  =  18,480 J   (18.5 kJ)
CP = 318 − 18,480/720         =  292 W
W'/CP = 63 J/W

Sixty-three. Rachel is the rider both standard multipliers were built for, and she can use whichever protocol she likes. The 0.95 rule will land within a watt of her CP.

Now the shortcut for anyone who doesn’t want to do the division: compare your best 3-minute power to your best 20-minute power. Under 25% higher means low W’. Over about 33% higher means high W’, and the ramp test will flatter you. Rachel’s ratio is 28%.

So, which one

If your W’/CP sits above 70 J/W (crit racers, pursuiters, CX, anyone who wins on 30-second accelerations): never take a ramp test seriously again. Run the 20-minute test, and replace the 0.95 multiplier with P20 − W'/1200 using the W’ you just calculated. For the cross rider above, that’s 322 − 21.7 = 300 W instead of the ramp’s 317 W.

Below 45 J/W, you have the opposite problem: the ramp robs you a little, the 0.95 rule robs you slightly less, and both understate what you can grind out. Use P20 − W'/1200 again, which for the diesel gives 313 − 12.5 = 300 W rather than 297 W. Diesels are also the riders for whom a 40-minute or full-course TT effort is the most honest test that exists, and UK club racing hands you one every Tuesday evening. If you ride a 10 in 24 minutes flat out, you don’t need a test protocol. You have the file.

Can’t pace? The 3-minute all-out test is the honest answer for riders who consistently blow up at minute six of a 20. Go from a rolling start, hold nothing back from the first pedal stroke, and take the mean of the final 30 seconds as CP. It is genuinely horrible and the result is more repeatable than a badly paced 20. The full derivation of why end-test power converges on CP, and how the two-parameter model relates to the three-parameter and Xert variants, is covered in our FTP estimation and fitness models pillar.

What to hand an LLM, and what it will get wrong

Every general-purpose model will cheerfully apply 95% of your 20-minute power without asking a single question about your power curve. That is the failure mode. Give it the phenotype data explicitly:

I'm a self-coached road/TT rider. From intervals.icu, best efforts in the
last 6 weeks (all fresh, all maximal, all on the turbo):

  1 min  445 W
  3 min  395 W
  5 min  358 W
  12 min 318 W
  20 min 307 W

1. Solve the 2-parameter CP model from the 3- and 12-min values.
   Show W' in kJ and CP in W.
2. Report W'/CP in J/W and tell me whether I'm above or below 63.
3. Compute what a Zwift 20 W/min ramp test would predict for me using
   peak = CP + sqrt(W' x r / 30), best 1-min = peak minus 10 W, x 0.75.
   State the error vs CP in W and %.
4. Recommend one protocol and the exact multiplier I should use, and
   tell me which of my listed efforts are contaminating the estimate.
Do not use 0.95 unless my W'/CP justifies it.

Two things to police in the answer. Models routinely fit CP to a 1-minute or 5-minute effort, which sits inside the anaerobic region and inflates W’ badly, so pin the durations yourself. And they will treat every number you paste as maximal: a paced 20-minute effort from a group ride will quietly drag CP down, while a 3-minute PB set during a race with a tow will drag it up.

Set your new FTP, then schedule the only verification that matters. Four weeks out, ride 45 minutes as hard as you can sustain on a road you know. If your average sits within 2% of the number in your app, the protocol you chose was the right one for the rider you actually are. If it comes in 15 W low, you already know which test lied to you, and you know why.