AI Cycling
010 AI Coaching Platforms, Tested 1,780 words · 8 min

TrainerRoad Adaptive Training: Six Weeks Of Watching It Change My Plan

I spent six weeks logging every single change TrainerRoad’s Adaptive Training made to my calendar. Not the marketing version of what it does, the actual diff: which workout was scheduled Monday, what it had become by Thursday, and why. Forty-one adjustments in forty-two days, tracked in a spreadsheet next to my intervals.icu compliance data and my Garmin HRV numbers.

The verdict up front, because this is a review and you came here for one: Adaptive Training’s progression logic is the best-behaved automated coach I’ve tested for riders training eight to ten hours a week, and it actively holds back riders doing fourteen or more. It is tuned to protect you from yourself. If you need protecting, that’s excellent. If you’re already a disciplined high-volume rider with a good base, you’ll spend your season fighting the ceiling it puts over your head.

The setup, so you can judge the sample

I’m 41, UK-based, racing regional TTs and doing a lot of gravel. FTP at the start of the block: 312 W (4.28 W/kg at 73 kg). I ran TrainerRoad’s Sweet Spot Base Mid Volume I into Sustained Power Build Mid Volume, deliberately picking mid volume so I could then push extra unstructured hours on top and watch how the system reacted. Power meter is a Favero Assioma Duo on the road bike, Wahoo Kickr Core indoors, all files double-syncing to intervals.icu so I had a second opinion on TSS and a proper eFTP estimate to check against TrainerRoad’s own AI FTP Detection.

Post-workout Survey answers were honest. That matters more than people admit: Adaptive Training is largely driven by those surveys plus its own analysis of how your power tracked the target, and if you tell it everything was “Moderate” out of vanity you’re just lying to a database.

The log: what actually changed

Here’s the shape of the six weeks in numbers.

WeekPlanned TSSActual TSSAT adjustmentsDirection
141246843 up, 1 down
243852175 up, 2 down
345561094 up, 5 down
4301 (recovery)38952 up, 3 down
547065583 up, 5 down
649258884 up, 4 down

Twenty-one increases, twenty decreases. Near perfect balance, which sounds like good control until you notice my actual TSS ran 19% above plan across the block because of the outdoor riding AT wasn’t counting properly. It was rebalancing a plan that had stopped describing my training.

Adjustment 1: the first “Productive” bump

Week 1, Tuesday. Scheduled Tunemah -2 (Sweet Spot, 3x12 min at 90% FTP). I did it at an average of 284 W against a 281 W target, all intervals held, surveyed it “Moderate”. Wednesday morning the calendar had changed: Thursday’s Antelope went from Level 4.4 to Level 5.1, and Saturday’s Tallac was replaced with Tallac +2, adding 6 minutes of work at threshold.

That’s a good adjustment. Fast, proportionate, driven by real evidence. Progression Levels moved Sweet Spot from 4.3 to 5.0. Note how granular that is: half a level, roughly 4-6% more work at the same intensity. This is the core of what TrainerRoad does well and what a generic LLM coaching prompt still can’t match, because the levels are anchored to a workout library with known ratings rather than a model guessing at “make it a bit harder.”

Adjustment 2: the honest downgrade

Week 3, Thursday. Scheduled Kaiser (VO2max, 5x3 min at 120%). I’d done four hours of hilly gravel on Wednesday evening, off-plan, and I bailed on interval four at 2:10. Surveyed “Very Hard”, flagged the failed interval.

Adaptive Training’s response that night: VO2max Progression Level dropped 6.1 to 5.4, Saturday’s Spencer +2 became plain Spencer, and it inserted a note that the following Tuesday’s VO2 session was now easier. Total reduction across the next seven days: about 38 TSS.

Correct call in isolation. Except the cause wasn’t that my VO2max capacity had regressed, it was that I’d smashed my legs the day before with a ride the system barely registered. AT reduced my quality work because of fatigue from volume work it had no structured view of. Do that repeatedly and you get a slow downward drift in your hard sessions while your easy hours climb. That’s the underloading mechanism, and it’s structural, not a bug.

Adjustment 3: AI FTP Detection undershooting

Week 4, I ran AI FTP Detection. Result: 318 W, up 6 W from 312.

intervals.icu’s eFTP over the same window, computed from my power duration curve: 329 W. A 20-minute all-out test on the Bealach-style climb near me two days later gave 341 W for 20 minutes, which puts a conventional 95% estimate at 324 W and my actual sustainable hour power realistically around 326-330 W.

So TrainerRoad’s number was 8-11 W light. On a 3x20 at 95% FTP session that’s the difference between 302 W and 313 W per interval: one is a solid tempo-threshold effort, the other is the session I actually needed. Across a build block, training 3% below your real threshold every single day is a meaningful amount of stimulus you simply never receive.

To be fair to the algorithm, it’s conservative by design, and it says so. AI FTP Detection is trained to produce a number you can complete workouts at, not your maximum. For someone who’s been burning out on over-inflated ramp test results, that’s a genuine kindness. For me it meant manually overriding to 326 W in week 5, after which everything felt correctly calibrated and, notably, Adaptive Training did not throw a fit. It just re-based the levels and carried on.

Adjustment 4: the outdoor ride problem

This is the one that decides the argument. Week 5, Sunday: 4h 48m gravel, 198 W normalised, 287 TSS, 2,100 m climbing. Not on the plan. Pushed from Garmin to TrainerRoad as an “unstructured ride”.

What Adaptive Training did with it:

Week 5 → Week 6 changes after 287 TSS unstructured ride
  Sweet Spot PL      5.9 → 5.9   (no change)
  Threshold PL       5.2 → 5.2   (no change)
  VO2max PL          5.4 → 5.4   (no change)
  Mon recovery       Taku       → Taku        (unchanged)
  Tue threshold      Mills +2   → Mills       (-14 TSS)
  Total week TSS     492        → 478

Fourteen TSS of acknowledgement for a five-hour ride. The Progression Levels didn’t move at all, because an unstructured ride doesn’t have interval structure to rate against the library. TrainerRoad will use outdoor rides to inform AI FTP Detection and it will nudge the next day’s session if you clearly dug a hole, but your progression is driven almost entirely by completed structured workouts.

For a time-crunched rider on six to eight hours where nearly everything is structured, that’s close to a complete picture. For a rider doing sixteen hours where nine of them are outdoor endurance and club runs, Adaptive Training is progressing you based on 40% of your training load. It will keep offering you Level 5.9 Sweet Spot in October when you have the base to be doing 7.5, because it never saw the evidence.

Where the conservatism is exactly right

I want to be specific about the upside, because it’s real and it’s the reason I’d still recommend this to most people reading.

Adaptive Training never once dug me a hole. Compare with what I’ve seen from the LLM-prompt approach, where you paste a fortnight of intervals.icu CSV into a model and ask for next week’s plan: it will happily write you four hard days, because nothing in the prompt encodes the empirical relationship between yesterday’s VO2 session and today’s repeatability. TrainerRoad’s ratings come from millions of completed workouts and a machine learning model that knows a Level 6.2 VO2 session after a Level 5.8 threshold session gets failed a lot.

The recovery week handling was also quietly good. Week 4 was scheduled at 301 TSS, I did 389, and it didn’t panic or try to claw the extra back. It let the week be a recovery week and adjusted the following Tuesday down slightly. A rules-based system would have tried to rebalance.

And the failure mode is the right one. When it’s wrong, it makes you slightly undertrained. Every other automated tool I’ve tested in this space, which I’ve written up properly in the AI coaching platforms comparison, errs toward giving you more when the signal is ambiguous. Slightly undertrained in June beats overcooked in April.

The workaround, if you’re high volume

Three things made this usable for me at 12-16 hours:

Override the FTP. Take AI FTP Detection’s number, compare it to your intervals.icu eFTP, and if the gap is more than 5 W trust the power duration curve. I run about 4% above what TrainerRoad suggests.

Log endurance rides as structured workouts. Tedious but effective: instead of a free ride, associate outdoor endurance with an actual TrainerRoad endurance workout and let it grade you. My Endurance PL finally started moving once I did this, from 4.1 to 6.3 in three weeks, and the plan stopped offering me 90-minute Petit as a “hard” Saturday.

Use Plan Builder’s volume setting one level above your instinct. I switched to High Volume in week 6 and immediately got sessions that matched what my legs were asking for. Mid Volume plus extra unstructured hours does not equal High Volume in the eyes of the algorithm, even if it equals it in your legs.

What I’d want in version next

A weighting for unstructured load in Progression Levels, even a crude one. If I do 287 TSS at 0.72 IF for nearly five hours, that is unambiguous evidence about my aerobic durability, and the system currently throws it away for progression purposes. Xert does attempt this and gets it wrong in a different, noisier direction, but at least it tries.

Six weeks, 41 adjustments, one FTP override and a lot of spreadsheet columns later: my threshold went from 312 to a tested 326 W, which is fine, and my confidence that any single day’s session was the right session went up more than the number did. That’s worth the £19.95 a month if your problem is consistency. If your problem is that you’ve got the hours and you need something to point them at, you’ll be doing what I did, quietly steering the thing from behind the wheel.