AI Cycling
009 Analysing Ride Data With LLMs 2,013 words · 9 min

A Prompt Library For Single-Ride Debriefs That Beats “Analyse This Ride”

Paste a ride summary into ChatGPT, type “analyse this ride”, and you get something like this back:

Great effort! Your normalised power of 243 W and average heart rate of
147 bpm show solid aerobic endurance. The intensity factor of 0.85
suggests a good tempo effort. Your cadence of 88 rpm is in the optimal
range. Consider adding more variety to your training and remember to
recover well. Keep up the great work!

Every number in that reply came from the file. Not one of them was judged against anything. The model has no idea whether you were meant to be riding at 243 W or at 190 W, whether rep four was a triumph or a capitulation, or whether the 147 bpm was cheap or expensive. It is describing your ride back to you in a friendly voice, which is the single most useless thing a coach can do.

The fix is not a better model. It is three inputs the model cannot possibly infer from a FIT file: what you intended to do, what your targets were in absolute watts, and how hard it felt. Give it those and the same model that produced the fluff above will tell you that you started rep one 13 W too hot, paid for it in rep four, and should move Thursday’s session to Friday. That is a chatgpt prompt to analyse a cycling ride properly, and it is mostly a structured-input problem rather than a clever-wording one.

Why the three missing inputs matter so much

Intent is the big one. A 2h30 ride at 0.72 IF is a perfect endurance ride or a badly blown tempo ride, and nothing in the data distinguishes them. Without your intended session, the model defaults to the safest interpretation, which is always flattery.

Zones have to be given in watts, not names. “Zone 2” means 56-75% of FTP under Coggan, roughly everything below the first ventilatory threshold under Seiler’s three-zone model, and something else again in the Zwift app. If you write “was I in zone 2?”, the model picks a definition and you never find out which. Write “target 175-205 W” and the ambiguity disappears.

RPE is the input that catches what power cannot. Power is an output measure and it lies about cost. 4 x 8 min at 305 W on fresh legs at RPE 7 and the same session at RPE 10 look identical in the file and mean completely different things for next week. Give the model both and it can flag the divergence.

There is a fourth input worth adding once you get serious: how your FTP was set and when. An intervals.icu eFTP of 302 W derived from a 5-minute effort in a chaingang is not the same number as 285 W from a 20-minute test on a turbo, and if the model anchors percentages to the wrong one, every verdict downstream is wrong by 6%.

Getting the data in without wrecking the analysis

Do not paste a 1 Hz CSV of a three-hour ride. That is roughly 10,800 rows, and even where the context window swallows it, arithmetic accuracy across that many values falls apart in ways you will not notice.

Paste the interval table instead. In intervals.icu, the Intervals tab on any activity gives you a copyable grid of each rep with average power, normalised power, HR, cadence and duration. That is 8 to 20 rows of already-aggregated numbers the model can actually reason over. For steady rides, split into 20-minute or 30-minute blocks and paste those. The broader question of which formats survive the trip into a language model, and which silently degrade, is covered in Analysing Ride Data With LLMs.

Always includeInclude when relevantLeave out
FTP in watts + date + test methodW’ / anaerobic capacity if you use itRaw 1 Hz streams
Target ranges in wattsCourse profile or gradient blocksGPS coordinates
Interval table (AP, NP, HR, duration)Wind and temperatureStrava kudos, segment PRs
RPE per rep or per blockSleep, illness, days since last hard dayCalories, “fitness” scores

Template 1: The Execution Verdict

This is the one to save first. It answers a single question: did I do the session I set out to do?

ROLE: You are my cycling coach. Be blunt. Do not praise anything
that was not executed well.

ATHLETE
FTP 285 W (20-min test x 0.95, 11 Sep 2026). 74 kg. Threshold HR 168.

INTENDED SESSION
4 x 8 min @ 299-314 W (105-110% FTP), 4 min recovery below 180 W.
Week 3 of 4 in a threshold build block.

HOW IT FELT
RPE per rep (1-10): 8, 8, 9, 10. Legs flat from the first rep.
5h40 sleep. Third hard session in six days.

WHAT ACTUALLY HAPPENED
[paste interval table]

TASK
1. For each rep, state HIT / PARTIAL / MISSED against the target range,
   with the actual average and the gap in watts.
2. Identify where in the session execution broke down, citing rep number
   and the number.
3. State whether the RPE pattern is consistent with the power pattern.
   If not, say which one you trust and why.
4. One sentence: what should have been done differently at the time.

CONSTRAINTS
- Absolute watts only. Never use zone names.
- If a question cannot be answered from the data given, write
  "insufficient data" and say what is missing.
- Finish with exactly one verdict from: EXECUTED AS WRITTEN /
  EXECUTED WITH DRIFT / CORRECTLY ABANDONED / OVERCOOKED / UNDERCOOKED.

Fed the reps 312, 308, 299, 281 W, a good model returns something closer to this:

Rep 1: 312 W  PARTIAL (in range but top end, +13 W over the 299 floor)
Rep 2: 308 W  HIT
Rep 3: 299 W  PARTIAL (bottom of range, HR 171 vs 164 in rep 1)
Rep 4: 281 W  MISSED (-18 W below target, RPE 10)

Breakdown point: rep 3. Power held at the floor only by pushing HR 7 bpm
above rep 1 for the same work. Rep 4 was the bill arriving.

RPE vs power: consistent. 8/8/9/10 against a 31 W decline from rep 1 to
rep 4 is a fatigue curve, not a pacing error in isolation, but rep 1 at
the top of the range accelerated it.

Should have: opened rep 1 at 302-305 W. On 5h40 sleep and day 6 of a
hard block, three reps at target beats four with one failed.

VERDICT: OVERCOOKED

That is a debrief. You can act on it tomorrow.

Template 2: The Fade Audit

For long endurance and gravel rides, where the question is durability rather than compliance. Split the ride into halves or thirds before pasting, and give the model the Pw:HR figures if your platform computes them (intervals.icu shows decoupling on the activity page).

INTENDED: 3h15 endurance, 175-205 W, no efforts above 260 W.
RPE: 4 for the first two hours, 6 in the last 45 min.
FUELLING: 62 g carbs/hour, 500 ml/hour.

DATA (30-min blocks): [paste]
Decoupling (Pw:HR): [paste]

TASK
1. Was the ride aerobically steady, or did it become a tempo ride by
   accident? Quantify with the block averages.
2. Compute or interpret the power:HR drift between first and second half.
   State the W/bpm for each half.
3. Attribute the drift: fatigue, heat, fuelling, or terrain. Say which
   evidence supports your attribution and which would contradict it.
4. Does the ride support or contradict the aim of building durability?

Answer in under 200 words. No encouragement.

With 232 W at 141 bpm in the first half and 228 W at 149 bpm in the second, the model will give you 1.645 vs 1.530 W/bpm, a 7.0% decoupling, and then argue about the cause using the fuelling line you supplied. That last part is why the fuelling goes in the prompt rather than being left out as irrelevant detail.

Template 3: The TT Pacing Post-Mortem

Time trials are the easiest ride type to analyse well, because the intent is unambiguous and the output is a single number you already care about.

EVENT: CTT 10-mile, sporting course, 22:14. FTP 285 W.
DATA: AP 297 W, NP 301 W, VI 1.013, IF 1.056, avg HR 173, max 181.
SPLITS: miles 1-5 avg 306 W. Miles 6-10 avg 288 W.
RPE: 9 at halfway, 10 from mile 7.
CONDITIONS: 12 mph crosswind, out-and-back, two roundabouts.

TASK
1. Grade the pacing: positive split, negative split, or even. Give the
   percentage difference and say whether it is within noise for a course
   with two roundabouts.
2. Was 306 W for the first five miles defensible given FTP 285 and a
   22-minute duration? Reference the expected sustainable percentage.
3. Estimate the time cost of the split pattern versus even pacing, and
   state your uncertainty in seconds.
4. One change for the next 10.

The 6.2% positive split gets named, the opening 107% of FTP gets challenged, and you get a specific instruction (hold 296-300 W to the first roundabout) instead of “pace yourself better”.

Template 4: The RPE Mismatch Flag

Run this when something felt wrong and the file looks fine, or vice versa. It is a diagnostic prompt, not an evaluation one.

CONTEXT: this session felt much harder than the power suggests.
Expected RPE 7. Actual RPE 9.

SESSION: [paste]
LAST 7 DAYS: [paste weekly TSS or hours, and RPE per session]
NON-TRAINING: sleep hours, resting HR, illness, alcohol, work stress.

TASK
1. List every plausible explanation for the RPE/power gap, ranked by how
   well the data supports each.
2. For each, name the one piece of data that would confirm or kill it.
3. Say which explanations you cannot test with what I have given you.
4. Do not diagnose overtraining. Flag patterns and tell me what to watch.

The ranked-hypotheses structure is the important bit. Ask an open question and you get one confident answer, usually the most dramatic one available.

Template 5: The Next 72 Hours

The debrief that changes a decision. Chain this after Template 1, in the same conversation, so it inherits the verdict.

Given the verdict above, and this plan for the next three days:
Tue: 4 x 8 threshold (done, see above)
Thu: 5 x 5 @ 320-335 W
Sat: 4h endurance with 3 x 20 tempo
Sun: club run, ~2h30, uncontrolled intensity

TASK
1. Should Thursday go ahead as written? If not, give the replacement
   session in watts and durations.
2. Which of the four days is most at risk of being junk?
3. If I can only complete two of the remaining three, which two?
4. Give your answer as a table. One line of reasoning per row, maximum.

Forcing the table matters more than it sounds. Left unconstrained, models hedge across three paragraphs and you end up making the call yourself, which was the thing you were trying to avoid.

Two habits that raise the hit rate

Keep the athlete block in a text file and paste it at the top of every conversation. FTP, weight, threshold HR, test date, current block, known limiters. Updating it every six weeks takes a minute and removes the most common source of nonsense output.

The other habit: make the model commit. A fixed verdict list, a required number, a word limit, a table. Models trained to be agreeable will find a way to be agreeable in any gap you leave open, and a bounded answer format is the cheapest way to close the gaps. “Say insufficient data” earns its place in the prompt for the same reason.

Pull up last Tuesday’s session tonight, run it through Template 1 with your honest RPE, and see whether the verdict matches what you told yourself at the time. That gap is the whole reason to do this.