Analysing Ride Data With LLMs
If you want to analyse Strava data with AI and get something better than a horoscope, the hard part isn’t the model. It’s the data handoff. Every large language model will happily tell you that your 3-hour endurance ride shows “good aerobic consistency” without ever having seen a single second of your power file. The difference between a useful analysis and expensive flattery comes down to what you put in the context window, what you compute outside it, and how tightly you constrain what the model is allowed to claim.
This page covers the parts that hold up when you check them: which tools genuinely read your files, where LLMs fail arithmetically, the prompt structures that survive scrutiny, and how to validate an AI’s reading of a ride against numbers you already trust from WKO5, intervals.icu or Golden Cheetah.
Why Uploading A FIT File Usually Fails
A one-hour ride recorded at 1 Hz from a Garmin Edge 840 with power, heart rate, cadence, speed, altitude, and left/right balance gives you roughly 3,600 records with 8-12 fields each. Export that to CSV and you’re looking at 250-400 KB of text, which tokenises to something like 90,000-150,000 tokens. A four-hour gravel ride triples that. A week of training is a quarter of a million tokens of mostly redundant numbers.
Claude and Gemini will accept that. GPT-5 with a large context will accept it too. What none of them do reliably is arithmetic across 14,000 rows. Ask for normalised power and you’ll get a number that looks plausible and is wrong, because computing NP requires a 30-second rolling average, raising each value to the fourth power, averaging those, and taking the fourth root. Transformers do not do that operation faithfully over long sequences. They pattern-match to a number in the neighbourhood.
Here’s the failure mode, from a real 2×20 threshold session (FTP 285 W):
Actual (intervals.icu): NP 268 W, IF 0.94, TSS 96, AP 254 W
Claude, raw CSV in prompt: NP 271 W, IF 0.95, TSS 94 (close, still guessed)
GPT-5, raw CSV in prompt: NP 259 W, IF 0.91, TSS 88 (5% low, cascades)
Either, with code execution: NP 268 W, IF 0.94, TSS 96 (correct)
The third row is the point. When the model writes and executes Python instead of reasoning in tokens, the numbers become exact. That single architectural choice separates tools that work from tools that produce confident fiction, and it’s the reason the ChatGPT FIT file workflow insists on code execution rather than direct file reading.
What Each Tool Actually Does With Your Data
The marketing language across this category is close to useless. Here is what the tools do mechanically, as of autumn 2026.
| Tool | Reads FIT directly | Computes or estimates | Model |
|---|---|---|---|
| ChatGPT (Plus/Pro, code interpreter) | Yes, via fitparse or fitdecode in sandbox | Computes in Python | GPT-5 family |
| Claude (Pro/Max, analysis tool) | CSV/JSON yes, FIT needs conversion | Computes in JS sandbox | Opus 5 / Sonnet 5 |
| Gemini Advanced | CSV yes, FIT unreliable | Mixed, often estimates | Gemini 2.5/3 Pro |
| intervals.icu (no LLM) | Yes | Computes, exactly | n/a |
| Athletica.ai | Via Strava/Garmin sync | Computes, own models | Proprietary + LLM layer |
| TrainerRoad AI FTP Detection | Internal | Computes | ML, not an LLM |
| Strava Athlete Intelligence | Internal | Summarises computed metrics | Undisclosed LLM |
Two things worth separating. Athletica and TrainerRoad are sports-science products with machine learning inside; the LLM, where present, sits on top as a narrator. ChatGPT and Claude are general reasoners with a Python sandbox, which makes them far more flexible and entirely dependent on you for the physiology.
Strava’s Athlete Intelligence deserves a specific note because so many people encounter it first. It reads metrics Strava already computed and writes a paragraph. On a ride where I deliberately soft-pedalled the final 20 minutes after a hard 90, it told me the effort was “well paced with strong consistency throughout.” It had the decoupling data available and didn’t use it. Treat it as a caption, not an analysis.
The Preprocessing Layer That Makes Everything Else Work
Stop sending raw time series. Send a summary that’s dense in information and cheap in tokens. For a single ride, this structure runs to about 600 tokens and supports more real analysis than 100,000 tokens of raw records:
{
"ride": "2026-09-22 Tuesday threshold",
"duration_s": 4920, "distance_km": 52.3,
"ftp_w": 285, "lthr_bpm": 168, "weight_kg": 71,
"summary": {"ap_w": 213, "np_w": 241, "if": 0.85, "tss": 98,
"hr_avg": 149, "hr_max": 176, "cad_avg": 88,
"kj": 1048, "elev_m": 612, "vi": 1.13},
"power_curve_w": {"5s": 742, "15s": 610, "30s": 498, "1m": 412,
"5m": 321, "8m": 306, "20m": 291, "60m": 258},
"intervals": [
{"n": 1, "dur_s": 1200, "ap_w": 289, "np_w": 292, "hr_avg": 161,
"hr_end": 167, "cad": 91, "pct_ftp": 1.01},
{"n": 2, "dur_s": 1200, "ap_w": 283, "np_w": 288, "hr_avg": 166,
"hr_end": 172, "cad": 87, "pct_ftp": 0.99}
],
"decoupling_pct": 4.1,
"time_in_zone_s": {"z1": 1180, "z2": 900, "z3": 420,
"z4": 2280, "z5": 140, "z6": 0, "z7": 0},
"context": {"ctl": 78, "atl": 94, "tsb": -16,
"temp_c": 19, "sleep_h": 6.2, "days_since_rest": 9}
}
Ninety-nine percent compression and almost no loss for the questions you actually want answered. You can get most of these fields straight out of intervals.icu’s activity API, or compute them once with a local script. The context block matters more than people expect: a model shown TSB of -16 and nine days without rest reads the same interval data very differently than one shown TSB of +4.
For multi-week questions, aggregate further. Per-ride rows with date, TSS, duration, IF, and zone distribution let you put three months into 3,000 tokens, which means the model can see your actual periodisation rather than one ride in isolation.
Prompts That Produce Checkable Answers
The generic prompt produces generic output. “Analyse my ride” gets you a restatement of your own numbers with adjectives attached. What works is forcing the model into a specific analytical frame, giving it the physiological rule you want applied, and requiring it to show the calculation.
Compare these two on the same file. First attempt:
Here’s my ride data. How did I do?
Response, paraphrased: strong session, good power, heart rate response suggests solid fitness, consider more recovery. Nothing here needed the file.
Second attempt:
You are analysing a threshold session. FTP 285 W, LTHR 168 bpm, set 3 weeks ago.
Do these four things and nothing else:
- For each interval, state mean power as % of FTP and mean HR as % of LTHR.
- Compute HR drift within each interval: mean HR of final 5 min minus mean HR of first 5 min, at matched power (±5 W). Report in bpm.
- Compare interval 2 to interval 1. Quantify the fade in watts and in bpm.
- Based only on 1-3, state whether 285 W is still the right FTP. If the evidence is insufficient, say so and name the test that would settle it.
Show your arithmetic. Do not give training advice.
That produced: interval 1 at 101% FTP with HR drift of 6 bpm, interval 2 at 99% FTP with drift of 9 bpm and mean HR 5 bpm higher at 6 W lower power. Conclusion: FTP is roughly correct but likely at the ceiling, since the second effort required disproportionate cardiac cost; a 20-minute maximal test or a 40-minute TT would confirm. Every claim traceable to a number in the file.
Four constraints do the heavy lifting. Name the session type so the model applies the right frame. Give it the physiological rule explicitly rather than hoping it knows. Demand arithmetic, which forces the code path. Add a falsification clause, because “if the evidence is insufficient, say so” is the single highest-value sentence in cycling prompts. Without it, models confabulate a conclusion from any input.
Five Analyses Where LLMs Genuinely Beat A Dashboard
Charts are better than prose for trends. LLMs win where the question needs cross-referencing, pattern description, or synthesis across sources that no single tool holds together.
Pacing forensics on a TT or hill climb. Split a 40 km TT into 5 km segments with power, speed, gradient and yaw-relevant heading, then ask which segments cost time relative to an even-pace model at the same NP. On my own 25-mile TT, the analysis identified that I was 11 W over target in km 3-8 and 14 W under in km 30-38, costing an estimated 38-52 seconds. Both ends of that range came from the model showing its speed-power assumptions, which I could then check.
Durability, properly framed. Ask for the 5-minute power curve computed separately for the first and last 1,000 kJ of every ride over 2.5 hours in the last 90 days. The answer for me: 321 W fresh, 289 W after 1,000 kJ, a 10% drop. That’s a real number I can track, and it’s the metric that decides whether a 4-hour gravel race goes well.
Cross-source contradiction hunting. Feed Oura or Whoop sleep and HRV alongside your training log. The model found that my four highest-TSS days in a 12-week block all fell on days following under 6.5 hours of sleep, which was a scheduling artefact I hadn’t noticed: Tuesday intervals after Monday late finishes.
Plan critique with a named method. Paste next week’s plan, give the model your CTL, ATL and TSB, and ask it to evaluate against a specific framework (polarised, Seiler’s 80/20, or a Coggan-style build) with the intensity distribution calculated as percentage of time, not sessions. It caught that my “polarised” week was 68/32 by time, not 80/20, because two supposed endurance rides drifted into tempo.
Race file debriefs. A 4-hour road race file with 200-plus surges over 400 W is genuinely hard to read on a chart. Ask for surge counts bucketed by duration and magnitude, plus where in the race they clustered, and you learn something about positioning. Mine: 43 surges over 500 W, 31 of them in the final 45 minutes, which told me I was riding too far back and closing gaps rather than holding a wheel.
Where This Breaks, And How To Catch It
Every analysis you get should be checkable against something. Five recurring failures, with the check for each.
Confident wrong arithmetic is the most common. Ask the same question twice in fresh sessions; if NP comes back as 268 then 259, neither was computed. Anything with a rolling window, a fourth power or a variable-slope calculation belongs in code, not in prose.
Zone inference without your zones is the second. Models default to a generic seven-zone Coggan model and will place 250 W in “zone 3” for an athlete with a 220 W FTP. State your zone boundaries in watts, every time, in the prompt.
Physiologically hollow explanations are harder to spot because they read well. “Your heart rate drift indicates inadequate aerobic base” is the sort of sentence that survives no follow-up. Push back with “what specific mechanism, and what would the data look like if that were false?” A model with nothing behind the claim will either soften it or contradict itself.
Sycophancy is real and measurable. Tell a model your FTP is 320 W on a file where you fade below 290 W in the second of two 20-minute efforts and many will accommodate you. Counter it by asking for the case against your own interpretation: “argue that my FTP is overstated, using only this file.”
Missing context invites invention. A model that doesn’t know you rode in 31°C heat will attribute your elevated heart rate to fatigue or declining fitness. Temperature, altitude, illness, sleep, and time since last rest day should all be in the context block. Without them the model fills the gap with a story.
One more discipline worth building: keep a small validation set. Three or four rides where you know the answers cold, computed in WKO5 or intervals.icu. Run any new tool, model or prompt against them first. It takes ten minutes and it has saved me from acting on a fortnight of plausible nonsense more than once.
A Workflow You Can Run This Week
Pick one focused question. Not “how’s my training”, but “am I recovering adequately between my Tuesday and Thursday intensity sessions?” Specific questions get answers you can act on or reject.
Export the relevant files. From intervals.icu, the activity list CSV plus per-activity summaries covers most multi-week questions. For a single-ride deep dive, pull the FIT and use the code-execution route described in the ChatGPT FIT file workflow, which handles parsing, interval detection and the rolling-window maths that models get wrong unaided.
Write the prompt with all five elements: your physiological constants (FTP, LTHR, weight, zone boundaries in watts), the question, the method you want applied, a requirement to show arithmetic, and permission to conclude that the data can’t answer it. Save the prompt. You’ll reuse it weekly, and a stable prompt means comparable answers over time.
Check three numbers by hand before you trust anything else in the response. If average power, duration and TSS are right, the parsing worked. If they’re off by more than 1%, the whole analysis is suspect regardless of how well-written it is.
Then interrogate the conclusion. Ask for the opposite case. Ask what data would change the answer. A model that can’t articulate what would falsify its own claim has told you something about the claim’s evidential basis.
Where does this leave the actual coaching decision? With you, which is the point of being self-coached. The model is a fast, tireless, arithmetically fragile analyst who has read every training text and has no idea how your legs felt on Thursday. Used as a way to interrogate your own files faster than a dashboard allows, it’s genuinely valuable. Used as an oracle, it will tell you what you want to hear at 3 a.m. with total conviction and no evidence at all, and your FTP will not move.
In this section
The supporting pages under this subject.