AI Cycling
003 Analysing Ride Data With LLMs 1,680 words · 8 min

What An LLM Actually Sees When You Paste A Ride File

Ask ten self-coached riders how they use AI on their training data and eight will describe the same ritual: finish the ride, open ChatGPT, upload the FIT file, type “analyse this”. The reply comes back fluent, structured, confident. Average power, a comment on your aerobic decoupling, a suggestion about next week. It reads like coaching.

Most of it is invented.

Not maliciously, and not because the model is stupid. It’s because the thing you handed over is roughly 11,000 rows of numbers, and almost nothing in the pipeline between your Wahoo and the model’s attention is designed to carry 11,000 rows intact. Somewhere in that chain it got truncated, sampled, silently re-read by a Python script you never saw, or turned into five example rows and a column list. The prose that came back was generated from whatever survived.

A three hour ride is bigger than it feels

Garmin, Wahoo and Zwift all default to 1 Hz recording on a modern head unit. That’s one row per second.

RideRecordingRows
1h Zwift session1 Hz3,600
2h club run1 Hz7,200
3h endurance ride1 Hz10,800
3h ride, Garmin “smart” recordingvariable~2,700
5h Dragon Ride day1 Hz18,000

Each of those rows carries a dozen or more fields: timestamp, latitude, longitude, altitude, distance, speed, power, heart rate, cadence, temperature, left/right balance, torque effectiveness, and whatever else your pedals volunteer. Fourteen fields across 10,800 rows is about 151,000 individual values. Exported as CSV at roughly 75 bytes per line, that’s an 800 KB text file.

Now the part that matters. Numeric CSV tokenises badly, far worse than English prose, because digit runs and commas fragment. A line like 2026-09-14T10:41:07Z,243,146,89,8.41,112,51.45231,-2.58734,19 lands somewhere around 25 to 35 tokens. Multiply out: your “quick ride file” is 270,000 to 380,000 tokens. That’s not a rounding error against a context window, it’s the whole meal.

What actually happens when you upload a FIT file to ChatGPT

Three different things, depending on how you did it, and the model will not tell you which.

You pasted CSV into the chat box. Both ChatGPT and Claude.ai silently convert long pastes into an attachment above the input field. Feels tidy. What you’ve actually done is hand the text to a file-handling path instead of into the conversation, and the model now reasons about a file rather than reading your data.

You uploaded a .csv. ChatGPT routes this to its Python sandbox (the tool formerly branded Advanced Data Analysis). The model does not read your 10,800 rows. It writes pandas code, runs it, and sees only what that code prints. Done well this is the best outcome available. Done carelessly it means the model saw df.head(), five rows, and started writing.

You uploaded a .fit. FIT is a binary format. The extension is frequently rejected outright as unsupported, and riders work around it by zipping the file or renaming it to .txt. Having done that, the model still needs to decode it, and the sandbox has no internet access, so pip install fitparse fails. What follows is a coin flip: sometimes it hand-rolls a FIT decoder that half works, sometimes it produces a plausible-looking table of numbers that came from nowhere in particular.

The df.head() illusion

Here is what the model sees on a good day, before it says a single word to you:

>>> df.shape
(10864, 14)
>>> df.head()
             timestamp  power  heart_rate  cadence  altitude
0  2026-09-14 08:57:12    142         104       71     58.2
1  2026-09-14 08:57:13    151         106       73     58.2
2  2026-09-14 08:57:14    149         107       74     58.4
3  2026-09-14 08:57:15    163         109       76     58.6
4  2026-09-14 08:57:16    158         111       77     58.8

Five rows from your warm-up in a Somerset lane. That’s the evidential basis for any sentence about your fourth threshold interval unless the model went and computed something specific, and you have no way of knowing whether it did.

There’s a test that settles it in one message. Before asking for any analysis, ask for three facts: the exact number of rows, the first and last timestamps, and the sum of the power column. A model working from real data answers in seconds with numbers that match your head unit. A model working from a truncated paste tells you the ride ended at 09:44, or produces a row count that isn’t 10,864, or explains that it can give you a qualitative overview instead. I’ve had that last answer arrive immediately after two paragraphs of confident interval-by-interval commentary.

Collapse 11,000 rows into 9

The fix is not a better prompt. It’s doing the aggregation yourself, locally, before anything reaches the model. Fifteen lines of Python with python-fitparse:

import fitparse
import pandas as pd

fit = fitparse.FitFile("2026-09-14-085712-elemnt.fit")

df = pd.DataFrame([{f.name: f.value for f in m}
                   for m in fit.get_messages("record")])
laps = pd.DataFrame([{f.name: f.value for f in m}
                     for m in fit.get_messages("lap")])

df["lap"] = pd.cut(df.timestamp,
                   bins=[df.timestamp.min()] + list(laps.timestamp),
                   labels=range(len(laps)))

summary = df.groupby("lap", observed=True).agg(
    secs=("power", "size"),
    avg_p=("power", "mean"),
    np=("power", lambda s: (s.rolling(30).mean().dropna() ** 4).mean() ** 0.25),
    max_p=("power", "max"),
    avg_hr=("heart_rate", "mean"),
    max_hr=("heart_rate", "max"),
    cad=("cadence", "mean"),
    climb=("altitude", lambda s: s.diff().clip(lower=0).sum()),
).round(0)

print(summary.to_markdown())

Output, from a real shape of session (4 × 10 min at threshold, 5 min recoveries, ride home):

lapsecsavg_pnpmax_pavg_hrmax_hrcadclimb
warm-up16801681814021281498496
effort 16002812833411631719148
recovery3001421582881411637912
effort 26002792813521681749051
recovery300139151241144165779
effort 36002742783661711778947
recovery3001331492621471687514
effort 46002622693811731798644
ride home588415417251213216881388

Nine rows. Around 450 tokens including headers. That is a 600:1 compression against the raw dump, and it is not lossy in any way a coach would care about, because every number a coach would actually use is still there.

Look at what those nine rows make unmissable. Power falls 281, 279, 274, 262 while heart rate climbs 163, 168, 171, 173. Effort four is 6.8% down on effort one for 10 bpm more cost, and cadence has dropped five rpm. That’s the session’s whole story: you went slightly too hard on the first two and paid for it, or you arrived underfuelled. An LLM given this table will find that pattern reliably, because comparing eight numbers is exactly what it’s good at. An LLM given 10,800 rows it can’t see will tell you about aerobic decoupling in general terms and pick a percentage that sounds right.

Where raw streams still earn their keep

Aggregate by default, not always. A 45 second sprint at 1 Hz is 45 rows and you should send every one of them, because the shape of the first 8 seconds is the point. Same for a single 20 minute test: 1,200 rows is fine, and a 5 second rolling mean brings it to 240 with nothing lost. For a whole ride when you genuinely want the shape, df.iloc[::10] gives a 10 second downsample: 1,086 rows, around 30,000 tokens, plottable and still cheap.

The mean maximal power curve deserves a mention too. Five numbers pulled from intervals.icu or WKO5 (best 5s, 1min, 5min, 20min, 60min for the ride, alongside your 90 day bests) tell a model more about what happened than the entire file does. Two lines of text, one clear verdict.

It’s worth noticing that the AI features shipping inside the platforms work this way already. Strava’s Athlete Intelligence and Garmin’s newer summary features don’t hand a language model your stream data, they hand it the derived metrics the platform has always computed. That’s not a limitation they’re working around. It’s the design. If you’re building out a fuller workflow, the same principle scales up, and the broader patterns are covered in our guide to analysing ride data with LLMs.

The prompt that goes with the table

Pre-aggregating changes what you should ask for, because you no longer need to beg the model to look properly.

Here is a lap summary from today's session. FTP 268 W, weight 71 kg,
threshold HR ~172, third week of a build block, 9h last week.

[paste the 9-row table]

1. Did I hit the target (4 x 10 min at 95-100% FTP)?
2. Quantify the fade across efforts and tell me the most likely cause,
   given the HR and cadence trend.
3. Was the 98 min ride home a problem for tomorrow's session?

Cite the specific numbers you used for each answer. If a number you
need isn't in the table, say so rather than estimating it.

That last sentence does more work than everything else combined. It converts a model that fills gaps into one that reports them, which is the difference between a tool and a very articulate liar.

Try the row-count test on your last ride file tonight, before you change anything else about how you work. Ask for the exact row count, first and last timestamp, and the sum of the power column. Whatever comes back will tell you how much of last month’s analysis was actually about your riding.