Reviewing A Whole Season With A Long-Context Model
Most people who try to analyse a year of Strava data start by pasting one ride into a chat window and asking what it means. The model says something agreeable about your normalised power and suggests you work on your endurance base. You close the tab. Nothing changes, because nothing was learned: you already knew that Tuesday’s intervals felt hard.
The interesting question is not “how was this ride”. It is “what has my training actually been, as opposed to what I think it has been”, and that question only has an answer at the scale of a season. I ran 312 rides from the 2025 season (11,400 km, 418 hours, FTP tested 268 W in January and 291 W in September) through a two-stage pipeline: a cheap model that compresses each ride into a fixed digest, then a long-context model that reads all 312 digests at once and looks for structure. The second stage is where everything useful happened. The first stage exists only to make the second stage possible.
The token arithmetic that forces the design
A three-hour ride sampled at 1 Hz is 10,800 rows. Keep power, heart rate, cadence, speed, altitude and temperature and you have roughly 65,000 numbers per ride. Across 312 rides that is somewhere north of four million rows of time series, which as CSV lands in the tens of millions of tokens. A 1M context window is large. It is not that large, and even if it were, dumping 1 Hz samples into a reasoning model is a waste of attention: you are asking it to do arithmetic it is bad at on data it cannot hold in working memory.
So you summarise first, then reason. Compress each ride to a dense, fixed-schema digest of about 250 tokens, and 312 rides becomes ~78,000 tokens. That fits in Sonnet 5 with room to spare, fits in Opus 5 with the 1M window with enormous room to spare, and leaves the reasoning model doing the thing it is genuinely good at: noticing that two things you thought were unrelated move together.
If you want the single-ride version of this (what to put in a prompt for one file, which metrics mislead, how to stop a model inventing a W’ balance it cannot compute), that groundwork is in Analysing Ride Data With LLMs. This piece assumes you have it and want to go a hundred times wider.
Getting a year of files out
Use Strava’s bulk export, not the API. Settings → My Account → Download or Delete Your Account → Request your archive gets you activities.csv plus every original upload as .fit.gz. The API’s rate limits (100 requests per 15 minutes and 1,000 a day, at time of writing) mean 312 activities plus 312 stream calls is a two-day job with retry handling, for data you can have as a zip in an hour.
intervals.icu is the better source if you already sync there, because it has done half the work. GET /api/v1/athlete/{id}/activities?oldest=2025-01-01&newest=2025-12-31 returns one JSON object per activity with load, intensity, decoupling, efficiency factor, interval detection and your own custom fields already computed. Parse .fit files yourself with fitparse or python-fitparse only for the fields intervals.icu does not expose, such as raw left/right balance or Di2 gear position if you log it.
The summarise pass
One model call per ride, no conversation, same schema every time. I used Haiku 4.5 because the task is extraction, not judgement, and because 312 calls of a reasoning model is money set on fire. Input is a downsampled stream (10-second averages, which cuts 10,800 rows to 1,080) plus the activity metadata and whatever interval structure the file carries.
The prompt that worked, after three rewrites:
You are converting one bike ride into a fixed digest. Output YAML only.
Do not interpret, judge, or advise. If a field cannot be computed from
the data given, write null. Never estimate a value you cannot derive.
Fields: date, duration_min, distance_km, work_kj, avg_power, np,
if_, vi, avg_hr, max_hr, pwhr_decoupling_pct, ef, avg_cadence,
elevation_m, avg_temp_c, best_5s/1min/5min/20min/60min power,
time_in_zone_min (z1..z7, Coggan), detected_intervals (list of
{reps, target_min, actual_mean_power, first_rep_power,
last_rep_power}), ride_type (one of: endurance, threshold, vo2,
sprint, tt_specific, race, commute, recovery, unstructured),
surface (road/gravel/indoor/mixed), notes_from_title (verbatim).
That last_rep_power versus first_rep_power pair turned out to be the highest-value field in the whole schema, and I nearly left it out. Forcing null rather than estimation is the other non-negotiable: without that line, Haiku cheerfully invented decoupling figures for rides with no heart rate strap, and those fabrications then got reasoned over downstream as if real.
Output per ride looks like this:
date: 2025-07-08
duration_min: 94
np: 243
if_: 0.84
vi: 1.09
pwhr_decoupling_pct: 8.2
ride_type: vo2
detected_intervals:
- reps: 5
target_min: 4
actual_mean_power: 331
first_rep_power: 352
last_rep_power: 309
avg_temp_c: 27
Cost: roughly 4.5M input tokens and 95k output across the batch. At Haiku 4.5’s published rates that came to about four pounds for the season. Check current pricing before you budget, but the order of magnitude is “cheaper than a cafe stop”, and it is a one-off because the digests are cached on disk forever.
The reasoning pass
Now hand all 312 digests to Opus 5 in one go, with the training blocks explicitly labelled. The labelling matters: on my first attempt I gave it dates and let it infer block boundaries, and it confidently discussed a “Block 6” that did not exist. Give it the structure you actually trained to.
| Block | Weeks | Intent | Hours | Load (TSS) | FTP at end |
|---|---|---|---|---|---|
| 1 | 1–6 | Base, polarised | 58 | 3,420 | 268 |
| 2 | 7–12 | Threshold build | 71 | 4,910 | 279 |
| 3 | 13–18 | VO2 + TT work | 64 | 4,780 | 282 |
| 4 | 19–26 | Race block, gravel | 97 | 6,340 | 281 |
| 5 | 27–34 | TT focus | 76 | 5,120 | 291 |
| 6 | 35–42 | Late road races | 52 | 3,290 | 286 |
The prompt I settled on:
Below are 312 ride digests for one athlete across 42 weeks, grouped
into 6 labelled blocks. Find patterns that hold across multiple rides
and multiple weeks. Rules:
1. Every claim must cite at least 4 ride dates as evidence. Claims
resting on 3 or fewer rides are noise: discard them.
2. Compare blocks against each other, not rides against rides.
3. Say explicitly where the stated intent of a block and the data
disagree.
4. Rank findings by how much changing them would alter outcomes.
5. Do not comment on individual sessions. Do not give me a plan yet.
Rule 1 killed about half of what it wanted to tell me, which is the point. Rule 5 is the one that turns a chat toy into an analyst.
What it found
Recovery weeks were not recovery weeks. My plan put a reduced week every fourth week at roughly 45% of the preceding week’s load. Across the 10 such weeks in the season, median actual load was 71% of the preceding week, and 8 of 10 exceeded 60%. The model traced it to ride titles: Zwift group rides logged as “easy spin” that carried IF 0.78 and 61 minutes above threshold heart rate. Nobody reading a single ride would flag that. The pattern is only visible as a ratio, repeated ten times.
Threshold work quietly stopped being threshold work. Total time in zone 4 held near-constant from week 14 onward, around 48 minutes per week, which is why it felt consistent. Mean interval duration fell from 12.4 minutes in Block 2 to 7.8 minutes in Block 5. Same minutes, different stimulus, delivered in shorter chunks because shorter chunks are more pleasant. FTP movement across Blocks 3 and 4 was +2 W over 14 weeks.
Heat, not form. Pw:Hr decoupling on rides over 90 minutes had a median of 4.1% in April and 7.9% in July. Reading July’s files alone, you would conclude aerobic decline and panic. Cross-referencing avg_temp_c across all 312 digests, every decoupling figure above 7% occurred on a ride averaging over 24°C, and 19 of 23 of those carried two bottles or fewer for over two hours.
Best efforts clustered where I did not expect. Of my top 14 twenty-minute powers, 11 came on the second or third day after a rest day, not the first. Day-one-back rides averaged 94% of the matched day-two figure. That is a scheduling fix, not a training fix, and it cost nothing to apply.
Gravel pacing was leaking watts. Variability index averaged 1.14 across 18 gravel events against 1.06 across 11 road races, with the surges concentrated in the first 40 minutes. Three of the five gravel results where I faded badly showed first-hour normalised power above 102% of eventual full-ride NP.
Where it was wrong, and the guardrails that caught it
It twice asserted a relationship between cadence and efficiency factor that evaporated when I pulled the cited rides and checked: two of the four citations had cadence sensors that had dropped out, logging implausible 42 rpm averages that the digest dutifully recorded as real. The citation requirement is what made this findable in ninety seconds. Without it you get prose you cannot audit, which is worse than no analysis because it feels like analysis.
A second failure mode: asked for findings, a model will produce findings, and it will produce them at whatever quantity you imply you want. I asked for ten. Four survived scrutiny. Ask for “as many as the evidence supports” and you will get a more honest count.
Spot-check the digests too, not just the conclusions. I re-derived best 20-minute power by hand for 15 randomly chosen rides against intervals.icu’s own figures. Fourteen matched within 2 W. The fifteenth was off by 31 W because the ride contained a paused segment that the downsampler had closed up, splicing two efforts into one. Sampling 5% of your digests takes an evening and tells you whether the other 95% can carry weight.
Next winter the pipeline runs weekly rather than annually, because the recovery-week finding would have been visible by March if anything had been looking for it, and 71% instead of 45% for ten weeks is most of a season’s adaptation left on the road.