AI Cycling
§6 Section 6 of 6 2,643 words · 12 min

Build Your Own Training Tools

Most cyclists who try to use AI for training analysis start in the wrong place. They paste a week of TSS numbers into a chat window, ask “am I overtraining?”, and get back a plausible-sounding paragraph that would have been equally plausible if the numbers were doubled. The model has no way to check itself. You gave it 40 tokens of context and asked it to reason about a physiological system it cannot see.

The fix is not a better prompt. It is a pipeline: get your ride files into a local dataset you control, compute the derived metrics yourself in Python, and use the model for the part it is genuinely good at, which is reading a structured summary and telling you something you would not have noticed. That workflow is what this page is about. It assumes you already know what CP, W’, normalised power and TSS are, and that you have at least two years of files sitting in Strava doing nothing.

Why The Data Layer Has To Come First

An LLM will happily tell you your FTP is 285W because you told it your FTP is 285W. It will not tell you that your last three 20-minute efforts all came in at 268, 271 and 266W normalised, which means your FTP estimate is 6% stale and every zone in your plan is drifting high.

That check costs about 15 lines of Python once you have the data. It is impossible without the data. So the first build is not a coaching assistant, it is an extract-and-store job: pull every activity’s summary from the Strava API, pull the time-series streams for the ones you care about, and write them somewhere you can query in under a second.

The practical shape of this, for a rider with roughly 800 activities over four years:

LayerWhat it holdsSize on disk
activities table (SQLite)id, date, type, moving_time, distance, avg/max power, avg HR, elevation, device~400 KB
streams/ (Parquet, one per ride)time, watts, heartrate, cadence, velocity, altitude at 1 Hz180–400 KB per ride, ~230 MB total
derived table5s/1m/5m/20m/60m maxima, CP/W’ fit, TSS, decoupling, VI~1 MB

Roughly 230 MB. That fits on a laptop, queries instantly, and belongs to you, which matters because Strava’s API terms restrict what you can do with data from other athletes and because rate limits make repeated pulls painful. The setup work (OAuth, refresh tokens, the 100-requests-per-15-minutes and 1,000-per-day limits, pagination through /athlete/activities) is covered in detail in Strava API With Python: Auth, Rate Limits And Your First Dataset. Build that first. Everything below assumes it exists.

One warning that trips up nearly everyone on their first backfill: GET /activities/{id}/streams counts as a separate request per activity. At 100 requests per 15 minutes, a 800-ride backfill is about two hours of wall clock with a sleep loop. Write it to resume from the last-fetched id, run it overnight, and never run it again.

The Derived Metrics That Actually Earn Their Place

Once streams are local, compute these on ingest and store them. Doing it at query time is slow and means the model gets inconsistent inputs.

Rolling power maxima. For each ride, the best mean power over 5s, 15s, 30s, 1m, 3m, 5m, 8m, 12m, 20m, 30m, 60m. A NumPy cumulative sum makes this trivial:

import numpy as np

def best_efforts(watts, durations):
    w = np.asarray(watts, dtype=float)
    w = np.nan_to_num(w)
    c = np.concatenate([[0.0], np.cumsum(w)])
    out = {}
    for d in durations:
        if len(w) < d:
            continue
        means = (c[d:] - c[:-d]) / d
        out[d] = float(means.max())
    return out

On a 3-hour ride at 1 Hz that runs in about 4 ms. Across 800 rides it’s a few seconds, so you can afford to recompute the whole set whenever you change a definition.

CP and W’ from a rolling 90-day window. Fit the two-parameter model P(t) = W'/t + CP over your best efforts between 3 and 20 minutes, using the maxima from the last 90 days rather than a single test. A linear regression of power against 1/t gives CP as the intercept and W’ as the slope. For a typical trained amateur you will see something like CP 271W, W’ 19.4 kJ. When W’ drops while CP holds, you have been doing threshold work and neglecting anaerobic capacity. When CP drops and W’ rises, you are fresh but detrained aerobically, which is exactly what six weeks of crits does.

Aerobic decoupling (Pw:Hr). Split any steady ride over 60 minutes at the halfway point, compute power-to-heart-rate ratio for each half, and report the percentage change. Under 5% is good aerobic durability. Above 8% on a zone 2 ride means either the ride was too hard, you were underfuelled, or your aerobic base has slipped. This is the single most useful number for base-phase riders and almost nobody computes it because Strava doesn’t show it.

Durability: power after 2,000 kJ. This one matters more than FTP for anyone racing over three hours. Take your best 5-minute power from the portion of each ride occurring after 2,000 kJ of work, and track it against your fresh 5-minute best. A rider with a fresh 5-min of 340W who drops to 268W after 2,000 kJ has a 21% fade. Tracking that number across a season tells you whether your long rides are doing anything.

Interval detection. Rather than trusting Strava’s laps (riders forget to press the button), detect efforts by finding contiguous blocks where power exceeds 88% of CP for at least 120 seconds with gaps under 60 seconds. Store each as a row with duration, average power, average HR, and the position in the ride. This turns an unstructured stream into something a model can reason about.

Which Models Hold Up Against Real Ride Files

Model choice matters less than most people think for text reasoning, and more than they think for numerical work and long context. Rough guidance based on what these tasks actually demand:

For weekly and block-level analysis, where you’re feeding a structured summary of 7 to 28 days and asking for interpretation, Claude Sonnet 5 is the sensible default. The input is a few thousand tokens, the reasoning is qualitative, and the cost is negligible: an entire season of weekly reviews runs to maybe 400,000 input tokens.

For anything that involves actually reasoning over the numbers, like reconciling a CP fit against a set of best efforts and deciding whether the fit is being dragged by a single anomalous effort, use Claude Opus 5 with extended thinking. The difference shows up specifically when the correct answer requires noticing an inconsistency rather than summarising. Ask a smaller model why your CP jumped 14W and it will tell you about training adaptation. Ask a stronger one with the underlying efforts in context and it is more likely to point out that the 3-minute effort in the fit came from a 4% descent with a tailwind, and the fit should be rerun without it.

For bulk classification, such as tagging 800 rides as endurance / tempo / threshold / VO2 / race / commute based on their interval structure, Haiku 4.5 is the right tool. At roughly 600 input tokens and 30 output tokens per ride, tagging the full archive costs well under a pound and takes a few minutes with modest concurrency.

The mistake worth avoiding: do not send raw 1 Hz streams to any model. A 3-hour ride at 1 Hz is 10,800 samples across five channels. Even downsampled to 30-second bins that’s 360 rows, and the model will still do arithmetic on it badly. Compute in Python, send the summary. The model’s job is interpretation, not addition.

A Prompt Structure That Survives Contact With Real Data

Here is the pattern that works, in three parts: a system prompt that fixes the model’s frame of reference, a structured data block, and a question with an explicit escape hatch.

SYSTEM = """You are analysing training data for a self-coached
amateur cyclist. You have the athlete's computed metrics; you do
not have the raw files.

Rules:
- Never estimate a number you were not given. If a metric you need
  is absent, say which one and stop.
- Distinguish clearly between what the data shows and what you are
  inferring from it.
- A single ride is almost never evidence of anything. Require at
  least three data points before calling a trend.
- Flag data-quality problems before interpreting: power spikes
  above 1500W, HR below 35 or above 210, rides where cadence is
  zero for more than 40% of moving time.
"""

The data block should be compact and typed. JSON works, but a plain table costs fewer tokens and models read it fine:

BLOCK SUMMARY 2026-08-04 to 2026-09-28 (8 weeks)

week  rides  hours  TSS   CTL  ATL  TSB  z2_hrs  >CP_mins
w1    5      8.2    412   61   58   +3   6.1     22
w2    6      10.4   561   64   71   -7   7.9     34
w3    6      11.1   598   68   76   -8   8.2     41
w4    3      4.1    198   64   48   +16  3.6     8
w5    6      10.8   574   67   73   -6   7.7     38
w6    6      11.6   631   71   80   -9   8.4     46
w7    5      9.9    522   72   72   0    7.1     31
w8    2      3.0    142   67   40   +27  2.7     4

BEST EFFORTS (rolling 90d, watts)
5s 1094 | 1m 512 | 3m 371 | 5m 336 | 12m 291 | 20m 279 | 60m 258

CP FIT (90d, 3-20min, n=11): CP 268W, W' 20.1 kJ, r2 0.987
PREVIOUS FIT (ending w1): CP 259W, W' 22.8 kJ, r2 0.991

DECOUPLING, z2 rides >90min (Pw:Hr % change, half 1 to half 2)
w1 4.1 | w2 5.8 | w3 6.9 | w5 5.2 | w6 7.4 | w7 4.4

DURABILITY: best 5min after 2000kJ, last 6 rides over 2500kJ
289 | 271 | 294 | 301 | 288 | 306

Then the question, narrow and answerable:

Given the above, identify the two most likely limiters for a
60-minute flat time trial in six weeks, and state what in the data
supports each. If the data does not support a confident answer,
say so and name the specific metric that would resolve it.

What you get back from a strong model on that input is specific: CP up 9W while W’ fell 2.7 kJ (the block was threshold-heavy, anaerobic capacity has been traded away, which is fine for a 60-minute TT); decoupling trending up from 4.1 to 7.4 across w1 to w6 despite growing fitness, which points to fuelling or heat rather than aerobic fitness; durability actually improving (289 to 306W), so the long rides are working. And a note that the 60-minute best of 258W sits below CP of 268W, meaning you have no recent effort at the actual duration of the event, so the CP figure is an extrapolation.

That last point is the kind of thing the pipeline earns you. It only appears because the 60-minute best and the CP fit were both in context and the model could compare them.

Building A Retrieval Layer Over Your Own History

Weekly summaries handle the recent picture. The harder question is comparative: “have I ever produced numbers like this before, and what happened next?”

Answer it with a simple similarity search rather than a vector database. Build a feature vector per week from eight to twelve normalised numbers (hours, TSS, CTL, TSB, z2 hours, minutes above CP, mean decoupling, best 20-min as a fraction of CP) and find the nearest historical weeks by Euclidean distance on the z-scored features. Then pull what happened in the following fortnight.

import numpy as np, pandas as pd

FEATURES = ["hours","tss","ctl","tsb","z2_hrs","mins_over_cp",
            "decoupling","best20_pct_cp"]

def similar_weeks(df, target_idx, k=5):
    X = df[FEATURES].to_numpy(float)
    z = (X - X.mean(0)) / X.std(0)
    d = np.linalg.norm(z - z[target_idx], axis=1)
    d[target_idx] = np.inf
    return df.iloc[np.argsort(d)[:k]].assign(distance=np.sort(d)[:k])

Feed the five nearest weeks plus their outcomes into the prompt alongside the current block. The model now has your personal history instead of a general prior about cyclists. In practice this catches patterns like: the three previous times you exceeded 45 minutes above CP in a week, your following week’s decoupling rose above 7% and you skipped a session. That’s a real signal from your own files, not received wisdom about a 10% weekly increase rule.

Four years of data gives you roughly 200 weeks, which is enough for this to be useful and nowhere near enough for anything statistical. Treat it as a way of finding comparable cases to look at, not as prediction.

Validating Model Output Against Your Own Files

Any tool you build needs a way of catching when the model is confidently wrong. Three mechanisms, in order of effort:

Structured output with required citations. Force the model to return JSON where every claim carries the metric names it rests on:

{
  "findings": [
    {
      "claim": "Aerobic durability improved across the block",
      "confidence": "high",
      "metrics_cited": ["durability_5min_post_2000kj"],
      "values": [289, 271, 294, 301, 288, 306]
    }
  ]
}

Then assert in code that every name in metrics_cited exists in the data block you sent, and that every number in values appears there too. A claim citing a metric you did not provide is a hallucination, and you catch it automatically rather than by reading carefully.

Held-out checks. Withhold one metric from the prompt, ask the model to predict it, compare to the real value. If you hide the 20-minute best and give it everything else, a decent model should land within about 5%. If it doesn’t, your data block is missing something informative and you should reconsider what you’re including.

Adversarial replay. Take a summary from a block where you know what happened (an illness, a two-week holiday, a genuine breakthrough) and check whether the analysis picks it up without being told. The holiday case is a good test because the naive reading is “detraining” while the correct reading depends on what came next. Run this on five or six historical blocks and you will get a fast read on whether your prompt is producing insight or horoscope.

What To Build, In Order

Ship the ingestion first: OAuth, activity backfill, stream fetch with resume, SQLite plus Parquet. A weekend. Then the derived metrics, starting with best efforts and CP fit because everything else depends on them. Then one prompt, one question, run weekly, with the validation assertions in place from day one rather than bolted on later.

Skip the chat interface. The temptation is to build something you talk to, but the version that changes your training is a scheduled Monday morning script that writes a markdown file to disk and, if you like, emails it. There is no session state to manage, the prompt is version-controlled, and you can diff this week’s output against last week’s.

A few things not to bother with. Don’t build a workout generator before you trust your analysis; generated sessions from a model that misreads your fitness are worse than the plan you already have. Don’t try to replicate WKO5’s model fitting, which represents a lot of years of work you will not match in a weekend. And don’t feed the model your subjective RPE notes without the power data next to them, because it will weight the prose far more heavily than the watts.

The riders who get real value out of this are the ones who treat the model as a reader rather than an oracle. It reads a well-constructed summary of your own numbers faster and more attentively than you do at 9pm on a Sunday, and it notices the 60-minute best sitting below the CP estimate. Build the thing that produces the summary, and the rest is a few hundred tokens.

In this section

The supporting pages under this subject.