AI Cycling
§4 Section 4 of 6 2,701 words · 12 min

Recovery, HRV And Readiness Scores

Your Whoop says 31% recovery. Your legs say they’re fine. You have 4×8 at 105% FTP on the plan, it’s Tuesday, and you have to decide in the next ten minutes whether to do it. That decision is the entire point of this page, and almost nothing written about HRV actually helps you make it.

The useful version of hrv guided cycling training is not “train hard when the number is green.” It’s a set of rules you’ve tested against your own training history, applied to a signal you’ve cleaned up, with an honest estimate of how often the signal is wrong. Below is how to build that, how to wire it into intervals.icu, and how to get an LLM to do the analytical grunt work without inventing conclusions your data doesn’t support.

What A Readiness Score Is Actually Made Of

Every consumer readiness score is a weighted composite with undisclosed weights. Knowing roughly what goes in changes how much you should trust the output.

Whoop computes recovery primarily from rMSSD measured during your last slow-wave sleep period, plus resting heart rate, respiratory rate, and sleep performance. Oura averages HRV across the whole night and folds in body temperature deviation, which makes it the best of the three at flagging an incoming illness. Garmin’s HRV Status uses overnight measurements to build a seven-day average, compared against a personalised range that takes roughly three weeks of consistent wear to establish; Training Readiness then blends that with sleep, acute load, recovery time and stress.

Three different sampling windows, three different baselining methods, three different composite formulas. That’s why the same night produces a Whoop recovery of 42% and a Garmin readiness of 74%. Neither is lying. They’re measuring different things and calling both of them “recovery.” If you’re deciding which device to actually buy, or you’re running two and want to know which one to believe, that comparison is its own page: Whoop, Oura and Garmin HRV compared for cyclists.

For training decisions, ignore the composite. Pull the raw HRV value out and work with it directly.

Baseline And Normal Range: The Only Two Numbers You Need

rMSSD is log-normally distributed, which means raw values are misleading at the tails. Take the natural log first. Everything that follows uses ln rMSSD.

Two derived figures do all the work:

The 7-day rolling mean of ln rMSSD. This is your actual state. Single days are noise.

Your normal range, defined as the 60-day rolling mean ± 0.75 SD. Some practitioners use ±0.5 SD (more sensitive, more false alarms) or ±1.0 SD (calmer, slower to react). Start at 0.75 and adjust after you’ve seen a season of data.

Here’s a real-shaped fortnight. Rider: 41, FTP 285 W, 11 hours a week, morning chest-strap measurement, supine, 60 seconds, same time each day. Sixty-day baseline mean ln rMSSD 4.12, SD 0.18, so the normal range is 3.99 to 4.26.

DaterMSSDln rMSSDRHRSession
Mon 14624.1344Rest
Tue 15584.06454×8 @ 105%
Wed 16493.8947Z2 90 min
Thu 17664.1944Rest
Fri 18644.1644Openers
Sat 19443.78485 h club ride, 3 pints after
Sun 20513.9346Z2 2 h

Rolling mean: 4.02. CV of ln rMSSD across the week: 3.9%. Inside the range, but bouncing around a lot.

DaterMSSDln rMSSDRHRSession
Mon 21453.81493×12 sweetspot
Tue 22413.7151Z2
Wed 23433.76505×5 VO2
Thu 24473.8550Rest
Fri 25423.7452Z2
Sat 26463.83514 h endurance
Sun 27443.7850Z2

Rolling mean: 3.78. Below the floor of 3.99, and it’s been below for five consecutive days. CV: 1.3%.

Notice the CV flipped direction. Week one had a high CV around a normal mean, which is day-to-day chaos from alcohol and a long ride, not a training problem. Week two is stable and suppressed: low variability around a depressed mean, with RHR up 5-6 bpm. That combination, tracked in endurance athletes by Plews and colleagues, is the pattern that precedes a bad third week. Week one needs no intervention. Week two needs the intensity pulled.

Why Your Daily Number Lies To You

A single morning reading responds to things that have nothing to do with training readiness. Two units of alcohol the night before will typically knock 10-20% off next-morning rMSSD. Eating within three hours of bed does something similar. Sleeping in a 24 °C room, flying, a blocked nose, needing the toilet, measuring at 06:10 on a work day versus 08:30 at the weekend: all of it moves the number more than a hard interval session does.

The trap that catches experienced riders is the opposite direction. Parasympathetic rebound means HRV can be elevated on the morning after a genuinely destructive session, and in very fit riders with low resting heart rates, rMSSD can saturate and stop tracking fitness gains. A rising seven-day mean is not automatically good news. Read it alongside RHR and whether your power is actually there.

Practical consequence: never act on one day. Act when the seven-day mean crosses the range boundary and stays there for two or more days, or when RHR and HRV move together.

Wiring It Into intervals.icu

intervals.icu is the right home for this because it already holds your power files, and its wellness endpoint gives you everything in one CSV.

curl -u API_KEY:$ICU_KEY \
  "https://intervals.icu/api/v1/athlete/i123456/wellness.csv?oldest=2026-04-01&newest=2026-09-30" \
  -o wellness.csv

The username is the literal string API_KEY; the password is the key from Settings → Developer. Useful columns that come back:

id,ctl,atl,rampRate,restingHR,hrv,hrvSDNN,sleepSecs,sleepScore,
avgSleepingHR,soreness,fatigue,stress,mood,motivation,respiration,
spO2,weight,readiness,comments

hrv is rMSSD. id is the date. comments is a free-text field, and it is the single highest-value column you own, because it’s the only place the reason for a dip gets recorded.

Compute the two numbers that matter:

import pandas as pd, numpy as np

w = (pd.read_csv("wellness.csv", parse_dates=["id"])
       .rename(columns={"id": "date"}).set_index("date").sort_index())

w["ln"]    = np.log(w["hrv"])
w["roll7"] = w["ln"].rolling(7, min_periods=5).mean()
w["cv7"]   = w["ln"].rolling(7, min_periods=5).std() / w["roll7"] * 100

base = w["ln"].shift(1).rolling(60, min_periods=30)   # shift so today isn't in its own baseline
w["lo"] = base.mean() - 0.75 * base.std()
w["hi"] = base.mean() + 0.75 * base.std()

w["state"] = np.select([w.roll7 < w.lo, w.roll7 > w.hi], ["LOW", "HIGH"], "NORMAL")
print(w[["hrv","ln","roll7","cv7","lo","hi","state","restingHR"]].tail(10).round(2))

Output for the fortnight above:

             hrv    ln  roll7  cv7    lo    hi   state  restingHR
2026-09-21  45.0  3.81   3.94  3.6  3.99  4.25     LOW       49.0
2026-09-22  41.0  3.71   3.90  3.9  3.99  4.25     LOW       51.0
2026-09-23  43.0  3.76   3.87  3.8  3.98  4.25     LOW       50.0
2026-09-24  47.0  3.85   3.83  2.0  3.98  4.24     LOW       50.0
2026-09-25  42.0  3.74   3.79  1.6  3.98  4.24     LOW       52.0
2026-09-26  46.0  3.83   3.78  1.4  3.97  4.24     LOW       51.0
2026-09-27  44.0  3.78   3.78  1.3  3.97  4.24     LOW       50.0

Seven consecutive LOW days with CV collapsing toward 1.3% is not a rest day. It’s a recovery week.

Rules That Survive A Real Season

Decision rules only work if they’re written down before you need them, because the whole problem with self-coaching is that you negotiate with yourself at 06:00.

IF roll7 inside range        -> execute plan as written
IF roll7 < lo for 1 day      -> execute plan, note it
IF roll7 < lo for 2-3 days   -> swap the next quality session for Z2, same duration
IF roll7 < lo for 4+ days    -> recovery week: cut weekly TSS 40%, zero above threshold
IF roll7 > hi AND cv7 > 4%   -> treat as unstable; hold intensity, don't add
IF roll7 > hi AND cv7 < 3%   -> green light; this is where you add a session
IF RHR > baseline + 7 AND respiration > baseline + 1.0 -> illness protocol, don't ride

That last line is where wearables genuinely earn their subscription. Detecting an infection 24-36 hours before you feel it, from resting heart rate and respiratory rate moving together, is worth more over a season than every readiness percentage combined.

Suspend the whole system in two situations. During a planned overreaching block, suppressed HRV is the intended outcome and the rules will sabotage you. During a taper, HRV often drops in the final week before it rebounds, and you’ll talk yourself into an unnecessary extra rest day at the worst possible moment.

Prove The Score Predicts Something For You

Before you let a number change your training, make it earn the right. This takes an evening.

Pull 180 days. For every day with a prescribed quality session, label the outcome: hit (completed the work at target) or missed (abandoned, or final interval more than 5% below target normalised power). Then cross-tabulate against the morning readiness band.

A real-shaped result from six months of data, 96 quality sessions:

Readiness bandSessionsMissedFailure rate
Red (<34%)14750%
Yellow (34-66%)471123%
Green (67%+)3539%

Base rate: 21 of 96, so 22%. There’s real signal here, with red mornings failing at 5.6× the rate of green ones.

Now price it. Skipping every red day over those 26 weeks would have avoided 7 bad sessions and binned 7 perfectly good ones, costing 14 quality sessions in total. That’s roughly 15% of your annual intensity to dodge seven sub-par workouts. For most self-coached riders the better rule isn’t “skip”, it’s “start the session and abandon it after rep two if the power isn’t there.” The score buys you permission to bail, not permission to skip.

Run the same table for your own data before adopting any rule. If your failure rate is flat across the three bands, the score is measuring your sleep and your commute, not your legs.

Getting An LLM To Read Your Wellness File

Paste 180 rows of CSV into a chat window, ask “am I overtrained?”, and you will get a confident, fluent, wrong answer. Three failure modes cause it.

Arithmetic over long tables. Models approximate when averaging 60 numbers. Never ask for a computed statistic from raw rows. Either compute it yourself with the script above, or make the model write the code and run it in an execution environment, then read the code.

Sycophancy. This is the big one, and it’s easy to test. Send the identical dataset twice, once framed “I feel fresh and think I should add a VO2 block”, once framed “I feel wrecked and think I need a week off.” If the analysis flips, the model is reading your framing and not your data. Most will flip. Build the test into your workflow permanently.

Missing context. The model cannot see that your HRV dip on 19 September was three pints after the club run. If it isn’t in the comments column, it doesn’t exist, and the model will attribute the dip to training load.

Here’s a prompt that works. It assumes you’ve attached the processed CSV with ln, roll7, cv7, lo, hi already computed.

You are analysing my cycling wellness data. I have attached a CSV with these
columns: date, hrv (rMSSD ms), ln, roll7, cv7, lo, hi, restingHR, respiration,
ctl, atl, comments. The lo/hi bounds are my 60-day baseline +/- 0.75 SD of
ln rMSSD, excluding the current day.

Do all arithmetic in code. Show me the code. Do not estimate from the table.

Produce exactly these four sections:

1. STATE. Is roll7 inside, above or below the range right now, and for how many
   consecutive days? Quote the dates and values.
2. CONFOUNDERS. For every day where ln rMSSD is more than 1 SD below baseline,
   check the comments field and tell me whether a non-training explanation is
   recorded. List the dates with no explanation separately.
3. COUNTER-CASE. Argue the position that nothing is wrong and this is normal
   variation. Use specific numbers. Make it the strongest version of that case.
4. DECISION. One of: PROCEED / SWAP NEXT QUALITY SESSION TO Z2 / RECOVERY WEEK.
   Give the two numbers that most drove the choice.

Do not give me general training advice. Do not tell me to listen to my body.
If the data is insufficient to decide, say INSUFFICIENT and say what is missing.

Section 3 is doing the heavy lifting. Forcing an explicit counter-argument is the cheapest available defence against a model that has already worked out which answer you want.

For the go/no-go call on a specific morning, a shorter version:

Context: planned session today is 4x8 min at 300W (FTP 285). Yesterday: 2h Z2,
TSS 95. CTL 78, ATL 94, form -16.
This morning: rMSSD 43, roll7 3.87, range 3.98-4.25, RHR 50 (baseline 45),
slept 6h10, respiration 14.8 (baseline 13.9).

Apply exactly these rules and nothing else:
[paste your rule block]

Output: the matching rule, the prescription, one sentence of reasoning.

Constraining the model to rules you wrote turns it from an oracle into a lookup function, which is what you actually want at six in the morning.

Where Readiness Scores Break For Cyclists

Strain and training stress measure different things, and the gap is worst for exactly the sessions cyclists care about. A five-hour endurance ride at 200 W produces something like 250 TSS and a Whoop strain near 18. A 75-minute session with 4×8 at threshold produces maybe 95 TSS and a strain around 13. The wearable is confident the long ride was the harder day. Your quads, and your next interval session, disagree. Duration-weighted heart rate systematically over-rates long steady rides and under-rates short, sharp, high-torque work.

Heat compounds it. A 26 °C ride in the Chilterns in July inflates heart rate at identical power, inflates strain, and suppresses next-morning HRV through thermoregulatory load rather than mechanical load. The readiness score reads it as training stress. It isn’t, and the adaptation you get from it is different.

Then there’s the measurement mismatch. Wrist optical sensors during sleep and a chest strap measurement taken supine at 06:15 are not the same construct, and they will not agree on the size of a change even when they agree on its direction. Pick one method and stay on it for at least a full training block; the device-by-device comparison covers which sensors hold up under which conditions.

What The Commercial Tools Actually Do

Worth being precise here, because the marketing isn’t.

TrainerRoad’s Red Light Green Light doesn’t use HRV or wearables at all. It reads your completed workouts and recent performance to estimate whether you’re absorbing load. Xert’s freshness model is built from power data and its MPA framework, again with no wellness input. Athletica does ingest HRV and wellness and adapts the plan around it. HRV4Training has the longest-running advice engine specifically built on morning rMSSD, and its baseline handling is more careful than most. Whoop Coach is an LLM over Whoop’s own data and has no idea what your FTP is or what’s on your calendar.

The structural gap: no product currently sees your wellness stream, your power files, and your actual season goals at the same time, with a sensible model of periodisation over the top. That gap is why a 400-word prompt over an intervals.icu export, run by someone who knows what a good 4×8 feels like, outperforms every in-app chatbot on the market right now.

Your next move is the validation table. Six months of data, 90 minutes of labelling, and you’ll know whether your readiness score deserves a vote on Tuesday mornings or whether it’s an expensive sleep tracker.

In this section

The supporting pages under this subject.