AI Cycling
006 Recovery, HRV And Readiness Scores 1,763 words · 8 min

HRV For Cyclists: What It Measures And What It Cannot

Your Oura ring says 42. Yesterday it said 58. You have 4x8 at threshold on the plan and a coffee going cold next to the turbo. The temptation is to read that 16ms drop as a verdict on how recovered you are, because that is the word every app prints next to the number. It is the wrong word, and the wrong frame, and it will make you a worse self-coach than you were before you bought the ring.

Heart rate variability measures one thing: the degree to which your vagus nerve is currently modulating the interval between heartbeats. That is it. High-frequency HRV, which is what the rMSSD figure in your app approximates, is essentially a proxy for parasympathetic outflow to the sinoatrial node during the window sampled. When you see 58, your vagus was braking your heart rate hard and variably at 04:30 this morning. When you see 42, it was braking less. The physiological process being indexed is parasympathetic reactivation: the return of vagal tone after it has been withdrawn by exercise, heat, alcohol, poor sleep, illness, caffeine at 6pm, a work deadline, or your kid’s school assembly.

Parasympathetic reactivation is genuinely interesting for endurance athletes, and it is genuinely useful for structuring hrv guided cycling training. But it is not “recovery”, because recovery is a composite. Glycogen resynthesis, muscle protein turnover, tendon remodelling, mitochondrial biogenesis, central fatigue, immune function: none of these are indexed by rMSSD. A rider can have fully reactivated vagal tone and completely empty legs. This happens constantly after a big week of low-intensity volume, where HRV often rises while the quads still feel like they belong to someone else.

What the number actually is, arithmetically

rMSSD is the root mean square of successive differences between R-R intervals. Take a string of beat-to-beat intervals in milliseconds:

R-R intervals (ms): 1042, 1088, 1031, 1097, 1054, 1102, 1048
Successive diffs:      +46,  -57,  +66,  -43,  +48,  -54
Squared:              2116, 3249, 4356, 1849, 2304, 2916
Mean of squares:      2798
rMSSD:                52.9 ms

Two things follow immediately. First, rMSSD is dominated by the largest successive differences, because of the squaring. One ectopic beat or one motion artifact in a 3-minute morning reading can shift the number by 10ms or more. Whoop, Oura and Garmin all apply artifact correction, and they do it differently, which is one reason your Garmin HRV Status and your Oura HRV do not agree and never will.

Second, most apps report the natural log, or a 0-100 score derived from it, because raw rMSSD is right-skewed. HRV4Training and EliteHRV both display transformed values. A 10-point move in an lnRMSSD-based score is not the same as a 10ms move in raw rMSSD. If you export from one platform and compare to another, check which you are looking at before you draw any conclusion at all.

Why a single day tells you almost nothing

Within-person day-to-day coefficient of variation for morning rMSSD sits in the region of 15-20% for most trained athletes. Take a rider whose 30-day mean rMSSD is 65ms. A CV of 18% gives a standard deviation of about 11.7ms. That means roughly two-thirds of normal, unremarkable, nothing-is-wrong mornings will land between 53 and 77ms, and about 95% of them between 42 and 88ms.

So the 42 that made you consider binning the intervals is, for this rider, at the bottom edge of ordinary variation. It is a data point, not a signal.

The standard fix is Kubios’s and Plews’s smallest worthwhile change approach: compute a rolling 7-day mean of your daily values, and compare it to a rolling 30- or 60-day baseline with a normal range of roughly ±0.5 to ±1.0 within-person SD. When the 7-day mean drops below the normal range, something systemic is happening. When one day drops, you had a glass of wine.

Here is a real-shaped week for a 42-year-old TT rider, 30-day mean 65ms, SD 11.7ms, normal range 53-77ms:

DayDaily rMSSD7-day meanIn range?Session done
Mon7166.1yes90min Z2
Tue4864.7yes5x5 VO2max
Wed5263.4yes60min recovery
Thu5962.0yes2x20 FTP
Fri4459.7yesrest
Sat4156.9yes4hr endurance
Sun3953.4borderline3hr endurance
Mon3745.7below?

Look at what happened. Tuesday’s 48 and Friday’s 44 both look alarming in isolation, and both were noise: the rolling mean barely moved. The actual signal arrives on the second Monday, when the 7-day mean falls out of the normal range after a big weekend. That is the day to swap threshold work for two hours easy, and it is the only day in eight where the data earned the right to change the plan.

Where AI coaching tools get this right and wrong

I’ve fed this kind of block into several tools and the failure mode is consistent.

Athletica AI handles it properly at the architecture level: it ingests HRV alongside training load and adjusts the next session’s prescription, rather than issuing a standalone daily verdict. intervals.icu does not editorialise at all, which is its virtue. You can pull HRV into a custom chart, plot it against your own 7-day and 30-day rolling means with the built-in formula fields, and read it yourself. That is the right level of interpretive restraint for a tool.

Whoop’s Strain Coach and Garmin’s Training Readiness are where it goes wrong. Both collapse HRV, sleep, RHR and load into one integer and then attach a verb to it. Garmin will tell you Readiness 34, “consider an easy day”, on the back of a single suppressed morning. If you follow that instruction across a 12-week build, you will skip roughly one in six quality sessions for no physiological reason, and the sessions you skip will cluster in exactly the weeks where hard training is producing adaptation.

The LLM layer is worse, because it will confabulate confidence you did not ask for. Paste a daily HRV figure into ChatGPT or Claude with no baseline and you get plausible, fluent, useless output. The fix is to change what you paste.

A prompt that survives contact with your own files

Bad prompt: “My HRV is 42 today, down from 58. Should I do my VO2max session?”

Better prompt, and one worth saving as a template:

You are analysing HRV as a proxy for parasympathetic reactivation, not
as a composite recovery score. Do not use the word "recovery" for rMSSD.

My data (morning supine rMSSD, ms, most recent last):
71, 48, 52, 59, 44, 41, 39, 37

Baseline: 30-day mean 65ms, within-athlete SD 11.7ms.
Normal range: 53-77ms (mean +/- 1SD).
Training load: CTL 78, ATL 94, TSB -16.
Last 7 days: 5x5 VO2max Tue, 2x20 @FTP Thu, 4hr Sat, 3hr Sun.
Planned today: 4x8 @105% FTP.

Tasks:
1. Compute the 7-day rolling mean and state whether it sits inside
   the normal range.
2. Distinguish single-day variation from a trend shift. Say explicitly
   which days are noise.
3. State what rMSSD cannot tell me about my readiness for 4x8.
4. Recommend: proceed, modify, or postpone, and give the reasoning
   in terms of the trend plus the load data, not the single value.

Run that and you get something checkable. It has to produce a rolling mean you can verify by hand, and it has to say out loud what the number cannot cover. If a model tells you to rest purely because today is 37, you have caught it reasoning from the wrong variable, and you can discard the output on the spot.

The measurement conditions matter more than most people admit

Morning supine, immediately on waking, before you look at your phone, same position, 2-3 minutes minimum after a 1-minute settling period. Deviate and you are measuring something else. Standing or seated readings shift rMSSD substantially downward because vagal tone drops with orthostatic challenge, so a seated reading compared against a supine baseline reads as suppression that does not exist.

Breathing rate is the other confounder nobody mentions. HRV in the high-frequency band is largely respiratory sinus arrhythmia, so if you consciously slow your breathing during the reading you will inflate rMSSD, sometimes dramatically. HRV4Training explicitly warns about this. Breathe normally, or you are measuring your breathing technique.

Wrist optical sensors sample less reliably than chest straps or finger PPG, and they sample during sleep rather than in a standardised morning window, which means they are picking up whatever autonomic state your 03:00 REM cycle happened to produce. Garmin’s overnight HRV and a Polar H10 morning supine reading are not interchangeable inputs to the same analysis. Pick one method, keep it for 60 days, then start interpreting.

What to do with the days HRV cannot speak to

The honest answer is that you need a second and third input, and the pillar piece on recovery, HRV and readiness scores works through how these stack. In practice: subjective session RPE against expected RPE, morning resting heart rate, and power at a fixed sub-threshold heart rate.

That last one is the most underrated tool available to a rider with a power meter. Ride 20 minutes at 140bpm and record the average power. Do it weekly, same route or same turbo resistance, same time of day. If that number drifts down 8 watts while your HRV trend sits comfortably in range, you have found fatigue that HRV structurally cannot see, because cardiac vagal tone and peripheral muscular function are different systems with different time courses.

Conversely, a suppressed HRV trend with unchanged power at fixed HR usually points somewhere outside the bike: a cold coming on, a bad fortnight at work, three nights of five hours’ sleep. Useful to know. Not necessarily a reason to change Thursday’s session.

The rider who gets value out of HRV is the one who spent 60 days collecting baseline before acting on anything, who computes a rolling mean rather than reading a daily verdict, and who treats a trend break as a prompt to go looking at the rest of their evidence. The rider who gets nothing out of it is the one who let an app turn a measure of vagal outflow into a permission slip.