A Morning HRV Protocol That Produces Signal Instead Of Noise
Most cyclists who quit HRV tracking didn’t quit because HRV doesn’t work. They quit because their numbers bounced 40ms in either direction for no reason they could name, the app told them to take an easy day after a rest day, and the whole thing started to feel like astrology with a Bluetooth chest strap.
Here’s the thing nobody selling you a readiness score wants to lead with: on any given morning, the conditions under which you measured explain more of the difference between today’s number and yesterday’s than your actual autonomic state does. Posture, time since waking, bladder, whether you spoke, caffeine timing, room temperature, measurement duration, and which algorithm smoothed the result. Get those wrong and you’re not measuring recovery. You’re measuring how you got out of bed.
So the question isn’t really how to measure hrv cycling adaptations at all. It’s how to hold everything else still enough that the autonomic signal is the only thing left moving.
What the variance actually looks like
Let me give you real numbers, because vague warnings about “confounders” don’t change behaviour.
Take a rider with a typical rMSSD baseline around 62ms. Here’s what a single morning can look like depending on nothing but protocol:
| Condition | rMSSD (ms) | Delta vs. supine baseline |
|---|---|---|
| Supine, 5 min after waking, no movement | 64 | — |
| Seated, immediately after standing up | 38 | −26 |
| Supine, but 40 min after waking (bathroom, kettle on) | 71 | +7 |
| Supine, 90 min after a double espresso | 51 | −13 |
| Seated, talking to partner during measurement | 44 | −20 |
| Supine, room at 12°C, no duvet | 55 | −9 |
Every one of those readings is “true.” The strap didn’t lie. But if you measure supine on Monday and seated on Tuesday, you’ve generated a 26ms swing that has nothing to do with Sunday’s 4-hour endurance ride. An app that sees −26ms will paint the day red, and you’ll shelve a session you were fully capable of doing.
Compare that against the size of a genuine training signal. A hard block that’s genuinely overreaching you tends to move a 7-day rMSSD average by something like 8-15%: for our 62ms rider, that’s a shift to roughly 53-57ms sustained over several days. That’s smaller than the posture artefact. Smaller than the caffeine artefact. Roughly the same size as the “I got up and made tea first” artefact.
The signal you’re hunting is quieter than the noise you’re generating. That’s the whole problem, and it’s a protocol problem before it’s an interpretation problem.
The protocol
This is what I’d have you run, and the ordering matters more than any individual item.
Wake up. Don’t get up. Reach for the strap or the phone before your feet touch the floor. If you need to use the bathroom, go, then get back into bed and lie still for 3 minutes before you start the recording. A full bladder measurably suppresses HRV, so emptying it and then re-settling is better than pushing through.
Stay supine, flat, one pillow. Not propped up on the headboard scrolling. Arms at your sides. The single biggest source of morning-to-morning junk is postural inconsistency, and supine is easier to replicate exactly than seated (there are a dozen ways to sit, one way to lie flat).
Don’t talk. Don’t look at your phone screen. Speech alters breathing rhythm, which directly drives respiratory sinus arrhythmia, which is most of what rMSSD picks up at rest. If you’re using a phone camera app, start it and put the phone face-down on your chest or look at the ceiling.
Breathe normally. Do not pace your breathing. This one gets argued about. Paced breathing at 6 breaths/min produces beautiful, high, stable numbers that are mostly a measurement of your breathing, not your nervous system. If you pace, pace every single day forever, or don’t pace at all. I’d say don’t: spontaneous breathing at rest is the condition most of the validation literature sits on.
Record for 60 seconds of clean data after a 60-second settle. Ultra-short recordings (60s) correlate strongly with the 5-minute gold standard for rMSSD specifically. Not for LF/HF, not for SDNN, and this is one reason rMSSD is the metric to use. HRV4Training defaults to 60s, Elite HRV to 2:30, Kubios to 5 min. Any is fine. Pick one and never change it.
Same time, same day-of-week-agnostic. Within a 30-minute window. If you normally wake at 06:15 and one morning you sleep to 09:00, record it and flag it, but know the reading isn’t comparable.
Log the conditions, not just the number. Alcohol the night before, illness, poor sleep, late meal, travel, a hot room. Every good app has a tag field. Use it, because those tags become the thing that makes your data interpretable six weeks later.
Chest strap or phone camera
Both work, neither mixes. A Polar H10 with HRV Logger or HRV4Training gives you beat-to-beat RR intervals with very low artefact rates. A phone camera PPG measurement (HRV4Training’s camera mode, Welltory) is genuinely validated against ECG for rMSSD at rest, and the convenience means you’ll actually do it daily, which beats a more accurate method you skip twice a week.
What you cannot do is alternate. PPG and ECG produce systematically different rMSSD values, typically offset by a few ms with different artefact behaviour. Switching devices mid-baseline resets your baseline whether you acknowledge it or not.
Wrist-worn overnight HRV (Whoop, Garmin, Oura) is a different animal again. It’s averaging across sleep stages and using a proprietary window, usually deep sleep or the last portion of the night. It’s not wrong, and the night-to-night consistency of a fixed automated window is genuinely an advantage over a manual morning test. But the numbers aren’t comparable to a morning rMSSD reading, and the readiness score sitting on top of them is a black box with sleep, respiratory rate and resting HR baked in. If you’re going to use a wearable, use the raw HRV number it exposes, not the composite score.
The baseline length nobody wants to hear
You need 60 days before you act on a single day’s reading.
Not 7. Not the two weeks most apps use to start showing you a “normal range.” Sixty.
Here’s why. To flag a meaningful deviation you need a normal range, and a normal range needs a standard deviation that isn’t itself noisy. With 7 days you’re estimating SD from seven points, and the confidence interval on that estimate is so wide that your “normal range” moves around more than your HRV does. Around 30 days the estimate stabilises enough to be useful. By 60 you’ve captured a training block, a rest week, a bad night’s sleep, a few beers, possibly a minor cold, and the normal variation of your own life.
In the meantime, measure daily and look at nothing. That’s the discipline. You’re building a reference, not reading a dashboard.
Once you’re past 60 days, here’s the decision rule that actually holds up:
7-day rolling mean rMSSD: 58.2 ms
60-day mean: 61.8 ms
60-day SD: 6.4 ms
Normal range (±1 SD): 55.4 – 68.2 ms
Today (single day): 49.1 ms → outside range, 1 day
Yesterday: 63.0 ms → inside range
7-day mean vs 60-day mean: -5.8% → inside normal fluctuation
Verdict: train as planned.
One day below the range means nothing. A single low reading is the most common thing in HRV data and the most common cause of unnecessary rest days. What you’re watching for is the 7-day rolling mean dropping below the lower bound of the normal range and staying there, or three-plus consecutive days below the range. That’s a pattern. One red morning is weather.
intervals.icu handles this reasonably well if you feed it HRV via the wellness API or the HRV4Training sync, and you can set the baseline period in the settings rather than accepting a 7-day default. It’ll also plot HRV against your CTL and ramp rate, which is where the interesting reading happens: a 7-day HRV mean sagging while CTL climbs hard is a different story from the same sag during a taper. The recovery and readiness pillar goes further into how those composite scores are constructed and where they break.
Where AI tools help and where they hallucinate
Feed an LLM your raw morning numbers and ask “am I recovered” and you’ll get confident nonsense. It has no baseline, no SD, no idea what your normal range is, and it will pattern-match to generic advice about listening to your body.
Feed it the summary statistics and a specific question and it becomes genuinely useful. Something like:
My 60-day rMSSD mean is 61.8ms with SD 6.4ms, measured supine, 60s, Polar H10, HRV4Training. Last 10 days: 63, 59, 64, 49, 58, 57, 55, 56, 54, 53. CTL has gone from 71 to 84 over those 10 days. Ramp rate 4.2 TSS/day/week. I have a 3x20min threshold session scheduled tomorrow. Based on the trend direction and the ramp rate, what’s the case for doing it as planned versus swapping to endurance, and what would I need to see tomorrow morning to make that call?
That prompt works because you’ve done the statistical work and you’re asking for reasoning about a decision, not asking the model to be your sensor. The data is structured, the protocol is stated (so the model knows the numbers are internally comparable), and the question has a defined answer space.
The failure mode to watch for: models will happily invent thresholds. If something tells you “rMSSD below 50 indicates overtraining,” that number came from nowhere. HRV is only meaningful relative to your own baseline. A 40ms rider and a 95ms rider can both be perfectly recovered, and the between-person spread is enormous. Ask any tool that gives you an absolute threshold to show you which baseline it’s relative to, and if it can’t, discard it.
The tools that hold up against your own ride files are the ones that show their working: the baseline window, the SD, the number of days in the calculation. HRV4Training does this. intervals.icu does this. A readiness score rendered as a single number out of 100 does not, and you can’t audit what you can’t see.
Start tomorrow morning. Strap on, flat on your back, sixty seconds, log it, close the app. Do that until late November and you’ll have something worth reasoning about.