AI Cycling
§4.1 Recovery, HRV And Readiness Scores 1,798 words · 8 min

Whoop, Oura And Garmin HRV Compared For Cyclists

Three devices, three wrists (or fingers), three different numbers claiming to represent the same physiological signal. If you’ve ever worn a Whoop 4.0 on one arm and an Oura Ring on the other while your Garmin Fenix logged overnight HRV from the same night’s sleep, you’ll have noticed the numbers don’t match. Not even close. One says 68ms, another says 41ms, the third gives you a “balanced” status with no number at all.

This page is about what those differences actually mean when you’re deciding whether to do Thursday’s threshold session. If you want the wider framing on how readiness scores fit into a training week, the recovery, HRV and readiness pillar covers that ground. What follows is narrower: the measurement methods, the numbers each device produces from the same body, and how to feed them into an AI coaching workflow without garbage propagating downstream.

The metric each device actually calculates

The first thing to understand is that “HRV” is not one number. It’s a family of metrics derived from the intervals between heartbeats, and the three devices don’t all report the same member of that family.

Whoop reports RMSSD (root mean square of successive differences), calculated across your last slow-wave sleep period. Not the whole night. Whoop specifically targets SWS because parasympathetic activity is most stable there and less contaminated by REM-related sympathetic surges. The window is typically 20 to 40 minutes long.

Oura also reports RMSSD, but as an average across the entire sleep period, sampled in 5-minute windows. Oura additionally exposes the full night’s curve in the app, so you can see HRV rise through the night (a good sign) or stay flat and depressed (less good). Because it’s a whole-night average including REM, Oura’s numbers tend to run lower than Whoop’s for the same person on the same night.

Garmin is the odd one out. Garmin’s HRV Status, available on Fenix 7/8, Forerunner 265/965/970, Epix and Venu 3, also uses RMSSD, measured in 5-minute windows across the whole night, then reports a 7-day rolling average rather than a single-night value. The number you see on the watch face is not last night. It’s the average of the last seven nights, plotted against a personal baseline built from 90 days of data.

That last point causes more confusion than anything else. A cyclist comparing “my Whoop said 52 and my Garmin said 38” is comparing a single-night SWS RMSSD against a seven-night whole-night rolling mean. They should be different.

What the numbers look like side by side

Here’s a representative week from a 38-year-old male cat 2 road racer, FTP 312W, wearing all three during a build block. Values in milliseconds.

NightWhoop (SWS RMSSD)Oura (night avg RMSSD)Garmin (7d avg)Previous day
Mon644744Rest
Tue5843443×12 @ 105%
Wed4131424×8 VO2max
Thu393040Endurance 2h
Fri554140Rest
Sat614541Endurance 3h
Sun483641Race, 4h

Notice three things. Whoop runs roughly 30 to 35 percent higher than Oura across every single night, which is exactly what the SWS-vs-whole-night methodology predicts. The ratio between them is remarkably stable, which is the useful part: both are measuring the same underlying signal with a consistent offset. Garmin barely moves, because a seven-day average is designed not to move. The Wednesday crash after VO2max work shows up as a 23ms drop on Whoop and a 12ms drop on Oura, but only a 2ms nudge on Garmin.

For day-to-day session decisions, Garmin’s rolling average is close to useless. For spotting a three-week accumulation of fatigue during a block, it’s arguably the most honest of the three, because it won’t spook you over one bad night’s sleep after a late dinner and two glasses of red.

Measurement site and why it matters on a bike

Oura measures from the finger, using infrared PPG against the palmar digital arteries. Whoop and Garmin measure from the wrist, using green LED PPG against capillary beds. The finger has denser perfusion and less tissue between sensor and artery, so signal quality is generally better, especially for someone with lower body fat or prominent wrist tendons.

In practice this means Oura loses fewer nights to bad data. Across a 90-day comparison, it’s common to see Oura return a usable HRV value on 88 to 90 nights, Whoop on 84 to 88, and Garmin on 70 to 80. Garmin’s gaps come from two sources: the watch has to be worn snugly overnight (many people don’t), and Garmin requires at least four hours of sleep detection to compute the night.

There’s a cycling-specific problem with wrist devices that nobody mentions in the marketing. If you’ve done a long ride in a stiff aero position, particularly TT work, wrist and forearm compression can degrade overnight PPG quality for hours afterward. If your Whoop returns a suspiciously low value the night after a two-hour TT block but you feel fine, check the sleep-stage confidence in the app before you take the number seriously.

The consistency test: run this before you trust any of them

Before you wire any of these into an AI coaching prompt, establish whether the device tracks your own baseline sensibly. Two weeks of data is enough for a first pass.

Take 14 consecutive nights. Calculate your mean and standard deviation. Then check where your HRV sits relative to that on nights following your three hardest sessions.

Whoop, 14 nights:
  mean  = 53.2 ms
  sd    =  8.9 ms
  CV    = 16.7%

  Post-hard-session nights: 41, 39, 44  →  z = -1.37, -1.60, -1.03
  Post-rest nights:         64, 61, 62  →  z = +1.21, +0.88, +0.99

Oura, same 14 nights:
  mean  = 39.8 ms
  sd    =  6.4 ms
  CV    = 16.1%

  Post-hard-session nights: 31, 30, 33  →  z = -1.38, -1.53, -1.06
  Post-rest nights:         47, 45, 46  →  z = +1.13, +0.81, +0.97

The z-scores are almost identical. That’s the signal you’re looking for: both devices, despite reporting different absolute values, place the same nights in the same relative position. If one device gives you z-scores that don’t correlate with training load at all, it isn’t measuring you well and no amount of AI analysis will fix that.

A coefficient of variation between 10 and 20 percent is normal for a trained endurance athlete. Below 8 percent usually means the device is smoothing more than it admits. Above 25 percent, look at your sleep consistency and alcohol intake before you blame the hardware.

Getting the data somewhere an AI can use it

This is where the practical differences bite hardest.

Oura has a genuinely good public API. Personal access tokens are free from cloud.ouraring.com, and /v2/usercollection/daily_readiness plus /v2/usercollection/sleep gives you RMSSD, the 5-minute HRV array, resting heart rate and temperature deviation as clean JSON. You can pull 90 days in one call.

Whoop’s API v2 requires an OAuth app registration at developer.whoop.com. Free, but more setup. The /v2/recovery endpoint returns hrv_rmssd_milli, resting_heart_rate and recovery_score. Rate limit is 100 requests per minute, which is ample.

Garmin is the painful one. There’s no free personal API for HRV Status. The Garmin Connect Developer Program is aimed at commercial partners. Your realistic routes are the unofficial garminconnect Python library, a Health API partnership, or exporting through intervals.icu, which ingests Garmin wellness data and exposes it via its own API at /api/v1/athlete/{id}/wellness. That last option is the one most self-coached riders should take, because intervals.icu also holds your power data, so HRV and TSS end up in the same place.

Once the data is somewhere queryable, the prompt matters more than the model. Here’s a version that works reliably with Claude or GPT-class models:

You are analysing HRV data for a self-coached cyclist.

Device: Whoop 4.0 (RMSSD from slow-wave sleep, single night)
Baseline: 60-day rolling mean 53.2 ms, SD 8.9 ms
Last 7 nights: 64, 58, 41, 39, 55, 61, 48
Last 7 days CTL/ATL (intervals.icu): 78/61, 79/72, 81/88,
  82/84, 80/71, 82/79, 85/96

Planned tomorrow: 4 x 8 min at 110% FTP

Tasks:
1. Convert each night to a z-score against the stated baseline.
2. Identify whether the 7-night trend is suppressed, stable or rebounding.
3. State whether the planned session should proceed, be modified, or
   be replaced, and say which of the three you are recommending.
4. Name explicitly what would change your recommendation.

Do not average across devices. Do not compare absolute ms values
to population norms. Use only the baseline supplied.

The last two lines carry a lot of weight. Left to itself, a model will happily tell you that 41ms is “below average for athletes,” which is meaningless, because population HRV norms span roughly 20ms to 120ms in trained cyclists and tell you nothing about your own state. Anchoring the model to your personal baseline is the single highest-value constraint in the prompt.

Which one for which rider

If you race and make session-level decisions most mornings, Whoop’s single-night SWS reading is the most responsive of the three, and the recovery score (0 to 100) is at least transparent about its inputs: HRV, resting HR, respiratory rate, sleep performance. The subscription is roughly £229 a year with no upfront hardware cost.

Riders who want the best raw data quality and don’t want another wrist strap should take Oura. The Ring 4 is about £349 plus £5.99 a month, the API is the best of the three, and finger PPG genuinely beats wrist PPG on signal integrity. The tradeoff: the readiness score leans heavily on sleep and temperature, which makes it excellent at flagging illness and mediocre at flagging training fatigue specifically.

Already own a Fenix or Forerunner and don’t want a second subscription? Garmin’s HRV Status is free and perfectly adequate for trend work. Pair it with the morning orthostatic test in HRV4Training or EliteHRV if you want a responsive daily number, and let the watch handle the seven-day picture. Two signals at two timescales is a better setup than one device trying to do both.

Worth saying plainly: buying a second device to cross-check the first rarely improves decisions. The z-score comparison above showed Whoop and Oura ranking nights identically. What actually improves decisions is 60 days of consistent data from one device, a personal baseline you trust, and a habit of writing down how your legs felt so you can check the number against the reality it’s supposed to predict.