Whoop, Oura Or A Chest Strap: Which HRV Source Survives Scrutiny?
Three devices, one wrist, one finger, one sternum, sixteen consecutive nights. I did this because I got tired of arguing about it in forum threads where nobody posts their raw data. The question sounds like a gear question. It isn’t. After you line up the numbers, the thing that actually determines whether your HRV trend means anything is when you measure, not what you measure with. But there’s a catch, and it lands hardest on exactly the sort of rider who reads this site: if you’re lean, your wrist-based optical data is the noisiest of the three by a wide margin.
Let me show you the numbers before I argue about them.
The setup
Whoop 4.0 on the left wrist, Oura Ring Gen 4 on the right index finger, Polar H10 chest strap worn overnight for eight of the sixteen nights and used every single morning for a five-minute seated capture through EliteHRV. All RMSSD values in milliseconds, ln-transformed where the platform does that natively (Whoop and Oura both report raw RMSSD; intervals.icu takes whatever you feed it). Rider: me, 71 kg, roughly 8% body fat in September, FTP 312 W, training 12 to 14 hours a week through a gravel block.
Here’s a representative week, overnight values only:
Night Whoop Oura H10 (overnight) Whoop-H10 delta
Mon 68 74 76 -8
Tue 91 88 87 +4
Wed 54 71 73 -19
Thu 77 79 81 -4
Fri 62 80 82 -20
Sat 104 95 93 +11
Sun 71 76 78 -7
Mean absolute error, Whoop against H10, across all sixteen nights: 9.4 ms. Oura against H10: 3.1 ms. Whoop’s worst single night was 21 ms low. Oura’s worst was 8 ms high, on a night I’d had three glasses of wine and slept badly, which is exactly when you’d expect any sensor to struggle with motion artefact.
That Whoop error isn’t random. It’s directionally biased low, and it clusters on the nights I slept coolest. Peripheral vasoconstriction in a cold bedroom at a low body fat percentage gives a wrist PPG sensor very little to work with. The signal-to-noise ratio at the radial artery through a lean forearm in a 16°C room is genuinely poor, and Whoop’s algorithm fills the gap with interpolation you never see.
Why the finger beats the wrist
Fingertips are dense with arteriovenous anastomoses. Blood flow there is high and, crucially for a PPG sensor, relatively stable across the sleep cycle. The wrist has none of that. You’re reading through skin, tendon and whatever fat layer you have (if you’re a cyclist in September, you don’t have much), at a site that moves every time you roll over.
Oura’s advantage in my data isn’t algorithmic sophistication. It’s anatomy. A friend of mine at 24% body fat ran the same comparison over ten nights and got a Whoop-to-H10 MAE of 4.8 ms. Half of mine. Same device, same firmware, different arm.
So if you’re a 62 kg climber asking about the best hrv device cycling setup, the honest ranking on raw measurement fidelity is: chest strap, then ring, then wrist, and the gap between ring and wrist widens the leaner you get. That’s the device part of the answer, and it’s the less interesting part.
The part that actually matters
Here’s the experiment that changed how I think about this. For eight days I took three morning readings: one immediately on waking while still lying down, one after standing and using the bathroom (roughly six minutes later), and one seated after making coffee (roughly eighteen minutes later). All three with the same H10, same app, same five-minute window.
Day Supine Post-void Seated+coffee
1 96 71 64
2 88 66 61
3 102 78 69
4 79 59 55
5 94 70 66
6 85 64 58
7 99 74 67
8 91 68 62
Look at the spread. Supine-to-seated difference averages 28 ms, and it’s consistent day to day (coefficient of variation on the difference: 7%). Now imagine you take the supine reading on Monday, the seated one on Tuesday because you overslept, and the post-void one on Wednesday. Your “HRV” swings by 30 ms across three days on which your actual autonomic state barely moved. You’d read that as a crash. You’d skip your Thursday intervals. You’d be wrong.
The day-to-day biological variation in true RMSSD for a trained cyclist sits around 10 to 15% of the mean. My posture-induced variation was over 30%. The measurement protocol is producing more variance than the physiology you’re trying to observe. That is the whole ballgame, and no device solves it for you.
What consistency buys you
Across the sixteen-night comparison, I ran the same trend analysis on all three data sources: seven-day rolling mean, plotted against a 28-day baseline, flagging any day where the 7-day dropped more than one standard deviation below baseline.
Whoop flagged four days. Oura flagged three. H10 overnight flagged three. Of those, all three sources agreed on two days: the Tuesday after a 4h 20m gravel ride with 2,100 m of climbing, and the Saturday I came down with a cold. The days they disagreed on were all days where the flag sat right at the threshold.
In other words, the noisy device and the clean device told me the same story about the two things that mattered. Whoop’s 9.4 ms nightly error mostly washed out in the rolling average, because the error was roughly symmetrical around a consistent bias once I stopped looking at single nights. A consistent bias is survivable. An inconsistent sampling window is not.
This is why I keep pushing people towards the trend rather than the number, and it’s covered in more depth in our pillar on recovery, HRV and readiness scores if you want the physiology behind why a single morning value tells you close to nothing.
Feeding it into a system that can use it
Raw HRV sitting in a proprietary app is a diary entry. Getting it next to your training load is where it becomes useful, and this is where the device choice has practical consequences beyond sensor quality.
intervals.icu ingests Oura natively through OAuth: connect it under Settings, Integrations, and your HRV, resting HR and sleep score land in the wellness table every morning without you touching anything. Whoop also connects directly. If you’re using a chest strap and EliteHRV or HRV4Training, you’re exporting CSV and using the intervals.icu wellness API, which looks like this:
curl -X POST \
"https://intervals.icu/api/v1/athlete/i12345/wellness" \
-u "API_KEY:your_key_here" \
-H "Content-Type: application/json" \
-d '{"id":"2026-09-29","hrv":76,"restingHR":42,"sleepSecs":27000}'
Once it’s in there, the AI-assisted part gets genuinely useful. I pull a 90-day wellness export alongside my activity data and hand both to Claude with a prompt like this:
Here are 90 days of wellness data (date, HRV RMSSD, resting HR, sleep hours) and 90 days of training data (date, TSS, duration, IF). Calculate the 7-day rolling mean and 28-day baseline for HRV. Identify every day where the 7-day mean dropped more than 1 SD below the 28-day baseline. For each of those days, report the preceding 3-day TSS total and whether a rest day followed. Do not interpret. Just give me the table.
The “do not interpret” instruction matters. Ask an LLM whether you’re overtrained and it’ll hedge and reassure. Ask it to compute a specific relationship across your files and it does arithmetic you’d otherwise do badly in a spreadsheet. On my data, that query surfaced something I’d missed: every one of my flagged suppression days followed a 3-day TSS block above 380, but only about half of my 380+ blocks produced a flag. The discriminator turned out to be whether the block contained a ride over three hours. Volume, not intensity. I’d have bet the other way.
What I’d actually tell you to buy
If you’re lean and you want data you can trust night to night, the ring wins on fidelity for roughly £350 plus a subscription. If you already own a Polar H10 or a Garmin HRM-Pro, you own the reference standard, and a five-minute morning protocol costs you nothing but discipline.
Whoop is not useless. My Whoop trend caught the same two meaningful events as the strap. It’s just that the individual nightly numbers, for a rider at my body composition, carry error of the same magnitude as the signal. If you look at your daily HRV number and make decisions from it, that’s a problem. If you look at your 7-day line and your 28-day baseline, it mostly isn’t.
The protocol I settled on: H10, seated, five minutes, after the bathroom and before caffeine, every morning, no exceptions. Logged to intervals.icu. Takes seven minutes including faff. My coefficient of variation on the morning reading dropped from 19% (when I was sloppy about timing) to 8%, without changing a single piece of hardware.
Pick whichever device you’ll actually wear or strap on every day without resenting it. Then be religious about the window. A cheap strap used identically at 06:40 every morning will out-inform a premium wearable whose sampling window drifts with your bedtime, and if you’re 68 kg with visible veins in your forearms, the wrist sensor is fighting your own physiology every night it tries.