Garmin Training Readiness And Body Battery: An Honest Audit
I spent eleven weeks logging what my Garmin told me every morning against what my legs actually did in the first interval of the day. The conclusion is narrower than the marketing and more useful than the cynicism: Training Readiness is a good sleep-debt detector wearing a training-load costume.
That distinction matters if you’re self-coached. You look at a number between 1 and 100 at 6:40am and decide whether today is 5x8 at threshold or a two-hour endurance ride. If the number is tracking one thing and you believe it’s tracking another, you’ll make the wrong call in a specific, repeatable direction: you’ll bin quality sessions after bad nights that your legs would have coped with fine, and you’ll green-light hard days after heavy blocks where your legs are cooked but you slept beautifully.
The setup, so you can argue with the method
Forerunner 965, worn 24/7 apart from charging during the post-ride shower. Assioma Duo pedals on the road bike, Favero Assioma on the TT bike, Zwift Hub for indoor. All files into intervals.icu, Garmin Connect left untouched so its own algorithms saw everything they expect to see.
Eleven weeks, mid-January to early April, building for a 25-mile TT and a spring gravel event. 47 sessions with a defined interval target. Weekly TSS ranged 380 to 720 (intervals.icu numbers, my FTP set at 291W then 298W from week seven).
Every morning I wrote down, before riding, four things from the watch: Training Readiness (the composite score), Body Battery at wake, HRV status (Balanced / Unbalanced / Low), and sleep score. Then after the session I graded execution on a three-point scale that I defined in advance:
- Hit: completed the prescribed intervals within 2% of target average power, no shortened reps.
- Partial: completed but had to drop a rep, or averaged 2 to 6% under target.
- Failed: abandoned the session or came in more than 6% under.
Crude, but it has to be crude to be honest. Anything more sophisticated and I’d be fitting the grading to the result.
What Readiness got right: sleep debt
Here’s the split by readiness band across the 47 target sessions.
Readiness n Hit Partial Failed Hit rate
1-25 6 2 2 2 33%
26-50 14 8 4 2 57%
51-75 19 15 3 1 79%
76-100 8 7 1 0 88%
At first glance that looks like a working score. Monotonic, sensible gradient, the low band genuinely predicting trouble. I nearly stopped there and wrote a much more flattering post.
Then I pulled apart why readiness was low on those days. Garmin exposes the contributing factors in Connect, and on every single one of the six sub-25 mornings the dominant negative factor was sleep or sleep history, not acute load. Week four, Thursday: readiness 19, sleep score 44, four hours fifty of sleep because a kid had a temperature. Week nine, Tuesday: readiness 22 after a 2:10am finish on a work deadline. The score was right, and it was right because I had slept badly, which I already knew without a wrist computer.
Strip the sleep-driven mornings out and the picture changes completely.
What it got wrong: training stress
Of the 47 sessions, I flagged 12 as “deep block” days in advance: third or fourth consecutive day of loading, or arriving on a seven-day rolling TSS above 600 with at least two quality days already in it. These are the days where you most want a score to tell you something.
Deep-block days (n=12), sleep score ≥ 75 on all of them
Morning readiness Execution
68 Failed (5x8 @ 300W, abandoned rep 3, avg 281W)
71 Partial (3x12 @ 295W, dropped to 2 reps)
64 Partial
77 Failed (40/20s, quit set 2 of 4)
59 Partial
81 Hit
66 Partial
73 Failed (TT sim 25min, avg 274W vs 291W target)
70 Hit
62 Partial
75 Partial
69 Hit
Three hits out of twelve, on an average morning readiness of 69.6. The 51-75 band overall had a 79% hit rate. Within deep-block days, the same band produced 25%. The score simply does not see accumulated training stress with anything like the weight it deserves.
The 73 is the one that annoyed me most. Week eight, Wednesday, fourth day of loading, 641 TSS in the prior seven days. Sleep score 82, Body Battery 87 at wake, HRV Balanced, readiness 73 with “Good” as the label and a little note suggesting I was ready for a productive session. I went out to do a 25-minute TT effort at 291W and averaged 274W with my heart rate 6bpm lower than normal for that power, which is the classic signature of a rider who is fatigued rather than unmotivated. Nothing in the morning data hinted at it. My intervals.icu form value that day was -31 and my legs knew it even if Garmin didn’t.
Body Battery is a worse instrument than readiness, dressed more convincingly
Body Battery has the most intuitive interface of anything on the watch. A phone-style battery, draining and charging. It’s also the metric I’d trust least for training decisions, and the reason is structural: it’s built from stress (which is built from HRV) plus activity plus sleep, and nothing in that stack has a memory longer than about a day.
The eleven-week correlation between wake Body Battery and session execution, scoring Hit=2, Partial=1, Failed=0, came out at r = 0.31. Readiness managed r = 0.48. Seven-day rolling TSS from intervals.icu, on its own, with no physiology in it at all, got r = -0.52. A number derived purely from what you did last week predicted today’s interval quality better than either wrist metric.
Worse, Body Battery recharges on the thing you least want it rewarding. Two examples from the same fortnight:
Week six, Saturday. Four-hour endurance ride, 3,100kJ, 218 TSS. Ate properly, got to bed at 22:15, slept 8h05. Sunday wake Body Battery: 91. Readiness 78. Legs on Sunday’s 3x15 tempo: fine for rep one, ragged by rep three, finished at 268W against a 275W target.
Week seven, Tuesday. One hour easy, 45 TSS, but a genuinely stressful day at work and 6h30 sleep. Wake Body Battery Wednesday: 54. Readiness 41. I ignored it and did 5x8 at 300W. Hit every rep, last one at 304W.
Both calls were wrong, and wrong in the same direction: good sleep after big training reads as recovered, bad sleep after nothing reads as wrecked. If you’re trying to understand how these scores relate to each other and to your own numbers, the broader picture in recovery, HRV and readiness scores is worth reading alongside this, because the failure mode I’m describing is not unique to Garmin.
Why this happens, mechanically
Garmin’s readiness is a weighted composite. Public documentation and the factor list in Connect give you: sleep (last night), sleep history (roughly three days), recovery time, HRV status, acute load, and stress history. Four of those six are sleep or autonomic. Only recovery time and acute load carry training, and recovery time is itself derived largely from the physiological effect estimate of your last activity, not from cumulative strain.
The practical consequence: a fourth consecutive loading day and a rest day can produce near-identical readiness if you slept the same both nights. I have that in my data. Week eight Wednesday (fourth loading day, readiness 73) and week ten Monday (day after a rest day, readiness 76). Three points apart. Actual capacity that morning was not three points apart.
HRV status is the part that could theoretically catch training stress, and it partly does, but Garmin’s seven-day-average-against-baseline implementation is slow. My HRV status went Unbalanced exactly twice in eleven weeks, both times three or four days after the block had already hurt. Useful for confirming you overdid it. Not useful for deciding whether to start the session.
What I’d actually do with it
Keep the watch. Change what you ask it.
Treat readiness below 35 as a real signal and check the factor breakdown in Connect: if sleep is the driver, that’s trustworthy and you should probably convert the day to endurance. Below 35 was 6 out of 47 mornings for me, all sleep-driven, and it was right or near enough right on all six.
Between 35 and 85, which was 36 of my 47 mornings, the score carries almost no information about training stress. In that band I now look at three other things instead:
- Seven-day rolling TSS and form in intervals.icu. Form below -25 on a day I’m asking for threshold work is a much better warning than anything on the watch.
- Consecutive loading days. Day four is day four regardless of how well I slept.
- The first three minutes of the warm-up. Sounds unscientific. It has a better hit rate than Body Battery. 290W in the warm-up ramp at a heart rate 5bpm below normal means the session is going badly and I now bail early rather than salvage.
The fix on the AI side, if you want one, is to stop asking a coaching model about your readiness score and start giving it the inputs. I’ve had far better results pasting a 14-day table of date, TSS, IF, duration, sleep hours and subjective legs score into Claude or ChatGPT and asking it to identify which days were genuinely compromised, than asking either to interpret a readiness number in isolation. The model has no idea what Garmin’s weightings are, so when you hand it a score you’re asking it to reason about a black box. Hand it the ride files and the sleep hours and it reasons about physiology.
One prompt that’s earned its place in my notes, run weekly against a CSV export from intervals.icu:
Here are my last 14 days: date, TSS, IF, duration, average power, normalised power, sleep hours, and a 1-5 subjective legs rating I wrote before each ride. Identify days where my power output was below what the prescribed intensity should have produced, and tell me which of the preceding variables best explains each shortfall. Do not use any wearable recovery score. Be specific about which days you’d have moved or cut.
It told me, unprompted, that my shortfalls clustered on the fourth day of loading blocks and had almost no relationship to my sleep hours, which is the same conclusion this audit reached the slow way.
The honest version of Garmin Training Readiness accuracy is this: it is a well-built sleep and autonomic monitor with a training-load label on the tin. Given that it was designed by a company selling to people who mostly want to know if they slept badly, that’s not a scandal. It’s just not the thing a self-coached cyclist in week four of a build needs, and the gap is widest exactly when the decision matters most.