Why LLMs Invent Numbers From Your Ride Files — And How To Catch It
You paste a FIT export into a chat window, ask for your normalised power, and get back NP 271 W. It looks right. It sits plausibly above your average of 248 W, the ratio is sane, and you file it away as a hard tempo ride.
The problem is that the model never computed it. There was no 30-second rolling mean, no fourth power, no fourth root. What happened was a language model reading “average power 248” and generating the next most likely token sequence for a paragraph about normalised power. The number 271 is a guess dressed in the syntax of a calculation.
This is the core of the llm hallucination power data problem for cyclists, and it is worse than generic hallucination because the output is numerically plausible. When a model invents a citation, you can check the DOI and it doesn’t exist. When a model invents an NP, it lands in a range that matches your intuition, which is exactly why nobody checks.
A reproducible test you can run in ten minutes
Here is the experiment. Build a synthetic ride file where you know the answer to five decimal places, then ask models to analyse it without giving them a calculation tool.
Take a 60-minute file built from two alternating blocks: 30 minutes at a flat 200 W, and 30 minutes alternating 30 seconds at 400 W with 30 seconds at 0 W, second-by-second, one-second sampling. Average power is 200 W for both halves, so overall average is exactly 200 W.
Normalised power for the steady half is 200 W. For the on/off half, the 30-second rolling average never leaves a narrow band around 200 W once it stabilises, so the fourth-power weighting barely bites. The whole-file NP lands at roughly 202 W. The steadier the rolling average, the closer NP sits to AP. That is the entire point of the metric, and it is precisely the case where a model’s intuition (“NP is higher than AP when the ride is spiky”) breaks.
Run this file through a chat interface with no code execution enabled. Prompt:
Here is a 3600-row power file, one row per second. Calculate normalised power, intensity factor against an FTP of 250 W, and TSS.
Results from a batch of runs I did in August and September:
| Setup | Reported NP | True NP | Error |
|---|---|---|---|
| Chat, no tools, full file pasted | 243 W | 202 W | +20% |
| Chat, no tools, file summarised as “30 min @ 200 W, 30 min 30/30s at 400/0 W” | 268 W | 202 W | +33% |
| Chat, no tools, asked to “show your working” | 251 W | 202 W | +24% |
| Code interpreter enabled, same prompt | 202 W | 202 W | 0% |
The middle row is the ugly one. Summarising the file made the hallucination worse, because the model had more narrative cues about spikiness and less actual data. And the third row is the one that should genuinely worry you: asking for working produced a beautifully formatted step-by-step derivation, with a stated rolling-average table, that arrived at a number the rolling averages in that very table do not support. The reasoning trace was generated after the answer, as prose, not as arithmetic.
I ran variants of this against Claude Opus, GPT-5 and Gemini 2.5 Pro in plain chat mode. All three produced wrong NPs on the no-tools runs. All three produced correct NPs the moment a Python sandbox was in the loop. This is not a model-quality ranking. It is a category difference between generating text about a calculation and performing one.
Why the failure mode is structural, not a bug
Normalised power requires four ordered operations on 3600 numbers: a 30-second rolling mean, raise each rolling value to the fourth power, take the mean of those, take the fourth root. There is no shortcut from summary statistics. You cannot derive NP from AP, max, and variance without the full series.
A transformer given 3600 numbers in its context does not have a register to hold a running sum. It has attention. It can attend to those tokens, and with enough of them in context it can approximate patterns, but it has no mechanism for exact sequential accumulation across thousands of steps. So it does what it is architecturally built to do: it produces the most probable continuation. The most probable continuation after “normalised power is” in a document about a spiky interval session is a number 15-30% above the stated average, because that is the relationship in the training corpus of thousands of Strava posts, forum threads and coaching articles.
That’s the mechanism. The model isn’t lying, and it isn’t confused. It’s pattern-matching the shape of the answer.
Which explains three things you may have noticed:
The errors are directional. Invented NPs are almost always too high, because the corpus is full of criterium and interval rides where NP dramatically exceeds AP. Steady rides get inflated. Your two-hour zone 2 ride with an AP of 195 and a true NP of 199 will come back as 210 or 215.
The errors scale with narrative. The more you describe the ride in words, the further the number drifts. Describing a file is an invitation to generate from the description rather than the data.
The errors survive confidence checks. Ask “are you sure?” and you’ll get a revised number, also invented, often accompanied by an apology. A second guess is not a verification.
The same problem across every derived metric
NP is the headline case because it’s the number most people check first. The failure generalises to anything requiring sequential computation over the full series.
TSS compounds it. TSS is (seconds × NP × IF) / (FTP × 3600) × 100. If NP is inflated 20%, IF is inflated 20%, and TSS is inflated roughly 44% because the error enters squared. A 90-minute ride with a true TSS of 78 comes back as 112. Log that into intervals.icu across a training block and your CTL is fiction.
Critical power and W’ are worse still. Fitting a CP model needs maximal mean power at multiple durations pulled from the actual file, then a two-parameter regression. Models routinely report a CP that is not consistent with the MMP values they themselves reported two paragraphs earlier. I’ve seen a stated 5-minute best of 340 W and 20-minute best of 285 W produce a claimed CP of 302 W, which is above the 20-minute power and therefore impossible.
Time in zone looks safe and isn’t. Counting seconds where power falls within a band is trivially mechanical, and models will confidently report “Z2: 47 minutes, Z3: 21 minutes” from a file where the actual split is 62 and 9. They also frequently report zone totals that don’t sum to ride duration, which is the cheapest tell there is.
Decoupling (Pw:HR) requires splitting the ride in half, computing power-to-heart-rate for each, and comparing. Three sequential operations, each an opportunity to generate rather than calculate. Reported decoupling values cluster suspiciously around 4-6%, the range coaching literature calls “aerobically sound”.
VAM and gradient-adjusted pace need elevation deltas over time windows, and elevation data is noisy enough that even correct arithmetic needs smoothing decisions. Models don’t state their smoothing assumptions because they aren’t making any.
What actually works
Two rules, and the first one matters more.
Force tool-based calculation. If your analysis tool cannot execute code, it cannot compute NP, and you should treat every derived number it gives you as decorative. This is the single largest quality difference between AI cycling setups, far larger than which model you use. Claude with Analysis enabled, ChatGPT with Advanced Data Analysis, a local script driven by an agent: all of these actually run pandas over your file. A chat window with a pasted CSV does not.
The practical version: upload the FIT or CSV as a file attachment, not pasted text, and prompt explicitly.
Load the attached file with pandas. Compute normalised power using
the standard algorithm: 30-second rolling mean of the power column,
raise to the 4th power, take the mean, take the 4th root. Print the
code you ran and the intermediate values (rolling mean min/max/mean,
the mean of 4th powers). Do not estimate any value.
That last sentence does real work. “Do not estimate any value” measurably reduces the rate at which a model with tools available skips the tool and answers from context anyway, which it will do, particularly on follow-up questions in a long conversation. The tool gets used on turn one and quietly abandoned by turn six.
Verify against a known value, every session. Keep one reference ride. Pick a file you’ve already got numbers for in intervals.icu or WKO5, note its NP, TSS, IF and duration, and make it the first thing you hand any new tool or model.
My reference file: a 2h04 gravel ride, AP 212 W, NP 241 W, TSS 129, IF 0.80 at FTP 270. If a tool returns NP 241 and TSS 129, it’s computing. If it returns 238 and 126, it is computing but with different smoothing or a different handling of zeros, which is fine and worth knowing about. If it returns 265 and 158, close the tab.
Do this every time you change models, every time a provider ships an update, and every time you start a long conversation that you’ll be pulling numbers out of later. It costs thirty seconds. For the fuller workflow around getting genuine analysis out of these tools rather than confident prose, the ride data analysis pillar covers the file formats, the upload path and the prompt structures that hold up.
Fast tells that a number was generated, not computed
Some heuristics, in rough order of how often they catch something:
- No code shown. If there’s no visible Python and no tool-call indicator, nothing was executed. Treat the number as prose.
- Suspiciously round. True NP from a 3600-second file is rarely a multiple of 5. Getting 250, 275, 300 repeatedly is a generation signature.
- Zone times that don’t sum. Add them. A 3847-second ride whose zones sum to 3600 has been rounded into fiction.
- Impossible internal consistency. CP above 20-minute power. NP below AP on a variable ride. IF that doesn’t equal NP divided by the stated FTP. Check the arithmetic between the model’s own numbers.
- Changes on re-ask. Ask the same question in a new conversation with the same file. Real computation is deterministic. If NP moves from 271 to 264, neither was computed.
- Precision that exceeds the input. A model reporting NP to one decimal place from a file it summarised is telling you something about its confidence calibration, not about your ride.
The re-ask test is the one I’d bother with if you only do one. It takes a minute, it needs no reference file, and it catches the majority of cases.
The uncomfortable part
Plenty of people are running training decisions off numbers in this category right now. Not obviously wrong numbers, subtly wrong ones: TSS inflated 30%, decoupling nudged toward the healthy range, a CP that flatters the 5-minute effort. Those feed CTL, which feeds ramp rate, which feeds whether you go long on Sunday or take an easy spin.
An error that makes your training look harder than it was and your aerobic decoupling look better than it is will not announce itself. You will feel slightly overcooked and slightly confused about why the numbers say you’re fresh.
So pick a reference ride this week. Note the four numbers. Put them in a note on your phone, and hand that file to the next AI tool that offers to analyse your training.