Wearable VO₂ Max Estimates

The number your watch calls "VO₂ max" or "Cardio Fitness" is a good trend tracker and a poor measurement. It is least accurate for exactly the people most likely to be watching it — the very fit and the very unfit — yet the direction it drifts over months is honest, and that is the part worth using.

Wearable VO₂ Max Estimates

A wearable cannot measure oxygen uptake. It infers it from the relationship between your heart rate and your work rate, then projects that line out to a predicted maximum. That inference is good enough to follow your own progress and bad enough that the absolute figure should never anchor a clinical decision or a training-zone calculation. This article covers how the estimate is built, how far it misses, what knocks it off, and how to get the cleanest reading your hardware allows.

How the watch guesses

Directly measuring VO₂ max — your maximal oxygen uptake — requires a mask, a maximal effort, and indirect calorimetry: analysing the air you breathe out, inside a cardiopulmonary exercise test.[1] A watch has none of that, so it models the relationship between heart rate and oxygen uptake from the sensors it does have: optical heart rate, speed from GPS (satellite positioning), a barometer for grade, and an accelerometer. During a qualifying effort it fits a submaximal heart-rate-to-pace slope and extrapolates it to your age-predicted maximum heart rate.[2]

The two dominant systems have different entry requirements:

  • Apple ("Cardio Fitness") updates from an outdoor walk, run, or hike on flat ground (grade under 5%), with heart rate raised to at least 30% of your heart-rate reserve (the gap between your resting and maximum heart rate). It uses your entered age, sex, and weight, and lets you flag heart-rate-lowering medication such as beta-blockers so the estimate isn't dragged down by a blunted heart rate.[3]
  • Garmin (Firstbeat) needs an outdoor run of at least ten minutes at a moderate intensity before it will revise the figure.[4]

Both of those vendor documents are old — Firstbeat's dates from 2005 with a 2012 revision, Apple's from 2021 — on a page whose own argument is that firmware quietly changes these algorithms.

Passive machine-learning estimates from resting heart rate and everyday movement are a further variation, correlating well with lab values at the population level but inheriting the same individual-level limits below. The large study of that approach used a research-grade chest sensor rather than a wrist device.[5]

How accurate is the number?

Evidence: Moderate — the pattern is consistent across multiple independent validation studies and two systematic reviews, but each study is small (mostly 12–44 participants), so the picture is reliable in shape and rough in precision.

One caveat runs through the device comparisons below: every one was done in young, fit volunteers — no validation in the table has a mean age above 32, and the newest excluded anyone taking medication that affects heart rate, which is the reader the beta-blocker note above is written for.

Across validation studies the pattern is consistent: a group average can look reasonable while agreement for any one person is not. Pooled across eight studies and 244 participants the average bias is close to zero — but the range within which an individual reading falls spans nearly 20 mL/kg/min either side.[6] Devices also fail strict statistical equivalence tests, though only one paper in this literature has run one.[7]

Device (method)Validation finding
Apple Watch Series 9 / Ultra 2 (vs maximal treadmill test)Underestimated by 6.1 mL/kg/min on average — a 12% shortfall, with individual errors spanning −6 to +18[8]
Apple Watch Series 10 (vs maximal treadmill test)Two hardware generations later the error is unchanged, and only 5 of 40 participants were placed in the correct fitness percentile[9]
Apple Watch Series 7 (vs cycle test)Overestimates the low-fit, heavily underestimates the high-fit — though the fitness strata hold only 3 and 5 people[10]
Garmin Fenix 6 (Firstbeat)Strong average agreement, but fails equivalence — individual error too large[11]
Garmin Forerunner 920XTSignificant underestimation, though within a 10% error limit[12]
Garmin Vivoactive 4s (vs cycle test, unfit cohort)Overestimated by about 14 mL/kg/min; its output tracked a formula built from age, height and weight more closely than it tracked the lab[13]
Polar V800Sex bias — accurate in women, not in men; random error 6.6–16.4 mL/kg/min in both[14]

One methodological caveat sits behind some of these numbers: a watch derives its estimate from running or walking, but several of the studies validate it against a cycle ergometer. The INTERLIVE expert consortium recommends matching the test to the activity the device was designed for, though it puts no number on the penalty for not doing so — and which way a cycle reference shifts the result is itself contested.[15]

Brand also matters more than a trend line can absorb. When two watches were worn by the same 24 people, one read about 3 mL/kg/min high and the other about 2 low — a 5 mL/kg/min gap between brands in identical subjects, which is why a trend line does not survive changing watches.[16]

It's most wrong at the extremes

Evidence: Moderate. The single most important pattern is a shrinkage toward the average: wearables overestimate VO₂ max in unfit people and underestimate it in fit people, with the error growing toward both ends of the spectrum. When the Apple Watch Series 7 was tested against a cycle test and participants were split by fitness, the poor-fitness group came out flattered (overestimated by about 4 mL/kg/min) while the excellent-fitness group came out badly short — underestimated by roughly 12 mL/kg/min, an error near 20%. Those strata held three and five people respectively, so read the direction rather than the size.[17]

The mid-fitness midlife adult may be where these estimates do best, which is the opposite of the impression an athlete-heavy literature gives. In 74 adults aged 35–64 with hypertension, prediabetes or metabolic syndrome, the estimate was essentially unbiased and its typical error about 10% — holding across sex, age and body size, and degrading only in type 2 diabetes. That was a chest-worn sensor rather than a wrist optical reading, and no wrist validation in adults over 65 exists at all.[18]

The cause is structural. The algorithms are anchored to demographic priors — most importantly the age-predicted maximum heart rate — so they pull outliers back toward the population average. A highly fit runner has a strikingly low heart rate at a given pace, but the model can't see stroke volume or oxygen extraction directly, so it attributes part of that efficiency to "typical" demographics and clips the estimate. The same priors inflate a sedentary user's number. The same compression shows up in Garmin's Firstbeat engine, where error stays small in moderately trained users and widens in the highly trained.[19] The reason offered here is partly the page's own inference: the demographic anchoring is documented by the vendors, but attributing a fit runner's efficiency specifically to "typical" demographics is a plausible reading rather than a measured one.

What throws the number off

Evidence: Moderate (mechanistic — each confounder has a well-established effect on heart rate, though few have been quantified against watch VO₂ specifically). Because the model assumes a clean, fixed heart-rate-to-work relationship, anything that raises heart rate without raising oxygen demand reads as a fitness change that isn't real:

  • Heat. In hot or humid conditions the body shunts blood to the skin to cool down, dropping stroke volume; heart rate climbs to compensate (cardiovascular drift). The watch reads the higher heart rate as lost fitness and underestimates.
  • Illness, poor sleep, stress, caffeine. All raise submaximal heart rate through sympathetic activation, again pushing the estimate down on the day — Apple names heat and caffeine directly as causes of underestimation.[20]
  • Cold. Peripheral vasoconstriction shrinks the optical pulse at the wrist, degrading the heart-rate signal the whole model depends on — Apple documents exactly this failure mode.[21]
  • Skin tone. Melanin absorbs green light — wavelengths under 650 nm penetrate darker skin less well — and the INTERLIVE expert group grades skin tone as one of the few artefact sources with solid evidence behind it.[22] The measured effect is narrower than it sounds: pooled wrist pulse-rate bias does not differ by pigmentation, but the error spread is about twice as wide in dark skin,[23] and the gap appears at exercise intensity rather than at rest — around 12 beats per minute of error above 60% of heart-rate reserve, which is exactly the range a VO₂ estimate is computed in.[24]
  • Motion artifacts. A loose band or hard arm-swing can let the sensor lock onto stride cadence instead of pulse.[25]
  • Stale body weight. VO₂ max is reported per kilogram, so an out-of-date weight in the companion app produces a wrong relative number even when nothing physiological has changed.
  • The software itself. A firmware change can freeze or alter the metric outright, independent of anything happening in your body — and this feature, unlike the ECG and irregular-rhythm alerts on the same watches, has never been through regulatory clearance by either Apple or Garmin. The hardware is capable of cleared medical functions; this is not one of them.

The trend is the honest signal

Evidence: Weak–Moderate — the case rests largely on one large longitudinal cohort.

Here is the redeeming feature: if your wearing habits and routes are consistent, the bias is consistent too. A watch that reads 10% low this month should read about 10% low next month, so the month-to-month change carries more signal than the absolute value. The best test of this is the Fenland Study, which estimated fitness from a submaximal treadmill test in 11,059 adults and re-tested roughly 2,675 of them about seven years later: across the 2,042 with matched data, the change a wearable-based model predicted did track the change in measured fitness — but only moderately, a correlation of about 0.57 where 1.0 would be perfect.[26]

Month-to-month repeatability has been measured once, by Apple, and looks reasonable — typical variation of about 1.2 mL/kg/min between readings four weeks apart, though for one user in ten it is more than twice that, which is the size of the bias itself. No independent group has replicated it, and nothing equivalent exists for Garmin.

That moderate number also hides a specific catch: the models were better at detecting improvement than decline, tending to under-predict when fitness was actually deteriorating — reported as an observation rather than a tested contrast. So a rising trend is worth trusting, but a flat or only slightly falling line does not rule out real loss of fitness. Treat an upward drift as encouraging and a sustained downward drift as worth investigating — but don't read a reassuring-looking line as proof that nothing is slipping. For longevity purposes, where the direction of change matters more than a one-off number, that still makes the watch a useful, low-cost complement to a periodic lab test rather than a replacement for one.

The one thing not to do is let a wrong absolute number set your training intensities. Neither Apple nor Garmin derives heart-rate zones from the VO₂ max estimate — both use maximum heart rate, which the watch also estimates and can get wrong, and a wrong maximum shifts every zone boundary with it. Where the VO₂ figure does propagate is Garmin's race predictor, pace guidance and suggested workouts, all of which inherit its error: an underestimate sets your Zone 2 targets too low and you under-stimulate, while an overestimate in a beginner pushes "easy" runs into anaerobic territory. Calibrate zones from how you actually feel and breathe, not from the watch's headline figure.

Getting a cleaner reading

You can't make a wrist estimate lab-accurate, but you can cut most of the avoidable error:

  • Feed it clean heart rate. Pair a chest strap (such as a Polar H10) for runs; it bypasses the wrist's optical weaknesses and is the same upgrade that helps heart-rate variability readings. It will not fix the demographic anchoring, which is the larger problem.
  • Wear it right. Snug band, sensor sitting just above the wrist bone, so it doesn't slide during arm-swing.
  • Give it a good data point weekly. One steady, flat outdoor run produces the clean pace-to-heart-rate segment the algorithm wants; the systematic review suggests 10–15 minutes is enough.
  • Compare like with like. Read trends across similar seasons rather than against a single hot or cold week.
  • Keep your weight current in the app.
  • If you're highly trained and comparing against a treadmill test, mentally add 10–20% to the absolute figure. The size of the correction depends on which reference test you have in mind and which watch you wear — against a cycle test the shortfall has run closer to 28%, and one Garmin cohort was already within 7%, where adding anything would move you away from the truth.
  • Anchor it once. A single lab test gives you a fixed reference to calibrate the trend against — though a laboratory number is itself protocol- and modality-dependent, as this whole literature demonstrates — worth doing once if you train enough to care about the difference. See VO₂ max for what that test involves.

Bottom line

Treat the wearable VO₂ max as a fitness trend line, not a fitness measurement. Watch its direction over months, ignore the decimal places, distrust the absolute value most when you're very fit or very unfit, and verify with a lab test if a real number matters to you. For scale: in the newest validation, only 5 of 40 participants were placed in the correct fitness percentile by their watch.

Further reading

  • Lambe R et al. Investigating the accuracy of Apple Watch VO₂ max measurements: A validation study. PLoS One 2025.[27]
  • Lambe R et al. Accuracy of VO₂ max Estimates From Apple Watch Series 10. Mayo Clin Proc Digit Health 2026.[28]
  • Caserman P et al. Assessing the Accuracy of Smartwatch-Based Estimation of Maximum Oxygen Uptake Using the Apple Watch Series 7: Validation Study. JMIR Biomed Eng 2024.[29]
  • Molina-Garcia P et al. Validity of Estimating the Maximal Oxygen Consumption by Consumer Wearables: A Systematic Review with Meta-analysis and Expert Statement of the INTERLIVE Network. Sports Med 2022.[30]
  • Železnik Mežan I. Accuracy of wearables for determining the maximal oxygen uptake and lactate threshold: a qualitative systematic review. Front Sports Act Living 2025.[31]
  • Carrier B et al. Validation of Aerobic Capacity (VO₂max) and Pulse Oximetry in Wearable Technology (Garmin Fenix 6). Sensors (Basel) 2025.[32]
  • Passler S et al. Validity of Wrist-Worn Activity Trackers for Estimating VO₂max and Energy Expenditure (Garmin Forerunner 920XT, Polar V800). Int J Environ Res Public Health 2019.[33]
  • Snyder NC et al. Comparison of the Polar V800 and the Garmin Forerunner 230 to Predict VO₂max. J Strength Cond Res 2021.[34]
  • Spathis D et al. Longitudinal cardio-respiratory fitness prediction through wearables in free-living environments — the Fenland Study. NPJ Digit Med 2022.[35]
  • Rissanen AE et al. Cardiorespiratory Fitness Estimation Based on Heart Rate and Body Acceleration in Adults With Cardiovascular Risk Factors: Validation Study. JMIR Cardio 2022.[36]
  • Mühlen JM et al. Recommendations for determining the validity of consumer wearable heart rate devices: expert statement and checklist of the INTERLIVE network. Br J Sports Med 2021.[37]
  • Singh B et al. Impact of Skin Pigmentation on Pulse Oximetry Blood Oxygenation and Wearable Pulse Rate Accuracy: Systematic Review and Meta-Analysis. J Med Internet Res 2024.[38]
  • Hung C et al. Validity of heart rate measurements in wrist-based monitors across skin tones during exercise. PLoS One 2025.[39]
  • Jamieson A et al. Comparison Between Smartwatch-Derived and CPET-Measured VO₂max. Computing in Cardiology 2024 — conference proceedings, not peer-reviewed as a journal paper.[40]

— § —