Research

What a recovery score actually measures

on August 18, 2026
Article cover: What a recovery score actually measures

Abstract

Recovery scores on consumer devices condense several physiological signals into a single daily figure using formulas and weights the manufacturers do not publish. Nocturnal measurement of heart rate variability, the main one of those signals, reaches high agreement with ambulatory electrocardiography: in a 2025 study of thirteen healthy adults across 536 nights, the best of five devices obtained a Lin’s concordance coefficient of 0.99 with a mean absolute percentage error of 5.96%. Equivalence between techniques has a limit documented in 2026: across 103 college athletes and three seasons, variability obtained by photoplethysmography fell below the electrocardiographic measure (rMSSD of 80.9 versus 103.9 milliseconds) and captured between 16.7% and 56.7% of autonomic deterioration events, with a delay of 1.8 to 5.5 days. The composite score built on that measurement lacks an equivalent validation, and a meta-analysis of training guided by variability found small, non-significant effects on performance and on maximal oxygen uptake, with g of 0.079 and 0.171. Checking the figure daily also has effects described in clinical sleep medicine, recorded under the term orthosomnia. The accuracy of the sensor does not vouch for the algorithm that interprets its signal, and the evaluation of the composite score awaits the scrutiny the measurement layer has already received.

Introduction

Consumer devices that record physiological signals during sleep present a recovery score each morning: a single figure derived from a model that combines heart rate variability with resting heart rate, sleep duration and structure, breathing rate and the load of previous days. Daily decisions on training intensity and the organization of rest are based on that figure.

A validated sensor and a validated score are two different things. The published validation covers the measurement of the signal; the interpretation layer that turns it into a decision figure is built with formulas and weights that are not public. The distance between the two validations bounds the weight the score can carry.

Mechanism of heart rate variability

The interval between consecutive beats is not constant: it shows differences of milliseconds, and that irregularity carries physiological information. Heart rate variability measures those differences and reflects the balance with which the autonomic nervous system acts on the sinus node, with a sympathetic branch that activates and a parasympathetic one that favors rest.

The most widespread nocturnal metric is rMSSD, the root mean square of successive differences between beats. It is recorded during sleep because in that window posture, physical effort, caffeine and conversation contaminate the signal less.

The same person records different figures from one night to the next. The sleep phase in which the measurement window falls has an effect, along with the hour, posture, breathing rate, the previous evening’s alcohol, room temperature, an infection or the point in the menstrual cycle. An isolated absolute value therefore provides little information; the useful reference is each person’s baseline and its trend.

Pulse and electrocardiogram as distinct quantities

Most wrist and finger devices do not record the electrical activity of the heart: they record the passage of the pulse wave through a peripheral vessel using photoplethysmography. The resulting quantity is pulse rate variability, and the 2026 literature separates it from heart rate variability measured by electrocardiogram.

A study published in Sports Medicine Open in June 2026 analyzed data from 103 male Division I collegiate American football players collected across three seasons, with beat-to-beat comparison between the two techniques. Heart rate and pulse rate proved equivalent (59.4 versus 59.7 beats per minute, with a bias of 0.24 to 0.44). The variability measures did not: pulse rate variability fell below the electrocardiographic measure in both rMSSD (80.9 versus 103.9 milliseconds) and SDNN (141.3 versus 167.9).

The operational consequence lies in event detection. In the sliding-window deviation analysis, pulse rate variability captured between 16.7% and 56.7% of the autonomic deterioration episodes the electrocardiogram identified through rMSSD, and between 16.8% and 52.0% of those identified through SDNN, with an average delay of 1.8 to 5.5 days. The authors conclude that variability calculated by photoplethysmography should not be called heart rate variability, given the confusion it introduces among scientists and consumers.

Four of the signatories are employees of the company that develops a competing device, a circumstance the article itself declares and one worth holding in mind when reading the conclusion. The analysis was performed double-blind on anonymized data. The sample is male collegiate athletes, so the absolute figures do not transfer to a general population.

Validation of nocturnal measurement

A study published in Physiological Reports in 2025, by Dial and colleagues, compared five consumer devices with a reference ambulatory electrocardiogram, worn simultaneously by thirteen healthy adults across 536 nights.

For nocturnal rMSSD, the best of the five reached a Lin’s concordance coefficient of 0.99 with a mean absolute percentage error of 5.96%, close to the clinical standard. Two more devices came in at 0.97 and 0.94.

The same study shows the distance between devices. Within the same night and against the same electrocardiogram, concordance fell to 0.87 and 0.82 in the remaining two, with percentage errors of 10.52% and 16.32%. The authors themselves note that thirteen healthy participants do not establish how a device performs in clinical populations.

Reading from the wrist is more fragile than from a chest strap. A study published in Sensors in 2024, by O’Grady and colleagues, with 39 healthy adults and 316 measurements over fourteen days, compared two Apple Watch models against a Polar H10 strap: the watch underestimated variability by 8.31 milliseconds on average, with a mean absolute percentage error of 28.88%.

Sleep structure, the other main input to the score, is measured with an accuracy of its own. A study published in Sleep Medicine in 2026 compared three consumer devices with home polysomnography and actigraphy in healthy adults aged 18 to 45: Fitbit Charge 5 (n = 20), Apple Watch Series 7 (n = 19), and Polar Vantage M2 (n = 15). Fitbit overestimated light sleep by 57.7 ± 80.5 minutes (P = .009); Apple Watch underestimated it by 73.8 ± 61.6 minutes (P < .001) and overestimated REM by 30.3 ± 36.3 minutes (P = .003); Polar overestimated REM by 32.3 ± 43.6 minutes (P = .026). Their authors consider the devices appropriate for longitudinal self-monitoring rather than for clinical-grade assessment of sleep architecture.

Validation of the composite score

On top of the measurement, each brand builds a composite score that combines variability with resting heart rate, sleep duration and structure, breathing rate and the load of previous days, using a formula and weights that go unpublished. That layer has not undergone the scrutiny the sensor has received, and the accuracy of the input does not vouch for the algorithm that interprets it.

An empirical check on that jump exists. A systematic review with meta-analysis on training guided by heart rate variability found a medium, significant effect on submaximal physiological parameters, with a g of 0.296. On performance and on maximal oxygen uptake the effects were small and not significant, with g of 0.079 and 0.171. Tuning training to each morning’s figure has not consistently outperformed a well-designed predefined plan.

The difficulty sits in the stretch separating a correct measurement from an interpretation prepared to support decisions.

Costs of self-monitoring and allocation of the decision

The habit of checking the score has described effects. Baron and colleagues coined the term orthosomnia in the Journal of Clinical Sleep Medicine in 2017, built on “ortho”, correct, and “somnia”, sleep, by analogy with orthorexia. They described patients who came to clinic with a sleep disorder self-diagnosed from their tracker data, and a perfectionist pursuit of ideal sleep that had itself become the problem.

The phenomenon also has a sector-level reading. The Global Wellness Summit places neurowellness among the trends of 2026, and the backlash against permanent optimization appears as a current of its own within the sector.

A measurement can support two different uses, and the difference lies in the allocation of the decision. On the ERL scale, measurement exists to take a concrete decision against a threshold: transfer a solution or hold it back. No solution is transferred below ERL-3. Once that decision is taken, the score leaves the path of whoever uses the solution, who receives a resource whose scope is already settled.

A daily score inverts that allocation. It is handed to the person without the baseline, without the device’s margin of error and without the weight of each variable in the formula, and the decision moves to whoever has the fewest elements for taking it.

A design criterion follows from that allocation. A recovery experience can use the measurement to adapt its length or its intensity without turning it into a verdict the person reads on waking. The hub’s audio-first principle and its work without screens operate in that direction: an experience that is listened to does not require checking a figure.

Limitations of the available evidence

The validation study against electrocardiography included thirteen healthy adults: the authors themselves note that this design does not establish how the devices perform in clinical populations, and the results correspond to the specific models evaluated, with concordances ranging from 0.99 to 0.82 within the same night.

The comparison of wrist reading is limited to two models from a single manufacturer against one chest strap, with 39 healthy adults and fourteen days of follow-up.

The meta-analysis of training guided by variability found small, non-significant effects on performance and on maximal oxygen uptake; the available evidence does not show a consistent superiority of that daily adjustment over a predefined plan.

The comparison between pulse rate variability and heart rate variability was carried out on 103 male collegiate athletes: the sample confines the absolute figures to that population, and four of the signatories declare employment with the manufacturer of a competing device.

The comparison of the three devices with polysomnography covered 54 healthy adults aged 18 to 45 split across three groups, with one reference night per participant.

The formulas and weights of composite scores are not published, and that opacity prevents subjecting them to an independent evaluation comparable to the sensor’s.

Orthosomnia proceeds from the description of patients in clinic; the sources reviewed do not quantify its prevalence.

Conclusions

Nocturnal measurement of heart rate variability on consumer devices reaches, in the best models evaluated, agreement close to ambulatory electrocardiography. The composite score built on that measurement combines variables with unpublished weights and lacks an equivalent validation. The evidence on training guided by variability does not show a consistent superiority over a well-designed predefined plan, and checking the score daily has effects described in clinical sleep medicine. A recovery score is an estimate derived from a model, and it describes what that model estimates about the body on that particular morning; the validated part of the system is the measurement of the signal, and the layer that turns it into a decision figure awaits the same scrutiny.

Open research lines: Arena Program

Arena Program is yeshcube’s applied research line on human performance in elite sport, and it addresses cognitive and emotional recovery, regulation under pressure and mental preparation. It has begun a laboratory study on the use of Somia Bloom in high performance, in its validation phase during 2026, with no conclusive results and nothing transferred. Until it concludes, there is no basis for claiming the experience improves performance or speeds recovery.

Collaborating on the validation of recovery measures

yeshcube develops this line within Allies, its scientific collaboration system, with four partner types and three principles: value for value, traceability and independence. No partner can veto a publication.

Separating a correct measurement from a usable interpretation calls for designs no single team assembles alone. The work is of interest to exercise physiology and sleep science groups, to clubs and high-performance centers holding longitudinal records, and to manufacturers willing to submit their interpretation layer to the same scrutiny as their sensor.

Discover our Allies program →

Fill in our contact form →

References

Frequently asked questions

What does heart rate variability measure?

It measures the difference in milliseconds between consecutive beats. That variation reflects the balance of the autonomic nervous system over the sinus node, between the sympathetic branch that activates and the parasympathetic one that favors rest. It is recorded at night because posture, activity and conversation contaminate the signal less.

Are the devices that measure heart rate variability validated?

Nocturnal measurement reaches high agreement. A study published in Physiological Reports in 2025 compared five devices with ambulatory ECG across 536 nights in thirteen healthy adults, and the best obtained a concordance coefficient of 0.99 with a mean absolute percentage error of 5.96%.

Is the recovery score shown in the app validated?

The composite score each brand builds combines variables with weights that are not public, and it has not undergone the scrutiny the sensor has. The accuracy of the input does not vouch for the algorithm that interprets it.

What is orthosomnia?

It is excessive preoccupation with achieving perfect sleep based on tracker data. Baron and colleagues coined the term in the Journal of Clinical Sleep Medicine in 2017, by analogy with orthorexia, describing patients who came to clinic with a sleep disorder self-diagnosed from their device.

Subscribe to our updates

By subscribing, you will receive yeshcube news and content by email. You can unsubscribe at any time. See our privacy policy.

Follow us