Entire Strength
Train hard. Think harder.

Recovery

How Wearables Estimate Sleep And Where They Go Wrong

Consumer sleep trackers infer sleep from movement and pulse rather than measuring brain activity, which sets a hard ceiling on how accurately they can report sleep stages.

Man in white tank top exercising outdoors on metal bars.
Man in white tank top exercising outdoors on metal bars. · Photo via Pexels
Health information notice. General information — not a substitute for professional advice. Read the full disclaimer.

Sleep trackers report a number of hours and a breakdown into stages, and both are inferences. Understanding what the device actually senses explains which parts of the output are worth attention.

What the sensors actually detect

A wrist or ring device typically contains an accelerometer for movement and an optical sensor that measures blood volume changes at the skin to derive pulse rate.

Some add skin temperature and blood oxygen estimates. None of them measure brain activity, which is what sleep staging is clinically defined by.

The device therefore produces stages by feeding movement and cardiac data into a model trained to predict what a laboratory would have scored.

Why sleep and wake are the easy part

Distinguishing sleep from wakefulness using movement alone works reasonably well, because people move considerably more awake than asleep.

The characteristic failure is lying still while awake, which the movement signal reads as sleep. This is exactly what people with insomnia do, so the devices tend to overestimate their sleep.

Total sleep time is nonetheless the most trustworthy figure such a device produces, and it is the figure with the clearest connection to how a person functions.

Why staging is the weak part

Deep and rapid-eye-movement sleep are defined by patterns in brain electrical activity that have no reliable external signature a wrist sensor can read.

The models exploit correlations, since heart rate variability and movement do differ between stages, but the relationship is loose and varies between individuals.

Comparisons against laboratory scoring generally find stage-level agreement well below what the confident presentation in an app implies.

The readiness score problem

Composite scores combine sleep estimates with resting heart rate and heart rate variability into a single figure describing how prepared someone is.

The inputs are noisy, the weighting is proprietary, and the underlying measures vary day to day for reasons unrelated to training, including position, room temperature and hydration.

Treating a low score as an instruction is therefore acting on a number whose construction is unpublished and whose components are individually unreliable.

Using the output for what it can support

Trends across weeks are more defensible than any single night, because averaging reduces the noise that dominates individual readings.

Bedtime and wake time are recorded rather than inferred, which makes schedule regularity one of the more genuinely useful things these devices track.

Suspected sleep disorders are diagnosed with clinical testing, not with consumer devices, and persistent poor sleep is a reason to see a physician regardless of what an app reports.

recoveryactive recoverymodalitiesevidence
Ruth Ostrowski
Physiotherapist, Entire Strength

Ruth is a musculoskeletal physiotherapist who works with lifters. She writes about pain without catastrophising it, which is rarer than it should be.

More from Ruth →

Also by Ruth Ostrowski

Strength Science

What we still do not know

A survey of the open questions in strength training, which is a more useful thing to read than another confident answer to a settled one.

Hiro Tanabe··3 min read