The sleep-tracking conversation has changed character over the past two years, and the direction is worth knowing before you buy anything. The threads used to be about which device measured stages most accurately. They are now largely about whether stage measurement means anything at all.
What the consensus has settled on
Total sleep time, timing and disruption count are useful. These are the outputs that hold up. If your problem is a bedtime that has drifted an hour later since spring, or you genuinely do not know whether you are getting six hours or seven and a half, tracking answers that and the answer is actionable.
Sleep staging is not. The light/deep/REM breakdown that dominates every interface is inferred from movement and heart-rate patterns, against a ground truth that requires electrodes and a laboratory. Two devices on the same person on the same night routinely disagree substantially about stage distribution. The threads have converged on treating those percentages as decorative.
The practical consequence is the important part: do not optimise against a number that is not measuring what it says. People restructure their evenings chasing a deep-sleep percentage that is an estimate with wide error bars, and there is no evidence the effort improves anything.
Phone versus wearable
A phone alone gets you bed timing and rough disturbance from microphone and accelerometer data. That is enough to fix a drifting schedule, it costs nothing, and for a lot of people it is the whole benefit available.
A wearable adds heart rate, heart-rate variability and in some cases respiratory rate and blood oxygen. That is a real capability difference, and it is where the more interesting outputs — readiness, recovery, strain — come from. Whether those composite scores are more than well-presented heuristics is exactly what the threads argue about, and there is no settled answer.
The gap between the two is large enough that phone-only tracking should be understood as a different, narrower product rather than a cheaper version of the same one.
The thing nobody should misunderstand
No consumer device diagnoses sleep apnoea.
This matters more than everything above, because it is the condition most people are quietly hoping a tracker will catch. Some devices flag breathing irregularities or desaturation patterns, and the resulting referrals are a genuine public-health benefit. But a normal-looking report is not reassurance, and treating it as such delays diagnosis of a condition with real cardiovascular consequences.
Heavy snoring, waking unrefreshed, or an observed pause in breathing is a clinical conversation. No app changes that.
Orthosomnia, which the threads take seriously
The most common regret posted in these communities is not about a device’s accuracy. It is about the score becoming a stressor — checking it on waking, feeling worse about a night that had felt fine, and carrying that into the next night.
It is a documented phenomenon and it is the clearest downside of the category. The community advice is consistent and sensible: if the number changes your mood before your first coffee does, stop looking daily and check weekly trends instead. The trend was always the useful part anyway.
Where we land
If your issue is schedule consistency, a phone app is sufficient and free. If you want physiological data and will read it as a weekly trend rather than a nightly verdict, a dedicated wearable earns its price. If you suspect a sleep disorder, skip all of this and get assessed.
And whatever you use, ignore the stage percentages. That is the one thing these threads now agree on almost without dissent.