Your Watch Is Guessing: In Defence of Not Measuring Everything

I had a bad night’s sleep in March, and I know it was bad because my watch told me so at 6:40 the next morning, with a score, in red.

The thing is, I’d woken up feeling perfectly fine. I’d have described it as a decent night. But then I saw the number, and within about ninety seconds I had revised my own experience — actually, now I thought about it, I was a bit foggy, wasn’t I, and that would explain why the morning felt slow, and I should probably take it easy today given the circumstances.

None of that was true before I looked at the screen. I had felt fine. A device strapped to my wrist, using an accelerometer and an optical heart rate sensor and a proprietary algorithm nobody outside the company has audited, had made an estimate, presented it as a fact, and I had overwritten my own experience with it.

That was the day I started wondering whether the whole thing was helping.

What these devices actually do, and don’t

It’s worth being precise about this, because the marketing is deliberately vague and the gap between what’s measured and what’s displayed is enormous.

What a modern wrist device genuinely measures is quite short: movement, via an accelerometer, and blood flow at the wrist, via light shone into your skin. Some add skin temperature and electrical conductance. That’s essentially the raw material.

Everything else — sleep stages, calorie burn, recovery score, “readiness,” stress level — is inferred from those signals by an algorithm, and the quality of those inferences varies wildly.

Heart rate during steady activity is generally decent. Step counts are broadly reliable, if a little generous with arm movement. That’s the strong end.

Energy expenditure is much weaker. Estimating how many calories someone burned requires knowing things a wrist sensor cannot know — actual body composition, metabolic efficiency, the specific mechanics of what you were doing — and validation studies have repeatedly found substantial errors in both directions. The number in the app is presented with the confidence of a measurement and is closer to an educated guess.

Sleep staging is weaker still. The clinical standard for determining sleep stages involves electrodes measuring brain activity. Your watch is inferring brain state from wrist movement and pulse. It’s usually reasonable at telling asleep from awake. It’s much shakier at telling you how much deep sleep you got, which is precisely the number people fixate on.

None of this makes the devices useless. It makes them instruments for spotting trends, which is a completely different job from the one most people use them for.

The ten thousand steps thing

Since we’re here: the ten thousand step target has no clinical origin whatsoever.

The most widely repeated account traces it to a Japanese pedometer marketed in the mid-1960s, whose name translated roughly as “ten thousand steps meter.” The number was chosen partly because it’s memorable and partly, apparently, because the character for 10,000 resembles a walking figure. It was a brand name.

Later research has been kinder to it than it deserved — more walking really is better, and the benefits are substantial — but the curve appears to flatten well before ten thousand for most health outcomes, with a lot of the benefit banked at considerably lower counts, particularly in older adults.

I mention this not because the target is bad. It’s a fine target. I mention it because millions of people have felt a small daily failure for missing a number that came from a 1960s marketing department, and that’s a decent illustration of how quantification acquires authority it hasn’t earned.

When the measurement becomes the problem

Sleep researchers have a term — orthosomnia — for people whose sleep gets worse because they’re anxious about their sleep tracking data. It emerged from clinicians noticing patients arriving with printouts, convinced they had a disorder, whose main problem appeared to be the tracking itself. Lying awake worrying about your sleep score is an outstanding way to sleep badly.

A milder version of this affects a lot of people across all these metrics, and it runs in two directions.

In one direction, the number overrides the body. You feel fine, the recovery score is low, so you skip a session you’d have enjoyed and benefited from. Or the reverse, which is worse: you feel wrecked, the app says you’re primed, so you train hard on a day your body was quietly asking for rest — and your body, unlike the algorithm, has access to every input.

In the other direction, the number becomes the goal. You start doing things to satisfy the device rather than because they’re useful: pacing the kitchen at 11pm to close a ring, picking a session type because it scores better, avoiding activities that don’t register properly. Measurement has a well-documented tendency to distort the thing being measured, and it applies here with full force.

There’s a subtler cost too. Interoception — reading your own internal state — is a skill, and skills atrophy when you outsource them. Spend two years asking a device whether you’re tired and you get slightly worse at knowing.

What they’re genuinely good for

I still wear one. The anti-tracking position gets as overstated as the pro one, so let me be fair.

Making the invisible visible. Almost everyone underestimates how sedentary they are. An honest account of a Tuesday — three thousand steps, mostly to the kettle — is clarifying in a way no amount of self-reflection achieves. That first shock is worth the price of the device on its own.

Trends over time. The absolute numbers are unreliable but the direction usually isn’t. If your resting heart rate has drifted up over three weeks, that means something even if each reading is imprecise, because the errors are fairly consistent. Comparing yourself to yourself is where these things shine.

Catching illness early. Resting heart rate and temperature deviations often appear a day or two before you feel unwell. That’s a real, practical use.

Structure for people who like structure. For some personalities a target is genuinely motivating, and there’s decent evidence that step counters increase activity. If that’s you, this isn’t a small benefit and most of the criticism doesn’t apply.

The occasional hard data point. Discovering that your “easy” runs are actually run at a hard effort, or that your heart rate takes a long time to settle afterwards, can be information you’d never have got by feel.

The test I’d apply

I’ve landed on a fairly simple question: is the number changing what I do in a way I’d endorse on reflection?

If a low step count makes me walk to the shop, good — that’s the tool working. If a low recovery score makes me anxious and cancel plans, that’s the tool working me, and it should go.

The follow-up test is whether you could stop wearing it for two weeks without discomfort. Not whether you’d want to — whether the idea of an unmeasured fortnight produces a small flicker of unease. If it does, that’s information about the relationship. A tool you can’t put down has stopped being a tool.

I ran that experiment last summer. Two weeks, watch in a drawer. For the first three days I reached for my wrist constantly and felt faintly like exercise didn’t count if it wasn’t recorded — a belief which, once said out loud, is obviously insane. By the end of the second week I was training entirely by feel, and I realised I’d been quietly ignoring how I felt for about two years.

I went back to wearing it. But I turned off nearly all the notifications, stopped looking at the sleep score in the morning, and now check the data roughly once a week instead of continuously. That’s the arrangement where it helps without intruding.

The wider point

There’s a broader instinct at work here that goes well beyond watches — the belief that if something matters it should be measured, and that what can’t be measured probably doesn’t count.