# PulseWatch: Teaching a Wearable to Notice It Isn't Being Worn

Engineering case study of a personal project, by Muhammad Usman Mateen.
Canonical article: https://usmanmateen.com/research/pulsewatch
Published 2026-09-29.

> The full text of the article. The figures on the page are not reproduced here.

I kept taking my Fitbit Air off to shower or charge it and forgetting to put it back on, sometimes for hours. PulseWatch is the small system I built so that something would notice. One Cloudflare Worker reads the Google Health API every minute, works out whether missing heart-rate data really means the tracker is off, and sends a Telegram reminder. It also compares sleep and recovery against my own 30-day baselines, and it runs for about £0 a month. This is how it's designed, and what its first two days in production changed.

A personal project. PulseWatch is informational: it is not a medical device and is not intended for diagnosis, treatment or medical monitoring.

## Missing data isn't proof

My previous tracker, a Whoop, nudged me when the strap had been off for too long. Fitbit doesn't. The tracker is good at recording and entirely passive about it: if it isn't on your wrist, the data simply has a hole in it, and you find the hole the next morning, when last night's sleep isn't there.

The obvious fix is a rule: no heart rate for thirty minutes, send an alert. It's wrong more often than you'd think. A Fitbit only uploads when it syncs through the phone, and between syncs Google's copy of your data is simply behind. So "no heart rate in the last thirty minutes" can mean the tracker is off, or that the phone is out of range, or that the app hasn't synced, or that the battery is flat. From the server's side, all four look the same.

What separates them is where you measure the gap from. Measured from *now*, a gap is an absence of evidence. Measured from the tracker's *last sync*, it is evidence: if the tracker synced at 14:36 and the newest heart rate it sent is from 14:00, it was demonstrably talking to the phone while recording nothing on skin.

Every decision about the wear reminder follows from that distinction. The rest of PulseWatch, the sleep and recovery rules, the personal baselines and the morning brief, sits on the same pipeline.

## The brief I gave myself

Beyond the reminder, I wanted a small personal health-event platform: rules that are independent and configurable, baselines built from my own history rather than population thresholds, and delivery that survives the realities of free-tier infrastructure. It had five constraints.

- *Unattended.* OAuth that keeps working for months, and a system that says so when it doesn't.
- *Private.* Read-only access, the least data possible, no raw samples stored and no health values in logs.
- *About £0 a month.* Inside Cloudflare's free plan and Telegram's free Bot API.
- *Honest.* A personal tool, not a medical device. "Unusual" always means unusual relative to my own recent history.
- *Demoable.* Every feature can be shown with synthetic data, through the real pipeline, without exposing anyone's health information.

The last two shaped the code as much as the first three. Honesty became alert wording that follows the evidence, and demoability became an emulator of the Google Health API that the whole test suite runs against.

## An explicit state machine

Each check classifies what it can see into one kind of evidence. The classifier is a pure function of the current time, the newest heart-rate timestamp, and the paired tracker's last sync and battery, so every case can be tested directly.

Two rows carry most of the weight. `off_wrist` needs a recent sync *and* a gap before it at least as long as the stale threshold. `awaiting_sync` is the case the naive rule gets wrong: the heart rate is old, but the tracker last synced only shortly after it, so nothing has been learned since. Battery is judged on the level as last reported, even when the sync is stale, because a flat tracker can't sync. An earlier draft required a recent sync, and would have reported every dead battery as a sync problem.

That evidence drives a state machine with seven states.

The transition function is pure too, and the full state-by-evidence table is tested. A few rules keep it honest.

- *Confirmation.* An episode needs a configurable number of consecutive stale checks. `awaiting_sync` counts as stale but can never confirm on its own.
- *Outages freeze it.* A failed API call is not evidence, so it neither starts nor ends an episode.
- *Check spacing.* Checks closer than five minutes apart don't count separately, so a manual run or a duplicate cron delivery can't fast-forward confirmation.
- *One alert per episode*, sent on entering a confirmed state. The explanation can change mid-episode, say from off the wrist to not syncing, without a second alert.
- *Recovery clears it.* Fresh heart rate moves any episode to `RECOVERED` and deletes the reminder from the chat. There is no "welcome back" message: the reminder disappearing is the signal.

The wording is hedged in proportion to the evidence. Only a confirmed off-wrist episode says "you may have forgotten to put it back on". A stale sync says plainly that PulseWatch can't tell whether the tracker is being worn, and a flat battery is explained as a flat battery.

## One Worker, every minute

PulseWatch is a single Cloudflare Worker with three entry points. `scheduled` runs the checks from a cron trigger, `queue` delivers notifications, and `fetch` serves a public `/health` endpoint alongside a token-protected `/status` and a few admin actions.

The cron fires every minute, but not every minute does the same work. Minutes 0, 10, 20 and so on run a *full check*: every rule, the daily sync, baselines and retention. The other minutes run a *wear check*, the off-wrist rule alone. While the tracker is worn, a wear check makes one API call, for the time of the newest heart-rate sample. The tracker's sync time and battery are fetched only once heart rate has gone quiet.

Each rule declares what it needs for the current check, and nothing else is fetched. At night the inactivity rule needs nothing, so its two roll-up calls are skipped, and daily data is loaded only when a daily rule is due. Every Google request also carries a field mask: the heart-rate query asks for the newest sample's time and never its value, and the device query never asks for the MAC address.

Crons fire in UTC, so everything local, from waking hours to the morning brief to where one day ends, is worked out in code. The morning the clocks change has its own test.

## Rules that never touch the network

Every rule is a small object with its own configuration, a validated persisted state, a `needs()` that declares its data, and a pure `evaluate()`: no I/O, no clock, no randomness.

A rule never calls a notification provider. It returns *intents*, each with a deterministic id, and delivery is someone else's problem. The engine validates each rule's stored state, falling back to the initial state if the schema has drifted, isolates exceptions per rule, and writes only the state that changed. In steady state, with the tracker worn and nothing happening, a check rewrites no rule rows at all.

The two trend rules come from one factory, so a respiratory-rate or SpO₂ trend would be one more call. A global budget caps alerts at 24 in a rolling day. Anything beyond that is stored as suppressed and counted, so a buggy rule can't spam the phone.

## Unusual for me, recently

The trend rules compare resting heart rate and heart-rate variability with a baseline built from my own last 30 days, never a population threshold. Wearable data is noisy, so the method is deliberately robust.

- *Median and MAD*, the median absolute deviation scaled by 1.4826, instead of mean and standard deviation. A night of alcohol, a cold or a sensor glitch barely moves a median.
- *A log scale for HRV*, which is right-skewed and changes multiplicatively, so a 30% drop means the same thing at any baseline level.
- *A spread floor*, so an unusually stable history can't turn a 1 bpm wobble into a dramatic score.
- *Two kinds of significance.* A day is unusual only if the robust z-score passes its threshold *and* the change is material: at least 3 bpm, or 15% for HRV.
- *Persistence.* Three consecutive days on the same side. A missing day breaks the run, because absence of data is never evidence.
- *A guard window.* The baseline ends before the days being judged, so a sustained change can't absorb itself into its own reference.
- *Patience.* Nothing is claimed until 14 of the 30 days have data.

Baselines are recomputed from stored daily aggregates only when new daily data arrives, a handful of times each morning, and never on the per-minute path. This is descriptive statistics about one person's history, not clinical anomaly detection, and the messages say so. Machine learning stayed out on purpose: a few months of one person's data don't justify a model, and every rule here can explain itself.

## An outbox between the rules and the phone

A notification has to survive the Worker being interrupted, the provider being down and the queue delivering a message twice. PulseWatch uses a transactional outbox. Rule state and new notification rows are committed to D1 in one batch, so an alert exists if and only if the state change that produced it was saved. The row's id is the rule's deterministic key, written with `ON CONFLICT DO NOTHING`, so the same event can never create two rows.

The queue carries only ids, `{v: 1, id}`. The notification text stays in D1 and is erased as soon as the row reaches a terminal state. The consumer claims a row with a 60-second lease in one conditional `UPDATE … RETURNING`. A duplicate or concurrent delivery finds either the lease, and retries later, or a finished row, and acknowledges it. Completion updates are fenced by the lease token, so a consumer that has lost its lease can't overwrite the row.

Retries back off exponentially from 30 seconds to a 30-minute cap with ±20% jitter, honouring `Retry-After`, for at most five attempts. Every alert has a time-to-live, three hours for the wear reminder, so a late reminder is dropped rather than delivered stale. Rows that never reached the queue are dispatched again after 15 minutes. No path retries forever.

One gap is worth stating plainly: delivery is at-least-once. If the Worker died after Telegram accepted a message but before D1 recorded it, the retry would send it again, and Telegram has no idempotency key. The window is narrow, and it's documented rather than papered over with a claim of exactly-once.

## Private by construction

PulseWatch reads personal health data, so the privacy controls are structural rather than promised.

- *Four read-only scopes*: health metrics, activity, sleep and device settings. No write access, and no location, ECG, profile or nutrition.
- *Field masks* mean heart-rate values and device MAC addresses are never downloaded in the first place.
- *Only daily aggregates are stored*, one number per metric per day, for 120 days. Intraday samples are used in memory for the current check and discarded. Notification rows are kept for 30 days and run records for 14, and notification text is erased once delivered.
- *No credentials in the database.* The refresh token is a Worker secret, and access tokens live only in memory.
- *Sanitised observability.* Logs carry event names, ids, states, counts and error codes. A test runs the whole pipeline and asserts that no durations, bpm values or message bodies reach them, and another asserts the same of `/status`.

OAuth follows the same principle. Consent is a one-off local command: the authorisation-code flow with PKCE, through a loopback callback on my own machine. The refresh token goes straight from Google's token response into `wrangler secret bulk` over stdin. It is never printed, logged, passed as an argument or written to disk, it never crosses a public endpoint, and the Worker never needs a Cloudflare API credential to store its own secret.

If Google later rejects the token, a circuit breaker stops every API call, one notification asks me to re-authorise, and checks resume by themselves once a new secret is stored. The breaker recognises a new secret by a truncated SHA-256 fingerprint, so the credential itself is never kept for comparison.

## Then production happened

Production changed three things in the first two days. All times are UK time.

### Night one: the API disagreed with its documentation

The first live check authenticated and computed a wear state, but two daily sources failed. A small diagnostics endpoint, `POST /admin/diagnostics`, replays each request the pipeline makes against the live API and returns only structure: statuses, counts, sizes and relative timings. It found both problems. `dailyRollUp` rejects any `pageSize` with `400 INVALID_ARGUMENT`, although the published schema lists the parameter. And sleep pages hold 11 to 14 sessions rather than 25, so a 60-day backfill needed five pages. The client had capped it at four and, correctly, refused to treat a truncated window as "no data".

Unit tests couldn't have caught either, because both were differences between the documentation and the service. After the fixes the backfill completed, and the emulator now reproduces both behaviours, so a regression fails the build.

### Day one: nine alerts, none delivered

On the first full day PulseWatch did its job: 144 checks, all successful, and nine alerts created, a short-sleep alert, the morning brief and seven inactivity nudges. None of them reached my phone. Delivery then went through public ntfy.sh, and 30 of 31 attempts came back `429` with ntfy's code 42908, "daily message quota reached". The other timed out.

The quota on ntfy.sh's free tier is per IP address, and Cloudflare Workers send from addresses shared with many other customers, so someone else's traffic had used it up. The first design did expect rate limiting: 42908 was classified separately, counted, and given a 15-minute retry floor in the hope of leaving from a different shared address. That wasn't enough. The quota resets only at midnight UTC, so retries at any spacing hit the same wall. The one test that got through went out at 01:09, just after the reset.

So I moved delivery to a Telegram bot, whose Bot API has per-chat rate limits but no per-IP daily quota. Providers sit behind one `NotificationProvider` interface, so the switch was one new class and a setting, and ntfy remains an option for a self-hosted server. A setup script reads the bot token at a hidden prompt, learns the chat id from my own `/start` message, and stores both as Worker secrets. The next test arrived on the first attempt, about ten seconds end to end.

The switch has a privacy cost, and it's documented rather than hidden. Telegram bot chats aren't end-to-end encrypted, so Telegram stores the message text until it's deleted. Resolved wear reminders are deleted automatically.

### Night two: making it fast

Two earlier removals hadn't produced an alert, so I measured how densely the Fitbit Air records. While it's worn, it logs about 400 heart-rate readings every 15 minutes, one every two to three seconds. Both removals had been short, about 15 and 17 minutes, under the 30-minute default, so PulseWatch had been right to stay quiet. The default was just more cautious than I wanted.

A false reminder costs a glance at my phone; a missed one can cost a night of sleep data. So I tuned for speed: a 5-minute threshold confirmed by a single check, reminders overnight as well, and wear checks every minute instead of every ten. Measuring the gap before the sync is what makes a 5-minute threshold usable without constant false alarms. A wear check costs 6 to 8 ms of CPU, inside the free plan's 10 ms limit, and one API call while the tracker is worn.

The faster cadence exposed a bug in one of my own guards. The five-minute spacing rule, there to stop duplicate runs fast-forwarding confirmation, was also delaying confirmation when a sync proved a removal a minute after an inconclusive check. An uncounted check now still applies new evidence once the count is already enough.

The first real alert came that night. The last reading was at about 01:46, the tracker synced at about 01:52, and the reminder arrived at 02:00, still on the old ten-minute schedule. With checks every minute, the reminder now follows the proving sync within about a minute.

## Demo mode swaps the network, not the code

Health data is personal, so nothing public can come from a real wearer: not the tests, not the screenshots, not this page. Demo mode handles that by replacing the network rather than the code. A deterministic synthetic wearer sits behind an HTTP emulator of the Google Health API that returns the real wire format, with int64 values as strings, protobuf durations and extra fields the client doesn't ask for. It paginates, validates filters, and reveals data only up to the tracker's last sync. A matching ntfy emulator can replay the 42908 rejection. Everything downstream, from the OAuth refresh to the queue consumer, is production code.

Twelve scenarios cover a normal day, a removal, low and empty batteries, stale sync, short sleep, both trends, inactivity, a notification retry and a duplicate queue delivery. They run from the command line, inside `wrangler dev` in real time, and in the test suite.

The suite has 212 tests in 17 files, and they run inside workerd, Cloudflare's own runtime, against a real local D1 built from the production migrations. They cover every state-by-evidence transition, the statistics, every rule, the Google client's pagination, 401 refresh and field-mask fallback, OAuth and PKCE, leases under concurrency, duplicate cron runs and queue deliveries, the clocks-change morning, the HTTP security headers, and the absence of health values from logs. CI runs type-checking, strict ESLint, Prettier and the tests, then builds the Worker bundle, runs two demo scenarios end to end and audits the dependencies.

## What it runs on

The runtime dependency list has one entry, Zod, which validates every Google response at the boundary, the configuration and every queue message. Everything else is platform or development tooling.

For one person, daily usage sits well inside Cloudflare's free allowances, which is how the whole thing costs about £0 a month.

## What it can't do

- *Latency depends on the sync.* A removal is only visible once the tracker syncs through the phone. The reminder then comes at the next check that sees a long enough gap: about 40 minutes after removal with the defaults, and usually 10 to 20 with my settings.
- *Naps in progress.* Google processes sleep sessions after the fact, so a long daytime nap can draw an inactivity nudge.
- *One person, one time zone, one tracker.* Multiple users would key everything by the Google Health user id. It's designed for, not built.
- *Telegram sees the text*, as above.
- *It is not medical.* Every threshold describes personal history, not health outcomes.

Next on the list: more trends through the same factory (SpO₂, respiratory rate, skin temperature), a weekly report, charging detection for a "fully charged, put it back on" reminder, and other sources behind the normaliser boundary.

## What stuck with me

Reading the data was the easy part. The work was deciding what missing data means. A gap measured from now is an absence; a gap measured from the last sync is evidence. Most of PulseWatch is that idea carried through: a state machine that won't confirm on weak evidence, outages that freeze it rather than move it, alert text that says exactly as much as it knows, and a system that tells me when it can't tell.

Missing data isn't an event. The design was deciding what it means, and being honest about when you can't tell yet.

---

The timeline, state-machine and baseline animations are illustrative and use no real readings. The evidence table, rule table, code, delivery lifecycle, configuration and demo output come from the source and its synthetic demo. Production figures (checks, alerts, delivery attempts, timings and CPU) were observed between 26 and 28 September 2026. No real health values appear on this page.

Fitbit, Google Health, Cloudflare, Cloudflare Workers, Telegram, TypeScript, Zod, Vitest, GitHub Actions and the other names and logos shown are trademarks of their respective owners. They are used only to identify the products and services PulseWatch works with. This is an independent personal project and is not affiliated with, endorsed or sponsored by any of them.
