01
Missing data isn't proof
My previous tracker, a Whoop, nudged me when the strap had been off for too long. Fitbit doesn't. The tracker is good at recording and entirely passive about it: if it isn't on your wrist, the data simply has a hole in it, and you find the hole the next morning, when last night's sleep isn't there.
The obvious fix is a rule: no heart rate for thirty minutes, send an alert. It's wrong more often than you'd think. A Fitbit only uploads when it syncs through the phone, and between syncs Google's copy of your data is simply behind. So "no heart rate in the last thirty minutes" can mean the tracker is off, or that the phone is out of range, or that the app hasn't synced, or that the battery is flat. From the server's side, all four look the same.
What separates them is where you measure the gap from. Measured from now, a gap is an absence of evidence. Measured from the tracker's last sync, it is evidence: if the tracker synced at 14:36 and the newest heart rate it sent is from 14:00, it was demonstrably talking to the phone while recording nothing on skin.
Every decision about the wear reminder follows from that distinction. The rest of PulseWatch, the sleep and recovery rules, the personal baselines and the morning brief, sits on the same pipeline.
02
The brief I gave myself
Beyond the reminder, I wanted a small personal health-event platform: rules that are independent and configurable, baselines built from my own history rather than population thresholds, and delivery that survives the realities of free-tier infrastructure. It had five constraints.
- Unattended. OAuth that keeps working for months, and a system that says so when it doesn't.
- Private. Read-only access, the least data possible, no raw samples stored and no health values in logs.
- About £0 a month. Inside Cloudflare's free plan and Telegram's free Bot API.
- Honest. A personal tool, not a medical device. "Unusual" always means unusual relative to my own recent history.
- Demoable. Every feature can be shown with synthetic data, through the real pipeline, without exposing anyone's health information.
The last two shaped the code as much as the first three. Honesty became alert wording that follows the evidence, and demoability became an emulator of the Google Health API that the whole test suite runs against.
03
An explicit state machine
Each check classifies what it can see into one kind of evidence. The classifier is a pure function of the current time, the newest heart-rate timestamp, and the paired tracker's last sync and battery, so every case can be tested directly.
- The heart-rate fetch failedunavailableNothing: the state freezes
- Newest heart rate is under 30 minutes oldfreshEnds any episode
- No sync information for the trackersync_unknownSYNC_STALE
- Battery last reported empty (5% or less)battery_emptyBATTERY_EMPTY
- No sync for 60 minutes, battery last reported low (15% or less)battery_lowBATTERY_LOW
- No sync for 60 minutessync_staleSYNC_STALE
- Synced recently, and the gap before that sync is 30 minutes or moreoff_wristCONFIRMED_OFF_WRIST
- Synced recently, but only shortly after the last heart rateawaiting_syncNever confirms on its own
Two rows carry most of the weight. off_wrist needs a recent sync and a gap before it at least as long as the stale threshold. awaiting_sync is the case the naive rule gets wrong: the heart rate is old, but the tracker last synced only shortly after it, so nothing has been learned since. Battery is judged on the level as last reported, even when the sync is stale, because a flat tracker can't sync. An earlier draft required a recent sync, and would have reported every dead battery as a sync problem.
That evidence drives a state machine with seven states.
↺ recovered → normal on the next fresh reading · an outage freezes every state
The transition function is pure too, and the full state-by-evidence table is tested. A few rules keep it honest.
- Confirmation. An episode needs a configurable number of consecutive stale checks.
awaiting_synccounts as stale but can never confirm on its own. - Outages freeze it. A failed API call is not evidence, so it neither starts nor ends an episode.
- Check spacing. Checks closer than five minutes apart don't count separately, so a manual run or a duplicate cron delivery can't fast-forward confirmation.
- One alert per episode, sent on entering a confirmed state. The explanation can change mid-episode, say from off the wrist to not syncing, without a second alert.
- Recovery clears it. Fresh heart rate moves any episode to
RECOVEREDand deletes the reminder from the chat. There is no "welcome back" message: the reminder disappearing is the signal.
The wording is hedged in proportion to the evidence. Only a confirmed off-wrist episode says "you may have forgotten to put it back on". A stale sync says plainly that PulseWatch can't tell whether the tracker is being worn, and a flat battery is explained as a flat battery.
04
One Worker, every minute
PulseWatch is a single Cloudflare Worker with three entry points. scheduled runs the checks from a cron trigger, queue delivers notifications, and fetch serves a public /health endpoint alongside a token-protected /status and a few admin actions.
The cron fires every minute, but not every minute does the same work. Minutes 0, 10, 20 and so on run a full check: every rule, the daily sync, baselines and retention. The other minutes run a wear check, the off-wrist rule alone. While the tracker is worn, a wear check makes one API call, for the time of the newest heart-rate sample. The tracker's sync time and battery are fetched only once heart rate has gone quiet.
full every rule, the daily sync, baselines and retention.
wear the off-wrist rule alone. The bottom row is API calls while worn.
Each rule declares what it needs for the current check, and nothing else is fetched. At night the inactivity rule needs nothing, so its two roll-up calls are skipped, and daily data is loaded only when a daily rule is due. Every Google request also carries a field mask: the heart-rate query asks for the newest sample's time and never its value, and the device query never asks for the MAC address.
Crons fire in UTC, so everything local, from waking hours to the morning brief to where one day ends, is worked out in code. The morning the clocks change has its own test.
05
Rules that never touch the network
Every rule is a small object with its own configuration, a validated persisted state, a needs() that declares its data, and a pure evaluate(): no I/O, no clock, no randomness.
export interface HealthRule<C extends { enabled: boolean }, S> { readonly id: RuleId; readonly severity: Severity; /** Validates persisted state; a schema mismatch resets to `initialState`. */ readonly stateSchema: z.ZodType<S>; initialState(): S; selectConfig(rules: RulesConfig): C; /** Data this check must fetch for the rule. */ needs(input: NeedsInput<C, S>): readonly DataRequirement[]; /** Pure: no I/O, no clock reads, no randomness. */ evaluate(context: RuleContext<C, S>): RuleOutcome<S>;}export interface RuleOutcome<S> { state: S; status: RuleStatus; /** Short machine-readable detail for /status and logs. Never a health value. */ detail?: string; notifications?: NotificationIntent[]; /** * Notification ids whose subject has resolved. Undelivered ones are * cancelled; delivered ones are cleared from the phone where supported. */ resolved?: string[];}A rule never calls a notification provider. It returns intents, each with a deterministic id, and delivery is someone else's problem. The engine validates each rule's stored state, falling back to the initial state if the schema has drifted, isolates exceptions per rule, and writes only the state that changed. In steady state, with the tracker worn and nothing happening, a check rewrites no rule rows at all.
- deviceOffWristevery checkA confirmed off-wrist, sync or battery episodeonce per episode
- inactivitywaking hours90 min without a 5-minute window of 100+ steps, while wornonce per still period
- sleepdaily, after wakingLast night's main sleep was under the 6h 30m targetonce per night
- restingHeartRatedaily3 consecutive days above the personal rangeonce per episode, plus a cooldown
- hrvdaily3 consecutive days below the personal rangeonce per episode, plus a cooldown
- morningBrief07:30–11:00Last night's data is ready, or the 11:00 deadline passesonce per day
- serviceHealthevery checkAuthorisation lost, or 6 checks in a row without a Google synconce per incident
The two trend rules come from one factory, so a respiratory-rate or SpO₂ trend would be one more call. A global budget caps alerts at 24 in a rolling day. Anything beyond that is stored as suppressed and counted, so a buggy rule can't spam the phone.
06
Unusual for me, recently
The trend rules compare resting heart rate and heart-rate variability with a baseline built from my own last 30 days, never a population threshold. Wearable data is noisy, so the method is deliberately robust.
- Median and MAD, the median absolute deviation scaled by 1.4826, instead of mean and standard deviation. A night of alcohol, a cold or a sensor glitch barely moves a median.
- A log scale for HRV, which is right-skewed and changes multiplicatively, so a 30% drop means the same thing at any baseline level.
- A spread floor, so an unusually stable history can't turn a 1 bpm wobble into a dramatic score.
- Two kinds of significance. A day is unusual only if the robust z-score passes its threshold and the change is material: at least 3 bpm, or 15% for HRV.
- Persistence. Three consecutive days on the same side. A missing day breaks the run, because absence of data is never evidence.
- A guard window. The baseline ends before the days being judged, so a sustained change can't absorb itself into its own reference.
- Patience. Nothing is claimed until 14 of the 30 days have data.
Baselines are recomputed from stored daily aggregates only when new daily data arrives, a handful of times each morning, and never on the per-minute path. This is descriptive statistics about one person's history, not clinical anomaly detection, and the messages say so. Machine learning stayed out on purpose: a few months of one person's data don't justify a model, and every rule here can explain itself.
07
An outbox between the rules and the phone
A notification has to survive the Worker being interrupted, the provider being down and the queue delivering a message twice. PulseWatch uses a transactional outbox. Rule state and new notification rows are committed to D1 in one batch, so an alert exists if and only if the state change that produced it was saved. The row's id is the rule's deterministic key, written with ON CONFLICT DO NOTHING, so the same event can never create two rows.
The queue carries only ids, {v: 1, id}. The notification text stays in D1 and is erased as soon as the row reaches a terminal state. The consumer claims a row with a 60-second lease in one conditional UPDATE … RETURNING. A duplicate or concurrent delivery finds either the lease, and retries later, or a finished row, and acknowledges it. Completion updates are fenced by the lease token, so a consumer that has lost its lease can't overwrite the row.
Retries back off exponentially from 30 seconds to a 30-minute cap with ±20% jitter, honouring Retry-After, for at most five attempts. Every alert has a time-to-live, three hours for the wear reminder, so a late reminder is dropped rather than delivered stale. Rows that never reached the queue are dispatched again after 15 minutes. No path retries forever.
One gap is worth stating plainly: delivery is at-least-once. If the Worker died after Telegram accepted a message but before D1 recorded it, the retry would send it again, and Telegram has no idempotency key. The window is narrow, and it's documented rather than papered over with a claim of exactly-once.
08
Private by construction
PulseWatch reads personal health data, so the privacy controls are structural rather than promised.
- Four read-only scopes: health metrics, activity, sleep and device settings. No write access, and no location, ECG, profile or nutrition.
- Field masks mean heart-rate values and device MAC addresses are never downloaded in the first place.
- Only daily aggregates are stored, one number per metric per day, for 120 days. Intraday samples are used in memory for the current check and discarded. Notification rows are kept for 30 days and run records for 14, and notification text is erased once delivered.
- No credentials in the database. The refresh token is a Worker secret, and access tokens live only in memory.
- Sanitised observability. Logs carry event names, ids, states, counts and error codes. A test runs the whole pipeline and asserts that no durations, bpm values or message bodies reach them, and another asserts the same of
/status.
OAuth follows the same principle. Consent is a one-off local command: the authorisation-code flow with PKCE, through a loopback callback on my own machine. The refresh token goes straight from Google's token response into wrangler secret bulk over stdin. It is never printed, logged, passed as an argument or written to disk, it never crosses a public endpoint, and the Worker never needs a Cloudflare API credential to store its own secret.
If Google later rejects the token, a circuit breaker stops every API call, one notification asks me to re-authorise, and checks resume by themselves once a new secret is stored. The breaker recognises a new secret by a truncated SHA-256 fingerprint, so the credential itself is never kept for comparison.
09
Then production happened
Production changed three things in the first two days. All times are UK time.
Night one: the API disagreed with its documentation
The first live check authenticated and computed a wear state, but two daily sources failed. A small diagnostics endpoint, POST /admin/diagnostics, replays each request the pipeline makes against the live API and returns only structure: statuses, counts, sizes and relative timings. It found both problems. dailyRollUp rejects any pageSize with 400 INVALID_ARGUMENT, although the published schema lists the parameter. And sleep pages hold 11 to 14 sessions rather than 25, so a 60-day backfill needed five pages. The client had capped it at four and, correctly, refused to treat a truncated window as "no data".
Unit tests couldn't have caught either, because both were differences between the documentation and the service. After the fixes the backfill completed, and the emulator now reproduces both behaviours, so a regression fails the build.
Day one: nine alerts, none delivered
On the first full day PulseWatch did its job: 144 checks, all successful, and nine alerts created, a short-sleep alert, the morning brief and seven inactivity nudges. None of them reached my phone. Delivery then went through public ntfy.sh, and 30 of 31 attempts came back 429 with ntfy's code 42908, "daily message quota reached". The other timed out.
- checks, all OK
- 144
- alerts created
- 9
- delivered
- 0
429 · ntfy code 42908, “daily message quota reached” (30) · t/o · timed out (1)
The quota on ntfy.sh's free tier is per IP address, and Cloudflare Workers send from addresses shared with many other customers, so someone else's traffic had used it up. The first design did expect rate limiting: 42908 was classified separately, counted, and given a 15-minute retry floor in the hope of leaving from a different shared address. That wasn't enough. The quota resets only at midnight UTC, so retries at any spacing hit the same wall. The one test that got through went out at 01:09, just after the reset.
So I moved delivery to a Telegram bot, whose Bot API has per-chat rate limits but no per-IP daily quota. Providers sit behind one NotificationProvider interface, so the switch was one new class and a setting, and ntfy remains an option for a self-hosted server. A setup script reads the bot token at a hidden prompt, learns the chat id from my own /start message, and stores both as Worker secrets. The next test arrived on the first attempt, about ten seconds end to end.
The switch has a privacy cost, and it's documented rather than hidden. Telegram bot chats aren't end-to-end encrypted, so Telegram stores the message text until it's deleted. Resolved wear reminders are deleted automatically.
Night two: making it fast
Two earlier removals hadn't produced an alert, so I measured how densely the Fitbit Air records. While it's worn, it logs about 400 heart-rate readings every 15 minutes, one every two to three seconds. Both removals had been short, about 15 and 17 minutes, under the 30-minute default, so PulseWatch had been right to stay quiet. The default was just more cautious than I wanted.
A false reminder costs a glance at my phone; a missed one can cost a night of sleep data. So I tuned for speed: a 5-minute threshold confirmed by a single check, reminders overnight as well, and wear checks every minute instead of every ten. Measuring the gap before the sync is what makes a 5-minute threshold usable without constant false alarms. A wear check costs 6 to 8 ms of CPU, inside the free plan's 10 ms limit, and one API call while the tracker is worn.
- Wear-check cadenceevery 10 minevery minute
- Stale after30 min5 min
- Confirming checks21
- Cooldown between reminders60 min15 min
- Overnightheld until 07:00sent
one wear check while worn · 1 API call · 6–8 ms CPU against the free plan's 10 ms
The faster cadence exposed a bug in one of my own guards. The five-minute spacing rule, there to stop duplicate runs fast-forwarding confirmation, was also delaying confirmation when a sync proved a removal a minute after an inconclusive check. An uncounted check now still applies new evidence once the count is already enough.
The first real alert came that night. The last reading was at about 01:46, the tracker synced at about 01:52, and the reminder arrived at 02:00, still on the old ten-minute schedule. With checks every minute, the reminder now follows the proving sync within about a minute.
last reading ~01:46 · tracker synced ~01:52 · reminder 02:00
10
Demo mode swaps the network, not the code
Health data is personal, so nothing public can come from a real wearer: not the tests, not the screenshots, not this page. Demo mode handles that by replacing the network rather than the code. A deterministic synthetic wearer sits behind an HTTP emulator of the Google Health API that returns the real wire format, with int64 values as strings, protobuf durations and extra fields the client doesn't ask for. It paginates, validates filters, and reveals data only up to the tracker's last sync. A matching ntfy emulator can replay the 42908 rejection. Everything downstream, from the OAuth refresh to the queue consumer, is production code.
$ npm run demo -- wearable-removed-recovered━━ Wearable removed, then recovered ━━━━━━━━━━━━━━━━━━━━━━━━━Off for a shower at 14:00, back on at 15:05. The reminder is cleared on recovery. 14:20 wear NORMAL 14:30 wear POSSIBLY_OFF_WRIST 14:40 wear CONFIRMED_OFF_WRIST 📲 ⌚ Fitbit reminder [high] No heart-rate data has been recorded for 41 minutes, even though your Fitbit Air synced 7 minutes ago. You may have forgotten to put it back on. 14:50 wear CONFIRMED_OFF_WRIST 15:00 wear CONFIRMED_OFF_WRIST 15:10 wear CONFIRMED_OFF_WRIST 15:20 wear RECOVERED 🧹 notification cleared from phone (deviceOffWrist-ep-1773153000000) 15:30 wear NORMAL
Twelve scenarios cover a normal day, a removal, low and empty batteries, stale sync, short sleep, both trends, inactivity, a notification retry and a duplicate queue delivery. They run from the command line, inside wrangler dev in real time, and in the test suite.
The suite has 212 tests in 17 files, and they run inside workerd, Cloudflare's own runtime, against a real local D1 built from the production migrations. They cover every state-by-evidence transition, the statistics, every rule, the Google client's pagination, 401 refresh and field-mask fallback, OAuth and PKCE, leases under concurrency, duplicate cron runs and queue deliveries, the clocks-change morning, the HTTP security headers, and the absence of health values from logs. CI runs type-checking, strict ESLint, Prettier and the tests, then builds the Worker bundle, runs two demo scenarios end to end and audits the dependencies.
11
What it runs on
The runtime dependency list has one entry, Zod, which validates every Google response at the boundary, the configuration and every queue message. Everything else is platform or development tooling.
FitbitAirrecords heart rate, sleep and steps; syncs through the phone
Google Health API v4four read-only scopes · OAuth 2.0 with PKCE · field masks
Workersone Worker: scheduled, queue and fetch
- Cron Triggersone trigger, * * * * *
- D1the only store: state, outbox, daily aggregates
- Queuesnotification ids only, one consumer
- Workers Logsstructured, sanitised, every invocation
Telegram Bot APIsendMessage and deleteMessage · the default
- ntfystill supported, for a self-hosted server
TypeScriptstrict · 48 files, ~7,200 lines
Zodthe only runtime dependency
Vitest212 tests inside workerd
GitHub ActionsCI, bundle build, demo smoke run, audit
For one person, daily usage sits well inside Cloudflare's free allowances, which is how the whole thing costs about £0 a month.
- Worker invocations / day~1,460 of 100,000
- D1 rows written / day~750 of 100,000
- D1 rows read / day~10,000 of 5,000,000
- D1 storage< 1 MB of 5 GB
- Queue operations / day< 100 of 10,000
- CPU per wear check(measured)6–8 ms of 10 ms
also: 1 of 5 cron triggers · ~1,500–3,000 Google API calls a day against 300 a minute · Telegram, no daily cap
12
What it can't do
- Latency depends on the sync. A removal is only visible once the tracker syncs through the phone. The reminder then comes at the next check that sees a long enough gap: about 40 minutes after removal with the defaults, and usually 10 to 20 with my settings.
- Naps in progress. Google processes sleep sessions after the fact, so a long daytime nap can draw an inactivity nudge.
- One person, one time zone, one tracker. Multiple users would key everything by the Google Health user id. It's designed for, not built.
- Telegram sees the text, as above.
- It is not medical. Every threshold describes personal history, not health outcomes.
Next on the list: more trends through the same factory (SpO₂, respiratory rate, skin temperature), a weekly report, charging detection for a "fully charged, put it back on" reminder, and other sources behind the normaliser boundary.
13
What stuck with me
Reading the data was the easy part. The work was deciding what missing data means. A gap measured from now is an absence; a gap measured from the last sync is evidence. Most of PulseWatch is that idea carried through: a state machine that won't confirm on weak evidence, outages that freeze it rather than move it, alert text that says exactly as much as it knows, and a system that tells me when it can't tell.
Missing data isn't an event. The design was deciding what it means, and being honest about when you can't tell yet.