Why the monitor watching your production needsits own monitor
How Beaam is watched by something that isn't Beaam: a heartbeat, silence detection, an independent judge off Cloudflare — and the 8 October gap that showed it checking the wrong thing.
9 min · conceptual · updated 11 October 2026
Beaam's second promise reads: "A monitor that's quiet because it's broken is worse than no monitor."
That's worth taking literally. When a monitoring tool fails loudly, someone notices. When it fails quietly, the quiet looks exactly like good news. You stop checking the dashboard because the tool is checking it for you, and the tool has stopped. Nothing in your inbox says so, because the thing that would say so is the thing that broke.
So Beaam can't be the only thing that knows whether Beaam is working. This guide covers how that is arranged, and the 8 October 2026 gap that showed the arrangement checking the wrong thing. The architecture page draws the same layers as a diagram.
The layers
There are two audiences here, and it matters which layer speaks to which.
For you, the customer:
- A daily heartbeat. About once a day (never less than 20 hours apart), Beaam sends an "all quiet" message to your default channels. Its absence is the signal: no heartbeat, and something is wrong with Beaam.
- Silence detection per connection. Each connection's last successful collection is timestamped. If it goes stale, Beaam tells you that it has stopped being able to see that provider. The check is on staleness, not on a count of failures. A collector that stops being scheduled never records a failure, and in July 2026 that is exactly what a Cloudflare Smart Placement outage produced.
For Beaam's operator — the person who fixes Beaam:
- The judge. One always-on machine on Fly.io in Frankfurt, in a separate Fly organization. It shares no company, account, runtime or deploy credentials with the app, which runs on Cloudflare Workers with Supabase in AWS us-east-1. Every minute it probes the app and the marketing site, judges what it finds, and pages the operator if something is down.
- The app watches the judge. Every minute, the app's own operator-alerts job asks the judge whether it is still completing checks, and pages if it isn't.
- An external dead-man's switch. The judge pings a third-party cron-monitoring service after every completed run. If Beaam and the judge die together, that service is the only thing left that speaks. The judge refuses a ping URL on any of Beaam's own domains or platforms, because a switch that dies with the thing it watches has not been added.
The rest of this guide is about the judge, because that's where the interesting decisions are.
Why it's off Cloudflare
The first watchdog was a Cloudflare Worker on a separate Cloudflare account. The account separation was real, and it wasn't enough. The app and the marketing site are both on Cloudflare, so a Cloudflare-wide incident would take the app dark and the watcher with it, and no alert would ever be generated. That is precisely the scenario the watcher exists for. So the judge moved to Fly, and the Worker was deleted on 31 August once the new judge had proved itself.
Two smaller choices follow from the same reasoning:
- Frankfurt, not Virginia. The database is in AWS us-east-1. A judge in the same region would share a failure domain with the thing it reports on.
- One machine, never two. One decision-maker means one alert per incident with no leader election and no shared state. The rule is plain: never scale the judge above one.
It judges; it doesn't relay
Until 23 August the watchdog fetched the app's health endpoint and reported
ok: response.ok. That made it a messenger for the app's opinion of itself.
Any bug in the app's own filtering blinded it completely. The endpoint had
already had one: when AWS was withdrawn, withdrawn providers were exempted on
the user-facing board and missed in the health endpoint, so the two disagreed
about the same fact.
Now the app publishes facts, not a verdict. For each active connection it publishes the raw timestamp of its last collection attempt, and a declared reason wherever a collection is deliberately not expected. The judge subtracts against its own clock and applies its own threshold. A stale collector fails the probe even when the app answers 200, and the page names the connections and how stale they are.
It refuses to assume health in three places:
- An unreadable or missing health payload is a fault. A Cloudflare error page where JSON should be is not a clean bill.
- An unparseable timestamp counts as stale.
- An app build that doesn't publish the facts falls back to the app's own verdict, and the probe's detail says so. Nobody reads a green probe as full coverage when it isn't.
The app can ask for more patience per connection (a connection checked hourly shouldn't be judged on a 20-minute clock). It can never ask for less, and nothing past four hours is honoured. A system being judged doesn't get to blind its own judge.
Who watches the watcher, honestly
Watching the judge turned out to be harder than building it.
Stopping the judge's machine doesn't test anything. Fly restarted the judge 40 seconds after it was stopped, and the request that woke it was the app's own liveness check. Checking the judge revived it.
The first version of that liveness check also produced a false alarm. On 9 September at 23:49:25 UTC, one request to the judge took longer than 10 seconds, and the operator was paged: "Beaam's watchdog has gone quiet". The recovery mail a minute later gave the judge's last completed check as 23:49:29 — 4.4 seconds after it had been declared dead. The judge was slow to answer because it probed the whole app before replying. Now the app asks for stored state only, and an unreachable judge gets a two-minute grace window before anyone is paged. Every episode is still recorded, paged or not, so "how often does this happen?" has an answer.
8 October: the probe checked the wrong thing
At 19:13 UTC on 8 October, Cloudflare stopped invoking the app's cron trigger. There was no deploy, no configuration change and no Cloudflare status incident. The schedule still existed. It simply stopped firing.
That trigger runs everything that turns data into alerts: detection, delivery, the heartbeat. None of it ran.
Nothing paged. The judge's probe stayed green the whole time, for a reason that is obvious only in hindsight. Since 23 August, collection has run in its own Worker: one Durable Object per connection, so one hanging provider can't starve the rest. Collection kept running, so every collector stayed fresh. The probe judged freshness, freshness was fine, and the judge reported healthy. Collection and alerting had been split into separate Workers for good reasons, and that split meant "collection is fresh" no longer implied "alerting is running". The watchdog had never been told.
The gap was found by chance. The phases were run by hand once a minute from 19:52, and Cloudflare began firing again at 19:57. Detection evaluates current data rather than a queue, so the cost was latency, not lost evaluations. Why Cloudflare stopped is still unexplained.
What changed the same evening:
- The health endpoint now publishes the alerting jobs too. It answers 503
if detection or delivery hasn't started within five minutes, and it
publishes each one's raw
last_started_at. - The judge ages those timestamps itself, against its own clock, exactly as it does for collectors. The default window is five minutes. The app may lengthen it but never shorten it, and nothing past 30 minutes is accepted. A job that has never run, an unreadable record or an empty list all fail.
- A healthy probe now says what it checked:
HTTP 200 (judged N collectors; judged 2 crons). Live on the judge that evening: 31 collectors, 2 crons.
The lesson isn't really about cron triggers. A watchdog checks what is easy to check, and it's tempting to treat that as proof of the rest. A probe's healthy answer is only as good as its list of things it judged, so the probe now prints that list.
What this still doesn't cover
In the same spirit, here is what the design does not do today:
- The judge is one machine in one zone. Fly restarts it after a crash, but losing the host or the zone takes it out until a human intervenes, and there is no measured recovery time yet. The dead-man's switch reports the loss; it doesn't shorten it.
- Both watchers send email through the same vendor, from separate accounts. The judge also posts to a webhook, the app also sends push, and the dead-man's switch is a third vendor, so no single vendor silences everything. But email, the channel actually read at 3am, has one supplier for now.
- The judge believes the app's facts. It doesn't read the database directly. An app that is up but publishing wrong timestamps would be believed. Giving the judge its own database credential would trade independence for a duplicated high-privilege secret, so this is a deliberate open question rather than an oversight.
- The judge pages one operator, in one time zone. It protects you by making sure Beaam's operator finds out. Your own signal that Beaam is unwell is the heartbeat and the per-connection silence alerts.