How Beaam decidesnot to alert
K-of-N detection, held signals, continuous history and the rule that Beaam's own blindness never pages — plus the four ways those rules were wrong in production.
10 min · conceptual · updated 11 October 2026
Beaam's first promise is "Quiet by default." Written out in full it reads "No flapping alerts. No noise from a known-flaky service." Detecting that something is wrong turns out to be the easy half of that. The hard half is deciding which of the things Beaam notices deserve a person, and being able to show the reasoning for the ones that don't.
This guide walks through the rules Beaam uses to stay quiet, and the four ways those rules have been wrong in production. Each of those failures shipped, ran against real accounts, and was found afterwards.
One reading is not an incident
Every detection rule in Beaam is a threshold plus two numbers: K of the last N readings must be bad before the rule fires. The check itself is a few lines of code. It takes the last N readings, counts the bad ones, and fires at K or more.
The default rules lean heavily on two shapes. Of the 220 defaults, 136 are 2-of-3 and 43 are 3-of-5. An HTTP check is the easiest one to picture:
http.up == 0is broken at 2-of-3. One timed-out request while a load balancer rotates is not an outage. Two of the last three is.http.response_time_ms >= 3000is degraded at 3-of-5. One slow response stays quiet. A slow minute or two does not.
That 3-second threshold used to be 10 seconds. It matched the fetch timeout exactly, so anything slow enough to trip it had already been aborted and recorded as down. The degraded band could never be reached. A threshold has to leave room for the state it describes.
Twenty-five rules are 1-of-1, and those are deliberate. They cover readings that cannot flap: a certificate with 14 days left, a database version near end of life, or a host scan that found malicious files. Waiting for a second reading of "Hostinger found 3 compromised files" buys nothing.
Recovery has the same shape. An open incident closes only when the window no longer holds K bad readings and the newest reading is good. For a 2-of-3 rule that is usually two good readings in a row, so a service that alternates good and bad does not open and close an incident every minute.
K-of-N is also recomputed from stored readings on every tick, rather than kept as a running counter. Evaluating the same window twice gives the same answer twice. That sounds like a pedantic property, and it isn't. Beaam's own watchdog once counted evaluations rather than observations. One failing sample, read on three consecutive judge cycles, crossed a three-failure threshold and paged on its own. Detection is structurally immune to that bug.
What it costs: time. Collection and detection both run every minute, so a 2-of-3 rule confirms a few minutes after the first bad reading, not seconds. That is the trade. A tool that fires on one reading is a tool that gets muted.
Held is not the same as quiet
When the newest reading is bad but the rule hasn't reached K-of-N, the signal is held. It isn't dropped. Beaam records it as a suppression with its reason attached — for example, "HTTP endpoint is not responding. — held: below 2-of-3." — and lists it under History → Near-misses: "Signals Beaam noticed but deliberately chose not to alert on."
That tab is the audit trail for "quiet by default". If Beaam held something back, you can see what it was and why.
A second kind of hold covers rates. A rate needs a denominator: one failed request out of one is a 100% error rate and not an outage. A rule can declare a companion volume metric, and below that volume it doesn't apply. The near-misses tab reads "held: fewer than N requests."
It used to work differently, and the old way failed quietly. Plugins withheld
the metric itself below a volume bar. A collector that has stopped sending what
detection watches is exactly what Beaam's silence guard looks for, so those
services read unknown. On 7 September 2026, every
Cloudflare Worker with fewer than three requests in
the window had been permanently unknown for that reason. It collected
perfectly, five metrics arrived, and no rule could see any of them. Moving the
denominator onto the rule kept the metric published and put "not enough traffic
to judge" where it belongs.
History has to be continuous
K-of-N assumes the last N readings exist. For a while, they didn't.
Detection used to read only the last three collection cycles of history, which kept about three readings. Any rule needing more could never fire: Hostinger's 10-of-10, MongoDB Atlas's 5-of-10, Supabase's 5-of-5. Nothing errored. Those rules just never had enough readings to fire.
The fix (4 October 2026) separates two questions that had been answered together:
- Is the latest reading fresh? That gate is unchanged. Pull data becomes
unknownafter three expected collection cycles, never less than three minutes, and a stale reading is never reused as current. - Which history counts? A reading counts while it is linked to the newest one by an unbroken chain, with no two consecutive readings more than three cycles apart. A missed cycle or two doesn't break the chain. A collection outage does, so a reading from before an outage can never combine with fresh ones to fire a rule.
Two more guards protect the read itself. PostgREST caps an RPC result at 1,000 rows and answers 200 with no error, so the bulk history read is chunked to fit under the cap and checked against an exact row count. A short answer throws, and detection falls back to reading one service at a time. Before that fix the services that reported least recently came back empty and evaluated as healthy. And a service whose history is missing its own latest check is skipped, not judged. An incomplete record that reads as quiet is how real incidents get closed as recovered.
A missed check is not a recovery
This is the newest rule, shipped on 8 October 2026, and it comes from the opposite direction. A false recovery creates noise just as surely as a false alarm, because the reopened incident pages again.
A Stripe sandbox connected to Beaam had charges disabled, and an incident had been open on it for 16 days. Then one collection failed. The failed check broke the continuous run and left a single reading, still "charges disabled". One reading can't meet 2-of-3, so the rule was held rather than triggered, and the service evaluated as quiet. The incident closed at 19:11 and sent its all-clear. At 19:13 the same fault opened a new incident and paged again.
It wasn't a one-off. In the 30 days before the fix:
- 32 of 72 closed incidents reopened on the same service within 10 minutes.
- Those reopenings sent 55 incident alerts and 20 false all-clear recaps.
- 29 of the 32 followed a failed check.
The fix is one condition in the alert gate. An open incident stays open while any rule's newest reading is still bad, even if there are too few readings to confirm it. In the code's own words, that is "not enough evidence", not "recovered". It closes once the newest reading is good. A test replays the production readings from that night.
So far: in the 2.5 days after the fix (8 October 20:45 UTC to 11 October 07:37 UTC), 1 incident closed and 0 reopened within 10 minutes. One closure is far too few to prove anything — the 30 days before had 32 reopenings out of 72 closures, and a sample of one can't be set against that. A real flapping service can still produce some reopenings, so the number will keep being measured rather than declared fixed.
Our failure is never their state
The last rule isn't about timing at all. It asks what a bad value can mean.
Between 10 and 20 September 2026 the same defect turned up four times, in four different plugins, and a fifth time on 25 September:
| What fired | What was true | Cost |
|---|---|---|
| Hostinger: "VPS telemetry is missing" | Beaam's own 6-second call budget was exceeded | 7 incidents, 14 emails in 48 hours |
| Neon: "no read-write compute ready to serve" | The branch was archived, which Neon unarchives on access | 12 incidents, one lasting 6.3 days |
| Supabase: "no recent completed backup" | A Free project, which has no automatic backups at all | 10 incidents, 22 emails |
| Cloudflare Pages: "the latest collection attempt failed" | The collector was healthy and the site is static | Every tick |
| Vercel and Resend supplements: degraded | Beaam's own Postgres read failed | Would have hit every project at once |
There are two mistakes in that table, and they're opposites. One reports Beaam's failure as the customer's state: a timeout, a provider 5xx, a failed read of our own database. All of those mean we could not look, not it is broken. The other reads a deliberate provider state as an outage: archived, scaled to zero, paused, static. That's a provider doing its job, usually to save the customer money.
The rule that came out of it is to ask, before giving a metric a rule, what a triggering value can mean. If it can mean either of those things, it must not page. In practice:
- Two facts, two metrics.
X_availableis published only on a tick the provider actually answered. A separateX_unreadablecarries Beaam's side. - Blindness gets its own rule.
hostinger.vps_metrics_unreadableis 10-of-10 and only degraded. An intermittent endpoint never produces ten failures in a row, and one that is genuinely gone produces nothing else. So it stays quiet through the noise and still speaks within ten minutes. - The distinction travels downstream. A rule about Beaam's own collection is marked as such. That mark is declared on the rule, not guessed from the metric's name. Splitting the metrics wasn't enough on its own. Beaam's correlation miner, which learns which services fail together, was still counting Beaam's blindness as the customer's episodes. 1,655 of those rows gave one VPS a support of 1,119 against 81 for the next service. No learned link could clear the bar, and the estate's most incident-prone service appeared in zero usable edges.
None of these raised an error. Each one looked like a monitoring tool doing its job.
What this adds up to
None of these rules is clever on its own. Together they mean that before Beaam interrupts anyone, it has checked four things:
- The evidence repeated: K of N, not one reading.
- The history it judged was continuous and complete.
- "Fewer readings" wasn't mistaken for "better readings".
- The bad value is about the customer's system, not about Beaam's view of it.
Whatever fails one of those checks is still recorded, and it's still visible in History. It just doesn't wake anyone.
After that, a 30-minute cooldown per service and cross-service correlation decide how many messages a real incident becomes. Correlation can only withhold a page on a strong link — the rest of that story is in how Beaam alerts.