Skip to main content
Silent Outage

Start free

Which of your destinations receive a page

The rules that decide who is contacted for an incident, why an unconfirmed destination is never one of them, and how escalation moves between them.

Silent Outage's alert path has three producers and one queue. The check that detects an outage computes that a check is Down, the escalation clock decides a later rung has come due, and the setup walkthrough sends the rehearsal alert you press a button for — and all three then have to answer the same question: given this page and the destinations this account has confirmed, which of them does it go to?

Each of them used to answer it separately. Three filters that agree today are three filters that stop agreeing the first time one is edited, and the way they stop agreeing is a page that silently never arrives. So the question is asked once, of one piece of code, and the answer names the purpose it was asked for.

The purposes

PurposeWho asksWhat it is
incidentthe check that saw the faultThe first page about an outage.
escalationthe escalation clockOne rung of the email → SMS → voice chain.
rehearsalthe setup walkthroughThe test alert a customer presses a button for.

The same account with the same destinations gets three different, correct answers.

The rules

1. Confirmed only

A destination is confirmed before it counts. An unconfirmed address is either a typo — whose outage goes to a stranger while the customer hears nothing — or an inbox somebody pointed us at without owning it.

Refused three times, independently:

  • the query that fetches the destinations selects only confirmed rows;
  • the routing decision drops a destination whose confirmed flag says false — the read returns the flag as well as filtering on it, so an edit to one cannot silently start paging strangers;
  • the step that unseals a destination to send to it throws rather than returning an address.

2. SMS and voice are reserved for the escalation chain

They are the later rungs of the chain, and the key that decides what "the same page" means is built from the incident, the channel kind and the state. A destination paged at the first rung would have its own escalation swallowed by that key fifteen minutes later. Excluding phone destinations from the first fan-out is not politeness about phone bills — it is what makes the escalation arrive at all.

incident and rehearsal both exclude them. escalation is how they are reached.

3. An escalation rung pages exactly its own kind

The chain is a ladder. A rung that fanned out to email again would be the previous rung repeating itself.

4. An unroutable page is recorded, never dropped — except a rehearsal

An incident with nothing confirmed to page still queues one row addressed to email with no destination on it. It fails to resolve, leaves a receipt and is abandoned: a recorded non-delivery.

Queue nothing instead, and "the customer configured no destination" and "the alert path was healthy" become the same reading of the same ≥99.9% figure. A gap the customer created must be visible in the same place as a gap we created.

Two purposes are exempt, for reasons that do not transfer:

  • `escalation` — SMS and voice are paid channels a customer may deliberately not have, so an account with email and no phone is a configuration and not a gap. The rung is claimed anyway so the chain keeps its schedule, and the timeline records the step as having paged nobody.
  • `rehearsal` — a test alert that could not be delivered measures nothing, and spending the alert path's availability figure on a rehearsal would make the figure mean less.

5. A rung pages exactly one escalation position

An account can hold more than one person, and with them a destination that belongs to a member rather than to the account, and a position saying where in the order that member is paged. The ladder therefore has two axes, and they compose kind-major:

email(0) → email(1) → sms(0) → sms(1) → voice(0) → voice(1)

so a co-founder is emailed before anybody is telephoned. For an account whose members all sit at position 0 — every account by default, and every one-person account for ever — the grid collapses to exactly the plain email → SMS → voice chain, which is what makes this an addition rather than a change.

Rule 2's argument applies to the second axis unchanged. The key for a member's destination carries that member, so a member at position 1 paged with the first alert would have their own rung swallowed one delay later — which is exactly the page they were added for. So the first fan-out selects position 0 only, and an escalation rung selects its own position only, for the same reason it selects its own kind only.

A destination that belongs to the account rather than to a person is therefore position 0. A rehearsal is exempt from the position filter: it is one page about an already-resolved incident that nothing escalates, so there is no later rung for it to collide with, and somebody who will be paged at rung 1 is exactly who should see the test alert.

This is a fixed list of people, not a schedule. There is no on-call rotation, no shift, no rota and no handover in Silent Outage, deliberately, and nothing in the code that stores one.

Whether an account's positions become separate rungs at all is what the plan decides. Paging one person after another is part of the Startup plan. On a plan that does not include it there is exactly one group however many positions the members hold — so the ladder is the plain email → SMS → voice chain, and no page ever reaches a second person one delay later.

Two doors hold it. The screen does not offer a rung above the first where the plan has none, and the step that stores a rung refuses one — so a request crafted by hand is refused as well as a press of the button. And both reads that place a destination in the order place it in the order the plan has.

An account that already holds somebody above the first rung — which is what moving to a smaller plan leaves behind — keeps it. That is the same ruling that governs everything else a downgrade touches: nothing is deleted and nothing is rewritten, so moving back up restores the order that was set without anybody typing it again. In the meantime that person is in the first group and is paged with the first alert rather than by nothing at all. The other reading — honour the position and let their rung simply never come — would make a plan change quietly stop paging somebody, which is the failure this product is named after. Losing the entitlement costs the ordering and never the page, exactly as running out of the included phone allowance costs the SMS rung and never the page.

6. A rung states the first page's detection facts, not its own

Every page carries when it went down and how long since the last good signal, and both travel with the queued page rather than being looked up when it is sent: the summary of a check's most recent heartbeat is empty for the six kinds of check that nothing pings.

A rung that looked those two facts up for itself when it came due would find nothing for the six kinds of check that nothing pings, and would render "Last good: never" and no gap under an incident whose first email had named the minute money last moved. Two pages about one incident disagreeing about when it was last good is the confusion escalation exists to end, and it lands on the rung that wakes somebody up.

So the rung inherits: it reads those two instants back off the incident's first queued page and copies them. It is a copy rather than a recomputation because the escalation is not a second judgement about the outage — it is the same page addressed to somebody else — and a re-derived instant could legitimately differ an hour later. Exactly those two facts are inherited; everything else on the page is either read from the incident or belongs to the rung itself. A read that fails costs the rung the field and never the page.

What the routing decision is not allowed to know

It is handed channel kinds and row ids. It never sees an address, a webhook URL or the sealed blob one of them is stored in; the destination is decrypted once, in the part of the system that sends, at the moment of the send. Routing decides which destinations, never where they point.

It also does not deduplicate. The fan-out collapses two destinations that would produce the same key; doing it twice would hide a read returning the same row twice.

One envelope per distinct destination

The deduplication key decides what "the same page" means, and it has to tell two account-level destinations of one kind apart. A shared team alias and a personal address are two envelopes; a key that could not distinguish them would write one row, deliver to one address, and count that as a complete fan-out — so the ≥99.9% figure would read clean while an address the customer deliberately added received nothing.

A cohort is (channel kind, owning member). The first destination in a cohort keeps the bare key and every further one is namespaced by its own destination id:

incident:email:triggered                             the oldest confirmed account-level address
incident:email~<destination id>:triggered            every further account-level address
incident:email@<member id>:triggered                 that member's oldest confirmed address
incident:email@<member id>~<destination id>:triggered   every further one of theirs

Three reasons it is shaped like that:

  • The first destination keeps the key it has always had. Namespacing every destination would change every key at once, and a key that changes shape is a second row for an incident that has already been paged about — a duplicate page for every open incident at the moment of the deploy.
  • The discriminator is the destination's id, never its ordinal. An ordinal moves when an earlier destination is deleted or stops being confirmed, so an account would be re-paged about an open incident for editing an unrelated destination. The id is the row's own identity: adding a destination adds a key and changes none.
  • A destination with no row cannot be namespaced. It takes the bare key and collapses if that is taken. The only such destination is rule 4's unroutable placeholder, which is by construction the only one in its fan-out.

The order destinations are returned in — oldest first — is therefore also what decides which one holds the bare key: the oldest confirmed one, which is the one already holding it.

Ordering

Destinations are returned in the order they were added. One incident becomes one row per destination, so a broken Slack delays Slack and nothing else.

Running out of an allowance is never a reason to go quiet

Every plan includes a number of SMS and voice alerts a month — 0 / 30 / 100 / 250 from the free plan upwards. That is what the plan covers, not a ceiling on alerting: a monitoring product that stops paging because a counter ran out is the failure it is named for. Past the included allowance the decision is to charge or to require credits; what is never allowed is silence.

The refusal takes the same shape the daily email allowance does:

  • it is terminal, so the alert is abandoned now rather than after nine minutes of bounded retries against a limit that cannot move inside the window;
  • the abandonment is a recorded non-delivery with a receipt, counted against the ≥99.9% alert-path figure like every other;
  • it un-confirms nothing — the number is fine and the allowance is ours; and
  • the same incident is re-routed through the same rules with every phone destination dropped, not only the one that was refused, because the allowance belongs to the account.

The consequence is the sentence worth remembering: an exhausted phone allowance costs the SMS rung and never the page. Rule 2 above already keeps phone destinations out of the first fan-out, so the page went to email and to the webhook before any of this was asked.

Email, Slack, Discord and the outbound webhook are never usage-metered, and nothing in the product may start metering one of them per account. The email provider's own 300-a-day ceiling is not a counter-example: it is one counter for the whole deployment with no account on it, it is our operational limit rather than your plan's, and its own refusal takes exactly the path above.

Nothing counts a phone alert yet, deliberately, because nothing can send one yet: SMS and voice are not switched on, so the meter will land with the ability to send rather than being written against a channel that cannot. What is enforced today is how many phone destinations a plan may hold — 0 / 1 / 3 / 10 — checked at the one place a destination is added.

Which of your destinations receive a page · Silent Outage