Skip to main content
Silent Outage

Start free

Telling your outage from your provider’s

How the verdict on an alert is reached, what evidence it rests on, and when the honest answer is that we cannot tell.

When something on your side goes wrong, Silent Outage looks back ±5 minutes across both of the sources it keeps about your providers — the provider's own status feed and Silent Outage's independent probe of it — and records one of three verdicts on the incident.

VerdictWhat it means
[PROVIDER] Likely providerAt least one provider the project watches was impaired inside the window.
[YOU] Likely youEvery provider the project watches was positively healthy inside the window.
[UNCLEAR] InconclusiveAnything else. This is a real answer, not a failure.

What is correlated against

The providers the project itself watches — the dependency checks it has created, and the provider behind any model canary it runs — and nothing else. Not the whole catalogue of roughly twenty: correlating an outage against a provider you have never told us you use is a coincidence presented as an answer.

A project that monitors no dependencies therefore always reads inconclusive. That is correct — we looked at nothing — and it is what onboarding's "dependency checks are proposed and confirmed" step exists to fix. If you want a webhook silence attributed to Stripe, the project needs a Stripe dependency check.

When it runs

Between the incident being recorded and the page going out, for exactly the events that page somebody. A breach that is still being held back to see whether it repeats pages nobody and is not correlated.

The window is centred on the instant the outage began, not on the moment we noticed it. An incident confirmed on the third breach is judged against the provider's series around the first.

The confidence label

high on a likely_provider requires two independent sources agreeing about one provider the anomalous check actually names: the provider's own feed published an outage and our own probe of it also found something wrong. That is the product's own thesis — vendor pages lag reality by 30–90+ minutes, which is why we probe as well as read.

high on a likely_you is a different rule, because that verdict is a claim about the whole watched set rather than about one provider: it needs every provider considered to carry at least three known samples inside the window. Below that the reading is still positive evidence of health — it is just not the strongest thing the product says — and the verdict is medium. The direct/monitored coincidence rule does not apply here: there is no single provider being blamed for it to be about.

medium is one source, or two sources on a provider the project merely watches (the coincidence rule: a provider the broken thing does not talk to cannot reach high).

low is everything weaker, and it is the only confidence an inconclusive ever carries.

The refusals

Each of these degrades toward inconclusive rather than toward a guess:

  • A feed we could not read is unknown, never down, and it blocks likely_you rather than merely failing to support it. likely_you needs positive evidence of provider health; "we did not look" and "we checked six providers and all six were green" are not the same claim.
  • A provider with no samples in the window — same.
  • Announced maintenance is neither the provider breaking nor a clean bill of health. It cannot page somebody about a provider outage and it cannot certify the provider green underneath a customer's real anomaly.
  • One failed probe is not an override. Two — the same floor that governs a probe contradicting a green status page anywhere else in the product.
  • A latency delta must be significant _and_ material — 2× and ≥250ms — and it is only ever decisive for a latency anomaly. A provider 800ms slower does not explain a customer's requests returning 500. It still blocks likely_you, though: "we checked and they were fine" is not a sentence to print about a provider answering four times slower than usual.
  • A delta with no baseline is `null`, never `x − 0`. A provider we have never measured is not infinitely slow.
  • An anomaly with no customer side cannot be `likely_you` however green the providers read. Which kinds of check have a customer side is decided once and for every kind, and a model-quality canary is the one that does not: it calls the provider with Silent Outage's own fixed prompt, on Silent Outage's own schedule, and touches nothing you wrote, so there is no "you" for the verdict to be about. The refusal is one-sided — likely_provider is untouched, because a canary failing while its provider publishes an outage is exactly the correlation this engine exists to make.

What is stored, and why

The verdict, the confidence, the instant it was computed and the width of the window it was computed over, plus one row of evidence per provider considered.

Two of the database's own constraints are the un-bypassable half:

  • a verdict that is not inconclusive cannot exist without a recorded run. A backfill, a hand-edited row or a future detector cannot assert "likely provider" without having correlated anything.
  • the three bookkeeping facts travel together or not at all.

The evidence is captured, not re-derived. A provider's status page is live: the incident that explained a 3am outage is resolved and off the page by 9am, and the raw samples behind it are swept on a retention window. So the provider's published headline, its short link and four timestamps are recorded beside every probe sample and kept with the incident.

What is not kept: the bodies of a provider's incident updates. They are parsed into memory for the length of a parse and nothing stores them — the list of what may be recorded is exhaustive, so "just the latest update, for context" is a build failure rather than a decision somebody makes quietly. A short link that is not https: is dropped when it is recorded and refused again by the database, because it reaches a link in an alert email.

Where the verdict appears

  • on the alert itself, carried by whatever detected the outage rather than worked out when the alert is sent;
  • on the incident, read by the dashboard and by the status-page configurator;
  • in the alert's evidence block — one line per provider that contributed, plus the window. Providers that were green are not printed under a likely_provider verdict: an alert listing twenty healthy vendors buries the one line that matters.

What it may never do

Attribution is evidence attached to a page. It can never be the reason a page does not arrive: every failure — a read that throws, a feed we could not read, an incident that has been deleted — is reported and stepped over, and the alert goes out with whatever verdict was already on it.

Telling your outage from your provider’s · Silent Outage