Problem
When the platform was down on 2026-10-04 and 2026-10-05, nobody outside could tell. The site, the console and the API failed together, and the only record of what happened was written afterwards. On 2026-10-06 the phone app showed "The server answered with an error (503)" in every section during a roll, because it had nothing better to show. A person who depends on us has no single place to look, and no way to be told.
Three other RFCs already promise a status page and leave it undesigned. RFC 0058 puts "a public status page and subscribers per account" in its phase 3. RFC 0067 says the page "reads the same windows" as maintenance. The phone app (RFC 0074) needs the same answer for its outage card. Each would otherwise build its own.
Our customers have the same need for their own systems. They already run monitors and record incidents on the platform (RFC 0037, RFC 0058), and then pay a second vendor for a status page that knows neither. Some already have a status page with a vendor and will not move it, but would show what our monitors measure on it.
A status page has one requirement that the rest of the platform does not: it must work when everything else is down. A page served by the same cluster that is failing shows nothing at the moment it is needed.
Proposal
What a page is
A page belongs to an account and has:
- Components, in groups: the parts a reader cares about (for us: API, Console,
Sign-in, Website, Docs, Mobile app). A component names the monitors that measure it,
zero or more. Its state is
operational,degraded,partial_outage,full_outageormaintenance. - History: for each component, 90 days of daily availability, computed from its monitors' runs. 90 days is how long runs are kept today; a longer history needs daily rollups, which this RFC adds (one row per monitor per day), so the page never asks for raw runs.
- Incidents with their public updates, and maintenance windows, past, current and scheduled.
- Subscribers, who are told when any of these change.
A component's state comes from two sources, and the page says which one it is:
- Measured: its monitors. Down monitors make it
partial_outage(some) orfull_outage(all); a monitor over its latency limit makes itdegraded. - Declared: a person's public incident update or an active maintenance window sets it, and wins over the measurement while it lasts.
A page never shows a guess. A component with no monitors and no declaration shows
operational only with the words "not measured", so a reader can tell an assertion from
a measurement (the same rule as stubs everywhere in this tree).
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Nothing private reaches a page
An incident in RFC 0058 holds internal detail: responders, root cause, a timeline with names, evidence. None of it is fit for a public page, and a draft must never leak.
So a page never reads incidents. It reads public updates: short texts a person
writes for the page, each with a state (investigating, identified, monitoring,
resolved) and the components it affects. A public update is its own record, linked to
the incident it describes (the "customer-facing summary" of RFC 0058, phase 2). Only an
owner or admin publishes one, the same rule as publishing a postmortem. A published
postmortem can be linked from the page; its internal fields stay off it.
Measured state needs no person: a monitor that goes down turns its component red within one monitor interval, with the words "detected by monitoring". This is what the page says at 3 AM before anyone is awake, and it is true because it was measured.
A page is a snapshot, published away from the platform
The page is not served by the cluster. On every change the status service renders the page as static files (HTML for readers, a JSON summary for programs, an Atom feed) and publishes them to object storage behind our CDN, outside the platform. Readers only ever hit the CDN. When the platform is down, the last snapshot is still there.
A stale snapshot must not look healthy. Every snapshot carries the time it was made, and the status service publishes a fresh one at least every 60 seconds even when nothing changed. The page's own script compares that time with the reader's clock. Older than 5 minutes, the page puts a banner over everything: "Our status system has not reported since 03:12. Treat this as an outage until it does." The JSON summary carries the same time, so the app and widgets apply the same rule.
That rule detects a dead platform from the reader's side. Telling people actively needs a check that runs elsewhere: an outside probe of the public endpoints that opens an incident on our page, and notifies subscribers, when the platform cannot do it itself. The first version uses our incident tool's heartbeat (the platform pings it every minute; a missed ping pages on-call). Running our own probe in a second location follows, and the page shows which check reported last.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Where pages live
- Ours at
status.<domain>. - A customer's, on any plan, at
<slug>.<sites domain>, the separate domain RFC 0039 chose for published labs. A different registrable domain keeps customer content away from our cookies and sign-in. - A custom domain (
status.customer.com): a CNAME to the sites domain, verified the way RFC 0039 verifies domains (RFC 0030 ownership, rechecked daily), with certificates issued by the CDN for that hostname only after the check passes.
The snapshot sets no cookies and needs none, so the CDN caches every file. A page is small: the history is a fixed 90 cells per component.
Telling people
- Email, with double opt-in: a reader enters an address, gets a confirmation mail, and is subscribed only after the click. Every mail has a one-click unsubscribe. Mail goes through the platform's existing mail relay, with its caps. A subscriber picks components, or all.
- Atom feed and webhook (the signed webhooks of RFC 0022) for programs.
- The platform's own surfaces: the console, the site and the docs show a banner from our page's summary during an incident or window. The phone app's outage card says what the page says, links to it, and shows the incident's last public update instead of a generic sentence.
- Slack and Teams follow, through the connectors of RFC 0044.
A subscriber's address is personal data. It is stored only for the subscription, shown
masked to the account, deleted on unsubscribe, and covered by erasure (EraseSubject).
Our page's subscribers are ours; a customer's page's subscribers are that customer's,
and we process them for that customer (the DPA covers it).
Embedding and existing status pages
A customer's status should reach readers where they already are.
- JSON summary at a stable URL, cacheable, readable from any origin. This is what everything below is built on, and what a customer's own code can read.
- Badge: an SVG with the overall state, for a README or a footer.
- Embed: one script tag that renders a small status line or a component list in the host page, in the page's own light or dark theme. It reads the JSON summary, so it costs the customer's site nothing when we are slow.
- Mirror into an existing page: for customers who keep their status page with a vendor, a connector (RFC 0044) pushes component states and incident updates into it through the vendor's API. A person maps our components to theirs once. The first vendors are the ones customers ask for; each is written against its own API documentation. Where a vendor's page accepts custom HTML, the embed works there too.
Mirroring is one-way and says so in the vendor page's text ("measured by InOrbit"). We never read from the vendor page, so the customer's other data there stays theirs.
Our page, first
Our own page is the first customer of all of this. Components: API, Console, Sign-in, Website, Docs, Mobile app, each measured by the monitors the platform account already runs. Maintenance comes from RFC 0067's windows; incidents from public updates written on our incidents. Until the outside probe is ours, our incident tool keeps paging on-call and runs the response; the status page is ours from the first version.
Plans
Every plan gets a page, so an account can start for free and buy more as it grows. Each plan adds to the one below it; nothing is taken away on the way up. What a plan allows is a feature (RFC 0061); how much of it is a limit the status service reads per plan from its configuration, in the same way usage units are read.
| Plan | Adds | Pages | Components | Subscribers |
|---|---|---|---|---|
| Free | A public page on the sites domain with the InOrbit footer, badge, JSON summary, Atom feed | 1 | 5 | 100 |
| Personal | Email subscribers by component, the embed, webhooks, a custom domain | 1 | 20 | 1,000 |
| Team | Mirroring into a vendor's page, a private page for the account's own people, no footer | 3 | 50 | 10,000 |
| Enterprise | Pages per audience (one per customer or region, each with its own readers), sign-in with the customer's own identity provider, SLA and DORA exports of the page's history | agreed | agreed | agreed |
The limits are a proposal, not a measurement; they are set in configuration and change
without a release. The features are status.page (every plan), status.subscribers,
status.embed and status.custom_domain (Personal and up), status.mirror,
status.private and status.unbranded (Team and up), and status.audiences and
status.exports (Enterprise). Latency charts are an option on every plan: off by default, and turned on by the
account per component, next to the availability history. Going over a limit never
hides what is already published:
the page keeps showing, and the console asks for the next plan before it adds more.
Security and compliance
- The public path cannot reach private data. The status service keeps its own tables of published state and builds snapshots from them alone. It is told about changes (monitor health, public updates, windows) through events; it has no read access to incidents or postmortems. A bug in a page can therefore leak only what was already public.
- Readers never reach the cluster. Pages, summaries and feeds are static files on the CDN; the only public write is the subscribe form, which is rate limited per address and per page, and stores nothing until the confirmation click.
- Customer content is isolated by domain. Customer pages are on a separate domain; a component name or update text is escaped in every output, and the embed renders text only, never HTML from the page.
- Custom domains are issued only after the ownership check, so nobody can point a domain they do not own at a page, or a page at a domain they do not own.
- Audit: every publish, every component change and every subscription change is logged with who and what, never the subscriber's address (a hash, as the sign-in mail does).
- Availability claims: a status page is evidence for a customer's own SLA and for DORA incident reporting. Every measured figure on it names its source and window, and a declared state names the person's role who declared it. Nothing on the page claims a compliance status we do not have.
Plan
- This RFC.
- The status service: pages, components, public updates, daily rollups, the snapshot renderer. Our page on the shared host, served from the cluster first, with our components measured by our monitors.
- The snapshot published to object storage behind the CDN, the stale-snapshot rule, and the heartbeat to our incident tool. The console, site and phone app read our summary.
- Customer pages on every plan: the console editor (components, monitor mapping, public updates, maintenance), email and feed subscribers.
- Badge, embed and webhooks; custom domains.
- Mirroring into vendor pages; private pages; our own outside probe in a second location.
Alternatives considered
- A vendor's status page for us, nothing for customers. Quick for us, and it is what most companies do. It leaves our customers paying a second vendor that knows nothing of their monitors, and RFC 0058 already decided incidents belong in the platform. We keep a vendor for paging until our own outside probe exists, not for the page.
- Serve pages from the cluster. Simpler, and fresh on every request. It fails with the platform, which is the one moment a status page exists for. Rejected.
- Pages inside the incidents service. One service fewer, but the public path would live next to drafts and internal fields, and one bad query would publish them. A separate service that only holds published state is the stronger wall.
- Pages read live from monitors in the browser. No snapshot to keep fresh, but every reader becomes a request to the platform, and an outage turns into a blank page.
Open questions
- The first vendors to mirror into (owner, from customer requests).
- Latency charts start off on a customer's page, because latency can reveal more about a customer's system than they intend. Whether a page may show percentiles other than p95 is open.
Decision
Open. The owner decided on 2026-10-06 that the status page is our own product, that our page is the first one, and that paid and enterprise accounts get their own page or embed their status in the page they already have. On the plans, the owner decided the same day that Free gets a real page and that features move down a plan, so an account can buy more gradually; the table above follows that.
Publication
Public. Deployment details (the storage, the CDN set-up, the heartbeat target) stay out of this document.
Status log
- 2026-10-06: Opened. Owner decided the status page is our own product; paid and enterprise accounts get their own page or embed their status into an existing one.
- 2026-10-06: Plans revised with the owner: Free gets a public page, features moved down a plan, and limits grow per plan. Latency charts are an option on every plan, off by default.
- 2026-10-07: Checked: nothing is built; #457 and #458 are text. Open.