Problem
On 2026-10-04 and 2026-10-05 the platform was down twice for storage, and once
partly down for a memory leak in the API gateway. The monitors saw all three: the
drops are in their charts. What happened is written up in docs/incidents/ (one
record per incident: the Markdown, an RFC 0046 evidence statement, a scenarios
file for tests). Those records are good, and they are in a repository: a
customer cannot read them, the monitor that saw the outage does not show them,
and nothing connects a postmortem's claims to the measurements that back them.
The usual answer is a third-party incident tool. It would hold our postmortems
and our customers' outside the platform, it would not know our monitors, runs or
rolls, and it would be one more vendor with personal data. And a postmortem is
the last step of the loop this platform is built around (docs/product-direction.md,
RFCs 0045 to 0047): understand, observe, change, verify, and keep the evidence.
An incident product that cannot point at measured evidence is a document store.
Proposal
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
The incident and its postmortem
An incident belongs to one account. Its header: title, severity (SEV1 to SEV4), status (investigating, identified, monitoring, resolved), when it started, was detected and was resolved, the monitors of the account it touched, and the people named as its responders. Its postmortem is one document saved whole:
- the impact in one line, and the sections a person writes or a record brings: summary, root cause, contributing factors, detection, diagnosis, response, repair, verification, what went well, what went badly, mistakes, open items;
- a timeline of entries (alert, monitor run, roll, note, status change), each with a time and, where the record gave one, the time as written ("10-04 ~18:00");
- action items with an owner, a due date, a status and the pull request, RFC or issue that tracks them;
- evidence links, typed (
pr:320,run:…, a record, a verdict), never pasted log content.
Impact is measured, not typed. Each incident carries figures computed from its monitors' runs between start and resolution: runs, failed runs, error rate. The duration comes from the times. A number in a postmortem's text is the author's; the figures beside it are the platform's.
Read where it was seen
The owner's acceptance test for phase 1: open a monitor's page and read the postmortems written for it, there. So the monitor page has:
- each incident that touched the monitor as a shaded band over the latency chart, in its severity's colour, titled;
- under the chart, "Incidents and postmortems": every incident that names the monitor, newest first, each a card that opens to the whole postmortem (Markdown rendered safely, readable on a phone), the newest open by itself;
- a click on a band opens its card and scrolls to it.
The account sees only its own incidents. Reliability gains an Incidents page (list, filters by status and severity, a new draft) and a page per incident.
Editing where it is read
An Edit button on each postmortem, on the monitor page and on its own page: the sections, the timeline (add, change, remove), action items, evidence, severity, status and times. Who may do what is decided by the service; a page only hides what would be refused:
| Who | Read | Edit a draft | Publish, edit what is published, name responders, delete |
|---|---|---|---|
| A member of the account | yes | when named a responder | no |
| The owner or an admin | yes | yes | yes |
| A key or API token (after launch) | incidents:read |
incidents:write |
no |
Every save names the version it was opened at. When someone saved since, the service refuses it (ABORTED, HTTP 409) and says who, so two people never overwrite each other unseen. Every save is a version with who, when and what ("edited", "draft → published", "restored from version 3"). Any version can be read again, each section diffed against the current one, and restored, which is a new version: nothing is ever overwritten.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Records from docs/incidents/ load unchanged
docs/incidents/ stays the format: the learning loop (RFCs 0045 to 0047) reads
those files, and their layout (summary, timeline, every check and how to read
it, repair, verification, mistakes, open items) is the postmortem's. A record
loads through ImportIncidentRecord (owner or admin) or the service's import
command:
- the front matter gives the header (id, title, status, opened, resolved, severity; "mitigated" reads as monitoring), the severity's words in brackets the impact line;
- each section maps to its field, a section the format does not name is kept under the diagnosis with its heading, the timeline table becomes entries that keep the record's own times, and the open items or action items become action items;
- the evidence statement is kept as written (in-toto JSON, unsigned until platform signing exists, and marked so), and the record, statement and scenarios file are the incident's evidence links;
- a record loaded before becomes a new version of the same incident, keeping who may edit it and whether it is published.
The first two loaded are the storage outage (INC-2026-10-04-storage) and the
API gateway's memory leak (INC-2026-10-05-protocol-memory), each linked to the
monitor that measured it, as drafts the owner reviews.
The operations journal (RFC 0053)
RFC 0053 keeps a team's operational knowledge: commands, facts about hosts,
runbooks. Incidents and postmortems are this product (accounts, monitors,
evidence, publishing); the journal does not store incidents of its own. A
journal entry links to an incident by its id (inc_…), and an incident links to
the journal's runbooks and commands as evidence. Both RFCs say so.
API and access
Every RPC has a REST route under /v1/accounts/orgs/{org_id}/incidents through
the gateway, so the SDKs and iohr get it; lists are paged one way (RFC 0033),
filtered by status, severity and monitor_id, with include_postmortem for
the monitor page. An incident keeps its monitors' category and tags from when it
opened (RFC 0037), and the list filters by category and tag too. Until launch the product is admin-only: a person needs the
platform's admin role, and keys and API tokens are refused. The scopes
incidents:read and incidents:write are written into the service now and
enter the key catalogue, the gateway's scope table and 's RBAC at launch,
so the public document does not list a product nobody can open yet.
Every call is audited (who, what, outcome), never a postmortem's content.
Erasure works end to end: an account's incidents go with it, and a person's
subject is replaced wherever it names them (responders, savers). Export lists
every incident of the account with its postmortem, and for a person also the
incidents in any account that name them. The accounts sweeper and "download your
data" reach it through [services] incidents (ACCOUNTS_INCIDENTS_URL).
Where it goes: what an enterprise customer needs
Phase 1 is shippable today. What follows is the target, in order.
Phase 2, the enterprise core.
- A review flow: draft, in review (named reviewers approve or ask for changes), published; comments on a section, mentions, notifications through the bell and mail.
- Templates per account (sections, required fields), custom fields (affected customers, regulatory reportable yes or no), severity definitions per account.
- Two faces of one incident: the internal postmortem and a customer-facing summary, published separately.
- Restricted incidents, visible only to named people (security incidents).
- Action items synced with GitHub issues through connections, so their status comes back; reminders when overdue.
- Metrics on a Reliability dashboard: time to acknowledge and to resolve, incidents by severity and month, overdue action items.
- Exports: a postmortem as PDF or Markdown, the account's incidents as CSV or JSON.
- Events and webhooks:
incident.opened,incident.updated,incident.resolved,postmortem.published, so customers wire their own tools. - Opening a draft incident from a monitor's alert (deduplicated per monitor while one is open), resolving suggested when it recovers, and rolls imported into the timeline from the commit each deployment records.
Phase 3.
- A public status page and subscribers per account: a customer's own status page, not only ours. Ours shows only published postmortems of the platform's own account, through a read-only endpoint that never serves a draft.
- A first draft of the postmortem from the timeline by the self-hosted model, labelled as AI-drafted (AI Act Art. 50), edited and published by a person; a customer's incident never reaches a third-party model.
- Retention settings per account.
- An export shaped for a DORA major ICT-related incident report, for customers in banking. A format, not a claim of compliance.
Controls
This product keeps incident records, which incident-management controls ask for (SOC 2 CC7.3 to CC7.5, ISO/IEC 27001, 2022 edition, A.5.24 to A.5.27). It can become the evidence the compliance programme's incident register points at; the register records that once the platform's own incidents are kept here. Nothing here makes the platform compliant with anything.
Alternatives considered
- A third-party incident tool. It would not know monitors, runs or rolls, and it would hold personal data and postmortems outside the platform. Rejected.
- Postmortems in the operations journal (RFC 0053). The journal is team knowledge; an incident is a record with an account, severity, monitors, versions, access and publishing. They link, they do not merge.
- The repository records only. They stay the learning loop's input and load into this product unchanged; a customer cannot read a repository.
Status log
- 2026-10-05: opened. Phase 1 built: the incidents service (schema
incidents, versions, measured figures, roles, stale saves refused, restore, records fromdocs/incidents/loaded with every section kept, export and erasure), the monitor page's bands and postmortems, the Incidents list and page, editing and version history with section diffs. Not yet: keys and scopes (at launch), and everything under phases 2 and 3. - 2026-10-05: phase 1b, the monitor page.
GetMonitorSeries(GET …/monitors/{id}/series,connections:read) reads a range in up to 500 buckets on the server: runs, passes, failures, uptime and the p50, p95 and p99 of the passing runs per bucket and over the range, and the failures by class. A monitor keeps an availability objective (slo_target, 0.9 to below 1). The page: a range from an hour to 90 days, an availability strip, the objective and its error budget, the latency percentiles with failures marked, failures by class, the runs, the incidents and the configuration. - 2026-10-06: erasure and export wired. Phase 1 had the incidents service's
EraseSubjectandExportSubjectbut the accounts sweeper and the person's export did not call it, so an erased account's incidents stayed. The sweeper now asks incidents with the others (ACCOUNTS_INCIDENTS_URL), a person's export includes their own account's incidents and those that name them, and a test erases an account through a real incidents service. - 2026-10-07: Checked: phase 1 (#336), the monitor page (#366, #399) and erasure (#392) run in incidents, connections, accounts and console-ui. Not built: the public incidents API (owner, 2026-10-07) and alerts that open incidents. Open.
- 2026-10-08: Alerts open incidents (RFC 0074.1, item 1), after the owner had two
monitors down for hours with no incident and no page. Until now a monitor going down
only published
monitor.down/agent.check.down; nothing turned that into an incident, so nothing paged, mailed or reached incident.io. The incidents service now runs a reconciler that reads the monitors' state (theiralertingflag afterfail_afterfailures in a row): one incident per down episode, opened byalert, SEV2, paging the account's phones, publishingincident.opened(mailed to owners and admins, on the bell) and opened in incident.io through the account's connection; recovery adds a timeline entry and resolves it, here and in incident.io, once no monitor on it is down. Decision on who gets this by default (owner, 2026-10-08): on for InOrbit's own accounts (the accounts service's internal accounts), off for customers until they turn it on withPUT …/incident-settings(open_on_alert); the per-checknotifyof an agent's checks file stays what it was, theagent.alert.*mail for that check, and does not decide incidents. Resolving on recovery is the owner's call ("resolving closes both");[alerts] resolve_on_recovery = falsekeeps RFC 0074.1's first wording (move to monitoring). inorbithr/core#683.