This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0058open2026-10-05

Incidents and postmortems, read on the monitor that saw them

An account's incidents and their postmortems as a product of the platform, in place of a third-party incident tool. Each incident names the monitors it touched; the monitor's page draws it over its chart and shows the whole postmortem under it, its impact measured from the monitor's own runs. Every save is a version, a stale save is refused, and an owner or admin publishes. The incident records already kept in the repository load into it unchanged. Phase 1 is admin-only; review, templates, customer-facing summaries, metrics, exports and a status page follow.

Problem

On 2026-10-04 and 2026-10-05 the platform was down twice for storage, and once partly down for a memory leak in the API gateway. The monitors saw all three: the drops are in their charts. What happened is written up in docs/incidents/ (one record per incident: the Markdown, an RFC 0046 evidence statement, a scenarios file for tests). Those records are good, and they are in a repository: a customer cannot read them, the monitor that saw the outage does not show them, and nothing connects a postmortem's claims to the measurements that back them.

The usual answer is a third-party incident tool. It would hold our postmortems and our customers' outside the platform, it would not know our monitors, runs or rolls, and it would be one more vendor with personal data. And a postmortem is the last step of the loop this platform is built around (docs/product-direction.md, RFCs 0045 to 0047): understand, observe, change, verify, and keep the evidence. An incident product that cannot point at measured evidence is a document store.

Proposal

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

The incident and its postmortem

An incident belongs to one account. Its header: title, severity (SEV1 to SEV4), status (investigating, identified, monitoring, resolved), when it started, was detected and was resolved, the monitors of the account it touched, and the people named as its responders. Its postmortem is one document saved whole:

  • the impact in one line, and the sections a person writes or a record brings: summary, root cause, contributing factors, detection, diagnosis, response, repair, verification, what went well, what went badly, mistakes, open items;
  • a timeline of entries (alert, monitor run, roll, note, status change), each with a time and, where the record gave one, the time as written ("10-04 ~18:00");
  • action items with an owner, a due date, a status and the pull request, RFC or issue that tracks them;
  • evidence links, typed (pr:320, run:…, a record, a verdict), never pasted log content.

Impact is measured, not typed. Each incident carries figures computed from its monitors' runs between start and resolution: runs, failed runs, error rate. The duration comes from the times. A number in a postmortem's text is the author's; the figures beside it are the platform's.

Read where it was seen

The owner's acceptance test for phase 1: open a monitor's page and read the postmortems written for it, there. So the monitor page has:

  1. each incident that touched the monitor as a shaded band over the latency chart, in its severity's colour, titled;
  2. under the chart, "Incidents and postmortems": every incident that names the monitor, newest first, each a card that opens to the whole postmortem (Markdown rendered safely, readable on a phone), the newest open by itself;
  3. a click on a band opens its card and scrolls to it.

The account sees only its own incidents. Reliability gains an Incidents page (list, filters by status and severity, a new draft) and a page per incident.

Editing where it is read

An Edit button on each postmortem, on the monitor page and on its own page: the sections, the timeline (add, change, remove), action items, evidence, severity, status and times. Who may do what is decided by the service; a page only hides what would be refused:

Who Read Edit a draft Publish, edit what is published, name responders, delete
A member of the account yes when named a responder no
The owner or an admin yes yes yes
A key or API token (after launch) incidents:read incidents:write no

Every save names the version it was opened at. When someone saved since, the service refuses it (ABORTED, HTTP 409) and says who, so two people never overwrite each other unseen. Every save is a version with who, when and what ("edited", "draft → published", "restored from version 3"). Any version can be read again, each section diffed against the current one, and restored, which is a new version: nothing is ever overwritten.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Records from docs/incidents/ load unchanged

docs/incidents/ stays the format: the learning loop (RFCs 0045 to 0047) reads those files, and their layout (summary, timeline, every check and how to read it, repair, verification, mistakes, open items) is the postmortem's. A record loads through ImportIncidentRecord (owner or admin) or the service's import command:

  • the front matter gives the header (id, title, status, opened, resolved, severity; "mitigated" reads as monitoring), the severity's words in brackets the impact line;
  • each section maps to its field, a section the format does not name is kept under the diagnosis with its heading, the timeline table becomes entries that keep the record's own times, and the open items or action items become action items;
  • the evidence statement is kept as written (in-toto JSON, unsigned until platform signing exists, and marked so), and the record, statement and scenarios file are the incident's evidence links;
  • a record loaded before becomes a new version of the same incident, keeping who may edit it and whether it is published.

The first two loaded are the storage outage (INC-2026-10-04-storage) and the API gateway's memory leak (INC-2026-10-05-protocol-memory), each linked to the monitor that measured it, as drafts the owner reviews.

The operations journal (RFC 0053)

RFC 0053 keeps a team's operational knowledge: commands, facts about hosts, runbooks. Incidents and postmortems are this product (accounts, monitors, evidence, publishing); the journal does not store incidents of its own. A journal entry links to an incident by its id (inc_…), and an incident links to the journal's runbooks and commands as evidence. Both RFCs say so.

API and access

Every RPC has a REST route under /v1/accounts/orgs/{org_id}/incidents through the gateway, so the SDKs and iohr get it; lists are paged one way (RFC 0033), filtered by status, severity and monitor_id, with include_postmortem for the monitor page. An incident keeps its monitors' category and tags from when it opened (RFC 0037), and the list filters by category and tag too. Until launch the product is admin-only: a person needs the platform's admin role, and keys and API tokens are refused. The scopes incidents:read and incidents:write are written into the service now and enter the key catalogue, the gateway's scope table and 's RBAC at launch, so the public document does not list a product nobody can open yet.

Every call is audited (who, what, outcome), never a postmortem's content. Erasure works end to end: an account's incidents go with it, and a person's subject is replaced wherever it names them (responders, savers). Export lists every incident of the account with its postmortem, and for a person also the incidents in any account that name them. The accounts sweeper and "download your data" reach it through [services] incidents (ACCOUNTS_INCIDENTS_URL).

Where it goes: what an enterprise customer needs

Phase 1 is shippable today. What follows is the target, in order.

Phase 2, the enterprise core.

  • A review flow: draft, in review (named reviewers approve or ask for changes), published; comments on a section, mentions, notifications through the bell and mail.
  • Templates per account (sections, required fields), custom fields (affected customers, regulatory reportable yes or no), severity definitions per account.
  • Two faces of one incident: the internal postmortem and a customer-facing summary, published separately.
  • Restricted incidents, visible only to named people (security incidents).
  • Action items synced with GitHub issues through connections, so their status comes back; reminders when overdue.
  • Metrics on a Reliability dashboard: time to acknowledge and to resolve, incidents by severity and month, overdue action items.
  • Exports: a postmortem as PDF or Markdown, the account's incidents as CSV or JSON.
  • Events and webhooks: incident.opened, incident.updated, incident.resolved, postmortem.published, so customers wire their own tools.
  • Opening a draft incident from a monitor's alert (deduplicated per monitor while one is open), resolving suggested when it recovers, and rolls imported into the timeline from the commit each deployment records.

Phase 3.

  • A public status page and subscribers per account: a customer's own status page, not only ours. Ours shows only published postmortems of the platform's own account, through a read-only endpoint that never serves a draft.
  • A first draft of the postmortem from the timeline by the self-hosted model, labelled as AI-drafted (AI Act Art. 50), edited and published by a person; a customer's incident never reaches a third-party model.
  • Retention settings per account.
  • An export shaped for a DORA major ICT-related incident report, for customers in banking. A format, not a claim of compliance.

Controls

This product keeps incident records, which incident-management controls ask for (SOC 2 CC7.3 to CC7.5, ISO/IEC 27001, 2022 edition, A.5.24 to A.5.27). It can become the evidence the compliance programme's incident register points at; the register records that once the platform's own incidents are kept here. Nothing here makes the platform compliant with anything.

Alternatives considered

  • A third-party incident tool. It would not know monitors, runs or rolls, and it would hold personal data and postmortems outside the platform. Rejected.
  • Postmortems in the operations journal (RFC 0053). The journal is team knowledge; an incident is a record with an account, severity, monitors, versions, access and publishing. They link, they do not merge.
  • The repository records only. They stay the learning loop's input and load into this product unchanged; a customer cannot read a repository.

Status log

  • 2026-10-05: opened. Phase 1 built: the incidents service (schema incidents, versions, measured figures, roles, stale saves refused, restore, records from docs/incidents/ loaded with every section kept, export and erasure), the monitor page's bands and postmortems, the Incidents list and page, editing and version history with section diffs. Not yet: keys and scopes (at launch), and everything under phases 2 and 3.
  • 2026-10-05: phase 1b, the monitor page. GetMonitorSeries (GET …/monitors/{id}/series, connections:read) reads a range in up to 500 buckets on the server: runs, passes, failures, uptime and the p50, p95 and p99 of the passing runs per bucket and over the range, and the failures by class. A monitor keeps an availability objective (slo_target, 0.9 to below 1). The page: a range from an hour to 90 days, an availability strip, the objective and its error budget, the latency percentiles with failures marked, failures by class, the runs, the incidents and the configuration.
  • 2026-10-06: erasure and export wired. Phase 1 had the incidents service's EraseSubject and ExportSubject but the accounts sweeper and the person's export did not call it, so an erased account's incidents stayed. The sweeper now asks incidents with the others (ACCOUNTS_INCIDENTS_URL), a person's export includes their own account's incidents and those that name them, and a test erases an account through a real incidents service.
  • 2026-10-07: Checked: phase 1 (#336), the monitor page (#366, #399) and erasure (#392) run in incidents, connections, accounts and console-ui. Not built: the public incidents API (owner, 2026-10-07) and alerts that open incidents. Open.
  • 2026-10-08: Alerts open incidents (RFC 0074.1, item 1), after the owner had two monitors down for hours with no incident and no page. Until now a monitor going down only published monitor.down/agent.check.down; nothing turned that into an incident, so nothing paged, mailed or reached incident.io. The incidents service now runs a reconciler that reads the monitors' state (their alerting flag after fail_after failures in a row): one incident per down episode, opened by alert, SEV2, paging the account's phones, publishing incident.opened (mailed to owners and admins, on the bell) and opened in incident.io through the account's connection; recovery adds a timeline entry and resolves it, here and in incident.io, once no monitor on it is down. Decision on who gets this by default (owner, 2026-10-08): on for InOrbit's own accounts (the accounts service's internal accounts), off for customers until they turn it on with PUT …/incident-settings (open_on_alert); the per-check notify of an agent's checks file stays what it was, the agent.alert.* mail for that check, and does not decide incidents. Resolving on recovery is the owner's call ("resolving closes both"); [alerts] resolve_on_recovery = false keeps RFC 0074.1's first wording (move to monitoring). inorbithr/core#683.

← Back to Platform