This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0078open2026-10-06

A broken monitor explains itself

A failing or silent monitor gets a first report and judgement from measured facts before anyone looks; routes send it on, to incident.io as deduplicated alerts.

Problem

On 2026-10-06 the owner's phone showed monitors as "not reporting" and others as failing, and nothing more. Not reporting can mean five different things: the monitor was paused, the account's agent that runs it is offline, its runs are refused because the domain is not verified, the scheduler is behind, or the target never answers. Failing can mean the target is down, slow, or answering wrong. A person has to open the console, read runs and guess. At night, on the phone, that is the expensive part.

The platform already holds the facts that tell these apart: every run with its error class (RFC 0037), the agent's sessions (RFC 0029), rolls and merged changes (RFC 0041, RFC 0077), past incidents on the same monitors (RFC 0058). Nobody puts them together until a person does it by hand.

And the alert stays inside InOrbit. A company that already runs incident.io, PagerDuty or Opsgenie wants the signal there, in the place its on-call rotation and routing live. Connections to those tools exist (RFC 0044), but they can only open and update incidents, which bypasses the company's own alert routing and is the fastest way to annoy a team that has tuned it for years.

Proposal

The first report

When a monitor turns down or silent, the platform writes a first report, and updates it on every run until the monitor is healthy again. It holds only facts the platform measured or recorded, each with its source:

  • Since when: the first failed or missing run, and the last good one.
  • What fails: the checks and error classes, how many runs, whether it is every run or some, and whether latency rose before it failed.
  • Never measured: runs that did not reach the target, and why: the agent was away (with when it was last connected), the domain is not verified, the monitor is paused and by what.
  • What changed shortly before: rolls and merged pull requests in the hour before, from RFC 0041 and RFC 0077.
  • Around it: other monitors that share its target or its agent, failing at the same time.
  • Last time: earlier incidents on the same monitors and how they ended.

A first judgement, with its evidence

From those facts, rules (not a model) sort the case into one judgement that names the facts it rests on:

Judgement When
The target is down runs reach the target and fail by connection, timeout or status
The target is slow runs pass but latency is over the monitor's limit, or rising
The target answers wrong status or contract failures with the target reachable
Our side cannot measure the account's agent is away, or runs are refused before they leave
Paused the monitor is paused, with why
Not yet known none of the above fits

"Our side cannot measure" matters most: it tells a woken person that the customer's system may be fine and the problem is the measurement, which is a different person's job. A judgement is labelled as preliminary and automatic, and it never closes or resolves anything. A later version may add a written summary by the platform's self-hosted model, labelled as AI-written (AI Act Art. 50); the facts stay the source.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Where it shows

  • The phone: the monitor's page opens with "What we know so far" (the judgement, its evidence, what to check first); the home's "needs attention" rows show the judgement in a word; a silent monitor shows why it is silent instead of only "not reporting".
  • The console: the same card on the monitor's page, and on the incident page when the monitor opens one.
  • The incident: when an incident opens from a monitor, the first report becomes the first entry of its timeline and the draft of the postmortem's detection section, both editable and labelled as automatic.

Routes: the team's own tools hear of it

A connection to a paging or chat tool gets routes. A route says which signals go where in that tool:

  • what: monitors (all, a group, a tag, named ones), which changes (down, silent, recovered, an incident opened or resolved), from which severity;
  • where: a destination inside the tool, chosen from what the tool offers: an incident.io alert source, a PagerDuty service, a Slack channel, an Opsgenie team;
  • how loud: the urgency the tool should give it, when the tool supports one.

An account can route the same signal to several places (the phone, incident.io, a Slack channel) and different monitors to different teams. A route can be sent a test from the console, and is off until the person who made it switches it on.

incident.io and tools like it: arrive as an alert, not an incident

For a company that already runs incident.io, the platform behaves like any other alert source it has. The route's destination is an HTTP alert source the company creates in incident.io (or picks, if it has one for InOrbit); the platform:

  1. sends firing when a monitor turns down or silent, with a stable deduplication key per monitor and kind, the first report's judgement as the title, its facts as the description, metadata (account, monitor, judgement, error classes) for their routing conditions, and a link back to the monitor;
  2. sends updates under the same key only when the judgement changes, never on every run;
  3. sends resolved with the same key when the monitor recovers;
  4. stays within the source's rate limit, groups a burst (many monitors behind one agent) into one alert, and backs off when refused.

The company's own alert routes decide whether an alert opens an incident, who is paged and how; the platform never opens, edits or closes an incident there unless the company also grants those actions (RFC 0044's grants). PagerDuty's Events API and Opsgenie's alert API follow the same shape: deduplication key, trigger, resolve.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Through the agent

Monitors run by an account's own agent (RFC 0029) can go silent because the agent is. The agent reports, with its session, what it could not run and why (no route to the target, its policy refused it, it is overloaded), so a silent monitor names its cause instead of waiting for a timeout. That report travels on the agent's existing session; the agent gains no new access.

Security and compliance

  • Least privilege at the tool. Sending alerts needs only the alert source's own token; no key with incident permissions is required, and the incident actions of RFC 0044 stay ungranted unless the company grants them.
  • Data minimisation. An alert carries monitor names, judgements, error classes, times and a link; never response bodies, credentials or customer content.
  • Off by default. A route sends nothing until a person switches it on, and every send is audited with the route, the destination and the outcome.
  • No surprises for the other tool. Deduplication, grouping, rate limits and a resolve for every firing, so a company's incident.io never fills with duplicates or stale alerts from us.
  • Disclosure. A judgement is labelled automatic; a future model summary is labelled as AI-written.

Plan

  1. This RFC.
  2. The phone: "What we know so far" from what the app already reads (runs, error classes, pause reasons, the agent's last run, incidents on the monitor), so it helps before the backend part lands.
  3. The first report and the judgement on the server, as an API, with the agent's report of what it could not run; the console shows it.
  4. Routes on connections, the incident.io alert-source destination with deduplicate, update and resolve, a test send; then PagerDuty and Slack.
  5. The first report as an incident's first timeline entry and detection draft.

Alternatives considered

  • Open incidents in the company's tool directly. Already possible (RFC 0044), and wrong as a default: it skips the company's alert routing and floods a tuned setup. Alerts let their rules decide.
  • A model writes the first report. Fluent, and unchecked: the person must trust it at 3 AM. Rules over measured facts first, a labelled summary later.
  • Send every run's result. Simple, and noise. Changes only, deduplicated.

Decision

Open. The owner asked on 2026-10-06 that broken and silent monitors explain themselves on the phone and the site before anyone investigates, that the explanation start the judging process, and that alerts reach the company's existing tools without disturbing them, through channels chosen per connection.

Publication

Public.

Status log

  • 2026-10-06: Opened.
  • 2026-10-07: Checked: only the app's card (#494) is merged. Open.

← Back to Platform