This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0037open2026-10-03

Monitors, a connection's check kept running

A monitor runs a connection's check on a schedule, from our cloud or the company's agent, keeps the results, says when a target goes down and when it recovers, and is the first part of Reliability.

Problem

A connection of kind endpoint can already be checked: from our cloud for a host inside a verified domain of the account, or from the company's own agent for anything its local policy allows (RFC 0018, 0029, 0030). The check answers whether the target replied, with which status, how fast, and why not when it did not. It answers once, when someone asks.

The question a team actually has is whether it is up now, whether it was up last night, and who hears about it when it is not. That needs the same check run on a clock, the answers kept, and a change of state turned into an event. We need it first ourselves: the platform's own hosts are checked by our own agent, by hand, today.

Proposal

A monitor is a connection's check with a schedule: every so many seconds, with the expectation the check already takes (a status, a time limit), and how many failures in a row count as down. It belongs to the connection, so it inherits everything the connection has: its account, its executor (our cloud or a named agent), its credential by reference, its audit and its erasure.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

What it keeps and says

  • Runs. Each run keeps the time, the outcome, the latency, the status code and the class of error. Never a body. Runs are kept for 90 days.
  • Health. up, down or unknown. A monitor goes down after its set number of failures in a row (one to five, two unless said), and back up at the first success. An agent that is offline makes the monitor unknown, not down: the target was not seen, which is not the same as the target failing.
  • Summary. Uptime over 24 hours, 7 days and 30 days; the median and the 95th percentile latency; the last failure.
  • Events. monitor.down and monitor.recovered, with ids only, delivered like every other event as webhooks, MQTT or a stream (RFC 0022).

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Category and tags

A monitor carries one category from a fixed list (availability, transport, contract, security, performance, synthetic, or none) and up to ten tags, key:value labels such as env:prod or transport:rest. Both are checked by a pattern, so free text, and with it customer data, does not get in. Lists filter by a category and by tags, every tag holding. The category is a bounded set and labels the runs metric; tags never become a metric label. An incident opened on monitors keeps the category and tags they had at that moment, so an incident list filters the same way. A monitor an agent declared takes its labels from the agent's checks file, like the rest of it.

When it stops by itself

A monitor whose setup became wrong pauses with the reason instead of failing forever: its domain is no longer verified, its connection is paused or gone, its agent was revoked. Deleting a connection deletes its monitors.

Bounded

  • At most one run a minute per monitor, and a set number of monitors per account.
  • Every run goes through the same path as a check someone asked for: the same target rules, the same verified-domain rule for our cloud, the same agent policy.
  • Runs are spread by a scheduler that any number of service replicas share; a slot is run once, and slots missed while the service was down are skipped, not run in a burst.
  • Runs from our cloud are light by construction: one request, no body kept, a time limit of at most 30 seconds.

API and console

Under the account's connections: create a monitor on a connection, list, read, change, pause and delete them, read their runs and their summary. Reads need connections:read, changes connections:write. Writes honour Idempotency-Key (RFC 0033).

The console gains Reliability, starting with Monitors (health, uptime, latency over time, the runs) and Agents. A connection that can be checked offers "Monitor this".

Alternatives considered

A separate reliability service now. Reliability will grow load, faults, scenarios and findings (RFC 0031), and those belong in a service of their own. A monitor is a connection's check on a clock; putting it in a new service now would need a new way for one service to call another's actions, and one more hop per run, for no gain yet. The routes are named so monitors can move behind them later.

A time-series database for runs. Ninety days of one run a minute is about 130,000 rows per monitor, which a relational table with one index answers comfortably. A column store comes when the numbers say so.

Charging every run. Every call through the API costs units from the account's one budget (RFC 0015). Runs started by the scheduler are not API calls; for now they are bounded by the ceilings above, and charging them is a follow-up.

Decision

Open. Proposed: monitors as part of connections, run by a shared scheduler through the existing check path, with runs kept 90 days, health with a failure threshold, events on a change of state, a one-minute floor and a per-account ceiling, and Reliability in the console starting with Monitors and Agents.

Publication

The developer docs gain "Monitor an endpoint" and the monitor routes in the API reference; the event catalogue gains monitor.down and monitor.recovered; the console gains Reliability.

Status log

  • 2026-10-03: opened as the first part of Reliability (RFC 0031), on connections (RFC 0018), the agent (RFC 0029) and domains (RFC 0030).
  • 2026-10-03: built in the connections service: the monitor routes, the shared scheduler, runs kept 90 days, health with a failure threshold, the two events, and pausing with the reason on a refusal. The console's Reliability area follows.
  • 2026-10-06: monitors carry a category and up to ten tags, set on create and update and taken from an agent's checks file for a declared monitor; the monitor and incident lists filter by them, and an incident keeps its monitors' labels from when it opened. The runs metric gains the category as a label. The console shows them as chips, filters by them in a rail beside the list, groups the overview by category and edits tags in the monitor dialog.
  • 2026-10-07: Checked: monitors (#168), Reliability in the console (#166) and the missed-events resend (#170) are live. Category and tags (#501) rolled on 2026-10-07: connections, agents, incidents, protocol and console-ui run dab02d1a. Open while the monitors sweep adds to it.

← Back to Platform