This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0040.1decided2026-10-03

Declared checks and self-tests

The agent's own file names what it watches, the platform turns it into monitors bound to that agent, and tests that must fail run beside them, so a guard that stops guarding is an alert and not a green dot.

Part of RFC 0040 Chaos as a service, our test bench in the customer's network

Problem

The first agent on our own server ran its first checks on 2026-10-03, each typed by hand. Monitors (RFC 0037) keep a check running, but someone has to create a connection and a monitor per target through the console or the API. On a fleet of agents that is a spreadsheet nobody keeps current, and it lives far from the machine that does the work.

Worse, every check we have asks one question: does it answer? A guard that silently stops guarding answers too. If the agent's policy stopped refusing a private address, if the platform stopped refusing a host outside the account's domains, if the gateway started answering without a token, every monitor would stay green.

Proposal

The file

Next to agent.toml and policy.toml, an optional checks.toml lists what this agent watches:

[[check]]
name = "api"
surface = "http"
target = "https://api.example.com/healthz"
every = "60s"
expect = { status = 200, max_ms = 2000 }
fail_after = 2
rfc = "0029"

[[check]]
name = "api-tls"
surface = "tls"
target = "api.example.com"
every = "1h"
expect = { valid_for_days = 14 }

[[refuse]]
name = "private-is-refused"
target = "https://db.internal/"
by = "policy"
every = "5m"

[[refuse]]
name = "outside-domains-is-refused"
target = "https://example.com/"
by = "platform"
every = "5m"

[[refuse]]
name = "api-needs-a-token"
surface = "http"
target = "https://api.example.com/v1/me"
expect = { status = 401 }
every = "5m"

A check may name the RFC it proves (rfc); monitors gain the same optional field, and the RFC's page then shows "proved since" with the check's history (RFC 0041).

iohr agent checks lint validates it offline: every target inside the policy (or meant to be refused by it), every interval at or above the floor, names unique, no secret in the file (credentials stay references, as everywhere in the agent).

How it becomes monitors

On every connection the agent sends the file's content hash and its checks in the hello. The platform reconciles: each declared check becomes a monitor managed by that agent (created, changed or removed to match), with the same rules as any monitor: inside the agent's verified domains, at or above the one-minute floor, within the account's ceiling. A check the platform will not accept is reported back with the reason and shown on the agent's page; the rest still run.

Managed monitors are read-only in the console and the API, marked "declared by" and the agent's name, so there is one source of truth: the file. Pausing one in the console is allowed and recorded; it resumes when the file changes or the person resumes it.

The agent gets no new credential. Declaring is part of its existing session, the platform decides what to accept, and a stolen agent can only declare checks inside the domains its account already proved, at the same ceilings.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Tests that must fail

[[refuse]] entries are checks whose success is a refusal:

by Passes when
policy the agent's own policy refuses the target (nothing is sent)
platform the platform refuses the job before it reaches the agent
(with expect) the target answers with the expected refusal, e.g. 401 or a gRPC UNAUTHENTICATED

A refusal check that stops being refused goes down and raises monitor.down like any monitor, with the class guard_open. Those are the alerts that matter most, so they are never folded into an uptime figure.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Self-tests the agent runs without a file

Every agent, with or without checks.toml, reports on itself:

  • Round trip. Time from the platform sending a job to the result arriving, as a monitor with its own latency history.
  • Clock. Its clock against the platform's server_time, alerting past a skew that would break short-lived tokens.
  • Policy hash. The policy it runs against the one it last reported; a change outside a restart is an event.
  • Version. The running version against the newest release the account allows.

Alternatives considered

Give the agent an API token to create monitors. Simple, and it turns a stolen machine into a way to create work on the account. Declaring through the session keeps the platform as the one that decides.

Keep monitors only in the console. Fine for five targets, wrong for fifty agents; the file lives with the machine and goes through the customer's own review.

Make negative tests a separate product. They are monitors whose success is a refusal; one model, one history, one alert path.

Decision

Decided 2026-10-07, built and live as proposed: checks.toml reconciled into agent-managed monitors through the session, [[refuse]] checks with guard_open alerts, and the four built-in self-tests. First proof: the agent on our own server declares www, api, docs and console, the TLS expiry of each, and the three refusals above, replacing the hand-made monitors.

Publication

The agent's docs gain the checks.toml reference and iohr agent checks lint; the developer docs' monitors page explains managed monitors and refusal checks; the console shows managed monitors under their agent.

Status log

  • 2026-10-03: opened, after the first hand-typed checks from our own agent.
  • 2026-10-04: built and live. The agent (0.1.0-alpha.4) sends checks.toml in its hello and lints it offline; the platform reconciles it into monitors declared by the agent, judges refusals (policy, platform, answer) and certificate expiry, and records clock skew, round trip and policy changes. The agent on our own server declares nine checks: four surfaces every minute, two certificates hourly, three guards every five minutes. All nine are up. The first run found a bug of ours: nine monitors ran at once, the agent's ceiling turned three away, and the platform paused them as refused; a busy agent now counts as unseen, and new monitors start spread over a minute.
  • 2026-10-07: Decided: built and live. Declared checks (#200), the console's report (#193) and the busy-agent fix (#215) run in agents, connections and console-ui at dab02d1a.
  • 2026-10-08: RFC 0100.1 proposes that the agent also schedule declared checks itself and keep their history locally, so monitors work with no platform; marked for the owner's decision (RFC 0100, D2). The platform's scheduling stays as built until then.

← Back to Platform