Problem
The first agent on our own server ran its first checks on 2026-10-03, each typed by hand. Monitors (RFC 0037) keep a check running, but someone has to create a connection and a monitor per target through the console or the API. On a fleet of agents that is a spreadsheet nobody keeps current, and it lives far from the machine that does the work.
Worse, every check we have asks one question: does it answer? A guard that silently stops guarding answers too. If the agent's policy stopped refusing a private address, if the platform stopped refusing a host outside the account's domains, if the gateway started answering without a token, every monitor would stay green.
Proposal
The file
Next to agent.toml and policy.toml, an optional checks.toml lists what this agent
watches:
[[check]]
name = "api"
surface = "http"
target = "https://api.example.com/healthz"
every = "60s"
expect = { status = 200, max_ms = 2000 }
fail_after = 2
rfc = "0029"
[[check]]
name = "api-tls"
surface = "tls"
target = "api.example.com"
every = "1h"
expect = { valid_for_days = 14 }
[[refuse]]
name = "private-is-refused"
target = "https://db.internal/"
by = "policy"
every = "5m"
[[refuse]]
name = "outside-domains-is-refused"
target = "https://example.com/"
by = "platform"
every = "5m"
[[refuse]]
name = "api-needs-a-token"
surface = "http"
target = "https://api.example.com/v1/me"
expect = { status = 401 }
every = "5m"
A check may name the RFC it proves (rfc); monitors gain the same optional field, and the
RFC's page then shows "proved since" with the check's history (RFC 0041).
iohr agent checks lint validates it offline: every target inside the policy (or meant to
be refused by it), every interval at or above the floor, names unique, no secret in the
file (credentials stay references, as everywhere in the agent).
How it becomes monitors
On every connection the agent sends the file's content hash and its checks in the hello. The platform reconciles: each declared check becomes a monitor managed by that agent (created, changed or removed to match), with the same rules as any monitor: inside the agent's verified domains, at or above the one-minute floor, within the account's ceiling. A check the platform will not accept is reported back with the reason and shown on the agent's page; the rest still run.
Managed monitors are read-only in the console and the API, marked "declared by" and the agent's name, so there is one source of truth: the file. Pausing one in the console is allowed and recorded; it resumes when the file changes or the person resumes it.
The agent gets no new credential. Declaring is part of its existing session, the platform decides what to accept, and a stolen agent can only declare checks inside the domains its account already proved, at the same ceilings.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Tests that must fail
[[refuse]] entries are checks whose success is a refusal:
by |
Passes when |
|---|---|
policy |
the agent's own policy refuses the target (nothing is sent) |
platform |
the platform refuses the job before it reaches the agent |
(with expect) |
the target answers with the expected refusal, e.g. 401 or a gRPC UNAUTHENTICATED |
A refusal check that stops being refused goes down and raises monitor.down like any
monitor, with the class guard_open. Those are the alerts that matter most, so they are
never folded into an uptime figure.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Self-tests the agent runs without a file
Every agent, with or without checks.toml, reports on itself:
- Round trip. Time from the platform sending a job to the result arriving, as a monitor with its own latency history.
- Clock. Its clock against the platform's
server_time, alerting past a skew that would break short-lived tokens. - Policy hash. The policy it runs against the one it last reported; a change outside a restart is an event.
- Version. The running version against the newest release the account allows.
Alternatives considered
Give the agent an API token to create monitors. Simple, and it turns a stolen machine into a way to create work on the account. Declaring through the session keeps the platform as the one that decides.
Keep monitors only in the console. Fine for five targets, wrong for fifty agents; the file lives with the machine and goes through the customer's own review.
Make negative tests a separate product. They are monitors whose success is a refusal; one model, one history, one alert path.
Decision
Decided 2026-10-07, built and live as proposed: checks.toml reconciled into agent-managed monitors through the session,
[[refuse]] checks with guard_open alerts, and the four built-in self-tests. First
proof: the agent on our own server declares www, api, docs and console, the TLS expiry of
each, and the three refusals above, replacing the hand-made monitors.
Publication
The agent's docs gain the checks.toml reference and iohr agent checks lint; the
developer docs' monitors page explains managed monitors and refusal checks; the console
shows managed monitors under their agent.
Status log
- 2026-10-03: opened, after the first hand-typed checks from our own agent.
- 2026-10-04: built and live. The agent (0.1.0-alpha.4) sends
checks.tomlin its hello and lints it offline; the platform reconciles it into monitors declared by the agent, judges refusals (policy, platform, answer) and certificate expiry, and records clock skew, round trip and policy changes. The agent on our own server declares nine checks: four surfaces every minute, two certificates hourly, three guards every five minutes. All nine are up. The first run found a bug of ours: nine monitors ran at once, the agent's ceiling turned three away, and the platform paused them as refused; a busy agent now counts as unseen, and new monitors start spread over a minute. - 2026-10-07: Decided: built and live. Declared checks (#200), the console's report (#193) and the busy-agent fix (#215) run in agents, connections and console-ui at dab02d1a.
- 2026-10-08: RFC 0100.1 proposes that the agent also schedule declared checks itself and keep their history locally, so monitors work with no platform; marked for the owner's decision (RFC 0100, D2). The platform's scheduling stays as built until then.