Problem
An autonomous agent that changes this platform needs to ask two questions on its own: "run the bench against what I am about to change" and "what is different now". Both need runs it can start without a person, read as data rather than as a page, and compare.
chaos serve has runs, and they were built for one person at one console. There is one
run slot for the whole server: a second POST /runs while one is active answers 409,
and everything else waits in one queue. Schedules are cron expressions in UTC, kept in a
JSON file; a schedule that falls due while its previous job is still queued or running is
skipped and counted, never stacked. A run streams as server-sent events (started,
phase, load, timeline, stress, finding, finished), and a client that connects
late gets everything from started on. Results are files on a volume. Notifications are
one Slack webhook. There are no accounts, no tokens, no units and no way to say what
changed between two runs.
Proposal
A run
Every piece of work in this family is a run of an account, in a new reliability service:
| Field | Contents |
|---|---|
kind |
checks, load, scenario, faults, capacity, replay |
scenario |
its id and the version hash it ran (RFC 0040.5), when it has one |
environment, executor |
where it ran: our cloud, or an agent by id |
started_by |
a person, an API token, a schedule or a trigger, with the change it was started for when given (change = "core#212") |
status |
queued, running, passed, failed, error, cancelled, refused (with the policy's or the plan's reason) |
result |
the report of its kind: checks, load figures, objective, sweep points |
findings |
ids of the findings it created or saw again (RFC 0040.6) |
units |
reserved and spent |
No body appears in a run, its frames or its events.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Started and read by an agent
The same operations on every surface, for the same caller and scopes:
| API | MCP tool | iohr |
|
|---|---|---|---|
| start | POST /v1/reliability/runs |
reliability_start_run |
iohr reliability run |
| read | GET /v1/reliability/runs/{id} |
reliability_get_run |
iohr reliability runs show |
| wait | the live stream below | reliability_wait_run (bounded, returns the record) |
--wait |
| compare | GET /v1/reliability/compare?before=&after= |
reliability_compare_runs |
iohr reliability compare |
Starting needs the scope reliability:run; reading needs reliability:read. Neither
lets a token confirm a fault in production (RFC 0040.4), and both are API-token scopes
(RFC 0016), so an autonomous agent holds exactly these and nothing more.
Did my change break anything
compare takes two runs of the same scenario version and target, one before a change and
one after, and answers per measure: surface checks that changed verdict, invariants that
broke or healed, findings new to the second run, and p50, p99 and the knee with their
intervals. A latency difference counts only where the intervals do not overlap; the
answer says "no measurable difference" otherwise. Two runs of different scenario
versions or environments are refused, not compared.
The answer is short enough for a pull request: regressed, improved or unchanged,
then one line per difference.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Queues, no global slot
Runs queue per account, with as many running at once as the plan allows, and per agent,
within its policy's max_concurrent_jobs. Load, fault and capacity runs are exclusive
per agent: two at once on one agent would measure each other. In our cloud, accounts take
turns, so one account's long run never delays another's.
Schedules
A schedule runs a scenario on a cron expression in UTC, as today. One that falls due while
its previous run is still queued or running is skipped and counted; missed fires while
the service was down are skipped, not run in a burst, as monitors do (RFC 0037). Each kind
has a floor per plan (load no more than hourly). A schedule may queue a fault in
production; it waits for a person and lapses as RFC 0040.4 says.
The live view
GET /v1/reliability/runs/{id}/events streams today's frames (started, phase, load
once a second, timeline, check, finding, finished), each with an SSE id. A
viewer who joins late, or reconnects with Last-Event-ID, gets every frame from the
start or from that id, then live ones; a finished run answers with its finished frame.
The service streams over gRPC and the protocol serves it as SSE and over the multiplexed
WebSocket like every other stream.
Events instead of Slack
reliability.run.finished (run id, kind, status, environment) and
reliability.finding.created (finding id, run id, signature) join the event catalogue,
ids only, delivered as webhooks, MQTT or a stream (RFC 0022). Slack becomes one
subscriber among others.
Units, audit and retention
- Units. Reserved at start from what the run may do, spent as measured, the rest returned (RFC 0015). Scheduled runs are charged too; monitor runs stay as RFC 0037 left them.
- Audit. Every start, cancel, schedule change and production confirmation is one line: who, which run or schedule, which targets, which faults, the outcome.
- Retention. Run records 90 days, as monitor runs; per-second frames 30 days, after which the record keeps its summary; findings until 90 days after they close; audit lines one year. Deleting an account deletes all of it.
The console
Reliability gains Runs (live and past, with compare), Scenarios, Findings and Schedules beside Monitors and Agents. It stays admin-only until the owner opens it.
Later, customers' agents and CI use the same runs on their own accounts.
Alternatives considered
Keep one slot and queue everything. Simple, and one nightly capacity sweep would block every check an autonomous change asks for.
Compare by eye in the console. Fine for a person; an agent needs the verdict as data.
Slack as the notification. One destination, one workspace, and a payload with content. Events with ids reach Slack and everything else.
Decision
Open. Proposed: runs of an account with the fields above, started and read through the
API, MCP and iohr with two narrow scopes, a comparison of two runs that judges latency
by intervals, queues per account and per agent, schedules that skip, a live stream that
replays from the start, two events in place of Slack, units for every run, audit and
stated retention. First proof: every run against our platform (the checks, the nightly
load, the confirmed staging faults) goes through the Engineering team's queue, and every pull request an
autonomous agent opens carries the comparison of a run before and after its change.
Publication
The API reference gains the run, schedule and compare routes and the two scopes; the MCP
reference the tools; iohr gains iohr reliability; the event catalogue gains the two
events; the console gains the Reliability pages above.
Status log
- 2026-10-03: opened, from
chaos serve's runs, queue, schedules and live stream. - 2026-10-07: Checked: No PR names this RFC and the status log records nothing built. Open.