This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0040.8open2026-10-03

Runs, schedules and the live view

Every check, load, fault, scenario, sweep and replay becomes a run of an account that an autonomous agent or a person can start and read through the API, MCP and iohr, queued per account and per agent, scheduled, streamed live, compared before and after a change, announced as events and kept for a stated time.

Part of RFC 0040 Chaos as a service, our test bench in the customer's network

Problem

An autonomous agent that changes this platform needs to ask two questions on its own: "run the bench against what I am about to change" and "what is different now". Both need runs it can start without a person, read as data rather than as a page, and compare.

chaos serve has runs, and they were built for one person at one console. There is one run slot for the whole server: a second POST /runs while one is active answers 409, and everything else waits in one queue. Schedules are cron expressions in UTC, kept in a JSON file; a schedule that falls due while its previous job is still queued or running is skipped and counted, never stacked. A run streams as server-sent events (started, phase, load, timeline, stress, finding, finished), and a client that connects late gets everything from started on. Results are files on a volume. Notifications are one Slack webhook. There are no accounts, no tokens, no units and no way to say what changed between two runs.

Proposal

A run

Every piece of work in this family is a run of an account, in a new reliability service:

Field Contents
kind checks, load, scenario, faults, capacity, replay
scenario its id and the version hash it ran (RFC 0040.5), when it has one
environment, executor where it ran: our cloud, or an agent by id
started_by a person, an API token, a schedule or a trigger, with the change it was started for when given (change = "core#212")
status queued, running, passed, failed, error, cancelled, refused (with the policy's or the plan's reason)
result the report of its kind: checks, load figures, objective, sweep points
findings ids of the findings it created or saw again (RFC 0040.6)
units reserved and spent

No body appears in a run, its frames or its events.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Started and read by an agent

The same operations on every surface, for the same caller and scopes:

API MCP tool iohr
start POST /v1/reliability/runs reliability_start_run iohr reliability run
read GET /v1/reliability/runs/{id} reliability_get_run iohr reliability runs show
wait the live stream below reliability_wait_run (bounded, returns the record) --wait
compare GET /v1/reliability/compare?before=&after= reliability_compare_runs iohr reliability compare

Starting needs the scope reliability:run; reading needs reliability:read. Neither lets a token confirm a fault in production (RFC 0040.4), and both are API-token scopes (RFC 0016), so an autonomous agent holds exactly these and nothing more.

Did my change break anything

compare takes two runs of the same scenario version and target, one before a change and one after, and answers per measure: surface checks that changed verdict, invariants that broke or healed, findings new to the second run, and p50, p99 and the knee with their intervals. A latency difference counts only where the intervals do not overlap; the answer says "no measurable difference" otherwise. Two runs of different scenario versions or environments are refused, not compared.

The answer is short enough for a pull request: regressed, improved or unchanged, then one line per difference.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Queues, no global slot

Runs queue per account, with as many running at once as the plan allows, and per agent, within its policy's max_concurrent_jobs. Load, fault and capacity runs are exclusive per agent: two at once on one agent would measure each other. In our cloud, accounts take turns, so one account's long run never delays another's.

Schedules

A schedule runs a scenario on a cron expression in UTC, as today. One that falls due while its previous run is still queued or running is skipped and counted; missed fires while the service was down are skipped, not run in a burst, as monitors do (RFC 0037). Each kind has a floor per plan (load no more than hourly). A schedule may queue a fault in production; it waits for a person and lapses as RFC 0040.4 says.

The live view

GET /v1/reliability/runs/{id}/events streams today's frames (started, phase, load once a second, timeline, check, finding, finished), each with an SSE id. A viewer who joins late, or reconnects with Last-Event-ID, gets every frame from the start or from that id, then live ones; a finished run answers with its finished frame. The service streams over gRPC and the protocol serves it as SSE and over the multiplexed WebSocket like every other stream.

Events instead of Slack

reliability.run.finished (run id, kind, status, environment) and reliability.finding.created (finding id, run id, signature) join the event catalogue, ids only, delivered as webhooks, MQTT or a stream (RFC 0022). Slack becomes one subscriber among others.

Units, audit and retention

  • Units. Reserved at start from what the run may do, spent as measured, the rest returned (RFC 0015). Scheduled runs are charged too; monitor runs stay as RFC 0037 left them.
  • Audit. Every start, cancel, schedule change and production confirmation is one line: who, which run or schedule, which targets, which faults, the outcome.
  • Retention. Run records 90 days, as monitor runs; per-second frames 30 days, after which the record keeps its summary; findings until 90 days after they close; audit lines one year. Deleting an account deletes all of it.

The console

Reliability gains Runs (live and past, with compare), Scenarios, Findings and Schedules beside Monitors and Agents. It stays admin-only until the owner opens it.

Later, customers' agents and CI use the same runs on their own accounts.

Alternatives considered

Keep one slot and queue everything. Simple, and one nightly capacity sweep would block every check an autonomous change asks for.

Compare by eye in the console. Fine for a person; an agent needs the verdict as data.

Slack as the notification. One destination, one workspace, and a payload with content. Events with ids reach Slack and everything else.

Decision

Open. Proposed: runs of an account with the fields above, started and read through the API, MCP and iohr with two narrow scopes, a comparison of two runs that judges latency by intervals, queues per account and per agent, schedules that skip, a live stream that replays from the start, two events in place of Slack, units for every run, audit and stated retention. First proof: every run against our platform (the checks, the nightly load, the confirmed staging faults) goes through the Engineering team's queue, and every pull request an autonomous agent opens carries the comparison of a run before and after its change.

Publication

The API reference gains the run, schedule and compare routes and the two scopes; the MCP reference the tools; iohr gains iohr reliability; the event catalogue gains the two events; the console gains the Reliability pages above.

Status log

  • 2026-10-03: opened, from chaos serve's runs, queue, schedules and live stream.
  • 2026-10-07: Checked: No PR names this RFC and the status log records nothing built. Open.

← Back to Platform