Problem
We test this platform with a chaos tool of our own. It checks every surface a service speaks (REST, gRPC, server-sent events, WebSocket, MQTT, GraphQL, MCP), including whether a call without a verified caller is refused; it drives load on a fixed clock, so a slow server cannot slow the test down and hide its own latency; it injects faults on a timeline; it runs campaigns that check every answer against a model of what should have happened, and turns a broken rule into a finding shrunk to the fewest steps that still break it, which can be replayed against another environment. The Break-it playground showed the same faults can be safe in strangers' hands when they come from a small menu, cost from a budget and expire on their own.
None of it can be offered as it stands. It reaches only our services, faults need hooks compiled into them, it trusts any address it is given, and its server has no accounts, no tenancy and no limits of its own. Meanwhile the companies who would use it spend their days on the same questions about their own systems, with tools that each do one part.
Proposal
Reliability is a new area of the console, the same objects over the API, the MCP
server and iohr. It runs against connections (RFC 0018): every service under test
is a connection of kind service or endpoint, and every check, load or fault is one of
its actions, with the action classes, grants, audit and units that come with it.
Where it runs
| From | May reach | May do |
|---|---|---|
| Our cloud | addresses inside a verified domain of the account (RFC 0030) | checks, light load within the plan's ceilings |
| The company's agent (RFC 0029) | what its local policy allows, private networks included | checks, load, faults, within the policy's ceilings |
Faults never run from our cloud.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
What it does
- Checks. Each surface a service speaks, checked the way the chaos tool checks ours: it answers, it answers in the right shape, it refuses a caller it cannot identify, an MCP server lists no tool that writes unless it should. Results per check, with latency, kept as history.
- Load, honestly. Requests on a fixed clock at a set rate or ramp, a weighted mix of operations, warm-up excluded. The report puts each limit next to what was measured, separates what was sent from what the service counted, and classes errors.
- Scenarios as code. A file that names targets, load, a timeline of actions and the
assertions that decide pass or fail, kept in the company's repository and checked
against the platform before it runs, as the chaos tool's
checkdoes today. The console edits the same file, checked as it is typed. - Faults with a budget. Through an agent acting as a proxy in front of a service: a small menu (added latency, a share of errors, dropped connections, later a restarted pod), each move with a cost from a budget the account sets and a lifetime after which it ends by itself. An objective stated up front (a success rate and a p99 latency over a window) turns a run into a pass, a fail and a time to breach.
- Findings. A broken assertion or contract becomes a finding: what was expected, what happened, the steps that led there. The same finding seen again is grouped, not repeated; it is shrunk to its smallest reproduction; and it can be replayed against another environment, which records whether it reproduced.
- Capacity. A sweep raises one parameter step by step, repeats each step, puts confidence intervals on the latency, and marks where the p99 breaks.
- Live, scheduled, connected. A run streams as it happens; runs can be scheduled;
results are events (
reliability.run.finished,reliability.finding.created) delivered as webhooks, MQTT or a stream (RFC 0022), and an incoming alert connection can start one.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Safe to hand to strangers
- Every parameter has a ceiling set by the plan and, on an agent, by its policy; nothing like the internal tool's unbounded hang or rate.
- Targets come only from connections, never from a typed address; the cloud's are inside a verified domain.
- Runs cost units from the account's one budget (RFC 0015) and queue per account; a run stops when its budget, its duration or its connection's grant ends.
- Every run, fault and finding is in the audit log as who, what and outcome, never request or response bodies.
- An environment named
productionneeds a person's confirmation before a fault, every time.
Console
Reliability holds Overview (recent runs, failing checks, open findings), Targets (the connections under test and their latest checks), Runs (live and past), Scenarios, Findings, Schedules and Agents (enrolled agents, their health, version and policy).
Alternatives considered
Offer the chaos tool as it is, hosted. It would reach only what the internet can, take any address and share between companies. Each of those is a reason it stays internal.
Faults only, as chaos products usually start. Faults without checks, honest load and findings tell a team that something broke, not what or whether it is fixed.
Integrate an existing open-source chaos project. Chaos Mesh and LitmusChaos act and do that well; neither checks surfaces, models load honestly or shrinks findings. arrive through the agent later either way.
Let tests run from our cloud against any address. The simplest product and an attack tool. Verified domains for the cloud and the agent's policy for everything else close it.
Decision
Open. Proposed: Reliability in the console, built on connections, run from our cloud only against verified domains and with checks and light load, and through the company's agent for everything else, with faults only through the agent's proxy, every parameter bounded, units for every run and findings that can be replayed.
Publication
The console gains Reliability. The developer docs gain the scenario reference, the
checks per surface, the fault menu with its costs and the events; the API reference gains
the Reliability endpoints and scopes; iohr gains iohr reliability.
Status log
- 2026-10-03: opened, with RFC 0028 (extensions), 0029 (the agent) and 0030 (domains), on top of connections (RFC 0018).
- 2026-10-03: the first part, monitors, is RFC 0037: a connection's check kept running
from our cloud or the company's agent, with health, uptime, latency and
monitor.down/monitor.recoveredevents. - 2026-10-07: Checked: monitors (RFC 0037) and declared checks (RFC 0040.1) are live; the rest of the family is not built. Open.