Problem
We test this platform with a tool of our own. It checks every surface a service speaks, drives load on a fixed clock so a slow server cannot hide its latency, injects faults on a timeline, checks every answer of a campaign against a model and shrinks a broken rule to the fewest steps that still break it. It found real bugs here before anyone else did.
Three things keep it from being a product:
- It only knows our services. Every check asserts our own contract (our paths, our RPCs, our topics); every load operation speaks our protos; every fault is a hook compiled into our binaries.
- It trusts any address it is given. A check, a load run or a stack call goes wherever the URL says.
- It has no accounts, limits or tenancy. One run slot, one results directory, rates and durations without a ceiling.
There is a fourth reason, and it is ours. We are building agents that change this platform on their own: write code, open pull requests, roll services. An autonomous change is only as good as the check that judges it. Today that check is a person reading a diff and a test suite that knows nothing about how the platform behaves when it runs. We need a bench that answers, continuously and without being asked, whether every surface still answers, still refuses what it must refuse, still holds its latency under load and still recovers from a fault - and that an agent can call before and after its own change.
Meanwhile the agent (RFC 0029) already runs in a customer's network, under a policy the customer owns, bound to domains the customer proved (RFC 0030). It answers checks when someone asks, monitors (RFC 0037) keep a check running, and since 2026-10-04 the checks an agent declares in its own file become monitors, refusals included (RFC 0040.1). It refuses load and faults. It can do far more, and it should prove that on itself before anyone else trusts it.
Proposal
Chaos as a service is the chaos tool rebuilt on three things that already exist: the agent as the only place work runs inside a network, connections and monitors as the only targets, and the account's units as the only budget. It is a family of RFCs; this one holds what is common, each child one part.
What stays the same as the internal tool
The parts that make the tool worth having: open-loop load on an absolute clock, warm-up excluded, what was sent beside what the service counted; checks per surface including the refusal of an unknown caller; faults from a small menu that cost a budget and expire; findings with a signature, shrunk and replayable; capacity with confidence intervals.
What changes
| Today (internal) | As a service |
|---|---|
| Any URL | A connection inside a verified domain, or what the agent's policy allows |
| Our contracts only | The customer's contract: OpenAPI, protobuf descriptors, GraphQL schema, MCP tool list |
| Faults through hooks in our binaries | Faults through the agent acting as a proxy in front of a service |
| One global run slot | Runs queued per account, bounded per plan and per agent |
| No ceilings | Every parameter capped by the plan, then by the agent's policy, whichever is lower |
| Results on a volume | Runs, findings and history stored per account, audited as who, what, outcome |
| Slack messages | Events (RFC 0022), delivered as webhooks, MQTT or a stream |
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
The rules every child keeps
- Work runs only on an agent or, for public checks and light load, from our cloud against a verified domain. Faults never run from our cloud.
- The agent's policy wins. A child may add a capability to the agent; the policy file decides whether it is on, and the agent announces it in its hello.
- The agent never holds a write credential to the API. What it does is declared on its machine or asked of it by the control plane; it never creates work on its own account through the API.
- Bounded everything. Rate, duration, in-flight requests, fault lifetime, budget per run, runs per hour.
- Never a body. Checks, load and findings keep status, timing, classes and the shape of a failure, not request or response content, unless a child states otherwise and why.
- Production needs a person. A fault in an environment named
productionneeds a confirmation every time, from someone with the right role. - The bench does not share the fate of what it watches. At least one vantage point runs on another machine, network and power; a bench that goes quiet is itself an alert.
The children
| Part | What it adds | |
|---|---|---|
| 0040.1 | Declared checks and self-tests | checks.toml on the agent's machine becomes monitors bound to that agent; tests that must fail (policy refusals, domain refusals, unauthenticated calls), so a passing negative test is an alert |
| 0040.2 | Surface checks on every protocol | Generic checks for HTTP, gRPC, SSE, WebSocket, MQTT, GraphQL and MCP, each with "answers", "answers in the declared shape" and "refuses an unknown caller"; MCP's tool list checked for write tools |
| 0040.3 | Honest load through the agent | The open-loop pacer in the agent, generic HTTP and gRPC operations from the customer's contract, sent beside counted, hard ceilings |
| 0040.4 | Faults through the agent's proxy | Latency, a share of errors, dropped connections and later a restarted pod, each with a cost, a lifetime and an objective that turns a run into pass, fail and time to breach |
| 0040.5 | Scenarios as code | The scenario file (targets, load, timeline, assertions) in the customer's repository, checked by iohr and the console before it runs |
| 0040.6 | Findings that replay | Assertions on the customer's contract, failures grouped by signature, shrunk to the fewest steps, replayed against another environment |
| 0040.7 | Capacity | One parameter raised step by step, repeated, p50 and p99 with confidence intervals, the knee marked |
| 0040.8 | Runs, schedules and the live view | Account-scoped runs, a queue per account, schedules, a live stream, events, units, the console's Reliability pages |
| 0040.9 | Vantage points that do not share our fate | At least one agent off the machine, network and power it watches; a silent bench is an alert, not unknown |
| 0040.10 | Hosts and clusters, read first | The agent reads what a person would check by hand: disk, memory, certificates, units, pods, restarts, events, pressure; read-only by default, any action a fault with its rules |
| 0040.11 | Who may see what | Tenant isolation and scopes probed on every route: one account's token never reads another's data, a token without a scope gets 403, limits answer 429 |
| 0040.12 | Consistency across surfaces and data | The same question over REST, gRPC, GraphQL and the stream gives the same answer; writes read back; counts and totals agree |
| 0040.13 | Journeys in a browser | Sign-in and the console's main paths run in a headless browser by the agent, timed per step, a screenshot only on failure |
| 0040.14 | The change verdict, for people and AI assistants | One answer per change from the repository's own tests, the bench runs and the findings; the bench as MCP tools, and plugins for Claude and the other major assistants |
Our platform is the first account, and for now the only one
The first purpose is ours: the bench our own autonomous work runs against. Every child is built for this platform first and judged by whether it catches what our agents and our people break. Customers get the same parts later, when they have been proved here; until then nothing in this family is sold, and Reliability stays admin-only.
The bench is something an agent can use, not only a dashboard: a run can be started and
read through the API, MCP and iohr, so an autonomous change can ask "what did I break"
before it asks for a merge, and again after it rolls.
The agent on our own server (Engineering team, bound to inorbit.hr) is to run every
child as it lands, against this platform, continuously. Today it runs the first: declared
checks every minute and the negative tests every five (RFC 0040.1). Surface checks on
every protocol we serve, a nightly load run within our own ceilings and faults in staging
come with their children. Environments stay part of the model
(a staging and a production per account, agents and rules per environment), but our own
staging is, for now, the same deployment as production: an environment that shares a
deployment with production follows production's rules, so our faults are small, short
and confirmed by a person every time, and test writes go only to probe accounts the
platform marks as such, never to anyone else's data. The mark exists: RFC 0051 added the
probe flag on an account, set by a platform administrator and audited.
Each child's status log records what it has proved here and since when. Nothing is
offered to a customer that has not run against us first.
Alternatives considered
Host the internal tool as it is. Any address, our contracts, one shared server: each is a reason it stays internal.
Put the work in our cloud and leave the agent to checks. The cloud cannot reach a private network, and faults from it would be an attack tool. The agent is already in the right place, under a policy the customer owns.
Adopt an existing chaos or load product. Chaos Mesh and LitmusChaos act inside a container cluster; k6 and Gatling drive load. None checks every surface, keeps sent and counted apart, or shrinks a finding. Cluster actions can still arrive through the agent later.
Let the agent create its own work through the API. Simpler to build, and a stolen agent would then be able to start work against anything its account can reach. Declared work on the machine, enforced by the platform, keeps the control with whoever controls that machine and that account.
Decision
Open. Proposed: the family above, built for this platform first and only, as the bench our autonomous work and our people run against; offered to customers later, on the same parts, once proved here. Order, by what that bench needs first: 0040.1 (declared checks and self-tests), 0040.2 (every surface, refusals included), 0040.8 (runs an agent can start and read), then 0040.3 (load), 0040.5 (scenarios as code), 0040.4 (faults), 0040.6 (findings), 0040.7 (capacity). The gaps a bench like this misses on its own are children too, built as soon as the first three run: 0040.9 (an off-machine vantage point) at once, then 0040.10 (hosts and clusters), 0040.11 (who may see what), 0040.14 (the change verdict and the assistants' plugins), 0040.12 (consistency) and 0040.13 (browser journeys). Reliability stays admin-only in the console until the owner opens it.
Publication
Each child names its own docs. The family adds the Reliability pages of the console, the developer docs' Reliability section and the agent's policy reference for every new capability.
Status log
- 2026-10-03: opened as the family under RFC 0031, after monitors (RFC 0037) and the first checks the agent ran against this platform.
- 2026-10-06: brought up to date with main. The parts are listed under this page as its
sub-RFCs. Checked what is built: 0040.1 is live on our own agent; the account
probeflag exists (RFC 0051); RFC 0047'schaos verifyposts a per-claim verdict on our pull requests, which 0040.14 builds on; RFC 0050 opened the MCP server to every assistant; RFC 0061 holds traffic seen on a host and RFC 0063 our host's own alerts. Nothing else in the family is built. - 2026-10-07: Checked: 0040.1 is built and live (now decided), 0040.2 partly (#489); 0040.14 has a private plugin; the other parts are not built. Open.