Problem
When an autonomous agent changes a service, the test of how that service behaves under load and faults should change in the same pull request, be reviewed with it, and run before the merge. Today the scenarios that express that live in the chaos tool's own directory, run only in the process that boots them, and describe a stack rather than the platform as it is deployed.
A scenario today is one TOML file with five tables, each rejecting unknown keys so a typo
fails at chaos check instead of doing nothing:
[scenario] name, description, skip
[stack] instances of our service kinds, booted on free loopback ports
[load] rate, duration, warm-up, pattern, a weighted mix of our operations
[[timeline]] set_behavior, set_store_behavior, stop, start, log, at an offset
[assertions] error rate, p50, p99, requests, throughput, per-service counters
chaos check also refuses references to instances that do not exist, a fault on a kind
without fault injection, and load assertions without [load]. That checking is the part
worth keeping. [stack] is the part that cannot survive: the services under test already
run, and a scenario that starts its own copies tests the copies.
Proposal
The file
[scenario]
format = 1
name = "engine-under-latency"
description = "The protocol keeps p99 under 400 ms while the engine is slow"
environment = "staging"
[targets]
api = { connection = "public-api-staging" }
engine = { connection = "engine-staging", proxy = "engine" }
[load]
target = "api"
rate = 100
duration = "5m"
warmup = "30s"
max_in_flight = 128
[[load.operations]]
openapi = "me"
weight = 1
[[timeline]]
at = "1m"
action = "fault"
target = "engine"
fault = { type = "latency", latency = "200ms", jitter = "50ms" }
for = "2m"
[assertions]
max_error_rate = 0.01
max_p99_ms = 400
checks = ["surfaces"]
[objective]
success_rate = 0.99
p99_ms = 400
window = "1m"
format is required and an unknown value is refused, so old files never change meaning.
From today's tables
| Today | In the file |
|---|---|
[scenario] |
the same, plus format and the environment it may run in |
[stack] |
gone; [targets] maps names to connections (RFC 0018) and, for faults, to a proxy the agent's policy declares (RFC 0040.4) |
[load] |
the same keys; operations from the target's contract (RFC 0040.3) |
set_behavior |
fault, from the menu of RFC 0040.4, with a mandatory for lifetime |
set_store_behavior |
no equivalent; drop_answer is the nearest |
stop / start |
restart of one pod, when the proxy gains cluster actions |
log |
the same |
[assertions] |
the same bounds; services.* only where the service's count is measured; checks runs the surface checks of RFC 0040.2 before and after |
| (none) | [objective] for fault runs (RFC 0040.4) |
Connections, never addresses
A target names a connection of the account. A URL, a host or an address anywhere in the file is refused at check time. The connection decides the executor (our cloud or an agent), the credential reference and the verified domain, so a scenario copied between repositories cannot reach anything its account has not already connected.
Checked before it runs
iohr reliability check FILE... runs the offline half: the file parses, every table
denies unknown keys, durations and offsets are sane, every timeline target exists and a
fault names a target with a proxy. It then asks the platform for the online half: every
connection exists, its agent's policy allows the work and the proxy, every value is under
the plan's and the policy's ceilings, and a fault in production is flagged as needing a
person. Output is ok FILE (NAME) or error FILE: MESSAGE per file, exit 1 on any error,
as chaos check prints today.
The console's scenario editor runs the same two halves as it is typed. There is one parser
and one set of rules; the console and iohr call the same check.
Versions
Every scenario saved to the platform, from the console or by a run, is kept as a version with its content hash. A run records the version it ran, so a result is always read against the file that produced it.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
From CI and from an autonomous agent
iohr reliability run scenarios/engine-under-latency.toml --wait
uploads the file as a version, starts a run, streams its progress and exits 0 on pass and
1 on fail, with the run's id printed for later reads. In CI or in an agent it uses an API
token (RFC 0016) with the scope reliability:run: it can start and read runs of the
account's scenarios and nothing else, and it cannot confirm a fault in production. The
same is the MCP tool reliability_run_scenario.
Later, a customer keeps scenarios in their own repository against their own connections.
Alternatives considered
Keep [stack] and boot copies. Useful for our local development, where chaos run
keeps it. For the bench it tests something other than what runs.
Scenarios only in the console. Easy to start, and nothing reviews a change to them beside the code they test.
YAML. More common in CI tooling; our scenarios, campaigns and agent files are TOML already, and one format keeps one parser.
Decision
Open. Proposed: the scenario file with format, [targets] of connections in place of
[stack], the timeline's faults from RFC 0040.4 with a lifetime, the same assertions,
an offline and an online check shared by iohr and the console, versions by hash, and
runs from CI or an agent with a token that can only run. First proof: our shipped
baseline, error_injection and latency scenarios rewritten against connections to
our staging environment (today the same deployment as production), kept in our
repository, checked by iohr on every change and
run before every merge an autonomous agent proposes.
Publication
The developer docs gain the scenario reference (generated from the parser, as the chaos
kinds page is today) and iohr reliability check and run; the console gains the
scenario editor.
Status log
- 2026-10-03: opened, from the chaos tool's scenario format and
chaos check. - 2026-10-07: Checked: No PR names this RFC and the status log records nothing built. Open.