This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0040.5open2026-10-03

Scenarios as code

The scenario file (targets, load, a timeline of faults and the assertions that decide pass or fail) lives in the repository beside the code it tests, names connections instead of addresses, is checked by iohr and the console before it runs, and runs from CI or an autonomous agent with a token that can only start runs.

Part of RFC 0040 Chaos as a service, our test bench in the customer's network

Problem

When an autonomous agent changes a service, the test of how that service behaves under load and faults should change in the same pull request, be reviewed with it, and run before the merge. Today the scenarios that express that live in the chaos tool's own directory, run only in the process that boots them, and describe a stack rather than the platform as it is deployed.

A scenario today is one TOML file with five tables, each rejecting unknown keys so a typo fails at chaos check instead of doing nothing:

[scenario]      name, description, skip
[stack]         instances of our service kinds, booted on free loopback ports
[load]          rate, duration, warm-up, pattern, a weighted mix of our operations
[[timeline]]    set_behavior, set_store_behavior, stop, start, log, at an offset
[assertions]    error rate, p50, p99, requests, throughput, per-service counters

chaos check also refuses references to instances that do not exist, a fault on a kind without fault injection, and load assertions without [load]. That checking is the part worth keeping. [stack] is the part that cannot survive: the services under test already run, and a scenario that starts its own copies tests the copies.

Proposal

The file

[scenario]
format = 1
name = "engine-under-latency"
description = "The protocol keeps p99 under 400 ms while the engine is slow"
environment = "staging"

[targets]
api = { connection = "public-api-staging" }
engine = { connection = "engine-staging", proxy = "engine" }

[load]
target = "api"
rate = 100
duration = "5m"
warmup = "30s"
max_in_flight = 128

[[load.operations]]
openapi = "me"
weight = 1

[[timeline]]
at = "1m"
action = "fault"
target = "engine"
fault = { type = "latency", latency = "200ms", jitter = "50ms" }
for = "2m"

[assertions]
max_error_rate = 0.01
max_p99_ms = 400
checks = ["surfaces"]

[objective]
success_rate = 0.99
p99_ms = 400
window = "1m"

format is required and an unknown value is refused, so old files never change meaning.

From today's tables

Today In the file
[scenario] the same, plus format and the environment it may run in
[stack] gone; [targets] maps names to connections (RFC 0018) and, for faults, to a proxy the agent's policy declares (RFC 0040.4)
[load] the same keys; operations from the target's contract (RFC 0040.3)
set_behavior fault, from the menu of RFC 0040.4, with a mandatory for lifetime
set_store_behavior no equivalent; drop_answer is the nearest
stop / start restart of one pod, when the proxy gains cluster actions
log the same
[assertions] the same bounds; services.* only where the service's count is measured; checks runs the surface checks of RFC 0040.2 before and after
(none) [objective] for fault runs (RFC 0040.4)

Connections, never addresses

A target names a connection of the account. A URL, a host or an address anywhere in the file is refused at check time. The connection decides the executor (our cloud or an agent), the credential reference and the verified domain, so a scenario copied between repositories cannot reach anything its account has not already connected.

Checked before it runs

iohr reliability check FILE... runs the offline half: the file parses, every table denies unknown keys, durations and offsets are sane, every timeline target exists and a fault names a target with a proxy. It then asks the platform for the online half: every connection exists, its agent's policy allows the work and the proxy, every value is under the plan's and the policy's ceilings, and a fault in production is flagged as needing a person. Output is ok FILE (NAME) or error FILE: MESSAGE per file, exit 1 on any error, as chaos check prints today.

The console's scenario editor runs the same two halves as it is typed. There is one parser and one set of rules; the console and iohr call the same check.

Versions

Every scenario saved to the platform, from the console or by a run, is kept as a version with its content hash. A run records the version it ran, so a result is always read against the file that produced it.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

From CI and from an autonomous agent

iohr reliability run scenarios/engine-under-latency.toml --wait

uploads the file as a version, starts a run, streams its progress and exits 0 on pass and 1 on fail, with the run's id printed for later reads. In CI or in an agent it uses an API token (RFC 0016) with the scope reliability:run: it can start and read runs of the account's scenarios and nothing else, and it cannot confirm a fault in production. The same is the MCP tool reliability_run_scenario.

Later, a customer keeps scenarios in their own repository against their own connections.

Alternatives considered

Keep [stack] and boot copies. Useful for our local development, where chaos run keeps it. For the bench it tests something other than what runs.

Scenarios only in the console. Easy to start, and nothing reviews a change to them beside the code they test.

YAML. More common in CI tooling; our scenarios, campaigns and agent files are TOML already, and one format keeps one parser.

Decision

Open. Proposed: the scenario file with format, [targets] of connections in place of [stack], the timeline's faults from RFC 0040.4 with a lifetime, the same assertions, an offline and an online check shared by iohr and the console, versions by hash, and runs from CI or an agent with a token that can only run. First proof: our shipped baseline, error_injection and latency scenarios rewritten against connections to our staging environment (today the same deployment as production), kept in our repository, checked by iohr on every change and run before every merge an autonomous agent proposes.

Publication

The developer docs gain the scenario reference (generated from the parser, as the chaos kinds page is today) and iohr reliability check and run; the console gains the scenario editor.

Status log

  • 2026-10-03: opened, from the chaos tool's scenario format and chaos check.
  • 2026-10-07: Checked: No PR names this RFC and the status log records nothing built. Open.

This document mentions

Mentioned in

← Back to Platform