Problem
An autonomous change that passes its tests can still make a route three times slower. To catch that, the bench has to put the same load on the platform before and after the change and compare, with numbers that a slow server cannot bend.
The chaos tool has the right pacer. Each request is due at a point on an absolute clock
and is sent when it is due, however long earlier ones take, so a slow backend shows up as
latency and not as a quietly lower rate. Warm-up runs the same load and is discarded;
max_in_flight bounds concurrency. Patterns are constant and ramp.
Three things keep it from being the bench:
- It speaks our operations only.
rest_evaluate,grpc_ping,ledger_appendand twenty more, each written by hand. A new route has no operation until someone writes it. - It runs anywhere and without ceilings.
rate,durationandmax_in_flighttake any value, from any machine, against any URL. - It hides its own stall. When
max_in_flightis reached the pacer waits for a permit, and the latency of the delayed request is timed from when it left, not from when it was due. The overdue requests then go out at once. That under-reports latency exactly when the service is slow, the error wrk2 calls coordinated omission (wrk2).
The agent (RFC 0029) refuses load today: load = false is the only value its policy
accepts.
Proposal
The pacer in the agent
The same algorithm, in the agent, with two changes. Every request records its latency
from the time it was due as well as from the time it was sent, and the report counts late
sends and the seconds the pacer spent stalled at max_in_flight. A run that could not
hold its rate says so instead of printing a lower throughput.
Patterns: constant and ramp first, step (a rate held for a while, then the next)
and burst (a short spike over a base rate) later. Warm-up is excluded from every number.
Operations from the target's contract
An operation is read from the contract attached to the connection (RFC 0040.2), not written as code:
[load]
target = "public-api"
rate = 200
duration = "5m"
warmup = "30s"
timeout = "5s"
max_in_flight = 128
pattern = { type = "ramp", start_rate = 50, end_rate = 200 }
[[load.operations]]
openapi = "me"
weight = 6
[[load.operations]]
grpc = "iohr.engine.v1.EngineService/Evaluate"
request = { subject_id = "load-{n}" }
weight = 2
An HTTP operation is an OpenAPI operation with its parameters from a small set of values written in the scenario; a gRPC operation is a unary method whose request is JSON turned into protobuf with the descriptors. The response is read only to classify it (status, gRPC code, and, when asked, the shape check of RFC 0040.2) and is dropped. Request values in a scenario are test values, never data copied from production.
What the report says
| Line | Contents |
|---|---|
| Asked and held | the rate asked, the rate sent, late sends, seconds stalled |
| Latency | p50, p90, p99, max, each from the due time and from the send time |
| Sent and counted | requests sent; requests the service counted, or "not measured" |
| Errors | by class: HTTP status, gRPC code, timeout, transport, contract |
| Limits | every ceiling that applied, beside the value used |
What the service counted is measured in one of two ways only: a counter the target exposes,
read through a connection, or the agent's proxy (RFC 0040.4) counting what passed through
it. Otherwise the report says "not measured", never "equal to sent". The gap between the
two is where a gateway drops or retries, which today's per_target beside services
already shows for our own stack.
No body is kept: snapshots once a second carry counts, classes and latency quantiles.
Hard ceilings, lower wins
A run's limits come from the plan and from the agent's policy; the lower of the two applies. A run that asks for more is refused with the limit it broke, not quietly cut, because a cut run measures something nobody asked for.
[work]
load = true
[load]
max_rate = 500
max_duration = "15m"
max_in_flight = 256
max_runs_per_hour = 4
The agent enforces the duration on its own clock and stops when its session to the platform drops. From our cloud, load is light and only against hosts inside a verified domain (RFC 0030), under low ceilings on rate and run length that are set before it ships.
Units
Each run reserves units from the account's budget (RFC 0015) for the requests it may send, spends what it sent, and returns the rest. A run that exhausts its reservation stops.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
How an agent uses it
A load run is a run (RFC 0040.8): started through the API, the MCP tool
reliability_run_load or iohr reliability load, read as the report above. An
autonomous change runs the same named load before and after; the comparison is RFC
0040.8's.
Later, customers get the same load against their own connections, under their own policy. None of it is offered before it has run against this platform.
Alternatives considered
k6, Gatling or wrk2 on the agent. Mature generators, and wrk2 avoids coordinated omission by design. None reads a gRPC descriptor and an OpenAPI document as one mix, keeps sent apart from counted, or obeys the agent's policy; wrapping one would leave the ceilings in a second program.
Closed-loop load (a fixed number of virtual users). Simple, and a slow service then
receives fewer requests and looks healthier than it is. Closed loop stays available as an
sweep of max_in_flight in RFC 0040.7, named as such.
Clamp a run to the ceiling. Friendlier, and the result answers a different question.
Decision
Open. Proposed: the open-loop pacer in the agent with latency from the due time and the stall reported, operations from OpenAPI and protobuf, sent beside counted only where it is measured, ceilings from the plan and the policy with refusal above them, light load only from our cloud, units per run. First proof: a nightly run from the agent on our own server against our public API at our own ceilings, with what the gateway counted beside what was sent, and the same run before and after every change our agents roll.
Publication
The developer docs gain load operations, the report's fields and the ceilings; the agent's
policy reference gains [load].
Status log
- 2026-10-03: opened, from the chaos tool's pacer and the agent's refusal of load today.
- 2026-10-07: Checked: No PR names this RFC and the status log records nothing built. Open.