Problem
A single load run says whether the platform held one rate. It does not say how far it is from falling over, and that distance is what an autonomous change most easily takes away: a query that grew a join, a lock held a little longer, a pool shrunk by a config default. The tests pass, the nightly run at 200 requests a second passes, and the point where p99 leaves its limit has moved from 500 to 300 without anyone seeing it.
The stress crate measures this for the ledger. A campaign with [sweep] runs once per
value of one parameter (owner workers, contention workers, max_in_flight and a few
more), repeat times each, every repeat with fresh workers, its own warm-up and its own
measured phase. The repeats of a point pool into one sample, and p50 and p99 come back
with the 95 % interval a percentile bootstrap gives (1000 resamples, seeded, so the same
measurements give the same interval). The knee is the first point whose p99 passes a
limit or grows by a factor over the previous point. Invariants are judged throughout, so
a point that is fast and wrong is still a finding.
It sweeps ledger workers only, against the ledger only, without ceilings. Scenarios (load on everything else) have no sweep at all.
Proposal
The sweep
A sweep is a load run (RFC 0040.3) repeated over the values of one parameter:
[capacity]
target = "api"
parameter = "rate"
values = [100, 200, 300, 400, 500, 600]
repeat = 3
point_duration = "2m"
warmup = "30s"
pause = "30s"
[capacity.knee]
p99_ms = 250
factor = 2.0
[capacity.stop]
error_rate = 0.05
parameter |
What changes | Loop |
|---|---|---|
rate |
requests per second | open: the rate is held, latency shows the strain |
max_in_flight |
requests in flight | closed: throughput shows the strain |
weight:<operation> |
one operation's share of the mix | open, at the scenario's rate |
A rate sweep times latency from each request's due time (RFC 0040.3), so a stalled pacer cannot hide the knee by sending less. Every point reports the rate it achieved beside the rate it asked for; a point that could not hold its rate is marked, not averaged in.
Statistics, as the stress crate does them
Repeats of a point pool into one sample; p50 and p99 carry a 95 % percentile-bootstrap interval from 1000 resamples seeded from the run. Two points whose intervals overlap did not measure differently, and the report says so instead of drawing a slope between them.
Two limits are stated with every result:
- A p99 rests on the slowest one per cent of its samples. Each point shows how many samples it has, and a point with fewer than 1000 cannot be the knee.
- The interval covers the noise inside this run, on this day. It says nothing about how the same sweep varies from one night to the next; that is what the history is for.
The knee and when to stop
The knee is the first point whose p99 passes knee.p99_ms or is knee.factor times the
previous point's, as today. The sweep stops raising the parameter at the first point
whose error rate passes stop.error_rate or whose objective (RFC 0040.4) breaks: pushing
a service that has already failed measures nothing new and costs its owner an outage.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Inside the ceilings
Every value is checked against the plan and the agent's policy before the sweep starts,
and a value over a ceiling is refused at check time, not cut. The whole sweep (points times
repeats times warm-up, measured phase and pause) must also fit the policy's
max_duration; a sweep that does not fit is refused with its computed length. Units are
reserved for the whole sweep and returned for points not run.
Invariants at every point
The contract invariants (RFC 0040.6) and the shape checks (RFC 0040.2) are judged throughout. A finding names the point that found it.
What a person reads
The table, one line per point, as chaos stress prints it today, and one sentence:
The knee is between 400 and 500 rps: p99 rose from 118 ms [109, 127] to 312 ms
[281, 349], over the 250 ms limit. 300 and 400 rps did not measure differently.
Every point held its rate. No findings.
The numbers above illustrate the format; they are not a measurement.
How an agent uses it
A sweep is a run (RFC 0040.8), started through the API, the MCP tool
reliability_run_capacity or iohr reliability capacity, and read as the points, the
knee and the sentence. An autonomous change that touches a hot path compares the knee of
its branch with the last knee of main; a knee that moved down by more than one point is
a regression the agent reports.
Later, customers sweep their own services within their own policy.
Alternatives considered
Report the mean of the repeats with a standard deviation. Latency is skewed and p99 is not normal; the bootstrap makes no assumption about the shape.
Find the knee by bisection. Fewer points, and every point depends on the last, so one noisy point sends the search the wrong way. A fixed list is predictable in cost and duration, which the ceilings need.
Run until it breaks. The usual load test. On a shared environment it is an outage by design; the stop rule ends the sweep at the first failed point.
Decision
Open. Proposed: sweeps over rate, max_in_flight or one operation's weight, repeats
pooled with bootstrap intervals and sample counts, the knee rule of the stress crate, a
stop at the first failed point, every value and the total length inside the ceilings,
invariants at every point, and a one-sentence result. First proof: a monthly rate sweep
of our public API's read mix from the agent on our own server within our own ceilings,
its knee recorded in this RFC's status log, and the sweep run before and after every
autonomous change that touches the protocol or the engine.
Publication
The developer docs gain the capacity reference and how to read an interval; the console shows a sweep as its table, its chart and its sentence.
Status log
- 2026-10-03: opened, from the stress crate's sweeps, bootstrap intervals and knee.
- 2026-10-07: Checked: No PR names this RFC and the status log records nothing built. Open.