This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0040.3open2026-10-03

Honest load through the agent

The open-loop pacer moves into the agent, drives operations read from the target's own OpenAPI document or protos, reports what was asked beside what was measured and what was sent beside what was counted, and stops at hard ceilings the plan and the agent's policy both set.

Part of RFC 0040 Chaos as a service, our test bench in the customer's network

Problem

An autonomous change that passes its tests can still make a route three times slower. To catch that, the bench has to put the same load on the platform before and after the change and compare, with numbers that a slow server cannot bend.

The chaos tool has the right pacer. Each request is due at a point on an absolute clock and is sent when it is due, however long earlier ones take, so a slow backend shows up as latency and not as a quietly lower rate. Warm-up runs the same load and is discarded; max_in_flight bounds concurrency. Patterns are constant and ramp.

Three things keep it from being the bench:

  • It speaks our operations only. rest_evaluate, grpc_ping, ledger_append and twenty more, each written by hand. A new route has no operation until someone writes it.
  • It runs anywhere and without ceilings. rate, duration and max_in_flight take any value, from any machine, against any URL.
  • It hides its own stall. When max_in_flight is reached the pacer waits for a permit, and the latency of the delayed request is timed from when it left, not from when it was due. The overdue requests then go out at once. That under-reports latency exactly when the service is slow, the error wrk2 calls coordinated omission (wrk2).

The agent (RFC 0029) refuses load today: load = false is the only value its policy accepts.

Proposal

The pacer in the agent

The same algorithm, in the agent, with two changes. Every request records its latency from the time it was due as well as from the time it was sent, and the report counts late sends and the seconds the pacer spent stalled at max_in_flight. A run that could not hold its rate says so instead of printing a lower throughput.

Patterns: constant and ramp first, step (a rate held for a while, then the next) and burst (a short spike over a base rate) later. Warm-up is excluded from every number.

Operations from the target's contract

An operation is read from the contract attached to the connection (RFC 0040.2), not written as code:

[load]
target = "public-api"
rate = 200
duration = "5m"
warmup = "30s"
timeout = "5s"
max_in_flight = 128
pattern = { type = "ramp", start_rate = 50, end_rate = 200 }

[[load.operations]]
openapi = "me"
weight = 6

[[load.operations]]
grpc = "iohr.engine.v1.EngineService/Evaluate"
request = { subject_id = "load-{n}" }
weight = 2

An HTTP operation is an OpenAPI operation with its parameters from a small set of values written in the scenario; a gRPC operation is a unary method whose request is JSON turned into protobuf with the descriptors. The response is read only to classify it (status, gRPC code, and, when asked, the shape check of RFC 0040.2) and is dropped. Request values in a scenario are test values, never data copied from production.

What the report says

Line Contents
Asked and held the rate asked, the rate sent, late sends, seconds stalled
Latency p50, p90, p99, max, each from the due time and from the send time
Sent and counted requests sent; requests the service counted, or "not measured"
Errors by class: HTTP status, gRPC code, timeout, transport, contract
Limits every ceiling that applied, beside the value used

What the service counted is measured in one of two ways only: a counter the target exposes, read through a connection, or the agent's proxy (RFC 0040.4) counting what passed through it. Otherwise the report says "not measured", never "equal to sent". The gap between the two is where a gateway drops or retries, which today's per_target beside services already shows for our own stack.

No body is kept: snapshots once a second carry counts, classes and latency quantiles.

Hard ceilings, lower wins

A run's limits come from the plan and from the agent's policy; the lower of the two applies. A run that asks for more is refused with the limit it broke, not quietly cut, because a cut run measures something nobody asked for.

[work]
load = true

[load]
max_rate = 500
max_duration = "15m"
max_in_flight = 256
max_runs_per_hour = 4

The agent enforces the duration on its own clock and stops when its session to the platform drops. From our cloud, load is light and only against hosts inside a verified domain (RFC 0030), under low ceilings on rate and run length that are set before it ships.

Units

Each run reserves units from the account's budget (RFC 0015) for the requests it may send, spends what it sent, and returns the rest. A run that exhausts its reservation stops.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

How an agent uses it

A load run is a run (RFC 0040.8): started through the API, the MCP tool reliability_run_load or iohr reliability load, read as the report above. An autonomous change runs the same named load before and after; the comparison is RFC 0040.8's.

Later, customers get the same load against their own connections, under their own policy. None of it is offered before it has run against this platform.

Alternatives considered

k6, Gatling or wrk2 on the agent. Mature generators, and wrk2 avoids coordinated omission by design. None reads a gRPC descriptor and an OpenAPI document as one mix, keeps sent apart from counted, or obeys the agent's policy; wrapping one would leave the ceilings in a second program.

Closed-loop load (a fixed number of virtual users). Simple, and a slow service then receives fewer requests and looks healthier than it is. Closed loop stays available as an sweep of max_in_flight in RFC 0040.7, named as such.

Clamp a run to the ceiling. Friendlier, and the result answers a different question.

Decision

Open. Proposed: the open-loop pacer in the agent with latency from the due time and the stall reported, operations from OpenAPI and protobuf, sent beside counted only where it is measured, ceilings from the plan and the policy with refusal above them, light load only from our cloud, units per run. First proof: a nightly run from the agent on our own server against our public API at our own ceilings, with what the gateway counted beside what was sent, and the same run before and after every change our agents roll.

Publication

The developer docs gain load operations, the report's fields and the ceilings; the agent's policy reference gains [load].

Status log

  • 2026-10-03: opened, from the chaos tool's pacer and the agent's refusal of load today.
  • 2026-10-07: Checked: No PR names this RFC and the status log records nothing built. Open.

This document mentions

Mentioned in

← Back to Platform