Problem
"What did I break" needs more than a red check. An autonomous agent that broke something has to be told what rule broke, on which call, after which steps, and it has to be able to run those steps again after its fix to see whether the rule holds. A person reviewing the change needs the same thing, short enough to read.
Our stress campaigns do this, for one service. Workers drive the ledger, a model of what
it should answer judges every reply, and a broken rule becomes a finding: the invariant,
what the model expected, what the ledger answered, and the trace of steps on that
subject. A signature over the invariant and the message, with ids and stamps blanked,
groups the same bug found twice. After the run each finding is shrunk by delta debugging
(ddmin) to the shortest trace that still breaks the same rule, with the steps a kept step
depends on (a cursor, a recorded_at) pinned back in. chaos stress replay runs a trace
against another ledger and records whether it reproduced.
Two things do not carry over. The model is the ledger's; every other service has none. And a trace step keeps the request and the response as JSON, which is fine for synthetic ledger data and wrong for anything holding real people's data.
RFC 0051 checks our own API's contract (paging, request ids, JSON errors, retried writes)
from our agent, and its platform side is live since 2026-10-06: a contract check whose
result names each rule it judged, kept on the monitor's run. That says a rule broke on
this call. It keeps no steps, groups nothing and replays nothing; this child does that
with the same rules.
Proposal
Contract invariants first
A generic model of an arbitrary service does not exist. What exists is the contract, and a set of rules any service with one can be held to:
| Invariant | Holds when |
|---|---|
status_declared |
every answer's status or gRPC code is one the contract declares for that operation |
shape |
every answer matches the declared schema (RFC 0040.2) |
no_server_error |
a valid request never draws a 5xx, INTERNAL or UNKNOWN outside a fault window |
clean_refusal |
a request made invalid from the schema (a missing required field, a wrong type, a value out of range) draws a 4xx or INVALID_ARGUMENT, never a 5xx and never success |
idempotent_retry |
a write retried with the same Idempotency-Key answers the same status and the same resource id; the same key with a different body is refused |
pagination |
walking a list by its cursor returns no item twice and, with no writes in between, the same set as one larger page |
refuses_unknown_caller |
every operation refuses a call without a credential |
read_your_write |
after a create that succeeded, reading the returned id succeeds |
idempotent_retry follows the IETF draft our own writes honour (RFC 0033,
draft-ietf-httpapi-idempotency-key-header).
pagination and read_your_write run only where the contract names the cursor and the
read. Writes run only when the scenario says so, and never in production.
These rules are shallow on purpose. They find what an autonomous change most often breaks: a status nobody declared, a field renamed, a retry that writes twice, a cursor that skips. A model of a service's meaning, as the ledger has, is the service owner's to write later; this child does not promise one.
What a finding keeps
| Kept | Not kept |
|---|---|
| the invariant, a message, the operation, the status or code | request and response values |
| expected and actual as shapes: field names, types, lengths, status | the content of any field |
| each step: operation, which parameters were present, status, timing | headers other than the status |
| generated inputs as the seed and the step index | |
| ids from earlier answers as references ("the id from step 3") | the ids themselves |
Generated inputs are reproduced from the seed, so they need not be stored; values that came from the service are stored as symbols, as the ledger's traces already are. The cost is stated: a bug that depends on one particular stored value may not replay, and the finding says "not reproduced on a fresh replay" rather than pretend.
Signature and grouping
The signature is the invariant, the operation and the message with ids, stamps and numbers blanked, as today. A finding seen again joins its group: first seen, last seen, how many times, in which runs.
Shrinking
The ddmin of the stress crate, unchanged in principle: replay candidates on fresh inputs, keep the shortest that breaks the same rule with the same signature, pin dependencies back in. Every replay is a request against the target, so shrinking is bounded by attempts and time, counts against the run's units and the load ceilings, and runs in the environment the finding came from.
Replay against another environment
A replay names a finding and a target of another environment with the same contract,
runs the trace attempts times (a race rarely recurs on the first try), and appends to
the finding: when, against which connection, reproduced or not, steps run. A finding is
closed only when a replay in the environment that found it does not reproduce.
How an agent uses it
Findings are created by runs (RFC 0040.8) and read through the API, the MCP tools
reliability_list_findings and reliability_replay_finding, and iohr reliability findings and replay. The loop an autonomous agent follows: run, read the shrunk
finding, fix, replay, and attach the replay's "not reproduced" to its pull request.
Later, customers' findings work the same way against their own contracts.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Alternatives considered
Keep bodies, encrypted. Easier replays and a store of other people's data we would have to protect, export and erase. Shapes and seeds are enough for most of what breaks.
Schemathesis. It generates requests from an OpenAPI document, including sequences through its links, and reports a failure with a request that reproduces it, body and all (Schemathesis). Closest to this; it keeps the body, knows nothing of gRPC, and has no notion of replaying in another environment under an agent's policy.
A model per service before anything ships. Thorough, and nothing would ship. Contract invariants come first; the ledger shows what a model adds.
Decision
Open. Proposed: the eight contract invariants, findings that keep shapes, seeds and
references and never values, signature grouping, the stress crate's ddmin with
dependency pinning, replay against another environment with the outcome recorded, and a
finding closed only by a replay that does not reproduce. First proof: the invariants run
nightly from the agent on our own server against our public API, writes only in probe accounts (our staging is today the same deployment as
production),
with idempotent_retry on every route that honours Idempotency-Key and pagination
on every list, and each finding replayed against a local stack.
Publication
The developer docs gain the invariants, what a finding keeps and the replay; the console gains Findings with groups and replays.
Status log
- 2026-10-03: opened, from the stress crate's findings, shrinking and replay, model-checked today only for our ledger.
- 2026-10-06: RFC 0051's contract check, live on the platform side, is named in the problem: it judges the rules per call; findings, shrinking and replay stay here and are not built.
- 2026-10-07: Checked: no PR builds findings, shrinking or replay. Open.