Problem
Our API serves one registry of RPCs over several transports: REST and server-sent events through the gateway's transcoding, gRPC, GraphQL, calls on the multiplexed WebSocket, MCP, and the events of RFC 0022 as webhooks, MQTT and a stream. One contract, rendered several ways. The renderings are generated, so they should agree. "Should" is the word this family exists to remove.
Each rendering has its own code where it can go wrong: how a 64-bit integer becomes a string, how an empty list or an unset field is written, how an enum's zero value is named, which fields GraphQL resolves lazily, how a page token is carried. An autonomous change to the transcoder, a proto or one service's mapping can make REST and gRPC answer differently for the same object, and every per-surface check (RFC 0040.2) still passes, because each answer matches its own schema.
The data has the same gap. A create that succeeded should be readable at once on every
surface that serves the object. A list's total_size should be the number of items its
pages hold. A write should produce exactly the events it promises, with its ids. A retry
with the same Idempotency-Key sent over REST and then over the WebSocket should still
write once. RFC 0040.6 checks read_your_write and idempotent_retry on one surface. None
of this is checked across them.
Proposal
What is compared
| Check | Holds when |
|---|---|
same_read |
one read, made as the same caller on every surface that serves it, gives the same normalised answer |
read_your_write |
after a write on one surface, the object reads back with the written values on every surface within the window |
list_total |
walking a list to its end yields exactly total_size items, where the route promises a total, and agrees with any summary that counts the same objects |
events_for_write |
each write produces the events the catalogue promises for it, carrying exactly its ids, within the window, and no others |
cross_surface_retry |
a write retried with the same Idempotency-Key on another surface answers as the first did, marked replayed, and writes once |
The surfaces are the ones the platform already serves: REST, gRPC, GraphQL, the WebSocket mux, MCP where the operation is a tool, and the event stream. A surface that does not serve an operation is skipped for it and counted, not failed.
Normalised, in memory, then dropped
A comparison has to read bodies, which rule 5 of the parent forbids keeping. This child states its exception the way RFC 0040.2 does: each answer is read (at most 1 MiB), reduced in memory to a normalised form, compared, and dropped.
The normal form is the shape plus the ids: field paths, types, which fields are present, list lengths, enum names, and the values of fields the contract marks as identifiers or timestamps. The rules that make two correct renderings equal are fixed and published: 64-bit integers compared as numbers whatever their encoding, an absent field equal to its default only where proto3 says so, timestamps compared as instants, enum values by name. Free text, amounts and anything else are compared by type and presence only, never value, so a run cannot hold a customer's content even in memory longer than one comparison.
That makes this contract-level consistency, and we say so. It catches a field one surface drops, a renamed key, a wrong type, an id that differs, a list that is shorter on one surface, a total that is wrong. It does not catch two surfaces that both return the same wrong amount.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Windows
Every check that waits has a stated window, part of the run's report:
read_your_write: 2 seconds by default; a service that promises less (or more) states it in its contract, and the window comes from there.events_for_write: 10 seconds on the stream and over MQTT, 60 seconds for a webhook delivery to the bench's own receiver, which answers at once, so a delivery that needed a retry (RFC 0022) shows up as late rather than hidden.
A miss is reported with the time the answer finally arrived, if it did. "Arrived after 4.1 seconds" is a different bug from "never arrived".
Writes only in probe accounts
Reads run as a probe account (RFC 0040.11) and read only its objects. Every write in this
child is made by a probe account on its own objects, in every environment. That matters
most where a mistake would cost most: in production, and in our staging environment,
which is today the same deployment as production and follows production's rules. A
write check creates its object, checks it, and deletes it; a
failed delete is a finding of its own. The probe accounts' webhook endpoints point at a
receiver the bench runs, which records event ids and drops everything else.
Findings
A broken check is a finding (RFC 0040.6) with the check, the operation, the two surfaces that disagree, and the first difference as a path and two types or two ids, never two values:
FAIL same_read getMonitor rest vs graphql /lastRun/latencyMs: number vs absent
Findings group by signature, and a replay runs the same calls again, so a fix in the transcoder is proved by a replay that agrees.
What it does not do
Deep models of a service's meaning stay out of scope. Our ledger has one, and it checks balances, ordering and history in a way no generic rule can (RFC 0040.6). This child does not try. It checks that the platform tells the same story on every surface and that the counts it reports add up, which is what a change to the shared machinery breaks.
How an agent uses it
consistency is a run kind (RFC 0040.8), started and read through the API, MCP and iohr.
An autonomous change to the gateway, the transcoder, a proto or an event's producer runs
it before and after, and compare lists any check that changed verdict.
Later, customers whose APIs speak more than one protocol get the same checks against their own contract.
Alternatives considered
Snapshot answers and diff them later. Easy to inspect, and a store of bodies. The comparison happens in memory or not at all.
Compare raw bytes. Two correct renderings differ in field order, number encoding and default fields. Byte comparison would fail on every run and be ignored.
Test consistency in the protocol's own test suite. We do, against services booted on free ports. That proves the transcoder on test data; it says nothing about the deployed edge, the event delivery and the windows in production.
Decision
Open. Proposed: five checks across every surface the platform serves, answers normalised to shapes and ids in memory and dropped, published normalisation rules, stated windows with the late arrival time reported, writes only by probe accounts on their own objects, findings that name two surfaces and a path, and no promise of service-level models. First proof: the Engineering team's agent, bound to our domain, runs the five checks nightly in production on every object of our public API that is served on more than one surface, starting with monitors, connections and webhook endpoints.
Publication
The developer docs gain the checks, the normalisation rules and the windows; the protocol's contract page links them; the console shows consistency runs with the surfaces that disagree.
Status log
- 2026-10-03: opened, because one registry rendered over six surfaces is checked surface by surface and never across them.
- 2026-10-07: Checked: No PR names this RFC and the status log records nothing built. Open.