Problem
RFC 0033 made every answer of the API follow one contract: a list pages with
page_size and an opaque page_token and links to its next page, every answer carries
x-request-id, every error is the same JSON envelope, and a create sent twice with one
Idempotency-Key answers once. On the day it was decided we proved it by hand, through
the public edge, with curl. Nothing proves it the day after.
The same day showed why that matters. Two sessions rolled the console a few minutes apart, the second from an older commit, and a fix that had shipped was gone from the live page without a single test failing. The unit and integration tests run before a change merges; they say nothing about what the edge serves an hour later. A contract that regresses in production is found by a customer, as a support question, unless something calls the API the way a customer does and keeps calling.
We have most of the parts. Our agent runs on the machine next to production and checks
the public hosts every minute, and every check it declares becomes a monitor in the
console's Reliability area. But an agent check today sends one GET and judges only the
status line: it never reads a header or a body, it cannot carry a value from one answer
into the next request, and it holds no credential that may write. When a declared check
goes down, the event reaches the stream and webhooks, but no person: it is not on the bell
and sends no mail.
Proposal
A new kind of agent check, the contract probe, declared in the agent's checks.toml like
any other check, reconciled into a monitor like any other, and run by the agent against
the public API with an API token (RFC 0016) of an account marked as a probe. It ships in
two steps: the read rules first, with a read-only token; the retried write second, once
the platform can issue a token that writes nothing but a test inbox.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
What one run checks
A run is a fixed sequence of at most twelve requests, each named after the rule it tests, so a failure says which rule broke, not only that something did:
| Invariant | Requests | Passes when |
|---|---|---|
identity |
GET /v1/me |
200; the account is marked probe; the token expires in more than 14 days |
request_id |
every request | every answer carries x-request-id, and every error body repeats it as request_id |
pagination |
the Radar digests with page_size=1, then its next_page_token |
the first page holds one item and a Link: rel="next" header and a /v1/ path; the second page answers 200 |
foreign_token |
the same list with a token from another list | 400 bad_request in the JSON envelope |
clean_refusal |
the digests without a credential; an unknown route under /v1/ |
401 unauthenticated and 404 not_found, both JSON, never plain text |
rate_headers |
the first list call | x-ratelimit-limit and x-ratelimit-remaining present |
idempotent_retry |
create a test inbox with a fresh key, then again with the same key; once with no key; once with the same key and the account id in the body | the second carries Idempotency-Replayed: true, the keyless one does not, and the last is 422 unprocessable |
cleanup |
delete the inbox | 200; a failed delete is a failure of its own |
The list it pages is the Radar's digests: public data, the same for every caller. When
fewer than two digests exist (a fresh environment), pagination and foreign_token
report skipped, not failed. No customer's data is read.
idempotent_retry judges only the replay header. An account holds one live test inbox, so
a second create returns the same inbox with or without a key, and "the same inbox twice"
would pass with idempotency broken. The key is random per run, since a key is remembered
for a day and a fixed one would replay an inbox that has since been deleted. The last
request sends the probe's own account id in the body, which changes the request's
fingerprint, where an empty body and an empty id would not.
Two exceptions to the agent's rules, both bounded
The agent was built on two rules: it never reads a body, and it never holds a credential that writes to the API. A contract cannot be checked without both, so the probe breaks them narrowly and nowhere else:
- It reads bodies, and keeps none. The probe reads at most 64 KiB of an answer, parses
it in memory, checks the fields the invariant names, and drops it. A result carries the
invariant, the request's method and route template, the status, the latency, the
request id and, on failure, the field path and what was expected
(
FAIL pagination GET /v1/radar/digests: next_page_token empty on page 1). Never a body, a header value other than the request id, a token, or the inbox id, which is the inbox's only secret. - It writes, and the platform decides how much. The write step's token carries a new
scope,
webhooks:inbox, which the edge lets through to creating and deleting a test inbox and nothing else: not endpoints, secrets, tests or retries, whichwebhooks:writeallows. The bound holds on the server, so a leaked token can do no more than open and close an inbox on a probe account, which expires by itself within a day. The agent also refuses the write step unlessidentityhas just confirmed a probe account, but that only catches a misconfigured file; it is not the security boundary.
The probe . Its target is pinned to the API in the agent's local
policy, redirects are off, and it follows no URL from an answer except a Link on the same
host. Everything else about the agent stays: its policy decides what it may call, the
token stays in its own secret store and the platform knows only the reference, and a run
has the job deadline of every check (30 seconds). The agent runs one probe at a time per
account, so a scheduled run and the one after a roll never share the inbox.
The tokens
Step one uses a token scoped identity:read and radar:read. Step two adds a second token
scoped identity:read and webhooks:inbox. Both live 30 days, the shortest lifetime that
leaves room for the 14-day warning in identity; the platform's administrators own them,
rotate them when identity warns, and revoke one at once if it is ever seen outside the
agent's secret store.
Probe accounts
The platform gains a flag on an account, probe, which a platform administrator sets
and clears, recorded in the account's audit log like any other change. A probe account:
- answers
probe: trueinGET /v1/me; - holds a plan with units enough for its own calls, about 3,500 a day, so its budget never runs out mid-month;
- is left out of the product's counts of accounts and customers;
- holds no person but its owner, and no objects but the probe's own short-lived ones.
We create one, owned by the platform's administrators, for production. Later checks that need more probe accounts (one with every read scope, one with none, one per environment) build on the same flag.
The monitor, and who hears it
The probe is declared with an interval of five minutes and fail_after = 2: one failed
run is noise, two in a row are worth a person's attention. Reconcile turns it into a
declared monitor in the console's Reliability area, managed by the agent, with each run's
invariant results.
Two changes make a failure reach a person:
- a monitor run records which invariant failed and the field path, so the console shows
paginationfailing, not onlydown; - a declared check may set
notify = true. Itsagent.check.down,agent.check.recoveredandagent.guard.openevents then ring the bell and send mail to the owners and administrators of the account the agent belongs to, the way the health alerts of RFC 0038 do. The default is off, so no customer starts receiving mail for checks they declared before this existed.
After every roll
A roll that changes the protocol, or a service behind the public API asks the agent to run the probe once at the end and prints the result. A failing probe does not undo the roll; it tells the person who rolled, at the moment they can still act.
In what order
- The platform: the
probeflag and its administration,probeinGET /v1/me, the contract probe as a check kind in the agent protocol, invariant results on monitor runs, andnotifyon declared checks. Rolled first, so the platform understands a probe before any agent sends one. - The agent: the probe's read invariants, released. An older agent that receives a probe job refuses it as an unknown kind; it never reports it failed.
- The declaration on our own agent, with the read-only token, and the probe account.
- The
webhooks:inboxscope across the edge, the gateway and the scope catalogue; then the agent's write step and its token.
Alternatives considered
Run the probe in CI. The checks would run before a merge, against a server the tests start. That is what the integration tests already do, and they passed on the day the console regressed. The point is to call the edge that customers call, after the roll.
A probe inside the cluster. A job next to the services could call them without a credential at all. It would skip , which is where the JSON envelope for 401, 403 and 404 and the rate-limit headers are produced, and it would not be a customer's view.
An external uptime service. Hosted checkers can assert on status, headers and JSON. They cannot hold our credential without a vendor review, they do not know our contract's invariants, and the result would not land in our own Reliability area, which is the bench the platform's own agents are meant to check their work against.
Keep the agent's rules and skip writes. Step one does exactly that, and it is worth shipping alone. But the rule most likely to break unnoticed, a retried write answered twice, needs a write, and a scope that can only open and close a test inbox costs little.
Use webhooks:write for the write step. It exists today, but it also creates endpoints
that deliver events to any address. A leaked probe token would then be a way to send our
events somewhere else. The narrow scope is the price of the exception.
Decision
Open. Proposed: the contract probe as a check kind of our own agent, in the order above:
the probe flag, invariant results on monitor runs and opt-in notification first, the
read invariants next, the narrow inbox scope and the retried write last. Built for our own
platform first; the same check kind is offered to customers later, pointed at their own
API and their own contract.
Publication
The console's Reliability area lists the probe's monitor with its last run per invariant. The developer docs gain nothing until customers can declare a probe of their own. When the probe has run for a week, its record (runs, failures, what each failure was) is added to this page.
Status log
- 2026-10-04: opened. Proposed by the owner after RFC 0033 was decided: the contract is checked continuously, by our own agent, and a monitor watches it. Reviewed before publication; the review moved the write step's bound from the agent to a server-side scope, made notification opt-in per check, and fixed what the retried write judges.
- 2026-10-04: step 1 built, the platform side. Accounts take a
probeflag that a platform administrator sets, audited, shown inGET /v1/meand left out of the platform's counts. The agent protocol accepts acontractcheck whose result carries each rule it judged; a monitor run keeps the rule that broke and where; a declared check may setnotify, and its alerts then ring the bell and are mailed to the account's owners and administrators. An agent that does not know a contract probe yet refuses it, and its monitor pauses instead of failing. The probe itself is step 2, in the agent. - 2026-10-06: step 1 live (#254, #266). The accounts
probeflag, thecontractcheck and its rules in run results, and opt-in alerts are running; no account is a probe yet and no agent runs a contract probe until step 2. - 2026-10-07: Checked: step 1 (#254, #266) is live; no account is a probe yet. The method suite (#537) is merged, not rolled; the probe account task is not verified here. Open.
- 2026-10-07: the probe account, built (#551). One command makes it, idempotently: a team
account owned by the platform's administrators, marked as a probe, on a plan whose units
are counted and never charged. Billing refuses a probe account a payment customer and
leaves it out of the overage report, and it is left out of the price preview. Its API
token holds exactly the scopes the method suite of RFC 0040.2 calls with, plus
probe:read, the scope that reads the suite and that only a probe account's keys may hold; it lives 30 days, is written only to the agent's local secret file, and is made again when it has less than 14 days left. Not yet run against production.