Problem
The worst bug an autonomous change can ship here is not a slow route or a broken one. It is a route that answers one account with another account's data. It passes every test that uses one account, every surface check, every load run. It is the first item on OWASP's API list for a reason (API1 in the 2023 edition, broken object level authorization).
Our design makes it unlikely. The gateway is the only place a token is verified; each route names the scope it needs; services take the verified caller and scope every query to its account. Unlikely is not checked. Today nothing asks, on every route, whether a token of account A can read account B's objects. The surface checks of RFC 0040.2 ask whether a call without a token is refused; that is the easy half.
The other guards we rely on have the same gap: a token without the right scope, a revoked token, an expired one, a rate limit, the origins CORS accepts, the security headers, the TLS versions offered, and the refusal of private addresses on every route that fetches a URL. Each was tested when it was built. None is tested after each change.
Proposal
Two probe accounts
An account can be flagged probe. The flag exists since RFC 0051 (set by a platform administrator, audited, left out of the platform's counts); no account carries it yet. A probe account is an ordinary account that the bench owns. It holds a fixed set of fixture objects of every kind the public API serves (a token, a connection, a webhook endpoint, a monitor, and so on), made once and kept. It is left out of usage summaries and billing, cannot be invited into, and is listed on the environment's page so nobody mistakes it for a customer.
Two probe accounts per environment, A and B. Each has tokens of its own: one with every read scope, one with none, one revoked, one expired. The tokens are secret references on the agent, resolved at the moment of the call like any other (RFC 0029).
The probes
Every operation in the public OpenAPI document is probed. The document says which operations take an object id and which scope each operation's security requirement names, so the probes are generated, not written per route:
| Probe | Request | Must answer | Never |
|---|---|---|---|
cross_account_read |
A's token, B's object id in the path | 404, or 403 | 200 |
cross_account_list |
A's token on a list | no id of B's in the page | B's id present |
cross_account_write |
A's token, B's object id, a write | 404 or 403 | 2xx |
missing_scope |
a token without the operation's scope | 403 | 2xx |
revoked_token |
a revoked token | 401 within the window | 2xx after it |
expired_token |
an expired token | 401 | 2xx |
rate_limit |
calls past a route's limit, under the load ceilings | 429 with Retry-After |
2xx past the limit, or 429 without when |
cors |
a preflight from an origin not allowed | no Access-Control-Allow-Origin for it |
the origin echoed back |
headers |
any answer | the security headers the edge promises | one missing |
tls |
a handshake per protocol version | TLS 1.2 and 1.3 only, no weak ciphers | TLS 1.0 or 1.1 accepted |
ssrf |
a private address, a link-local or metadata address, a redirect to one, on every route that fetches a URL | refused before any connection | fetched |
404 is preferred over 403 for another account's object, because 403 says the object
exists. Both pass; 403 is noted. A probe that finds a guard open goes down with the class
guard_open, as the refusal checks of RFC 0040.1 do, and is never part of an uptime
figure.
The revocation window is the one RFC 0016 states and measured: a revoked token refused
within one reporting interval of the gateway, a few seconds (2.3 seconds on 2026-10-02).
The probe revokes a fresh token of a probe account, calls once a second, and reports the
time to the first 401; past ten seconds is a failure. The rate-limit probe follows the
headers RFC 0033 promises: the limit, what remains, when the window resets, and
Retry-After on the 429.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Coverage is part of the answer
Generated probes cover only what the document describes and the fixtures provide. Each
run says so: how many operations were probed, and which were not and why ("no fixture of
this kind in account B", "no scope named in the contract"). An operation with no scope in
its security requirement is itself reported, because the contract should say it. Coverage
that drops between two runs is a regression like any other (RFC 0040.8 compare).
Findings, never bodies
Each failure is a finding (RFC 0040.6): the operation, the probe, the expected and actual status, the probe accounts involved, the time. Not the body that came back, even when the body is the leak. A body from a cross-account read could hold the probe account's fixture and nothing else, but a probe that keeps bodies would one day keep a real one. The finding says "200 with a body of 1.2 KiB", and the person fixing it reproduces it with the probe tokens.
Where it runs
The probes run in every environment, production included, because that is where a leak
matters. Our staging environment is today the same deployment as production, and an
environment that shares a deployment with production follows production's rules. In
every environment the probes touch only probe accounts: every read asks for a probe account's
object, and every write, cross_account_write included, names an object of probe account
B with a token of probe account A. If that guard is open, the worst case is a changed
fixture of a probe account; the run records which, and the fixtures are restored from
their declared state before the next run. No write ever names an object of an account
that is not marked probe, and the platform refuses a probe token's call that would.
Everything runs through the agent or from our cloud against our own verified domain,
under the load ceilings of RFC 0040.3; the rate-limit probe is the only one that sends
more than a few calls a route.
What it is not
This is not a penetration test. It does not look for injection, does not fuzz parameters (RFC 0040.6 generates invalid input, within the contract), does not test the sign-in pages' own defences, does not chain bugs, and does not try what the document does not describe: an undocumented route is invisible to it. It checks, after every change, that the guards we designed are still where we put them. A penetration test by people who did not build the platform is a separate thing, and this does not replace it.
How an agent uses it
The probes are a run kind, guards, started and read through the API, MCP and iohr
like any run (RFC 0040.8), with reliability:run and reliability:read. An autonomous
change that adds or changes a route runs them before it asks for a merge; a new
guard_open blocks its verdict (RFC 0040.14).
Later, customers name their own probe accounts and get the same probes on their own API.
Alternatives considered
Write an isolation test per route in each service. We do, in places. Tests run against a service with its own fixtures, not through the gateway and the deployed edge, which is where a scope or a route rule goes wrong.
Use real accounts with consent. No fixtures to keep, and the bench would then read real data whenever a guard broke. Probe accounts make the worst case a leak of test data.
A commercial API security scanner. Good at finding classes of bugs in an unknown API. Our question is narrower and must be answered after every change, against our own contract, with no bodies leaving the run.
Decision
Open. Proposed: a probe flag on accounts, two probe accounts per environment with fixtures and four tokens each, probes generated from the OpenAPI document for cross-account reads, lists and writes, scopes, revocation, expiry, rate limits, CORS, headers, TLS and SSRF, coverage reported with every run, findings with statuses and never bodies, and every read and write confined to probe accounts in every environment. First proof: the Engineering team's agent, bound to our domain, probes every public route of our API in production nightly with the two probe accounts, and the revocation probe measures the window against RFC 0016's.
Publication
The developer docs gain the probe list and what each one means; the console marks probe accounts and shows coverage per run; RFC 0016's status log gains the window as measured by the probe.
Status log
- 2026-10-03: opened, because tenant isolation and scopes are designed in and tested once, not after every change.
- 2026-10-06: the
probeflag this RFC asked for was built and is live with RFC 0051; the text says so. The rest is not built. - 2026-10-07: Checked: the
probeflag is live (#254, RFC 0051); nothing else is built. Open.