Problem
Every check in this family talks to the API. People do not. They open the console, are sent to the sign-in pages, come back with a session, click through pages that call the API from the browser, and leave if a button does nothing.
Most of what breaks there is invisible to an API check: a sign-in script that fails to load, a redirect that loops, a cookie set for the wrong domain after a gateway change, a content security policy that blocks the page's own script, a page that throws on the first click. Every surface check stays green, because every route still answers.
We test these by hand after a change that touches them. An autonomous change cannot look at the console; the bench can.
Proposal
A browser on the agent
A new capability, browser, off until the policy names it. The agent embeds no browser:
it starts a pinned headless Chromium as a separate process, drives it over the DevTools
protocol, and ends it after every journey:
- Bounded. One journey at a time per agent, each in its own fresh profile, under a memory and CPU ceiling the policy sets, with a time limit per step and per journey (proposed: 30 seconds a step, 3 minutes a journey). A journey that passes a limit is stopped and fails with the step it was on.
- Sandboxed. Chromium runs with its own sandbox. Where the machine cannot provide what that sandbox needs, the agent refuses to start the browser rather than drop the sandbox, unless the policy names a stronger boundary it runs inside.
- Under the policy. Every request the page makes goes through the agent's own
resolver and the same
networksanddomainsrules as a check (RFC 0029). A page that tries to load something from a private address, or from a host outside the bound domains and the policy'sbrowser.allow_hosts(fonts, a CDN), is refused, and the refusal is part of the result.
[work]
browser = true
[browser]
max_memory = "1GiB"
max_cpu = 1.0
allow_hosts = ["fonts.example.com"]
screenshots = "on_failure"
Journeys as steps
A journey lives in checks.toml (RFC 0040.1) and becomes a managed monitor. Steps are
few and typed; there is no script.
[[journey]]
name = "console-sign-in"
every = "15m"
account = "probe-a"
[[journey.step]]
open = "https://console.example.com/"
[[journey.step]]
fill = { label = "Email", secret = "env:PROBE_EMAIL" }
[[journey.step]]
fill = { label = "Password", secret = "env:PROBE_PASSWORD" }
[[journey.step]]
click = { role = "button", name = "Sign in" }
[[journey.step]]
totp = { label = "Authentication code", secret = "env:PROBE_TOTP_SEED" }
[[journey.step]]
expect = { role = "heading", name = "Overview" }
wait = "network_idle"
| Step | Does |
|---|---|
open |
navigates to a URL inside the bound domains |
click |
the element with that accessible role and name, or label |
fill |
a field by its label, from a literal test value or a secret reference |
totp |
a field by its label, with a code computed on the agent from a seed held as a secret reference |
expect |
a role and name, or text, present (or absent = true) within the step's limit |
wait |
network_idle, or a fixed time up to 5 seconds |
Elements are found by accessible role, name and label, as a person or a screen reader finds them, not by CSS selectors. A journey that breaks because a button lost its label has found an accessibility bug, which is a bug. Secrets are resolved when the step runs and typed into the field; they never appear in the result, the log or a screenshot.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
What a result keeps
Per step: its name, its outcome, its time, the status of every document and script
request the step caused (counts by class, not URLs with their query strings), console
errors counted by kind. Per journey: the total time and two of the page's vitals as
measured in the lab, Largest Contentful Paint and Cumulative Layout Shift, plus the time
from each click to the next paint. Google's thresholds for these
(Web Vitals) are for field data at the 75th percentile
of real visits; one synthetic visit is not that, so each journey is judged against its own
history and its declared limits, and the console says "lab measurement" next to them.
Screenshots
A screenshot can show anything the page shows, including a customer's data. So:
- only on failure, of the step that failed;
- only when the policy says
screenshots = "on_failure"(the default isnever); - kept on the agent by default; sent to the platform only when the policy also says
upload_screenshots = true; - kept seven days wherever they are, then deleted;
- never taken of a page outside the bound domains.
Our journeys run as probe accounts (RFC 0040.11), so a screenshot should show probe data. "Should" is why the list exists.
Our journeys
| Journey | Steps |
|---|---|
console-sign-in |
open the console, sign in through our identity pages with password and second factor, land on the overview |
api-token |
from a signed-in console, create an API token, see it once, revoke it, see it listed as revoked |
docs-try-it |
open the API reference, make a test token with the credential bar (RFC 0016), send GET /v1/me, expect a 200 in the playground |
The writes these make (a token created and revoked) are a probe account's own, in every
environment. In production, and in our staging environment, which is today the same
deployment as production and follows its rules, nothing else is written.
How an agent uses it
A journey run is a run (RFC 0040.8) started and read through the API, MCP and iohr. An
autonomous change to the console, the docs, the sign-in pages or the gateway's cookie and
header rules runs the journeys before a merge and after it rolls. Later, customers
declare journeys on their own sites, with their own probe accounts.
Alternatives considered
Playwright or Puppeteer inside the agent. The best tools for writing journeys, and a Node runtime and its dependency tree inside a program built to stay small. The step format is narrow enough to drive Chromium directly; journeys written for Playwright can be translated to it.
Journeys from our cloud. No agent needed for public pages. Our cloud shares the platform's fate today (RFC 0040.9), and a browser there would compete with the platform it is timing.
Screenshots of every step. Easier debugging, and an archive of whatever the pages showed.
Decision
Open. Proposed: a browser capability that drives a pinned, sandboxed, bounded Chromium
under the agent's network policy, journeys as typed steps by role and label in
checks.toml, secrets and TOTP from references, per-step timings and lab vitals,
screenshots only on failure, opt in and kept seven days, and journeys run only as probe
accounts. First proof: the Engineering team's agent, bound to our domain, runs
console-sign-in, api-token and docs-try-it against production every 15 minutes.
Publication
The agent's policy reference gains [browser]; the developer docs gain the journey steps
and what a result keeps; the console shows journeys with their steps, timings and, where
allowed, the failing screenshot.
Status log
- 2026-10-03: opened, because sign-in, the console and the docs are checked by hand after every change that touches them.
- 2026-10-07: Checked: nothing built (#271 only mentions it). Open.