This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0047open2026-10-04

The first complete loop, a change to our own platform proved end to end

The first workflow that runs every step of RFC 0045's loop - an agent's change to this platform, from a machine-checkable claim through bench runs before and after, a verdict per dimension and planted regressions as ground truth, to a signed evidence record - and the acceptance test that decides when we may say it works.

Problem

RFC 0045 asks for one complete, verified workflow before many shallow AI features. It does not say which workflow, and "complete" is easy to claim and hard to check.

The choice matters. Incident response looks like the obvious first workflow, and it is the one every SRE-agent vendor already ships. It is also the hardest to grade on a live system: a real incident's cause is often not known until much later, if ever. A change we make ourselves has none of that problem. We know what it was meant to do, we can measure the system before and after it, and we can plant a regression on purpose and see whether it is caught.

We also already have the actors. Agents change this platform today, and we are its first customer by rule (RFC 0045, rule 7).

Proposal

The workflow

A change to this platform, made by one of our own agents or by a person, goes through every step of the loop:

  1. Understand. The change names the RFC it implements (the Implements: line of RFC 0041), or says it implements none.
  2. Hypothesize. The change states at least one claim a machine can check in the form RFC 0046 defines: what must improve, and what must not change (a contract, a latency bound, an error rate under a given load).
  3. Observe. The bench runs against the current deployment: the surface checks on every protocol, the declared checks of the services the change touches, and their tests that must fail.
  4. Experiment. Where the claim is about behaviour under stress, the bench runs the load or the budgeted fault the claim names, before the change.
  5. Change. The pull request is opened, reviewed and merged by our normal rules. An agent works under its autonomy level: it may open the pull request, it may not merge it.
  6. Verify. After the roll, the same bench runs again. The verdict is one answer per dimension: each claim held or not, nothing that passed before fails now, no new finding, every test that must fail still fails.
  7. Grade. The verifier itself is graded against ground truth we control (below).
  8. Evidence. One signed record per change (RFC 0046), linked from the pull request and from the RFC it implements.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Ground truth: planted regressions

A verifier that never fails is indistinguishable from one that is broken. So we plant regressions: changes made on purpose to break one claim (an added delay on a route, a contract field renamed, an error rate raised, a check silenced) and sent through the same workflow, labelled as planted only in a place the verifier cannot read. Each planted regression is a cause we know. The verifier is graded on whether it caught it, how long it took and whether it blamed the right claim. The planted change never reaches production: it runs against a deployment built for it.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Claims

A change states its claims in its pull request description, in one fenced block marked claims, as TOML with one [[claim]] table per claim:

```claims
[[claim]]
kind = "surface"
check = "http_healthz"

[[claim]]
kind = "contract"
change = "none"
```

Each claim has one kind and the fields that kind needs:

Kind Says Fields
surface a named check of the bench passes (or fails) once the change runs check, expect (pass or fail)
contract the public API changes only as stated change (none or declared), operations, and for declared the added, removed and changed operation ids
latency an operation's percentile stays under a bound at a given rate operation, percentile, under_ms, rps, duration
errors an operation's share of server errors stays under a rate at a given load operation, max_rate, rps, duration
monitor a monitor stays healthy from before the change to after it org, monitor
refuse a check that must fail still fails org, monitor
text anything else, in words text

The block is refused as a whole when a claim is wrong in shape: an unknown kind or field, a missing field, a check that does not exist. A text claim is allowed and listed, never measured and never counted as a pass. A contract counts an operation as changed when its own definition or any schema it refers to changes, so a field renamed in a shared type changes every operation that uses it.

Load and faults run only on a deployment built for the change, never on production. Until that deployment exists, latency and errors claims are reported as not measured, with that reason.

The verdict gives each claim and each dimension one of pass, fail, not applicable or not measured. The dimensions every change is held to are: nothing that passed before fails now, no new finding, and every check that must fail still fails. The whole is a pass only when everything that applies passed and at least one claim was measured.

The order of work

Each step builds on what exists:

  1. Claims in RFC 0046's form, accepted in pull request descriptions and checked for shape.
  2. The bench runs that observe and verify, as runs an agent can start and read through the API, MCP and iohr.
  3. The verdict per change, posted on the pull request.
  4. The evidence record, signed and stored, linked from the pull request and the RFC.
  5. Planted regressions and the verifier's grade.
  6. Agents required to attach a passing verdict before they ask for a merge.

Acceptance test

We may say the loop is complete, for this workflow only, when all of these hold over four consecutive weeks on this platform:

  • every merged change to a service has an evidence record, with at least one machine-checkable claim and a verdict;
  • at least twenty planted regressions have gone through the workflow, covering each kind listed above, and the verifier caught all of them on the right claim;
  • no change was rolled back for a cause its verdict had passed, or each one that was is a published finding with what the verifier missed;
  • every number above is on a public page, measured, with its date.

Until then the site describes this workflow as being built, step by step, as RFC 0045 requires.

Alternatives considered

Incident response first. It is the workflow the market is buying, and it will come. It comes second because its ground truth (injected faults on a live system) needs the same bench, verdicts and records this workflow builds, and because grading an agent's diagnosis is only credible once the verifier itself has been graded.

A customer's system first. We would learn faster about their systems and slower about whether the loop works. Our own platform gives us full access, real changes every day and no one else's production at risk.

Grade only real failures. Real regressions are rare and unlabelled. Planted ones are frequent, known and cheap; real ones are kept as findings beside them.

Decision

Open. The workflow, planted regressions as ground truth and the acceptance test are proposed together with RFC 0045 and RFC 0046.

Publication

The workflow's state (which steps run, how many changes carry a record, the planted regressions and how many were caught) is published on this page's status log as each step lands. Nothing claims the loop is complete before the acceptance test passes.

Status log

  • 2026-10-04: Written, beside RFC 0045 and RFC 0046. Nothing is built for this workflow yet; the bench, checks and monitors it reuses exist in part.
  • 2026-10-04: Steps 1 to 3 started. The claim grammar above, the bench run before and after a change rolls (every check of the bench against the deployment, and the monitors a claim names), the verdict per claim and per dimension, and the verdict posted on the pull request. The record is written in RFC 0046's form but unsigned and marked as a draft. Not built yet: load and faults on a deployment built for the change, signing and storing records, planted regressions, and the rule that agents attach a passing verdict.
  • 2026-10-07: Checked: chaos verify (#251) runs in chaos at e3b6c779. Not built: load and faults on a change's own deployment, signed records, planted regressions and the required verdict. Open.

← Back to Platform