This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0040.14open2026-10-03

The change verdict, for people and AI assistants

One answer per change (pass, fail, regressed or incomplete), computed by rules over the per-claim verdict of RFC 0047 together with the repository's own CI result, the bench runs before and after, new findings and the guard checks, posted on the pull request with its evidence; required of our autonomous agents before they ask for a merge and after they roll; and the bench offered as MCP tools on the server RFC 0050 opened to every assistant, with thin plugins kept in a separate public repository that publishes nothing until the owner has read a review of each assistant vendor's terms.

Part of RFC 0040 Chaos as a service, our test bench in the customer's network

Problem

The rest of this family produces evidence: surface checks, refusals, load, faults, findings, capacity, guards, consistency, journeys. A person reviewing a pull request will not open nine pages, and an autonomous agent should not decide for itself which of them matter. Both need one answer per change, with the evidence one click away.

The same evidence has to reach the place the work happens. Our autonomous agents, and the people using AI assistants in their editors, ask "did my change break anything" in a chat, not in a console. Since 2026-10-04 chaos verify (RFC 0047) measures a pull request's claims before and after it rolls and posts a verdict per claim and per dimension on the pull request; that covers our own repository and the checks it runs, not load, faults or findings, and no assistant can ask the bench for any of it.

Proposal

The verdict

A verdict belongs to a change: a pull request (core#212) or a roll (a service and the version it moved to). It is computed from four inputs, each linked:

Input From Counts as failing when
The repository's tests the CI result for the change's commit, attached by iohr reliability verdict attach from the CI job CI failed
Bench runs a run of the same scenario before and after the change, compared (RFC 0040.8 compare) regressed on any measure
Findings findings created by the after-run that the before-run did not have (RFC 0040.6) any new finding
Guards refusal checks, surface refusals and the probes of RFC 0040.11, in the after-run any guard_open

The verdict is one of three:

  • pass: CI passed, no regression, no new finding, every guard closed.
  • regressed: CI passed and the guards are closed, but compare found a measurable regression or the after-run has a new finding. A person decides.
  • fail: CI failed, or a guard is open. Nobody merges on it.

This verdict does not replace the one RFC 0047 built. That one judges each claim and each dimension (pass, fail, not applicable, not measured) and stays on the pull request beside this one, in full: RFC 0046 rules out a single score that hides a failed dimension. The answer here is a rule over those results and the four inputs above, and every line that decided it is shown with it.

A missing input is not a pass. A change with no before-run, or a CI result nobody attached, gets the verdict incomplete with the missing input named. The verdict is computed by rules, not by a model, and stored with the hashes of the scenarios it used.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

On the pull request

The verdict is posted as a check on the pull request, through a GitHub connection of the account (RFC 0018) whose app holds permission to write checks and nothing else. The check's summary is the verdict and one line per difference, as compare writes it, with links to the runs and findings. A branch protection rule can require it like any other check.

Our autonomous agents need it first

This is the first use, and the reason for the order of this family. An autonomous agent working on this platform:

  1. starts the bench's before-run when it opens its branch;
  2. attaches its CI result and starts the after-run when its change is ready;
  3. asks for a merge only with a pass, or with a regressed and its reasoning written on the pull request for a person to judge;
  4. after the change rolls, runs the bench again against what is now live and posts a second verdict on the same pull request; a fail there makes it open a revert pull request and tell a person at once.

Our staging environment is today the same deployment as production, and follows production's rules: the runs a verdict needs are checks, guards, consistency, journeys and bounded load, and any fault in them waits for a person (RFC 0040.4).

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

The bench as MCP tools

The bench is offered on our existing MCP server (RFC 0005), behind the same gate, as the caller, with the same audit line. RFC 0050 made that server an OAuth protected resource any assistant can connect to on a person's behalf, with consent on our own sign-in, tokens bound to the server and scopes per tool; it has been live since 2026-10-05, with the RFC tools as its first product. The reliability tools join it under their own scopes:

Tool Scope Does
reliability_list_targets reliability:read the account's environments, connections and scenarios
reliability_get_run, reliability_wait_run reliability:read a run's record; wait bounded
reliability_compare_runs reliability:read the comparison of two runs
reliability_list_findings, reliability_explain_finding reliability:read findings, and one finding's invariant, shrunk steps and replays in prose built from the record
reliability_get_verdict reliability:read the verdict of a change, with its inputs
reliability_start_run reliability:run starts a run of a declared scenario; never a fault run

Read-only by default: a token with only reliability:read sees no tool that starts anything. reliability_start_run is the one tool that triggers work, and it refuses the faults kind outright, so RFC 0005's rule that no MCP tool injects a fault holds. No tool confirms anything.

reliability_explain_finding assembles its text from the finding's record, not from a model. An assistant may add its own words around it; those are the assistant's.

Plugins for the assistants people use

Each plugin is a thin wrapper over the same MCP server, holding no logic of its own. It signs the person in through RFC 0050's flow, where the assistant supports it, and asks for reliability:read by default; where it does not, the person makes an API token (RFC 0016) with the scopes they choose and the assistant keeps it where it keeps secrets.

Assistant Package
Claude Code a plugin with the MCP server and a command that asks for the verdict of the current branch (plugins)
Claude Desktop the remote MCP server, or a desktop extension that configures it (MCP on Claude Desktop)
ChatGPT an app on the MCP server (Apps SDK)
Gemini CLI an extension naming the MCP server (MCP servers)
GitHub Copilot the MCP server for the editor and the coding agent (MCP for Copilot)
Cursor an mcp.json entry (Cursor MCP)

Revoking the app or the token in the console stops the plugin: an API token at once, an app's access token at the end of its short lifetime (RFC 0050).

The owner decided on 2026-10-06 where these plugins live and when they may be published:

  • A new, separate repository for customers. The customer plugins and their marketplace listing live in a repository of their own, meant to be public and shared with customers. It is not our internal kit: the private repository of plugins, guards, team rules, agents and skills that our own sessions use stays private and is never shipped to a customer, in whole or in part.
  • Nothing internal in it. That repository holds the manifests, the commands and their docs. No internal guard, team rule, private RFC, prompt or other private material is copied into it; a plugin that needs one of those is not a customer plugin.
  • No vendor's mark in a name. Neither the repository nor any plugin is named or drawn with "Claude", "Anthropic" or another assistant vendor's mark. The table above names the assistants a plugin works with; that is a description, not a name.
  • Each vendor's terms are reviewed first. Before a plugin for an assistant is named or published, a due-diligence review of that vendor's terms is written and read by the owner: plugin and marketplace distribution, trademark and branding, usage policy, and commercial terms. A vendor whose terms do not allow it is left out.
  • Private until the owner says so. The repository is created private and becomes public only when the owner has read the review and says so. Until then nothing is listed in a marketplace or submitted to a vendor's directory.

The review for Claude Code is done, and on 2026-10-06 the owner decided the first plugin's names and rules:

Name
Repository inorbithr/inorbit-plugins
Marketplace inorbit-plugins, different from the name of our private internal marketplace
Plugin inorbit, display name "InOrbit"
Description, and the README's opening line "InOrbit plugin for Claude Code"

The identifiers and the display name stay free of any mark; the description names the assistant in plain prose, which Anthropic's documentation allows. No identifier and no logo carries claude or anthropic. The command line refuses a plugin name that starts with claude-, anthropic- or cc-plugin-, and Anthropic's legal terms forbid its marks in another product's name.

  • The customer brings their own plan. We never resell or route Claude usage, never touch a person's Claude credentials and never bundle Claude Code. The plugin and its docs give no advice on sharing accounts or getting around usage limits.
  • The README says what it is. Directly under its opening line it carries "Claude and Claude Code are trademarks of Anthropic PBC. InOrbit is not affiliated with or endorsed by Anthropic." It lists exactly what data reaches InOrbit, and for each tool whether it only reads or can change or delete, and the scopes it needs. It says "aligned", never "compliant".
  • Our models' words are labelled. Anything a model of ours writes through the plugin is marked as AI output (AI Act Article 50, below).

A listing in Anthropic's directory on claude.com is optional and comes later. It asks for separate read and write tools with their annotations, OAuth 2.0 and Streamable HTTP (RFC 0050 serves both), a privacy policy, a test account, three example prompts, no tool that moves money, and an indemnity from us to Anthropic; four questions go to Anthropic before we apply, the first asking it to confirm the wording "InOrbit plugin for Claude Code". Plugins for the other assistants follow the same pattern, each after its own vendor's review.

After the review and the owner's word, plugins are installed from that public marketplace and, where each vendor's process allows, from the vendor's own directory. AWS Marketplace is noted only as a possible later channel; nothing is planned for it.

What an assistant may never do

Through any plugin, tool or token: confirm a fault in production or in an environment that shares its deployment; read a secret or a secret reference's value; change an agent's policy or widen a ceiling; close a finding without a replay; change a verdict. None of these exists as a tool, so none can be talked into.

Telling people it is AI

Where an assistant speaks on the bench's behalf, such as a pull request comment an autonomous agent writes around a verdict or a summary in a chat, the text says it was written by an AI system, as Article 50 of the AI Act requires since 2026-08-02 (Regulation (EU) 2024/1689). The verdict itself is computed, not written by a model, and is labelled as computed.

Later, customers use the same verdict on their own repositories and the same plugins on their own accounts.

Alternatives considered

Let each input post its own check. Nine checks on a pull request, and every reader deciding which ones matter. One verdict, with the evidence linked, is what gets read.

A model writes the verdict. Better prose, and an answer that cannot be reproduced from its inputs. Rules decide; a model may explain.

A separate plugin per assistant with its own logic. Six code bases that drift. One MCP server, six manifests.

Decision

Open. Proposed: a verdict of pass, regressed, fail or incomplete from CI, compared runs, new findings and guards, computed by rules and posted as a pull request check; our autonomous agents required to have it before a merge and after a roll; the bench as MCP tools, read-only by default, one tool to start a non-fault run; thin plugins for the assistants above, kept in a separate public repository with no internal material and no vendor's mark in its names, and published only after the owner has read the review of each vendor's terms; a fixed list of what an assistant may never do; AI-written text labelled. First proof: every pull request our own autonomous agents open in this repository carries the verdict, computed from runs of the Engineering team's agent, bound to our domain, against production, beside RFC 0047's per-claim verdict, and our own assistants read it through the MCP tools.

Publication

The developer docs gain the verdict, its inputs and iohr reliability verdict; the MCP reference gains the tools; each plugin gets an install page once it is published; RFC 0005's status log records the reliability tools and the one that starts work.

Status log

  • 2026-10-03: opened, so the bench's answer reaches the pull request and the assistants where changes are made.
  • 2026-10-06: brought up to date with what is built. The verdict sits beside RFC 0047's per-claim verdict (chaos verify, started 2026-10-04) and never replaces it; the MCP tools join the server RFC 0050 opened to every assistant, and plugins sign in through it. The owner's decisions on the plugins: a new, separate public repository shared with customers, distinct from our private internal kit, with no internal material in it; no "Claude", "Anthropic" or other vendor's mark in any name or logo; each vendor's terms reviewed and read by the owner before a plugin is named or published; the repository created private and made public only on the owner's word; AWS Marketplace only as a possible later channel. The Claude Code review is done: inorbithr/inorbit-plugins, marketplace inorbit-plugins, plugin inorbit shown as "InOrbit", description and README opening line "InOrbit plugin for Claude Code" with the trademark line under it, the repository private until the owner's word, with its rules (own plan, no resale, the README's disclaimer and data list, AI output labelled) and the directory listing as an optional later step.
  • 2026-10-06: the Claude Code plugin built, in the private plugins repository: the marketplace, the plugin with the MCP server entry (sign-in through Claude Code's own flow, an API token as the alternative), and three commands. One reads RFCs, one reads the platform's status, and one reads the verify verdict already posted on the branch's pull request, saying plainly that this RFC's single verdict is not built yet. The README lists every tool with its scopes and whether it writes, and what reaches InOrbit. Checked by the plugin validator, an install into a clean configuration and the live server's tool list; the first browser sign-in is the owner's to try. Still private.
  • 2026-10-07: the Claude Code plugin in the private plugins repository moved to the session that publishes the plugin and the MCP server to each assistant's directory; plugins for the other assistants stay with the session that built the MCP tools. Nothing changed in what is built: the repository is still private, and the first browser sign-in is still the owner's to try.
  • 2026-10-07: Checked: the Claude Code plugin is in the private plugins repository (#423, #440); the verdict itself is not built. Open.

← Back to Platform