Problem
Writing code has stopped being the hard part of changing a system. Knowing whether the change is right, and being able to show it, has not.
The evidence on this is better than the opinions:
- In METR's randomised trial (July 2025), experienced open-source developers working on their own repositories took 19% longer with AI tools available, while believing they had been 20% faster (METR). In February 2026 METR reported that developers were probably sped up more by then, but that it could no longer measure it cleanly: too many developers refused to work without AI for a control group to exist (METR). The size of the effect is open. The gap between how fast it feels and what was measured is not.
- Google's DORA 2025 report found AI adoption now linked to higher delivery throughput, while 30% of respondents trust AI "a little" (23%) or "not at all" (7%), and named "ensuring software works as intended before it's delivered to users" as the challenge that remains (Google).
- Benchmarks wear out. OpenAI stopped reporting SWE-bench Verified in 2026 after auditing it: in the problems it audited, at least 59.4% had tests that reject correct solutions, and models reproduced reference fixes they had seen in training (OpenAI).
- Operations is harder than coding for these systems. On ITBench-AA (Artificial Analysis with IBM, 2026-05-27), which asks an agent to find the root cause incidents, every frontier model scored below 50%; the best was 47% (Artificial Analysis).
Meanwhile AI agents are being given production access. Datadog's Bits AI SRE has been generally available since 2025-12-02 (Datadog), Azure SRE Agent since 2026-03-10 (Microsoft) and AWS DevOps Agent since 2026-03-31 (AWS). AWS reports that preview customers saw up to 94% root-cause accuracy. That is the vendor's measure of its own agent on its customers' incidents; the independent benchmark above puts the best model below 50% on incidents with a known cause. Both can be true, and a team cannot tell from either which applies to its own system. What no vendor gives it is a record, independent of the agent, of what the agent did and whether it was right.
The obligations point the same way. The EU AI Act's transparency duties under Article 50 apply from 2026-08-02, and the Cyber Resilience Act's 24-hour reporting of actively exploited vulnerabilities from 2026-09-11. Both ask for evidence, not confidence.
Our own platform has the same problem in miniature. Agents already change it. Its pieces (the chaos tool, the sandbox, the agent in a customer's network, monitors, signals, the Lab) each answer part of "is this right?", but nothing joins them, and a reader of this site could reasonably take us for seven unrelated products.
Proposal
The thesis
AI engineering that has to prove its work.
InOrbit is a platform where AI does engineering work on real systems, and every piece of that work ends in evidence a person, an auditor or another system can check without trusting the model that produced it.
The loop
The unit of work is not a prompt or a ticket. It is one pass through a closed loop:
| Step | Question it answers | What exists today |
|---|---|---|
| Understand | What is this system and what is it for? | RFCs and studies in the Lab (RFC 0035), the RFCs product in the console (admin only) |
| Observe | What is it doing right now? | Monitors (RFC 0037), signals (RFC 0038), the agent in the customer's network (RFC 0029), traces, metrics and logs on our own platform |
| Model | How do its parts depend on each other? | Architecture diagrams in the Lab; no model derived from the running system yet |
| Hypothesize | What do we expect to happen if X? | Written by hand in RFCs; nothing machine-checkable yet |
| Experiment | Does it? | The chaos tool on our own platform, load and faults through the agent (RFC 0031), the sandbox for code (RFC 0010) |
| Change | Make the change | Our own agents through the API, MCP and iohr; pull requests |
| Verify | Did the change do what it claimed, and break nothing else? | Partly: checks and monitors; no verdict per change yet |
| Grade | How good was the work, against something known to be true? | Not built for agents on systems; the idea is RFC 0008's grader, for code |
| Evidence | Can someone else check all of this later? | RFCs show their implementation (RFC 0041, open); no evidence record yet (RFC 0046) |
The right-hand column is the honest state on 2026-10-04. Most rows are partial, and the last three are the ones that make this a platform rather than a set of tools. RFC 0046 defines the evidence record; RFC 0047 builds the first loop that runs from start to finish on our own platform.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
What InOrbit is not
Each of these is a workload or a part of the platform, not what it is:
- An AI SRE. Incident response is one workload of the loop (observe, hypothesize, change, verify). Others sell agents; we make any agent's work, ours included, checkable.
- Chaos engineering. Faults are how the loop gets ground truth: a fault we injected is a cause we know, so a diagnosis can be graded against it.
- Observability. Telemetry is an input to the loop. We read it; we do not try to be the place it is stored.
- A coding agent. Writing the change is one step in nine.
- Integrations. Connections (RFC 0018) are how the loop reaches a system.
- Avatars. An avatar (RFC 0017) is a persona that does work in the loop and is held to the same evidence as any other agent.
Rules this sets for everything we build
- Measured evidence over model confidence. A model's own statement that it succeeded is never the evidence. A check, a measurement, a replayed finding or a person's decision is.
- One complete, verified workflow over many shallow AI features. A new AI feature is worth building when it closes a step of the loop for a workflow that runs end to end.
- Known ground truth before a score. Grade against faults we injected, tests that must fail, or outcomes a person confirmed. Never against another model's opinion alone.
- Separate scores, never one "AI score". Detection, diagnosis, wrong causes named, unsafe actions, recovery verified, success declared too early: each one is measured and reported on its own.
- Autonomy is earned and enforced. What an agent may read, change or run is set by policy the platform enforces, raised only on graded evidence, and lowered when the evidence turns. A prompt is not a permission.
- Nothing unbuilt is described as built. Every surface (the site, the docs, the console, a sales conversation) says what runs today and what is planned, as the table above does.
- We are the first customer. Each step of the loop runs against our own platform before a customer's.
How this changes the existing RFCs
Nothing is cancelled. Each open RFC in the platform lab is read against the loop: which step it closes and for which workflow. RFCs added from today state their step of the loop in the Problem section. The order of work follows RFC 0047: the steps that one complete loop needs come first.
Alternatives considered
Lead with an AI SRE agent. That market is crowded with large players, and an agent grading itself is the problem this RFC is about. We keep our own agents, held to the same evidence, but we do not lead with one.
Stay a set of separate developer tools (chaos, sandbox, Lab, connections). Each is useful alone, but alone each competes with a mature product, and none answers whether AI work on a system was right.
Grade with models (LLM-as-judge). Cheap and broad, but it moves the trust problem instead of solving it. A model's judgement can be one signal among others, never the ground truth.
Wait for public benchmarks to settle it. Public benchmarks are contaminated, saturate, and measure someone else's system. The evidence that matters to a team is about their own system.
Decision
Decided. The owner set the direction on 2026-10-04: the thesis, the loop and the rules above. What stays open is how fast each step is built, and RFCs 0046 and 0047 are the first two.
Publication
This page is the public statement of the direction. The site's product pages and the console describe each product as a step of the loop, and say what is live and what is planned. Nothing on any surface may say the loop is complete until RFC 0047's acceptance test passes, and then only for the workflow it covers.
Status log
- 2026-10-04: Written. The direction, the loop and the rules; RFC 0046 (the evidence record) and RFC 0047 (the first complete loop) opened beside it.
- 2026-10-07: Decided: the direction is adopted (#227);
docs/product-direction.mdand the rule in CLAUDE.md hold it. Building the loop is RFCs 0046 and 0047, both open.