This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0040.4open2026-10-03

Faults through the agent's proxy

The agent becomes a reverse proxy in front of a service and injects latency, a share of errors and dropped connections from a small menu, each fault paid from a budget, ended by the agent itself when its lifetime runs out, and judged by an objective stated before it starts.

Part of RFC 0040 Chaos as a service, our test bench in the customer's network

Problem

A change can keep every surface answering and still break what happens when a neighbour is slow: a retry loop that multiplies load, a timeout longer than the caller's, a missing idempotency key that turns a lost answer into a double write. Our autonomous work needs to ask that question in staging before a change reaches production.

Today faults are hooks compiled into our binaries. Each service holds a FaultHandle and calls it at the start of every request; the chaos tool flips it to slow, error, hang or delayed_failure. That has three costs. Every service carries chaos code in its production build. A fault exists only where someone wired the hook. And two of the behaviours are unbounded: hang parks a request until the client gives up, and delayed_failure nests another behaviour inside itself without a limit.

The Break-it playground showed the safe shape. Four moves, each with a cost (1 to 3) from a budget of ten tokens that refills one every three seconds, each ending by itself after 20 or 30 seconds, error shares capped at a half, hang and delayed_failure unreachable, and a score that is the time to breach an objective.

Proposal

The agent as a proxy

The agent listens on an internal address and forwards to one upstream, named by a connection. A service is routed through it by pointing its callers there (a cluster service, a sidecar, a client's base URL). With no fault active it is a passthrough.

  • L7 first: HTTP/1.1, HTTP/2 and gRPC, so a fault can carry a status or a gRPC code.
  • L4 TCP later: latency and dropped connections for databases and brokers.

The listener answers only the internal network; the agent still opens no port to the internet. Each proxy is declared in the policy, so nothing is intercepted that the machine's owner did not route:

[work]
faults = true

[[proxy]]
name = "engine"
upstream = "http://engine.internal"
protocol = "grpc"
max_latency = "2s"
max_error_share = 0.5
max_fault_ttl = "10m"

The menu

Fault Parameters Bounds Cost per minute
latency latency, jitter latency up to the proxy's max_latency; jitter at most the latency 1
errors share, an HTTP status (429, 500, 502, 503, 504) or a gRPC code (UNAVAILABLE, INTERNAL, RESOURCE_EXHAUSTED, DEADLINE_EXCEEDED) share up to max_error_share 2
drop share of connections reset without an answer up to a half 3
drop_answer share of requests forwarded, then the answer discarded up to a half 3
restart (later) one pod, through the cluster's API, in a namespace the policy names one per run 3

drop_answer is the proxy's form of the lost acknowledgement our store faults produce: the upstream did the work, the caller never heard. It is what idempotency keys are for.

There is no hang and no delayed_failure. A hang holds a connection and its memory in the proxy until the caller's timeout, and a caller without one never lets go. A delayed failure takes effect later than it is confirmed, possibly as something else, so the person confirming it cannot see what they confirm. Long latency up to the bound covers the first; a timeline entry at a later time (RFC 0040.5) covers the second.

Budget, lifetime and the agent's own clock

Every fault is paid from a budget per account: a token bucket with a capacity and a refill rate the account sets within its plan, as Break-it's is. A fault costs its price for each started minute of its lifetime, taken when it starts and refunded for minutes it did not use.

Every fault has a lifetime, at most the proxy's max_fault_ttl. The agent computes the end on its own monotonic clock when it applies the fault and removes it at that time whether or not the platform can be reached. When the session to the platform drops, every fault ends at once, as all work does (RFC 0029). The earlier of the two wins.

The objective

A fault run states, before it starts, what must hold:

[objective]
success_rate = 0.99
p99_ms = 300
window = "1m"
stop_on_breach = true

The proxy measures every request passing through it, not only ours, over a sliding window. The run ends as pass (held throughout), fail with the time to breach from the first fault, and the time to recover after the last fault ended. stop_on_breach removes every fault at the first breach; in production it cannot be switched off.

Production needs a person

A fault in an agent whose environment is production waits for a confirmation in the console from a person with the owner or admin role, every time, recorded in the audit log with their name. A schedule or an autonomous agent can queue one; it never starts until that person confirms, and the request lapses after 15 minutes. Faults never run from our cloud.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

What it replaces

Our faults move to the proxy, and the platform no longer needs FaultHandle for anything the bench runs. The hooks stay for local chaos run scenarios until the proxy covers them. The store faults (a write that commits and then fails inside the service) have no proxy form; drop_answer is the nearest, and we say so rather than pretend.

How an agent uses it

A fault run is a run (RFC 0040.8), started through the API, the MCP tool reliability_run_faults or iohr reliability faults, and read as its verdict, time to breach and time to recover. An autonomous agent may start one only in an environment other than production.

Later, customers route their own services through their own agent the same way.

Alternatives considered

Keep hooks in the services. No proxy to deploy, and chaos code in every production binary with faults only where someone wired them.

Kernel-level faults (tc, eBPF). More realistic, and root on the machine from day one. The proxy needs no privileges.

A service mesh's fault injection. Istio and Linkerd can inject delays and aborts; using them requires a mesh, and their faults carry no budget, lifetime or objective.

Decision

Open. Proposed: an L7 proxy in the agent declared in its policy, a menu of latency, errors, drop and drop_answer with restart later, a per-account token bucket, lifetimes the agent enforces on its own clock and ends early when the session drops, an objective that turns a run into pass, fail, time to breach and time to recover, and a person's confirmation for every fault in production. First proof: the agent on our own server proxies the engine in our staging environment, which today is the same deployment as production, so every fault there follows production's rules: a nightly run, confirmed by a person each time, adds small latency and a small share of errors for at most a few minutes under the light load of RFC 0040.3, and the objective of 99 % success with p99 under 300 ms, judged over sixty seconds, is reported with its time to breach.

Publication

The developer docs gain the fault menu with costs and bounds and the objective; the agent's policy reference gains [[proxy]]; the console shows live faults with their remaining lifetime and the budget.

Status log

  • 2026-10-03: opened, from FaultHandle, the Break-it playground's moves and budget, and the agent's refusal of faults today.
  • 2026-10-07: Checked: No PR names this RFC and the status log records nothing built. Open.

← Back to Platform