Problem
A change can keep every surface answering and still break what happens when a neighbour is slow: a retry loop that multiplies load, a timeout longer than the caller's, a missing idempotency key that turns a lost answer into a double write. Our autonomous work needs to ask that question in staging before a change reaches production.
Today faults are hooks compiled into our binaries. Each service holds a FaultHandle
and calls it at the start of every request; the chaos tool flips it to slow, error,
hang or delayed_failure. That has three costs. Every service carries chaos code in
its production build. A fault exists only where someone wired the hook. And two of the
behaviours are unbounded: hang parks a request until the client gives up, and
delayed_failure nests another behaviour inside itself without a limit.
The Break-it playground showed the safe shape. Four moves, each with a cost (1 to 3)
from a budget of ten tokens that refills one every three seconds, each ending by itself
after 20 or 30 seconds, error shares capped at a half, hang and delayed_failure
unreachable, and a score that is the time to breach an objective.
Proposal
The agent as a proxy
The agent listens on an internal address and forwards to one upstream, named by a connection. A service is routed through it by pointing its callers there (a cluster service, a sidecar, a client's base URL). With no fault active it is a passthrough.
- L7 first: HTTP/1.1, HTTP/2 and gRPC, so a fault can carry a status or a gRPC code.
- L4 TCP later: latency and dropped connections for databases and brokers.
The listener answers only the internal network; the agent still opens no port to the internet. Each proxy is declared in the policy, so nothing is intercepted that the machine's owner did not route:
[work]
faults = true
[[proxy]]
name = "engine"
upstream = "http://engine.internal"
protocol = "grpc"
max_latency = "2s"
max_error_share = 0.5
max_fault_ttl = "10m"
The menu
| Fault | Parameters | Bounds | Cost per minute |
|---|---|---|---|
latency |
latency, jitter |
latency up to the proxy's max_latency; jitter at most the latency |
1 |
errors |
share, an HTTP status (429, 500, 502, 503, 504) or a gRPC code (UNAVAILABLE, INTERNAL, RESOURCE_EXHAUSTED, DEADLINE_EXCEEDED) |
share up to max_error_share |
2 |
drop |
share of connections reset without an answer |
up to a half | 3 |
drop_answer |
share of requests forwarded, then the answer discarded |
up to a half | 3 |
restart (later) |
one pod, through the cluster's API, in a namespace the policy names | one per run | 3 |
drop_answer is the proxy's form of the lost acknowledgement our store faults produce:
the upstream did the work, the caller never heard. It is what idempotency keys are for.
There is no hang and no delayed_failure. A hang holds a connection and its memory in
the proxy until the caller's timeout, and a caller without one never lets go. A delayed
failure takes effect later than it is confirmed, possibly as something else, so the
person confirming it cannot see what they confirm. Long latency up to the bound covers
the first; a timeline entry at a later time (RFC 0040.5) covers the second.
Budget, lifetime and the agent's own clock
Every fault is paid from a budget per account: a token bucket with a capacity and a refill rate the account sets within its plan, as Break-it's is. A fault costs its price for each started minute of its lifetime, taken when it starts and refunded for minutes it did not use.
Every fault has a lifetime, at most the proxy's max_fault_ttl. The agent computes the
end on its own monotonic clock when it applies the fault and removes it at that time
whether or not the platform can be reached. When the session to the platform drops, every
fault ends at once, as all work does (RFC 0029). The earlier of the two wins.
The objective
A fault run states, before it starts, what must hold:
[objective]
success_rate = 0.99
p99_ms = 300
window = "1m"
stop_on_breach = true
The proxy measures every request passing through it, not only ours, over a sliding
window. The run ends as pass (held throughout), fail with the time to breach
from the first fault, and the time to recover after the last fault ended.
stop_on_breach removes every fault at the first breach; in production it cannot be
switched off.
Production needs a person
A fault in an agent whose environment is production waits for a confirmation in the
console from a person with the owner or admin role, every time, recorded in the audit
log with their name. A schedule or an autonomous agent can queue one; it never starts
until that person confirms, and the request lapses after 15 minutes. Faults never run
from our cloud.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
What it replaces
Our faults move to the proxy, and the platform no longer needs FaultHandle for
anything the bench runs. The hooks stay for local chaos run scenarios until the proxy
covers them. The store faults (a write that commits and then fails inside the service)
have no proxy form; drop_answer is the nearest, and we say so rather than pretend.
How an agent uses it
A fault run is a run (RFC 0040.8), started through the API, the MCP tool
reliability_run_faults or iohr reliability faults, and read as its verdict, time to
breach and time to recover. An autonomous agent may start one only in an environment
other than production.
Later, customers route their own services through their own agent the same way.
Alternatives considered
Keep hooks in the services. No proxy to deploy, and chaos code in every production binary with faults only where someone wired them.
Kernel-level faults (tc, eBPF). More realistic, and root on the machine from day one. The proxy needs no privileges.
A service mesh's fault injection. Istio and Linkerd can inject delays and aborts; using them requires a mesh, and their faults carry no budget, lifetime or objective.
Decision
Open. Proposed: an L7 proxy in the agent declared in its policy, a menu of latency,
errors, drop and drop_answer with restart later, a per-account token bucket,
lifetimes the agent enforces on its own clock and ends early when the session drops, an
objective that turns a run into pass, fail, time to breach and time to recover, and a
person's confirmation for every fault in production. First proof: the agent on our own
server proxies the engine in our staging environment, which today is the same deployment
as production, so every fault there follows production's rules: a nightly run, confirmed
by a person each time, adds small latency and a small share of errors for at most a few
minutes under the light load of RFC 0040.3, and the objective of 99 % success with p99 under
300 ms, judged over sixty seconds, is reported with its time to breach.
Publication
The developer docs gain the fault menu with costs and bounds and the objective; the
agent's policy reference gains [[proxy]]; the console shows live faults with their
remaining lifetime and the budget.
Status log
- 2026-10-03: opened, from
FaultHandle, the Break-it playground's moves and budget, and the agent's refusal of faults today. - 2026-10-07: Checked: No PR names this RFC and the status log records nothing built. Open.