Problem
On 2026-10-03 the root disk of the machine that runs our platform filled up, mostly with container images and build cache. The kubelet saw disk pressure and evicted every pod on the node, as it is designed to. Our site answered 502 for about nineteen minutes.
Every check stayed honest and useless: up until all down at once. Nothing watched the disk. A warning at 80 % used would have fired about an hour before the eviction, when the fix was a cache prune, not a recovery.
The bench checks what a caller sees. What takes a platform down before any caller notices sits one level lower: a disk, a clock, a certificate file, a stopped unit, a pod restarting every few minutes. An engineer looks at those by hand once something feels wrong; the agent can look all the time.
Proposal
Two capabilities
Two capabilities, announced in the hello and off until the policy names them (parent rule 2):
host:read, for an agent running on the machine it reads (as a system service, or
in a cluster as one pod per node with the host's paths mounted read-only):
| Check kind | Reads | Typical threshold |
|---|---|---|
host_disk |
used share and free inodes per mount | warn 80 %, fail 90 % |
host_memory |
available memory, swap in use | fail under 5 % available |
host_load |
load per core over 5 and 15 minutes | warn above 2 per core |
host_clock |
whether the clock is synchronised, and the offset | fail past 500 ms |
host_cert_file |
expiry of a certificate file by path, never its key | warn under 14 days |
host_unit |
a service unit's active state and restarts | fail when not active |
host_fd |
open files against the process or system limit | warn 80 % |
host_images |
the container runtime's image and build cache size, reclaimable share | warn 20 GiB reclaimable |
kube:read, through a read-only service account in the cluster:
| Check kind | Reads |
|---|---|
kube_node |
node conditions: Ready, DiskPressure, MemoryPressure, PIDPressure |
kube_pods |
pods not ready, pending past a limit, in CrashLoopBackOff, restart counts over a window |
kube_evictions |
pods evicted, by node and reason |
kube_events |
events of type Warning, counted by reason and object kind |
kube_workloads |
deployments and stateful sets not at their desired replicas |
kube_volumes |
persistent volume claims' used share |
kube_certs |
certificate expiry from the certificate resource's status, never from a stored secret |
The kubelet starts evicting by default when a node's filesystem has less than 10 % free (); 80 % is well before it.
Thresholds in the file
Each kind is a check in checks.toml (RFC 0040.1) and becomes a managed monitor:
[[check]]
name = "root-disk"
surface = "host_disk"
target = "/"
every = "60s"
expect = { used_below = 0.80, inodes_free_above = 0.10 }
[[check]]
name = "node-pressure"
surface = "kube_node"
target = "all"
every = "60s"
expect = { conditions_false = ["DiskPressure", "MemoryPressure", "PIDPressure"] }
A breach is a failed check like any other: monitor.down, a class (disk_used,
node_pressure, crashloop) and the measured value, a number about the machine.
What the policy names
The policy lists every mount, unit, certificate file and namespace by name (the names
below are an example, not ours); anything not listed is refused like a host outside
networks.allow:
[work]
host_read = true
kube_read = true
[host]
mounts = ["/", "/data"]
units = ["kubelet", "chronyd"]
cert_files = ["example.pem"]
[kube]
namespaces = ["platform", "auth"]
The service account the agent's chart installs for kube:read gets get, list and
watch on nodes, pods, events, deployments, stateful sets, persistent volume claims and
certificate resources in those namespaces. It gets nothing on Secrets, nothing on
pods/log, pods/exec or nodes/proxy: the last one reaches the kubelet's own API,
which is far more than a read. Volume usage comes from the host side or the cluster's
existing metrics.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
What it never reads
- The values of secrets. Certificate expiry comes from the certificate resource's status, a certificate file, or the served endpoint's TLS check (RFC 0040.1), never a stored secret.
- Log lines. Logs are counts and classes by default: warnings by reason, crash loops by container, lines per level where logs are structured. A line can hold anything a service was given, including a customer's data.
- Process arguments and environment. Both are where credentials end up.
Traffic is read elsewhere
What crosses the network, which process talks to whom, how much and how well, is not a host read here. It is RFC 0061's: a separate, privileged companion of the agent that keeps what it sees on the host and ships as an extension with its own declared privileges. A check here may name a process or a volume; it never captures a packet.
Actions are not reads
Cordoning a node, restarting a pod, pruning images and scaling a deployment are not part
of host:read or kube:read. Each is a fault or a remediation under RFC 0040.4's rules: declared in the policy, off by default, paid from
the budget, with a lifetime where one makes sense. In production, and in any environment
that shares a deployment with production, an action is small (one pod, one node, one
prune), short, and confirmed by a person every time, never by a schedule or an assistant
alone. "Prune images past 85 %" is a reasonable rule; it is still an action, and waits
for its own decision.
No shell
The agent offers no shell and no exec into a container, not even read-only. A shell on the agent is a shell for whoever controls the account that sends it work, and "read-only" is not a property a shell can keep. What is offered instead is a fixed allowlist of named reads, each implemented in the agent or by running one fixed program with fixed arguments, each returning typed fields. A new read is a release of the agent, reviewed like code, not a string typed into a console.
The reads are read-only MCP tools on our MCP server (RFC 0005):
reliability_host_status and reliability_cluster_status return an agent's latest host
and cluster results and run nothing new, so an assistant asked why the site is down reads
the disk and the evictions instead of guessing. Later, customers point the same checks at
their own machines and clusters.
Alternatives considered
Alert from the cluster's metrics stack. On 2026-10-03 no alert there watched the disk; since 2026-10-06 RFC 0063 raises alerts on our own host's disks, filesystems, drive temperature and containers, sent to the operators' channel. That closes the gap for our machine and is the right first step. It does not replace this RFC: those alerts live on the machine whose failure they report (RFC 0063 leaves the probe from outside for a later decision), they cover our own machines rather than any machine an agent is on, and their results stay out of the bench. From the agent the result joins it: vantage points (RFC 0040.9), events, history, and the same checks for a customer's machines.
Give the agent a read-only shell. The most flexible option, and the one nobody can review. Typed reads answer the questions we know; a new question is a new read.
Decision
Open. Proposed: host:read and kube:read with the check kinds above, thresholds in
checks.toml, the policy naming every mount, unit, certificate file and namespace, a
service account with no access to Secrets, logs, exec or the kubelet, every action left to
RFC 0040.4's rules, no shell, and read-only MCP tools. First proof: the Engineering
team's agent, bound to our domain, reads the root disk, node pressure, pod restarts,
evictions and warning events of our own cluster every minute, and warns at 80 % disk.
Publication
The agent's policy reference gains [host] and [kube]; the agent's docs gain the check
kinds and the service account's exact rules; the MCP reference gains the two tools; the
console shows host and cluster checks on the agent's page.
Status log
- 2026-10-03: opened, after a full root disk evicted every pod on our node and the site answered 502 for about nineteen minutes.
- 2026-10-04: a second reason, and a wider scope. Our own server froze twice with nothing in its logs. Finding why took a person and an assistant two hours of reading by hand: one drive of a striped array dropped off its bus six minutes after every boot, the array went read-only, and the sign-in database stopped. The heat behind it came from one of our own services reading about 450 MB/s from that array around the clock for a week, which no dashboard showed. Every fact needed was on the machine: the storage layout and which volumes live on it, which devices sit on which bus, each drive's temperature and health, disk throughput history per device, bytes read per process, kernel messages. The agent should discover this model of the host by itself (devices, buses, arrays, filesystems, what each workload stores where, who reads and writes how much) and keep it current, so that understanding the system comes before guessing at it, for people and for the assistants that work on it. Read-only, as this RFC says; to be specified here.
- 2026-10-06: checked against what is built. Our host's own alerts are now RFC 0063's, and the alternative above says what they cover and what they do not; traffic seen on a host is RFC 0061's, referenced from the host model. Nothing of this RFC is built.
- 2026-10-07: Checked: nothing built (#257 only mentions it). Open.