This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0040.10open2026-10-03

Hosts and clusters, read first

The agent reads what an engineer checks by hand on a machine and a cluster (disk, inodes, memory, clocks, certificates, units, node pressure, pods, restarts, warning events), each as a check kind with thresholds in checks.toml, read-only and named in the policy, with every action treated as a fault under its rules and no shell offered.

Part of RFC 0040 Chaos as a service, our test bench in the customer's network

Problem

On 2026-10-03 the root disk of the machine that runs our platform filled up, mostly with container images and build cache. The kubelet saw disk pressure and evicted every pod on the node, as it is designed to. Our site answered 502 for about nineteen minutes.

Every check stayed honest and useless: up until all down at once. Nothing watched the disk. A warning at 80 % used would have fired about an hour before the eviction, when the fix was a cache prune, not a recovery.

The bench checks what a caller sees. What takes a platform down before any caller notices sits one level lower: a disk, a clock, a certificate file, a stopped unit, a pod restarting every few minutes. An engineer looks at those by hand once something feels wrong; the agent can look all the time.

Proposal

Two capabilities

Two capabilities, announced in the hello and off until the policy names them (parent rule 2):

host:read, for an agent running on the machine it reads (as a system service, or in a cluster as one pod per node with the host's paths mounted read-only):

Check kind Reads Typical threshold
host_disk used share and free inodes per mount warn 80 %, fail 90 %
host_memory available memory, swap in use fail under 5 % available
host_load load per core over 5 and 15 minutes warn above 2 per core
host_clock whether the clock is synchronised, and the offset fail past 500 ms
host_cert_file expiry of a certificate file by path, never its key warn under 14 days
host_unit a service unit's active state and restarts fail when not active
host_fd open files against the process or system limit warn 80 %
host_images the container runtime's image and build cache size, reclaimable share warn 20 GiB reclaimable

kube:read, through a read-only service account in the cluster:

Check kind Reads
kube_node node conditions: Ready, DiskPressure, MemoryPressure, PIDPressure
kube_pods pods not ready, pending past a limit, in CrashLoopBackOff, restart counts over a window
kube_evictions pods evicted, by node and reason
kube_events events of type Warning, counted by reason and object kind
kube_workloads deployments and stateful sets not at their desired replicas
kube_volumes persistent volume claims' used share
kube_certs certificate expiry from the certificate resource's status, never from a stored secret

The kubelet starts evicting by default when a node's filesystem has less than 10 % free (); 80 % is well before it.

Thresholds in the file

Each kind is a check in checks.toml (RFC 0040.1) and becomes a managed monitor:

[[check]]
name = "root-disk"
surface = "host_disk"
target = "/"
every = "60s"
expect = { used_below = 0.80, inodes_free_above = 0.10 }
[[check]]
name = "node-pressure"
surface = "kube_node"
target = "all"
every = "60s"
expect = { conditions_false = ["DiskPressure", "MemoryPressure", "PIDPressure"] }

A breach is a failed check like any other: monitor.down, a class (disk_used, node_pressure, crashloop) and the measured value, a number about the machine.

What the policy names

The policy lists every mount, unit, certificate file and namespace by name (the names below are an example, not ours); anything not listed is refused like a host outside networks.allow:

[work]
host_read = true
kube_read = true
[host]
mounts = ["/", "/data"]
units = ["kubelet", "chronyd"]
cert_files = ["example.pem"]
[kube]
namespaces = ["platform", "auth"]

The service account the agent's chart installs for kube:read gets get, list and watch on nodes, pods, events, deployments, stateful sets, persistent volume claims and certificate resources in those namespaces. It gets nothing on Secrets, nothing on pods/log, pods/exec or nodes/proxy: the last one reaches the kubelet's own API, which is far more than a read. Volume usage comes from the host side or the cluster's existing metrics.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

What it never reads

  • The values of secrets. Certificate expiry comes from the certificate resource's status, a certificate file, or the served endpoint's TLS check (RFC 0040.1), never a stored secret.
  • Log lines. Logs are counts and classes by default: warnings by reason, crash loops by container, lines per level where logs are structured. A line can hold anything a service was given, including a customer's data.
  • Process arguments and environment. Both are where credentials end up.

Traffic is read elsewhere

What crosses the network, which process talks to whom, how much and how well, is not a host read here. It is RFC 0061's: a separate, privileged companion of the agent that keeps what it sees on the host and ships as an extension with its own declared privileges. A check here may name a process or a volume; it never captures a packet.

Actions are not reads

Cordoning a node, restarting a pod, pruning images and scaling a deployment are not part of host:read or kube:read. Each is a fault or a remediation under RFC 0040.4's rules: declared in the policy, off by default, paid from the budget, with a lifetime where one makes sense. In production, and in any environment that shares a deployment with production, an action is small (one pod, one node, one prune), short, and confirmed by a person every time, never by a schedule or an assistant alone. "Prune images past 85 %" is a reasonable rule; it is still an action, and waits for its own decision.

No shell

The agent offers no shell and no exec into a container, not even read-only. A shell on the agent is a shell for whoever controls the account that sends it work, and "read-only" is not a property a shell can keep. What is offered instead is a fixed allowlist of named reads, each implemented in the agent or by running one fixed program with fixed arguments, each returning typed fields. A new read is a release of the agent, reviewed like code, not a string typed into a console.

The reads are read-only MCP tools on our MCP server (RFC 0005): reliability_host_status and reliability_cluster_status return an agent's latest host and cluster results and run nothing new, so an assistant asked why the site is down reads the disk and the evictions instead of guessing. Later, customers point the same checks at their own machines and clusters.

Alternatives considered

Alert from the cluster's metrics stack. On 2026-10-03 no alert there watched the disk; since 2026-10-06 RFC 0063 raises alerts on our own host's disks, filesystems, drive temperature and containers, sent to the operators' channel. That closes the gap for our machine and is the right first step. It does not replace this RFC: those alerts live on the machine whose failure they report (RFC 0063 leaves the probe from outside for a later decision), they cover our own machines rather than any machine an agent is on, and their results stay out of the bench. From the agent the result joins it: vantage points (RFC 0040.9), events, history, and the same checks for a customer's machines.

Give the agent a read-only shell. The most flexible option, and the one nobody can review. Typed reads answer the questions we know; a new question is a new read.

Decision

Open. Proposed: host:read and kube:read with the check kinds above, thresholds in checks.toml, the policy naming every mount, unit, certificate file and namespace, a service account with no access to Secrets, logs, exec or the kubelet, every action left to RFC 0040.4's rules, no shell, and read-only MCP tools. First proof: the Engineering team's agent, bound to our domain, reads the root disk, node pressure, pod restarts, evictions and warning events of our own cluster every minute, and warns at 80 % disk.

Publication

The agent's policy reference gains [host] and [kube]; the agent's docs gain the check kinds and the service account's exact rules; the MCP reference gains the two tools; the console shows host and cluster checks on the agent's page.

Status log

  • 2026-10-03: opened, after a full root disk evicted every pod on our node and the site answered 502 for about nineteen minutes.
  • 2026-10-04: a second reason, and a wider scope. Our own server froze twice with nothing in its logs. Finding why took a person and an assistant two hours of reading by hand: one drive of a striped array dropped off its bus six minutes after every boot, the array went read-only, and the sign-in database stopped. The heat behind it came from one of our own services reading about 450 MB/s from that array around the clock for a week, which no dashboard showed. Every fact needed was on the machine: the storage layout and which volumes live on it, which devices sit on which bus, each drive's temperature and health, disk throughput history per device, bytes read per process, kernel messages. The agent should discover this model of the host by itself (devices, buses, arrays, filesystems, what each workload stores where, who reads and writes how much) and keep it current, so that understanding the system comes before guessing at it, for people and for the assistants that work on it. Read-only, as this RFC says; to be specified here.
  • 2026-10-06: checked against what is built. Our host's own alerts are now RFC 0063's, and the alternative above says what they cover and what they do not; traffic seen on a host is RFC 0061's, referenced from the host model. Nothing of this RFC is built.
  • 2026-10-07: Checked: nothing built (#257 only mentions it). Open.

← Back to Platform