Problem
On 2026-10-04 and 2026-10-05 the platform had three kinds of outage:
- A storage array lost a drive twice, and every database stopped with it.
- The API gateway leaked memory until the kernel killed it over and over.
- The host froze five times. The last freeze lasted about nine hours.
Each time, a person noticed first, sometimes hours later. The monitors (RFC 0037) watch the public endpoints from outside the stack, but they say that something is down, not why. And no alert at all looked at the layer underneath: the drives, the arrays, the filesystems and the containers' memory.
Compliance asks for the same thing: control LOG-03 in our compliance programme requires that security and availability events reach a person without anyone having to look. Until now it was a gap.
Proposal
What is watched
Three new metric sources feed the existing metrics store:
- The host. The distribution's node_exporter reports per array the disks it needs and the disks that are active or failed; per drive its model, serial and bus state, and its temperature against the drive's own warning threshold; per filesystem whether it is read-only, its size and its free space. It listens only on the cluster's private network, never on the host's LAN or public address, and the scrape keeps only the series the alerts read.
- The containers. Each node's kubelet reports every container's memory working
set, its limit and its OOM kills. The metrics store reads it with the
nodes/metricspermission only. It never getsnodes/proxy, which would reach far more of the kubelet than metrics. - The chaos tool's schedules. Every finished scheduled run sets two gauges: its verdict (1 passed, 0 failed or errored) and when it finished, labelled with the schedule's name and id and, for scenario runs, the scenario. The chaos tool runs copies of the services inside its own process, so these gauges live on a recorder of their own with its own listener; the copies' request metrics never appear under the chaos tool's name.
The rules
evaluates every rule once a minute, and the rules live in the repository (read-only in the UI):
| Rule | Fires when | Severity |
|---|---|---|
| Array missing a disk | an array runs with fewer active disks than it needs, for 1 min | critical |
| Array member failed | an array marks a member failed, for 1 min | critical |
| Drive left the bus | an drive is anything but live, for 1 min | critical |
| Drive hot | a drive is within 3 °C of its own warning temperature for 10 min | warning |
| Filesystem read-only | the kernel remounted a filesystem read-only, for 1 min | critical |
| Disk filling | a filesystem is over 85% full for 10 min | warning |
| Disk full | a filesystem is over 95% full for 2 min | critical |
| Container OOM-killed | any OOM kill in the last 10 min | critical |
| Container near its limit | above 90% of its memory limit for 15 min | warning |
| A source went quiet | the host exporter or every kubelet stops answering | critical / warning |
| End-to-end check failing | the arena's every-10-minutes check of every way in has failed for 15 min | critical |
| End-to-end check stopped | no run of that check finished in 30 min, or chaos is not scraped | warning |
| Scenario failing | a scenario failed or errored on its last scheduled run under "all scenarios" | warning |
Every alert names the array, drive serial, filesystem or container, and links to the runbook step that applies: what to check, and what not to do, such as remounting a filesystem before the database dumps exist.
Drives are named by serial and model, because device names change on every boot.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Where alerts go
Alerts go to the operators' private Discord channel through 's notification policy, which already exists for the dns and radar services: grouped by rule and service, repeated every 12 hours while firing. The webhook lives in 's database and in Discord, never in a file.
What it never carries
An alert's labels and text are infrastructure names only: the array, drive, filesystem, namespace, pod and container. No request, payload, header, account or person ever reaches a rule, a label or a message. The metrics behind them hold no customer data either.
What it does not cover yet
The host stopping as a whole. Every rule here runs on the host it watches. When the host freezes, the rules freeze with it, and that is how the nine-hour outage went unnoticed. Catching it needs something outside the host. Options:
- A scheduled worker at the DNS provider we already use. It probes the public endpoints every minute and posts to the same channel after three failures. No new vendor.
- A dead-man's switch service. The host checks in every minute and the service raises the alarm when the check-ins stop. This adds a vendor, which sees only the check-ins.
- Both. The probe also catches edge and network failures; the switch catches a silent host behind a cached page.
The owner decided to leave this for later (2026-10-06).
Request-level rules. Error rate, latency by route and upstream health are on the dashboards but have no rules yet.
Security and compliance
- Least privilege. Read-only metrics permissions. The host exporter is reachable only from the cluster's network. The scrape keeps an allow-list of series.
- No customer data. See "What it never carries".
- Control LOG-03 in the compliance registry moves from gap to partial with this change: the host and container events are covered, and the whole-host and request-level events are still open. That is recorded in the compliance repository, not claimed here.
- The platform aligns with SOC 2 and ISO 27001 criteria. It is not certified, and this RFC does not change that.
Alternatives considered
- vmalert beside . A second rule engine and a second notifier. The platform already provisions rules for two services and routes them to the same channel. One engine is simpler to operate.
- kube-state-metrics for restarts and OOM reasons. Another deployment with cluster-wide list rights. cAdvisor's OOM counter answers the question the incident asked. Restart storms can follow if they turn out to matter.
- Running the exporter inside the cluster. A pod sees the kernel's arrays and sensors, but not the host's own filesystems: the root disk that filled on 2026-10-03 would stay invisible.
Status log
- 2026-10-06: opened. Built: node_exporter on the host (cluster network only), the
kubelet container metrics with
nodes/metrics, the eleven rules above, Discord delivery through the existing policy. Every host query was tested against live data before merge. Not yet: the outside probe (owner: later), chaos schedule metrics, request-level rules. - 2026-10-06: step 2 built, not yet rolled: the chaos tool's schedule gauges on a recorder of its own, scraped from the chaos pod, and three rules on them (the arena's end-to-end check failing or stopped, a scenario failing under "all scenarios"). A test proves the verdicts and labels reach the scrape and that none of the in-process services' metrics do.
- 2026-10-07: Checked: step 1 is live (#360); step 2 (#401) is merged, not rolled. Open.