This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0040.9open2026-10-03

Vantage points that do not share our fate

Several agents per environment, each declaring the machine, network, power and provider it shares fate with, so a monitor counts as covered only when someone outside the target's fate watches it, goes down only when independent vantage points agree, and a bench that goes quiet is raised as an alert by a path that does not run on the platform.

Part of RFC 0040 Chaos as a service, our test bench in the customer's network

Problem

Today the only agent that watches this platform shares the platform's fate: what takes the platform down can take the agent with it. Every monitor turns unknown, because RFC 0037 rightly says an agent that is offline has not seen the target fail. And then nothing alerts. The events that would say so, monitor.down and the rest, are produced and delivered by the platform that just went away. The worst failure we can have is the one the bench is built not to report.

Our cloud is no help either: checks "from our cloud" share the fate of what they check. Today it is not a second vantage point, whatever the console calls it.

The smaller problem runs the other way: one vantage point that loses its own route for thirty seconds sees every target fail at once, and with one agent that blip and a real outage look the same.

Proposal

Fate groups

Each agent declares what it shares fate with, in its own agent.toml, as labels the person who installed it chooses:

[fate]
machine = "app-1"
network = "site-a-lan"
power = "site-a-feed-1"
provider = "example-colo"

The labels are names, not addresses (the values above are an example). The platform compares them and nothing else. A target's fate group is the fate group of the machine it runs on, declared once per environment in the console (for us: the platform's environment has the same four labels as the agent on our own server).

An undeclared label counts as shared. A vantage point that does not say where it runs is treated as if it ran next to the target, so leaving the file empty never makes a monitor look covered.

Covered, or not

A monitor is covered when at least one vantage point that runs it shares none of the target's declared labels. The console and the API show every monitor's coverage, and the environment's page shows one line: how many of its monitors are covered. A monitor that is not covered is marked, not hidden, because it will go unknown at the moment it matters. With one agent on our own server, none of ours is covered today, and the console will say so.

Several agents, one file

Several agents may serve one environment, each with its own enrollment, policy and fate. A checks.toml (RFC 0040.1) can be the same file on every one of them; the platform reconciles one monitor per check with one result stream per vantage point, not one monitor per agent. A vantage point's results stay its own: latency from another network is a different number and is never averaged with the local one.

Quorum

A monitor's state comes from its vantage points together:

Vantage points report State
failure from two that share no label with each other down
failure from every vantage point that reported, none of them independent of another down, marked "not confirmed independently"
failure from some, success from an independent one up, with the class vantage_disagrees on the failing one
success from at least one up
nothing within two intervals unknown

vantage_disagrees is a finding about the path, not about the target: a vantage point's network, its DNS, or a route between it and the target. It is shown on the vantage point's page and raises monitor.vantage_disagrees when it lasts longer than the monitor's failure threshold, so a broken uplink is noticed without waking anyone for an outage that did not happen.

Results count together when they fall in the same two intervals. A refusal check (RFC 0040.1, class guard_open) needs no quorum: a guard seen open from any vantage point is open.

The bench that goes quiet

unknown is not an alert today. It becomes one, twice over. Inside the platform. An agent that has sent nothing for three of its shortest check intervals raises agent.silent; every vantage point of an environment silent at once raises bench.silent. These are events like any other (RFC 0022) and reach whoever subscribes. They cover an agent that died while the platform lives.

Outside the platform. When the platform itself is what died, its events die with it. So the path that reports a silent bench must not run on the platform:

  • The off-machine agent speaks for itself. An agent may hold a fallback in its policy: a mail relay or a webhook URL, with the credential as a secret reference resolved on its own machine like any other (RFC 0029). When it cannot reach the platform for five minutes, or when its quorum says the platform's own surfaces are down, it sends one message: which environment, which checks failed, since when. Names and classes, never a body. One message per state change, not one per check run.
  • A heartbeat somewhere else. Each agent can also ping a third-party heartbeat service on a fixed interval while its session to the platform is healthy, such as Healthchecks.io, which alerts when pings stop. That catches the case the agent cannot report: the agent itself is gone.

Neither path holds a credential to our API (parent rule 3). The fallback is off by default and named in the policy, so a company decides where its agent may send.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Our cloud as a vantage point

Our cloud counts as a vantage point exactly when it declares fate labels that differ from the target's. Today its labels are the platform's, so it covers nothing. When it moves elsewhere, the labels change and the console starts counting it, with no special case in the code.

How an agent uses it

Coverage and quorum are fields of a monitor and of a run (RFC 0040.8): covered, vantage_points with each one's verdict, and the quorum's decision. An autonomous change that touches the gateway or the edge reads them after it rolls: a check that fails only from outside is exactly what a change to public routing breaks. Later, customers place agents across their own sites and providers and get the same rule.

Alternatives considered

Rely on our cloud for the outside view. Correct once our cloud runs elsewhere. Today it does not, and a bench that pretends otherwise is worse than one that says it is blind.

Mark the target down when its only agent goes silent. It would alert, and it would alert on every agent restart and every lost uplink, which teaches people to ignore it. Silence is its own alert, with its own name.

A commercial uptime service in place of a second agent. Several exist and they work. They see only public surfaces, cannot run our refusal checks or checks.toml, and would be a second source of truth. We use one only as the heartbeat, the smallest job that must live off the platform.

Quorum by majority of all agents. Simple to explain, and three agents in one rack would outvote one elsewhere. Independence is the point, so the rule counts fate, not heads.

Decision

Open. Proposed: fate labels per agent with undeclared meaning shared, coverage per monitor shown in the console, several agents per environment from one checks.toml, quorum by independent vantage points with vantage_disagrees for the rest, agent.silent and bench.silent inside the platform, and a fallback message plus a third-party heartbeat outside it. First proof: a second agent of the Engineering team, bound to our domain, on a small machine at another provider with its own network and power, runs the same checks.toml as the agent on our own server; every monitor of our platform shows as covered; blocking the second agent's route to the platform in a drill produces its mail within five minutes, and stopping it produces the heartbeat service's alert.

Publication

The agent's docs gain [fate] and the fallback in the policy reference; the developer docs' monitors page explains coverage, quorum and the silence events; the event catalogue gains agent.silent, bench.silent and monitor.vantage_disagrees; the console shows coverage per monitor and per environment.

Status log

  • 2026-10-03: opened, because every check of our platform shares the platform's fate, and a failure that takes both down would turn every monitor unknown with nobody told.
  • 2026-10-07: Checked: No PR names this RFC and the status log records nothing built. Open.

← Back to Platform