Implements: PRD 0001
Problem
The agent (RFC 0029) is the only way InOrbit reads or tests anything inside a customer's network. PRD 0001 makes it more than a check runner: it reads a system's repositories, docs and sites, runs builds and tests in its sandbox, and calls the model providers the customer chose. Each of those needs configuration, and today the agent is configured for checks only.
What a person meets today:
- Three files, all on the machine. One for the agent itself, one for its policy (what it may reach and do), one for its declared checks (RFC 0040.1). Each has its own format rules and hash. None of them describes sources to read, schedules, model providers, a proxy or logging.
- A restart for every change. The agent reads its files once, at start. A changed policy reaches the platform only with the next connection, after the restart.
- No way in from the platform or from CI. The console can show a policy's hash, not the policy. A company that keeps its infrastructure in a repository has nothing to put there and no step that applies it.
- No page that says what an agent is doing. The console shows an agent's status, version, last contact, policy hash and declared checks. It does not show the machine it runs on, what it is reading, how far it got, its errors or its own log lines. Nothing on it can pause the agent or start an update.
- Installation stops at the binary.
iohr ext install agentfetches the signed agent (RFCs 0028 and 0073). Running it as a background service on a laptop or a server is left to the person.
Two rules from RFC 0029 stand, and the design has to keep both:
- The machine's policy wins. Whoever runs the agent owns its policy. The platform can propose work; the agent refuses anything outside the policy and says why. Only a person with access to the machine widens what the agent may reach.
- Secrets stay where they are. The agent works with references to secrets and resolves them on the machine, when it needs them. InOrbit never holds a customer's secret values.
Proposal
One file, one schema
An agent is configured by one YAML document. It has these sections:
| Section | What it holds | Who may set it |
|---|---|---|
host |
The agent's own identity and plumbing: the API it talks to, its name, where its key and state live, the local status page, the proxy it uses, where its telemetry goes, and who manages the rest of the file (management) |
The file on the machine only |
policy |
The bounds: verified domains, the networks it may reach, which kinds of work are on, which credential references it may resolve and the hosts each may be sent to, which model providers it may call, how it may be updated, ceilings on concurrency and duration | The machine sets them; the platform may only narrow them |
environments |
One block per environment (production, staging, and so on), each with its own sources, checks, schedules and limits; an agent uses the block of the environment it was enrolled in, over the top-level defaults | Whoever manages the agent |
sources |
What to read: repositories, sites, documentation, granted connections, each with an optional credential reference | Whoever manages the agent |
schedules |
When each source is read again and how often each check runs, never more often than the plan allows | Whoever manages the agent, within the plan |
checks |
The declared checks of RFC 0040.1, unchanged in meaning | Whoever manages the agent |
sandbox |
Builds, tests and benchmarks the agent may run, and their budgets | Whoever manages the agent, within the plan |
models |
Which allowed provider each kind of reading work uses: the account's own model (RFC 0059) or an endpoint the customer runs, with a credential reference | Whoever manages the agent, within policy |
limits |
Resource limits: memory, disk for the index, parallel reads; defaults at the top, overridden per environment | Whoever manages the agent, within the plan |
logging |
Local level and format, and which levels are sent to the platform (or none) | Whoever manages the agent |
The agent only ever dials out. Nothing in the file opens a port to the outside; the local status page stays on the machine's loopback interface unless the machine's file says otherwise.
Where a destination matters, the machine decides it. A stolen console session or CI key
must not be able to send the agent's traffic through a proxy of its choosing, its telemetry
to a collector of its choosing, a credential to a host of its choosing, or code to a model
provider the company has not approved. So the proxy and the telemetry destination live in
host, and model providers and the hosts each credential may be sent to are allow-lists in
policy.
A short example (names invented):
version: 1
host:
name: build-runner-eu
management: platform
policy:
domains: [example.com]
networks:
allow: ["*.example.com", "github.com"]
work: {checks: true, reading: true, sandbox: false}
secrets:
allow:
- {ref: "vault:kv/prod/github#token", for: ["github.com"]}
- {ref: "env:IOHR_CHECK_*", for: ["*.example.com"]}
models:
allow: [account]
environments:
production:
sources:
- key: core
type: repository
url: https://github.com/example/core
credential: "vault:kv/prod/github#token"
- key: docs
type: site
url: https://docs.example.com
schedules:
core: {every: 6h}
docs: {every: 1d}
checks:
- name: api-health
target: https://api.example.com/health
every: 5m
models:
default: {provider: account}
logging:
level: info
to_platform: warn
Secrets are references, never values. Every field that carries a credential takes a
reference in the grammar the agent already resolves: an environment variable, a file, the
cluster's secret store or a vault path, each allowed in policy together with the hosts it
may be sent to. macOS adds the login keychain. The schema marks those fields as references.
The validator refuses a literal where a reference belongs, and it refuses any field whose
value looks like a key or a token, wherever it sits. The platform runs the same scan before it
stores a document, and refuses rather than stores; a refusal names the field, never its
value, and a refused document is never written to the platform's logs. The PRD also lists the
platform's own secret store as a place a reference may point. We leave it out: InOrbit holds no
customer secret values, and a store on the platform would be one.
A published, versioned JSON Schema. The schema is generated from the agent's own types,
so it cannot drift from what the agent accepts. It is published on the docs site and in the
SDK repository, one file per major version, so editors complete and check the file as it is
typed. version: 1 names the major version. A new optional field is a minor change, which
old agents ignore only when the file says it may be ignored (x-optional-since); anything
else is a new major version, and an agent refuses a major it does not know, with the
version it supports in the error.
iohr agent config validate checks a file against the schema, then the meaning it can
see (names are unique, schedules are well formed, and, when the file has a policy, that each
environment's sources and references fit it), and prints every error with its line and column,
all of them, not only the first. It runs without enrolment and without network access, so a
CI job can use it on any pull request. A file meant for agents the platform manages carries no
machine policy, so only plan checks it against the real policies, by asking the agents.
iohr agent config convert writes the YAML from today's three files. iohr agent config show
prints the effective configuration the agent is running.
Today's agents keep working
- An agent that finds no YAML file reads its three files exactly as now. Nothing is required of an existing installation.
convertmaps the files one to one: the agent file tohost, the policy file topolicy, the checks file to the environment'schecks(renaming nothing a person wrote). The policy hash is computed from the same normalised form as today, with the environment the agent is enrolled in, so a converted policy keeps its hash and the console shows no change.- The agent refuses to start with both forms present, so nobody edits the file that is not read.
Who manages an agent: one source of truth
Each agent has exactly one manager, recorded on the platform and confirmed by the agent:
| Manager | Where the configuration lives | What the console shows |
|---|---|---|
local |
The file on the machine | The configuration the agent reports, read-only |
console |
Versions on the platform, edited in the console | A form and a YAML view, editable by owners and admins |
ci |
A file in the customer's repository, applied by a CI step | Read-only, with a link to the repository, the file and the commit it came from |
host.management on the machine decides whether the platform may manage the agent at all.
With local, every attempt to apply configuration from the console or CI is refused and the
refusal names the reason. With platform, the machine's file keeps only host and policy,
and the rest comes from the console or CI.
Switching between console and ci is a deliberate act by a person. An owner or admin hands
an agent to CI (in the console, or with iohr signed in as themselves) and names the
repository allowed to manage it; a CI key can never take an agent over by itself, and an apply
from CI to an agent it does not manage is refused. "Manage here" in the console takes a
CI-managed agent back, and the next CI run is refused with a message saying so. Every
hand-over is recorded with who did it. The two never write over each other silently.
A managed document may carry its own policy, but only to narrow: it is stored with the
document, intersected with the machine's policy, and refused if it would widen anything.
Every write carries the version it was based on. A console save or a CI apply whose base is no longer current is refused with the difference, so two people never lose each other's changes.
The bounds: the machine's policy still wins
The configuration the agent runs is the intersection of three layers:
- The machine's
policy, which only someone on the machine changes. - The plan's ceiling: how many sources, how much code read in full, which sandbox checks, how often schedules may run (PRD 0001's plan table, set by the pricing RFCs).
- The managed configuration from the file, the console or CI.
A managed configuration that asks for more than the machine's policy allows (a host outside the allowed networks, a reference outside the allow-list, a kind of work that is off) is refused as a whole, with every violation listed. Refusing all of it, rather than applying the parts that fit, means the agent never runs a half-configuration nobody wrote. The plan on a pull request shows the same violations before anything is applied.
When the person on the machine narrows policy while a managed configuration is running,
the narrower policy applies at once. Any part of the managed configuration it no longer
allows is suspended and reported; it is never kept running because it was there first. When
the policy is widened again, the agent checks the suspended parts once more and resumes those
that now fit.
Reload without a restart
The agent applies a new configuration while it runs, from two places: its file, when the file changes or when the service manager asks it to reload, and the platform, over the agent's existing connection.
Over the connection, an apply is four steps:
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
- Offer. The platform sends the configuration with its version (a hash of its canonical form), its serial and its manager. It is sent only to an agent that announced it can apply configuration, and only after the connection is established. The agent refuses an offer whose serial is not newer than the one it runs, so an old or repeated offer is never applied.
- Check. The agent checks the schema and the meaning, intersects it with its policy and the plan's ceiling, and checks that every reference resolves on the machine (that it exists and is allowed; the value is read only when it is used).
- Apply atomically. The agent builds the whole new runtime beside the running one, then swaps it in one step. Work already running finishes under the configuration it started with; new work uses the new one. The agent writes the applied configuration and keeps the previous one as the last good configuration.
- Confirm. The agent answers with the version it applied and the hash of the effective configuration. The platform records the version as applied only on that answer. Until then, the console shows it as offered.
If any step fails, the agent keeps running on what it had, answers with the reason (a class and the place in the file, in fixed words that never repeat a value from the file), and the dashboard shows it. If the agent restarts during an apply and does not come up on the new configuration, it starts on the last good one. A version the agent refused or rolled back is marked failed on the platform and is not offered again; the next attempt is a new version. A person can roll back to the previous version from the console, through the same four steps.
A change of proxy is the one change that can cut the connection the confirmation travels
over. It is made on the machine (host), and the agent tries it before it commits: it opens a
test connection to the platform through the new proxy, one the platform answers and does not
count as the agent's session, and keeps the old proxy if the test fails.
Declared checks that change with a configuration reach the platform's monitors the same way they do today: the confirmation carries the agent's checks, and the platform reconciles the monitors from what the agent confirmed, not from what it offered.
The same path serves a dry run: the platform can ask a connected agent to check a configuration without applying it. That is how a pull request's plan learns what the agent itself would say.
The platform keeps every agent's desired and applied versions in its database, not in memory, and offers the desired one again after every reconnect when the two differ. An agent that went offline gets it when it next connects. An agent too old to apply configuration announces nothing, and the console says so instead of showing "pending" forever.
Old platforms and old agents. Neither side sends a message the other did not say it accepts. The agent lists what it can do in its greeting; the platform lists the message types it accepts in its welcome. Without that list, the agent sends only today's messages. A new agent connected to an older platform therefore behaves like today's agent instead of being disconnected.
Configuration as code in CI
The file lives in the customer's repository. It may name its targets: agents by name, by label, or by environment. Names and labels are the ones on the platform's record of the agent: labels are attached to the one-time code when it is made or set later by an owner, never by the agent itself, so a machine cannot move itself into another group's configuration.
iohr agent config plansends the file and its targets to the platform and prints, for each agent it matches, what would change (a structured difference), whether the agent is managed by CI, the console or the machine, and the agent's own check of it when the agent is connected. It changes nothing.iohr agent config applyapplies it. Applied from a plan, it refuses if any agent changed since the plan was made, and if the targets now match a different set of agents (one newly enrolled or labelled, say).- Ready-made steps for GitHub Actions and GitLab CI run
validateandplanon every pull request, post the plan as a comment, and runapplyon the default branch. They installiohrand the signed agent extension at pinned versions. Any other CI system runs the same two commands. Every apply records the repository, the file, the commit and the run. - The comment on a pull request is a summary by default: which agents, which sections change, and whether each agent accepted it, with a link to the full difference in the console. A public repository's pull requests should not list a company's internal sources and allow-lists.
A CI job authenticates with an API key (RFC 0016) that holds only the new configuration scope, and an owner can limit that key to agents of one environment or label, so a staging pipeline cannot configure production. Short-lived identity from the CI system itself, with no stored key, is a later step.
Editing in the console
On an agent the console manages, its page shows the configuration as a form for the common sections and as YAML, checked as it is typed against the same schema. Saving shows the difference first. Every version is kept with who saved it, when, and a note; any two versions can be compared, and any earlier version can be applied again (it becomes a new version, so history only grows).
The agent's page
Every agent gets a page, and the agents list becomes a dashboard. What it shows:
| Part | Shown | From |
|---|---|---|
| Status | Online, offline, paused or revoked; connected since; last seen; reconnects | The connection |
| Machine | Host name, operating system, architecture, how it was installed | The agent's greeting, refreshed on every connection |
| Version | The version it runs, whether it is behind the newest signed release, and how to update | The greeting and the extension catalogue |
| Policy | The policy hash, and work it refused with the reasons | The greeting and job results |
| Configuration | Version, manager, last apply and its result, rollback | The reload protocol |
| What it reads | Each source with its state and progress, a bar only when the total is known | A report the agent sends when something changes |
| Checks | Declared checks, monitors and self-tests, as today | Unchanged |
| Errors and logs | Failed sources and error lines first, then recent log lines | Lines the agent sends, bounded |
| Plan | Agents and sources used against the plan's limits | The plan's entitlements |
| Controls | Pause and resume, update, run a check, revoke | Owners and admins |
Every part the agent has to report is shown only when the agent announced that it reports it. An older agent's page says "not reported by this version yet", never an empty list or a zero that looks like a measurement.
Pause stops the platform sending the agent work, and tells an agent that understands it to stop its own scheduled reading too. The connection stays open, so the agent is still seen and can be resumed or revoked. A monitor whose check could not run during a pause is recorded as not seen, not as down.
Revoke already ends the agent's credential and connection; the page offers it beside the other controls.
Updates are not built today; this follows RFC 0029's design. The console shows the newest
signed version from the extension catalogue, a person asks for the update, and the agent
applies it under its own policy (policy.updates: a channel, a window, or never), keeps the
previous version to roll back to, and confirms by reporting the new version. The agent checks
the signature itself and refuses a version older than the one it runs unless its policy allows
a downgrade, so the platform cannot push an old signed release with known faults. An agent that
cannot update itself gets the command to run on the machine instead.
Logs are the agent's own operational lines: what it connected to, what it read and how
far, what failed and why. Never file contents, code, model prompts or answers, or document
text. A line names hosts, not full addresses: paths and query strings are cut. The agent
redacts each line before sending it, the platform redacts it again before storing it, and the
platform keeps a bounded number of recent lines per agent for a short time. The levels sent are
set in the file, and none sends nothing.
The host name is personal data on a laptop. It is stored only on the agent's record,
shown only to the account's members, included in the account's data export, erased with the
agent and with the person who enrolled it, and never written to the platform's own logs or
metrics. host.report_hostname: false sends none, and the page shows
the agent's name instead.
Install on every common system
iohr ext install agent already fetches and verifies the signed agent. Setup becomes one
more command with a one-time code from the console:
iohr ext install agent
iohr agent setup --code-file code.txt
setup writes the agent's files with a safe policy (checks and reading on, sandbox off, nothing
reachable beyond the verified domains), enrols with the one-time code, installs and starts the
background service, waits for the first connection, and prints the agent's page. The code
works once and expires within the hour; no long-lived secret is written. It can come from a
file or the terminal, never only the command line, where other users of the machine could
read it. Until the one-file format ships (slice 3), setup writes today's three files;
after it, the YAML file.
| System | How it runs |
|---|---|
| macOS, Apple silicon and Intel | A per-user background service, no administrator rights needed |
| Linux | A system service from the existing packages, or a per-user service when installed without root |
| Containers and clusters | The published image and chart, the file mounted from the cluster's configuration |
| Windows | Later |
iohr agent service status|stop|start|logs|uninstall manage the service the same way on each
system. The console's enrolment wizard and the agent's entry in the extension catalogue show
the same two commands, for each system. The first demonstration, an agent on a Mac enrolled
and visible on its page, needs only slices 1, 2 and 5; reading a repository there also needs
the reading work, which its own RFC describes.
Plan limits
Enrolment, plan and apply ask the platform's entitlements for the account's limits (the
number of agents, sources, code read in full, sandbox checks, the shortest schedule). The
offer to the agent carries the ceiling, so the agent enforces it on the machine as well. The
values belong to the pricing RFCs; this RFC defines where they are checked and what a
refusal says: the measured figure, the plan's limit, and which plan covers it.
Slices
Each slice ships on its own; nothing waits for a later one to be safe.
| # | Slice | Where | Ships alone because |
|---|---|---|---|
| 0 | The platform ignores and counts message types it does not know, instead of closing the connection, and lists in its welcome the types it accepts | Platform | Today's agents send nothing new |
| 1 | The greeting gains the machine (host name, system, architecture, install method), update channel, configuration summary and the capability names; the platform stores and returns them | Platform, agent | Optional members inside the greeting, which today's platform already ignores; old agents send none |
| 2 | Install on every common system: iohr agent setup, the service commands, the macOS per-user service, the wizard and catalogue cards |
Agent, console | Uses today's enrolment and today's files |
| 3 | One file on the machine: schema version 1 published, validate, convert, show, reload from the file with last good kept |
Agent | Reports the result only to a platform whose welcome accepts it (slice 0); otherwise in the next greeting |
| 4 | The page's backend: pause and resume, source progress, log lines, update requests, labels, agent actions in the account's audit log | Platform | Each part shown only when the agent announced it |
| 5 | The dashboard pages in the console | Console | Reads slices 1 and 4; shows "not reported yet" for the rest |
| 6 | The agent's side of slice 4: pause, source progress, log lines, applying an update under policy.updates |
Agent | Sends only what the welcome accepts |
| 7 | Configuration managed from the platform: versions, the console editor, the reload protocol, rollback, managers | Platform, agent, console | Needs 0 and 3; agents without the capability are left alone |
| 8 | Configuration as code: plan, apply, targets, keys limited to an environment or label, hand-over to CI, the CI steps |
Command line, platform | Needs 7 |
| 9 | Plan limits checked at enrolment, plan, apply and on the agent | Platform, agent | Allows everything until the pricing values exist |
Alternatives considered
Let the platform own the policy. The simplest console experience: edit everything, including what the agent may reach, from the browser. Rejected. It breaks RFC 0029's first rule: a stolen console session or API key could then point an agent inside a company's network anywhere. Keeping the policy on the machine costs one thing: widening reach still needs a person there. The dashboard shows what a refused configuration would need, so that person knows what to change.
Apply the parts of a configuration that fit, skip the rest. Friendlier when one source is out of bounds. Rejected: the agent would run a configuration nobody wrote or reviewed, and the plan on the pull request would not describe what runs. A refusal that lists every violation is one more edit away from correct.
Restart the agent to apply, managed by the service manager. Already how the cluster chart works. Rejected as the main path: a restart drops running reads and checks, a bad file leaves the agent down instead of on its last good configuration, and the confirmation only arrives with the next connection. A restart still works and reads the same file.
Pull the configuration on a timer instead of offering it. Simpler on the platform. Rejected: it delays every change by the interval and adds traffic from every agent, when the connection that would carry it is already open.
Keep TOML. The agent's files are TOML today. The PRD asks for YAML, and the deciding reason is the customer's side: CI systems, clusters and most configuration repositories are YAML, and editors check YAML against a JSON Schema out of the box. The agent keeps reading TOML for the installations that have it.
Labels set by the agent. Convenient for fleets built from one image. Rejected for now: a compromised machine could label itself into a group whose configuration names sources and references it should not see. Labels attached to the one-time code cover image-built fleets: every agent enrolled with it gets them.
Decision
Open. Slice 0 can ship now. The rest waits for the owner's review of the open questions in the design note.
Publication
- The agent's install, configuration, schema and CI pages in the developer docs, and the schema files themselves, once slice 3 ships.
- The console's agents list, agent pages, configuration editor and the catalogue entry's install card as their slices ship.
- This page's status log records each slice.
Status log
- 2026-10-07: Opened. Implements PRD 0001. Slice 0 found during the research; the plan revised after an independent critique.
- 2026-10-08: RFC 0100 proposes that the plan's ceiling come from the signed licence (RFC 0061.1)
rather than the platform's offer, and that edits in the agent's local console be manager
local; marked for the owner's decision (RFC 0100, D5).