This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0053open2026-10-04

Operations journal

A team's operational knowledge (reusable commands without secrets, facts about its hosts and systems, and runbooks, linked to the incidents they came from) kept on our own platform, internal by default, written by people, assistants and the agent, and read over MCP, the API, the command line and our own models under scopes the owner grants.

Problem

What we know about running this platform lives in the wrong places. The commands that fixed the last disk-pressure outage are in one person's shell history. Which storage array holds which workload's data, which device runs hot, which service writes the most: an engineer finds out by hand during an incident, and the next session, a person's or an assistant's, finds out again. Some of it ends up in an assistant's own memory, on a vendor's servers, where we cannot query it, audit who read it or delete it.

The owner put the requirement plainly on 2026-10-04: our own system to store this and read it back, not handed to the assistant's vendor, with the end goal that an assistant (Claude among them) can query it on its own, and so can the site's and console's own models, from the start.

The loop of RFC 0045 needs the same thing. Observe, model and evidence all produce knowledge about a running system that outlives the session that found it. Today that knowledge is either a log line (no meaning attached) or a sentence in a chat (no provenance). Neither is something the next step of the loop can build on, or something a reviewer can check.

What we have:

  • The memory service (RFC 0002) keeps a person's own entries with embeddings in its database and recalls them by similarity. It is strictly personal: every RPC a person reaches works on the verified caller's own entries, and an entry is at most 500 characters.
  • The RFCs product (RFC 0035) already knows how to keep our own accounts internal: spaces in an account listed in [access] internal_accounts exist only for the platform's admins; anyone else gets NOT_FOUND, never a listing.
  • The MCP server (RFC 0050) is an OAuth protected resource with scopes per tool, so a person can connect an assistant and grant it exactly what it may read.
  • The agent (RFC 0029) runs on our own host and reports over a session it dials out; RFC 0040.10 says it should discover the host's own model (devices, arrays, filesystems, which workload stores what) and keep it current.

What is missing is a store owned by a team rather than a person, with entries that say who wrote them, from where, on what evidence, and a gate that keeps it ours.

Proposal

An operations journal: entries owned by a team account, kept by the memory service, internal by default, written and read through the same surfaces as the rest of the platform. The first journal is the Engineering account's.

What the journal is not: incidents

Incidents and their postmortems (the record, the timeline, the evidence, publishing them, the monitors that opened them) live in the incidents service, which has its own RFC. The journal does not store an incident. It holds what a team learns that outlives one: the command that fixed it, the fact it surfaced about a host, the runbook for next time. Each of those links to the incident it came from by the incident's id, so a reader goes from a runbook to the incident that taught it, and a search by incident id lists what the team kept from it.

Entries

Three kinds:

Kind What it holds Example
command A reusable command as a template with named placeholders, never a value that is a secret smartctl -a {device}; du -sh {path}
fact A statement about a host or a system, observed or stated a device's model and temperature; an array's level and members; which workload reads and writes a volume, and how much
runbook Steps for a known situation, which may reference command entries "Root disk above 85 %"

Every entry carries:

  • Author: who or what wrote it. A person (their subject), an assistant acting for a person (the person's subject and the OAuth client id), the agent (its id), or one of our services. Stored as kind plus id, so the record says "Claude for Nevio" or "the agent on host X", never only "someone".
  • When: created_at, and for a fact observed_at, the time the observation was made.
  • Source: where it came from (cli, mcp, agent, console, hook).
  • Links: typed references, each one of rfc (0040.10), pr (inorbithr/core#251), incident (an incident's id in the incidents service), entry (another journal entry), evidence (an RFC 0046 record id) or url. A link to an incident is a reference only; the journal never copies the incident's text.
  • Evidence: what the entry rests on. A check result id, a command entry plus the sha256 of its output, a measurement with its unit, window and time. Never the output itself unless it passes the write check below.
  • Key (facts only): a stable name for the thing observed, such as host:<name>/disk/<serial>. A new fact with the same key supersedes the previous one; the superseded versions stay as history up to a bound, so a temperature or a usage figure has a series behind it.
  • Tags and a title.

A command entry's placeholders are declared ({device}, {age}) with a short description each. The journal stores the template, not a filled-in line. Running a command stays the person's act: iohr ops cmd show <name> --set device=… prints the filled line; nothing in the journal executes anything.

Secrets and customer data are refused at write

Every write passes a check before it is stored, embedded or logged. It works the way the Lab's redaction check does (docs/lab/rules.json): rules as data, regular expressions over one line plus a few structural checks, a conformance suite that pins what they find, and the same rules run on both sides.

  • Where it runs. In the memory service on every write, whoever the caller is. The command line and the session hook (phase 4) run the same rules first, so a secret normally never leaves the machine; the server's check is the one that counts.
  • What it refuses. Private key blocks; tokens with a known shape (provider prefixes, JWTs, bearer headers); credentials in a URL (scheme://user:password@); an assignment of a literal to a name that means a secret (password=, token=, secret=, api_key= and their flag forms such as --password) unless the value is a declared placeholder; long high-entropy strings; e-mail addresses, card numbers (by Luhn), IBANs and phone numbers. A command template must take every value after a secret-bearing flag as a placeholder.
  • How it answers. INVALID_ARGUMENT with each finding's rule id and position, never the matched text, so the refusal does not repeat the secret into an error envelope, a log or an assistant's context. The audit line records the rule ids.
  • What it does not do. It does not mask and keep. A write that fails is not stored in any form; the writer fixes it and writes again. The one exception is the session hook, which masks on the person's machine before anything is sent (phase 4).

Host details such as device names, mount points or internal addresses are allowed in an internal journal: they are the point of a fact. The Lab's public rules (addresses, ports, paths) do not apply, because nothing in the journal is published. The secret and personal data rules apply everywhere.

Internal by default

The journal follows the same rule as internal spaces in the RFCs product. The memory service's configuration lists InOrbit's own accounts ([ops] internal_accounts, the same ids as the labs service's list) and the role they need (internal_role, the platform's admin). For a journal owned by one of those accounts:

  • A caller with the platform's admin role reads and writes.
  • A person an admin approved reads, or reads and writes. Approvals are rows in the store (account, subject, read or write, who granted it, when, optional expiry), granted and revoked with iohr ops grant and iohr ops revoke by an admin, each one audited. They are not configuration, so approving a person needs no deploy and puts no person's id in a file.
  • Anyone else gets NOT_FOUND on every call, on a listing and on a direct id, the same answer as for an account that has no journal. No member, invitation, API key, OAuth app or MCP client of an unapproved person reaches it.

Journals for other accounts (customers) are not opened by this RFC. The store is built for them (every row has an account), but the service refuses any account outside the internal list until a later decision says what their rules are.

Write paths

  1. The agent on a host. Under RFC 0040.10's host:read capability the agent discovers the host's model (devices, buses, arrays, filesystems, temperatures, which workload stores what and how much it reads and writes) and reports it as a facts frame on its existing session. The agents service writes the facts as itself (svc:agents), naming the agent's account and id. The memory service accepts that service for fact entries only, with source agent and the agent as author. The agent never holds a token that reads the journal.
  2. Any session, a person or an assistant. Through the API, the command line (iohr ops note, iohr ops cmd add, iohr ops runbook), and MCP write tools. A write needs ops:write; over MCP it also needs mcp:write. An assistant writes as the person who connected it, and the entry records both.

Read paths, from day one

  1. Assistants over MCP. Tools ops_search (by meaning and by words, filtered by kind, tag, key or time) and ops_get (one entry with its links, evidence and, for a fact, its recent history), under mcp:read and a new ops:read scope. Runbooks are also offered as MCP resources once the gateway serves resources (RFC 0050, phase 3); until then ops_get reads them.
  2. Our own models. The model service (RFC 0002's tool table) gains ops_search and ops_get as tools that call the journal as the caller. An agent file names them; the console's assistant uses them first. They run on our own engines, so reading the journal through the site or console sends nothing to a third party. The site guide does not get them: an anonymous visitor would only ever see NOT_FOUND, and the tool in its file would be an invitation to probe.
  3. The command line. iohr ops search, iohr ops get, iohr ops ls.

Whether a third-party assistant may read is a grant, nothing else. The owner connects it and consents to ops:read; without that scope /mcp answers 403 insufficient_scope and the tools are not listed. Nothing is pushed to any model. An assistant reads what it asks for, when it asks, and every read is audited.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Search is hybrid. An embedding (the model service's embedding tier, as for memory) finds entries by meaning; the database's full-text search finds the exact words that embeddings handle badly, such as device names, error codes and unit names. The two rankings are merged by reciprocal rank fusion. Each answer is capped in count and bytes, says which part matched, and says when the embedding came from a stub engine.

Audit, retention, erasure

  • Audit. One line per call: caller, OAuth client if any, account, operation, kind, entry count, bytes, outcome and, on a refused write, the rule ids. Never the text, the query or a placeholder's description. Approvals and revocations also go to the account's audit log (the audit.event stream), since they change who may read.
  • Retention. Entries stay until deleted. Superseded fact versions are bounded per key (oldest go first). Deleting an entry deletes it; there is no soft delete. Deleting is a person's act through the API or iohr ops rm, never an MCP tool.
  • Erasure. The memory service's EraseSubject also covers the journal: entries a person authored in a team's journal stay (they are the team's), with the author replaced by erased:<sha256 of the subject>, and the person's approvals are deleted. ExportSubject includes the entries a person authored. Deleting a team account deletes its journal.

One deployment

Staging and production are the same deployment for now, so there is one journal. A fact names the host it is about; nothing in an entry depends on an environment split. When staging becomes its own deployment, its journal is its own too.

Where it sits in the loop

The journal is where the loop's results stay. Observe produces facts (the agent's and a person's), model reads them, a runbook or a command links to the incident (in the incidents service) that taught it, and a runbook links to the evidence record (RFC 0046) that showed it works. An entry is not evidence on its own: a fact is an observation with a time and a source, and its weight is the evidence it links to. A model's sentence with no evidence link is recorded as exactly that.

Plan

Phase 1 is the store, the command line and MCP reads. Phase 2 is the agent's facts, phase 3 our own models, phase 4 capture from sessions. Each phase is its own pull request (or a short series), with its line in the status log.

Phase 1: store, API, command line, MCP

Contract. A new package proto/iohr/ops/v1/ops.proto, service OpsService, served by the memory binary next to MemoryService:

  • Record (POST /v1/ops/entries): kind, title, body, key, placeholders, tags, links, evidence, optional account. Empty account is the caller's own (org), else the service's [ops] default_account for a caller who passes the internal gate.
  • Get (GET /v1/ops/entries/{id}): the entry, its links and a fact's last versions.
  • List (GET /v1/ops/entries): paged per RFC 0033, filtered by kind, tag, key, time.
  • Search (POST /v1/ops/search): query, filters, k.
  • Delete (DELETE /v1/ops/entries/{id}).
  • Grant, Revoke, ListReaders (/v1/ops/readers): admin only.
  • RecordFacts: no route, svc:agents only (phase 2, declared now so the contract is complete).

The memory service's own EraseSubject and ExportSubject are extended rather than duplicated; the accounts sweeper keeps calling one service. buf lint holds the package to proto/iohr/ops/v1/. The gateway serves the package as its own backend, ops, at the memory service's address (the gateway shares channels to one URL), so the MCP tools are ops_search, ops_get, ops_record.

Migration. mise run db:new ops_journal: schema ops, granted to the memory service's role.

  • ops.entries: id, account, kind (check in the three), title, body, key (null except for facts), placeholders (jsonb), tags (text[]), links (jsonb), evidence (jsonb), source, author kind and id, client id, created and observed times, superseded by, a generated tsvector with its GIN index, embedding vector(1024). Indexes on (account, created_at desc) and (account, key, observed_at desc).
  • ops.readers: account, subject, access (read or write), granted by, granted at, expires at; primary key (account, subject).

Exact scans of one account's rows, as memory does; an HNSW index waits for counts that need it.

Service (crates/memory):

  • src/ops/mod.rs, service.rs (the OpsService impl, admit and audit as in service.rs), store.rs (every query takes the account), access.rs (the internal gate and approvals, modelled on crates/labs/src/service.rs hidden and allowed), search.rs (hybrid ranking), check.rs (the write check).
  • src/config.rs: [ops] with internal_accounts, internal_role, default_account, limits (title, body per kind with runbooks the largest, tags, links, entries per account, versions per key, k, answer bytes), and the services allowed to record facts.
  • CLAUDE.md and docs/memory/README.md gain the journal's invariants.

The write check. Rules as data in docs/ops/rules.json, the same format as docs/lab/rules.json, with cases in docs/ops/conformance/. The line checker in crates/labs/src/checks.rs (compile, check) moves to tbd-common as a transport-free module both services use, so there is one implementation of "run these rules over this text"; the structural checks (Luhn, IBAN, entropy, placeholders after secret flags) live next to it. The rules file is served on the docs site for the command line, as the Lab's is.

Gateway, edge, scopes.

  • The gateway's configuration: [services.ops] (the memory service's address), [mcp.tools] entries ops_search and ops_get (mcp:read, scopes = ["ops:read"], read only, closed world) and ops_record (mcp:write, scopes = ["ops:write"], not destructive, not idempotent), [openapi.scopes] for both.
  • The accounts service's scope catalogue: ops:read and ops:write; consent (ui/auth) may grant them; the protected resource metadata lists them.
  • The edge: /v1/ops gated like every other API route; internal gRPC calls for iohr.ops.v1.OpsService reach the memory service.
  • PROTOCOL_OPS_URL registered where every backend URL is.

Command line (inorbithr/sdk, cli/crates/iohr/src/commands/ops.rs): note, cmd add|ls|show, runbook, search, get, ls, rm, grant, revoke. Each write runs the rules locally first and says which rule and where, without printing the match.

Tests.

  • crates/memory/tests/it/ops.rs on the shared test database and the llm's stub engine: an admin records and finds an entry by meaning and by an exact device name; a non-admin, a member of the account, an API key bound to the account and an OAuth client of an unapproved person each get NOT_FOUND on list, get and search; an approved reader reads and cannot write; an expired approval reads nothing; svc:agents writes facts and nothing else; a fact with the same key supersedes and keeps a bounded history; every conformance case is refused with its rule id and the error and audit line hold no part of the matched text; an entry's text never reaches a log; EraseSubject pseudonymises authorship and removes approvals; ExportSubject includes authored entries.
  • crates/protocol/tests/it/mcp.rs: the three tools listed only to a caller holding their scopes, 403 insufficient_scope without ops:read, annotations as configured.
  • The conformance suite runs in both checkers (service and command line).

Phase 2: the agent's facts

inorbithr/dataplane: under host:read, a discovery pass at start and on an interval (devices and their health and temperature, arrays and their state, filesystems and usage, which workload's volumes sit where, read and write rates per device), sent as a facts frame. Keys are stable per host. crates/agents: accept the frame, bound it, call RecordFacts as svc:agents. Tests: a frame over the limits is refused; a fact from an agent of another account cannot land in ours.

Phase 3: our own models

crates/llm/src/tools.rs: targets for ops_search and ops_get calling OpsService as the caller, [tools] ops_url (LLM_OPS_URL, registered everywhere a variable is). A console agent file names them; the site guide's does not. A test refuses a site guide file that names them. The answer is cut to [agents] max_result_bytes like every tool.

Phase 4: commands captured from sessions

An opt-in hook for a person's shell or an assistant's session (a Claude Code PostToolUse hook on its shell tool is the first) pipes the command line, never its output, to iohr ops cmd capture. The command line masks it with the same rules, replaces literals after secret-bearing flags with placeholders, and posts it as a command entry marked as captured, for a person to name and keep or delete. Off unless the person turns it on per repository. Offline, captures queue locally, bounded, and are masked before they are queued.

Compliance notes

This adds a new store of operational data and a new read path to third-party assistants. In the compliance repository: the controls for access control (the internal gate, approvals), audit logging and data classification gain the journal; the vendor register must name each assistant provider before the owner grants it ops:read. We align with the criteria; nothing here is a claim of compliance.

Alternatives considered

  • An assistant's own memory. What we do today by accident. The vendor holds it, we cannot audit reads, other assistants and our own models cannot use it, and deleting it is the vendor's process. The owner ruled it out.
  • Files in a repository (a runbooks/ directory, a facts file the agent commits). Versioned and reviewable, and good for runbooks we publish. Bad for facts that change every few minutes, unsearchable by meaning, and readable by anyone with the repository, with no per-read audit. A runbook that should be public can still become an RFC or a doc page and link back.
  • A new service. Clean boundaries, and one more deployment, database role, chaos adapter and health check for a store that is the memory service's store with an owner column and a few more fields. The memory service already owns embeddings, recall, erasure and export. A second gRPC service in the same binary keeps the personal notebook's invariants untouched (its schema and RPCs do not change) while sharing everything else. If the journal grows past that, it moves out behind the same contract.
  • The RFCs product's spaces. Already team-owned and internal by default. Its unit is a long, versioned document under review, not a fact observed a minute ago; facts and commands would be documents by force.
  • Mask secrets and keep the entry. Friendlier, and a masked secret is still a hint (its length, its prefix, where it was used). Refusing teaches the writer to use a placeholder, and leaves nothing to leak. Masking stays on the person's own machine (phase 4).
  • Push the journal into an assistant's context (a project file, a system prompt). Every session would carry all of it to the vendor whether it needed it or not. Reads on demand, under a scope, audited, send only what a question needs.

Decision

Open. Three questions for the owner, each answered with one word; the recommendation is first:

  1. Who may read an internal journal besides admins? Rows (recommended): approvals are rows an admin grants and revokes, audited, no deploy. Or config: a list of people in the service's configuration.
  2. Does a third-party assistant need more than the ops:read grant? No (recommended): the owner granting ops:read is the decision. Or allowlist: also a list of OAuth clients allowed on internal accounts.
  3. May the public site guide ever read the journal? Never (recommended): an anonymous visitor would only see NOT_FOUND, and the tool would invite probing. Or later: decide again when the journal has public parts.

Proposed with those answers: the journal as above, phase 1 first.

Publication

The Lab shows this RFC under the platform. The developer documentation gains an operations journal page (the contract, scopes, the write check and its rules) and the MCP page lists the tools; the changelog says when each phase shipped. Nothing about the journal's content is published.

Status log

  • 2026-10-04: opened, with the owner's requirements of the same day and the plan. Nothing built.
  • 2026-10-05: the boundary with the incidents service agreed. Incidents and postmortems live there; the journal keeps commands, facts and runbooks and links to incidents by id. The incident entry kind is gone, and the three open questions are framed for a one-word answer each.
  • 2026-10-07: Checked: nothing built (#257 and #411 are the text). Open.

← Back to Platform