Problem
What we know about running this platform lives in the wrong places. The commands that fixed the last disk-pressure outage are in one person's shell history. Which storage array holds which workload's data, which device runs hot, which service writes the most: an engineer finds out by hand during an incident, and the next session, a person's or an assistant's, finds out again. Some of it ends up in an assistant's own memory, on a vendor's servers, where we cannot query it, audit who read it or delete it.
The owner put the requirement plainly on 2026-10-04: our own system to store this and read it back, not handed to the assistant's vendor, with the end goal that an assistant (Claude among them) can query it on its own, and so can the site's and console's own models, from the start.
The loop of RFC 0045 needs the same thing. Observe, model and evidence all produce knowledge about a running system that outlives the session that found it. Today that knowledge is either a log line (no meaning attached) or a sentence in a chat (no provenance). Neither is something the next step of the loop can build on, or something a reviewer can check.
What we have:
- The memory service (RFC 0002) keeps a person's own entries with embeddings in its database and recalls them by similarity. It is strictly personal: every RPC a person reaches works on the verified caller's own entries, and an entry is at most 500 characters.
- The RFCs product (RFC 0035) already knows how to keep our own accounts internal:
spaces in an account listed in
[access] internal_accountsexist only for the platform's admins; anyone else getsNOT_FOUND, never a listing. - The MCP server (RFC 0050) is an OAuth protected resource with scopes per tool, so a person can connect an assistant and grant it exactly what it may read.
- The agent (RFC 0029) runs on our own host and reports over a session it dials out; RFC 0040.10 says it should discover the host's own model (devices, arrays, filesystems, which workload stores what) and keep it current.
What is missing is a store owned by a team rather than a person, with entries that say who wrote them, from where, on what evidence, and a gate that keeps it ours.
Proposal
An operations journal: entries owned by a team account, kept by the memory service, internal by default, written and read through the same surfaces as the rest of the platform. The first journal is the Engineering account's.
What the journal is not: incidents
Incidents and their postmortems (the record, the timeline, the evidence, publishing them, the monitors that opened them) live in the incidents service, which has its own RFC. The journal does not store an incident. It holds what a team learns that outlives one: the command that fixed it, the fact it surfaced about a host, the runbook for next time. Each of those links to the incident it came from by the incident's id, so a reader goes from a runbook to the incident that taught it, and a search by incident id lists what the team kept from it.
Entries
Three kinds:
| Kind | What it holds | Example |
|---|---|---|
command |
A reusable command as a template with named placeholders, never a value that is a secret | smartctl -a {device}; du -sh {path} |
fact |
A statement about a host or a system, observed or stated | a device's model and temperature; an array's level and members; which workload reads and writes a volume, and how much |
runbook |
Steps for a known situation, which may reference command entries |
"Root disk above 85 %" |
Every entry carries:
- Author: who or what wrote it. A person (their subject), an assistant acting for a person (the person's subject and the OAuth client id), the agent (its id), or one of our services. Stored as kind plus id, so the record says "Claude for Nevio" or "the agent on host X", never only "someone".
- When:
created_at, and for a factobserved_at, the time the observation was made. - Source: where it came from (
cli,mcp,agent,console,hook). - Links: typed references, each one of
rfc(0040.10),pr(inorbithr/core#251),incident(an incident's id in the incidents service),entry(another journal entry),evidence(an RFC 0046 record id) orurl. A link to an incident is a reference only; the journal never copies the incident's text. - Evidence: what the entry rests on. A check result id, a command entry plus the sha256 of its output, a measurement with its unit, window and time. Never the output itself unless it passes the write check below.
- Key (facts only): a stable name for the thing observed, such as
host:<name>/disk/<serial>. A new fact with the same key supersedes the previous one; the superseded versions stay as history up to a bound, so a temperature or a usage figure has a series behind it. - Tags and a title.
A command entry's placeholders are declared ({device}, {age}) with a short
description each. The journal stores the template, not a filled-in line. Running a
command stays the person's act: iohr ops cmd show <name> --set device=… prints the
filled line; nothing in the journal executes anything.
Secrets and customer data are refused at write
Every write passes a check before it is stored, embedded or logged. It works the way the
Lab's redaction check does (docs/lab/rules.json): rules as data, regular expressions
over one line plus a few structural checks, a conformance suite that pins what they find,
and the same rules run on both sides.
- Where it runs. In the memory service on every write, whoever the caller is. The command line and the session hook (phase 4) run the same rules first, so a secret normally never leaves the machine; the server's check is the one that counts.
- What it refuses. Private key blocks; tokens with a known shape (provider prefixes,
JWTs, bearer headers); credentials in a URL (
scheme://user:password@); an assignment of a literal to a name that means a secret (password=,token=,secret=,api_key=and their flag forms such as--password) unless the value is a declared placeholder; long high-entropy strings; e-mail addresses, card numbers (by Luhn), IBANs and phone numbers. Acommandtemplate must take every value after a secret-bearing flag as a placeholder. - How it answers.
INVALID_ARGUMENTwith each finding's rule id and position, never the matched text, so the refusal does not repeat the secret into an error envelope, a log or an assistant's context. The audit line records the rule ids. - What it does not do. It does not mask and keep. A write that fails is not stored in any form; the writer fixes it and writes again. The one exception is the session hook, which masks on the person's machine before anything is sent (phase 4).
Host details such as device names, mount points or internal addresses are allowed in an internal journal: they are the point of a fact. The Lab's public rules (addresses, ports, paths) do not apply, because nothing in the journal is published. The secret and personal data rules apply everywhere.
Internal by default
The journal follows the same rule as internal spaces in the RFCs product. The memory
service's configuration lists InOrbit's own accounts ([ops] internal_accounts, the
same ids as the labs service's list) and the role they need (internal_role, the
platform's admin). For a journal owned by one of those accounts:
- A caller with the platform's
adminrole reads and writes. - A person an admin approved reads, or reads and writes. Approvals are rows in the
store (account, subject,
readorwrite, who granted it, when, optional expiry), granted and revoked withiohr ops grantandiohr ops revokeby an admin, each one audited. They are not configuration, so approving a person needs no deploy and puts no person's id in a file. - Anyone else gets
NOT_FOUNDon every call, on a listing and on a direct id, the same answer as for an account that has no journal. No member, invitation, API key, OAuth app or MCP client of an unapproved person reaches it.
Journals for other accounts (customers) are not opened by this RFC. The store is built for them (every row has an account), but the service refuses any account outside the internal list until a later decision says what their rules are.
Write paths
- The agent on a host. Under RFC 0040.10's
host:readcapability the agent discovers the host's model (devices, buses, arrays, filesystems, temperatures, which workload stores what and how much it reads and writes) and reports it as afactsframe on its existing session. The agents service writes the facts as itself (svc:agents), naming the agent's account and id. The memory service accepts that service forfactentries only, with sourceagentand the agent as author. The agent never holds a token that reads the journal. - Any session, a person or an assistant. Through the API, the command line
(
iohr ops note,iohr ops cmd add,iohr ops runbook), and MCP write tools. A write needsops:write; over MCP it also needsmcp:write. An assistant writes as the person who connected it, and the entry records both.
Read paths, from day one
- Assistants over MCP. Tools
ops_search(by meaning and by words, filtered by kind, tag, key or time) andops_get(one entry with its links, evidence and, for a fact, its recent history), undermcp:readand a newops:readscope. Runbooks are also offered as MCP resources once the gateway serves resources (RFC 0050, phase 3); until thenops_getreads them. - Our own models. The model service (RFC 0002's tool table) gains
ops_searchandops_getas tools that call the journal as the caller. An agent file names them; the console's assistant uses them first. They run on our own engines, so reading the journal through the site or console sends nothing to a third party. The site guide does not get them: an anonymous visitor would only ever seeNOT_FOUND, and the tool in its file would be an invitation to probe. - The command line.
iohr ops search,iohr ops get,iohr ops ls.
Whether a third-party assistant may read is a grant, nothing else. The owner connects it
and consents to ops:read; without that scope /mcp answers 403 insufficient_scope
and the tools are not listed. Nothing is pushed to any model. An assistant reads what it
asks for, when it asks, and every read is audited.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Search
Search is hybrid. An embedding (the model service's embedding tier, as for memory) finds entries by meaning; the database's full-text search finds the exact words that embeddings handle badly, such as device names, error codes and unit names. The two rankings are merged by reciprocal rank fusion. Each answer is capped in count and bytes, says which part matched, and says when the embedding came from a stub engine.
Audit, retention, erasure
- Audit. One line per call: caller, OAuth client if any, account, operation, kind,
entry count, bytes, outcome and, on a refused write, the rule ids. Never the text,
the query or a placeholder's description. Approvals and revocations also go to the
account's audit log (the
audit.eventstream), since they change who may read. - Retention. Entries stay until deleted. Superseded fact versions are bounded per
key (oldest go first). Deleting an entry deletes it; there is no soft delete.
Deleting is a person's act through the API or
iohr ops rm, never an MCP tool. - Erasure. The memory service's
EraseSubjectalso covers the journal: entries a person authored in a team's journal stay (they are the team's), with the author replaced byerased:<sha256 of the subject>, and the person's approvals are deleted.ExportSubjectincludes the entries a person authored. Deleting a team account deletes its journal.
One deployment
Staging and production are the same deployment for now, so there is one journal. A fact names the host it is about; nothing in an entry depends on an environment split. When staging becomes its own deployment, its journal is its own too.
Where it sits in the loop
The journal is where the loop's results stay. Observe produces facts (the agent's and a person's), model reads them, a runbook or a command links to the incident (in the incidents service) that taught it, and a runbook links to the evidence record (RFC 0046) that showed it works. An entry is not evidence on its own: a fact is an observation with a time and a source, and its weight is the evidence it links to. A model's sentence with no evidence link is recorded as exactly that.
Plan
Phase 1 is the store, the command line and MCP reads. Phase 2 is the agent's facts, phase 3 our own models, phase 4 capture from sessions. Each phase is its own pull request (or a short series), with its line in the status log.
Phase 1: store, API, command line, MCP
Contract. A new package proto/iohr/ops/v1/ops.proto, service OpsService, served
by the memory binary next to MemoryService:
Record(POST /v1/ops/entries): kind, title, body, key, placeholders, tags, links, evidence, optional account. Empty account is the caller's own (org), else the service's[ops] default_accountfor a caller who passes the internal gate.Get(GET /v1/ops/entries/{id}): the entry, its links and a fact's last versions.List(GET /v1/ops/entries): paged per RFC 0033, filtered by kind, tag, key, time.Search(POST /v1/ops/search): query, filters,k.Delete(DELETE /v1/ops/entries/{id}).Grant,Revoke,ListReaders(/v1/ops/readers): admin only.RecordFacts: no route,svc:agentsonly (phase 2, declared now so the contract is complete).
The memory service's own EraseSubject and ExportSubject are extended rather than
duplicated; the accounts sweeper keeps calling one service. buf lint holds the package
to proto/iohr/ops/v1/. The gateway serves the package as its own backend, ops, at
the memory service's address (the gateway shares channels to one URL), so the MCP tools
are ops_search, ops_get, ops_record.
Migration. mise run db:new ops_journal: schema ops, granted to the memory
service's role.
ops.entries: id, account, kind (check in the three), title, body, key (null except for facts), placeholders (jsonb), tags (text[]), links (jsonb), evidence (jsonb), source, author kind and id, client id, created and observed times, superseded by, a generatedtsvectorwith its GIN index,embedding vector(1024). Indexes on (account, created_at desc) and (account, key, observed_at desc).ops.readers: account, subject, access (readorwrite), granted by, granted at, expires at; primary key (account, subject).
Exact scans of one account's rows, as memory does; an HNSW index waits for counts that need it.
Service (crates/memory):
src/ops/mod.rs,service.rs(theOpsServiceimpl,admitandauditas inservice.rs),store.rs(every query takes the account),access.rs(the internal gate and approvals, modelled oncrates/labs/src/service.rshiddenandallowed),search.rs(hybrid ranking),check.rs(the write check).src/config.rs:[ops]withinternal_accounts,internal_role,default_account, limits (title, body per kind with runbooks the largest, tags, links, entries per account, versions per key,k, answer bytes), and the services allowed to record facts.CLAUDE.mdanddocs/memory/README.mdgain the journal's invariants.
The write check. Rules as data in docs/ops/rules.json, the same format as
docs/lab/rules.json, with cases in docs/ops/conformance/. The line checker in
crates/labs/src/checks.rs (compile, check) moves to tbd-common as a transport-free
module both services use, so there is one implementation of "run these rules over this
text"; the structural checks (Luhn, IBAN, entropy, placeholders after secret flags) live
next to it. The rules file is served on the docs site for the command line, as the Lab's
is.
Gateway, edge, scopes.
- The gateway's configuration:
[services.ops](the memory service's address),[mcp.tools]entriesops_searchandops_get(mcp:read,scopes = ["ops:read"], read only, closed world) andops_record(mcp:write,scopes = ["ops:write"], not destructive, not idempotent),[openapi.scopes]for both. - The accounts service's scope catalogue:
ops:readandops:write; consent (ui/auth) may grant them; the protected resource metadata lists them. - The edge:
/v1/opsgated like every other API route; internal gRPC calls foriohr.ops.v1.OpsServicereach the memory service. PROTOCOL_OPS_URLregistered where every backend URL is.
Command line (inorbithr/sdk, cli/crates/iohr/src/commands/ops.rs): note,
cmd add|ls|show, runbook, search, get, ls, rm, grant, revoke.
Each write runs the rules locally first and says which rule and where, without
printing the match.
Tests.
crates/memory/tests/it/ops.rson the shared test database and the llm's stub engine: an admin records and finds an entry by meaning and by an exact device name; a non-admin, a member of the account, an API key bound to the account and an OAuth client of an unapproved person each getNOT_FOUNDon list, get and search; an approved reader reads and cannot write; an expired approval reads nothing;svc:agentswrites facts and nothing else; a fact with the same key supersedes and keeps a bounded history; every conformance case is refused with its rule id and the error and audit line hold no part of the matched text; an entry's text never reaches a log;EraseSubjectpseudonymises authorship and removes approvals;ExportSubjectincludes authored entries.crates/protocol/tests/it/mcp.rs: the three tools listed only to a caller holding their scopes,403 insufficient_scopewithoutops:read, annotations as configured.- The conformance suite runs in both checkers (service and command line).
Phase 2: the agent's facts
inorbithr/dataplane: under host:read, a discovery pass at start and on an interval
(devices and their health and temperature, arrays and their state, filesystems and
usage, which workload's volumes sit where, read and write rates per device), sent as a
facts frame. Keys are stable per host. crates/agents: accept the frame, bound it,
call RecordFacts as svc:agents. Tests: a frame over the limits is refused; a fact
from an agent of another account cannot land in ours.
Phase 3: our own models
crates/llm/src/tools.rs: targets for ops_search and ops_get calling OpsService
as the caller, [tools] ops_url (LLM_OPS_URL, registered everywhere a variable is). A
console agent file names them; the site guide's does not. A test refuses a site guide
file that names them. The answer is cut to [agents] max_result_bytes like every tool.
Phase 4: commands captured from sessions
An opt-in hook for a person's shell or an assistant's session (a Claude Code
PostToolUse hook on its shell tool is the first) pipes the command line, never its
output, to iohr ops cmd capture. The command line masks it with the same rules,
replaces literals after secret-bearing flags with placeholders, and posts it as a
command entry marked as captured, for a person to name and keep or delete. Off unless
the person turns it on per repository. Offline, captures queue locally, bounded, and
are masked before they are queued.
Compliance notes
This adds a new store of operational data and a new read path to third-party
assistants. In the compliance repository: the controls for access control (the internal
gate, approvals), audit logging and data classification gain the journal; the vendor
register must name each assistant provider before the owner grants it ops:read. We
align with the criteria; nothing here is a claim of compliance.
Alternatives considered
- An assistant's own memory. What we do today by accident. The vendor holds it, we cannot audit reads, other assistants and our own models cannot use it, and deleting it is the vendor's process. The owner ruled it out.
- Files in a repository (a
runbooks/directory, a facts file the agent commits). Versioned and reviewable, and good for runbooks we publish. Bad for facts that change every few minutes, unsearchable by meaning, and readable by anyone with the repository, with no per-read audit. A runbook that should be public can still become an RFC or a doc page and link back. - A new service. Clean boundaries, and one more deployment, database role, chaos adapter and health check for a store that is the memory service's store with an owner column and a few more fields. The memory service already owns embeddings, recall, erasure and export. A second gRPC service in the same binary keeps the personal notebook's invariants untouched (its schema and RPCs do not change) while sharing everything else. If the journal grows past that, it moves out behind the same contract.
- The RFCs product's spaces. Already team-owned and internal by default. Its unit is a long, versioned document under review, not a fact observed a minute ago; facts and commands would be documents by force.
- Mask secrets and keep the entry. Friendlier, and a masked secret is still a hint (its length, its prefix, where it was used). Refusing teaches the writer to use a placeholder, and leaves nothing to leak. Masking stays on the person's own machine (phase 4).
- Push the journal into an assistant's context (a project file, a system prompt). Every session would carry all of it to the vendor whether it needed it or not. Reads on demand, under a scope, audited, send only what a question needs.
Decision
Open. Three questions for the owner, each answered with one word; the recommendation is first:
- Who may read an internal journal besides admins? Rows (recommended): approvals are rows an admin grants and revokes, audited, no deploy. Or config: a list of people in the service's configuration.
- Does a third-party assistant need more than the
ops:readgrant? No (recommended): the owner grantingops:readis the decision. Or allowlist: also a list of OAuth clients allowed on internal accounts. - May the public site guide ever read the journal? Never (recommended): an
anonymous visitor would only see
NOT_FOUND, and the tool would invite probing. Or later: decide again when the journal has public parts.
Proposed with those answers: the journal as above, phase 1 first.
Publication
The Lab shows this RFC under the platform. The developer documentation gains an operations journal page (the contract, scopes, the write check and its rules) and the MCP page lists the tools; the changelog says when each phase shipped. Nothing about the journal's content is published.
Status log
- 2026-10-04: opened, with the owner's requirements of the same day and the plan. Nothing built.
- 2026-10-05: the boundary with the incidents service agreed. Incidents and postmortems
live there; the journal keeps commands, facts and runbooks and links to incidents by
id. The
incidententry kind is gone, and the three open questions are framed for a one-word answer each. - 2026-10-07: Checked: nothing built (#257 and #411 are the text). Open.