Implements: PRD 0001
Problem
PRD 0001 (InOrbit Atlas, approved 2026-10-07) asks for two things this RFC designs.
Templates. Evaluating a Rust blockchain client, a multi-service SaaS platform and a twenty-year-old monolith should not be three projects. The PRD lists what a template sets (sources, map shape, question bank, document skeletons, checks), names the first ones (Rust node client for reth, multi-service SaaS for InOrbit, legacy monolith, generic) and says every evaluation records the template version it used. It does not say what a template is on disk, who may change it, or how a version is pinned.
A grade. The PRD's main risk is "plausible but wrong documents", and its answer is to grade Atlas on the one system whose truth we know: our own. Its success measure is the share of Atlas's statements about InOrbit that our RFCs and code confirm, and the share they contradict, per template version. Nothing says yet what counts as truth, how a statement is matched to it, who decides a contradiction, or how a grade can be run again and give the same answer.
Two constraints make both harder:
- Every language and access point from the start (owner, 2026-10-07). A template cannot pick its languages. Every one InOrbit has an SDK for or is building one for, and every access point InOrbit's own API gateway serves, is read from the first release. Lists like that drift: the SDK gains a language, the gateway gains a surface, and Atlas's readers fall behind without anyone noticing.
- Our truth is not clean. Some of our RFCs describe things that are not built. Some status logs lag the code. On 2026-10-07 a pass over 96 RFCs found several whose text and code disagreed. A grade that treats every RFC sentence as truth would punish Atlas for being right about our own stale documents.
This RFC covers the templates and the grade. It does not design the document kinds (RFC 0081), how the agent reads (Atlas A2), the agent's configuration (RFC 0082, being written), the system map's model (Atlas A4) or how questions are routed (Atlas A5). It uses each of them through the interface it names.
Proposal
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
1. A template is one YAML file with a published schema
A template is a declarative YAML file checked against a versioned JSON Schema
(https://inorbit.hr/atlas/template/v1), the same convention the agent's configuration
uses (RFC 0082): editors complete it, and iohr atlas template check <file> explains each
error with its line. A template holds no code. Its only commands are the checks in
section 1.6, as argument lists that the customer's own agent policy may refuse.
Top level:
| Key | What it holds |
|---|---|
schema |
atlas.template/v1. A new major schema is a new key value; old templates keep validating against the schema they name. |
id |
Stable slug: rust-node-client, multi-service-saas, legacy-monolith, generic. |
version |
Semver (section 1.8). |
extends |
Another template by id@version. Every template but generic extends generic@<version>; lists merge by id, a child entry replaces its parent's. |
applies_when |
Signals that suggest this template: files present, manifest fields, directory shapes. Atlas proposes the best match; the owner confirms the pick before reading starts. generic always matches. |
sources |
What to read and in what order. |
map |
The map's expected shape: layers, node kinds, required edges. |
questions |
The question bank. |
documents |
Document skeletons per kind. |
checks |
Builds, tests, lints, benchmarks and probes the agent may run, with consent. |
evidence |
What each kind of statement needs before it may be written, and before it may be accepted. |
Languages and access points are not keys of a template. They come from the coverage
catalogue (section 2), which every template reads in full. A template may add hints
(a language to read first, a framework to look for) but the schema has no way to switch
a language or an access point off.
1.1 Sources
Each source has an id, a kind (repository_paths, docs_site, book, ci,
releases, benchmarks, changelog, connection), match rules (globs, a manifest
field, a site path), a priority and required: true|false. A required source that is
missing is a finding ("no CI configuration found"), never a silent gap. Connections
(tickets, incidents, chat) are named by the catalogue kind RFC 0044 uses; the template
says which ones help, the customer decides which ones are granted.
1.2 Map shape
map.layers names the layers a system of this type has, each with match rules over what
the readers found (crate or package names, paths, imports, manifest fields). map.expect
lists nodes and edges a system of this type should have ("an RPC surface", "a storage
layer reached from execution"). Something expected and not found becomes a question, not
an assumption. The map's data model is Atlas A4's; the template only sets the expected
shape and the layer names.
1.3 Question bank
Each question has an id, the text, kind (intent, decision, ownership,
security_privacy, cost, status), an answer type (yes_no, choice, free,
person), a when condition over the map ("more than one storage backend found"), and
settles: which statements or map nodes the answer confirms or changes. The kind is
what Atlas A5's routing policies route by; the template never names people. A question
that would ask for a secret value is rejected by the schema's linter (it may ask where a
secret is kept, never what it is).
1.4 Document skeletons
Per document kind of RFC 0081 (prd, decision, rfc), a skeleton says how many to
expect and what each one covers: for_each (a product surface, a major dependency, a
layer), a title pattern, and per required section a look_for list. The required
sections themselves are RFC 0081's per-kind template; a skeleton never redefines them,
it says where in this type of system to find their content. RFC kinds that fit this type
are listed under rfc.kinds (for a node client: a sync performance change, a storage
migration, a new network upgrade).
1.5 Evidence expectations
The PRD's three states (inferred, confirmed, measured) and its rules hold for every template. A template adds what is specific to the type, per statement class:
| Field | Meaning |
|---|---|
min_sources |
How many independent sources an inferred statement needs ("read in 2 places"). |
accept_needs |
The state a statement must reach before its document can be accepted: confirmed for intent in a PRD, for example. |
measured_only |
Statement classes that may only be written as measured: throughput, latency, resource use, sync time. With no measurement they are not written at all, only listed as "not measured". |
reason_sources |
For a decision record: what may count as its reason (a commit message, a doc page, an owner's answer). With none, the record says "reason not found", never a guessed one. |
1.6 Checks
Each check has an id, a kind (build, test, lint, bench, probe), run as an
argument list (no shell), a working directory, a timeout, network: none unless the check
says why it needs more, the plan tier it needs (PRD: Personal builds and tests, Team adds
benchmarks), and turns: the statement classes a pass or a measurement makes measured.
Checks run only in the agent's sandbox, only after the customer consents to the exact
list shown with its commands, and only within the plan's budget. The agent's policy
(RFC 0082) wins over the template: a check it denies is skipped and named in the report.
A probe reads a running system the customer names (a health endpoint, a public API
description) and never writes.
1.7 Identity and pinning
A released template version is immutable. Its digest is the SHA-256 of the file's canonical JSON form (RFC 8785). Every evaluation records:
template: id, version, digest, and the ids, versions and digests of the templates it extends;catalogue: the coverage catalogue's version and digest (section 2);- the agent's version and its readers' versions.
The evaluation room and every exported document show the template id and version. A grade (section 4) is always about one template version.
1.8 Versions
- Major: the map shape or the statement classes change, so grades are no longer comparable with the previous major.
- Minor: questions, skeletons, sources or checks added or changed; grades stay comparable.
- Patch: wording only.
A new version of a template we grade (section 4.6) is released only after its grade on the same truth snapshot is at least as good as the version it replaces.
2. Languages and access points, from one list
The owner decided on 2026-10-07 that every template reads every language and every kind of access point InOrbit supports, from the first release. This section makes that a list a machine keeps, not a sentence someone has to remember.
The coverage catalogue is one generated JSON file with two parts.
- Languages come from the SDK generator's own list, the
Languageenum in the SDK repository's code generator. Today it holds the six languages it renders: Rust, TypeScript (with JavaScript), Python, Go, Java and C#. It already models a language that is known and not yet rendered (is_built). The PRD's list adds the languages InOrbit is building SDKs for: Dart (its runtime has started, outside the generator), Swift, Kotlin (on the JVM, read with Java's build files plus its own), C and C++, and Zig through the C layer. They join the enum as not-built entries, so the generator refuses to render them, as it does now for anything unbuilt, and the catalogue lists them.iohr sdk languages --jsonprints the list with each entry's state. - Access points come from the API gateway's own surface list. Today that list is a
table in a code comment. It becomes data in the gateway's crate, printed by a
subcommand into a JSON file beside the OpenAPI document and checked in CI the same way
(the file must equal what the code prints). It holds REST, server-sent events, the
WebSocket bridges, MQTT over WebSocket, GraphQL, gRPC and MCP. The SDK repository
already syncs the OpenAPI document from the platform into its
spec/folder; the surface file travels with it. Three kinds the PRD names are not gateway surfaces and are added by the catalogue's generator, each citing PRD 0001: webhooks (the platform's event deliveries, RFC 0022), queues, and scheduled jobs.
iohr atlas coverage --json builds the catalogue from those two inputs. CI in the SDK
repository regenerates it and fails when the checked-in file differs, so a new language
or surface cannot land without changing the catalogue in the same pull request.
Readers follow the catalogue. The agent (Atlas A2) is built against a catalogue version. Its CI fails when an entry has no reader, and each reader is held to a fixture: a small repository per language and per access point whose expected map is written by hand. Per language that covers what the PRD asks: index the code, resolve build and dependency files, find entry points, tests and configuration, and build and test in the sandbox when consent allows. Per access point: where it is defined, who calls it, how it is protected. The fixtures are the smallest truth sets in this RFC, and the readers' results on them are graded like any other evaluation (section 4).
Anything else is mapped and named. Code in a language outside the catalogue is still mapped: files, size, where it sits, what calls it when the caller is in a covered language. The report lists it under "Read at map level only" with its file count and size, and no statement about its behaviour is written. Nothing is skipped silently.
3. The first four templates
All four are written by us, kept in one folder with the schema and the validator
(section 5), and released as 1.0.0 when Phase 0 starts.
3.1 Rust node client (rust-node-client), first on reth
Shapes checked on 2026-10-07 against
paradigmxyz/reth at 1e8eb0c18202
(Apache-2.0 or MIT): a Cargo workspace with crates/net, crates/stages,
crates/storage, crates/rpc, crates/engine, crates/evm, crates/transaction-pool,
crates/consensus, crates/exex and crates/node, a book under docs/vocs, CI
workflows for unit, integration, sync and benchmark runs, and a Makefile with build,
test-unit and lint targets. The template's patterns are tested against a pinned reth
commit (slice 3); a pattern that stops matching is a template defect.
schema: atlas.template/v1
id: rust-node-client
version: 1.0.0
extends: generic@1.0.0
applies_when:
all:
- file: Cargo.toml
has: "[workspace]"
any:
- crate_name_matches: "*-network|*-p2p|*-net"
- crate_name_matches: "*-rpc*"
- path: "crates/stages/**"
sources:
- { id: workspace, kind: repository_paths, match: ["Cargo.toml", "crates/*/Cargo.toml", "crates/**/Cargo.toml"], required: true, priority: 1 }
- { id: book, kind: book, match: ["docs/**", "book/**"], required: false, priority: 2 }
- { id: ci, kind: ci, match: [".github/workflows/*.yml"], required: true, priority: 3 }
- { id: benches, kind: benchmarks, match: ["**/benches/**", "bin/*bench*/**"], required: false, priority: 4 }
- { id: releases, kind: releases, required: false, priority: 5 }
- { id: changelog, kind: changelog, match: ["CHANGELOG.md", "docs/**/release*.md"], required: false, priority: 5 }
map:
layers:
- { id: networking, match: { paths: ["crates/net/**"], crate_names: ["*-network*", "*-discv*", "*-eth-wire*", "*-p2p"] } }
- { id: sync, match: { paths: ["crates/stages/**", "crates/engine/**"] } }
- { id: storage, match: { paths: ["crates/storage/**", "crates/trie/**", "crates/static-file/**"] } }
- { id: execution, match: { paths: ["crates/evm/**", "crates/revm/**"] } }
- { id: rpc, match: { paths: ["crates/rpc/**"] } }
- { id: mempool, match: { paths: ["crates/transaction-pool/**"] } }
- { id: extension_points, match: { paths: ["crates/exex/**", "crates/node/**"] } }
expect:
- { node: rpc_surface, layer: rpc, why: "a node serves JSON-RPC and an engine API" }
- { edge: [sync, storage], why: "the sync pipeline writes state" }
- { edge: [execution, storage], why: "execution reads state" }
- { edge: [networking, mempool], why: "peers gossip transactions" }
questions:
- { id: networks, kind: intent, type: free, text: "Which networks is this node meant to run on?" }
- { id: storage_backends, kind: decision, type: choice, when: "map.layer(storage).backends > 1", text: "Several storage backends are present. Which are supported for production, and which are experimental?" }
- { id: stability, kind: status, type: free, text: "Which crates are a stable API for others to build on, and which are internal?" }
- { id: sdk_surface, kind: intent, type: yes_no, when: "map.layer(extension_points).found", text: "Is building your own node from these crates a supported use, with its own users?" }
- { id: perf_targets, kind: intent, type: free, text: "Is there a sync time or throughput target the project holds itself to?" }
documents:
prd:
for_each: product_surface # the node binary; the SDK for building nodes, if confirmed
look_for: { who_uses_it: [book, README.md], success_measures: [benches, ci] }
decision:
for_each: major_dependency # the database, the EVM, the networking stack, the async runtime
reason_sources: [commit_messages, book, owner_answer]
rfc:
kinds: [sync_performance, storage_migration, network_upgrade, rpc_compatibility, test_gap]
checks:
- { id: build, kind: build, run: ["cargo", "build", "--workspace", "--locked"], timeout: 60m, network: fetch_dependencies, tier: personal, turns: [builds] }
- { id: unit, kind: test, run: ["cargo", "test", "--workspace", "--lib", "--locked"], timeout: 90m, network: none, tier: personal, turns: [tests_pass] }
- { id: lint, kind: lint, run: ["cargo", "clippy", "--workspace", "--locked"], timeout: 60m, network: none, tier: personal, turns: [lint_clean] }
- { id: bench, kind: bench, run: ["cargo", "bench", "--workspace", "--locked", "--no-run"], timeout: 60m, network: none, tier: team, turns: [benches_build] }
evidence:
min_sources: { default: 1, architecture: 2 }
accept_needs: { prd.intent: confirmed }
measured_only: [sync_time, throughput, disk_use, memory_use]
The benchmark check above only builds the benchmarks. Running them is a separate check a customer adds with their own budget, because a node's benchmarks can need chain data that a sandbox does not have. Phase 0 runs reth's checks on our own runner, with reth's licence and credit shown on every page that uses the result.
3.2 Multi-service SaaS (multi-service-saas), first on InOrbit
schema: atlas.template/v1
id: multi-service-saas
version: 1.0.0
extends: generic@1.0.0
applies_when:
any:
- count: { service_entry_points: ">= 3" }
- path: "proto/**/*.proto"
- file_matches: "**/openapi*.json"
sources:
- { id: code, kind: repository_paths, match: ["**"], required: true, priority: 1 }
- { id: contracts, kind: repository_paths, match: ["proto/**", "**/openapi*.json", "**/schema.graphql"], required: false, priority: 1 }
- { id: deploy, kind: repository_paths, match: ["**/deploy/**", "**/*.yaml", "**/compose*.y*ml"], required: false, priority: 2 }
- { id: migrations, kind: repository_paths, match: ["**/migrations/**"], required: false, priority: 2 }
- { id: ci, kind: ci, match: [".github/workflows/*.yml", ".gitlab-ci.yml"], required: true, priority: 3 }
- { id: decisions, kind: repository_paths, match: ["docs/adr/**", "docs/adrs/**", "docs/rfcs/**", "docs/decisions/**"], required: false, priority: 2 }
- { id: site, kind: docs_site, required: false, priority: 4 }
- { id: incidents, kind: connection, connection: incident_tool, required: false, priority: 5 }
map:
layers:
- { id: edge, match: { roles: [reverse_proxy, load_balancer, auth_gateway] } }
- { id: api_gateway, match: { roles: [protocol_translation, api_router] } }
- { id: services, match: { roles: [rpc_server] } }
- { id: data_stores, match: { roles: [database, cache, object_store, analytics_store] } }
- { id: messaging, match: { roles: [queue, event_bus, outbox, webhook_sender] } }
- { id: identity, match: { roles: [identity_provider, token_issuer] } }
- { id: clients, match: { roles: [web_app, mobile_app, cli, sdk] } }
- { id: jobs, match: { roles: [scheduled_job, worker] } }
- { id: operations, match: { roles: [metrics, logs, traces, alerts] } }
expect:
- { node: authentication_point, layer: edge, why: "someone decides who a caller is" }
- { edge: [services, data_stores], per: service, why: "which service owns which data" }
- { edge: [clients, api_gateway], why: "how clients reach the services" }
- { node: tenancy_boundary, why: "how one customer's data is kept from another's" }
questions:
- { id: data_owner, kind: ownership, type: choice, when: "map.data_store.writers > 1", text: "Several services write to this data store. Which one owns it?" }
- { id: public_route, kind: security_privacy, type: choice, text: "Is this route meant to be reachable without signing in?" }
- { id: personal_data, kind: security_privacy, type: yes_no, text: "Does this data store hold personal data?" }
- { id: deprecated, kind: status, type: yes_no, when: "map.service.callers == 0", text: "Nothing calls this service. Is it still used?" }
- { id: tenancy, kind: decision, type: free, text: "How is one customer's data kept apart from another's?" }
- { id: secrets_location, kind: security_privacy, type: choice, text: "Where are this service's secrets kept?" }
documents:
prd:
for_each: product_area # what customers buy or use, from the site and the clients
decision:
for_each: [service_split, data_store_choice, authentication_design, messaging_choice]
reason_sources: [decision_docs, commit_messages, owner_answer]
rfc:
kinds: [missing_monitor, missing_test, unowned_service, contract_drift, authentication_gap]
checks:
- { id: build, kind: build, run: ["$build"], timeout: 60m, network: fetch_dependencies, tier: personal, turns: [builds] }
- { id: test, kind: test, run: ["$test"], timeout: 90m, network: none, tier: personal, turns: [tests_pass] }
- { id: contract_lint, kind: lint, run: ["$contract_lint"], timeout: 10m, network: none, tier: personal, turns: [contracts_valid] }
- { id: health, kind: probe, target: owner_named_hosts, path_from: map.health_endpoints, tier: team, turns: [service_running] }
evidence:
min_sources: { default: 1, data_ownership: 2 }
accept_needs: { prd.intent: confirmed, data_ownership: confirmed, security_privacy: confirmed }
measured_only: [latency, error_rate, throughput, availability]
$build, $test and $contract_lint are filled from what the readers found (the build
tool of each language in the catalogue, a contract linter when contracts exist) and shown
to the customer before consent, resolved, never as placeholders.
What this template should find on InOrbit, and where our truth for it is (section 4):
| The template expects | Our truth for it |
|---|---|
| An authentication point at the edge | The edge's gate configuration and the RFCs that define it; a live check that an unauthenticated call is refused |
| An API gateway that translates and forwards, with no business logic | The gateway's surface list (section 2) and its RFCs |
| Each service and the data it owns | Service crates, their migrations, and the decided RFC per service |
| Messaging: an outbox, event deliveries, webhooks | The event catalogue and RFC 0022 |
| Clients: consoles, the phone app, the CLI and six SDKs | The client folders, the SDK repository and its language list |
| Tenancy | The access RFCs (RFC 0065 for documents) and the services' own checks |
3.3 Legacy monolith (legacy-monolith)
Written for the first customer pilot (PRD Phase 1) and refined there in minor versions. It assumes one deployable unit with a large single codebase, and spends its effort on what such systems lack: written intent and module boundaries.
| Part | Contents |
|---|---|
| Sources | The codebase with its full history (commit messages are often the only reasons), build files, database schema and migrations, scheduled jobs, any wiki or ticket connection granted |
| Map shape | Modules found by package and import clusters, the database tables each module reads and writes, entry points (web routes, jobs, queue consumers), dead code candidates |
| Questions | "Is this module still used?", "Why are there two tables for the same thing?", "Who owns this area?", "Which jobs may never be stopped?" |
| Skeletons | A PRD per business capability found; a decision per framework, database and integration; RFCs for module seams, missing tests, unowned areas |
| Checks | Build and the existing tests; a schema dump against a disposable copy only when the customer provides one |
| Evidence | accept_needs: confirmed for every "still used" and every ownership statement; dead code is never stated, only "no caller found in the code read" |
3.4 Generic (generic)
The base every other template extends, and the template used when nothing better
matches. Its sources are the whole repository, its docs and CI. Its map has no expected
layers: it maps what the readers find by language, entry point and access point. Its
questions are the ones every system gets (who owns it, what it is for, what is deprecated,
where secrets are kept). Its skeletons are one PRD for the system and a decision per
major dependency. Its checks are the build and tests the readers found. An evaluation on
generic says so on every page, so a reader knows no type-specific knowledge was used.
4. Grading on our own system
4.1 Where it sits in the loop
Grading is the "grade" step of the loop in RFC 0045, applied to Atlas's own work: understand (the agent reads) → model (the map) → hypothesize (each statement is a claim) → verify (the template's checks turn some claims into measurements) → grade (against our truth) → evidence (a record). It follows the rules of the product direction: ground truth we know, never another model alone; dimensions reported separately, never one score.
It reuses what exists:
- The statement is RFC 0035's claim: text, a state, and sources (a file and lines at a commit, a contract operation, a migration, a doc page, an owner's answer).
- The record is RFC 0046's evidence record, with its own predicate type
https://inorbit.hr/evidence/atlas-grade/v1. The first step of RFC 0047 already writes unsigned draft records of that format (chaos verify); the grader's record is written the same way, a draft until the evidence service signs records, and labelled so. - Per-item results follow the grader of RFC 0008: per statement, the expected (what the truth says), the actual (what Atlas wrote), and the outcome.
4.2 What counts as truth
A grading run reads one truth snapshot, pinned when the run starts:
| Source | Pinned as | Counts as truth for |
|---|---|---|
| Code | One commit per repository | What exists: services, crates and packages, contracts, routes, migrations, dependencies and their versions, deploy manifests. Extracted by deterministic tools (manifest and lock files, contract files, the gateway's surface list, the OpenAPI document, migration names), never by a model. |
| RFC text | Each document's version id in the RFCs product | Intent and decisions: what a part is for, why it was built so, what was decided and when. Only RFCs whose stage says it is built (partly_live, live) count for what exists; an open RFC counts for intent only. Decision records (RFC 0081) count for decisions. |
| Live checks | Results captured when the run starts: health and surface checks, monitor results (RFC 0037), the served OpenAPI document | What runs and answers now. |
When they disagree, live checks outrank code for what runs, code outranks RFC text for what exists, and RFC text and decision records are the only truth for intent and reasons. A disagreement between our own sources is recorded as a finding about us (section 4.5), not as Atlas's error.
Private RFCs and internal documents are part of the truth snapshot. The grade record holds references and hashes only, so a published grade shows the shares and never what a private document says.
4.3 Matching a statement to the truth
- Structured statements (a map node, an edge, an access point, a dependency and its version, a data store's owner) are matched by key against the extracted facts. The matcher is code with tests, versioned, and the same inputs give the same outcome.
- Prose statements (what a part is for, why a decision was made) are matched in two
steps. A matcher finds the truth passages that speak to the statement and may propose an
outcome; a model may help it find them. A proposed outcome counts only when a person
accepts it. Until then the statement is
proposed, reported in its own column, and counted as unknown in the grade. - A person reviews every proposed contradiction, and a random sample of proposed confirmations in every run (the sample size is a grader setting, recorded in the run). The share of sampled confirmations the person overturns is reported as the matcher's error rate, beside the grade.
4.4 The grade
Each statement ends in one outcome: confirmed (a truth source supports it), contradicted (a truth source with standing on that kind of statement says otherwise), or unknown (no truth source speaks to it). A grade reports these dimensions separately, per template version, per evaluation:
| Dimension | Meaning |
|---|---|
| Confirmed share | Confirmed statements over all graded statements. |
| Contradicted share | Contradicted over all graded. The PRD's main risk, and the number that gates a template release. |
| Unknown share | Statements no truth source speaks to. High means our own documentation has gaps, or Atlas wrote things nobody can check. |
| Coverage of the truth | The share of extracted facts (services, access points, data stores, decisions) that Atlas stated at all. Without it, saying little would score well. |
| Evidence coverage | Statements with at least one source. Anything below 100% is a bug (PRD). |
| Citation validity | Statements whose cited source, opened at its commit or version, says what the statement says. A confirmed statement with a wrong citation is still a defect. |
| State honesty | Statements marked confirmed or measured that have no owner answer or measurement behind them. Must be zero. |
| Conflicts surfaced | Disagreements our truth snapshot holds (code against RFC) that Atlas showed as conflicts instead of picking a side. |
| Matcher error | From section 4.3. |
This grade is the PRD's success measure "Accuracy on known systems": the confirmed
and contradicted shares of the multi-service-saas template on InOrbit, tracked per
template version. No target is set before the first run; the first run is the baseline,
and the owner sets the Phase 1 bar from it (open questions).
4.5 Triage of a contradiction
Every contradicted statement opens a triage item with the statement, its sources and the truth passage. A person decides one of:
| Outcome | What follows |
|---|---|
| Atlas was wrong | A defect against the template version (a map rule, a missing question, a skeleton), a reader, or the model's writing, by category. It stays contradicted in this run. |
| Our truth is stale | A defect on that RFC through RFC 0080's facts, and the RFC is fixed. The statement is regraded confirmed in a triaged view of the run, with who decided and when. |
| Both partly right | The statement is split; each part is graded on its own. |
| The truth snapshot was wrong | An extractor bug; fixed, and the run is graded again on a corrected snapshot, with both runs kept. |
A run's grade is never edited. Triage produces a new record that supersedes the first by reference (RFC 0046's rule), so both "as graded" and "after triage" stay visible, and how much triage moved the number is itself reported.
4.6 Reproducible runs
A grading run stores, by reference and digest: the evaluation's output snapshot, the
template id, version and digest, the coverage catalogue, the agent and reader versions,
the model and its version and the prompts' digest, the truth snapshot (commits, document
versions, captured live results), the matcher version and settings, and every person's
decision with who and when. grade replay <run> recomputes the grade from those stored
inputs and must produce the same numbers; live checks are replayed from their captured
results, never probed again. CI runs a replay of the latest run on every change to the
grader.
A template release is gated by its grade. A new version of multi-service-saas is
graded on the same truth snapshot as the version it replaces. It is released when its
contradicted share is no higher and its evidence coverage is 100%. Other dimensions are
reported, not gated, until we have enough runs to know their noise.
4.7 The other systems
- reth. We hold no truth for reth. Its evaluation is reviewed statement by statement before it is published (PRD question 4). Each review outcome is a person's confirmation or correction, recorded like triage, and reported as "reviewed by InOrbit", never as reth's maintainers' view and never mixed with InOrbit's grade.
- A customer's system. The owner's confirmations and acceptances are the truth. The pilot's measure is the PRD's: the customer confirms the map and accepts at least one PRD and one RFC. Their statements and answers never enter our truth snapshots or our templates.
- Fixtures. The per-language and per-access-point fixtures of section 2 are graded on every agent build; a fixture's contradicted share must be zero.
5. Where templates live, and who writes them
The schema, the validator and the four templates live in one folder in the SDK
repository, beside iohr and the generator whose language list they depend on. CI there
validates every template, tests the reth template's patterns against the pinned reth
commit, and refuses a change to a released version's file (a change is a new version
file). The platform serves released versions by id@version with their digests, read-only;
the agent fetches a template through the platform and checks its digest before use.
Templates hold no customer data. Nothing from a customer's index, answers or documents crosses into a template (PRD Security and privacy); a template changes only by a person writing it in a pull request. The schema's linter refuses literal hosts, addresses and secrets in any field.
Who writes them, by PRD phase:
| Phase | Templates |
|---|---|
| 0 · Ourselves and reth | We write all four. rust-node-client runs on reth, multi-service-saas on InOrbit; the first grading run and the release gate are built here. |
| 1 · First customer pilot | legacy-monolith refined on the pilot in minor versions. The pilot's lessons enter the template only as text we write. |
| 2 · Self-serve | The templates library: every released template listed in the console with its version history and what changed; the suggested template shown at the start of an evaluation with why it was suggested. |
| 3 · Rooms at scale | Template authoring for customers: an account's own templates, private to it, same schema, validated by iohr, extending ours. Graded only against the customer's own truth, if they provide one. |
6. One engine for every target, and what a run produces
The owner approved on 2026-10-08 that a study is what a template run produces (ADR 0033, PRD 0001's second addendum of that day). This section widens sections 1 to 5 to match; nothing in them is withdrawn.
6.1 The template, as a package
The YAML of section 1 is the package's template.yaml. A released template is one signed
package (RFC 0092): a manifest.yaml beside it names the id, version, target, the
publisher, the licence, the templates it extends by digest, its permissions (what it
may read, fetch and run), its outputs and the minimum agent and coverage catalogue
versions. The manifest's signature covers the template's canonical digest (section 1.7),
and the agent checks both before every run. Versions follow section 1.8; a released
version is immutable; the grade gates a release (section 4.6).
6.2 Targets other than a system
target says what a template runs against. The schema allows each target only the source
and check kinds that make sense for it:
| Target | Sources | Checks | First templates |
|---|---|---|---|
system |
repositories, docs, CI, releases, connections (section 1.1) | build, test, lint, bench, probe | rust-node-client, multi-service-saas, legacy-monolith, generic |
site |
pages reachable from the agent, its sitemap and robots file | crawl, lighthouse, probe (read only) |
seo-audit |
documents |
a set of documents or a documentation connection | none; extraction only | vendor-terms |
vendor |
a vendor's published terms, pricing and DPA pages | none | vendor-terms, vendor-selection |
pipeline |
a build or deploy the agent can observe, with its logs and timings | time (run a named step N times and record each duration) |
pipeline-timing |
api |
an API's served description and the clients that call it | contract_lint, probe |
api-parity |
A sketch of seo-audit (not released; the shape it will take):
schema: atlas.template/v1
id: seo-audit
version: 0.1.0
target: site
extends: generic@1.0.0
sources:
- { id: pages, kind: site_pages, follow: internal, respect_robots: true, required: true, priority: 1 }
- { id: sitemap, kind: sitemap, required: true, priority: 1 }
checks:
- { id: crawl, kind: crawl, max_pages: 3000, network: target_hosts, tier: free, turns: [page_facts] }
- { id: lighthouse, kind: lighthouse, sample: [front_page, one_per_section], forms: [mobile, desktop], tier: personal, turns: [page_performance] }
questions:
- { id: indexed_on_purpose, kind: intent, type: yes_no, when: "page.noindex", text: "This page asks not to be indexed. Is that intended?" }
- { id: languages, kind: intent, type: free, text: "Which languages should search engines see, and at which URLs?" }
evidence:
min_sources: { default: 1 }
measured_only: [page_performance, page_weight]
6.3 What a run produces: a study
Every run writes one study (ADR 0033) through the RFCs API, in the space the run names,
from the study's findings template: question, context, method, findings, recommendation,
open questions. The template's parts land in fixed places:
| Template part | In the study |
|---|---|
id, version, digest, target and its snapshot |
Context: the run record (section 1.7 plus the target) |
sources, checks and what was skipped |
Method: what was read and run, with dates, and what the agent's policy refused |
| What the readers and checks found | Findings: each one a claim with its evidence and its state (inferred, confirmed, measured) |
questions answered |
Findings marked confirmed, with who answered |
questions not answered |
Open questions |
documents skeletons |
The documents the study informed (an RFC, an ADR, a PRD draft), each naming the study in its context documents |
Findings are the claims of RFC 0086 (support, conflict, validity, observability), so a second run on a changed target marks a finding stale, not false, and two runs are compared finding by finding.
6.4 Grading templates that are not about a system
Section 4 grades against a truth we hold. Other targets have truths too, and the same rules apply (deterministic extraction decides, a person decides every contradiction):
site: the site's own build output (titles, canonicals, sitemap entries from the export) is the truth for what the crawl reports;api: the served API description and the checked-in views map (RFC 0090) for parity findings;pipeline: the per-step timings the roll scripts record for timing findings;vendoranddocuments: a person's review of each finding against the cited page at its fetched date; no automatic grade.
Each first runs on InOrbit, like section 4.
Controls
- Customer data. Templates hold none. Grade records hold references and hashes only (RFC 0046). A customer's evaluation is never part of our truth snapshot or our template changes. The matcher sends nothing to a third-party model that the evaluation's model choice (PRD question 2) does not allow; for our own system the truth snapshot stays on our infrastructure.
- Least privilege. Checks run only in the customer's sandbox, after consent to the exact commands, with no network unless the check says why, and the agent's policy can refuse any of them. Probes read and never write.
- Change management. Template versions change only through pull requests with required checks; a released version is immutable; a release is gated by its grade.
- Audit. Every grading run, review decision and triage outcome is recorded with who, when and the outcome, never the statement's text in a log line.
- This RFC weakens no existing control.
Alternatives considered
Templates as code (a plugin per system type). A Rust or WASM module per template could detect layers more precisely than match rules. Rejected for now: a template would become code we ship into a customer's agent, with its own supply chain and review, and customers could never read or write one in Phase 3. Declarative files with a fixed vocabulary are enough for the four types; a reader that needs code belongs in the agent (A2), where code already lives.
Templates as documents in the RFCs product. They would get versions, review and access for free. Rejected as the source: a template is data the agent executes against and CI must validate and test against reth; the product is the right place to show templates and their history (Phase 2), not to keep them.
Grading with a model as the judge. Cheaper and fast: a second model compares Atlas's documents with our RFCs and scores them. Rejected as the grade, because it breaks the platform's first rule: a model's opinion is not ground truth. A model may still propose matches (section 4.3); only deterministic extraction and people decide outcomes.
Grading against RFC text alone. Simple and closest to the PRD's wording. Rejected because our RFCs describe unbuilt work and lag the code; Atlas would be penalised for reading the code correctly. Code and live checks outrank text for what exists.
One accuracy score. One number per template version is easy to chart. Rejected: a template that writes few, safe statements would score well while missing half the system. Coverage of the truth and the contradicted share must be read together.
A hand-kept language list in Atlas. One list in the agent, edited when someone remembers. Rejected for the drift it invites; the generator's list and the gateway's surfaces are already maintained because the SDKs and the gateway break without them.
Trade-offs
- Match rules over paths and names will miss systems that do not follow their type's conventions. The owner confirms the template pick and the questions fill gaps; a template's misses show up as unknowns and as low coverage of the truth.
- Human review of contradictions and samples costs our time on every grading run. It is the price of not grading with a model, and it shrinks only as the matcher's measured error shrinks.
- Adding not-built languages to the SDK generator's list makes the generator carry entries it cannot render. It already models that state; the alternative is a second list.
- A release gate on the contradicted share can hold back a template version that is better on coverage. That is deliberate in Phase 0, while wrong statements are the risk the PRD names first.
Open questions
- Are the template files public? They hold no customer data, and publishing them lets anyone see what Atlas reads and asks before installing an agent. The SDK repository is public, so keeping them there publishes them. Proposed: yes. The alternative is a private folder in the platform repository. Owner's call.
- The Phase 1 bar. Proposed: no target before the first graded run on InOrbit; after it, the owner sets the contradicted share a template must stay under before a customer pilot uses it.
- One template per evaluation, or per area? A system can hold a node client and a SaaS around it. Proposed: one per evaluation in Phase 0 and 1, with a per-area template as a minor schema change if the pilot needs it.
- Who reviews in triage. Proposed: the owner for intent and decisions, any staff engineer for structural statements, both recorded. With one person approving all changes today, separation of duties is limited, and the record says who decided.
- The sample size for reviewing proposed confirmations: proposed 10% with a floor of 20 statements per run, revisited after the first three runs.
Slices
Each slice is its own pull request, recorded on this RFC when it lands.
- Schema and validator. The
atlas.template/v1JSON Schema,iohr atlas template check, the canonical digest, and thegenerictemplate. Acceptance: a template with a language switch-off, a literal host or a secret-shaped value is refused with the line. - The coverage catalogue. The generator's not-built languages,
iohr sdk languages --json, the gateway's surface list as data with its CI check, the sync into the SDK'sspec/,iohr atlas coverage --jsonand its CI check. Acceptance: adding a surface to the gateway without regenerating the file fails CI in the platform repository; adding a language to the enum without the catalogue fails CI in the SDK repository. - The first four templates at
1.0.0. With the reth pattern test against a pinned commit. Acceptance: every layer ofrust-node-clientmatches at least one crate in the pinned reth. - The truth snapshot for InOrbit. The extractors (manifests, contracts, the surface list, migrations, RFC stages and decision records) and the live capture. Acceptance: two snapshots of the same commits and versions are byte-identical apart from the capture times.
- The grader and its record. Structured matching, the proposed column, review and
triage, the draft evidence record,
grade replay. Acceptance: a replay of a stored run gives the same numbers; a triage produces a superseding record. - The first grade. A Phase 0 evaluation of InOrbit (Atlas A10) graded; the result written on this RFC and in the evaluation room. Acceptance: every dimension of 4.4 is reported, none of them as a stub.
- The release gate. CI grades a new
multi-service-saasversion on the pinned snapshot before release. - Targets and the study output (section 6): the
targetkey and its allowed source and check kinds in the schema, runs writing studies through the RFCs API, and the first non-system templates (seo-audit,api-parity,pipeline-timing,vendor-terms), each first run on InOrbit against the hand-written studies of 2026-10-08. - Packaging and signing of templates, built in RFC 0092.
Slices 1 to 3 need no other Atlas part. Slices 4 and 5 need the evaluation's statement format (A2 and A4). Slice 6 needs A10's run.
Decision
Open. Proposed: declarative YAML templates with a published schema, immutable versions pinned by digest in every evaluation; languages and access points from one generated catalogue built from the SDK generator's list and the gateway's surfaces; a grade of separate dimensions against a pinned truth snapshot where code and live checks outrank RFC text for what exists, people decide every contradiction, and a template version is released only when its contradicted share does not rise. Since 2026-10-08 (owner): every run produces a study whose findings are claims with evidence (ADR 0033), templates exist for targets other than systems (section 6), and they are signed packages the agent verifies before it runs them (RFC 0092).
Publication
Public on the Lab with this text. The template schema and the four templates are published with slice 3 if the owner answers open question 1 with yes. No surface says Atlas is graded, or shows a grade, until slice 6 has produced one.
Status log
- 2026-10-07: opened for Atlas A6 (PRD 0001: templates and grading). Nothing is built.
The reth layout above was read from its repository at
1e8eb0c18202. - 2026-10-08: section 6 added after the owner approved templates producing studies: the package and its manifest, targets other than a system, the study as every run's output, grading for the new targets; slices 8 and 9. Distribution, the catalogue and company portals are RFC 0092. Nothing is built.