Problem
The service in RFC 0001 answers one conversation at a time and forgets it: nothing is stored but who asked and what it cost. That is right for a demo and wrong for anything that has to do work across days. An assistant worth using remembers what it was told, recalls it when it matters and says where it came from; and it can act, not only talk: read a file, query a service of this platform, start a measurement, and look at the result before it answers.
Both are easy to do badly. Memory that is a pile of transcripts is a privacy problem and a retrieval problem at once; tools that the caller declares freely are a way to make the platform call anything; an agent loop without a budget and a trace is a way to spend a day's tokens on a question nobody can reconstruct afterwards.
Proposal
Memory is a service of its own, not a table inside the model service. It keeps three kinds of entry per subject, each with the text, where it came from and when:
- facts: short statements the subject or an agent chose to keep ("the ledger's facts are append-only"), with the source they were taken from;
- episodes: what happened in a session, summarised when the session ends, never the raw transcript;
- preferences: how the subject wants to be answered.
Each entry is embedded by a tier whose engine declares an embedding model (RFC 0001's capability rule) and stored beside its vector in . Recall is a query: the conversation's last turn, embedded, the nearest entries of that subject above a threshold, and a hard cap on how many and how many tokens. What was recalled goes into the prompt as a block the model is told to cite, and comes back to the caller beside the answer, so every answer that leaned on memory says which entries it leaned on. A subject can list, correct and forget their entries; forgetting is deletion, not a flag.
Tools are the platform's, not the caller's. The contract gains tool calling: the model may ask to call one of the tools the service offers for the tier and the caller, with arguments that must validate against the tool's schema. The first tools are the platform's own RPCs, the same set RFC 0005 exposes over MCP, filtered to the caller's rights; a tool runs as the caller, never as the service. A caller cannot hand the model a new tool in the request.
The agent loop lives in the service and is bounded: a number of steps, a number of tokens (counted against the same daily budget), a wall-clock limit, and a trace per step (what the model asked, what the tool returned, how long each took) recorded beside the generation row without the text of either. The caller sees every step as it happens, as chunks of their own kind in the same stream.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Alternatives considered
Keep the transcripts and search them. The most recall for the least design, and the most exposure: every question ever asked, verbatim, in a database. Rejected; the record stays free of text, and memory keeps what was chosen to be kept.
A vector database . Rejected for now: the platform already , the entries are counted in thousands per subject, not billions, and one store is one thing to back up, migrate and reason about. A study says when that stops being true.
A multi-model database in its place (SurrealDB). One engine for documents, graph edges, vectors and full-text search, written in Rust, embeddable, and sold for exactly this: agent memory as a knowledge graph with vector recall. Rejected for now, for four reasons. It would be a second stateful system to run, back up and test beside that already has vector indexes and serves the rest of the platform. Its ready-made memory layer is the part this RFC means to build and understand, not buy. Its third major version is months old and its index layouts have changed between minor releases. Its licence restricts offering it as a service until it becomes Apache 2.0 in 2030. It would win if memory became a dense graph whose main question is several hops deep; the memory study measures how deep recall actually goes, and this is revisited if the answer is "often more than one".
Caller-declared tools, as the model APIs offer them. Rejected: a platform that runs whatever tool a request describes runs whatever the request wants. The tools are the platform's, listed and schema-checked.
Decision
Open. Fixed so far: memory as its own service with the three kinds, no transcripts, recall with sources returned to the caller, forgetting as deletion; tools as the platform's RPCs run as the caller; the loop bounded by steps, tokens and time and traced without text.
Decided on 2026-09-24, and built:
The tools are a table in the model service, and an agent's file names the ones it may call. The first four are
platform_status(the arena's snapshot, summarised),list_models(the tiers),radar_latest(the Radar's latest published digests) andrun_code(a program in the sandbox, RFC 0010); the fifth,open_page, is below. Each remote tool is one call through the platform's internal gateway, carrying the caller's verified identity, so the service behind it decides what the caller may do.Who has which. The code reviewer has
run_codeandplatform_status. It is the only agent that may run code on its own, and every run counts against the caller's daily runs. The site guide hasplatform_statusonly.The step bound is four tool calls a turn. After that the model answers without tools. Every round is checked against the caller's daily token budget, and the tier's deadline covers the whole turn.
The trace. Each call reaches the caller as a step of its own in the stream. The record keeps the number of steps and the tools' names, never what was sent or answered.
The embedding model and its tier. The fast tier embeds with Qwen3-Embedding-0.6B: 1024 dimensions, multilingual including Croatian, and small enough to run on the processor, so the models on the graphics card are untouched.
Memory is its own service. It keeps facts and preferences, each embedded as the caller, and recalls the nearest above a threshold of 0.40 cosine similarity, chosen by study 0004 on 50 notes and 86 questions: the right note shown to 94% of the questions against 84% at the first value of 0.45, for 2.7 other notes shown instead of 1.3. A turn is shown at most 5 entries. Every call is the caller's own entries, never another's, and forgetting deletes.
Which agents remember. The code reviewer, for preferences such as a language or how short to be. Before each turn it is shown what it remembers, each entry with its id, and it cites the ones it uses. It can keep and forget entries at the person's request. The site guide keeps nothing about a visitor.
Opening pages (decided 2026-09-29). The person can ask the reviewer to open any web page or link and read it.
open_pageis a tool of the table backed by a web service of its own: one GET of the URL, http or https, public names and addresses only (never the cluster, the host or a metadata endpoint; the name resolved and every address checked before a socket, the connection pinned to what was checked, each redirect checked again), nothing of the caller sent with it, a deadline and a size cap, the page reduced to readable text with its links numbered, read a window at a time so the model can page through and follow a link by its href. Counted against the caller's pages for the day, audited by host and size, never by path or text. Only the reviewer has it; it is never callable on its own or over MCP, because a page fetcher open to an MCP client is an open proxy. Not read: pages that draw themselves with JavaScript (a headless browser is its own RFC), PDFs.Episodes (decided 2026-09-29). When: the person ends the session, in the workbench; nothing ends one by itself, because the service holds no transcript (a client sends the whole conversation on every turn) and because a note in a person's memory should be their act, not a side effect. Who writes: the model service, as itself, through the one memory RPC that names a subject,
WriteEpisode, which has no route and takes only a service the memory service is told to trust; a person cannot write an episode, so the kind means what it says. Which tier: the fast one, a config dial; the summary is at most 80 words and 500 characters, in the third person and the person's language, without code, commands, addresses or credentials, and it is the person's own generation, budgeted and recorded. The deep tier would write a better summary in ten seconds on the one slot the Radar's digests share; the trade is the dial. A recalled episode is shown with its date. The transcript is read once and never stored.
Still to decide: a recall threshold per language or a query instruction in the note's language (study 0004 found short Croatian notes the weak spot), whether episodes, being longer than notes, want a threshold of their own (the study's follow-up), and whether memory is shared across subjects ever (the default is never).
Publication
The workbench (RFC 0004) shows recalled entries beside each answer and every tool call as a card in the transcript. Study 0004 measured recall: how often the right entry comes back, how often a wrong one does, and what it costs per turn; it moved the threshold to 0.40 and committed the corpus and the instrument.
Status log
- 2026-09-24: opened.
- 2026-09-24: a multi-model database considered in ; rejected for now, with the condition that would reopen it.
- 2026-09-24: agents (RFC 0011) are the first users of tools and memory once they exist; until then an agent is grounded by a brief and the page it is on.
- 2026-09-24: tools built: four platform tools run as the caller, at most four calls a turn, each streamed as a step and recorded without text; the code reviewer may run code in the sandbox, no other agent may.
- 2026-09-24: memory built as its own service, on the fast tier's new embedding model; the code reviewer recalls, remembers and forgets as the caller; episodes deferred.
- 2026-09-29: the table is callable on its own (
ListTools,CallTool: a reviewed list of its reads, neverrun_codeor the memory tools) and mirrored over MCP under the same names (RFC 0005), through the loop's own path, so one table reads the same in an agent's step, in the audit log and to an MCP client. - 2026-09-29: study 0004 measured recall; the threshold is 0.40 with the trade stated.
- 2026-09-29: episodes built: the person ends a session, the model service summarises it on the fast tier and keeps it as itself through memory's one service-only RPC.
- 2026-09-29:
open_pagebuilt on a web service of its own: the reviewer opens any public page or link the person names, a window at a time, within the service's bounds; never over MCP. - 2026-10-07: Checked: tools, memory, episodes and
open_pageare on main; llm runs 932c4445, while the memory and web deployments record no commit, so their running build is not verified. Open: the decision lists what is fixed so far, not the whole proposal.