This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to LLM

RFC 0010open2026-09-24

The sandbox

Running a stranger's Go, Rust or Solidity program on the platform's own machine and giving back what it printed, with every way it could do harm named first and closed by a layer that does not trust the one before it.

Problem

The workbench's models write code, and the next thing anyone wants is to run it. Running a program someone else wrote, on the machine that also serves the platform, is the most dangerous thing the platform could offer: a program can do anything the machine lets it do, and a model will happily write one that tries. The same need exists elsewhere in the lab: grading a learner's program on the platform rather than on their own machine was deferred as "a sandboxing project of its own" (RFC 0008). This is that project.

What a hostile program could try

Every one of these is a test in the escape suite before anything is offered, and every defence below names which of them it stops.

  1. Reach a network: the internet, the machine, or the platform's own services.
  2. Read or write the machine: its files, its other programs' memory, its devices.
  3. Exhaust the machine: memory, processors, processes, disk, open files.
  4. Outlive its turn: run forever, sleep forever, leave a child behind.
  5. See another run: anything left by an earlier program, anything of a concurrent one.
  6. Attack the compiler: code that is slow or huge to compile, or that asks the build to run something before the program itself does.
  7. Flood the answer: endless output, or output shaped to break the page that shows it.

Proposal

Two layers, the pattern the model service already uses. The engine is a small daemon on the machine, the only thing that starts a sandbox. The service in front of it, the runner, is a platform service like any other: it knows who is asking, how many runs are in flight, what each caller has left today, and it writes the audit line. A request says three things (the language, the source, the input) and nothing in it ever becomes an option of the sandbox.

One throwaway sandbox per run, compiling and then running inside the same one, so no compiled program ever crosses back to the machine. It is a container whose kernel is : a kernel written in a memory-safe language that runs in user space and answers the program's system calls itself, so the program never talks to the machine's kernel directly.

Defence Stops
's own kernel between the program and the machine's 2, 4, and the kernel attacks behind them
No network at all: none requested, and the runtime configured to refuse one even if asked 1
The image read-only; the only writable place a small in-memory scratch space, gone with the sandbox 2, 5
An unprivileged user, every capability dropped, no way to gain one 2
Limits on memory, processors, processes and file size 3
A deadline for compiling and one for running, enforced from outside by removing the sandbox 4, 6
Standard library only: no packages fetched, no Cargo (so no build scripts or macros from elsewhere), no C 6, 1
Output capped per stream; the page renders it as text, never as markup 7
The runner: a caller required, a bounded number at once, a daily allowance per caller 3 at the platform's scale
One audit line per run: who, which language, how big, how it ended, never the code or its output the record SOC 2-style controls ask for

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Who may run. Four kinds of caller, each with its own daily allowance: admins; any signed-in visitor, with a smaller one; the platform's own services, which run code only to check what they publish (the Radar runs every example once before it ships) and are not counted; and, since 2026-09-29, anyone at all, so a reader of the Radar can run its examples without an account. An anonymous visitor is counted per network address, and the address is kept only as a salted hash, in memory, for the day; visitors share a fixed part of the sandbox and never the whole of it, get a smaller program than a signed-in caller, and are refused outright when the gateway could not say where they came from. A switch turns anonymous runs off without a release. Agents over MCP do not get the tool: running code on behalf of an agent is a separate decision, made later and on purpose.

Alternatives considered

microVMs. Each run in its own virtual machine on the processor's virtualisation: a stronger wall than 's, and what large hosted sandboxes use. More to build and keep (a kernel, a root filesystem per language, the jailer that confines the VM monitor itself, which had a file-overwrite flaw of its own early in 2026). Not first; a study compares the two on this machine, and replaces if the cost is small.

Namespaces and seccomp alone (bubblewrap, a plain container). Much cheaper, and the program talks to the machine's real kernel, so one kernel bug is a way out. Rejected as the only wall; it is what adds a second one to.

WebAssembly. Both languages compile to it, and a WebAssembly runtime is a strong, small sandbox. Rejected for now: the program the page runs would not be the program the learner would ship, the compiler itself still has to run somewhere, and Go's support is the newest of its targets.

Decision

Open. Fixed so far: two layers, the engine on the machine and the runner in the platform; one sandbox per run for compiling and running; no network; the standard library only; every limit enforced from outside the sandbox; admins only while the lab is private; no MCP tool. Still to decide: after its study; a vetted, offline set of packages; a longer-lived shape for programs that serve a port (the hosted grader of RFC 0008), which is a different sandbox and a later RFC.

Publication

The workbench runs Go and Rust code blocks and shows what they printed. A study, "What it takes to run a stranger's code", reports the escape suite's results, how long a run takes cold, and what a run costs the machine.

Status log

  • 2026-09-24: opened. The kernel is installed on the machine: a sandbox sees , not the machine's kernel, and has no network even when one is asked for.
  • 2026-09-24: the engine runs. The daemon compiles and runs a Go or Rust program in one throwaway sandbox in one to two seconds cold, and the escape suite contains every hostile program it has (study 0003). A sandbox that reaches its memory or process limit is ended whole, and the daemon says so in words.
  • 2026-09-24: the runner is the platform's way in: a verified caller with the role, the bounds, a daily allowance, a bounded queue, then the sandbox; one audit line per run without its code, and a test that fails if the code ever reaches the log.
  • 2026-09-24: the workbench runs Go and Rust code blocks through it, as an admin; the page's "what this page sends" says the code is run once and not kept.
  • 2026-09-24: the sandbox is on the live view. The arena reads the runner as it reads the model tiers, and the workbench shows its slots, runs a minute, how long a run takes and how often the sandbox did not answer, beside the models.
  • 2026-09-29: anyone may run a Radar example. The runner knows four kinds of caller (admin, signed in, platform service, anonymous), each with its own allowance; an anonymous visitor is counted by a salted hash of the address, never the address, in a fixed share of the sandbox, and a test fails if the address or the code reaches the log.
  • 2026-09-29: Solidity is a third language (RFC 0013). The compiler checks the contract; the program the sandbox then runs is not the learner's but a harness the image carries, which compiles the contract with the tests it reads on its input, runs them in an embedded EVM and prints a report. Every defence above applies unchanged; a test's gas allowance ends a loop long before the deadline does.
  • 2026-09-29: a run says how big the program was. After a compile that wrote a file the daemon measures it, one stat in the sandbox, and the compile step carries artifact_bytes; the runner passes it on and the workbench's card shows it beside the times. Nothing more of the build is shown: linking is not a step of its own in either toolchain, and the verbose output that would show it names the image's layout, which stays inside.
  • 2026-10-07: Checked: the sandbox, the runner and the Solidity harness are on main; the runner deployment runs with no recorded commit, so the running build is not verified. Open: the decision lists what is fixed so far.

Mentioned in

← Back to LLM