Problem
The loop the platform is built around (RFC 0045) starts with observing a system and ends with evidence someone else can check (RFC 0046). The agent (RFC 0029) is how the platform observes systems inside a company's network, and today it sees exactly one thing: the checks it was asked to run. It knows whether a request it sent came back in time. It knows nothing about the traffic the server was actually serving when it failed.
That is the gap a post-mortem falls into. The questions after an incident are about the
wire: which clients were connected, which endpoint slowed first, whether the kernel was
retransmitting, whether the accept queue overflowed, which process owned the connection
that hung. Logs answer some of that, after the fact and only where someone thought to
log. Metrics answer it in aggregates chosen in advance. A packet capture answers it, if
someone happened to be running tcpdump on the right interface at the right time, which
nobody ever is.
Three constraints make it harder than attaching a capture tool:
- The agent has no privileges, on purpose. It runs with an empty capability set and its service unit blocks privileged system calls and raw sockets. That is part of why a security team lets it in (RFC 0029). Reading traffic needs the kernel's cooperation, and that needs privileges.
- Traffic is personal data. Addresses, host names, DNS queries and request paths identify people. Courts have held that even a dynamic IP address can be personal data (CJEU, C-582/14, Breyer). A capture that ships what it sees to us would make us a processor of every customer's users' traffic.
- Most of an edge is encrypted. Whatever reads packets before TLS terminates sees handshakes, not requests. A design that promises request paths from the wire has to say where it cannot deliver them.
There are two more gaps on the same path. The agent arrives on a laptop as an iohr
extension (RFC 0028), and the extension model has no way to say "this one needs root".
And the console is a web page tied to : a person working with an agent
on their own machine has no console there, and nothing tells them which plan, features
and extensions their team may use.
Proposal
A privileged companion, not a privileged agent
Capture is a separate program, iohr-capture, in the same repository as the agent
(inorbithr/dataplane). The agent stays exactly as unprivileged as it is today.
- Privileges only to load, then none. The companion starts with
CAP_BPF,CAP_PERFMONandCAP_NET_ADMIN(capabilities(7)), loads and attaches its eBPF programs, then drops every capability before it parses a single byte. The ring buffers it already holds stay readable. A flaw in the parser runs as a process that can no longer change the kernel. - Its own service. It runs as its own system service, sandboxed, with a unit that allows only what loading needs. It is never started by the agent and never by the platform.
- Two local sockets, no listener. An aggregates-only socket the agent reads, open to the agent's own user and checked against the caller's credentials by the kernel; and a control socket only root can open, for the one request that touches payloads (a packet file, below). Neither listens on a network. No job sent from the platform can reach the control socket, so the platform can never start a packet capture.
- Off until two people's switches say on. The agent's local policy gains
[work] capture = falseand a[capture]section (interfaces, layers, bounds), and the companion has its own configuration. The local policy wins, as it always has: if either says no, nothing is captured. - Kernel. Linux 5.8 or newer with BTF, for compile-once-run-everywhere programs and the BPF ring buffer (kernel docs). An older kernel is refused with a message that says so.
The programs are written in Rust with aya, the toolchain RFC 0003 already chose for the shield, so the eBPF code, its loader and the agent share one language and one build.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
What it sees, layer by layer
The owner chose these layers for the first build on 2026-10-05. Each has a proof, and every proof is a count.
| Layer | What it reads | What it proves | Where it stops |
|---|---|---|---|
| 1. Headers | traffic-control programs on ingress and egress of the configured interfaces: per-CPU counters and the first header bytes of sampled packets | packets and bytes per protocol, port and TCP flag; drops | Packets the kernel has merged on receive (GRO) are labelled as merged, never presented as wire packets |
| 2. Protocols | the first bytes of each flow's payload | HTTP/1 request line and host, TLS ClientHello server name and ALPN, DNS queries, HTTP/2 and gRPC paths, , command names; requests per protocol, method, normalised path or host | Behind TLS only the handshake is visible. HTTP/2 paths come from a connection's first request; later streams are counted and their path is "unknown" |
| 3. Full packets | a second ring buffer, only when its own switch is on, rate-capped | packets copied and dropped; a packet file made on request | Kept in memory; a file only when a person asks, below |
| 4. Owner | socket diagnostics first, then probes at accept and connect time: socket to process and cgroup | flows per process, container and pod | Packets that arrive before a connection is accepted are counted as unowned |
| 5. TCP health | the kernel's own TCP state, retransmit events and listen-queue counters | retransmits, resets, round-trip time distribution, accept-queue overflows | Only what the kernel tracks; nothing inferred |
| 7. Request timing | layer 2's per-flow bytes, request paired with response | latency histogram and status counts per method and path template (numbers and ids replaced) | Plaintext only: behind TLS, timing is per connection, not per request |
Layer 6, decrypted TLS, is left out by the owner's decision. It needs keys or hooks in the application's TLS library, and deserves its own RFC.
Layers 4 and 5 start with what needs no probes: socket diagnostics, tcp_info and the
kernel's network counters. Probes and tracepoints come after, only for what those miss
(retransmit events, ownership at the moment of accept), with a version check where the
kernel changed a function's signature.
Bounded everywhere. A flow table with a limit, fixed ring-buffer sizes, a per-CPU token bucket in the kernel that caps samples per second, and a CPU budget in user space. When a bound is hit, data is dropped and the drop is counted. Nothing queues, so a flood costs the host a counter, not memory.
Full packets, on request, never sent
Layer 3 is the one that carries payloads, so it has its own switch and its own rules:
iohr capture pcap --for 30s --filter …asks the control socket for a capture file. It is written readable by its owner only, capped in size and duration, deleted after a set time, and never sent anywhere.iohr capture dissectruns the host's owntsharkon that file, as the person who asked, never from the companion. Wireshark's dissectors are the best there are and are licensed under the GPL (Wireshark); runningtsharkas a separate program on the host keeps that licence where it belongs. We do not bundle it, and the package only suggests it.
Nothing stored, nothing leaves the host
In the first phases the aggregates live in the companion's memory and reset when it
restarts. What leaves the host is one fact per layer: in the agent's session hello, the
capability strings capture:headers, capture:protocols, capture:packets,
capture:owners, capture:tcp and capture:timing. They tell the platform what this
host can do, never what it saw. The agent reports them only when the companion's socket
answers and the local policy allows that layer.
The numbers themselves are read on the host: iohr capture status (and the companion's
own stats) show live counters per layer with drops and the time window, and the agent's
read-only admin page gains a Traffic section with the same counts.
Addresses, server names, DNS names, host headers, paths and pod names are kept only as bounded top-K tables in memory. They are never used as labels on the agent's own OpenTelemetry export, because a label is a stored, indexed copy. A privacy test fails the build if any of them reaches a result sent to the platform or an exported metric.
The extension declares its privileges
iohr-capture ships two ways from one signed release: as a system package (deb and rpm)
with its service unit, and as an iohr extension, capture, built and signed by the
same pipeline as the agent (RFC 0028).
The package comes first, because iohr runs as the person and cannot grant
capabilities. The extension carries the commands a person types (iohr capture status|pcap|dissect) and controls the system service through the root-only socket and
the package manager.
The extension manifest gains a field, privileges, for example
["CAP_BPF", "CAP_PERFMON", "CAP_NET_ADMIN"]. iohr ext install shows it next to the
scopes, the way it shows scopes today, and iohr will not run an extension that declares
privileges until the person has confirmed them. This is a change to the command line's
security requirements and is recorded there.
On the platform side, the agents service stores the capture:* capabilities from each
hello and the console shows them on the agent's page, so a team can see which hosts can
answer a question about the wire before an incident asks it.
The console on your own machine
iohr console serves the console on the person's machine:
- It serves the console's static build on a random port on loopback and opens the browser.
- It forwards API calls with the access token of the active
iohrprofile.iohrrefreshes the token; the browser never holds it. No cookies are set and nothing changes in the platform's cross-origin rules. - It checks the Host header on every request, so a page elsewhere cannot rebind a name to loopback and borrow the session.
- It ships as an
iohrextension,console, so the console's version follows the platform's, not the command line's.
The console gains a local mode at build time: its API base is the local origin, and an
expired session says "run iohr login" instead of sending the browser to the sign-in
host. Switching teams is switching iohr profiles.
Entitlements. No API says today what an account may use: plans grant units and
nothing else (RFC 0015). The accounts service gains one read,
GET /v1/accounts/orgs/{org}/entitlements: the plan, its features and the extensions the
team may install. The local console shows it, iohr ext install checks it, and the
extension marketplace will use the same model.
Switching plans opens Stripe Checkout in the browser, as billing does today (RFC 0052). The return address is built from the request, so it comes back to the local console as well as the hosted one. The plan still changes only on Stripe's signed event.
Where this leads: analysis by a model, on paid plans
The end goal, and none of it is built: the capture's aggregates become something a model can read and reason over, inside the loop (RFC 0045), so a post-mortem starts from the wire instead of a guess. The owner set the direction on 2026-10-05:
- An entitlement,
capture.ai. Granted through the entitlements API above to every paid plan (Personal, Team and Enterprise), and always to InOrbit's own internal accounts and to platform admins whatever their plan. The Free plan never has it. - Read tools over MCP, behind new scopes. Tools that read capture aggregates are
added to the platform's MCP server (RFC 0050) behind new scopes such as
capture:read. Default deny: no token has the scope until a person grants it, and every call checks the account's own plan forcapture.aiat the moment of the call, so a downgrade takes effect on the next call. - The customer decides what leaves the host. Nothing reaches a model unless the customer chooses to send it, and then summaries and aggregates first. Raw packets and packet files are never given to a model. Our self-hosted models are the default (and the account's own model, RFC 0059, where one is set). A third-party assistant the person connects, such as Claude, reads only through the scope the customer granted, under the customer's own agreement with that provider: a data processing agreement, and a business associate agreement if health data is ever in scope.
- Said plainly. A person reading an analysis is told it was written by a model, as the EU AI Act's transparency rules require (Regulation (EU) 2024/1689, Article 50). The terms exclude medical use and credit scoring.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Later, in separate RFCs
- Sending capture data to the customer's own connections (RFC 0018): their collector, their bucket, their SIEM. That makes the data leave the host, so it needs its own rules first: a lawful basis on the customer's side, a data processing agreement where we touch it, and minimisation written into what is sent.
- Decrypted TLS (layer 6).
- The first build is a system service. A per-node DaemonSet whose aggregates reach the cluster's agent over an authenticated in-cluster channel is its own design step, before a chart.
Operating it
This section is for the person who installs and runs the companion. The full install guide and the developer guide live in the dataplane repository (the dataplane's install guide); this is the summary they expand.
Requirements
| Requirement | Why | If it is missing |
|---|---|---|
| Linux 5.8 or newer, built with BTF | compile-once programs and the BPF ring buffer | refused at start, with the kernel version found |
| TCX on 6.6 or newer | attachments that disappear with the process | on older kernels netlink filters are used, cleaned at start and stop |
CAP_BPF, CAP_PERFMON, CAP_NET_ADMIN |
load programs, read kernel state, attach to interfaces | refused at start, naming the missing capability; all three are dropped after attaching |
| Locked memory without a limit on 5.8 to 5.10 | those kernels charge BPF maps to the locked-memory limit; 5.11 and later charge them to the cgroup | refused at start on those kernels, with the setting to change |
iohr-capture doctor
One command checks every row above on the host it runs on and prints pass or fail per requirement, each failure with the exact fix: the kernel option, the line for the service unit, the capability to add to the container. It changes nothing. Support starts with its output.
Installing
- Packages and . The deb and rpm install the companion with its own service
unit, which grants the three capabilities and nothing else, allows the
bpfsystem call and netlink, and sets the locked-memory limit. Enabling capture is then the agent's policy switch and the companion's own configuration, both off by default. - A container image. Run with exactly the three capabilities added, host networking,
and read-only mounts of the kernel's BTF and of the host's process and cgroup
information (for owners), plus the seccomp profile the guide gives for the
bpfsystem call. Never--privileged: the guide shows the exact command, and the doctor fails a container that has more than it needs. - comes later, as a per-node DaemonSet, after its own design step (above).
Uninstalling
Removing the package stops the service, detaches every program and deletes any netlink filter it left on an interface, then checks that none remains. On TCX kernels the attachments are gone the moment the process exits. A crash leaves nothing a restart or an uninstall does not clean.
Plan
Each phase is a pull request with its checks; a release still needs the owner's two approvals.
| Phase | Where | What |
|---|---|---|
| 0. Toolchain | dataplane | an empty eBPF program built, loaded and run in the virtual-machine test; the eBPF crate outside the workspace members with a pinned nightly and bpf-linker; CI, the release build, packaging and the SBOM all green before any feature code |
| 1. First layers | core, dataplane | this RFC; a decision record in the dataplane repository for the privileged companion and its unsafe exceptions; the companion with layers 1, 2, 4 and 5; [work] capture and [capture] in the policy; iohr capture status, the admin page's Traffic section and the hello capabilities; the service unit; core's agents service storing and showing capture:* |
| 2. Payloads and timing | dataplane | layer 3 with packet files on request and tshark dissection; layer 7 |
| 3. Packaging | dataplane, sdk | deb and rpm; the capture extension with privileges in its manifest and the install prompt in iohr |
| 4. Local console | core, sdk | the entitlements API; the console's local mode; iohr console; team and plan switching |
| 5. Off the host | core | the later RFCs above |
| 6. Analysis by a model | core, sdk | capture.ai in the entitlements (every paid plan, InOrbit's internal accounts and platform admins; never Free); MCP read tools over aggregates behind capture:read, default deny, the plan checked on every call; summaries sent only by the customer's choice, self-hosted models by default; the AI notice in every surface that shows an analysis |
Phase 4 design: the entitlements API
The first part of phase 4, written before it was built (the repository's rule for a change to the contract across services):
- Features per plan are data. The accounts service's configuration holds the features
(a name and a sentence each) and, per plan, the features it grants. Every plan the
platform sells or falls back to (Free, Personal, Team, Enterprise) is named, even with
none; a plan or feature name the service does not know stops it at start. The first
feature is
capture.ai: Personal, Team and Enterprise, never Free. A plan made later by the platform's admin grants nothing until it is named here (default deny). - Why accounts and not billing's catalogue. The plan an account is on lives in the accounts service, which billing sets on Stripe's signed event, and the answer is needed where the plan is read, on every call that checks a feature. Billing's catalogue is the price proposal. A test keeps the two lists of plan names equal, so a plan cannot be sold without its features being decided.
- Our own accounts and the platform's admins get everything. The accounts that are InOrbit's own are the list the RFCs product already keeps (RFC 0035's internal spaces): the accounts service now reads the same list, and a test fails when the two differ. An admin's view of any account shows every feature, because an admin has every feature wherever they are.
- One read, one check.
GET /v1/accounts/orgs/{org_id}/entitlementsanswers the plan, the features (sorted) and the extensions the account may install (none yet; the field is there for the marketplace), and why: the plan, an internal account, or the platform's admin. A member of the account, a key or token of the account holdingaccount:read, or the platform's admin may read it; anyone else is answeredNOT_FOUND.HasFeatureanswers one feature for one account to the platform's own services only (named in configuration, never over HTTP), for the MCP tools and the capture to check on every call. - The console shows the features under the current plan on the billing page.
The agent that understands [capture] ships before the docs describe it. Older agents
refuse a policy with fields they do not know, and the release notes say so.
Repositories
inorbithr/dataplane: three crates beside the agent.iohr-capture-ebpfholds theno_stdkernel programs;iohr-capture-commonthe record types both sides share;iohr-capturethe user-space daemon. Also the service unit, the packages and the extension artifact.inorbithr/sdk: theprivilegesfield, its install prompt and security requirement, andiohr console.inorbithr/core: this RFC,capture:*in the agents service and its docs, the entitlements API and the console's local mode.
Verification
- The capture proves itself. In a network namespace joined by a virtual pair, a scripted client sends a known number of HTTP/1 requests, TLS handshakes with a known server name, DNS queries and gRPC calls. The companion must report exactly those counts per layer, with no drops at that rate.
- The bounds hold. A flood past every limit shows drops counted, memory flat and no crash.
- Every kernel it claims. Privileged tests run in virtual machines on 5.15 and a current 6.x kernel, never on a production host. Kernels from 6.6 attach through TCX links, which go away with their file descriptor; older ones use netlink filters, with stale ones removed at start and on stop, and the test checks nothing is left behind.
- The install guide works as written. It is followed step by step on a clean virtual
machine for each supported distribution, through
doctor, a capture and an uninstall that leaves no filter behind. - The agent stays unprivileged. Its unit does not change. The companion's unit is scored with .
- On our own edge. With the owner's agreement, capture runs on the host in front of the platform. TLS handshakes and server names are compared with 's own handshake counters, and requests on the plaintext hop behind with its request counters. The numbers are published in this RFC's log, including where they disagree.
Security and compliance
- A new privileged component. The companion is the first part of the agent's family that holds capabilities. It holds them only while loading, drops them before parsing, has its own threat model and security requirements in the dataplane repository, and is never reachable from the network or from a platform job. Under the EU Cyber Resilience Act it is a product component like the agent (Regulation (EU) 2024/2847): the same signed releases, SBOM, vulnerability statements and advisories (RFC 0029).
unsafe, by exception. The dataplane repository forbidsunsafe. The kernel program crate needs it for eBPF's context API (and the kernel's verifier checks every program before it runs), and the shared record types need one narrow exception to be read from a ring buffer. Both are recorded in the decision record, and nothing else gains one.- Licences. The kernel lets a program call many of its helpers only when the program
declares a GPL-compatible licence
(BPF licensing), so the kernel
program crate is
MIT OR GPL-2.0; the rest stays Apache-2.0.tsharkis never linked and never bundled: it is a separate program the person installs and runs. - Personal data. Capture runs only where the customer runs it, under the customer's policy, and in these phases stays on their host. We receive capability names, not traffic. Payloads exist only behind layer 3's own switch and only as files a person asked for, capped and deleted. Personal data never becomes a telemetry label.
- No claim. None of this makes anyone compliant with anything. It is designed so that a customer's own assessment has less to cover: data that does not leave the host does not need a transfer basis.
- The compliance registry needs entries in
inorbithr/compliancebefore the first host runs it: the companion as a new privileged component (access control, change management, vulnerability handling) and the processing of personal data on customer hosts (minimisation, retention, the later off-host RFC as the point where a processing agreement is needed).
Alternatives considered
Give the agent the privileges. One program, one service, nothing to coordinate. Every check, every proxied fault and every parser of untrusted input would then run in a process that can load kernel programs. The agent's lack of privileges is what a security team approves; we keep it.
Raw sockets or libpcap, without eBPF. Simpler and portable, and it copies every packet to user space before deciding anything, so the cost scales with the traffic. It also cannot see the owning process or the kernel's TCP state, which is half of what a post-mortem asks. It stays an option for hosts too old for the ring buffer.
Link Wireshark's dissector libraries. The best protocol coverage there is, inside our
process, and the GPL would then apply to the program that links them. A separate tshark
gives the same dissectors with the licence kept apart.
tshark only, no eBPF. Every layer from one tool. It is a capture-and-dissect tool
for a person at a terminal, not a bounded always-on counter: it has no owning process, no
kernel TCP state, and no budget that holds under a flood.
A kernel module. Everything is reachable from a module, and a bug in one takes the machine down. The eBPF verifier refuses a program that could, which is the property a production host needs.
Decision
Open. The owner decided on 2026-10-05:
- capture lives in a privileged companion; the agent stays unprivileged;
- the first build has layers 1, 2, 3, 4, 5 and 7, with full packets in memory and packet files only on request; decrypted TLS is left out;
- dissection uses
tsharkfrom the start, as a separate program, so the GPL never touches our code; - nothing is stored and nothing leaves the host in the first phases;
- the console runs on a person's own machine through
iohr console; - the full scope stays, with every edge connected: the extension declares its
privileges, the agent reports its capture capabilities, the platform stores and shows
them, and entitlements are one API the console,
iohrand the marketplace share; - the direction beyond these phases: capture data open to analysis by a model and over
MCP on paid plans only, through the
capture.aientitlement (every paid plan, InOrbit's internal accounts and platform admins, never Free), with what leaves the host chosen by the customer and never raw packets.
Still to decide: the default bounds (from the flood test), the retention of packet files, and the .
Publication
The dataplane repository carries the companion, its decision record, threat model,
security requirements, the install guide and the developer guide. The developer docs' agent section gains capture: enabling it,
each layer and what it proves, packet files and tshark, and what the agent reports.
The command line's reference gains iohr capture and iohr console, and the extension
docs the privileges field. The console shows each agent's capture capabilities and,
later, entitlements. Nothing on the site describes a layer as available before an agent
release that has it.
Status log
- 2026-10-05: opened, with the owner's decisions above. Phase 0 (the toolchain spike) started in the dataplane repository; nothing is built or released yet.
- 2026-10-06: the platform stores and shows an agent's capture capabilities. The agents
service keeps the latest hello's capability strings on the agent's row (at most 64, each
at most 64 characters from a small charset; the rest dropped and counted),
Agentanswers them inGetAgentandListAgents, and the console's agent page shows a "Traffic capture" row with the layers reported, "not available" when none. Unknown strings are kept and never used to authorize anything. No agent release reports them yet. - 2026-10-06: phase 4 started with the entitlements API, as designed above. The accounts
service answers
GET /v1/accounts/orgs/{org_id}/entitlements(the plan, its features, the extensions, none yet, and why) to the account's members, its keys and tokens withaccount:read, and the platform's admin; anyone else is answered as if the account did not exist.HasFeatureanswers one feature to our own services.capture.aiis on Personal, Team and Enterprise, never Free, and on our own accounts and for the platform's admin whatever the plan. The billing page lists what the plan includes. Nothing checkscapture.aiyet: the MCP tools and the analysis it will gate are not built. - 2026-10-06: phases 0 to 2 built in the dataplane repository (dataplane #13, #14, #17).
Phase 2 adds layer 3 and layer 7. Whole packets are off unless the companion runs with
its own switch; a packet file is made only when root asks, kept at mode 0600, deleted
after an hour by default and never sent anywhere;
iohr-capture dissectruns the host's owntsharkas the person who asked, which the packages suggest and never bundle. Request timing pairs requests with responses for HTTP/1, cleartext HTTP/2 and gRPC, per method and route template. Version 2 of the aggregates socket adds a lookup of one key, numbers only, rate limited per user. The security review's findings were fixed before the merge. The first release carrying the companion, v0.1.0-alpha.5, failed to build: its lock files did not follow the capture crates' version (fixed in dataplane #18 and #19). v0.1.0-alpha.6 shipped the same day, signed, with the checksums and the build provenance verified for the tarball and the image, and the image public. - 2026-10-06: the platform side rolled: agents store and show capture capabilities (#378),
and the entitlements API with features per plan (#422). The compliance registry records
the companion's data control as implemented with the signed release; its access control
stays partial (kernels before 6.6 and the per-node deployment are open);
tsharkis recorded as a suggested dependency that runs as the person who asked and updates with the host. - 2026-10-07: next are phase 3 (the
captureextension foriohrwith its declared privileges, the per-node deployment for clusters, a deb and rpm repository) andiohr console. Nothing checkscapture.aiyet. - 2026-10-07: Checked: #378 and #422 are live; the capture companion is not verified as built. Open.
- 2026-10-07: the first numeric entitlement. Each plan now gives numbers as well as
features:
[entitlements.limits]in the accounts service's configuration, asked by our services withGetAccountLimits(never over HTTP). The first is agents per account (Free 1, Personal 3, Team 10, Enterprise by contract, 50 for now), which the agents service enforces (RFC 0085). Our own accounts and the platform's admin get the largest number any plan gives. - 2026-10-08: Part RFC 0061.1 written: the entitlements answer carried to the agent as a signed
licence (ADR 0057, ADR 0062), for RFC 0100. Phase 4's
iohr consoleis marked for the owner's decision (RFC 0100, D3): the agent now serves the console.