Problem
A company running a program on our platform today can see it from two sides, and neither side is connected to the other.
- The host agent (RFC 0029, RFC 0061) sees the wire. It knows the packets, the protocols, the owning process, TCP retransmits and resets. It does not know which function was running, which user action started the request, or what the program believed happened.
- The browser recording (Trails, RFC 0055) sees the person. It knows the clicks, rage clicks, page timings and errors. It records no network requests at all, so it cannot say which backend call made a page slow.
The program itself is invisible to us. Our SDKs (Go, Rust, TypeScript, Python, Java, C#, release 0.2.2 on 2026-10-05) already create OpenTelemetry spans, but only for their own calls to our API. OpenTelemetry is optional and never a hard dependency of the core package. The SDKs have no instrumentation for frameworks, no exporter, no runtime metrics, and no worker. Their security requirements say "the SDK sends nothing anywhere except the requests the caller makes" (SDK requirement SR-16).
So the question an on-call engineer asks at 2 AM, "what broke, for whom, and why", has no single place to be answered. Each witness tells part of the story. When they disagree, nobody notices: the program logs a 200, the person sees a spinner, and the wire shows resets.
The loop the platform is built around (RFC 0045) needs the observe step to produce evidence a second party can check. One witness is an opinion. Three independent ones that agree, or visibly disagree, are evidence.
Proposal
The three witnesses and their join keys
| Witness | Where it runs | What it knows | Join keys it carries |
|---|---|---|---|
| Browser (Trails) | the person's browser | actions, page timings, errors, request timing | trace id per request, route template, time |
| Program (SDK) | inside the company's process | spans, function calls, exceptions, runtime state | trace id, route template, service, process and container id, time |
| Wire (agent + capture) | on the host, beside the program | packets, protocols, TCP health, owning process | owning cgroup/container/unit, port, route template, time |
There are three join keys, and every component must produce them the same way:
- The W3C trace id (
traceparent, Trace Context Level 1). The SDK starts or continues the trace. The browser learns the trace id from the response (below). The wire reads it only where it can see it in plaintext. - The route template:
GET /orders/{id}, never/orders/8812?coupon=X. The SDK, the capture companion and Trails normalise paths with one shared rule set. The rules ship as conformance vectors in the SDK repository, and the capture companion's tests read the same file. That is how wire timing per endpoint and program spans per endpoint land on the same row. - The owner: the cgroup, container or service unit that owns the process. The SDK
reports
container.id, the pod UID (never the pod name),process.pidand the service name as resource attributes. The capture companion already maps sockets to owners by cgroup id, container id and pod UID (RFC 0061, layer 4). It has no container or pod names, which would need the runtime's API.
The route template rules are exact, because three implementations (six SDKs, the capture companion, the Trails recorder) must agree byte for byte:
- the query string and fragment are dropped;
- the path is percent-decoded once and compared case-sensitively;
- a trailing slash is kept as written;
- a segment becomes
{id}when it is all digits, a UUID, hex of 16 characters or more, a date (2026-10-06), a token-like run of 20 or more characters from the base64url alphabet containing at least one digit, or contains@(an e-mail address is personal data and never part of a route); - after 8 segments the rest becomes
{rest}.
Where a framework knows the route it matched (/orders/{orderId}), the SDK uses that
route as-is and the rules apply only where it does not. The vectors (input, expected
template) live in the SDK repository's conformance suite as the source. The dataplane
repository vendors them with a checksum and a CI check that the copy matches.
Time alone is never a join key. It only narrows a join made on one of the three.
SDK: from API client to instrumentation
Each SDK gains an instrument entry point in a companion package, beside the core
package and separate from it. The core package keeps no OpenTelemetry dependency, as the
SDK's configuration ADR (ADR 0015, decision 10) requires. The model is the Go module
github.com/inorbithr/sdk/go/otel, which already connects the client's tracing and
metrics to OpenTelemetry as a module of its own; each language gets the equivalent
package. A program that only calls our API installs nothing new and is unchanged.
inorbit.instrument(service="checkout", environment="production") # proposal, not released
It sets up the OpenTelemetry SDK for the process. It is built on the upstream OpenTelemetry SDK and its maintained instrumentation libraries in each language, not on a fork. We add five things upstream does not have:
Instrumentation for frameworks, as presets. We use "instrumentation for frameworks" for the server side, because in the SDK "middleware" already means the outbound client pipeline (ADR 0015, decision 7). Upstream instrumentation is used where it exists (for example
otelhttpin Go, the Flask/FastAPI/Django instrumentations in Python, ASP.NET Core in .NET). We wrap it so it applies our route normalisation and privacy rules, and we write our own only where upstream has nothing. The first set covers the server side of HTTP and gRPC, the outbound HTTP client, database drivers (statement shape only, never parameters), queue consumers and producers, and scheduled or background jobs. The full list per language is in the plan, and every preset has a conformance case.Function tracing. A decorator, attribute or wrapper that turns one business function into a span with its outcome, for the code no framework sees. Examples:
@tracedin Python,#[inorbit::traced]in Rust,inorbit.Trace(ctx, "name", fn)in Go,[Traced]in C#. Arguments are never recorded. A function may name attributes explicitly, and those go through the same privacy filter.A privacy processor that is always on. It runs as the last span and log processor before export and cannot be removed by configuration. No request or response bodies. No query values. No headers except an allowlist. Database statements as shapes only. Attribute values cut to a fixed length. Exception messages scrubbed with the same rules as Trails error messages. A company can add rules; it cannot loosen the floor. This is SR-13 to SR-15 and SR-30 extended from our client to everything the SDK exports.
The worker.
instrumentstarts one background worker per process, andshutdownflushes and stops it. It owns:- the export queue, bounded in memory with a counted drop when full, never blocking the program;
- runtime metrics per language (heap, GC pauses, goroutines, threads or event-loop lag, open file descriptors);
- a heartbeat every 30 s with the service, version and deploy id, so a silent service is distinguishable from a quiet one;
- uncaught exception and panic hooks;
- the bridge from the language's standard logger to OpenTelemetry logs, carrying trace and span ids.
The worker is a thread or task inside the program, not a separate process. Programs that cannot run one (short-lived functions, CLIs) flush on exit instead.
Join attributes and response timing. Resource attributes carry the owner keys above. The server-side instrumentation for frameworks adds one
Server-Timingresponse header entry,traceparent;desc="00-…", so a page can read the trace id of each request from the Performance API without wrappingfetchand without reading headers (W3C Server Timing, Working Draft of 2026-04-07). It is off by default, on per service, and it only answers requests from the company's own origins (Timing-Allow-Origin). Trace Context Level 2 warns against leaking trace ids to cross-origin callers.
Configuration follows the SDK's existing rule (ADR 0015, decision 3): the SDK reads
the command line's config.toml, and INORBIT_CONFIG_FILE overrides its location. The
keys that are ours (destination, privacy additions, presets, the worker, Server-Timing)
go in its [sdk] table. The same keys exist as INORBIT_* environment variables,
following the SDK's existing precedence (code, environment, file, defaults).
iohr sdk config prints the effective result. OpenTelemetry's declarative configuration
(OTEL_CONFIG_FILE, YAML) is a separate and optional input. Where a program sets it, it
configures the OpenTelemetry pipeline (exporters, sampling, processors) and nothing of
ours. A company
that already runs its own telemetry collector points the SDK at it and uses none of our
destinations.
Requirement SR-16 changes from "no telemetry" to "no telemetry the caller did not
configure". The core package still sends nothing on its own. instrument exports only
to a destination the program names, and the default is none. This needs a new SDK ADR,
ADR 0016, which amends SR-16 and adds requirement SR-33 for the companion package's
export. In the same change, the SDK's CRA mapping (controls.md, Annex I part I.2(b),
secure by default) lists SR-33 beside SR-16, and the threat model row "SDK reports usage
or data somewhere else" names both.
Release. The SDK releases per component: release-please opens a separate release
pull request for each language, and the Go module go/otel has its own. Today Rust pins
opentelemetry 0.32 and Java pins Jackson 2.x. Version 0.3.0 moves them to the current
majors, and under the SDK's semver rules before 1.0 such a breaking change bumps the
minor version, so 0.3.0 comes from those dependency breaks in each component that has
one. The instrument packages themselves are additive. C# and Java are tagged but not
yet published to a registry, because they have no registry users yet; they are kept in
step with the other four all the same. Migration notes ship with each component that
breaks.
Where the data goes: two paths, one format
Everything is OTLP (OpenTelemetry Protocol 1.11): traces, metrics and logs, protobuf over gRPC or HTTP.
SaaS path. The SDK exports to our API (OTLP's standard paths /v1/traces,
/v1/metrics and /v1/logs under a telemetry prefix, or OTLP/gRPC) with an API key
holding the new scope telemetry:write. A new telemetry service on the platform:
- implements the OTLP collector services (
TraceService/Exportand its metrics and logs siblings) over gRPC. The edge routes gRPC by service name, so this is one more route. OTLP/HTTP protobuf gets a dedicated handler in the gateway that decodes the body and forwards the same gRPC call (translate and forward, no logic). OTLP/HTTP JSON is accepted only once the OTLP specification marks it stable; - re-applies the privacy floor on arrival. A client we did not write, or an old SDK,
must not be able to store what ours would have removed. Rejected items are counted in
OTLP's
partial_success, never silently dropped; - stores per account in our column store, the one we already run for high-volume rows with a time to live (the request log, the ledger's events). Not in our own internal trace store, which has no tenancy;
- keeps spans and logs 30 days and metrics 90 days by default, with erasure through
EraseSubjectlike every service, and export throughExportSubject; - meters by volume. Today units are priced per call or per 1,000 tokens. Telemetry needs a per-volume basis (per million spans, per GB of logs). That is a pricing decision for the owner before launch, not something this RFC sets.
Agent path. On a host that runs our agent, the SDK finds the agent's local socket (a Unix socket on Linux and macOS, a named pipe on Windows) and exports there instead.
This path changes the agent's founding rule. Today nothing but timings, codes, classes
and counts leaves the machine (RFC 0029, the agent's ADR 0001). Our processing record
says anything more needs its own RFC and a data processing agreement (Art. 28 GDPR)
first (processing record PRV-01). For program telemetry, this RFC is that RFC. It
amends ADR 0001, widens control DAT-09 (what the agent sends from a customer's network)
and adds a control entry for telemetry minimisation and forwarding. [work] traces may
not be switched on for any customer before both exist. What crosses the host boundary
depends on the policy's mode:
[traces] mode |
What leaves the host | Old rule still holds | Needs the DPA |
|---|---|---|---|
metrics (the default once traces are on) |
rate, errors and duration per service and route template, as numbers | no: route templates and per-route numbers from a customer's program are new content, so DAT-09 is widened first | no |
spans |
spans after the privacy floor, names and routes included; never the attributes joined from capture | no | yes |
local |
nothing; spans go only to the company's own collector | yes | no |
The receiver:
listens on a Unix socket only, mode 0660, owned by a dedicated group (
iohr-telemetry). It is not the agent's own group, which can read the agent's configuration; the capture companion moved its socket to its own group for the same reason. Peers are checked withSO_PEERCRED. The agent's rule "no new listener" is kept for the network: there is still no new TCP port;runs as its own unprivileged process, not inside the agent. It parses protobuf from any local program, so it is untrusted input next to the process that holds the platform credentials. It is bounded like the capture sockets:
- 4 MiB per message, a cap on attributes per span and on spans per batch;
- protobuf decode limits, at most 16 connections at once, read and idle timeouts;
- a counted drop when any limit is reached.
It hands the agent minimised data over the same local channel the agent already uses for extensions;
applies the company's
policy.tomlbefore anything leaves: the new[work] tracesswitch (off by default) and a[traces]table with the mode and drop rules. Agents older than the release that adds it refuse the key and do not start, the same note[work] capturecarries in the policy documentation. This is the company's control point, and the reason a regulated customer may accept the agent path where it would not accept direct export.
Joining program spans with what the wire saw does not let the agent read the capture companion's tables. The companion refuses names, paths and addresses to the agent's user, and stays the enforcer of the host-boundary rule. Instead the capture aggregates socket gains a keyed lookup (protocol version 2). The agent sends an owner key and a route template taken from a span. The companion answers with counts for that key only: requests, retransmits, resets, RTT and a latency histogram. It never answers with a list. Counts are already allowed to the agent.
The attributes this lookup adds (RTT, retransmits, resets for an owner and route) are traffic captured on the host, and control DAT-10 keeps captured traffic on that host. We therefore keep capture-joined attributes host-only in every mode: the agent uses them for the join and for the company's own collector, and strips them from anything it forwards to the platform. We recommend this over widening DAT-10, which would end the capture companion's central promise. The comparisons that need the wire's view therefore run on the host or in the company's own collector; whether any result of them may reach the platform is an open question below.
Forwarding goes to the platform's telemetry service, to the company's own collector, or
to both. The forwarder buffers on disk, because a host losing its uplink is exactly when
telemetry matters:
- mode 0600 in the agent's state directory;
- capped in size and age, with counted drops;
- erased on uninstall.
It holds customer data at rest on the customer's host, so it gets its own row in the agent's threat model.
The agent reports trace:otlp in its hello capabilities when the receiver runs, and
reports nothing about traces when the policy has them off. Like capture:*, this is what
the host can do, never used for authorization.
The agent path is for hosts first: service units and containers that share a kernel with the agent. On a container cluster the agent runs once per cluster today, so a per-node socket cannot reach it. That needs the per-node design RFC 0061 also defers.
An SDK that finds no socket and no configured destination exports nothing. It never guesses.
Programs we cannot change. For a language we have no SDK for, or a binary nobody
will rebuild, OpenTelemetry eBPF Instrumentation (OBI) is a later option. If it is
adopted, it is its own component and decision, never part of the capture companion:
it needs uprobes, and it injects traceparent into traffic, while the companion
promises never to modify or drop a packet. OBI derives HTTP and gRPC spans from the kernel and propagates traceparent. It
was beta and at v0.x in 2026, and its maintainers expect breaking changes before 1.0.
OBI cannot inject trace context into TLS traffic except between two OBI-instrumented
services. This RFC names it as a later option behind its own decision, not part of the
first plan.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Trails: the browser learns the trace id
Trails gains one event kind:
req { m: method, r: route template, s: status, d: duration ms, tid: trace id }
It is built from PerformanceResourceTiming entries and their serverTiming, through
the recorder's existing performance observer. It needs no fetch wrapper, no request
headers and no bodies. The rules:
- Origins. Only entries whose origin is on an explicit list of the company's own
origins are kept. The site configuration passes the list, as it passes element names
today. Everything else is dropped before any field is read. A third party may send
Timing-Allow-Origintoo, so the list is ours, not theirs. - Fields.
ris a route template with query keys only, through the same scrubber as page events, never the full URL.dis the duration.sis the response status where the browser exposes it.tidcomes only from a Server-Timing entry named exactlytraceparentthat matches the W3C format (00-32 hex-16 hex-2 hex). It is reduced to the trace id; the span id and flags are dropped.
- Volume. A page can make hundreds of requests, so
reqhas its own per-kind cap within the recorder's byte budget. Drops are counted like every other kind. - Browsers. Current Chromium, Firefox and Safari expose
serverTiming. On a cross-origin API host (a site onwww.callingapi.) the entries are empty unless the API sendsTiming-Allow-Originfor that site, so the SDK's instrumentation for frameworks sends it for the configured origins only. The trace id is visible to any script on the page. That is acceptable, because a random trace id is not a secret and identifies no one, but the consent text says so.
This changes RFC 0055's rule "never reads headers" into "reads the Server-Timing entry of our own responses". It needs these changes, in the same change and behind the same consent as the rest of Trails, whose legal basis is still pending the owner's decision:
- the consent text: "timings of requests to our own servers and their trace ids";
- the browser's scrub rules and their tests;
- the trails service's never-recorded check on upload for
req: no query values, no third-party origins; - the extension, refreshed to the same rules;
- the Trails event table in its documentation;
- the processing record for Trails.
Disagreement is the signal
The platform answers four questions from the joined data. Each one is a comparison between witnesses, not a guess by a model:
| Disagreement | Likely meaning |
|---|---|
| The browser saw a slow or failed request; the program's span for the same trace id was fast and fine | the time went somewhere between: CDN, proxy, load balancer, queueing before the handler |
| The program answered; the wire for that owner shows resets or retransmits in the same minute | network or kernel trouble under a healthy program |
| The wire counts requests to a route the program has no spans for | an uninstrumented path, a sidecar, or traffic that never reached the handler |
| Monitors (RFC 0037) say up; the browser shows rage clicks and errors on one route | a partial outage that synthetic checks do not exercise |
Detection works per route template and per minute, from rates and percentiles, against thresholds the company sets, and against its own baseline once one exists. A model may describe a finding in words, on our self-hosted engine, labelled as AI-written (AI Act Art. 50). It never decides one.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
From signal to the phone at 2 AM
A finding (a disagreement, or a plain threshold such as an error rate per route) opens a draft incident (RFC 0058, phase 2). It is deduplicated per service and route, and resolution is suggested on recovery.
The incident carries its evidence as links, never pasted content:
- the trace ids that show it;
- the route and time window;
- the deploy id from the heartbeat (what changed);
- the wire's counts for the owner;
- the Trails recordings that hit it.
This is RFC 0058's typed evidence, extended with
traceandtrailkinds.incident.openedgoes out on the events catalogue (RFC 0022). From there it reaches:- the company's paging tool through connections we already have as data: incident.io
alert events (
POST /v2/alert_events/http/{source}, deduplicated by key) and PagerDuty Events (trigger, acknowledge and resolve by dedup key); - our own push to the mobile app, which does not exist yet (no push, no incident screen; a separate RFC);
- mail and the console bell, which exist today.
- the company's paging tool through connections we already have as data: incident.io
alert events (
The person opens the incident on the phone and sees one page:
- what is broken, for which route, since when, and how many people it hit (counted from Trails and spans, not estimated);
- what each witness saw, and where they disagree;
- what changed just before.
The incident record holds what an incident report needs. Its timeline is built from the events, not typed afterwards. It has severity, impact counted from evidence, the classification fields DORA asks of a financial entity's major-incident report, and an export. It becomes an evidence record (RFC 0046) when signing exists. We do not file reports, and the record does not make the customer compliant.
Access and limits
- Default deny.
telemetry:writeandtelemetry:readare new scopes. Writing is limited to API keys made for it; reading is limited to the account's owners and admins and keys with the read scope. Both are gated at the edge like every scope. - Bounded. Per-request body limit (4 MiB, OTLP batch size), per-account rate and volume limits, and the plan's entitlement (RFC 0061's entitlements API, which is not built yet). Free plans get no telemetry ingest.
- Hidden first. Admin-only until the owner opens it, like every new product.
Plan
Each step ships on its own and is useful without the next. It runs on our own platform first: our services are instrumented with our SDK before any customer is.
This RFC, and SDK ADR 0016 amending SR-16 and adding SR-33, with the SR-16 rows in the SDK's
controls.md(CRA I.2(b)) andthreat-model.mdupdated. Shared route-template vectors in the SDK conformance suite, read by the capture companion's tests.SDK 0.3.0, released per component:
- the breaking dependency updates (Rust
opentelemetrypast 0.32, Java Jackson past 2.x); - the
instrumentcompanion package in each language, with the privacy processor, the worker, function tracing, server and client presets for HTTP and gRPC, and declarative configuration; - export to any OTLP endpoint the program names.
It works with a company's own collector on day one and needs nothing from our platform. Conformance cases extend to the instrumentation for frameworks and export: a local OTLP sink on the replay server, parent and child links, span kinds, exceptions.
- the breaking dependency updates (Rust
Platform ingest: the
telemetryservice, the gateway's OTLP/HTTP handler, the scopes, column-store storage, retention, erasure and export. A console page per service (routes, rates, errors, durations, traces) for admins only.Agent path, hosts first:
- before any of it: the ADR 0001 amendment and the control entry for telemetry forwarding;
- the receiver as its own process on a Unix socket, the
[work] tracesand[traces]policy, the disk buffer, forwarding; - the capture socket's keyed lookup (protocol v2) for the join;
- the
trace:otlpcapability.
Clusters wait for the per-node design.
Trails
reqevents, Server-Timing in the SDK's instrumentation for frameworks, and the RFC 0055 and consent updates.Findings and incidents: the four disagreements and threshold rules, draft incidents with evidence, and delivery through incident.io and PagerDuty connections. Mobile push follows in its own RFC.
Presets beyond the first set (more frameworks, queues, ORMs per language), and the OBI decision.
Security and compliance
Customer data. Telemetry from a customer's program is personal data more often than not: user ids, IP addresses, paths with names in them. We act as processor under the customer's DPA. The privacy floor is enforced twice (SDK and ingest) and on the agent path a third time by the company's policy. It is never a label in our own metrics, never in our logs, and never sent to a third-party model.
Health and card data stay out of scope by design. Program traces from a health or payments customer could carry PHI or card data despite the floor (a patient id in a path segment the normaliser did not catch, for example). The terms exclude both. The privacy floor's defaults (no bodies, no query values, normalised paths) are what keep that exclusion true in practice. A customer who needs either is a decision for the owner (a BAA, PCI scope) before onboarding, not something this RFC opens.
Controls. Existing controls are amended where they already cover the ground, and only what is new gets a new id:
- the SDK's promise to send no telemetry is control DAT-04. It is amended to match SR-16 as amended and SR-33: the core package sends nothing, and the companion package exports only to a destination the program names;
- ingest maps to DAT-03 (retention: 30 days for spans and logs, 90 for metrics),
DAT-04 (no customer data in our own logs, URLs, errors or metric labels) and DAT-06
(erasure through
EraseSubject); - the agent path widens DAT-09 for
metricsmode, because per-route metrics from a customer's program are new content leaving the host; - DAT-10 is unchanged: capture-joined attributes stay on the host in every mode;
- one new entry, for telemetry minimisation and forwarding on the agent path (the
receiver, the policy modes, the disk buffer). It is required before
[work] tracesis on for any customer, and thespansmode needs the customer's DPA.
The record of processing (control PRV-01) gains entry PA-03 "program telemetry" in
registers/processing.md, with the customer as controller and us as processor.The CRA. The SDK and the agent are products with digital elements. The new exporter, the worker, the agent's receiver (a protobuf parser fed by any local program) and its disk buffer (customer data at rest on the customer's host) enlarge the attack surface. Each gets a threat-model row, and the receiver's decoder is fuzzed before release.
Alternatives considered
- Our own tracing format and protocol. Rejected. OpenTelemetry is what enterprise programs already emit, OTLP is what their collectors already speak, and a company must be able to leave with its instrumentation intact.
- Wrap only upstream instrumentation and add nothing. Rejected. Without the always-on privacy floor, the shared route normalisation and the join attributes, the three witnesses cannot be joined, and customer data reaches us unfiltered.
- Write all framework instrumentation ourselves. Rejected for frameworks upstream already covers: six languages times dozens of frameworks is a maintenance load we cannot carry, and upstream follows the semantic conventions closely.
- A TCP OTLP receiver on the agent, on the standard OTLP ports on the loopback address. Rejected for the first version. Any local process could write to it, and it breaks the agent's "no new listener" rule. A Unix socket gives peer credentials and file permissions for free. Containers that cannot mount the socket export to the company's collector or directly to SaaS instead.
- Inject
traceparentfrom the page in Trails. Rejected for the first version. It changes CORS preflights for every API call the page makes. Reading Server-Timing needs nothing from the page's own code. - eBPF-only (OBI) instead of SDKs. Rejected as the main path. It is beta, it loses context across TLS, and it cannot see functions, exceptions or business attributes. It stays the fallback for programs nobody can change.
Open questions
- Pricing basis for ingest (per million spans, per GB) and how it maps to units. This is the owner's decision.
- Whether the agent's disk buffer has one size for every company or a policy key with a ceiling.
- Data residency: one region today. Enterprise and DORA customers will ask, and this RFC does not answer it.
- Whether a finding computed on the host from the wire's view (which disagreement, for which service, route and minute) may be forwarded to the platform, and under which control, given that DAT-10 keeps captured traffic on the host.
- Whether the platform's own services move from our internal trace store to the
telemetryservice, as the first customer of it.
Status log
- 2026-10-06: RFC opened, nothing built. Reviewed in draft by the agent and capture owner (the off-host rule, the keyed lookup in place of tables, a separate receiver process, exact route rules, hosts before clusters, OBI outside the companion) and the Trails owner (origin list, field rules, per-kind cap, browser caveats); both reviews are folded in. Research on the current SDKs, the agent and capture, Trails, incidents and connectors is the basis of the Problem section.
- 2026-10-07: Checked: nothing is built (#431 is text). Open.