Problem
Every generation on the platform runs on models we host ourselves (RFC 0001): a fast tier and a deep tier, both open-weight models on our own hardware. That is the right default. It keeps prompts on our machines, and the cost per token is known. It also has limits a team notices quickly. The deep tier runs one answer at a time, and a team that already pays for a frontier model through its own Anthropic, OpenAI or Google account cannot use it for the agents, the workbench, the RFC drafting or the MCP tools that generate.
Since RFC 0044 an account can connect those providers with its own key. The connection works, and the platform cannot use it for any of that:
- The only way in is
CallAction. A connector action is one HTTP request with a ten-second deadline and a capped body. A provider'sgenerateaction answers one prompt, without streaming, without tools and without the conversation as turns. A generation from the deep tier takes minutes, streams, and an agent calls tools between rounds. - The model service has no engine for it. The llm service builds one engine per tier at start (one of our self-hosted engines or the test stub) from a reviewed configuration file. Nothing chooses an engine per request or per account.
- Nothing records whose model answered. A generation's record and its answer carry the engine kind and the model name. A model that runs under a customer's own account at a provider needs to say so, both for the person reading the answer and for the record.
- The meter would charge for it. The gateway prices a generation by the tokens in its
answer's
usage. Tokens the customer already paid their provider for would be charged again as our units.
The owner decided on 2026-10-04 how this must behave. The platform's AI features run on an account's own provider only by explicit choice. Connecting a key changes nothing. An owner of the account picks "Use for the platform's AI features" and a default model. Our self-hosted models stay the default. The platform's own callers never move: our services, anonymous visitors, the site guide, the Radar and every embedding.
Proposal
Four parts: streaming through connections, the choice, the model service, and disclosure.
1. Streaming through connections
A new internal server-streaming RPC, ConnectionsService/StreamGenerate, with no REST
route. It speaks each provider family's streaming API and answers the same chunk shape
for all of them.
- Who may call it. Only the model service, speaking for
product:llmthrough the product delegation RFC 0044 built. The connection must hold a live grant toproduct:llmnaming itsgenerateaction, and it must belong to the account the request names. The model service forwards the verified claims of the person it acts for, and the connections service checks that their account is the connection's. Anything else isNOT_FOUNDorPERMISSION_DENIED, as for any action. - The request. The conversation as turns (system, user, assistant with its tool calls, tool answers), the tools offered with their JSON Schemas, the model (empty: the default the owner chose), the longest answer, the temperature and whether the model may reason.
- The answer. A stream of chunks: text, reasoning text, one whole tool call with its arguments assembled, and a final chunk with the model the provider says it used, the stop reason and the token counts the provider reported. Never invented: a provider that reports nothing leaves zeros, and the record says so.
- Three dialects, chosen by the connector:
| Dialect | Connectors | Streaming | Tools | Reasoning | Tokens |
|---|---|---|---|---|---|
| OpenAI chat completions | OpenAI, Azure OpenAI, Mistral, an OpenAI-compatible endpoint | stream: true, SSE data: lines ending in [DONE] |
tools with function schemas; calls arrive as delta.tool_calls fragments by index |
reasoning_effort where the model takes it |
stream_options.include_usage: one last chunk with usage |
| Anthropic Messages | Anthropic | stream: true, SSE events message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop |
tools with input_schema; arguments arrive as input_json_delta |
thinking: {type: "adaptive"} and output_config.effort on current models |
message_start carries the input tokens, message_delta the output tokens |
| Gemini | Google Gemini | streamGenerateContent?alt=sse |
functionDeclarations; a call is a whole functionCall part |
thinkingConfig, thought parts marked thought |
usageMetadata on the chunks, the last one final |
Each dialect is a parser fed one line at a time, the same shape as the llm engines' parsers, so each is tested on recorded fixture lines without a socket. A model that refuses a parameter (several current Anthropic models refuse a temperature, and refuse to turn reasoning off) has it dropped or mapped from what the provider's models API says about that model, not from a list we keep by hand. "Reasoning off" on such a model becomes the lowest effort, which is how our own gpt-oss tiers already behave. Sources: OpenAI's chat completions streaming and its OpenAPI document; Anthropic's streaming and extended thinking; Google's generateContent and thinking pages.
- Bounds. The ten-second limit stays on
CallAction.StreamGeneratehas its own: a deadline set by the caller and capped in the connections service's configuration, a separate limit on the wait for the first byte, a cap on the bytes one stream may carry, and the host allow-list, address pinning and no-redirect rules of every connector call. Hosts are the connector's own, as before. - Nothing stored. The prompt and the answer pass through memory once. They are never
sealed for a retry, written to the database, logged or put in an error. The use is
recorded like any action's (who, which connection,
generate, outcome, latency) with the input and output tokens the provider reported, which the history already has columns for, and the llm generation's id, so an account can match the two records. - Failure. A provider's 401 or 403 marks the connection
needs_reauth, exactly as RFC 0044 does for every connector. A 429 becomesRESOURCE_EXHAUSTEDthat says the provider is limiting this key. A 5xx or a dropped stream isUNAVAILABLEbefore the first chunk, or an error item after it, matching the engine contract in the model service.
2. The choice
The choice belongs to an account, not to a connection, and there is at most one.
- In the console, an AI connection's page offers "Use for the platform's AI
features", with a model picked from the connection's
list_models. The dialog says what changes before anything does: from then on the conversation, the agent's instructions, what its tools read and what it recalls from memory go to that provider, under the account's own agreement with it, for every person and key of the account. The site guide, the Radar and embeddings stay on our models. Choosing again moves the choice; "Use InOrbit's models" ends it. - On the command line,
iohr connections use-for-ai <name> --model <id>,iohr connections use-for-ai --offandiohr connections ai-providerto show it, with--jsonlike every command (repoinorbithr/sdk). - Who may choose. The account's owner: the person of a personal account, the owner of a team. Not a key.
- What it does. One call,
SetModelProvider, does three things in one transaction: it creates aproduct:llmgrant on the connection'sgenerateandlist_modelsactions, with no expiry; it records the connection and the model as the account's provider; and it writes an audit line (who, which connection, which model, never a key). Ending the choice revokes that grant. Revoking the grant anywhere, deleting the connection or pausing it ends the choice too, so the grant and the choice never disagree. The grant goes through the rule against the trifecta like every grant:generatereads untrusted text and sends data out, andproduct:llmholds nothing that reads private data, so it passes. - Reading it.
ResolveModelProvider(account), internal and for the model service only, answers the connection, its connector, the model and whether it is active. The model service caches each answer for 60 seconds. AStreamGeneraterefused because the choice or the grant is gone drops the cached answer at once. The 60 seconds bound how long a change takes to be seen, not how long a revoked grant keeps working: the grant is checked live on every stream. - On the wire, a connection gains
model_provideranddefault_model, so the console andiohr connections showsay which connection the platform's AI runs on.
3. The model service
A connection engine.
EngineKindgainsconnection. It implements the sameEnginetrait overStreamGenerateand passes the same conformance suite as the stub and our self-hosted engines. It declares tools. It never embeds.Built per request, not from configuration. The code says otherwise than the plan assumed: engines are built once per tier at start by
engine::build, the only place a kind is matched, from a reviewed file whosekindno flag may override. A connection engine cannot come from that file.buildrefusesconnectionas a configured kind, and the service constructs one per request from the resolved provider, a cheap value that holds a channel to the connections service and the provider's details.The rule for which engine answers. A generation runs on the account's provider only when all of these hold:
- the caller is a person or a key whose verified claims name an account, and that account has an active choice;
- the caller is not one of our services. The model service reads callers today without
the list of our own service subjects; it gains that list, so
svc:radar,svc:llmand the chaos tool's load subject are recognised as ours; - the request did not come through the public route (an anonymous visitor);
- the agent spoken to is not pinned to our models. An agent file gains
platform_model = true, and the site guide sets it; - the request did not ask for our models.
GenerateRequestgainsplatform_model, so a person can run one turn on our tiers, and a feature that must stay on ours says so.
Everything else runs on our tiers as it does today.
Embed, and the embedding behind an agent's memory recall, never use a connection.Tier, admission and the deadline. A request on the account's provider still names a tier. The tier picks the deadline and is recorded, but its engine and its admission slots are not used: our GPU's queue is not the provider's. A generation on a provider waits in a bounded line of its own, per account, so one account cannot hold every stream the service runs. The provider's own rate limits apply on top and are reported as the provider's.
The record.
llm.generationsgainsprovider(the connector id,anthropic) andconnection_id.engineisconnection, andmodelis what the provider said it used. A row's build and weights come from the engine, not from the tier's probe, which the service reads today for every row (unknownfor the weights of a hosted model, never invented).The answer.
GenerateResponsegainsproviderandconnection_idon every chunk, besideengineandmodel.ListModelslists the caller's account provider after the tiers, so a client can show it before the first turn.The daily token ceiling (a per-person abuse bound, not the budget) counts only tokens on our models. A provider's tokens are recorded and shown, not summed into it.
Units. The gateway's meter prices a generation by the
usageit reads from the answer, by field name, without naming an RPC. It learns one more rule the same way: a message whoseconnection_idis set used someone else's model, so its tokens are not priced as our units. The call itself is still counted like any call.What follows without more work. Agents and their tool loop, the workbench, MCP's generation tools, the RFC drafting of RFC 0035 and an agent's episode summaries all call
Generateas the person or key. They run on the account's choice with no change of their own.No silent fallback. A provider that fails fails the request, and the error says whose provider and what to do: "Your Anthropic connection needs a new key. Reconnect it, or switch the platform back to InOrbit's models." Answering from a different model than the one the account chose, even one of ours, is a different answer, and the account did not ask for it.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
4. Disclosure and compliance
- Every answer says who answered. The workbench, the console, MCP tool results and the
API show the model and, when it is not ours, the provider and the connection: "Answered
by claude-sonnet-5-5 through your Anthropic connection
prod". The AI Act's transparency rule, Art. 50(1), asks that a person is told they are dealing with an AI system, and it applies since 2026-08-02. Naming the model and whose provider ran it goes further than the article asks. We do it because a record that cannot say what produced it is not evidence (RFC 0046). - Prompts are customer data. They are never logged, traced, put in a metric label, a URL or an error, on our side of the call. What the provider keeps is governed by the account's own agreement and settings with that provider (retention, training use, the region), and the choice dialog says so.
- The roles. As RFC 0044 reads control AI-03 of our compliance programme ("customer data reaches only self-hosted models or providers under a DPA"): the customer chooses a provider it contracts with itself, and we send to it at the customer's instruction. We act as the customer's processor. The provider is the customer's, not a sub-processor of ours, and is not added to our sub-processor list. Models we choose and pay for stay under the control as written. The control's note in the registry gains this reading, and our data processing terms gain the clause for customer-directed transfers before the choice is offered to anyone outside our own accounts.
- Out of scope stays out of scope. The terms exclude health data and medical use, and credit scoring. The choice does not change either, and the dialog repeats it.
- Text that changes. Wherever we say generation is self-hosted only, it becomes "self-hosted unless your account chose its own provider": RFC 0035's "where the code is read", the console's privacy page, the developer docs' models page. The site guide's paragraph on the legal page stays true as written, because the guide never moves; a new paragraph covers a signed-in account's choice. None of these texts says the platform is compliant with anything; it is aligned with the criteria, and an auditor decides the rest.
Plan
Proto.
iohr/connections/v1:StreamGenerate(server streaming, no HTTP option) withStreamGenerateRequestandStreamGenerateResponsedefined in the package, not imported fromiohr/llm/v1, so neither package depends on the other;ResolveModelProvider(internal);SetModelProvider,ClearModelProviderandGetModelProvider(REST underconnections:writeandconnections:read);Connection.model_providerandConnection.default_model;Use.generation_id.iohr/llm/v1:GenerateRequest.platform_model;GenerateResponse.providerandconnection_id;ModelInfo.providerandconnection_id.buf lintwith STANDARD rules: the streaming RPC has its own request and response.
Migrations, each made with mise run db:new:
model_provider:connections.model_providers(the account as the key, the connection withon delete cascade, the model, who chose and when) andgeneration_idon the uses table.llm_generation_provider:provider(not null, default empty) andconnection_idonllm.generations; the budget's sum leaves out rows with a connection. Erasure and export carry the new columns.
Code.
crates/connections: aconnectors/streammodule with the three dialect parsers and the streaming request under the existing host and address rules;StreamGenerate, the model-provider RPCs and their store; connector files gain the dialect and the streaming URL beside the existinggenerateaction.crates/llm:engine/connection.rs; the engine resolved per request inservice.rs(Generate, the agent loop andEndSession); the provider cache; the services list; the per-account line; records and responses;platform_modelon agent files, set on the site guide.crates/protocol: the meter's rule forconnection_id; the MCP generation tools carry the provider line in their result.- The console: the choice on an AI connection's page, admin-only first like every new
product, and the provider line on answers. The command line in
inorbithr/sdk. - Docs in the same change as the behaviour: the connections and models pages, the API reference from the OpenAPI document, the privacy text, RFC 0035's line, and the status lines of RFCs 0001 and 0044.
Tests, against real servers on an ephemeral port and a fake provider, never a mock of our own services:
- A fake provider for each dialect replays recorded SSE through
StreamGenerate: text, reasoning, a tool call split over many fragments, usage at the end, a stream cut in the middle, a 401 (the connection becomesneeds_reauth) and a 429. - Only the model service may call
StreamGenerate; a connection without a liveproduct:llmgrant is refused; a connection of another account isNOT_FOUND. - Nothing is stored: a sentinel in the prompt and in the answer appears in no table and no captured log line, while the use row carries the token counts.
- The connection engine passes the model service's engine conformance suite.
- End to end, the model service and the connections service on ephemeral ports with the fake
provider: a person of an account with a choice gets the provider's chunks, labelled,
and the record carries
providerandconnection_id; the same person withplatform_modelgets our stub tier. - Platform callers never reach the provider:
svc:radar, an anonymous visitor through the site guide,Embedand memory recall leave the fake provider's request count at zero. - Ending the choice or revoking the grant stops the next stream, without waiting for the cache.
- The meter prices a labelled answer's tokens at zero units and still counts the call; the daily token ceiling ignores provider rows.
mise run ci:changedbefore every push;roll:checkbefore each roll.
Phases
Each phase is its own pull requests, with a line in the status log.
- Streaming through connections.
StreamGenerate, the three dialects, the model-provider RPCs and their migration. Nothing a person sees changes. - The model service. The connection engine, the rule for which engine answers, the records, the answer's fields, the meter and the ceiling. The console's choice and the command line, admin-only, used first on our own account (we are the first customer).
- Disclosure and the text. The provider line on every surface, the docs, the privacy text, the data processing terms and the registry's note. Then the choice is opened to account owners.
- Later. One model per tier, Amazon Bedrock once its SigV4 signing is built, and an OpenAI-compatible model server reached through the account's own agent in its network (RFC 0029) instead of the public internet.
Alternatives considered
CallAction with a longer timeout. One request, one answer, no stream. Every person
would wait for a whole deep answer with nothing on screen, tools would need a second
protocol, and the ten-second limit that protects every other connector would become a
setting per action. A streaming RPC of its own keeps CallAction simple.
The model service calls the provider itself. The llm service would hold the
provider's key for the length of a call. That puts a second place that opens customer
credentials beside the vault RFC 0018 built, with its own host rules and its own audit.
Connections already has the vault, the host allow-list, the address pinning,
needs_reauth and the history; it gains a streaming path instead.
On whenever a provider is connected. Simpler, and it is what the owner ruled out.
Connecting a key to let one agent call generate as a tool must not move every answer
the account gets to a third party.
Choose per person, or per request. A per-person choice lets one member send the
team's conversations to a provider the owner never approved. The owner decides for the
account. A person can still go the other way and keep a turn on our models
(platform_model), which sends less, not more.
A hosted router (OpenRouter, a provider's own gateway). One more company that sees every prompt, a sub-processor for every customer, and a second bill. The customer's own provider account is the point.
Fall back to our models when the provider fails. Answers would keep coming, from a different model than the one chosen, which the label would show but the account did not ask for. An error that names the provider and says what to do is the honest answer.
Decision
Open. The owner's rules (an explicit choice by an owner, our models as the default, the platform's own callers never moved) are fixed; the rest is proposed, built in the phases above. Questions for the owner:
- Units. Should a generation on the account's provider be free of units, as proposed, or carry a small charge per call for what we still run (the agent loop, its tools, the records)? A charge per call is easy to add later through the existing categories; removing a token charge after launch is harder to explain.
- One model or two. One default model for the account, as proposed, or one for the fast tier and one for the deep tier?
- The token advisor. It forwards the person's identity, so under the rule above it follows the choice, and RFC 0016's "our own model" line changes. Pin it to our models instead (recommended: it proposes token scopes, a security decision we would rather grade on a model we run and measure)?
- Our agents' instructions. On a provider, an agent's instructions travel with every turn, and a customer whose provider account keeps request logs can read them there. Treat our agents' instructions as not secret (recommended), or pin every agent whose instructions we would rather keep?
- Who chooses. Only the owner of a team, as proposed, or its admins as well?
- Events. Publish
ai_provider.chosenandai_provider.clearedto webhooks, so a customer's own compliance tooling sees the change, in phase 2 or later?
Publication
When phase 3 ships: the developer docs' connections page gains "Run the platform's AI on your provider", the models page lists what moves and what never does, and the API reference gains the model-provider routes and the new answer fields. The console's privacy page and the legal page gain the paragraph on an account's own provider. RFC 0035's text on where code is read changes in the same commit. The changelog says the day the choice opened. Each phase adds its line below.
Status log
- 2026-10-05: opened. Measured before: a connector action has a ten-second deadline and
no streaming; the model service builds its engines once per tier from configuration;
GenerateResponseand the generation record name the engine and the model, not a provider; the gateway prices a generation by theusagein its answer. - 2026-10-07: Checked: No PR names this RFC and the status log records nothing built. Open.