This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0059open2026-10-05

The account's own model

An account owner can choose to run the platform's AI features on the account's own AI provider connection and a default model. Connecting a key changes nothing on its own, our self-hosted models stay the default, and the platform's own callers never leave them. Every answer says which model and whose provider produced it.

Problem

Every generation on the platform runs on models we host ourselves (RFC 0001): a fast tier and a deep tier, both open-weight models on our own hardware. That is the right default. It keeps prompts on our machines, and the cost per token is known. It also has limits a team notices quickly. The deep tier runs one answer at a time, and a team that already pays for a frontier model through its own Anthropic, OpenAI or Google account cannot use it for the agents, the workbench, the RFC drafting or the MCP tools that generate.

Since RFC 0044 an account can connect those providers with its own key. The connection works, and the platform cannot use it for any of that:

  • The only way in is CallAction. A connector action is one HTTP request with a ten-second deadline and a capped body. A provider's generate action answers one prompt, without streaming, without tools and without the conversation as turns. A generation from the deep tier takes minutes, streams, and an agent calls tools between rounds.
  • The model service has no engine for it. The llm service builds one engine per tier at start (one of our self-hosted engines or the test stub) from a reviewed configuration file. Nothing chooses an engine per request or per account.
  • Nothing records whose model answered. A generation's record and its answer carry the engine kind and the model name. A model that runs under a customer's own account at a provider needs to say so, both for the person reading the answer and for the record.
  • The meter would charge for it. The gateway prices a generation by the tokens in its answer's usage. Tokens the customer already paid their provider for would be charged again as our units.

The owner decided on 2026-10-04 how this must behave. The platform's AI features run on an account's own provider only by explicit choice. Connecting a key changes nothing. An owner of the account picks "Use for the platform's AI features" and a default model. Our self-hosted models stay the default. The platform's own callers never move: our services, anonymous visitors, the site guide, the Radar and every embedding.

Proposal

Four parts: streaming through connections, the choice, the model service, and disclosure.

1. Streaming through connections

A new internal server-streaming RPC, ConnectionsService/StreamGenerate, with no REST route. It speaks each provider family's streaming API and answers the same chunk shape for all of them.

  • Who may call it. Only the model service, speaking for product:llm through the product delegation RFC 0044 built. The connection must hold a live grant to product:llm naming its generate action, and it must belong to the account the request names. The model service forwards the verified claims of the person it acts for, and the connections service checks that their account is the connection's. Anything else is NOT_FOUND or PERMISSION_DENIED, as for any action.
  • The request. The conversation as turns (system, user, assistant with its tool calls, tool answers), the tools offered with their JSON Schemas, the model (empty: the default the owner chose), the longest answer, the temperature and whether the model may reason.
  • The answer. A stream of chunks: text, reasoning text, one whole tool call with its arguments assembled, and a final chunk with the model the provider says it used, the stop reason and the token counts the provider reported. Never invented: a provider that reports nothing leaves zeros, and the record says so.
  • Three dialects, chosen by the connector:
Dialect Connectors Streaming Tools Reasoning Tokens
OpenAI chat completions OpenAI, Azure OpenAI, Mistral, an OpenAI-compatible endpoint stream: true, SSE data: lines ending in [DONE] tools with function schemas; calls arrive as delta.tool_calls fragments by index reasoning_effort where the model takes it stream_options.include_usage: one last chunk with usage
Anthropic Messages Anthropic stream: true, SSE events message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop tools with input_schema; arguments arrive as input_json_delta thinking: {type: "adaptive"} and output_config.effort on current models message_start carries the input tokens, message_delta the output tokens
Gemini Google Gemini streamGenerateContent?alt=sse functionDeclarations; a call is a whole functionCall part thinkingConfig, thought parts marked thought usageMetadata on the chunks, the last one final

Each dialect is a parser fed one line at a time, the same shape as the llm engines' parsers, so each is tested on recorded fixture lines without a socket. A model that refuses a parameter (several current Anthropic models refuse a temperature, and refuse to turn reasoning off) has it dropped or mapped from what the provider's models API says about that model, not from a list we keep by hand. "Reasoning off" on such a model becomes the lowest effort, which is how our own gpt-oss tiers already behave. Sources: OpenAI's chat completions streaming and its OpenAPI document; Anthropic's streaming and extended thinking; Google's generateContent and thinking pages.

  • Bounds. The ten-second limit stays on CallAction. StreamGenerate has its own: a deadline set by the caller and capped in the connections service's configuration, a separate limit on the wait for the first byte, a cap on the bytes one stream may carry, and the host allow-list, address pinning and no-redirect rules of every connector call. Hosts are the connector's own, as before.
  • Nothing stored. The prompt and the answer pass through memory once. They are never sealed for a retry, written to the database, logged or put in an error. The use is recorded like any action's (who, which connection, generate, outcome, latency) with the input and output tokens the provider reported, which the history already has columns for, and the llm generation's id, so an account can match the two records.
  • Failure. A provider's 401 or 403 marks the connection needs_reauth, exactly as RFC 0044 does for every connector. A 429 becomes RESOURCE_EXHAUSTED that says the provider is limiting this key. A 5xx or a dropped stream is UNAVAILABLE before the first chunk, or an error item after it, matching the engine contract in the model service.

2. The choice

The choice belongs to an account, not to a connection, and there is at most one.

  • In the console, an AI connection's page offers "Use for the platform's AI features", with a model picked from the connection's list_models. The dialog says what changes before anything does: from then on the conversation, the agent's instructions, what its tools read and what it recalls from memory go to that provider, under the account's own agreement with it, for every person and key of the account. The site guide, the Radar and embeddings stay on our models. Choosing again moves the choice; "Use InOrbit's models" ends it.
  • On the command line, iohr connections use-for-ai <name> --model <id>, iohr connections use-for-ai --off and iohr connections ai-provider to show it, with --json like every command (repo inorbithr/sdk).
  • Who may choose. The account's owner: the person of a personal account, the owner of a team. Not a key.
  • What it does. One call, SetModelProvider, does three things in one transaction: it creates a product:llm grant on the connection's generate and list_models actions, with no expiry; it records the connection and the model as the account's provider; and it writes an audit line (who, which connection, which model, never a key). Ending the choice revokes that grant. Revoking the grant anywhere, deleting the connection or pausing it ends the choice too, so the grant and the choice never disagree. The grant goes through the rule against the trifecta like every grant: generate reads untrusted text and sends data out, and product:llm holds nothing that reads private data, so it passes.
  • Reading it. ResolveModelProvider(account), internal and for the model service only, answers the connection, its connector, the model and whether it is active. The model service caches each answer for 60 seconds. A StreamGenerate refused because the choice or the grant is gone drops the cached answer at once. The 60 seconds bound how long a change takes to be seen, not how long a revoked grant keeps working: the grant is checked live on every stream.
  • On the wire, a connection gains model_provider and default_model, so the console and iohr connections show say which connection the platform's AI runs on.

3. The model service

  • A connection engine. EngineKind gains connection. It implements the same Engine trait over StreamGenerate and passes the same conformance suite as the stub and our self-hosted engines. It declares tools. It never embeds.

  • Built per request, not from configuration. The code says otherwise than the plan assumed: engines are built once per tier at start by engine::build, the only place a kind is matched, from a reviewed file whose kind no flag may override. A connection engine cannot come from that file. build refuses connection as a configured kind, and the service constructs one per request from the resolved provider, a cheap value that holds a channel to the connections service and the provider's details.

  • The rule for which engine answers. A generation runs on the account's provider only when all of these hold:

    • the caller is a person or a key whose verified claims name an account, and that account has an active choice;
    • the caller is not one of our services. The model service reads callers today without the list of our own service subjects; it gains that list, so svc:radar, svc:llm and the chaos tool's load subject are recognised as ours;
    • the request did not come through the public route (an anonymous visitor);
    • the agent spoken to is not pinned to our models. An agent file gains platform_model = true, and the site guide sets it;
    • the request did not ask for our models. GenerateRequest gains platform_model, so a person can run one turn on our tiers, and a feature that must stay on ours says so.

    Everything else runs on our tiers as it does today. Embed, and the embedding behind an agent's memory recall, never use a connection.

  • Tier, admission and the deadline. A request on the account's provider still names a tier. The tier picks the deadline and is recorded, but its engine and its admission slots are not used: our GPU's queue is not the provider's. A generation on a provider waits in a bounded line of its own, per account, so one account cannot hold every stream the service runs. The provider's own rate limits apply on top and are reported as the provider's.

  • The record. llm.generations gains provider (the connector id, anthropic) and connection_id. engine is connection, and model is what the provider said it used. A row's build and weights come from the engine, not from the tier's probe, which the service reads today for every row (unknown for the weights of a hosted model, never invented).

  • The answer. GenerateResponse gains provider and connection_id on every chunk, beside engine and model. ListModels lists the caller's account provider after the tiers, so a client can show it before the first turn.

  • The daily token ceiling (a per-person abuse bound, not the budget) counts only tokens on our models. A provider's tokens are recorded and shown, not summed into it.

  • Units. The gateway's meter prices a generation by the usage it reads from the answer, by field name, without naming an RPC. It learns one more rule the same way: a message whose connection_id is set used someone else's model, so its tokens are not priced as our units. The call itself is still counted like any call.

  • What follows without more work. Agents and their tool loop, the workbench, MCP's generation tools, the RFC drafting of RFC 0035 and an agent's episode summaries all call Generate as the person or key. They run on the account's choice with no change of their own.

  • No silent fallback. A provider that fails fails the request, and the error says whose provider and what to do: "Your Anthropic connection needs a new key. Reconnect it, or switch the platform back to InOrbit's models." Answering from a different model than the one the account chose, even one of ours, is a different answer, and the account did not ask for it.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

4. Disclosure and compliance

  • Every answer says who answered. The workbench, the console, MCP tool results and the API show the model and, when it is not ours, the provider and the connection: "Answered by claude-sonnet-5-5 through your Anthropic connection prod". The AI Act's transparency rule, Art. 50(1), asks that a person is told they are dealing with an AI system, and it applies since 2026-08-02. Naming the model and whose provider ran it goes further than the article asks. We do it because a record that cannot say what produced it is not evidence (RFC 0046).
  • Prompts are customer data. They are never logged, traced, put in a metric label, a URL or an error, on our side of the call. What the provider keeps is governed by the account's own agreement and settings with that provider (retention, training use, the region), and the choice dialog says so.
  • The roles. As RFC 0044 reads control AI-03 of our compliance programme ("customer data reaches only self-hosted models or providers under a DPA"): the customer chooses a provider it contracts with itself, and we send to it at the customer's instruction. We act as the customer's processor. The provider is the customer's, not a sub-processor of ours, and is not added to our sub-processor list. Models we choose and pay for stay under the control as written. The control's note in the registry gains this reading, and our data processing terms gain the clause for customer-directed transfers before the choice is offered to anyone outside our own accounts.
  • Out of scope stays out of scope. The terms exclude health data and medical use, and credit scoring. The choice does not change either, and the dialog repeats it.
  • Text that changes. Wherever we say generation is self-hosted only, it becomes "self-hosted unless your account chose its own provider": RFC 0035's "where the code is read", the console's privacy page, the developer docs' models page. The site guide's paragraph on the legal page stays true as written, because the guide never moves; a new paragraph covers a signed-in account's choice. None of these texts says the platform is compliant with anything; it is aligned with the criteria, and an auditor decides the rest.

Plan

Proto.

  • iohr/connections/v1: StreamGenerate (server streaming, no HTTP option) with StreamGenerateRequest and StreamGenerateResponse defined in the package, not imported from iohr/llm/v1, so neither package depends on the other; ResolveModelProvider (internal); SetModelProvider, ClearModelProvider and GetModelProvider (REST under connections:write and connections:read); Connection.model_provider and Connection.default_model; Use.generation_id.
  • iohr/llm/v1: GenerateRequest.platform_model; GenerateResponse.provider and connection_id; ModelInfo.provider and connection_id.
  • buf lint with STANDARD rules: the streaming RPC has its own request and response.

Migrations, each made with mise run db:new:

  • model_provider: connections.model_providers (the account as the key, the connection with on delete cascade, the model, who chose and when) and generation_id on the uses table.
  • llm_generation_provider: provider (not null, default empty) and connection_id on llm.generations; the budget's sum leaves out rows with a connection. Erasure and export carry the new columns.

Code.

  • crates/connections: a connectors/stream module with the three dialect parsers and the streaming request under the existing host and address rules; StreamGenerate, the model-provider RPCs and their store; connector files gain the dialect and the streaming URL beside the existing generate action.
  • crates/llm: engine/connection.rs; the engine resolved per request in service.rs (Generate, the agent loop and EndSession); the provider cache; the services list; the per-account line; records and responses; platform_model on agent files, set on the site guide.
  • crates/protocol: the meter's rule for connection_id; the MCP generation tools carry the provider line in their result.
  • The console: the choice on an AI connection's page, admin-only first like every new product, and the provider line on answers. The command line in inorbithr/sdk.
  • Docs in the same change as the behaviour: the connections and models pages, the API reference from the OpenAPI document, the privacy text, RFC 0035's line, and the status lines of RFCs 0001 and 0044.

Tests, against real servers on an ephemeral port and a fake provider, never a mock of our own services:

  • A fake provider for each dialect replays recorded SSE through StreamGenerate: text, reasoning, a tool call split over many fragments, usage at the end, a stream cut in the middle, a 401 (the connection becomes needs_reauth) and a 429.
  • Only the model service may call StreamGenerate; a connection without a live product:llm grant is refused; a connection of another account is NOT_FOUND.
  • Nothing is stored: a sentinel in the prompt and in the answer appears in no table and no captured log line, while the use row carries the token counts.
  • The connection engine passes the model service's engine conformance suite.
  • End to end, the model service and the connections service on ephemeral ports with the fake provider: a person of an account with a choice gets the provider's chunks, labelled, and the record carries provider and connection_id; the same person with platform_model gets our stub tier.
  • Platform callers never reach the provider: svc:radar, an anonymous visitor through the site guide, Embed and memory recall leave the fake provider's request count at zero.
  • Ending the choice or revoking the grant stops the next stream, without waiting for the cache.
  • The meter prices a labelled answer's tokens at zero units and still counts the call; the daily token ceiling ignores provider rows.
  • mise run ci:changed before every push; roll:check before each roll.

Phases

Each phase is its own pull requests, with a line in the status log.

  1. Streaming through connections. StreamGenerate, the three dialects, the model-provider RPCs and their migration. Nothing a person sees changes.
  2. The model service. The connection engine, the rule for which engine answers, the records, the answer's fields, the meter and the ceiling. The console's choice and the command line, admin-only, used first on our own account (we are the first customer).
  3. Disclosure and the text. The provider line on every surface, the docs, the privacy text, the data processing terms and the registry's note. Then the choice is opened to account owners.
  4. Later. One model per tier, Amazon Bedrock once its SigV4 signing is built, and an OpenAI-compatible model server reached through the account's own agent in its network (RFC 0029) instead of the public internet.

Alternatives considered

CallAction with a longer timeout. One request, one answer, no stream. Every person would wait for a whole deep answer with nothing on screen, tools would need a second protocol, and the ten-second limit that protects every other connector would become a setting per action. A streaming RPC of its own keeps CallAction simple.

The model service calls the provider itself. The llm service would hold the provider's key for the length of a call. That puts a second place that opens customer credentials beside the vault RFC 0018 built, with its own host rules and its own audit. Connections already has the vault, the host allow-list, the address pinning, needs_reauth and the history; it gains a streaming path instead.

On whenever a provider is connected. Simpler, and it is what the owner ruled out. Connecting a key to let one agent call generate as a tool must not move every answer the account gets to a third party.

Choose per person, or per request. A per-person choice lets one member send the team's conversations to a provider the owner never approved. The owner decides for the account. A person can still go the other way and keep a turn on our models (platform_model), which sends less, not more.

A hosted router (OpenRouter, a provider's own gateway). One more company that sees every prompt, a sub-processor for every customer, and a second bill. The customer's own provider account is the point.

Fall back to our models when the provider fails. Answers would keep coming, from a different model than the one chosen, which the label would show but the account did not ask for. An error that names the provider and says what to do is the honest answer.

Decision

Open. The owner's rules (an explicit choice by an owner, our models as the default, the platform's own callers never moved) are fixed; the rest is proposed, built in the phases above. Questions for the owner:

  1. Units. Should a generation on the account's provider be free of units, as proposed, or carry a small charge per call for what we still run (the agent loop, its tools, the records)? A charge per call is easy to add later through the existing categories; removing a token charge after launch is harder to explain.
  2. One model or two. One default model for the account, as proposed, or one for the fast tier and one for the deep tier?
  3. The token advisor. It forwards the person's identity, so under the rule above it follows the choice, and RFC 0016's "our own model" line changes. Pin it to our models instead (recommended: it proposes token scopes, a security decision we would rather grade on a model we run and measure)?
  4. Our agents' instructions. On a provider, an agent's instructions travel with every turn, and a customer whose provider account keeps request logs can read them there. Treat our agents' instructions as not secret (recommended), or pin every agent whose instructions we would rather keep?
  5. Who chooses. Only the owner of a team, as proposed, or its admins as well?
  6. Events. Publish ai_provider.chosen and ai_provider.cleared to webhooks, so a customer's own compliance tooling sees the change, in phase 2 or later?

Publication

When phase 3 ships: the developer docs' connections page gains "Run the platform's AI on your provider", the models page lists what moves and what never does, and the API reference gains the model-provider routes and the new answer fields. The console's privacy page and the legal page gain the paragraph on an account's own provider. RFC 0035's text on where code is read changes in the same commit. The changelog says the day the choice opened. Each phase adds its line below.

Status log

  • 2026-10-05: opened. Measured before: a connector action has a ten-second deadline and no streaming; the model service builds its engines once per tier from configuration; GenerateResponse and the generation record name the engine and the model, not a provider; the gateway prices a generation by the usage in its answer.
  • 2026-10-07: Checked: No PR names this RFC and the status log records nothing built. Open.

← Back to Platform