This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0038open2026-10-03

Signals, the platform tells you before you look

A notification bell, health alerts by mail, a per-call request log and webhook recovery, on the events and counts the platform already keeps.

Problem

The console's dashboards (RFC 0034) answer questions, but only when someone opens them. An account whose error rate tripled overnight, a webhook endpoint the platform turned off after five days of failures, a key that started timing out: each is visible on a chart and told to nobody. When a developer does find a bad number, the next question is "which call?", and today nothing below a daily count can answer it. And an endpoint that comes back after an outage has no way to get what it missed: events published while it was off were never queued for it.

The pieces are already there. Every account's events are kept for 30 days with ids only, and webhooks, MQTT and the stream already carry them. Every call is counted with its outcome and answer time. Every answer carries a request id. What is missing is the last step from data to a person.

Proposal

Notifications

A bell at the top of the console shows the active account's notable events, newest first, with an unread count: units at 80 and 100 % of the allowance, a health alert, a webhook endpoint turned off, a member who joined or left, a key or token revoked, a domain that lost its proof, an agent revoked, a connection deleted. Routine events (a call through a connection, an agent connecting) are never in it. Each notification is a sentence built from the event's type and ids, with a link to the page it is about.

The notifications are the account's events, so there is no second store of what happened: GET /v1/events/notifications lists them (paged as every list is), and a read marker per person and account says what that person has seen. The bell follows new ones live over the same stream integrations use.

Health alerts

Once a day at most for each, per account and per key, the platform raises usage.health when today's calls cross a line: an error rate above 5 % over at least 50 calls, a p95 answer time above 5 seconds, or ten failures of the platform's own. The event goes where every event goes (webhooks, MQTT, the stream, the bell), and the account's owners and admins get a mail that says what crossed and links to the usage page. Those are the defaults; each account sets its own lines, the calls a day needs before a rate counts, which kinds may fire, and whether its owners and admins are mailed at all (off, the bell and the webhooks still hear of it).

When the platform turns a webhook endpoint off (five days of failures, or a 410 Gone), it publishes webhook.endpoint.disabled and mails the owners and admins. Neither mail carries anything from the calls themselves.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

The request log

Every counted call is kept for 30 days as one row: when, which key or person, which operation, how it ended (ok, the caller's error, the platform's failure, with the status code), how long it took, and its request id. Never a request or answer body. The usage page gains a request log filtered by key, outcome, operation or request id, and the error figures link straight into it. A support question can name one request id, and the developer can find that call themselves.

The rows live in a column store built for this volume, not in the accounts database, and are deleted with the account or the person.

Recovery

POST /v1/webhooks/endpoints/{endpoint_id}/recover with a time in the last 30 days sends an endpoint everything it missed since then: every delivery that failed for good is due again, and every event of its types that was never sent to it, because it was off, is sent once, oldest first, with its original event id. At most 10,000 a call; the answer says when to call again. The endpoint page has a "Send what it missed" button, and a disabled endpoint's banner turns it on and recovers in one step.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

Alternatives considered

  • A separate notification store. A table of rendered messages per person would duplicate what the events already say and need its own retention and erasure. Reading the account's events, with one read marker per person, keeps one source.
  • Alerts on every metric, configurable from day one. Thresholds per account and per key, quiet hours and channels are where this ends up; a small fixed set first shows which alerts people act on.
  • The request log in the accounts database. One row per call grows with traffic, not with accounts; a column store with a time-to-live keeps 30 days cheaply and drops the rest by itself.
  • Retry only what failed. Simple, and wrong after an outage: events published while an endpoint was off were never queued for it, so a retry of failures would miss them.

Decision

Open. Recovery is built first, then the bell and the alerts, then the request log.

Publication

The console gains the bell, a notifications page, the request log and the recover button. The developer docs gain the notifications route, the two new event types, the request log route and recovery. The API reference gains the routes under the scopes their neighbours use.

Status log

  • 2026-10-03: opened; recovery built.
  • 2026-10-03: the bell, the notifications page and health alerts by mail built.
  • 2026-10-03: the request log built.
  • 2026-10-04: alert lines and mail are per account settings.
  • 2026-10-07: Checked: the bell and health alerts (#194), the request log (#204, #216) and alert settings (#211) run in accounts, protocol and console-ui. Since a database upgrade on 2026-10-06 (#321) the request log drops its batches; the fix (#472) is merged, not rolled: accounts runs 932c4445, which predates it, and ledger records no commit. Open.

← Back to Platform