Problem
The console showed what the platform counted, but not what a developer wants to know. The usage page had a meter and a strip of bars for the month; webhooks showed two counts for the last day; events were a live tail that forgot everything when the tab closed. The numbers were right, but they did not answer the questions people open these pages with:
- Will this account run out of units before the month ends?
- Which key, person or operation made yesterday's spike?
- Which credentials are dead weight, unused for a month or never used at all?
- Is my webhook endpoint healthy, how fast does my server answer, and why did deliveries fail?
- Which event types does this account actually produce?
The data for most of these was already kept: every call is counted per day, key, person and operation, and every delivery records its status, the answer's status code, how long the answer took and why it failed. Only month totals reached the console. Charts were hand-made, and the five chart colours were shades of one blue, which reads as a ramp, not as different things.
We looked at how the developer platforms people already know answer the same questions: the allowance meter that projects the end of the cycle, daily bars stacked by key or model with a range picker and an export, credentials flagged by when they were last used, webhook endpoints with a success rate, answer times and failures grouped by cause. None of them uses a donut or a chart with two y axes.
Proposal
A chart kit
One kit for every chart in the console, on the chart library the console already shipped: thin columns with rounded tops, hairline grids, a tooltip for every day or hour, a legend whenever there is more than one series, and a table view beside every chart, so no number is reachable only by hovering. The series colours are a categorical palette (blue, orange, aqua, yellow, magenta) checked with a colour-vision validator in light and dark mode. The brand amber marks one thing: today, and where the month is heading. Delivered and failed use status colours, which mean state and nothing else.
Answers from the data the console already had
- The month's forecast. From the third day, the meter shows where the month ends at the pace so far, the day the allowance runs out if it would, and the days left.
- Idle credentials. A key or token never used a week after it was made, or unused for 30 days, is marked; each one shows the units it cost this month.
- Webhook health. Each endpoint's last day is a success rate with a ratio bar.
- The overview leads with three numbers: credentials needing attention, endpoints failing, and the last day's delivery rate.
Three read-only series
| Route | Scope | Answers |
|---|---|---|
GET /v1/accounts/orgs/{org_id}/units/series |
usage:read |
calls, tokens and units per day over at most 92 days, split by who made them (a key or a person), category or operation |
GET /v1/webhooks/stats |
webhooks:read |
delivered and failed per hour (24 hours) or day (7 or 30), the server's answer time at p50 and p95, failures by cause and status, each event type's record; for every endpoint or one |
GET /v1/events/stats |
events:read |
events published per day and type over 7 or 30 days |
A series names its four largest groups and sums the rest as "everything else", so a stacked chart never needs a colour the palette does not have. All three read only what the account may already read, cost what the other reads cost, and carry counts, never content.
On top of them: the usage page charts the last 7, 30 or 90 days by key, category or operation, as units, calls or tokens, with a CSV download; every webhook endpoint, and the webhooks page for all of them, gets delivery health over 24 hours, 7 or 30 days; the events page shows the last 7 or 30 days by type above the live tail.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
Errors and answer times
Every counted call also keeps how it ended and how long it took at the gateway, per key, person, operation and day: the caller's errors (a bad request, not found, refused), which are answered and billed as before, and the platform's failures (unavailable, internal, a timeout), which are kept but never billed. Times go into eleven fixed buckets from 10 milliseconds to above 10 seconds, so p50 and p95 can be read for any group and any range without keeping a row per call. Only the class of an error is kept, never its message.
Everyone signed in sees it for their own accounts, with nothing to switch on: the usage page shows the error rate, the platform's failures, the caller's errors and the slowest day, and charts errors and answer time by key, category or operation; every operation in the month's table shows its errors and p95; every key and token shows its error rate; the overview shows the month's errors. The platform's admin sees the whole platform: errors and answer times per day across every account, the operations that fail or are slow, and the accounts with the most errors.
Alternatives considered
- Read the monitoring system. The platform's own metrics have answer times and error rates, but not per account, and putting an account on every metric would multiply what the monitoring stores. The accounts and events services already keep what an account needs, by account.
- Ship every raw row to the browser and chart there. Simple, but it sends a month of rows to draw thirty bars, and an API caller would have to do the same aggregation. Aggregating where the data lives answers both.
- Name every group. Twelve keys in twelve colours cannot be told apart, and a generated thirteenth colour is indistinguishable from one of the twelve for many readers. Four named groups and the rest summed keeps the chart honest; the table and the export keep every number.
- A row per call, for exact percentiles. Exact, but it keeps a row for every call of every account. Fixed buckets on the existing daily counts cost eleven numbers per key, operation and day, and the percentile is accurate to its bucket.
Decision
Decided 2026-10-07. Built as described above on 2026-10-03, in three steps: the chart kit with the answers the console already had the data for, the three series and the charts on them, then errors and answer times per key and operation, with the admin's view of the whole platform.
Publication
The console's overview, usage, keys, webhooks and events pages change as described. The three routes appear in the public API reference and in the playground, admitted by the scopes the related reads already use. Nothing changes for a key that does not call them.
Status log
- 2026-10-03: opened; the chart kit is live, the series in review.
- 2026-10-03: the series are live; errors and answer times per key built.
- 2026-10-07: Decided: built and live. The series (#131) and errors and answer times per key (#149) run in accounts, protocol and console-ui since 2026-10-03.