Problem
Every schema migration, storage move and upgrade on the platform so far has been done live, or by stopping things by hand. The incidents of 2026-10-04 and 2026-10-05 show what that costs:
- No clean way to say "closed". When a service was down for work, callers got whatever the failure produced: a generic error page, a connection error, a 500. The platform has copy and artwork for "down for maintenance", in both languages. Nothing ever shows it.
- No way to stop writes and keep reads. A migration that needs a quiet table takes the whole service down, though customers could have kept reading their data.
- No way to pause background work. Monitors, webhook delivery, sweepers, metering and sync keep writing to the shared database while someone works on it.
- Every switch needs a roll. The edge's configuration is static, so closing anything means changing configuration and restarting pods. That is slow, and it is itself a change made during an incident.
- Nobody is told. No customer learns of planned work in advance. Our own alerts fire for work we planned.
- A roll can run ahead of its migration. A service started on a schema it does not expect answers "internal error" instead of refusing to start.
Proposal
One source of truth, outside the database
A maintenance window is a small record:
- the scope: a service, such as
billing, or a host, such as the console; - the mode:
closed,read_onlyorworkers_paused; - a reason, a start and an end;
- an allow-list of accounts that keep full access.
The windows live in one configuration object in the cluster, not in the database. The work a window protects may be on the database, and the switch must keep working while it is.
One writer changes it: the platform's command line, used through mise run ops:*. Every
change is audited (who, scope, mode, reason). In the first version the console shows the
windows but cannot change them, so no pod in the cluster holds the power to close the
platform. A small service with exactly that one permission comes later, as its own step.
Services enforce the details
Each service reads the windows from a mounted file, through a new transport-free type in the shared library beside the fault handle:
- It reads the file every 5 seconds and compares a content hash.
- A missing file means open.
- An unreadable file keeps the last good state and raises a metric and an alert.
- Each service applies start and end itself from the timestamps, so a window ends on time even if nothing rewrites the file.
Every service already passes each call through one admission step before doing any work. That step learns about windows:
- Closed: the call is refused with
UNAVAILABLE, a retry delay (RetryInfo) and a reason (ErrorInfo,MAINTENANCE). The gateway already turns that into HTTP 503 withRetry-Afterand a machine-readable reason on every surface. - Read-only: a write is refused the same way, with the reason
MAINTENANCE_READ_ONLY. A read goes through. It is deliberately a 503 and not a 400, so clients and SDKs wait and retry instead of treating it as their own mistake. - Bypass: a platform admin, or an account on the window's allow-list, passes. Every bypassed call is audited.
Which calls write. Each service gets a table generated from its API annotations:
GET is a read, everything else is a write. A short override list covers the few reads
sent as POST. A call with no annotation counts as a write, so the table fails closed.
The table is keyed by the full method name, which covers the same calls arriving over the
multiplexed socket, MQTT and MCP. Calls between our own services, such as usage metering
and erasure, get an explicit decision per window mode. They are never exempt by accident.
Background work. Each background loop asks before every tick whether it may run. It finishes its current unit of work, never half of it, and reports paused or running as a metric.
Knowing a window has taken effect. Configuration reaches every pod within about a minute, not at once. Each pod reports the generation of the windows it has applied. Anything that depends on a window, a migration above all, waits until every pod reports the new generation.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
The edge closes hosts
For a whole host (the console, the documentation, the API), the edge can close it before any request reaches a service:
- A switch per host takes effect without a restart.
- Platform admins pass; everyone else gets the maintenance page in their language, or the
JSON error envelope with
Retry-After. - The account allow-list is not enforced at the edge, which cannot read it. A window that must let a named customer through closes services, not the host.
- A rule alerts when an edge switch is still set after its window's end.
Before this RFC calls the edge switch working, it is proven on the local cluster: both edge replicas pick up a change, how long that takes is timed, and the edge still starts where the switch file is absent.
Telling people
- Ahead of a window: a notice endpoint returns the scheduled and active windows. The
console, the site and the documentation show a banner from it. Events
platform.maintenance.scheduled,.startedand.endedgo to the accounts that subscribe. They are sent before a window that affects the database begins, because the event store lives in that database. - During a window: browsers see the "down for maintenance" page with the planned end.
API clients get
Retry-After. Our SDKs already honour it up to 60 seconds. A longer window ends a call quickly with the reasonmaintenance, and the documentation says so. - Notice period. Planned work is announced at least 5 business days ahead; security fixes 24 hours ahead. This goes into our incident and continuity policies, because contracts with financial entities under DORA ask for it.
Migrations inside a window
mise run db:migrate --window billing,accounts takes an explicit list of services, never
guessed from the migration's text. It runs four steps:
- Open a read-only window for those services and pause their workers.
- Wait until every pod reports the new generation.
- Migrate, with a lock timeout so a migration waiting on a lock fails instead of freezing reads.
- End the window.
Services also check at start that the database schema is the one they were built for, and refuse to start otherwise. That makes "migrate before roll" a rule the roll command enforces, not a habit.
A diagram is drawn here in the RFCs product; this page does not show diagrams yet.
The rest of the switchboard
Each item below is its own change under this RFC, in this order:
- The console can change windows, through a small service whose only permission is to update that one configuration object.
- Silenced alerts. A window silences the alerts for its scope's labels, and only those. Alerts about the host and its disks are never silenced.
- Change freeze. A window mode in which the roll command refuses to roll unless it is an emergency change, which is logged and named as such (our change management policy, emergency changes).
- Kill switches. Named switches for risky features (anonymous runs, sign-up, generation), in the same file, with no roll.
- Runbooks. One runbook per alert rule, linked from the rule.
- On-call. A rotation and an escalation rule that pages the on-call person's phone through the PagerDuty or incident.io connections.
- Service levels. Availability and latency objectives per surface from the request metrics we already record, an error budget, and maintenance windows excluded from it.
- A deploy log. Every roll writes an event: who, what, which commit.
- Clients. The console and the mobile app honour
Retry-Afterand say "back at HH:MM". The agent receives a maintenance frame and pauses its jobs instead of reconnecting in a loop. - Health and load balancing. Stricter host ejection only for services with more than one replica, after a drill. Never globally: today it would turn one backend's outage into a 503 for the whole API.
The public status page stays with incidents (RFC 0058, phase 3) and reads the same windows.
Plan
- This RFC.
- The shared windows type, the generated read/write tables, admission in every service, worker pause points, the generation metric, and the command line.
- The edge switch, proven on the local cluster first, and the maintenance page and envelope.
- The notice endpoint, the banner, and the events.
- The migration helper and the start-time schema check.
- A drill on a low-risk service: close it for 5 minutes and check from outside the page, the envelope, the banner, the events, every pod's generation, and that nothing else is touched.
- The switchboard items above, one change each.
Security and compliance
- Least privilege. The first version gives no pod the right to change windows. The later console service may change exactly one configuration object and nothing else.
- Audit. Every window change and every bypassed call is logged with who, what and outcome, never content.
- Bounded. A window has an end. Services enforce it themselves, and an alert covers an edge switch left behind.
- Controls. This strengthens emergency change management, availability alerting and continuity in our compliance programme, and adds a planned-maintenance notice period to the incident and continuity policies.
Alternatives considered
- Windows in the database. Rejected. The switch would fail during exactly the work it exists for.
- An external authorisation call from the edge to a service per request. Rejected. The edge would depend on a service that may itself be under maintenance.
- Restarting with a different configuration. That is what happens today. It is slow, it is a change in itself, and it cannot schedule.
- A feature-flag product. Rejected for now. It is another vendor and data processor for a need that one file and one type meet.
Status log
- 2026-10-06: RFC opened, nothing built. Research on the edge, the gateway, the services, the workers and the clients is the basis of the Problem section, and an internal review of the plan is folded in.
- 2026-10-07: Checked: nothing is built; #394 is the text. Open.