Owner: workbench-platform on-call (rota:workbench-platform). Last reviewed:
2026-08-14. Version: 1.
The primary model provider stops answering. The standby rests on the same gateway, so this is a test of the gateway boundary as much as of the provider: a plan that fails over to the standby has to say why the standby is reachable when the primary is not.
Related: provider-failover.md is the estate-wide procedure for a degraded AI, search, payment, messaging or media provider. This runbook is narrower and answers to a registered scenario: it names the dependency the drill injects into, the capabilities that stop, and the observation that would refute a pass.
1. Detection and declaration#
Detection signal. dep:model-provider-primary carries circuit-breaker
among its registered mitigations, which is the detector and the mitigation at
once — and S10.5's finding applies here: a breaker's error rate falls to near
zero on a dependency that is one hundred percent down, because the calls stop
being made. Read slo:provider-backed-success-rate and the caller-failure rate,
not the dependency's own error rate.
The three capabilities that stop are "run a model-backed gate", "run an assisted action" and "generate a suggestion". None of them is a read of stored state, so nothing looks broken to an actor who is only browsing.
- Detection authority: workbench-platform on-call, contact
rota:workbench-platform. - Declaration authority: production operations duty officer, contact
rota:prod-ops. - Gap, stated.
docs/operations/on-call.mddescribes a weekly primary/secondary rotation and names no person and no rotation config; none of the register'srota:identifiers appears in it or indocs/operations/alert-routing.md. The name above is registered; the route from it to whoever is carrying the pager is not.
2. Communications#
| Audience | Within | Content |
|---|---|---|
| Actors | 15 min | that assisted actions and suggestions are unavailable, and that stored work is untouched |
| Operators | 5 min | that this runbook is running, and by whom |
| Provider | 60 min | the failing request ids and the observed onset, where the provider takes them |
All travel over dep:notification-channel, which this scenario leaves standing.
3. Dependency assumptions#
dep:provider-gatewayis healthy. This is the assumption the scenario is designed to test. Bothdep:model-provider-primaryanddep:model-provider-standbyrest on the gateway in the S10.5 substrate, so a failover to the standby is only a recovery if the gateway is not the thing that failed. If it is, the standby is already gone and this runbook does not apply.dep:primary-storeis healthy: the record of what was requested and what came back is stored there, and it is what makes a duplicate dispatch detectable.
4. Recovery infrastructure#
In place. The failover is a routing decision at the gateway. Nothing is provisioned and nothing moves.
5. Secrets#
- The standby provider's credentials, which must be present before the incident. A failover to a provider whose key was never configured fails closed at the worker, which is correct behaviour and a longer outage.
Recoverable out of band: the configuration-and-secrets state class is in the
off-box backup (infra/hetzner/backup/offbox-backup.sh).
6. Restore order#
One dependency: dep:model-provider-primary comes back when the provider does,
or when traffic is routed back to it from the standby. No reference edge touches
it, so there is no ordering constraint against the graph — the constraint is
§10's: do not route back while duplicate dispatch is unresolved.
7. Validation#
scripts/isis/run_provider_outage_drill.sh is the real instrument and it now
does what its name says. Two properties matter, both repaired in the same pass
that wrote this register:
SUBMIT_SYNTHETIC_JOBsubmits a real job (SYNTHETIC_JOB_TYPE, defaulttext-to-image) and polls it to a terminal state, so the check is a completion rather than a health endpoint;- the read-only gate honours an explicit
false. It previously readjq -r '.isReadOnly // true', and jq's alternative operator treats a literalfalseas absent, so the expression returnedtruefor every input and the drill could never inject its fault.
Run it with DRY_RUN=false. In dry run the mutations are not applied and every
verification records skip — and a skipped check is not a pass.
8. Fencing — NOT APPLICABLE, and the reason is specific#
The standby holds no authoritative state, so there is no old primary whose writes must be rejected: two providers answering is not split-brain. The real risk of this failover is duplicate dispatch — an in-flight request the primary already accepted, reissued to the standby — and the control for that is an idempotency key on the request, not a fencing token on the provider. Naming this section "fencing" and filling it would be a false signal; the risk is recorded in §10 where it belongs.
9. Failover#
- Confirm
dep:provider-gatewayis healthy. If it is not, stop: §3. - Route model traffic to
dep:model-provider-standby. - Record the instant. Recovery time runs from the fault, not from this step.
run_provider_outage_drill.sh performs the flag mutations that stand in for
this and reverses them from an EXIT/INT/TERM trap, so a failure between
injection and restoration no longer exits with the fault applied.
10. Reconciliation#
The question this outage leaves is whether a request was dispatched twice: once to the primary before it stopped answering, once to the standby after.
scripts/operations/v1-restore-reconcile.mjs compares restored provider effects
against what the outside world observed, and is the control that answers it.
providerEffects is one of the seven collections it validates, and it is keyed
by idempotencyKey — which is the field that decides whether the second
dispatch was the same request or a new one. It is also the reason
safeToResumeAutonomy exists.
Give it expected counts. An empty collection reconciles clean.
11. Failback#
Route traffic back to the primary once it has answered a synthetic job successfully, and not before: a provider that has recovered its health endpoint before its capacity fails the first real request. Restore the flags the drill mutated — the drill's trap does this, but a manual failover has no trap and the flags are the estate's memory of the incident.
12. What this runbook does not cover#
- A gateway outage. Both providers rest on it, so that is a different scenario and this procedure makes it worse.
- Provider credential rotation —
scripts/isis/rotate_runpod_credentials.shcovers that path. - Cost or quota exhaustion, which presents as a provider outage and is not one.
13. Machine-checkable statement#
{
"id": "runbook:v1-model-provider-outage",
"scenarioId": "gd:model-provider-outage",
"version": "1",
"detectionAuthority": {
"role": "service on-call",
"name": "workbench-platform on-call",
"reachableOver": ["dep:notification-channel"]
},
"declarationAuthority": {
"role": "operations duty officer",
"name": "production operations duty officer",
"reachableOver": ["dep:notification-channel"]
},
"communications": [
{
"audience": "actors",
"channelDependencyId": "dep:notification-channel",
"withinMs": 900000
},
{
"audience": "operators",
"channelDependencyId": "dep:notification-channel",
"withinMs": 300000
},
{
"audience": "provider",
"channelDependencyId": "dep:notification-channel",
"withinMs": 3600000
}
],
"assumes": ["dep:provider-gateway", "dep:primary-store"],
"recoveryInfrastructureDependencyId": null,
"secrets": [
{
"id": "standby provider credentials",
"heldInDependencyId": "dep:model-provider-standby",
"recoverableOutOfBand": true
}
],
"restoreOrder": ["dep:model-provider-primary"],
"fencing": null,
"promotesAStandby": false,
"validation": [
"run_provider_outage_drill.sh with DRY_RUN=false submits a synthetic job and polls it to a terminal state on the standby",
"slo:provider-backed-success-rate recovers, read from the caller side rather than from the dependency error rate"
],
"failback": "route back to the primary only after it completes a synthetic job, and restore the feature flags the drill mutated"
}