Egbe Companions · Guides & deep dives

V6 Token-Cost Contingency Plan — Price Shocks, Outages, Kill Criteria

Anthropic list prices (per MTok, as published in the Claude platform docs, cached 2026-06-04; re-verify quarterly):

5sections10 minread5tables

On this page

Status: Planning gap-fill per V1_V7_PLAN_SET_AUDIT_2026-06-12.md §6.2 ("no model-price-shock or provider-outage contingency for a product whose stated make-or-break bet is token cost"). Date: 2026-06-12. Owners: Moirai kernel owner (tiering levers), model-ops lead (routing, caching, dashboard), V6 product lead (player-visible degradation decisions, kill-criteria escalation), finance partner (weekly cost readings).

Canonical inputs: per-tier token budgets and concurrency caps in V6_DEPENDENCIES.md §19 (:400–437) — Clotho ≤ 50k tokens/agent/active-minute (typical 15–25k with caching) on Opus-class; Lachesis ≤ 10k/agent/reflection (≈ ≤1,200/agent/game-minute amortised) on Sonnet-class; Atropos ≤ 15k/agent/ game-day and Vac ≤ 2k/parse on Haiku-class; Clotho scene cap design 16 / hard 24; Lachesis resident world 150–400; Solo-world hourly cognition cap with "homestead rest". The dependency doc deliberately leaves the dollar figure floating ("the dollar figure floats with provider pricing", dep:404–406); this document is where the dollars get computed and stress-tested.


1. Reference prices and unit-cost model#

Anthropic list prices (per MTok, as published in the Claude platform docs, cached 2026-06-04; re-verify quarterly):

Class (V6 routing, dep:426–428) Model tier Input Output Cache read (≈0.1× in) Cache write 5-min (1.25× in) Batch discount
Opus-class (Clotho) Opus $5.00 $25.00 $0.50 $6.25 n/a (interactive)
Sonnet-class (Lachesis) Sonnet $3.00 $15.00 $0.30 $3.75 −50% on everything
Haiku-class (Atropos, Vac) Haiku $1.00 $5.00 $0.10 $1.25 −50% (Atropos only; Vac is interactive)

1.1 Reference scenario — cost per player-active-hour (all mix figures are planning assumptions adopted 2026-06-12, to be replaced by staging telemetry per the §19 cost gate)#

Clotho. Full-cognition agent-minute at the typical 20k tokens (mid of the doc's 15–25k): split 90% input / 10% output ⇒ 18k in / 2k out; input mix 70% cache-read / 10% cache-write / 20% fresh:

text
reads  12,600 × $0.50/M = $0.0063      writes 1,800 × $6.25/M = $0.0113
fresh   3,600 × $5.00/M = $0.0180      output 2,000 × $25.0/M = $0.0500
                              per full-cognition agent-minute ≈ $0.0856

Player attention is serial: assume dialogue/active-direction occupies 30% of active minutes with on average 1.5 agents in full cognition ⇒ 27 full-cognition agent-minutes per player-hour ⇒ $2.31. Ambient co-present Clotho agents (≤16 scene, plans cache-served, "cognition recomputed only on material change", arch:1156–1158): 12 agents × 42 min × 0.8k tokens, ~90% cache-read input, sparse output ⇒ ≈ $0.58. Clotho ≈ $2.89/h (input-side $1.34, output-side $1.55).

Lachesis. 250 resident agents (mid of 150–400); 40% have a due reflection per cycle (the rest resolve from cached plans); 5 reflections/h (12-game-min cadence, 1 game-min = 1 real-min online); typical reflection 4k tokens (vs the 10k cap), 85/15 in/out ⇒ 2.0M tokens/h. Batched Sonnet with 50% of input as cache reads:

text
in  0.85M × $0.15 + 0.85M × $1.50 = $1.40      out 0.30M × $7.50 = $2.25
                                            Lachesis ≈ $3.65/h

Atropos. 500 distant/offline agents advancing per player at typical 6k/game-day (cap 15k), 1 game-day ≈ 2 real-h online ⇒ 1.5M tokens/h, batched Haiku, 50% input cached ⇒ in $0.35 + out $0.56 ⇒ ≈ $0.91/h.

Vac. 30 voice parses/h × 1.2k typical (cap 2k) ⇒ ≈ $0.05/h. Clio. Chronicle/arc-watch amortisation ⇒ ≈ $0.50/h (0.30 in / 0.20 out).

Tier $/player-active-hour Input-side Output-side
Clotho (Opus-class) 2.89 1.34 1.55
Lachesis (Sonnet-class, batch) 3.65 1.40 2.25
Atropos (Haiku-class, batch) 0.91 0.35 0.56
Vac (Haiku-class) 0.05 0.03 0.02
Clio (mixed, batch) 0.50 0.30 0.20
Reference total $8.00 $3.42 $4.58

Two readings of this table matter:

  1. The headline risk is not the Opus dialogue minutes — it is the 250 quietly reflecting residents. Lachesis is the largest line despite the cheaper model, which is why the existing [P2] distillation note (V6_DEPENDENCIES.md:336) is the single biggest lever in §3.
  2. Budget caps ≠ expected spend. If every tier ran at its cap (16 × 50k × 60 Clotho; 400 × 1,200 × 60 Lachesis; Atropos at 15k), the ceiling is ≈ $205 + $53 + $2 ≈ $260/player-active-hour — two orders of magnitude above reference. The token budgets are correctly framed as the contract (dep:431–434); the dashboard target below is the economic control, measured, not derived from caps.

Dashboard target (planning assumption adopted 2026-06-12): launch at the $8.00 reference, trending to ≤ $5.00 by GA+2 quarters via §3 rungs 1–2. The target is recorded on the model-ops cost dashboard the dependency doc already designates as the live tracker.


2. Model-price-shock scenarios#

The audit asks specifically about input-price moves; cache read/write prices scale with input price (they are multipliers of it), so they shock together. Output prices held constant in S1/S2.

Scenario Input-side Total $/player-active-hour Delta
S0 — today $3.42 $8.00
S1 — input +50% $3.42 × 1.5 = $5.13 $9.71 +21.4%
S2 — input +100% $3.42 × 2.0 = $6.84 $11.42 +42.8%
S3 — full reprice +50% in and out $12.00 +50%
S4 — full reprice +100% in and out $16.00 +100%

Observations: (a) V6 is output-heavy ($4.58 of $8.00) because agents generate dialogue, reflections, and beats — input-only shocks hurt less than intuition suggests, and output-token discipline (shorter reflections, schema-constrained beats) is a real lever, not just caching; (b) batch-tier work (Lachesis+Atropos+Clio ≈ $5.06 of $8.00) reprices with the batch discount — if a provider shock ever removed the 50% batch discount instead, the hit would be +$5.06/h (+63%), worse than S2; the contingency below treats "batch-discount loss" as equivalent to S2-and-a-half and triggers the same ladder.


3. Mitigation ladder — in order, each rung quantified against S0#

Rungs are ordered cheapest-first in player-experience cost; each is independently deployable and gated by the existing eval posture (a routing or cap change ships only with V6/evals/* suites green at their minimumPassRates, per arch:1307–1327).

Rung Lever Mechanics Δ vs S0 Running total
1 Cache hit-rate improvements (current Clotho typical is already "15–25k with caching") Deterministic prompt assembly per shared-prefix rules; persona/policy/Ori-dossier prefixes on 1-h TTL caches keyed per agent; plan-cache reuse on the Ori (arch:1156–1158). Clotho input mix 70→85% reads (full-cog input $0.0356→$0.0223/agent-min); Lachesis batch reads 50→75% (in $1.40→$0.83) −$0.98 $7.02
2 Lachesis/Atropos distillation — the existing P2 note (dep:336) promoted to P1 upon any shock Distilled/fine-tuned Haiku-class for reflection and summary workloads; Lachesis post-rung-1 $3.08 → $1.03; Atropos −30% −$2.32 $4.70
3 Scene-cap tightening 16 → 8 (design target; hard cap 24 → 12) Halves ambient Clotho upkeep (−$0.27) and, more importantly, halves the tail: cap-spend ceiling $205 → $103/h; p95 scenes are where Clotho blowouts live −$0.30 mean $4.40
4 Solo-world hourly-cap reduction + homestead rest default-on Lachesis due-reflection fraction 40% → 27% (−$0.33 post-rung-2); Atropos batch −33% (−$0.21); the cap and rest multiplier already exist (dep:423–424, arch:1150–1155) −$0.54 $3.86

Net: the full ladder reaches ≈ $3.86/h (−52%). Under S2 (+100% input) applied to the post-ladder mix, cost ≈ $5.4/h — still below today's unmitigated $8.00. The ladder absorbs a full +100% input-price shock; under S4 (full +100% reprice) post-ladder cost ≈ $7.7/h, i.e., S4 consumes the entire ladder. Anything beyond S4 forces §5 structural moves.

Deliberately not on the ladder: silently faking cognition (hardcoded "reflections"), shrinking safety/eval coverage, or cutting the Isis gate — forbidden under the quality standard; degradation must be honest (arch: 1167–1169: "the player is told honestly if their world is running degraded").


4. Provider-outage degradation — player-visible contract and recovery#

The BT/HTN fallback is already specced: cheap execution is co-located with the world server and survives a cognition outage (arch:657–660, 704, 728, 759–760); "a cognition outage costs richness, never the world" (arch:422–424). This section defines the contract — what the player sees, tier by tier.

Degradation tiers#

State Trigger Player-visible contract
D0 normal Full fidelity.
D1 — Opus-class unavailable or 529-saturated Provider incident on the Clotho route Clotho reroutes to Sonnet-class. Honest banner: "Your companions are thinking a little more simply right now." Dialogue continues; behavior evals for the Sonnet route are pre-certified in CI (a standing clotho-sonnet-fallback eval profile kept green at all times, so the reroute is push-button, not a scramble).
D2 — all interactive cognition unavailable Full provider outage / network partition World never pauses: agents keep schedules, navigation, gestures, and cached barks via BT/HTN. Free conversation is declined honestly in-fiction and in-UI ("Abeni can't find her words right now — she'll remember what you did"). Vac voice→intent parsing is down, but the structured intent builder still issues objectives (non-voice parity, arch:864–870) executed by BT/HTN. No Ori writes are fabricated: perception-derived events buffer durably (see ori-dr-and-compaction.md §3.5); reflections/summaries defer. Crossroads holds extend; irreversible transitions freeze — no Departed/Transcended/Died resolution while cognition is degraded (the Ereshkigal state machines simply do not advance, arch:983–1024).
D3 — batch tier also unavailable Extended outage As D2, plus Chronicle/Book generation pauses; on the player's next return Clio covers the gap from the buffered events ("while the threads were tangled…") — which is exactly its narrative-reconciliation job (arch:894–897).

Recovery contract#

  • Playability RTO: 0 (the world never stopped).
  • Full-fidelity restore: ≤ 15 min after provider recovery (tier reassignment is recomputed every tick; rehydration is the normal escalation path, arch:1171–1173).
  • Backlog drain: deferred Lachesis/Atropos work drains via batch at a catch-up budget ≤ 1.5× the normal hourly tier budget until clear; drain SLO 6 h. Catch-up never exceeds per-tier token caps — an outage must not become a cost incident.
  • No retroactive fabrication: drained reflections process buffered episodes with their original timestamps; if the player witnessed degraded behavior, Clio writes one reconciliation beat, logged as such.
  • Multi-provider hedging is out of scope at launch (V6 consumes V1 model-ops routing; adding a second provider is a V1-platform decision). This plan's posture: degrade honestly within one provider, and keep the D1 Sonnet profile + distilled-model rung warm so the dependency is on a capable model, not on one price point.

5. Kill criteria#

Measured quantity: blended cognition cost per player-active-hour, weekly reading from the model-ops dashboard (the metric dep:431–434 already tracks), trailing-4-week average, evaluated every Monday by model-ops lead + finance partner; product lead owns escalation.

Trigger (X, Y): trailing cost > $12.00 (1.5× the $8.00 reference) for 6 consecutive weekly readings with ladder rungs 1–2 already deployed.

Then (Z), staged:

  1. Within 7 days: deploy rungs 3–4 (scene cap 8, Solo hourly cap −33%, homestead rest default-on) with the §4 honesty banner and a player comms note. These are player-visible; shipping them is a product decision made now, in this document, so the on-call isn't negotiating it mid-incident.
  2. Within 30 days: re-route Clotho default Opus-class → Sonnet-class behind the full agent-behavior + safety eval gates (all minimumPassRate: 1 suites must stay at 1; value-refusal, persona, crisis, minor-protection are non-negotiable). If the Sonnet route cannot hold the eval bar, this step is rejected and we proceed to 3.
  3. If cost remains > $12.00 for 8 further weeks, or step 2 fails evals: executive go/no-go on V6 live-service posture — halt regional expansion, freeze Commons cognition growth (interest-management radius down, wild- agent foundry seeding paused), and a formal decision memo on continuing, re-scoping (Solo-only product), or sunsetting, with the unit economics in this document attached.

Emergency override: trailing cost > $20.00 (2.5×) for 2 consecutive weeks at any time ⇒ rungs 3–4 plus Clotho hard cap 8 immediately, without waiting for the 6-week window.

Standing review: this plan's prices (§1) re-verified quarterly against the provider price list; the reference model re-derived from staging telemetry at every release via verify:v6 moirai-cost-load-readiness (V6/release/moirai-cost-load-readiness.v6release.json), which remains the binding release gate. All scenario arithmetic above is reproducible from the tables in §1 — anyone disputing a number should be able to re-derive it on one page, which is the point.