Metis · Reference & analysis

V9 Product Review — Metis: A Curious Ape's Guide to Reality

V9 productizes the Metis education domain into a consumer flagship whose front door is not a course catalog but a question: *what do you wonder about?* A free-text wonder resolves against a unified concept graph (the

11sections21 minread1table

On this page

Status: independent product review — revised and expanded in a second meticulous pass on 2026-07-07 (§9, Additional considerations) and a third pass on 2026-07-07 covering monorepo integration, shared functionality, and TODOS/ roadmap alignment (§10) Reviewer: Claude (product-perspective deep review) Date: 2026-07-07 (first, second, and third passes) Scope: V9/V9_features.md + all V9/features/* pages, V9_PRODUCT_ANALYSIS.md, V9_TODOS.md, V9_GAP_ANALYSIS.md, V9_SOTA_RESEARCH.md, README.md, architecture/overview.md, plus the implementing packages under libs/v9/*. Method: full deep-read, then critique against the shipped competitive set (NotebookLM incl. Video Overviews, ChatGPT Study Mode, Gemini Guided Learning, Claude Learning Mode, Khanmigo, Duolingo/Max, Brilliant.org, PhET, 3Blue1Brown/Manim, Kurzgesagt, Bret Victor's explorables, Synthesia/ HeyGen avatars, Universe Sandbox, Outer Wilds) and industry SOTA in AI education. Market facts reflect knowledge through early 2026.


1. Executive summary#

V9 productizes the Metis education domain into a consumer flagship whose front door is not a course catalog but a question: what do you wonder about? A free-text wonder resolves against a unified concept graph (the Atlas of Reality), a solve-first pipeline (Prometheus) forges a lesson whose every fact is pinned to a vetted source and every STEM number recomputed from a real kernel (Nyx ephemeris, Kalika CAS/cosmology), a truth gate (Aletheia) blocks anything ungrounded, a voiced Socratic mentor (Chiron) teaches it, a computed explorable (Hephaestus) lets the learner do the idea (drag the universe's age and watch Olbers' paradox resolve), and a mastery engine (real FSRS-v4 spaced retrieval) makes it stick. Seven release gates define "done"; engagement-harvesting is banned in code (BannedOptimizationTargetError); the north star is durable understanding per learner per week.

V9 is the inverse of its siblings. V2–V5 are magnificent plans on absent substrates; V9 is real substrates (production-grade Metis backend with ~174 operations, real psychometrics, real science kernels, real learning science) missing exactly one thing: the consumer app. The front door, lesson player, and star-map exist as tested view-models with no rendered screens; the richest experience layer (Theia: emotional arcs, films, threads) is built and imported by nothing. GAP_ANALYSIS says it plainly: what's missing is "almost entirely connective tissue and consumer experience."

The thesis is the right one for 2026. ChatGPT Study Mode, Gemini Guided Learning, and Khanmigo all teach from a model's memory and inherit its fabrications; NotebookLM grounds but doesn't compute, assess, or remember; Brilliant is interactive but hand-authored and slow to expand; free YouTube is beautiful but passive and unassessed. "Grounded like NotebookLM, computed-interactive like PhET, mastery-tracked like Khan, awe-driven like the best science film" is a real hole in the market — and V9's zero-fabrication architecture (claims that fail to parse without source pins; numbers that cannot ship unless recomputed) is the only credible answer I've seen to the AI-tutor trust problem.

Three risks dominate: (1) instant-answer gravity — the front door competes with the ChatGPT habit; if a lesson takes minutes to mint while chat answers in seconds, wonder loses to convenience; (2) thin computed coverage under expansive language — two explorable builders and six kernel refs ship beneath "manipulate any idea" prose; the kernel/explorable expansion is the content roadmap and it is unpriced; (3) the soulful differentiator is structural, not populated — the bridges between myth, meditation, and mathematics that justify "a curious ape's guide" are edge types awaiting editorial content.

Verdict: build the screen, make the first answer instant, expand the kernels on a public cadence, and market the trust architecture as loudly as the wonder. This is the most shippable product in the V-series and the one whose moat — verified truth with computed interactivity — nobody in the market can quickly copy.


2. What V9 is (product identity)#

  • Five primitives: Wonder → Lesson → Thread → Atlas, on a concept-graph spine; seven wonder axes (cosmos, laws, mind, meaning, deep-time, living-world, made-world) powered by the monorepo's science constellation (Nyx, Kalika, Nisaba, Demeter, Seshat, Saraswati, Sophia, Mnemosyne).
  • Five pillars: True (the moat), Beautiful, Interactive, Memorable, Soulful — with the fifth explicitly anti-nihilist ("wonder, made navigable") and the fourth ("the quiet pillar") separating it from edutainment.
  • Chiron: one disclosed synthetic persona at P1 (warm generalist, Hathor personality, Psyche voice), four academic-integrity modes including do-not-complete-for-me (refuses to do graded work), richest-first delivery fallback with disclosed downgrades; grounded historical reconstructions at P2 (Nisaba-bounded, consent-gated, labeled).
  • Audience: the 16–25 late-night wonderer first; returning adults second; institutions explicitly not at launch ("procurement would smother the consumer magic").
  • Monetization: generous free tier (5 mints/period) as the awe funnel; "Netflix for your mind" subscription for embodied Chiron, unlimited lessons, full workshop, personal Atlas; deny-by-default safety gating that a subscription can never buy past.
  • Honest completion: 46 [x] / 11 [~] / 0 open; 173 tests green; adversarial stub-scan clean; the [~] items are precisely the consumer screens, provider-gated embodiment, and P3 motions.

3. What is genuinely strong#

  1. The zero-fabrication architecture is the category answer. Facts that fail to parse without a Sophia pin; STEM values bound to callable kernel refs ("you cannot ship a number you cannot reproduce"); a gate that recomputes cosmology through a real Friedmann solver before a lesson can exist; fabrication as a sev-1 with a zero target. Every AI tutor on the market asks for trust; V9 is built to prove it. This is the marketing, the moat, and the regulatory hedge in one.
  2. Computed explorables over generated video. The Nyx sky slider and Kalika orbit integrator are Bret-Victor-grade interactivity generated on demand — with successState.reachable verified by actual numerical sweeps before the widget ships. NotebookLM's Video Overviews are passive and occasionally wrong; V9's films render only from verified derivations. "She did the paradox. She didn't read it" is the right product sentence.
  3. The anti-engagement stance, enforced in code. Banning time-on-app as an optimization target via a thrown error, and measuring durable-understanding-after-delay as the north star, is a position no incumbent (least of all streak-driven Duolingo) can adopt without self-harm. In the 2025–26 climate of AI-edtech backlash, this is a brand asset — publish it.
  4. Mastery as consumer delight, not LMS chore: FSRS-driven forgetting frontiers feeding a personal star-map where mastery is brightness and the frontier is "the inviting next star" — the best visual metaphor for spaced repetition I've seen; it turns review into tending a sky.
  5. De-risked foundations. Unlike every sibling, the hard engines exist and are production-grade (the Metis backend's IRT/DIF psychometrics, live-voice tutoring machinery, LMS interop held in reserve for P3; Mnemosyne's real FSRS; the kernels). The remaining build is "an ordinary product build" — GAP_ANALYSIS's own words, and correct.
  6. Honest self-caveating: Bloom's 2σ cited as aspiration with the real 0.3–0.4σ literature attached; awe-pedagogy marked suggestive; SOTA entries flagged [POST-CUTOFF]/[UNVERIFIED]; the vision docs actively distinguish themselves from the more modest shipped code.

4. Product gaps#

4.1 The instant-answer problem (front-door latency)#

The wonder front door competes head-on with the reflex of asking ChatGPT. No lesson-mint latency budget appears anywhere in the corpus. If forging (ground → plan → write → gate) takes minutes, the magic moment dies in a spinner. Three mitigations, all compatible with the architecture:

  • Progressive assembly: answer the wonder immediately with the grounded skeleton (the resolved concept, the one-paragraph grounded answer with pins — cheap, mostly retrieval), then stream the lesson around it (Socratic beats, explorable, film) as stages complete. The learner should be reading a true first answer in ≤5 seconds.
  • Cache-first serving: determinism + profile-class caching already means popular lessons are minted once; precompute the top ~10k wonders per axis so the median first-time query is instant.
  • Publish the latency SLO alongside the truth SLO; both are trust.

4.2 Free-tier design fights the growth loop#

Five mints per period is a fine cap on novel generation cost — but the docs' own economics (mint-once, serve-many; cached lessons cost ~nothing) argue that shared and cached lessons should never count against quota. A shared metis.oshun/share/... link that hits a paywall kills the "the free tier is the marketing" thesis at its strongest moment. Recommended free tier: unlimited playback of any cached/shared lesson + the daily communal wonder (below) + 5 novel mints; subscription buys unlimited novel mints, embodied Chiron, the full workshop, and the Atlas.

4.3 No daily ritual (the retention gap)#

Mastery scheduling exists; a consumer habit surface does not. Two cheap, on-thesis constructs:

  • The Daily Wonder: one communal question per day, same for everyone (deterministic minting makes this free), with the global "how many apes wondered this today" reveal and a shareable answer card. The Wordle slot for curiosity — and the top of the funnel.
  • Fading stars: the star-map's brightness already models retrievability; let stars visibly dim as the forgetting frontier approaches, and make "re-kindle three stars" the two-minute daily loop. Review becomes tending, not homework — Duolingo's streak with none of its coercion (and honest to the banned-metrics posture).

4.4 Computed coverage vs. promised breadth#

Two explorable builders (sky, orbit), six kernel refs, one wired search substrate — under "any astronomy concept becomes a slider" and Schrödinger language. The kernel/explorable registry is the content pipeline of this product, the way cases are V8's and asanas are V3's. It needs: a public expansion cadence ("a new computed kernel every week," announced like content drops), the Demeter binder (admitted missing), the generative-widget runner (verifier real, runtime absent), and the TS↔Python Manim adapter (both halves exist; the bridge doesn't — a small seam blocking the single most shareable artifact, the computed-correct explainer film).

4.5 The soulful pillar needs an editorial program#

Pillar 5 is the differentiator against Brilliant/Khan (who own rigor) and YouTube (who own awe): the braid — Hathor's Orion myth beside the ephemeris, the meditation session beside the neuroscience of attention. Today it is an edge type (bridges) with priority in thread-walking and no populated content. This is not a code gap; it is a curation program (commission the first 100 bridge pairings across the seven axes) and it should be resourced like content, not backlog. Without it, V9 is an excellent physics tutor; with it, it is the only product that teaches the cosmos and what the cosmos means in one thread.

4.6 Industry SOTA checklist#

Capability (2026 bar) Bar-setter V9
Grounded answers w/ citations NotebookLM ✅ + structural (unpinned facts fail to parse)
Computed-correct STEM values (nobody at consumer scale) ✅✅ category-defining (kernel recomputation gate)
Generated interactive explorables PhET (hand-authored) ✅✅ novel — 2 builders shipped; breadth gap (§4.4)
Computed-correct explainer films NotebookLM Video Overviews (ungated) ✅ designed better; TS↔Python bridge missing
Socratic integrity modes Khanmigo ✅ + do-not-complete-for-me enforced in code
Spaced retrieval w/ real FSRS Anki (unfriendly) / Duolingo (gamified) ✅✅ star-map metaphor is best-in-class (ship it)
Anti-engagement metrics posture (nobody) ✅✅ enforced in code — publicize
Embodied AI tutor Synthesia-class avatars ⚠️ P2, provider-gated seams
Learn-by-playing bridges (open niche) ❌ contract-only; blocked on Bellona cook path
Daily habit surface Duolingo / Wordle ❌ gap (§4.3)
Consumer app shipped table stakes ❌ view-models only — the P0 build
Efficacy evidence Harvard-class AI-tutor RCTs emerging ⚠️ instrumented, unrun (§5.6)
Offline/equity Khan Lite ⚠️ P2; Nous local inference designed

5. Ideas that would make the product better#

  1. The Daily Wonder + fading stars (§4.3) — the habit spine.
  2. Progressive answer-first assembly with a ≤5s first-truth SLO (§4.1).
  3. Shared lessons always open free (§4.2) — let the social object do its job; the marginal cost is a cache hit.
  4. A public correction ledger. When a shipped lesson is found wrong, publish the correction like a newspaper — visible, dated, linked from the lesson. Zero-fabrication is the goal; visible accountability is the brand. No AI product does this; the provenance ledger makes it nearly free.
  5. Kernel-drop cadence as live-ops (§4.4): "This week: wave interference — slit width is now a slider." Content marketing that is literally the product roadmap.
  6. Run the efficacy study at beta, not P3. A pre-registered, controlled durable-understanding result — even a modest one — plus the zero-fabrication record is an unanswerable marketing position against Study Mode/Khanmigo, and the north-star instrumentation already exists.
  7. The braid editorial program (§4.5): 100 commissioned bridge pairings; poach science-YouTube writers — the people who already speak awe fluently — as the first Agora creators with guaranteed rates.
  8. Bundle position: V9 is the natural premium anchor for the Oshun membership (the V1 review's "Oshun+"); cross-product grants (meditation streak → neuroscience lesson; V8 case solved → logic lesson) already encode the bundle logic — surface them as delightful surprises ("your meditation practice just unlocked a lesson").
  9. "Explain it at my level" slider on every lesson (the learner-model scoping already computes the prerequisite bridge; expose it as a visible control — kid/curious/undergrad/expert) — the single most requested feature of every explainer product, and V9 can do it without changing the facts, only the bridge.
  10. Classroom mode later, wedge now: keep the institutional "no" at launch (correct call), but ship a lightweight share-to-teacher artifact (lesson + sources + mastery evidence) so teachers become the free distribution channel students bring them, not a procurement problem.

6. Criticisms and tweaks#

  1. Ship the screens. Four [~] items — front door, lesson player, mastery view, closed beta — are the entire product as far as a human is concerned. Everything else in this review is secondary to a rendered, latency-budgeted web app wrapping the already-tested view-models. Risk is marked "low — ordinary product build"; treat it with extraordinary urgency anyway.
  2. Unify the seven gates. Two conflicting G1–G7 numberings (contract framing vs release-gate suite) — the docs themselves call it "the single easiest mistake to make here," and V8 has the same disease. One canonical numbering, one alias table, one CI check, across both products.
  3. Resolve the small honesty drifts: 12 vs 13 packages, ten-vs-nine- vs-eight pipeline stages, Agora-as-subsystem-but-not-package, two ununified "surprise me" implementations, the built-but-unimported Theia. None is damning; together they blur a corpus whose credibility is its precision.
  4. Wire Theia or stop counting it. The most differentiating layer (emotional arc, films, threads, Agora) is imported by nothing. The experience-layer build (§6.1) should consume Theia's thread engine rather than the player's simpler precomputed nextWonders, or the richer engine will rot.
  5. G5 "delight" is beat-counting. The structural proxy is honest but Pillar 2 deserves a real judge: an LLM panel calibrated on human ratings from the beta (the champion-challenger machinery exists), and the emotional-arc labeler should be decoupled from Prometheus's exact beat phrasing before either changes.
  6. Historical personas (P2) need a press strategy, not just consent architecture. "Galileo teaches you optics, reconstructed from his letters" is either the best marketing beat this product will ever have or a "digital necromancy" headline; the labeling/Nisaba-bounding is right — pair it with an opt-in public methodology page and pick figures with estates/scholars engaged.
  7. C2PA is hashing, not signing (c2paSigned: false by default) — inherit V1's G0 SDK-wiring gate; a trust-branded product should not ship provenance that stops one step short of cryptographic.
  8. Kids (12–15 supervised) as tertiary audience is correctly modest; the deny-by-default gating is right. Do not let growth pressure invert it — the companion-app regulatory climate applies to tutors with faces too.

7. Risks#

  1. Instant-answer gravity (§4.1) — losing to "good enough, right now" chat is the demand-side risk that dwarfs everything technical.
  2. Free-YouTube substitution — Kurzgesagt is free and gorgeous; V9's answer must be doing and remembering (explorables + mastery), which makes §4.4's breadth and §4.3's ritual load-bearing, not optional.
  3. Kernel-expansion economics — each new computed domain is real engineering; without the cadence commitment (§5.5) coverage stalls at astronomy-plus-orbits and the "guide to reality" promise reads hollow.
  4. Platform giants copying the surface — Google can bolt citations and a slider onto Gemini faster than V9 can build distribution; the durable moats are the kernel registry, the mastery graph, the monorepo's cross-product flywheel, and the trust record. Invest in what compounds (correction ledger, efficacy evidence, Atlas depth).
  5. Efficacy shortfall — if the study lands near zero, the north star becomes a liability; run it early (§5.6), size expectations honestly (the docs already do), and let mastery-retention telemetry guide pedagogy before the public claim.
  6. Sibling-dependency creep — embodiment (Psyche/Isis), films (Manim bridge), games (Bellona) are all external seams; the P1 posture (web-first, text+explorable, one persona) is correctly independent — protect that independence in planning.

8. Prioritized recommendations#

P0 — ship the product

  1. Build and beta the consumer web app (front door, player, star-map) with a ≤5s first-truth latency SLO and progressive assembly (§6.1, §4.1).
  2. Free-tier redesign: shared/cached lessons free; Daily Wonder; fading stars (§4.2, §4.3).
  3. Manim TS↔Python bridge — unlock the shareable film (§4.4).
  4. Unify gate numbering + resolve doc drifts (§6.2, §6.3).
  5. Kernel-drop cadence committed and public (§5.5).

P1 — the trust wave 6. Public correction ledger (§5.4); C2PA signing (§6.7). 7. Pre-registered efficacy study at beta (§5.6). 8. Braid editorial program + first Agora creator cohort (§4.5, §5.7). 9. "Explain it at my level" control (§5.9). 10. Real G5 judge calibrated on beta ratings (§6.5).

P2 — compounding 11. Wire Theia's thread engine into the player (§6.4). 12. Embodied Chiron pilot; historical-persona methodology page (§6.6). 13. Cross-product grant surprises surfaced in-product (§5.8). 14. Share-to-teacher artifact (§5.10). 15. Mobile/offline motion for the equity claim.


9. Additional considerations (second pass)#

9.1 Source licensing and the pin schema#

"Show your sources" panels that quote pinned excerpts need a licensing posture, not just a citation format: open-licensed corpora first (Wikipedia, OpenStax, arXiv, public-domain classics, government data), publisher licensing later, and license metadata in the Sophia pin schema now so per-source display rules (link-only vs excerpt vs full quote) are enforceable structurally — the same move that made grounding un-skippable should make licensing un-skippable.

9.2 The poisoned-source threat model#

Grounding relocates the attack surface from model hallucination to source manipulation: a compromised or adversarial source produces a wrong-but-cited lesson — the subtle failure mode that damages a trust-branded product most. Inherit V6's instruction/data-separation threat model (source text is data, never instructions), and add provenance tiers to Aletheia (peer-reviewed > reference > web) with per-tier trust weighting and G1 thresholds. The kernel-recomputation gate already immunizes the numbers; the claims need the tiered version of the same immune system.

9.3 Duolingo is coming sideways#

Duolingo's expansions (math, music, chess) show the gamified giant with 100M+ MAU distribution moving into adjacent learning verticals, and its AI-first content pipeline scales faster than hand authoring. V9 cannot out-distribute it; it can out-trust it (verified truth, computed interactivity, anti-engagement posture — the exact axes Duolingo's model monetizes against). This sharpens §4.3's urgency: the habit surface must exist before the giant's science course does, because "Duolingo but true" is a positioning V9 must claim first.

9.4 Wonder data as a public asset#

Aggregate, k-anonymized "what humanity wonders" trends are marketing gold (an annual Wonder Report — the Spotify-Wrapped of curiosity, individually shareable and collectively newsworthy) and the internal curriculum compass (the most-wondered unanswerable queries = the next kernels to build, §4.4's cadence prioritized by demand). Privacy cost is near zero if designed in from the start; retrofitted, it becomes a consent problem.

9.5 Spoken wonders#

Wonder capture is bedtime-and-walking behavior; typing is friction. The Metis backend's live-voice machinery (barge-in, VAD) should power a speak-your-wonder front door from P1 on mobile web — and it doubles as the accessibility path and the kids-tier input (12–15 supervised users type less than they talk).

9.6 Offline as an equity strategy, not a feature#

The kernels are local math; an offline "field edition" (cached Atlas slice + kernels + downloaded lessons, no live LLM) serves low-connectivity learners at near-zero marginal cost and substantiates the equity language the product analysis leans on. Khan Lite is the precedent; V9's version is better offline because computation, not video, is its medium. Pairs with the Nous local-inference roadmap.

9.7 Teachers as channel, institutions as later revenue#

The launch "no" to institutional procurement is right; add the free middle path — teacher accounts with class-share links and a share-to-teacher artifact (lesson + sources + mastery evidence, already recommended §5.10). Students bring tools to teachers years before districts buy them; the free tier should make that frictionless while LMS interop waits in reserve (it's already built — a P3 switch to flip, not a build).

9.8 App-store positioning#

V9's stances — on-device where possible, no engagement-metric optimization (enforced in code), privacy-respecting learner records — are precisely what platform education-category featuring teams reward. Build the privacy nutrition label and the "no dark patterns" story into the launch plan as assets; being featurable is distribution V9 can't otherwise buy (§9.3).

9.9 Learner-record rights#

Mastery histories are sensitive educational records. Beyond DSAR boilerplate: self-serve export (Open Badges already), a stated retention policy, and an explicit "forget my learning history" control distinct from account deletion (a learner may want a fresh Atlas without losing their account). Trust-branded products get held to standards general apps don't; pre-empt.

9.10 Cross-portfolio notes#

V9 is the intellectual spine of the bundle: Nyx feeds V3's Observatory nights and V5's sky; V8 solves grant logic lessons (make the grants reciprocal — mastery unlocks a themed case, a solved case unlocks the lesson); meditation streaks grant neuroscience threads (V1). Two structural notes: V9's subscription and V1's Oshun+ must be one SKU (V1 §9.1 — two overlapping "premium learning/wellness" subscriptions in one account graph is self-competition), and V9's Aletheia/kernel registry is portfolio infrastructure the same way V8's fairness engine is — Veritas (V1) should eventually cite through the same computed-claims machinery, one truth stack for the whole platform.

10. Monorepo integration & shared functionality (third pass)#

Audited against DOMAINS/, the 185-phase TODOS/ roadmap, V_SERIES.md, V_SERIES_PLATFORM_CONSOLIDATION.md, and package-dependency evidence.

10.1 V9's substrate breadth is exemplary — one wire is missing#

Dependency evidence: libs/v9 genuinely imports the knowledge constellation — @mnemosyne/core (17 sites), @kalika/cosmology and @kalika/symplectic, @nyx/ephemeris and @nyx/constants, @sophia/semantic-search, plus the shared release-gates and quality-judge. Among all nine products, V9 best embodies "compose the monorepo, don't rebuild it." The missing wire is provenance: V9 consumes the shared gates but not @oshun/content-signing — consistent with its honest c2paSigned: false default. The portfolio consolidation (V1 §10.3, V8 §10.1) picks content-signing as the canonical signer; V9 should adopt it in the same change that closes its own G7 gap.

10.2 The kernel roadmap already exists: Kalika Phases 99–131#

The second pass asked for a public "kernel-drop cadence" (§5.5) without noticing the roadmap already contains its supply line: thirty-three Kalika phases (99–115: CAS foundation, pure math, theoretical/advanced physics, computational engine, formal verification, frontier math, classical/continuum physics; 116–131: DFT, crystallography, many-body, phonons, molecular dynamics, ML potentials, defects, transport, spectroscopy, CALPHAD, functional materials, multiscale, autonomous experimentation, reproducible compute). V9's explorable/kernel expansion should be scheduled directly off this ladder — each completed Kalika capability is a candidate computed-kernel explorable and a G2 registry binder. The cadence marketing writes itself ("this week Kalika learned phonons; so did your sky"). Two adjacent phases extend the ceiling: Phase 178 (autonomous research / agentic scientist) for frontier-topic lessons with live-research epistemic status, and Phase 139 (sovereign scholarly/collab suite) for learner notebooks — both should be consumed, not duplicated.

10.3 The Demeter binder and the seven axes#

The V9KernelSchema admits four kernels (nyx, kalika, nisaba, demeter) but only three have binders — while Demeter (Phase 29) is a completed domain (plant databases, phenology, soil). The living-world wonder axis is therefore the cheapest breadth win available: one binder unlocks a whole axis whose substrate already exists. The same audit should walk the remaining axes: Seshat (made-world/craft) and Saraswati (EV/battery/electronics — the technology domain, not V3's concert venue; the naming collision is documented portfolio-wide) both exist as completed domains awaiting kernel adapters.

10.4 The game-bridge stays honest — because Bellona's game side is 0%#

V9 correctly treats learn-by-playing as contract-only. The roadmap confirms the caution: Bellona's adapters (Phase 8) are complete, but the engine-side phases the bridge actually needs — Phase 137 (platform/live-service), Phase 140 (cloud gaming), and the UE cook path every V-game lacks — are unbuilt, and the V2–V5 destinations themselves have no cooked content. The cross-product grant seam (podium→physics-lesson etc.) is the right near-term integration; note that grants presuppose the shared entitlement graph (consolidation Problem 2, unstarted) — V9 should be named as one of its first consumers so the account/entitlement work is scoped to real cross-product use cases, not hypothetical ones.

10.5 Learner data and the ML flywheel#

Phases 85–96 plan passive training-data harvest explicitly including Metis. For V9 this crosses educational-record territory (§9.9): mastery histories, misconception patterns, and wonder streams are exactly the data a flywheel wants and exactly the data learners must control. Inherit V1's training-consent gate (V1 §10.6) with an education-specific clause — no learner data trains models without explicit, revocable, age-appropriate consent — and treat FSRS/mastery telemetry as the most conservative tier. The anti-engagement code enforcement (BannedOptimizationTargetError) shows V9 knows how to make policy structural; do the same for training consent.

10.6 Shared-functionality notes#

  • Nous (Phase 47, complete) — quantized on-device execution is the offline "field edition" (§9.6) and cost floor; V9's docs cite Nous in-principle — add the package edge when the runtime lands.
  • Iris consolidation (finding F2) — V9's own gap analysis flags "seven agent loops, one needed"; V9 should consume the unified Iris loop rather than letting Theia become the eighth (pairs with the adapter-convergence work in V1 §10.2/V6 §10.4).
  • V8 reciprocity — the logic-lesson grants are conceptual today (no V9↔V8 package edge); when the entitlement graph exists, make them data (§9.10's reciprocal unlocks).
  • GPS tier — Daily Wonder leaderboards/streaks-free stats and share infrastructure are GPS domains (telemetry, seasons); V9, like V8, currently has no service stack of its own — keep it that way and consume the shared tier.

11. Closing note#

Every AI-education product of this era makes the same promise — learn anything — and hides the same flaw: you cannot be sure it's true, and you will not remember it anyway. V9 is the first design I've reviewed that attacks both flaws structurally: facts that cannot exist without sources, numbers that cannot ship without being recomputed, and a memory engine disguised as a night sky. The engines are real; the philosophy is enforced in code; what's missing is only the door. Build the screen, answer in five seconds, let the stars dim and be re-kindled — and the curious ape gets the guide it has been improvising around campfires for a hundred thousand years.