Source: AGENTIC_CONTENT_GENERATION_SOTA_REPORT_2026-06-12.md §6. Rules: one
task at a time; [x] only after the code is read, built, and its tests pass in
this conversation; adversarial stub scan before every commit.
Honesty note (2026-06-13): "framework complete" ≠ "quality verified." The Phases A–E below deliver the integration plumbing (cook pipeline, Iris agentic loop, Yemaya pipeline-runner, AgentRun governance, release gates, portfolio wiring) — and those claims are honestly scoped (note the
[~]markers and fail-loud seams). What they do not establish is that the generated content is compelling: there is no calibrated LLM-judge taste signal, no best-of-N / self-refine inference-time quality, and no corpus-level diversity gate. Those are tracked separately inAGENTIC_CONTENT_QUALITY_TODOS_2026-06-13.md. No[x]in this ledger attests to best-of-N / preference-learning / reward-model capability (those were never claimed here — they lived in the fabricated@nous/*stack, remediated in that ledger's Phase 0), so none required reverting.
Phase A — Unblock the chain#
- A1.1 Bellona UE cook pipeline: RunUAT
BuildCookRuncommand construction (platform-aware paths, typed options) inlibs/bellona/unreal/src/cook/—uat.ts; 33 tests green 2026-06-12 - A1.2 Bellona cook runner: headless execution with UAT log parsing (errors,
warnings, cook progress, known exit codes), fail-loud when engine/RunUAT
absent —
cook-runner.ts+log-parser.ts - A1.3 Bellona headless asset import:
ImportAssetscommandlet + generated UE Python (unreal.AssetImportTask) viaUnrealEditor-Cmd -run=pythonscript—headless-import.ts; success is read from editor marker output only, empty imports fail - A1.4 Content-addressed cook manifest + cook lineage records per
V2/docs/integration/hathor-narrative-contract.md:143-146(sha256 artifact hashing, canonical serialization, verify/tamper detection) —manifest.ts. NOTE: on-engine end-to-end run (real RunUAT on the UE box) is part of C1's vertical slice; library mechanisms are complete. - A2.1 Iris agentic tool-call loop (
@iris/agents-corev2): LLM → tool calls → results → repeat, max-iteration guard, streaming hooks —libs/iris/agents/core/src/agentic/loop.ts; 18 tests green 2026-06-12 - A2.2 Normalized cross-provider tool protocol (Anthropic/OpenAI/Google)
over the existing conversation providers —
protocol.ts+provider-adapters.ts(tool_use blocks ↔ tool_calls ↔ functionCall parts, id synthesis for Gemini, malformed-arguments → recoverable error results). Local providers ride the OpenAI adapter (OpenAI-compatible inference servers). - A2.3 Budget/cost constraints enforced at call time —
budget.ts(llm calls, tool calls, total tokens, USD cost via pricing table, wall-clock deadline; fail-loudUnpricedModelErroron cost caps over unpriced models) - A2.4 Migration shims — two-way bridge
libs/iris/agents/core/src/agentic/oshun-ai-bridge.ts(wrapSharedProvider: any@oshun/aiprovider powers the agentic loop;asSharedProvider: Iris-routed providers serve anyLLMProviderInterfacecall site, which covers Yemaya's CreativeAgent base) + HathorNPCLLMGeneratorinterface lift withSharedProviderNPCClient(libs/hathor/llm-npc/src/platforms/shared-provider-client.ts) so DialogueGenerator accepts Iris-routed backends. 6 bridge tests + 4 hathor tests green 2026-06-12; hathor dialogue/platform suites (19) green.
Phase B — Hathor end-to-end narrative spine#
- B1 V2 reciprocal records + emission. Records with full contract validation
in
libs/hathor/domain-models/src/fighting-game/(22 tests green 2026-06-12: append-only canon events, unordered-pair uniqueness, challenge-loop-aware cycle detection, terminal-path-to-boss analysis, faction coverage, deterministic shot order, package cardinality). V2 emission contract inlibs/hathor/lore-compiler/src/v2/(6 tests green: full artifact set underV2/ue/Content/Generated/, dialogue banks per surface, codex from canon events, UE string-table exports, content-addressed manifest; refuses invalid packages;bellonaCookLineageIdsappended only by the real cook).SignatureStageContextomitted — no field contract exists in any V2 doc; flagged for spec work rather than invented. - B2 LLM narrative generation pipelines — new lib
libs/hathor/narrative-generation/(9 tests green 2026-06-12):generateStoryBeats/generateQuestNarrative/generateNpcBackstorywith NeithNarrativeGenerator/QuestGeneratoroutput as the deterministic planner (bridged via dependency —@neith/integration-hathoris dependency-free so no cycle; relocation unnecessary), canon constraints injected verbatim into every prompt, strict-JSON parsing that fails loud with raw output preserved, full provenance (model, prompt sha256, usage, latency). Writers mount any@oshun/aiprovider, including Iris-routed via A2.4's bridge. - B3 Lore-consistency validation on generated output —
lore-consistency.ts(part of 18 green tests): deterministic layer (entity-attribution resolution, literal prohibition screening) + independent skeptical LLM judge at temperature 0 whose contradiction/ violation verdicts only block when grounded in a verbatim quote that actually appears in the text (hallucinated evidence downgrades to warning). Structured-record validation (the "four validators") runs in the B1 compiler gate; this layer covers prose. - B4 Review workflow —
review-workflow.ts: submitted → in_review → approved → published state machine with append-only audit trail; approval hard-blocked while consistency errors exist; request_changes requires direction and emits aRegenerationDirective; revisions chain to their predecessors; storage behind aReviewStoreinterface (in-memory impl for tests/tooling, persistence owned by services).
Phase C — Yemaya pipeline end-to-end#
- C1 Vertical slice —
tests/integration/flows/agentic-content-slice.integration.test.ts(green 2026-06-12): composes the REAL implementations end to end in-process — narrative generation → lore-consistency judge → review submission → pipeline runner pausing at theyemaya:approvegate → human approval → publish gate → fighting-game record validation → V2 package compilation → content-addressed cook manifest over real files + real RunUAT command construction. Doubles only at true external boundaries (scripted LLM; absent UE engine, where the cook runner is asserted to fail loud rather than fabricate).[~]remaining: the RunUAT execution itself on the UE build box (external operation; runpnpm test:integration:flowsthere with an engine present, then attach the cook lineage id to the package manifest). - C2 Mid-pipeline HITL —
pipeline-runner.tsinlibs/yemaya/orchestration/src/execution/(16 tests green 2026-06-12):yemaya:approvesteps pause the run with a checkpoint; resume requires an explicit decision; approve continues, reject cancels, request_changes injects reviewer direction into the configs of every step transitively downstream of the gate (and only those). - C3 Recovery wiring — same runner: every completed step checkpointed
through a pluggable
CheckpointStore; failed/budget-stopped runs resume from checkpoint re-running only incomplete steps; terminal step failures consult a replanning hook (retry via the existingRetryManager, skip-with-reason, replace-with-subgraph rewiring dependents to the replacement sinks, abort); per-step budget check via the existingBudgetManagerBEFORE dispatch; honestbudget_exceededstop with preserved work. - C4 HTN → dispatcher steps —
htn-translator.ts: structuralPlannedTaskinput (no agents-package cycle), explicit step-type resolution (declaredmetadata.stepTypevalidated against the schema, or caller resolver; unmappable tasks are translation ERRORS, never emitted), dependency remapping, cost annotations feeding the runner's budget checks.
Phase D — Governance plane#
- D1 C2PA cryptographic signing —
libs/isis/3d-asset-library/src/provenance/claim-signing.ts(7 tests green 2026-06-12): completes the embedding planner's deferred-to-export seam with sha256 hard bindings over the real payload bytes and ed25519 signatures over canonical claim envelopes; verify requires digest AND signature, naming every failure. Documented plainly as c2pa-aligned, not the spec COSE container;ClaimSigner/ClaimVerifierinterfaces mount a COSE/X.509 signer without call-site changes. Release-gate hardening is delivered through D3's unified gate service (Isis suites register their quality gates there). - D2 Agentic AI Studio runtime —
libs/iris/agents/core/src/agentic/agent-run.ts(9 tests green): auditable AgentRun envelopes (append-only events + cost ledger), kill switches scoped agent/family/tenant/provider (block new runs, abort in-flight at loop boundaries, audited activation/release), child budget reservation bounded by the parent's remaining cap with unused reservation returned, authority-subset enforcement at spawn, anomaly quarantine (identical-failure and failure-storm thresholds) with audited operator release. - D3 Unified content release gates — new lib
libs/shared/content-release-gates/(13 tests green): one gate schema (validator/eval/manifest/human_signoff) with adapters matching the shapes of Yemaya QA validators, Isis quality scores, and V3/V6/V7 manifest checks; suite evaluation runs every gate (no short-circuit), required gates cannot skip, evidence summaries are mandatory; promotion requires a passing report bound to the artifact's content hash plus named human signoff where demanded; champion-challenger ramping with deterministic traffic routing and one-sided two-proportion z-test readiness (verified against a hand-computed case), settling experiments so evidence cannot be re-litigated.
Phase E — Portfolio rollout#
- E1 V5 side-quest + dialogue generation —
libs/hathor/narrative-generation/src/batch.ts(batch driver: write → consistency-judge → review-queue with per-item failure isolation, no overall success unless every item succeeded) +tests/integration/flows/v5-sidequest-generation.integration.test.ts(2 green): all four audited V5 patterns (talk-to-npc, escort, hunt, scavenge) with the 1930s period-authenticity canon bundle, through release gates with human signoff and champion-challenger readiness. UE-side V5Procgen consumption of compiled artifacts is on-box ([~]). - E2 V4 Calliope commentary — new lib
libs/calliope/match-commentary/(8 tests green, hand-verified constants): deterministic moment scoring (multi-kill chains, clutch context), logistic win-probability turning points, reel selection with separation; two-host per-cell personas (the five spec presets) with structure PINNED by scored data (skipped moments / invented decisive plays fail loud); BiasReview.AICommentary broadcast gate — grounded-quote gender/region/playstyle findings block, ungrounded verdicts drop, failed review fails closed. The V4 Rust service mounts this lib from its boxes ([~]). - E3 V3 concert/scene export —
consent-registry.tsin@oshun/content-release-gates(19 lib tests green: scoped grants with evidence refs, immediate audited revocation, per-surface verification, consent gate failing artifacts that cannot name their voices) +tests/integration/flows/v3-concert-export.integration.test.ts(green): real Ed25519 C2PA signing over export bytes, signed-manifest verification as a gate, named-signoff promotion, and revocation blocking the very next export. - E4 V6 cognition gateway —
libs/iris/agents/core/src/agentic/cognition-gateway.ts(5 tests green): transport-agnostic JSON contract for V6's Rust services (Moirai/Orun) onto the governed run manager — structural validation that never lets malformed requests reach a model, per-tier iteration ceilings (clotho 2 / lachesis 8 / atropos 24), ori-scoped envelopes, kill switches and quarantine inherited from D2. V6's HTTP mount of the handler is product wiring ([~]). - E5 V7 forge AI-assist — new lib
libs/maya/forge-assist/(10 tests green): capability policy making the platform/realm trust boundary executable (realm allowlist with prefix matching, platform-owned namespaces forbidden regardless of declaration, undeclaredoshun.invokesmuggling detected by source scan) and the assist service whose generated artifacts pass the policy or are REJECTED with named violations — never proposed to the Rust forge-sandbox. Sandbox compile/determinism validation remains in the existing forge-sandbox crate ([~]integration on V7's side).