Disciplines · Audits

Eve — SOTA Gap Audit (2026-09-01)

Every closed gate in Eve's record closes "MET WITH GAPS" with the gaps named (SMX exit gate: judge provisional, embedding honestly unbound, human-label-blocked tournament).

6sections13 minread

On this page

An honest answer to the operator's question: "is Eve truly perfectly designed and SOTA based on current industry-leading practice, or are there gaps?" — extended, at the operator's request, to cover computer-use capability and the ability to operate creative suites (Blender, Unreal, and the wider DCC estate).

This audit compares Eve as she is against industry state of the art as known at audit time (the design doc's own mid-2026 research plus this session's knowledge). It follows the estate's status vocabulary: a claim below is either verified live this audit, sampled this audit (file read, shape confirmed, no adversarial pass), or cited from a dated record. Nothing is graded from taste.

Charter context: the Eve handbook mission (recorded 2026-09-01, V1/docs/eve/README.md) — Eve is the agent who helps build the rest of the full product vision, V1 and every release beyond. Gaps below are ranked by leverage against that charter.

Execution ledger: EVE_SOTA_GAP_CLOSURE_TODOS_2026-09-01.md (repo root) carries the phased task breakdown that closes this register; its checkboxes, not this audit, are the source of truth for closure state.


Verdict#

Not perfectly designed — and the estate's own method forbids the claim. Every closed gate in Eve's record closes "MET WITH GAPS" with the gaps named (SMX exit gate: judge provisional, embedding honestly unbound, human-label-blocked tournament). The accurate three-part statement:

  1. Eve's operating method — honesty as wire-level mechanism, regression- locked measurement, governed attributable agency, behavioral cost science — is ahead of any published industry practice this audit can name.
  2. A set of capability ceilings are deliberate, recorded tradeoffs, not oversights.
  3. A set of real gaps are behind current industry SOTA, including two large dormant estates — computer use and DCC/creative-suite agency — where the machinery exists in the repo but no path connects it to Eve.

Part 1 — Ahead of industry (keep, and protect)#

  1. Honesty as wire property (cited: polish gate 2026-08-15, V1/OMNIPRESENT_ASSISTANT_DESIGN_2026-08-03.md §Polish gate). The completion-checker family holds claims off the stream until the server proves them (announce-before-act hold; lookup-claim, audit-claim, interior-location, deferred-room checks); verified is machine-only via the artifact-diff verifier; unsupported shipped claims become recorded ship-verify-gap events. Industry enforces truthfulness with prompts and post-hoc evals; Eve enforces it in the transport.
  2. Measurement discipline (cited: SMX design + scorecard). Per-family regression-locked decks, Wilson floors, negative controls beside every harness, Fisher-exact arm comparisons, prompt-byte ratchets, drift alarms, and the doctrine that a lock that cannot fail proves nothing.
  3. Governed agency (cited: product-graph restructure; verified live 2026-08-22 batteries). Event-sourced intent ledger, per-request attribution (x-workbench-agent), lease protocol, confirm-first cards, independent verification. Industry agent frameworks log; Eve's govern.
  4. Behavioral cost science (cited: docs/agents/eve-smx-maintenance.md). Cached-run billed cost over sticker price, fp8 route pinning, byte-identical HEAD controls attributing failures to the serving route — ~100x under frontier with measured floors.
  5. Accessibility-symmetric tours (cited: design doc §3.5). LLM-planned, deterministically played on the a11y anchor registry; the same tour serves sighted and screen-reader users.

Part 2 — Deliberate variance (recorded tradeoffs, re-examine on schedule)#

  • Smallest-passing serving model, not frontier. The ~100x moat, with residuals recorded (EVE-VIS-280 tour re-authoring; the held-out drift miss where the champion asserts docs-absence without searching — recorded, not repaired). Re-examine at each downshift review (glm-4.7-flash is the named candidate).
  • Async voice, not realtime speech-to-speech. ~90% of value at ~10% of complexity; upgrade path preserved through the psyche provider seam.
  • A11y-tree/anchor grounding, not vision. More precise and radically cheaper on an owned DOM; a real limitation on unregistered content.
  • Confirm-first autonomy, not autonomous fleets. A governance stance this audit endorses; its throughput cost is G1 below.

Part 3 — Gap register (ranked by leverage against the charter)#

G1 · Governed agent lanes exist; executable orchestration does not (S1)#

Two attributed lanes (eve-codex, claude-code) already lease, brief, report, ship, and submit work for machine-only verification. A documented serial queue-drain recipe also exists and was rehearsed end to end on two real work items. What is still missing is an executable, supervised, recoverable drain or planner/orchestrator: selection and dependency ordering remain operator-driven, there is no durable continuous loop, and sustained throughput has not been measured. Seed: implement the existing recipe as a governed drain with local concurrency 1 and cloud concurrency N only after operator spend approval; add verifier-triaged retry and surface the queue, dependencies, leases, cost, and failures in the drawer. The governed substrate is already stronger than what many fleet systems start from; the gap is safe orchestration and proof, not policy from zero.

G2 · Retrieval is sparse-only; the dense leg is declared and unbound (S1)#

Docs search is BM25/BM25F-lite (tools/build-assistant-docs-search.mjs). The embedding/reranker leg exists in the model registry as a fail-loud seam that is honestly unbound (apps/oshun/bff/src/assistant/model-registry.ts:143 — "when the reranker lands, price the route first and bind it here"). Industry SOTA is hybrid sparse+dense with rerankers. Clearest single technical gap. Seed: price an embedding route per the cost rule, bind the leg, A/B the deck's grounding families against BM25-only before promoting.

G3 · The second brain is built and unfused (S1)#

CORRECTION 2026-09-01, same day: this audit first shipped claiming the checklist sat at "23 open / 0 done" — a misread of a search result. The full read says the opposite. docs/audits/HERMES_PARITY_TODOS.md is complete, 23/23: spine wired, real I/O tools (Playwright browser, media, de-stubbed executors), autonomous scheduling (durable stores, NL→cron, unattended runner), the Voyager-style skill library, and breadth (Telegram/Slack/Discord/ email channels, SSH/Modal/Singularity exec backends, governed subagents) — 352 tests, with its own honest seams named (Signal live leg and the Singularity live run infra-blocked; Qdrant semantic recall explicitly deferred at 0.3).

The gap is therefore not construction but fusion: SMX was explicitly designed so this stack could consume the same deck harness ("two brains, one measurement doctrine") and that consumption never started — no task families, no floors, no misroute audit cover the autonomous brain; none of its channels or watchers is wired into any Eve plane; and whether any channel runs live for the builder is unverified. A complete, tested second brain with no measurement and no product surface is dormant capability, not capability. Seed: background watchers (design-doc feature 7) are the natural first fusion — its scheduling subsystem is already built; wire it to a drawer tool and notifications, measured by the existing deck discipline.

G4 · Workbench breadth: 3 of 57, new generation off-kit (S1)#

workbench_kit_read covers tara-workbench, hathor, isis — 31 views, hardwired seam union (workbench-kit-read.ts:250); kit mutations cover tara only (5 card-gated commands). The 7.3 sweep's zero-view verdicts for the other 54 were correct when recorded (no durable server state to read) — but Phase M is now building exactly such state (item banks, clarity/readability/alignment gates) on apps/metis/service, outside the kit seams, with no Eve path. Veritas, Yemaya, Euterpe, Aja, and Bellona will repeat this as their phases start. Seed: per the sweep's own re-audit discipline — when Phase M surfaces stabilize, delta-survey and register a metis seam (views + ratchet re-stamp + battery extension, the 7.3.3 recipe); make "Eve seam registered" a standing exit-gate line for every future workbench phase.

G5 · DCC/creative-suite agency: estate exists, no Eve path (S1; sampled)#

The bellona domain holds ~40 libraries: a real Blender RPC bridge (libs/bellona/blender-agent/src/blender-rpc-bridge.ts — spawns Blender, WebSocket/stdin transports, execute_python/invoke_operator, headless and interactive modes) under a typed action schema with an execution audit log, automation validation rules, geometry-nodes authoring/diff/merge, staging presets, kitbash and packaging workflows; an MCP gateway (47 files, stdio-server.ts, remote-control resources, dry-run planners with cost/time estimators and impact summaries, autonomous stop criteria, capability memory); adapters for Unreal (bridge/cook/onbox/plugin-transport/metahuman), Unity (+unity-agent), Godot, Houdini, Maya, 3ds Max, DaVinci, OpenUSD, mocap, virtual production, XR, text-to-3d; Python and C++ SDKs. Beside it, iris owns a voice→DCC delegation lane (libs/iris/conversation-intent/: dcc-intent-classifier, dcc-agent-delegation-router, dcc-availability-awareness, dcc-voice-launch with confirmation decisions, dcc-status-voice-streamer), and yemaya a Blender film-capture integration.

Verified this audit: no reference to any of it in Eve's tool surface (the only psyche import in the BFF assistant is voice providers). The V2–V10 charter releases are engine/content products; an Eve that cannot operate the DCC estate cannot help build most of the vision's back half.

Three honesty notes before anyone wires this:

  • Bellona has no recorded fabrication audit. The 2026-07 audit ledger covered hathor/yemaya/maya-loom and V2–V9, not bellona. Shapes sampled here are real transports, not paper — but the estate predates the .spec.ts ratchet and MUST get the full adversarial pass (.claude/rules/task-checkbox-verification.md) before Eve claims any of it.
  • Unreal legs are hardware-gated locally (no UE on this Mac — recorded across the V4/V5 audits). Blender headless RPC is the cheapest first light and runs here.
  • The Phase B workbench console (5/1264 done) is the console, not the libs — the libs are the asset.

Seed: wire the bellona MCP gateway into the agent plane (leased, attributed, dry-run-first, verified), never the copilot drawer — the four-planes authority rule already decides this. First light: a leased work item drives one headless-Blender execute_python round trip through the existing dry-run + audit-log machinery, verified by artifact diff.

G6 · Computer use exists as a library and is unwired (S2; sampled this#

audit)

libs/psyche/computer-use-core is a full agent loop — screenshot, screen analyzer, action executor, safety checker (agent.ts) — beside libs/psyche/browser-automation. Industry has shipped OS-level computer-use agents since 2025; the 2026 differentiator is governed GUI agency, which is exactly the estate's house strength. Two blockers, both honest: the core binds @anthropic-ai/sdk directly, which violates the test/harness model-binding cost rule the moment it is exercised — it must be rebound through the provider seam first; and like bellona it has no recorded fabrication audit. Seed: rebind through @oshun/ai, adversarial-audit, then admit it as the agent plane's highest-risk tool class — lease + kill-switch + budget guards, same as the agentic-studio runtime. Note the recorded environment landmine: the Chrome-extension window occlusion on this Mac means live verification runs through the Playwright harness, not the extension.

G7 · The judge is provisional and human-label-blocked (S2, human-gated)#

The entire deck tower rests on a judge pin whose validation tournament needs the operator's ≥40-transcript labeled set (SMX 7.2, open by design — the only reopen key is human labels). Until then "the deck decides" has an unvalidated decider. Seed: the pre-registered protocol in docs/agents/eve-smx-maintenance.md; this is operator work, not Claude work.

G8 · Seven release-scoped cases await automatic activation, not re-authoring (S2)#

The six family-deck cases plus one golden case already grade honest V1.0 behavior as advisory evidence: current expectations forbid the withheld Metis/Veritas tools and accept only the named boundary or an explicitly grounded adjacent-room answer. Each carries executable releaseBlocked.restoreExpectation metadata for its V1.2 semantics, and all seven have k=10 measurements. The remaining gap is that expectation selection does not yet consume the authoritative runtime release scope automatically. Seed: make the grader activate the stored V1.2 expectation when the runtime scope opens those rooms, exercise both scopes provider-free, then remeasure. The V1.0 annotations remain active until that release change actually occurs.

Task 10.1 reverified this exact substrate on 2026-09-12 and bound the seven current expectations, exact restore expectations, advisory state, authoritative release constants, and retained k=10 receipt in docs/audits/eve-sota-evidence/phase-10/task-10-1.json. That verification does not narrow the remaining gap: automatic scope activation is still Task 10.2.

G9 · Interaction breadth ships; depth, accessibility, and parity lag (S2/S3)#

The builder/admin plane already ships three curated tours on the shared tour contract, drawer voice input and per-reply synthetic output with honest fallbacks, scoped selection-ask for crash/incident/review rows, and contextual invocation from crash and incident workspaces. Those paths have focused unit and Playwright coverage. The remaining gaps are different: mobile still uses the buffered non-SSE turn mode; review-row and wider-workspace invocation need a fresh totality inventory; capability metadata understates some shipped voice behavior; long-turn steering and cross-client parity are incomplete; and typed generative UI remains deferred while the industry converges on AG-UI/A2UI-like patterns. Seed: re-baseline every shipped surface first, then close typed registry totality, streaming/control parity, WCAG 2.2 AA, and assistive-tech coverage before deciding whether one allowlisted generative-UI island earns admission.

G10 · Injection cases exist; capability isolation is incomplete (S2)#

The core adversarial deck already contains eight direct/indirect injection cases spanning page headings, page titles, selected text, documentation, work-item/thread content, a demanded write tool, a member role override, and a forged tool-result transcript. They are distributed across existing task families, however: there is no dedicated injection family or Wilson floor, no trust-label/taint-propagation contract across prompt boundaries, and no complete agentic-security crosswalk. Data-not-instructions framing, redaction, scoped grants, and confirm-first remain useful controls, but wiring DCC or native computer use multiplies the untrusted tool-result, screenshot, OCR, and project content entering the loop. Seed: promote and extend the cases into a dedicated injection family, then pair measured floors with centralized effective authority, provenance/taint propagation, isolation, egress controls, and failure-injected negative tests before admitting G5/G6.

G11 · Operator memory is default-on when durable storage exists (S3)#

Operator memory is already default-on when an admin/V1 database is configured; OSHUN_ASSISTANT_OPERATOR_MEMORY=0 is the explicit kill switch, while an unconfigured database degrades honestly to session-only memory. The durable path has integration coverage for persistence across sessions, same-subject isolation, confirm-carded writes/deletes, recall disclosure, and subject-wide deletion without adjacent-subject loss. The remaining gap is memory quality and data-rights depth: operator-visible inspection/correction/forget-all controls, provenance and supersession, dedupe/salience/expiry policy, poisoning resistance, semantic recall, and measured usefulness are incomplete. Seed: build those quality and rights gates over the existing default-on port; do not describe already-shipped durability as a future flag promotion.


Method and verification depth#

  • Verified live this audit: kit-read seam union and view counts; absence of bellona/psyche-computer-use references in the BFF assistant; the unbound embedding leg's registry comment; workbench TODOS per-phase checkbox counts.
  • Source-verified in the correction pass: the two attributed agent lanes, serial drain recipe, and recorded two-item rehearsal; Hermes 23/23 ledger and recorded 352-test gate; seven executable release-blocked contracts and their k=10 evidence; three admin tours plus drawer voice, row selection, and crash/incident invocation implementations; eight core adv-injection-* cases across non-injection families; and operator-memory default/kill-switch, durable, isolation, confirmation, disclosure, and deletion contracts. This pass re-read source and committed evidence; it did not relabel those dated provider/browser runs as newly executed.
  • Sampled this audit (shape confirmed, no adversarial pass): Blender RPC bridge, bellona MCP gateway, computer-use-core agent, iris DCC voice lane.
  • Cited: everything date-stamped above (design doc, SMX design/TODOS, Everywhere register, polish gate, maintenance runbook, 2026-07 audit ledger).
  • Nothing in Parts 1–3 rests on a claim this audit could not point at.

The shortlist, if the next initiative starts tomorrow#

  1. G4 Metis seam survey and registration once its current route/store/authz boundary is stable enough to bind without creating a shadow API.
  2. G1 executable, supervised serial drain with dependency/retry/recovery semantics; the written policy and two-item rehearsal are inputs, not the missing deliverable.
  3. G2 measured dense retrieval leg and hybrid comparison against the current BM25 baseline; promote only on relevance, grounding, security, latency, and cost evidence.
  4. G10 dedicated injection family and capability-isolation controls, followed by the Bellona fabrication/security audit and only then G5 leased headless-Blender first light.
  5. G3 background watchers as the first measured, governed fusion of the already complete Hermes construction with Eve's product and evaluation planes.