# Eve — SOTA Gap Audit (2026-09-01)

An honest answer to the operator's question: _"is Eve truly perfectly designed
and SOTA based on current industry-leading practice, or are there gaps?"_ —
extended, at the operator's request, to cover computer-use capability and the
ability to operate creative suites (Blender, Unreal, and the wider DCC estate).

This audit compares Eve as she is against industry state of the art as known at
audit time (the design doc's own mid-2026 research plus this session's
knowledge). It follows the estate's status vocabulary: a claim below is either
**verified live this audit**, **sampled this audit** (file read, shape
confirmed, no adversarial pass), or **cited** from a dated record. Nothing is
graded from taste.

**Charter context:** the Eve handbook mission (recorded 2026-09-01,
`V1/docs/eve/README.md`) — Eve is the agent who helps build the rest of the full
product vision, V1 and every release beyond. Gaps below are ranked by leverage
against that charter.

**Execution ledger:**
[`EVE_SOTA_GAP_CLOSURE_TODOS_2026-09-01.md`](../../EVE_SOTA_GAP_CLOSURE_TODOS_2026-09-01.md)
(repo root) carries the phased task breakdown that closes this register; its
checkboxes, not this audit, are the source of truth for closure state.

---

## Verdict

**Not perfectly designed — and the estate's own method forbids the claim.**
Every closed gate in Eve's record closes "MET WITH GAPS" with the gaps named
(SMX exit gate: judge provisional, embedding honestly unbound,
human-label-blocked tournament). The accurate three-part statement:

1. Eve's **operating method** — honesty as wire-level mechanism, regression-
   locked measurement, governed attributable agency, behavioral cost science —
   is **ahead of any published industry practice** this audit can name.
2. A set of capability ceilings are **deliberate, recorded tradeoffs**, not
   oversights.
3. A set of real gaps are **behind current industry SOTA**, including two large
   dormant estates — computer use and DCC/creative-suite agency — where the
   machinery exists in the repo but no path connects it to Eve.

---

## Part 1 — Ahead of industry (keep, and protect)

1. **Honesty as wire property** (cited: polish gate 2026-08-15,
   `V1/OMNIPRESENT_ASSISTANT_DESIGN_2026-08-03.md` §Polish gate). The
   completion-checker family holds claims off the stream until the server proves
   them (announce-before-act hold; lookup-claim, audit-claim, interior-location,
   deferred-room checks); `verified` is machine-only via the artifact-diff
   verifier; unsupported shipped claims become recorded `ship-verify-gap`
   events. Industry enforces truthfulness with prompts and post-hoc evals; Eve
   enforces it in the transport.
2. **Measurement discipline** (cited: SMX design + scorecard). Per-family
   regression-locked decks, Wilson floors, negative controls beside every
   harness, Fisher-exact arm comparisons, prompt-byte ratchets, drift alarms,
   and the doctrine that a lock that cannot fail proves nothing.
3. **Governed agency** (cited: product-graph restructure; verified live
   2026-08-22 batteries). Event-sourced intent ledger, per-request attribution
   (`x-workbench-agent`), lease protocol, confirm-first cards, independent
   verification. Industry agent frameworks log; Eve's govern.
4. **Behavioral cost science** (cited: `docs/agents/eve-smx-maintenance.md`).
   Cached-run billed cost over sticker price, fp8 route pinning, byte-identical
   HEAD controls attributing failures to the serving route — ~100x under
   frontier with measured floors.
5. **Accessibility-symmetric tours** (cited: design doc §3.5). LLM-planned,
   deterministically played on the a11y anchor registry; the same tour serves
   sighted and screen-reader users.

## Part 2 — Deliberate variance (recorded tradeoffs, re-examine on schedule)

- **Smallest-passing serving model, not frontier.** The ~100x moat, with
  residuals recorded (EVE-VIS-280 tour re-authoring; the held-out drift miss
  where the champion asserts docs-absence without searching — recorded, not
  repaired). Re-examine at each downshift review (glm-4.7-flash is the named
  candidate).
- **Async voice, not realtime speech-to-speech.** ~90% of value at ~10% of
  complexity; upgrade path preserved through the psyche provider seam.
- **A11y-tree/anchor grounding, not vision.** More precise and radically cheaper
  on an owned DOM; a real limitation on unregistered content.
- **Confirm-first autonomy, not autonomous fleets.** A governance stance this
  audit endorses; its throughput cost is G1 below.

---

## Part 3 — Gap register (ranked by leverage against the charter)

### G1 · Governed agent lanes exist; executable orchestration does not (S1)

Two attributed lanes (`eve-codex`, `claude-code`) already lease, brief, report,
ship, and submit work for machine-only verification. A documented serial
queue-drain recipe also exists and was rehearsed end to end on two real work
items. What is still missing is an executable, supervised, recoverable drain or
planner/orchestrator: selection and dependency ordering remain operator-driven,
there is no durable continuous loop, and sustained throughput has not been
measured. **Seed:** implement the existing recipe as a governed drain with local
concurrency 1 and cloud concurrency N only after operator spend approval; add
verifier-triaged retry and surface the queue, dependencies, leases, cost, and
failures in the drawer. The governed substrate is already stronger than what
many fleet systems start from; the gap is safe orchestration and proof, not
policy from zero.

### G2 · Retrieval is sparse-only; the dense leg is declared and unbound (S1)

Docs search is BM25/BM25F-lite (`tools/build-assistant-docs-search.mjs`). The
embedding/reranker leg exists in the model registry as a fail-loud seam that is
**honestly unbound** (`apps/oshun/bff/src/assistant/model-registry.ts:143` —
"when the reranker lands, price the route first and bind it here"). Industry
SOTA is hybrid sparse+dense with rerankers. Clearest single technical gap.
**Seed:** price an embedding route per the cost rule, bind the leg, A/B the
deck's grounding families against BM25-only before promoting.

### G3 · The second brain is built and unfused (S1)

_CORRECTION 2026-09-01, same day: this audit first shipped claiming the
checklist sat at "23 open / 0 done" — a misread of a search result. The full
read says the opposite._ `docs/audits/HERMES_PARITY_TODOS.md` is **complete,
23/23**: spine wired, real I/O tools (Playwright browser, media, de-stubbed
executors), autonomous scheduling (durable stores, NL→cron, unattended runner),
the Voyager-style skill library, and breadth (Telegram/Slack/Discord/ email
channels, SSH/Modal/Singularity exec backends, governed subagents) — 352 tests,
with its own honest seams named (Signal live leg and the Singularity live run
infra-blocked; Qdrant semantic recall explicitly deferred at 0.3).

The gap is therefore not construction but **fusion**: SMX was explicitly
designed so this stack could consume the same deck harness ("two brains, one
measurement doctrine") and that consumption never started — no task families, no
floors, no misroute audit cover the autonomous brain; none of its channels or
watchers is wired into any Eve plane; and whether any channel runs live for the
builder is unverified. A complete, tested second brain with no measurement and
no product surface is dormant capability, not capability. **Seed:** background
watchers (design-doc feature 7) are the natural first fusion — its scheduling
subsystem is already built; wire it to a drawer tool and notifications, measured
by the existing deck discipline.

### G4 · Workbench breadth: 3 of 57, new generation off-kit (S1)

`workbench_kit_read` covers tara-workbench, hathor, isis — 31 views, hardwired
seam union (`workbench-kit-read.ts:250`); kit mutations cover tara only (5
card-gated commands). The 7.3 sweep's zero-view verdicts for the other 54 were
correct when recorded (no durable server state to read) — but Phase M is now
building exactly such state (item banks, clarity/readability/alignment gates) on
`apps/metis/service`, **outside the kit seams**, with no Eve path. Veritas,
Yemaya, Euterpe, Aja, and Bellona will repeat this as their phases start.
**Seed:** per the sweep's own re-audit discipline — when Phase M surfaces
stabilize, delta-survey and register a `metis` seam (views + ratchet re-stamp +
battery extension, the 7.3.3 recipe); make "Eve seam registered" a standing
exit-gate line for every future workbench phase.

### G5 · DCC/creative-suite agency: estate exists, no Eve path (S1; sampled)

The bellona domain holds ~40 libraries: a **real Blender RPC bridge**
(`libs/bellona/blender-agent/src/blender-rpc-bridge.ts` — spawns Blender,
WebSocket/stdin transports, `execute_python`/`invoke_operator`, headless and
interactive modes) under a typed action schema with an execution audit log,
automation validation rules, geometry-nodes authoring/diff/merge, staging
presets, kitbash and packaging workflows; an **MCP gateway** (47 files,
`stdio-server.ts`, remote-control resources, dry-run planners with cost/time
estimators and impact summaries, autonomous stop criteria, capability memory);
adapters for **Unreal** (bridge/cook/onbox/plugin-transport/metahuman), Unity
(+unity-agent), Godot, Houdini, Maya, 3ds Max, DaVinci, OpenUSD, mocap, virtual
production, XR, text-to-3d; Python and C++ SDKs. Beside it, iris owns a
**voice→DCC delegation lane** (`libs/iris/conversation-intent/`:
dcc-intent-classifier, dcc-agent-delegation-router, dcc-availability-awareness,
dcc-voice-launch with confirmation decisions, dcc-status-voice-streamer), and
yemaya a Blender film-capture integration.

Verified this audit: **no reference to any of it in Eve's tool surface** (the
only psyche import in the BFF assistant is voice providers). The V2–V10 charter
releases are engine/content products; an Eve that cannot operate the DCC estate
cannot help build most of the vision's back half.

Three honesty notes before anyone wires this:

- **Bellona has no recorded fabrication audit.** The 2026-07 audit ledger
  covered hathor/yemaya/maya-loom and V2–V9, not bellona. Shapes sampled here
  are real transports, not paper — but the estate predates the `.spec.ts`
  ratchet and MUST get the full adversarial pass
  (`.claude/rules/task-checkbox-verification.md`) before Eve claims any of it.
- **Unreal legs are hardware-gated locally** (no UE on this Mac — recorded
  across the V4/V5 audits). Blender headless RPC is the cheapest first light and
  runs here.
- The Phase B workbench console (5/1264 done) is the _console_, not the libs —
  the libs are the asset.

**Seed:** wire the bellona MCP gateway into the **agent plane** (leased,
attributed, dry-run-first, verified), never the copilot drawer — the four-planes
authority rule already decides this. First light: a leased work item drives one
headless-Blender `execute_python` round trip through the existing dry-run +
audit-log machinery, verified by artifact diff.

### G6 · Computer use exists as a library and is unwired (S2; sampled this

audit)

`libs/psyche/computer-use-core` is a full agent loop — screenshot, screen
analyzer, action executor, safety checker (`agent.ts`) — beside
`libs/psyche/browser-automation`. Industry has shipped OS-level computer-use
agents since 2025; the 2026 differentiator is _governed_ GUI agency, which is
exactly the estate's house strength. Two blockers, both honest: the core binds
`@anthropic-ai/sdk` **directly**, which violates the test/harness model-binding
cost rule the moment it is exercised — it must be rebound through the provider
seam first; and like bellona it has no recorded fabrication audit. **Seed:**
rebind through `@oshun/ai`, adversarial-audit, then admit it as the agent
plane's highest-risk tool class — lease + kill-switch + budget guards, same as
the agentic-studio runtime. Note the recorded environment landmine: the
Chrome-extension window occlusion on this Mac means live verification runs
through the Playwright harness, not the extension.

### G7 · The judge is provisional and human-label-blocked (S2, human-gated)

The entire deck tower rests on a judge pin whose validation tournament needs the
operator's ≥40-transcript labeled set (SMX 7.2, open by design — the only reopen
key is human labels). Until then "the deck decides" has an unvalidated decider.
**Seed:** the pre-registered protocol in `docs/agents/eve-smx-maintenance.md`;
this is operator work, not Claude work.

### G8 · Seven release-scoped cases await automatic activation, not re-authoring (S2)

The six family-deck cases plus one golden case already grade honest V1.0
behavior as advisory evidence: current expectations forbid the withheld
Metis/Veritas tools and accept only the named boundary or an explicitly grounded
adjacent-room answer. Each carries executable
`releaseBlocked.restoreExpectation` metadata for its V1.2 semantics, and all
seven have k=10 measurements. The remaining gap is that expectation selection
does not yet consume the authoritative runtime release scope automatically.
**Seed:** make the grader activate the stored V1.2 expectation when the runtime
scope opens those rooms, exercise both scopes provider-free, then remeasure. The
V1.0 annotations remain active until that release change actually occurs.

Task 10.1 reverified this exact substrate on 2026-09-12 and bound the seven
current expectations, exact restore expectations, advisory state, authoritative
release constants, and retained k=10 receipt in
`docs/audits/eve-sota-evidence/phase-10/task-10-1.json`. That verification does
not narrow the remaining gap: automatic scope activation is still Task 10.2.

### G9 · Interaction breadth ships; depth, accessibility, and parity lag (S2/S3)

The builder/admin plane already ships three curated tours on the shared tour
contract, drawer voice input and per-reply synthetic output with honest
fallbacks, scoped selection-ask for crash/incident/review rows, and contextual
invocation from crash and incident workspaces. Those paths have focused unit and
Playwright coverage. The remaining gaps are different: mobile still uses the
buffered non-SSE turn mode; review-row and wider-workspace invocation need a
fresh totality inventory; capability metadata understates some shipped voice
behavior; long-turn steering and cross-client parity are incomplete; and typed
generative UI remains deferred while the industry converges on AG-UI/A2UI-like
patterns. **Seed:** re-baseline every shipped surface first, then close typed
registry totality, streaming/control parity, WCAG 2.2 AA, and assistive-tech
coverage before deciding whether one allowlisted generative-UI island earns
admission.

### G10 · Injection cases exist; capability isolation is incomplete (S2)

The core adversarial deck already contains eight direct/indirect injection cases
spanning page headings, page titles, selected text, documentation,
work-item/thread content, a demanded write tool, a member role override, and a
forged tool-result transcript. They are distributed across existing task
families, however: there is no dedicated injection family or Wilson floor, no
trust-label/taint-propagation contract across prompt boundaries, and no complete
agentic-security crosswalk. Data-not-instructions framing, redaction, scoped
grants, and confirm-first remain useful controls, but wiring DCC or native
computer use multiplies the untrusted tool-result, screenshot, OCR, and project
content entering the loop. **Seed:** promote and extend the cases into a
dedicated injection family, then pair measured floors with centralized effective
authority, provenance/taint propagation, isolation, egress controls, and
failure-injected negative tests before admitting G5/G6.

### G11 · Operator memory is default-on when durable storage exists (S3)

Operator memory is already default-on when an admin/V1 database is configured;
`OSHUN_ASSISTANT_OPERATOR_MEMORY=0` is the explicit kill switch, while an
unconfigured database degrades honestly to session-only memory. The durable path
has integration coverage for persistence across sessions, same-subject
isolation, confirm-carded writes/deletes, recall disclosure, and subject-wide
deletion without adjacent-subject loss. The remaining gap is memory quality and
data-rights depth: operator-visible inspection/correction/forget-all controls,
provenance and supersession, dedupe/salience/expiry policy, poisoning
resistance, semantic recall, and measured usefulness are incomplete. **Seed:**
build those quality and rights gates over the existing default-on port; do not
describe already-shipped durability as a future flag promotion.

---

## Method and verification depth

- **Verified live this audit:** kit-read seam union and view counts; absence of
  bellona/psyche-computer-use references in the BFF assistant; the unbound
  embedding leg's registry comment; workbench TODOS per-phase checkbox counts.
- **Source-verified in the correction pass:** the two attributed agent lanes,
  serial drain recipe, and recorded two-item rehearsal; Hermes 23/23 ledger and
  recorded 352-test gate; seven executable release-blocked contracts and their
  k=10 evidence; three admin tours plus drawer voice, row selection, and
  crash/incident invocation implementations; eight core `adv-injection-*` cases
  across non-injection families; and operator-memory default/kill-switch,
  durable, isolation, confirmation, disclosure, and deletion contracts. This
  pass re-read source and committed evidence; it did not relabel those dated
  provider/browser runs as newly executed.
- **Sampled this audit** (shape confirmed, no adversarial pass): Blender RPC
  bridge, bellona MCP gateway, computer-use-core agent, iris DCC voice lane.
- **Cited:** everything date-stamped above (design doc, SMX design/TODOS,
  Everywhere register, polish gate, maintenance runbook, 2026-07 audit ledger).
- Nothing in Parts 1–3 rests on a claim this audit could not point at.

## The shortlist, if the next initiative starts tomorrow

1. G4 Metis seam survey and registration once its current route/store/authz
   boundary is stable enough to bind without creating a shadow API.
2. G1 executable, supervised serial drain with dependency/retry/recovery
   semantics; the written policy and two-item rehearsal are inputs, not the
   missing deliverable.
3. G2 measured dense retrieval leg and hybrid comparison against the current
   BM25 baseline; promote only on relevance, grounding, security, latency, and
   cost evidence.
4. G10 dedicated injection family and capability-isolation controls, followed by
   the Bellona fabrication/security audit and only then G5 leased
   headless-Blender first light.
5. G3 background watchers as the first measured, governed fusion of the already
   complete Hermes construction with Eve's product and evaluation planes.
