# Eve chat-surface and generated-media benchmark preparation — 2026-09

Task 12.11 now has a deterministic, schema-validated benchmark candidate over
the dated chat-parity audit, the ratified Task 12.1 taxonomy, and the Task 12.2
evaluation design. It remains blocked because Task 8.12 has not ratified the
source-owned per-surface parity register. The expected Task 8.12 artifact is
represented by a null digest, the dependency is named in the limitations, and
the verifier refuses a preregistered status while that contract is absent.

The benchmark defines 11 independent case classes:

- attachment understanding across documents, spreadsheets, images, audio, and
  video;
- rich reply rendering;
- versioned artifact editing;
- turn controls and branching;
- conversation management and privacy;
- separate image, video, audio, and 3D generation classes;
- cross-surface accessibility; and
- failure/degraded-state honesty.

Every class requires at least 200 independent identities on each applicable
decision denominator. Model-dependent repeats use at least three fixed seeds
nested under the identity rather than counted as additional evidence. Every
class must cover the admin drawer, member web panel/dock, and mobile sheet and
remain visible by surface, risk, fresh/corrected input, measured provider route,
and assistive-technology state. Each class has two planted, outcome-changing
negative controls rather than empty or cosmetic negatives.

The 109 class criteria bind claims to their deciding boundary: typed attachment
and modality oracles, source regions/timecodes, the typed reply AST and rendered
accessibility tree, semantic artifact diffs and save/reopen, durable branch and
turn readback, all declared persistence sinks for temporary/delete tests,
artifact signatures/native readers for generated media, real assistive
technology, and complete job/provider attempt ledgers. Cross-tenant disclosure,
active-content escape, post-cancel/duplicate side effects, unopenable edits,
temporary-chat persistence, deletion residue, false completion, partial jobs
labelled complete, and missing failed-attempt cost are exact-zero hard locks.

Cost uses every priced runtime receipt and a preregistered per-class budget;
latency runs from request receipt to a verified outcome against the per-class
SLO. A headline generation duration is not end-to-end latency. Missing or
unreceipted cost is an exact-zero hard-lock criterion. This internal budget/SLO
measurement stays executable even when a frontier product is inaccessible.

ChatGPT and Claude.ai are separately preregistered as paired comparison products
on the same independent case and acceptance oracle. Executable comparisons
require an authorized product session in the same dated window, plan tier,
locale, device class, prompt/source, time budget, and stopping rule, with failed
attempts retained. Paired outcome-rate difference, blinded quality difference,
cost ratio, and request-to-verified-outcome latency ratio have fixed numeric
targets and floors. No authorized product session was provided in this
environment, so both rows are explicitly
`not-executed-no-authorized-session-provided`. They contribute no numeric ratio
and block any corresponding frontier-parity claim. Published pages, screenshots,
remembered behavior, product labels, and Eve simulations cannot substitute for
execution.

Subjective media grading requires at least three blinded humans per case.
Automated judge use requires 100 independent calibration cases per criterion and
class, Krippendorff alpha at least 0.67, Spearman correlation at least 0.80,
false accepts no greater than 0.05, and false rejects no greater than 0.10.
Humans remain mandatory for failed/missing/drifting calibration and for privacy,
accessibility, provenance, and hard-lock decisions.

The decision is conjunctive: every criterion target and floor must pass at its
minimum independent n in every required stratum; every negative must pass; every
hard lock must be zero; and all 11 classes must pass on all three Eve surfaces.
Missing cases, surfaces, strata, labels, artifacts, prices, receipts, attempts,
or external access are `INCOMPLETE`, never imputed. Changes are prospective,
independently approved, content-hashed, and locked before affected outputs or
comparison results are visible.

The focused suite proves deterministic regeneration and observes 19 real-CLI
negative controls, including class/surface omission, seed pooling, undersized
samples, vacuous negatives, missing dimensions, weakened quality and hard-lock
thresholds, unpriced work, headline latency, weakened paired parity,
unblinded/uncalibrated grading, fabricated external access or ratification,
screenshot substitution, inaccessible comparisons promoting a claim, stale
source bytes, and output-visible amendment.

Verify the retained blocked candidate with:

```bash
node tools/eve-everywhere/verify-chat-media-benchmark.mjs
```

Regenerate it with:

```bash
node tools/eve-everywhere/generate-chat-media-benchmark.mjs
```

This artifact contains no Eve run, external product run, attachment, surface
journey, generated artifact, provider receipt, human label, quality score, cost,
latency, or passing result. It does not close Task 12.11, Phase 12, G7, or G14.
