Task 12.11 now has a deterministic, schema-validated benchmark candidate over the dated chat-parity audit, the ratified Task 12.1 taxonomy, and the Task 12.2 evaluation design. It remains blocked because Task 8.12 has not ratified the source-owned per-surface parity register. The expected Task 8.12 artifact is represented by a null digest, the dependency is named in the limitations, and the verifier refuses a preregistered status while that contract is absent.
The benchmark defines 11 independent case classes:
- attachment understanding across documents, spreadsheets, images, audio, and video;
- rich reply rendering;
- versioned artifact editing;
- turn controls and branching;
- conversation management and privacy;
- separate image, video, audio, and 3D generation classes;
- cross-surface accessibility; and
- failure/degraded-state honesty.
Every class requires at least 200 independent identities on each applicable decision denominator. Model-dependent repeats use at least three fixed seeds nested under the identity rather than counted as additional evidence. Every class must cover the admin drawer, member web panel/dock, and mobile sheet and remain visible by surface, risk, fresh/corrected input, measured provider route, and assistive-technology state. Each class has two planted, outcome-changing negative controls rather than empty or cosmetic negatives.
The 109 class criteria bind claims to their deciding boundary: typed attachment and modality oracles, source regions/timecodes, the typed reply AST and rendered accessibility tree, semantic artifact diffs and save/reopen, durable branch and turn readback, all declared persistence sinks for temporary/delete tests, artifact signatures/native readers for generated media, real assistive technology, and complete job/provider attempt ledgers. Cross-tenant disclosure, active-content escape, post-cancel/duplicate side effects, unopenable edits, temporary-chat persistence, deletion residue, false completion, partial jobs labelled complete, and missing failed-attempt cost are exact-zero hard locks.
Cost uses every priced runtime receipt and a preregistered per-class budget; latency runs from request receipt to a verified outcome against the per-class SLO. A headline generation duration is not end-to-end latency. Missing or unreceipted cost is an exact-zero hard-lock criterion. This internal budget/SLO measurement stays executable even when a frontier product is inaccessible.
ChatGPT and Claude.ai are separately preregistered as paired comparison products
on the same independent case and acceptance oracle. Executable comparisons
require an authorized product session in the same dated window, plan tier,
locale, device class, prompt/source, time budget, and stopping rule, with failed
attempts retained. Paired outcome-rate difference, blinded quality difference,
cost ratio, and request-to-verified-outcome latency ratio have fixed numeric
targets and floors. No authorized product session was provided in this
environment, so both rows are explicitly
not-executed-no-authorized-session-provided. They contribute no numeric ratio
and block any corresponding frontier-parity claim. Published pages, screenshots,
remembered behavior, product labels, and Eve simulations cannot substitute for
execution.
Subjective media grading requires at least three blinded humans per case. Automated judge use requires 100 independent calibration cases per criterion and class, Krippendorff alpha at least 0.67, Spearman correlation at least 0.80, false accepts no greater than 0.05, and false rejects no greater than 0.10. Humans remain mandatory for failed/missing/drifting calibration and for privacy, accessibility, provenance, and hard-lock decisions.
The decision is conjunctive: every criterion target and floor must pass at its
minimum independent n in every required stratum; every negative must pass; every
hard lock must be zero; and all 11 classes must pass on all three Eve surfaces.
Missing cases, surfaces, strata, labels, artifacts, prices, receipts, attempts,
or external access are INCOMPLETE, never imputed. Changes are prospective,
independently approved, content-hashed, and locked before affected outputs or
comparison results are visible.
The focused suite proves deterministic regeneration and observes 19 real-CLI negative controls, including class/surface omission, seed pooling, undersized samples, vacuous negatives, missing dimensions, weakened quality and hard-lock thresholds, unpriced work, headline latency, weakened paired parity, unblinded/uncalibrated grading, fabricated external access or ratification, screenshot substitution, inaccessible comparisons promoting a claim, stale source bytes, and output-visible amendment.
Verify the retained blocked candidate with:
node tools/eve-everywhere/verify-chat-media-benchmark.mjs
Regenerate it with:
node tools/eve-everywhere/generate-chat-media-benchmark.mjs
This artifact contains no Eve run, external product run, attachment, surface journey, generated artifact, provider receipt, human label, quality score, cost, latency, or passing result. It does not close Task 12.11, Phase 12, G7, or G14.