# Eve creative-workflow benchmark preparation — 2026-09

Task 12.9 now has a deterministic, content-bound benchmark candidate for four
separate creative families: reference-to-editable motion, original
brief-to-editable motion, measured-house reconstruction, and interactive
architectural walkthrough. It remains blocked, rather than preregistered,
because Task 17.7 has not supplied and ratified the authoritative creative
delivery contracts and coverage map. The absent contract is represented as a
null digest, the limitation is explicit, and the verifier refuses a promoted
status while that dependency is absent.

Each family fixes its independent case identity before variants or output
inspection, prohibits derivatives and practiced capstones crossing splits, and
requires at least 200 independent cases with at least 40 per case class. Three
fixed model seeds are nested under a case rather than counted as extra evidence;
deterministic runtime paths require two replays. Decisions must remain visible
by case class, complexity, native/runtime format, provider route or no-model
route, and fresh versus corrected input.

The benchmark measures quality at the boundary that can establish it. Reference
motion uses decoded frames and native-timeline readback for layout, timing, and
Bezier fidelity. Original motion adds blinded brief, brand, composition, and
originality review. Reconstruction compares native geometry and topology with
independent dimensional ground truth and hard-locks unsupported load-bearing
claims. Walkthroughs are exercised in the actually served runtime, including
input-to-motion latency, frame rate, navigation containment, reconnect behavior,
accessible alternatives, and dimensional readback after engine import. Native
editability is checked after save and reopen; an unopenable project or unserved
build is a zero-tolerance failure.

Every family also carries source-to-verified-delivery measures for success,
failure or abandonment, paired human labor, revisions, wall-clock latency, and
provider/render budget. Route identity comes from runtime receipts joined to the
current approved model registry. Asset metadata, UI/marketing labels, and a
headline clip duration are expressly inadmissible. Model, rendering, storage,
streaming, retry, queue, and revision legs remain in their respective cost and
latency boundaries, and any unpriced provider or render leg hard-locks a family.
Ratios use the same held-out case completed by an independently assigned
qualified professional in conventional native tools without Eve or model help.
Lane order is randomized across at least three professionals, acceptance inputs
and limits are identical and blinded, and a preregistered stopping rule retains
all attempts instead of selecting the best output after inspection.

Visual-judge calibration is itself held out by source identity and requires at
least 100 independent cases per criterion and case class, three blinded human
raters per case, Krippendorff alpha of at least 0.67, Spearman correlation of at
least 0.80, false accepts no greater than 0.05, and false rejects no greater
than 0.10. The model judge is advisory. Humans remain required for missing or
failed calibration, drift, unresolved appeals, dimensional decisions, high-risk
cases, and hard-lock-relevant decisions. Human ratings use fixed anchors for
unusable, substantively revision-requiring, and delivery-ready work.

The decision rule is conjunctive: every applicable target and floor must pass in
every preregistered stratum, all hard locks must remain zero, and all four
families must pass before release. Missing cases, prices, receipts, native
readbacks, labels, repeats, failed attempts, or denominators are `INCOMPLETE`
and block. Amendments are prospective, independently approved, content-hashed,
and made before affected outputs are visible. Criterion/stratum denominators are
locked before execution; Wilson 95% intervals govern binary measures, while
continuous measures use a separately held non-promoting pilot and at least
10,000 case-clustered BCa resamples.

The schema and verifier enforce the exact family membership, case classes,
sample and repeat rules, criteria, numeric thresholds, measurement boundaries,
judge calibration, routing and costing rules, source hashes, dependency status,
and record digest. The focused suite includes a deterministic regeneration check
and real-CLI negative controls for a missing family, pooled repeats, undersized
samples, missing dimensions, weakened thresholds, nonzero hard locks,
unvalidated judges, metadata attribution, unpriced cost, hidden dependency
status, stale source hashes, fabricated ratification, and output-visible
amendments, an assisted comparator, and unanchored human scores.

Verify the retained blocked candidate with:

```bash
node tools/eve-everywhere/verify-creative-workflow-benchmark.mjs
```

Regenerate it deterministically with:

```bash
node tools/eve-everywhere/generate-creative-workflow-benchmark.mjs
```

This artifact contains no held-out creative case, human label, model run, native
project, runtime receipt, quality score, cost, labor, latency, or operator
decision. It is preregistration machinery awaiting Task 17.7, not evidence that
any creative family, Task 12.9, Phase 12, G7, or G14 passes.
