Disciplines · Audits

Eve release-scoped evaluation audit — Task 10.1

The seven release-scoped eval cases already exist.

4sections3 minread

On this page

Date: 2026-09-12

Decision#

The seven release-scoped eval cases already exist. Task 10.1 verifies that substrate; it does not re-author the cases and does not claim that automatic release-scope activation exists.

The runtime source of truth still declares V1.0, with Veritas and Metis deferred to V1.2. The complete eval deck contains exactly six family cases and one golden case with executable releaseBlocked metadata. Every case is advisory, every current expectation forbids its withheld tool, and every exact restore expectation calls that tool.

Case V1.0 withheld tool V1.2 restore Retained raw k=10
veritas-tool-selection veritas_top_claims call claims; do not substitute Tara recommendations 10/10
family-md-empty-saved-articles veritas_saved_articles call saved articles; require honest empty result and no fixture headline 4/10
family-md-course-continue metis_continue_learning call continue learning 10/10
family-md-metis-search metis_search_catalog call catalog search 6/10
family-md-metis-recommend metis_recommended_courses call recommendations and ground the fixture course 10/10
family-md-empty-course-search metis_search_catalog call catalog search; require honest empty result and no fixture course 8/10
family-md-empty-recommend metis_recommended_courses call recommendations; require honest empty result and no fixture course 10/10

Retained measurement#

The latest retained production-binding measurement is the 2026-08-30 Phase-11 receipt. It contains 70 runs (seven cases at k=10), the raw 58/70 split above, and a current-grader classification of 59/70 after one narrow, logged boundary phrase repair. It records zero provider retries, $0.0295 billed cost, 8.1 s median turn latency, serving through DeepInfra and Baidu, 83.6% cache-read over 63 reporting runs, and complete teardown of both run-unique databases. No withheld tool was admitted.

This is retained sanitized aggregate evidence, not raw transcript evidence. The record preserves per-case counts, providers, usage, cost, latency, retry, and cleanup fields; it does not preserve prompts or response text. Task 10.1 does not rerun the model and moves no full-deck, family, or builder-Wilson floor.

The focused BFF TypeScript project for this audit passes. The complete BFF typecheck is not a green branch-wide signal at this checkpoint: a 6 GiB run reached its heap cap, and a supervised 8 GiB retry completed with existing diagnostics in unrelated Arete/Yemaya libraries and none in the files touched by this task. Those upstream diagnostics are not relabelled as Task 10.1 failures, and the smaller project is retained as the exact changed-source typecheck.

Executable proof#

release-blocked-eval-audit.ts derives the seven cases from the complete deck, executes the deck's existing admission validator, and independently checks the exact V1.2 expectations. The deterministic capture binds that projection, the authoritative release constants, the retained measurement, its verifier, and the scorecard by SHA-256. The semantic verifier rechecks the current files and rejects inventory narrowing, an advisory case promoted early, a withheld tool allowed in V1.0, weakened restore behavior, a fabricated automatic selector, k below 10, inflated measurement, invented raw transcripts, stale sources, and widened floor claims. Each control is observed red before the clean regression.

Boundary#

releaseBlocked.restoreExpectation remains metadata. The evaluator still does not select it from the authoritative runtime release scope; that is Task 10.2. Task 10.3 owns provider-free dual-scope behavior and inverted/missing mapping controls. Task 10.4 owns a fresh V1.0 k≥10 rerun and the eventual V1.2 rerun. Therefore this audit closes Task 10.1 only; Phase 10 and G8 remain open.

Machine-readable audit: docs/audits/eve-sota-release-scope-eval-audit/2026-09-12.json

Evidence manifest: docs/audits/eve-sota-evidence/phase-10/task-10-1.json