<!-- GENERATED by renderEvaluationGovernance() in
     libs/yemaya/study-workspace/src/evaluation/evaluation-governance.ts.
     Edit the register, not this file: the spec compares them byte for byte. -->

# Evaluation governance — study-evaluation-governance@1.0

Approved by @GreyChimp on 2026-08-14. The approval is the governing act;
the writing is not.

The eleven subjects YSD-18001 names, each as a rule that refuses something.
A rule phrased only as an intention cannot be violated and so cannot be
complied with; every entry below therefore says what becomes impossible.

## gold-set-versioning

Every gold set is a named version whose manifest states work, edition and build identity, licences, allowed uses, source hashes, slices, annotation version, adjudication and known limitations.

**Refuses.** scoring against "the gold set" as though it were one thing. Two runs quoting one score must be able to prove they read the same version, and a set with no version cannot be re-read after it moves.

**Accountable role.** evaluation owner

**Carried by.** libs/yemaya/study-workspace/src/evaluation/dataset-manifest.ts, libs/yemaya/study-workspace/src/evaluation/gold-set-slices.ts

## annotation-guides

Every label cites the guide version it was made under, and a revision declares per task whether it changed a meaning or said the same thing better.

**Refuses.** pooling labels across a meaning change. No diff of prose distinguishes a clarification from a reversal, so the declaration is required and pooling across an undeclared revision is refused rather than averaged.

**Accountable role.** evaluation owner

**Carried by.** libs/yemaya/study-workspace/src/evaluation/annotation-guide.ts

## adjudication

An item is adjudicated to one label only when its task's answer model is single-answer; an interpretive item's disagreement is preserved rather than resolved.

**Refuses.** adjudicating an interpretive item to one label, which makes the score a measure of conformity to whichever reading won — it rises as a system gets narrower, and every system giving the other supported reading is counted wrong.

**Accountable role.** evaluation owner

**Carried by.** libs/yemaya/study-workspace/src/evaluation/annotation-guide.ts

## legitimate-disagreement

A bounded disagreement holds the readings the material supports and states what puts a reading outside them; any supported reading scores correct.

**Refuses.** treating a beginner who chose a different supported reading as having made an error. Inside the bounds it is a difference; only outside them is it teachable, and an evaluation without that distinction reports the corpus’s interpretive richness as learner failure.

**Accountable role.** domain expert panel

**Carried by.** libs/yemaya/study-workspace/src/evaluation/bounded-readings.ts

## licenses

No item enters a gold set before its licence is recorded and its allowed uses are stated; evaluation is an allowed use or the item is not evaluated against.

**Refuses.** acquiring evaluation material on the strength of it being available. Reachable is not licensed, and a set assembled that way cannot be published or shared without re-clearing every item.

**Accountable role.** rights owner

**Carried by.** libs/yemaya/study-workspace/src/evaluation/dataset-manifest.ts, libs/yemaya/study-workspace/src/corpus/open-corpus-catalog.ts

## consent

Material of a private performance, a minor, or an identifiable person enters evaluation only under recorded consent naming that use, and consent is withdrawable.

**Refuses.** reading consent for one purpose as consent for evaluation. A performance shared for study is not thereby a benchmark item, and withdrawal after the fact must remove the item rather than being recorded as regrettable.

**Accountable role.** rights owner

**Carried by.** libs/yemaya/study-workspace/src/policies/minor-consent.ts, libs/yemaya/study-workspace/src/policies/audience-consent.ts

## privacy

Evaluation records carry no identifier that is not needed to score, and protected attributes are slices for measuring fairness rather than inputs a product may infer.

**Refuses.** turning a fairness slice into a product feature. Skin tone, age, disability presentation and voice type exist in the manifest so performance can be measured across them, and a model reading them as inputs is the harm the slicing exists to detect.

**Accountable role.** privacy owner

**Carried by.** libs/yemaya/study-workspace/src/policies/telemetry-privacy-filter.ts

## retention

Evaluation material is retained under the class its rights basis implies, and an item whose basis lapses leaves the set rather than being kept for comparability.

**Refuses.** keeping a lapsed item so that a historical score stays reproducible. The score becomes irreproducible and that is the correct outcome; the alternative is holding material with no basis in order to protect a number.

**Accountable role.** rights owner

**Carried by.** libs/yemaya/study-workspace/src/policies/object-namespaces.ts

## access

Reading evaluation material is a named act with a stated purpose, and every read is re-decided against the live rights and consent state.

**Refuses.** a standing grant that outlives the reason it was given. A window opened for one review does not authorize the next one, and a read that does not say what it is for is refused before anything is returned.

**Accountable role.** evaluation owner

**Carried by.** libs/yemaya/study-workspace/src/policies/support-access.ts

## publication

A published evaluation report states failures, uncertainty, limitations, slice performance, reviewer disagreement and the release decision it supported.

**Refuses.** publishing a headline number without the slices under it. An aggregate that hides a slice it performs badly on is the specific dishonesty the slicing was built to prevent.

**Accountable role.** evaluation owner

**Carried by.** libs/yemaya/study-workspace/src/evaluation/model-card.ts

## deletion

Deletion of an item or a participant’s contribution propagates to every derived record, and a deleted item leaves a tombstone rather than a gap.

**Refuses.** a deletion that leaves derived scores, embeddings or projections behind. A gap with no tombstone is indistinguishable from an item that never existed, which makes an executed deletion unprovable.

**Accountable role.** privacy owner

**Carried by.** Nothing yet — declared and unenforced, which is recorded rather than hidden.

## What this document does not yet hold

1 rule(s) are declared and enforced by nothing: deletion. They govern acts that cannot yet occur or mechanisms not yet built, and are stated so the gap is visible rather than absent.
