Checklist: YSD-22058 (deletion), YSD-3063, YSD-3064 Decision owner
role: the rights owner for the study workspace Machinery:
libs/yemaya/study-workspace/src/use-cases/deletion-saga.ts,
apps/yemaya/svc-study-workspace/src/persistence/deletion-saga-store.ts
This exists because procedure-exercise-register.json filed the deletion
procedure against incident-runbooks.md#stuck-deletion. That document is right
about what it covers and it is not this: a deletion that got stuck is an
incident, and deleting a source because somebody asked is an ordinary Tuesday.
Pointing the procedure at the incident meant the register held no document for
the ordinary case, and — because the incident runbooks are also the incident
procedure's document — any duration written into them would have become both
procedures' objective at once.
Running this document IS the drill. Record the result in
procedure-exercise-register.json; a drill that leaves no mark is a drill
nobody can cite.
1. What the system does on its own — read this first#
A deletion is a durable saga (YSD-3063), not a DELETE statement, and it is
already ordered correctly. Knowing that changes what you should be watching.
| Order | Store | What happens |
|---|---|---|
| 1 | object-storage |
Stored bytes go first — originals, sidecars, and export documents embedding the work. |
| 2 | authoritative-records |
The dependency closure is tombstoned in ONE transaction, grants included so nothing re-authorizes. |
| 3 | search-projection |
Acks only after verifying zero live rows remain. |
| 4 | vector-projection |
Same. |
| 5 | graph-projection |
Same. |
Bytes before records is deliberate. If the bytes cannot be removed, no record may claim they were. And the closure is re-discovered at step 2 rather than trusted from the plan, so a record created after the saga started — a new anchor on a doomed edition — dies with it instead of surviving as an orphan pointing at nothing.
The saga resumes from its last acknowledged store, so a crash mid-deletion is
recoverable and re-execution is a no-op rather than a duplicate-tombstone
failure. A saga that stops advancing is the incident, and that is where
incident-runbooks.md#stuck-deletion takes over from this document.
2. Before you start#
- Confirm the request is the subject's, not an operator's. An operator's authorization to run the procedure is not the subject's request to have it run; they are different facts and only one of them makes the deletion lawful.
- Confirm no legal hold covers the work. A hold outranks the retention clock and outranks this request (YSD-3023). The sweep already skips held copies; a manual deletion must too.
- Read the impact first.
GET /api/study/works/:workId/impactreturns the closure the saga will act on WITHOUT deleting anything. If the subject is going to lose notebooks, exports or decisions built on this source, that is knowable now and not after.
3. The steps#
| # | Action | Owner role | Expected |
|---|---|---|---|
| 1 | Confirm the requester, the work, and that no legal hold applies | rights owner | 10 min |
| 2 | Read the deletion impact and confirm what the subject will lose with it | rights owner | 10 min |
| 3 | DELETE /api/study/works/:workId with the project and a written reason |
rights owner | 5 min |
| 4 | Read back every surface and confirm each refuses (§4) | rights owner | 10 min |
Deletion time objective: 45 minutes, from accepting the request to §5 being
satisfied. Set 2026-08-15: the four steps are budgeted at 10 + 10 + 5 + 10 = 35,
and the objective is that plus ten minutes of headroom rounded up to the next
five — the same rule yemaya-study-rollout.md, yemaya-study-rollback.md and
yemaya-study-recovery.md use.
Almost none of that is the machine. scripts/time-deletion.ts drives this
procedure over HTTP against a running deployment — build a subject, ingest a
real film, read every surface, delete, read every surface again — and reports
what it cost. Against the deployment on 2026-08-15, tombstoning 11 record
families and removing the stored object each time:
| rows in the study schema | DELETE |
|---|---|
| ~77,000 | 0.22s |
| ~154,000 | 1.38s |
Twice the estate, six times the duration. A deletion's cost is a function of what else is in the database, not of the work being deleted, because the projection steps have to establish that zero live rows remain — and that is the thing a floor measured on a small estate cannot extrapolate. The objective is for steps 1, 2 and 4, which are a person deciding and a person checking, and it has room for a machine part that grows.
4. Verify — and this is the part that is easy to fake#
A DELETE that returns 200 has confirmed that the call returns. The
register requires two things of a deletion drill, and both are about what is
true afterwards:
- A read-back from every downstream holder returns absent. Not the one you
happened to check.
GET /works/:workIdmust 404,GET /works/:workId/playbackmust refuse, and the work must be gone fromGET /sources— three different read paths over three different stores, and the search projection is the one that historically survives a deletion because nothing asked it. - Derivative and index copies went with the primary object. The response's
deletedObjectKeysmust contain the original's key AND the key of every export document that embedded it. A proxy that outlives its original outlives the rights basis it was made under.
scripts/time-deletion.ts performs all of these and fails loudly if any surface
still answers, so a drill can be run against a deployment without anybody having
to remember the list.
5. Done means#
- The saga reached every store:
deletedObjectKeysandtombstonedRecordsboth non-empty, and no store left awaiting manual review. - All three read surfaces refuse.
- The subject has been told what went with the source (§2), and told that it is done.
- Recorded in the procedure exercise register with the real elapsed time — which is the number the 45-minute objective is a claim about.
Communication#
| Audience | Within |
|---|---|
| The subject who asked, once the read-backs pass | 24 hours |
| The learners whose notebooks or decisions cited the deleted source | 24 hours |
The second row is the one that is easy to skip, and the impact read in §2 is what makes it possible: a deletion that quietly removes the evidence under somebody else's work leaves them holding a claim they can no longer support.
Known shortfall#
One unexplained outlier, recorded rather than smoothed. The first run of
scripts/time-deletion.ts against this deployment took 82.07 seconds for
the DELETE — the service's own request log records it at
durationMs: 82069.9, status 200, no error logged. Four subsequent runs at the
same estate size took 0.21–0.25 seconds, including the first run after a service
restart, so it is not a per-process cold start; and the growth to 1.38s above is
estate size, which had not changed when the outlier happened. Nothing here
explains it. Reporting the 0.22-second figure without it would be reporting the
warm case as the measurement, and if a deletion ever misses the objective on the
machine part, this is the thing to look for first.
The estate is small either way. 100 tables and under 160,000 rows on one box, with a handful of objects in the bucket. What the two measurements do show is that the machine part grows with what else is in the database; what they cannot show is where it stops being negligible against a 45-minute objective.