Moderators staff the
queue_moderationlane: triage user-generated content, message-center abuse, persona disputes, and reactive flags against published content. Distinct from reviewers (who certify content for publication) and from support (who handle individual customer tickets).
Audience and prerequisites#
- Audience: trust-and-safety moderation staff and contracted partners staffed against the moderation rota.
- Prerequisites: T&S onboarding (V1-TS-onboarding), policy-foundations module, well-being check-in protocol, RBAC and identity-verification current.
- Refresh cadence: full re-cert every 6 months because policy evolves; monthly micro-drill of the active policy delta.
Learning objectives#
By certification, a moderator can independently:
- Apply the canonical content-policy classes against an unfamiliar item and record the decision with audit-grade evidence.
- Recognize and route crisis-tagged items into the senior-reviewer lane within the documented SLA.
- Use the moderation tooling: queue navigation, bulk actions inside the per-class allowlist, evidence-attachment, decision logging, appeal gating.
- Recognize manipulation patterns: prompt injection, brigading, scraping, evasion via persona-impersonation, watermark removal/reapplication.
- Coordinate with reviewers when a flagged item touches certified evidence or persona-bound content.
- Apply well-being protocols when content load exceeds personal thresholds; rotate before disposition quality degrades.
Curriculum modules#
| # | Module | Duration | Format | Assessment |
|---|---|---|---|---|
| 1 | Policy foundations and canonical classes | 180 min | seminar + class-by-class card decks | 30-item policy classification (≥ 28/30) |
| 2 | Moderation tool walkthrough | 60 min | hands-on against staging queue | tool-fluency observation |
| 3 | Evidence handling and audit-grade decisions | 90 min | guided decisions + peer review | rubric pass on 10 decisions with audit packet |
| 4 | Crisis routing (joint with support, reviewers) | 120 min | scenario rehearsal + senior shadow | 4 crisis cases routed within SLA |
| 5 | Manipulation patterns and adversarial content | 120 min | curated adversarial set + red-team drill | adversarial-classification test (≥ 85% precision/recall) |
| 6 | Persona-impersonation and identity disputes (joint w/ T&S) | 90 min | scenario rehearsal | 3-case persona-dispute decision pass |
| 7 | Appeals coordination and rollback awareness | 60 min | walkthrough of appeals lane | 2-case appeals-coordination drill |
| 8 | Well-being protocol | 60 min | guided + clinician-led discussion | self-report adherence + supervisor sign-off |
Decision rubric#
Every moderation decision is graded on:
- Policy fit: the chosen class is the correct, narrowest applicable class.
- Evidence: the decision packet contains the items needed for an appeals review (the item, the surrounding context, the policy line applied, prior decisions on the same actor, the operator's reasoning).
- Action proportionality: the action (warn, remove, hide, suspend) is proportional to the harm and the actor's history.
- Customer-visible response: the templated notice fired carries the correct reason code and appeals link.
Scenario rehearsal#
- Crisis content with grounded source: a user posts a self-harm request alongside a quoted Veritas claim. Moderator routes to crisis lane, suspends the item, escalates to senior reviewer for the grounded element, and ensures the message-center crisis-routing path runs.
- Persona impersonation: a third party claims a Lilith-controlled persona is misrepresenting them. Moderator captures the dispute, suspends customer-visible publication on the contested persona surface, routes to the persona-dispute lane.
- Coordinated brigade against a Veritas claim: detect coordinated reports, separate good-faith reports from brigade noise, escalate abuse to security with the cohort packet.
- Watermark-evasion suspicion: a redistributed asset surfaces with the Oshun watermark stripped or reapplied. Moderator routes to the watermark-verification lane, freezes promotion if the source is still live.
- Bulk action under policy allowlist: hold a 200-item bulk decision on a clear policy class; verify per-item evidence; perform the bulk action with the documented authorization step.
Certification criteria#
- All 8 modules complete with passing assessment.
- All 5 scenarios passed; failures require remediation, 30-day re-test, and shadow shifts with a senior moderator before re-certification.
- 40 hours of shadow shifts with a certified moderator.
- 20 hours of live-supervised disposition with rubric pass.
- T&S lead and well-being supervisor sign-off.
Tabletop drills (post-certification)#
- Monthly: active policy-delta micro-drill (new classes, new edge cases).
- Quarterly: moderation-surge drill alongside the support team.
- Semiannually: coordinated abuse-wave drill alongside security.
Well-being protocol#
- Mandatory rotation off content load every 90 minutes for high-exposure classes (graphic, abuse, minors).
- Confidential clinician access on-call during all moderation shifts.
- Self-report tooling for elevated load; supervisors enforce rotation even when self-report is silent.
- Quarterly well-being review; mandatory time-off compliance audited by T&S lead.
Owner#
The T&S lead owns this training program. Privacy lead, security lead, and a contracted well-being clinician review the well-being and adversarial-content modules. Updates are version-controlled and distributed via the policy-delta micro-drill.