Overview#
Metis is the education and course-generation domain of the Oshun platform. Its purpose is to give learners structured, AI-augmented access to a curated curriculum — centred on the core six disciplines of philosophy, religion, psychology, neuroscience, anthropology, and astronomy — while giving instructors the tools to author courses, assess learning, and review AI-generated content before it reaches students.
The real-world problem Metis solves is the gap between raw AI content generation and trustworthy educational delivery. A plain LLM can draft lesson text, but without source-grounded review, item calibration, academic-integrity enforcement, and credential issuance that gap in trust cannot be closed. Metis owns every step of that pipeline: from ingesting reviewed source material, through course authoring and learner progress tracking, to live AI tutoring and the issuance of cryptographically signed credentials.
The production implementation is a Python backend service located at
services/metis/. It is built with FastAPI (async REST API), SQLAlchemy 2.0
(async ORM over PostgreSQL), and Celery (background task processing over Redis).
The service handles course authoring, adaptive assessment, AI tutoring, learner
analytics, source-grounded content generation, credentialing, and LMS
interoperability.
A separate, large TypeScript library tree exists at
libs/metis/*. This document describes the deployed, runnable Python backend underservices/metis/, which is the canonical Metis service.
Process Topology#
Metis is not a single process — it runs as three cooperating process types all built from the same Docker image. This split keeps synchronous HTTP traffic (the API process) decoupled from slow, fallible work like course export and analytics aggregation (the worker and beat processes), so an expensive background job can never block a learner's page load.
┌──────────────────────────────────────────────────────────┐
│ Clients / LMS │
└───────────────────────────────┬───────────────────────────┘
│ HTTPS (bearer JWT)
┌─────────▼─────────┐
│ metis-api │ gunicorn + UvicornWorker ×4
│ FastAPI / :8000 │ metis.main:app
└────┬─────────┬────┘
SQLAlchemy 2.0 │ │ Celery enqueue (Redis broker)
(asyncpg) │ │
┌──────────────▼──┐ ┌──▼──────────────────┐
│ PostgreSQL │ │ Redis │
│ (12 tables) │ │ cache + broker + │
└─────────────────┘ │ result backend │
└──┬───────────┬──────┘
│ │
┌──────────▼──┐ ┌────▼─────────┐
│ metis-worker│ │ metis-beat │
│ Celery │ │ Celery beat │
│ worker │ │ scheduler │
└──────┬──────┘ └──────────────┘
│
┌────────▼────────┐
│ S3 / MinIO │ exports, uploads,
│ object store │ rendered assets
└─────────────────┘
The three process types and their roles:
metis-api—gunicorn metis.main:appwith fourUvicornWorkerprocesses. Serves all REST endpoints under/api, plus root/healthand/ready.metis-worker—celery ... workerconsuming six queues (metis.default,metis.export,metis.notifications,metis.analytics,metis.cleanup,metis.ai). Executes all long-running background work.metis-beat—celery ... beat, the periodic-task scheduler. Triggers hourly analytics aggregation and daily cleanup jobs on a cron-style schedule.
The production container (services/metis/Dockerfile) is a two-stage Python
3.11-slim build: uv installs dependencies and gunicorn; the runtime stage
adds curl, tesseract-ocr, and tini, runs as a non-root metis user,
exposes port 8000, and declares a /health HEALTHCHECK.
Source Layout#
The entire service lives under services/metis/. The internal Python package is
at src/metis/ and follows a conventional layered structure — routers depend on
schemas, schemas are consumed by services, and services read and write ORM
models. Helper modules in services/ hold pure functions that can be tested
without touching FastAPI or the database.
services/metis/
├── pyproject.toml # metis-backend package (FastAPI, SQLAlchemy, Celery)
├── Dockerfile / Dockerfile.dev
├── docker-compose.yml # metis-api / metis-worker / metis-beat
├── alembic.ini / alembic/ # schema migrations (versions/001_initial.py)
├── openapi/ # exported metis.openapi.json (contract artifact)
├── scripts/export_openapi.py
├── tests/ # pytest suite (contract + helper tests)
└── src/metis/
├── main.py # FastAPI application factory + lifespan
├── config.py # pydantic-settings (METIS_-prefixed env vars)
├── deps.py # DI: db session, Redis, S3 client, auth (CurrentUser)
├── openapi_contracts.py
├── api/ # FastAPI routers (16 domain routers)
├── models/ # SQLAlchemy ORM models (12 tables)
├── schemas/ # Pydantic request/response schemas
├── services/ # domain service + helper modules
└── tasks/ # Celery app + task modules
API Routers (src/metis/api/)#
There are sixteen domain routers, each scoped to a distinct feature area. All
sixteen are aggregated by api/__init__.py into api_router and mounted at
/api. The routers cover: auth, courses, assessments, tutoring,
analytics, admin, curriculum, source_ingestion, credentials,
agent_runtime, lti, caliper, qti, scorm, xapi, and oneroster.
Layering#
The service follows a conventional layered architecture where each layer has a single, well-defined responsibility. This makes it easy to test domain logic in isolation and to change HTTP concerns (status codes, pagination shapes) without touching business rules.
api/ — FastAPI routers: HTTP concerns, auth dependencies, status codes
│ depends on
schemas/ — Pydantic models: request validation + response serialization
│ depends on
services/ — domain logic: CourseService, AssessmentService, TutoringService,
│ AnalyticsService, AdminService, AgentRuntimeService, the source-*
│ pipeline services, credential/LMS services, and helper modules
│ depends on
models/ — SQLAlchemy ORM models over PostgreSQL
config / deps — cross-cutting: settings, DB engine, Redis, S3, auth
Routers obtain a service instance, pass a DB session and (where needed) the
Settings object, and serialize results through Pydantic schemas. Several
feature areas keep service logic split into a *_service.py orchestrator and
one or more *_helpers.py modules holding pure, testable functions.
Persistence#
Metis stores all relational data in PostgreSQL via SQLAlchemy 2.0 async with the
asyncpg driver. Twelve tables are created by the single Alembic migration
alembic/versions/001_initial.py: users, courses, modules, lessons,
assessments, questions, submissions, tutoring_sessions,
tutoring_messages, enrollments, progress, and achievements. All twelve
inherit id, created_at, and updated_at from TimestampMixin. The content
hierarchy is three levels: Course → Module → Lesson.
The async engine, session factory, and Redis pool are module-level singletons
initialized in the FastAPI lifespan (metis.deps). get_db yields a
transactional session that commits on success and rolls back on exception,
ensuring no partial writes escape to the database.
Some feature areas — source ingestion, lecture packages, credentials, and
agent-runtime artifacts — produce large JSON documents that are better suited to
object storage than relational tables. SourceIngestionService writes these to
a configurable storage directory and can mint presigned S3 upload URLs against
the metis-uploads bucket.
Key Design Patterns#
Application Factory + Lifespan#
metis.main:create_app builds the FastAPI app in a single place, making it easy
to test with different settings: CORS middleware, an X-Request-Time timing
middleware, structured logging, optional OpenTelemetry instrumentation,
exception handlers, health endpoints, and all routers. The async lifespan
context manager opens and closes the database engine and Redis pool, so startup
failures surface immediately rather than at the first request.
Typed Validation at the Boundary#
Every request and response body is a Pydantic model in src/metis/schemas/.
Field constraints (regex patterns, numeric bounds, length limits) and
model_validator hooks enforce domain rules at the HTTP boundary, before any
code touches the database. RequestValidationError is translated into a
structured 422 response with per-field detail, so clients get actionable error
messages rather than generic 500s.
Dependency-Injected Auth#
metis.deps exposes three auth dependencies: get_current_user (any
authenticated request), get_current_instructor_user (instructor or admin; else
403), and get_current_admin_user (admin only; else 403). They decode a JWT
bearer token (python-jose) into a CurrentUser and gate by role. Endpoints
additionally enforce resource-level ownership — for example, tutoring sessions
are accessible only by the owning learner, and assessment calibration endpoints
require the course author or an admin.
Service + Helper Split#
Domain logic lives in services/. Larger feature areas pair a stateful service
class (taking a DB session and/or Settings) with pure helper modules — keeping
algorithmic logic (e.g. item-calibration formulas, concept-graph traversal)
unit-testable in isolation from FastAPI and the database.
Asynchronous Background Work#
Long-running work that should not block an HTTP response runs as Celery tasks
routed to named queues. This includes course export, notification delivery,
analytics aggregation, and session/upload cleanup. Celery configuration — queue
routing, retry policy, and the beat_schedule of periodic tasks — is produced
by Settings.get_celery_config() so it stays in the same place as all other
service settings.
Direct LLM Provider Calls#
TutoringService calls LLM providers directly through their official SDKs
(openai / anthropic), selected by Settings.ai_provider. When no provider
key is configured it falls back to a structured pedagogical response. There is
no provider-abstraction package and no GPU-dispatch layer — the service keeps
its AI dependency surface minimal and explicit.
Correctness Verification & Agentic Media#
Education is a verifiable domain — most of what a lesson asserts can be
checked against ground truth — so Metis runs generated content through a
verification gate before release and generates teaching media through an
author→critic loop rather than emitting it blindly. Both subsystems live in
TypeScript libraries (@metis/verification, @metis/agents,
@metis/multimedia) with the heavy/dangerous work (Manim rendering) isolated in
a Python worker under services/metis.
Verification gate (@metis/verification)#
A VerificationGate composes N independent Verifiers and aggregates their
per-claim and per-criterion results into a single release decision —
pass | needs-human | block. The gate is fail-loud: a required verifier
that is not_configured (no NLI/embedding/PRM backend wired) makes the gate
refuse rather than silently pass. composeP0Gate assembles the default P0
panel; runVerifiedGeneration wraps any generator in a
generate → verify → (regenerate|refine) → re-gate loop, and
runVerifierGuidedGeneration does best-of-N with pessimistic selection. The
verifiers:
- Factuality — atomic claim decomposition (FActScore/SAFE-style) → per-claim evidence retrieval + entailment → aggregation, with claim→source span linking.
- Faithfulness — TRACe (groundedness / relevance / completeness) over RAG context.
- Citation sufficiency — every load-bearing statement carries adequate support.
- STEM correctness — a CAS-lite identity checker for math, a sandboxed JavaScript runner for code, and a worked-solution process verifier that checks each step transition.
- Pedagogical quality — a judge panel with per-criterion reliability diagnostics (ICC consistency, Spearman alignment, item-total discrimination) and a conformal selective threshold (Clopper–Pearson / Hoeffding / RCPS) that abstains and escalates low-confidence items.
- Contradiction detection and uncertainty calibration (ECE + histogram calibrator) round out the panel.
Supporting machinery: a hash-bound evidence ledger stamping every verified
artifact, an eval harness + gold set with recorded baselines
(EVAL_BASELINES.md), a HITL active-learning loop that routes
low-confidence / high-disagreement items to humans and feeds labels back into
recalibration, telemetry + per-mode compute budgets with a kill-switch, and
production hardening — p95 latency/cost ceilings and Welch-z drift +
two-proportion-z champion-challenger promotion gating.
Agentic teaching media (@metis/agents + @metis/multimedia)#
Following Code2Video / TheoremExplainAgent, a lesson becomes video through a
Planner → Coder → Critic loop (runAgenticMediaLoop): the Planner turns a
verified lesson into a temporally coherent scene plan; the Coder compiles
each scene to Manim/Python; the scene is rendered; the Critic inspects the
rendered layout for overlap/clutter/out-of-frame issues and re-checks any
on-screen equation with the Phase-1 math verifier; the loop repairs until a
quality+correctness bar or a render budget. Two correctness gates are
load-bearing: a block from the lesson verification gate aborts before any
media is produced, and the per-scene critic keeps a scene with a wrong on-screen
equation from being accepted.
Rendered scenes are assembled into a finished segment by composeMediaSegment:
each scene becomes a full-frame background layer on the video/* compositor,
synced to a real TTS narration track, with a WebVTT/SRT caption track and a
timestamped transcript (and an optional avatar/lip-sync overlay).
Manim render worker (services/metis/src/metis/media/)#
The renderer executes generated Manim/Python, so it is deliberately not
mounted in the main API process — arbitrary code execution belongs in an
isolated, resource-limited worker. ManimRenderService runs each job in its own
working directory via the manim CLI subprocess with a wall-clock timeout +
process-group kill, POSIX RLIMIT_CPU/RLIMIT_AS limits, ffprobe duration
probing and optional ffmpeg frame extraction, classifying any failure into
syntax | runtime | timeout | unknown and failing loud
(RendererNotConfiguredError) when no Manim binary is present. It is exposed
two interchangeable ways the TS ManimClient can call: a standalone FastAPI job
API (manim_render_app, POST /render) and a one-shot stdin/stdout CLI
(manim_render_cli, spawned by the multimedia
createChildProcessManimTransport). This is distinct from the Yemaya
batch-render path used for prerecorded lecture packages — the Manim worker
renders agentic teaching animations on the verification path, not GPU-bound
lecture batches.
Technology Stack#
The table below lists every significant technology dependency and why it is
used. Python 3.11 is required for its asyncio improvements and type-narrowing
features that make the strict-mypy configuration practical.
| Layer | Technology |
|---|---|
| Language | Python 3.11+ |
| Web framework | FastAPI + Uvicorn (gunicorn UvicornWorker in prod) |
| ORM / database | SQLAlchemy 2.0 (async, asyncpg) over PostgreSQL |
| Migrations | Alembic |
| Cache / broker | Redis (redis.asyncio) |
| Task queue | Celery (celery[redis]) |
| Object storage | S3 / MinIO via boto3 |
| Validation | Pydantic 2 + pydantic-settings |
| Auth | JWT via python-jose; bcrypt via passlib |
| Document parsing | pypdf, python-docx, pillow, pytesseract (OCR) |
| AI providers | openai, anthropic SDKs (optional ai extra; also langchain, chromadb) |
| Observability | structlog, OpenTelemetry (API/SDK/OTLP, FastAPI + SQLAlchemy instrumentation) |
| Email (dev) | smtplib → Mailpit |
| Packaging / build | setuptools; uv for fast install in the Docker build |
| Tooling | pytest + pytest-asyncio, ruff, mypy --strict |
Deployment#
Metis ships as Docker images. services/metis/docker-compose.yml defines the
three process types and wires them to the shared Oshun infrastructure
(docker-compose.dev.yml: PostgreSQL, Redis, MinIO, Mailpit). The production
container runs gunicorn with four Uvicorn workers behind a /health
healthcheck. There is no Metis-specific Terraform, no AWS ECS/CloudFront
configuration, and no managed-GPU infrastructure in the repository.
Cross-Domain Integration#
Metis deliberately keeps its external dependencies narrow. It calls LLM providers directly (rather than routing through a shared AI gateway) because the tutoring context window, model selection, and fallback logic are too tightly coupled to pedagogical concerns to be abstracted away. It delegates voice infrastructure, media rendering, and academic-integrity adjudication to specialist domains — Psyche, Yemaya, and Themis respectively — each of which has its own scaling and governance concerns.
- AI providers — OpenAI / Anthropic, called directly from
TutoringServicevia their SDKs. The boundary exists here because no other Oshun domain shares the same tutoring prompt structure or fallback behaviour. - Psyche — provides the live-voice tutoring runtime. Metis provisions a
short-lived voice bridge via
psyche_*settings and degrades gracefully to text when the runtime is unavailable or breaches latency/fidelity budgets. Psyche owns audio transport and VAD; Metis owns the pedagogical session and the turn-taking state machine. - Yemaya — handles batch media rendering. When Metis generates a prerecorded
lecture package it labels render batches with the
yemaya-batch-renderprovider and themetis-lecture-render-pipelinepipeline tag.LectureGenerationServiceproduces slide/diagram/notes assets; Yemaya is responsible for the actual rendering workers. The boundary means Metis never runs GPU-bound rendering code directly. - Themis — supplies
AcademicIntegrityVerdictadjudications for assessment submissions. Metis is the consuming domain: it stores the verdict assubmissions.academic_integrity_verdict_jsonand validates incoming payloads against the canonical Themis contract inschemas/assessment.py(source_of_record = "themis",consuming_domain = "metis"). Themis owns detection and adjudication; Metis owns the learner-facing presentation and appeal routing.
Observability#
Metis is instrumented at three levels so operators can diagnose problems without adding debug logging to production:
- Structured JSON logging in staging and production (console renderer in
development), produced by
structlogwith request-scoped context vars so a single log line always carries the request ID and user context. - Optional OpenTelemetry tracing with an OTLP gRPC exporter and a
TraceIdRatioBasedsampler, enabled viaMETIS_OTEL_ENABLED. FastAPI and SQLAlchemy are auto-instrumented; health and readiness endpoints are excluded from traces to avoid noise. - Operational endpoints: an
X-Request-Timeheader on every response,/health(liveness check returning service name/version/environment), and/ready(readiness check that verifies PostgreSQL and Redis connectivity, returning 503 if either is unreachable).