Oshun Platform · Planning

V1 Hosting Move: AWS ECS/Amplify → Single Hetzner Box — 2026-07-09

Everything lives in deploy/v1/hetzner/ (compose stack, Caddy edge,

6sections4 minread1table

On this page

Directive. Deploy everything V1 that targeted ECS (and, before the 2026-07-04 consolidation, Amplify) onto one Hetzner Cloud CCX23 — staging now, production later — with GitHub → Hetzner CI/CD.

This supersedes the hosting half of V1/planning/IAC_CICD_AUDIT_2026-07-04.md (§8 "single plane" — the plane moves clouds; the single-plane idea stays). RunPod GPU inference is unchanged.


1. Shape#

Concern Was (AWS) Now (Hetzner CCX23)
Compute ECS Fargate, 21 services docker-compose project oshun-<stack>, 20 services
Ingress/TLS shared ALB + ACM Caddy edge (auto-ACME), host-routing preserved
Postgres RDS pg16 pgvector/pgvector:pg16 container + nightly pg_dump
Redis ElastiCache redis:7-alpine (AOF, volatile-lru)
Object storage S3 MinIO (public S3 API at s3.<domain>; path-style)
Registry ECR GHCR ghcr.io/greychimp/oshun/<service>
Deploy deploy-ecs.yml deploy-hetzner.yml → SSH → deploy.sh on the box
Cost ≈ $84/mo/env (+NAT/ALB floor) CCX23 flat ≈ €25/mo total (all envs it can hold)

Everything lives in deploy/v1/hetzner/ (compose stack, Caddy edge, setup-server.sh, deploy.sh, backup sidecar scripts, env templates, runbook README). Host-routing map, service env contracts, and per-service memory limits mirror the Terraform services map they replace.

2. CI/CD#

deploy-hetzner.yml keeps deploy-ecs.yml's plan semantics: main-push → staging with affected-services-only builds; v* tags/releases → production behind the production-approval environment; dispatch for manual/single-service. Builds push immutable sha-XXXXXXX (or version) tags plus a mutable per-env tag to GHCR with the workflow's own GITHUB_TOKEN; deploys rsync the bundle + Prisma schema to /opt/oshun/bundle, then run deploy.sh <stack> <tag> [services…] over SSH — registry login (token via stdin), edge sync + caddy reload, data tier up, Prisma schema apply (prisma migrate diff --script | psql, the repo's documented path — a step the ECS workflow never had), rolling up -d, a container health gate that dumps failing logs, an on-box smoke through Caddy, and a runner-side public smoke (real DNS + TLS).

3. Real defects found and fixed en route#

  1. docker/Dockerfile.web ignored the NEXT_PUBLIC_* build args. deploy-ecs.yml passed NEXT_PUBLIC_APP_ENV / NEXT_PUBLIC_OSHUN_BFF_URL, but the Dockerfile never declared the ARGs — every front-end image baked the localhost:4010 defaults. Now declared (empty ⇒ unset, so ??-style app defaults still work) and passed per environment.
  2. telegram-bot was never runnable. @oshun/telegram-bot is a library (no serve target / bot.start()); as an ECS service it would crash-loop. The live inbound path is the BFF webhook (POST /telegram/webhook). Dropped from the deployable set; changes under apps/oshun/telegram-bot now trigger a bff rebuild instead.
  3. Psyche images could not build from the repo-root context. Their Dockerfiles COPY pyproject.toml / COPY src from the service directory; deploy-ecs.yml built all of them with context: .. The new workflow uses the service directory as context.
  4. psyche-api-gateway peer defaults point at wrong ports (orchestrator :8003 vs real :8007, knowledge-base :8004 vs real :8008) and at localhost/"serwaa" DB creds — all pinned explicitly in the compose file.

4. Capacity & the production story (revised same-day: co-host at 6 GiB each)#

Each environment is budgeted at a hard 6 GiB of container memory limits (sum of mem_limits = exactly 6,144 MiB), so staging and production co-host on the one 16 GiB CCX23: 12 GiB of ceilings + Caddy edge + OS headroom + a 4 GiB swapfile as spike absorber. Achieved by sizing each ceiling to its workload class (estimates, not measurements — validate with docker stats once live) and dropping the two 4-worker psyche services (api-gateway, orchestrator) to 2 uvicorn workers via compose command: overrides. Limits are ceilings, not reservations. Per-stack isolation: separate compose projects, networks, volumes, env files, loopback Postgres ports; the shared Caddy edge routes by hostname.

CI blast-radius containment (same-day addition): one deploy user, but one SSH key per environment, installed with restrict,command="/usr/local/lib/oshun/ssh-forced-command.sh <stack>". The wrapper (root-owned, installed by setup-server.sh — deliberately outside the CI-writable bundle) permits exactly two operations per key: rsync writing under /opt/oshun/bundle, and oshun-deploy <its own stack> … (a root-owned shim that takes GHCR credentials as two stdin lines, never argv). A leaked staging key cannot open a shell or deploy production. Documented honestly: both stacks share one rootful Docker daemon, so this contains accidents and casual key abuse — hard isolation still means a second box (the upgrade path: run setup-server.sh there and repoint the production environment's HETZNER_HOST; nothing else changes).

5. AWS path status#

deploy-ecs.yml is dispatch-only (auto triggers removed, deprecation header added); infrastructure/terraform stays valid for a deliberate return to AWS. Launch-gate scripts (verify-phase-12/49) that assert the ECS artifacts still pass — verified, not assumed. V1/DEPENDENCIES.md §5/§20 updated (Hetzner primary, AWS dormant).

6. Known gaps (pre-existing, carried over knowingly)#

  • metis-api-gateway needs a Metis Python backend (METIS_PYTHON_BACKEND_URL) that is not one of the V1 services; proxied routes 502 (fail-loud) until it exists (DEPENDENCIES §21 gap).
  • apps/metis/mobile still has no release workflow (mobile gap, not hosting).
  • Off-box DR is operator-level: Hetzner server backups/snapshots and/or syncing /var/backups/oshun to a Storage Box (runbook §5).