This runbook provides guidance for scaling Oshun services to handle increased load or optimize costs during low-traffic periods.
Overview#
Oshun uses multiple scaling mechanisms:
| Component | Scaling Method | Metric | Default Config |
|---|---|---|---|
| ECS Services | Auto Scaling | CPU Utilization | 60-80% target |
| RunPod Workers | Auto Scaling | Queue Depth | 0 min, varies max |
| Database | Manual (instance) | CPU/Connections | As needed |
| Cache (Redis) | Manual (node type) | Memory/Connections | As needed |
ECS Service Scaling#
Current Configuration#
# View current scaling configuration
aws application-autoscaling describe-scalable-targets \
--service-namespace ecs \
--resource-ids service/oshun-production/oshun-api
# View scaling policies
aws application-autoscaling describe-scaling-policies \
--service-namespace ecs \
--resource-id service/oshun-production/oshun-api
Auto Scaling Settings#
| Service | Min | Max | Target CPU | Scale Out Cooldown | Scale In Cooldown |
|---|---|---|---|---|---|
| API | 2 | 10 | 60% | 60s | 300s |
| Worker | 2 | 10 | 60% | 60s | 300s |
| Frontend | 2 | 6 | 70% | 60s | 300s |
Manual Scaling#
Scale up immediately:
# Increase desired count
aws ecs update-service \
--cluster oshun-production \
--service oshun-api \
--desired-count 5
# Verify scaling
watch -n 5 'aws ecs describe-services \
--cluster oshun-production \
--services oshun-api \
--query "services[0].{Running:runningCount,Desired:desiredCount,Pending:pendingCount}"'
Adjust auto scaling limits:
# Increase max capacity
aws application-autoscaling register-scalable-target \
--service-namespace ecs \
--resource-id service/oshun-production/oshun-api \
--scalable-dimension ecs:service:DesiredCount \
--min-capacity 2 \
--max-capacity 20
Scale down (cost optimization):
# During known low-traffic periods
aws ecs update-service \
--cluster oshun-production \
--service oshun-api \
--desired-count 2
# Adjust min capacity for sustained low traffic
aws application-autoscaling register-scalable-target \
--service-namespace ecs \
--resource-id service/oshun-production/oshun-api \
--scalable-dimension ecs:service:DesiredCount \
--min-capacity 1 \
--max-capacity 10
Scaling Based on Custom Metrics#
Scale based on SQS queue depth:
# For worker services processing queues
aws application-autoscaling put-scaling-policy \
--service-namespace ecs \
--resource-id service/oshun-production/oshun-worker \
--scalable-dimension ecs:service:DesiredCount \
--policy-name sqs-queue-scaling \
--policy-type TargetTrackingScaling \
--target-tracking-scaling-policy-configuration '{
"TargetValue": 100,
"CustomizedMetricSpecification": {
"MetricName": "ApproximateNumberOfMessagesVisible",
"Namespace": "AWS/SQS",
"Dimensions": [{"Name": "QueueName", "Value": "oshun-production-generation-requests"}],
"Statistic": "Average"
},
"ScaleOutCooldown": 60,
"ScaleInCooldown": 300
}'
Scheduled Scaling#
For predictable traffic patterns:
# Scale up for business hours (9 AM UTC)
aws application-autoscaling put-scheduled-action \
--service-namespace ecs \
--scheduled-action-name scale-up-morning \
--resource-id service/oshun-production/oshun-api \
--scalable-dimension ecs:service:DesiredCount \
--schedule "cron(0 9 ? * MON-FRI *)" \
--scalable-target-action MinCapacity=4,MaxCapacity=15
# Scale down for night (10 PM UTC)
aws application-autoscaling put-scheduled-action \
--service-namespace ecs \
--scheduled-action-name scale-down-night \
--resource-id service/oshun-production/oshun-api \
--scalable-dimension ecs:service:DesiredCount \
--schedule "cron(0 22 ? * * *)" \
--scalable-target-action MinCapacity=2,MaxCapacity=6
RunPod Scaling#
Endpoint Configuration#
RunPod endpoints scale automatically based on queue depth.
View current configuration:
curl -X POST https://api.runpod.io/graphql \
-H "Authorization: Bearer $RUNPOD_API_KEY" \
-d '{"query": "{ myself { endpoints { id name workersMax workersMin gpuIds } } }"}'
Scaling Parameters#
| Endpoint | Min Workers | Max Workers | Idle Timeout | GPU Type |
|---|---|---|---|---|
| ComfyUI SD1.5 | 0 | 5 | 60s | A40 |
| ComfyUI SDXL | 0 | 3 | 120s | A100 |
| Flux | 0 | 3 | 120s | A100 |
| SD | 0 | 10 | 30s | 4090 |
Adjusting Max Workers#
Via RunPod Dashboard:
- Go to RunPod Console
- Select endpoint
- Click "Edit"
- Adjust "Max Workers"
- Save
Via GraphQL API:
curl -X POST https://api.runpod.io/graphql \
-H "Authorization: Bearer $RUNPOD_API_KEY" \
-d '{
"query": "mutation { updateEndpoint(input: { id: \"ENDPOINT_ID\", workersMax: 10 }) { id workersMax } }"
}'
Scaling for Expected Load#
Before a known high-traffic event:
- Increase max workers on all endpoints
- Consider keeping minimum workers warm (min > 0)
- Monitor queue depth during event
# Set minimum workers to avoid cold starts
curl -X POST https://api.runpod.io/graphql \
-H "Authorization: Bearer $RUNPOD_API_KEY" \
-d '{
"query": "mutation { updateEndpoint(input: { id: \"ENDPOINT_ID\", workersMin: 2, workersMax: 10 }) { id workersMin workersMax } }"
}'
After event:
- Reset minimum workers to 0
- Reduce max workers if not needed
- Review costs
Database Scaling#
RDS Scaling Options#
| Scaling Type | When to Use | Downtime |
|---|---|---|
| Vertical (resize) | CPU/memory consistently high | Yes (minutes) |
| Storage autoscale | Storage approaching capacity | No |
| Read replicas | Read-heavy workload | No |
| Aurora Serverless | Variable/unpredictable workload | Migration req. |
Vertical Scaling#
Check current utilization:
# CPU utilization
aws cloudwatch get-metric-statistics \
--namespace AWS/RDS \
--metric-name CPUUtilization \
--dimensions Name=DBInstanceIdentifier,Value=oshun-production \
--start-time $(date -d '7 days ago' -u +%Y-%m-%dT%H:%M:%SZ) \
--end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
--period 3600 \
--statistics Average Maximum
# Memory
aws cloudwatch get-metric-statistics \
--namespace AWS/RDS \
--metric-name FreeableMemory \
--dimensions Name=DBInstanceIdentifier,Value=oshun-production \
--start-time $(date -d '7 days ago' -u +%Y-%m-%dT%H:%M:%SZ) \
--end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
--period 3600 \
--statistics Average Minimum
Resize instance:
# Schedule resize during maintenance window
aws rds modify-db-instance \
--db-instance-identifier oshun-production \
--db-instance-class db.r6g.xlarge \
--apply-immediately false
# Or apply immediately (causes downtime)
aws rds modify-db-instance \
--db-instance-identifier oshun-production \
--db-instance-class db.r6g.xlarge \
--apply-immediately true
Adding Read Replicas#
For read-heavy workloads:
# Create read replica
aws rds create-db-instance-read-replica \
--db-instance-identifier oshun-production-read1 \
--source-db-instance-identifier oshun-production \
--db-instance-class db.r6g.large
# Update application to use read replica for reads
# Configure in connection string or ORM settings
Connection Pooling#
If hitting connection limits:
- Enable connection pooling (PgBouncer)
- Adjust application pool settings
- Increase
max_connections(requires restart)
# Check current connections
aws cloudwatch get-metric-statistics \
--namespace AWS/RDS \
--metric-name DatabaseConnections \
--dimensions Name=DBInstanceIdentifier,Value=oshun-production \
--start-time $(date -d '1 day ago' -u +%Y-%m-%dT%H:%M:%SZ) \
--end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
--period 300 \
--statistics Maximum Average
Redis/ElastiCache Scaling#
Scaling Options#
| Scaling Type | When to Use | Downtime |
|---|---|---|
| Vertical (resize) | Memory consistently high | Brief |
| Horizontal (shards) | High throughput needed | Depends |
| Read replicas | Read-heavy cache workload | No |
Check Current Utilization#
# Memory utilization
aws cloudwatch get-metric-statistics \
--namespace AWS/ElastiCache \
--metric-name DatabaseMemoryUsagePercentage \
--dimensions Name=CacheClusterId,Value=oshun-production \
--start-time $(date -d '7 days ago' -u +%Y-%m-%dT%H:%M:%SZ) \
--end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
--period 3600 \
--statistics Average Maximum
# CPU utilization
aws cloudwatch get-metric-statistics \
--namespace AWS/ElastiCache \
--metric-name CPUUtilization \
--dimensions Name=CacheClusterId,Value=oshun-production \
--start-time $(date -d '7 days ago' -u +%Y-%m-%dT%H:%M:%SZ) \
--end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
--period 3600 \
--statistics Average Maximum
Resize Cache Node#
# Modify node type
aws elasticache modify-cache-cluster \
--cache-cluster-id oshun-production \
--cache-node-type cache.r6g.large \
--apply-immediately
Scaling Checklist#
Before Scaling Up#
- Identify bottleneck (CPU, memory, connections, etc.)
- Check if issue is load-related or code bug
- Review recent deployments for potential issues
- Estimate cost impact of scaling
- Notify team in #ops channel
After Scaling#
- Verify metrics improved
- Monitor for 30 minutes
- Document scaling action
- Schedule review for potential permanent change
Cost Considerations#
| Action | Cost Impact | Reversibility |
|---|---|---|
| ECS scale up (auto) | Hourly cost increase | Auto |
| ECS scale up (manual) | Hourly cost increase | Manual |
| RunPod max workers increase | Per-second GPU cost | Manual |
| RDS instance resize | Hourly cost change | Manual |
| RDS read replica | Hourly cost add | Manual |
| ElastiCache resize | Hourly cost change | Manual |
Emergency Scaling Procedures#
Traffic Spike#
-
Immediately increase ECS desired count:
bashaws ecs update-service --cluster oshun-production --service oshun-api --desired-count 8 -
Increase auto scaling max:
bashaws application-autoscaling register-scalable-target \ --service-namespace ecs \ --resource-id service/oshun-production/oshun-api \ --scalable-dimension ecs:service:DesiredCount \ --max-capacity 20 -
If RunPod queues backing up, increase max workers
-
Monitor CloudWatch dashboard closely
-
Communicate status in #incidents
Resource Exhaustion#
- Scale horizontally (more instances) rather than vertically (larger instances)
- Enable request throttling if needed
- Consider circuit breaker patterns
- Shed non-critical load (e.g., analytics)