Scaling
Temps runs your application in Docker containers on your own server. Scaling means giving those containers more resources or running more of them. This guide covers when and how to scale, and what to optimize before adding hardware.
Looking for the step-by-step dashboard and API walkthrough for changing CPU, memory, and replica count? See Scale for Traffic. This page focuses on the strategy and trade-offs behind those settings.
Scaling options
Vertical scaling (bigger server)
- More CPU cores and RAM
- Simplest change — upgrade your VPS
- No application changes needed
- Best for database-heavy workloads
Horizontal scaling (more replicas)
- Multiple container instances per environment
- Built-in to Temps via replica configuration
- Requires stateless application design
- Best for CPU-bound web traffic
Multi-node scaling — if you need more capacity than a single server provides, you can add worker nodes to distribute containers across multiple machines; see that guide for how the scheduler picks which node runs a given replica.
Before scaling, check whether optimization can solve the problem for free. A slow database query or missing index can make an application feel overloaded when the server has plenty of capacity.
When to scale
Signs your application needs more resources:
| Symptom | Likely bottleneck | First action |
|---|---|---|
| Response times increasing gradually | CPU or memory pressure | Check resource usage in dashboard |
| Occasional timeouts under load | Too few connections or CPU saturation | Add replicas or optimize queries |
| Out-of-memory errors | Container memory limit | Increase memory limit or optimize memory usage |
| Build times increasing | Build-time CPU/memory | Upgrade server for faster builds |
| Database queries slow | Database, not app server | Optimize queries, add indexes, connection pooling |
Replicas
Temps supports running multiple container instances (replicas) for each environment. Traffic is distributed across all healthy replicas.
Configuring replicas
Set the replica count at the project or environment level:
Project-level (applies to all environments by default):
- Go to your project > Settings > Deployment Config
- Set Replicas to the desired number
Environment-level (overrides project default):
- Go to your project > Environments > select environment > Settings
- Set Replicas to the desired number
The environment setting takes priority over the project setting.
How replicas work
When you deploy with multiple replicas:
- Temps creates N containers (e.g.
myapp-1,myapp-2,myapp-3) - Each container gets its own host port
- All containers run the same image with the same environment variables
- The proxy distributes incoming requests across all healthy replicas
- Health checks run independently per replica
- If any replica fails during deployment, all replicas are cleaned up and the deployment fails
Requirements for replicas
Your application must be stateless — it cannot rely on local files, in-memory sessions, or container-local state that is not shared:
- Sessions: Store in Redis or a database, not in memory
- File uploads: Store in S3 or blob storage, not the local filesystem
- Cache: Use Redis or an external cache, not in-process memory
- WebSocket connections: Each client connects to one replica — use a pub/sub system (Redis) if clients need to communicate across replicas
Stateless session example
import session from 'express-session';
import RedisStore from 'connect-redis';
import { createClient } from 'redis';
const redisClient = createClient({ url: process.env.REDIS_URL });
await redisClient.connect();
app.use(session({
store: new RedisStore({ client: redisClient }),
secret: process.env.SESSION_SECRET,
resave: false,
saveUninitialized: false,
}));
Resource limits
Each container can be given CPU and memory limits. New projects start with a conservative default profile rather than being fully uncapped — a hard memory limit, so a runaway app can't OOM a small single-node host, plus recorded CPU and memory request values. You can raise, lower, or uncap these per environment at any time.
| Setting | Default | Description |
|---|---|---|
| CPU request | 0.5 CPU | Recorded, but not currently enforced (see below) |
| CPU limit | Unset (uncapped) | Maximum CPU, if configured |
| Memory request | 128 MB | Recorded, but not currently enforced (see below) |
| Memory limit | 512 MB | Maximum memory |
Only the limits are applied to the container. The request values are stored and shown in the dashboard and CLI, but Temps doesn't currently pass them to Docker as reservations or use them when choosing a node, so they don't guarantee any minimum CPU or memory. On a shared host, other workloads can still contend for those resources.
Set these per environment from the dashboard (Project → Environments → select environment → Settings), which takes plain CPU core counts (e.g. 0.5, 1, 2) and converts them correctly.
The CLI's environments resources command has a known unit bug in its --cpu/--cpu-request flags: they're documented and displayed as millicores (1000 = 1 core), but the value is passed straight through to a field the API stores in microcores (1_000_000 = 1 core), with no conversion. Setting --cpu 1000 therefore asks for a 0.001-core limit — too small for Docker to apply, so the container fails to start and the deployment fails. Until this is fixed, pass the raw microcore value instead of the number the flag's help text suggests:
# Sets a 1 CPU core limit (1_000_000 microcores) and a 0.5 CPU core request (500_000 microcores)
bunx @temps-sdk/cli environments resources production -p my-project \
--cpu 1000000 --cpu-request 500000 --memory 1024 --memory-request 128
- Name
--cpu- Type
- microcores, despite the flag's own help text saying millicores
- Description
Maximum CPU allocation.
1_000_000= 1 full CPU core,500_000= half a core.
- Name
--memory- Type
- MB
- Description
Maximum memory allocation in megabytes, e.g.
512.
- Name
--cpu-request- Type
- microcores, despite the flag's own help text saying millicores
- Description
CPU request. Recorded only — not currently enforced or used for scheduling.
- Name
--memory-request- Type
- MB
- Description
Memory request. Recorded only — not currently enforced or used for scheduling.
If you do set a memory limit and your container exceeds it, Docker kills it and Temps restarts it (restart policy is always). If this happens repeatedly, raise the memory limit or investigate memory leaks. A CPU limit is not set by default, so a CPU-bound container can use all available cores unless you cap it explicitly.
Vertical scaling
Upgrade your VPS to get more CPU and RAM. This gives more resources to all containers.
Steps
- Check current usage — Look at CPU and memory utilization in the Temps dashboard or via
htopon the server - Upgrade the VPS — Use your provider's upgrade option (DigitalOcean, Hetzner, Linode, AWS, etc.)
- Temps continues running — Most VPS upgrades preserve the disk. Temps and your containers restart automatically after the server reboots
Recommended server sizes
| Traffic level | Server spec | Monthly cost (approx) |
|---|---|---|
| Hobby / side project | 2 CPU, 4 GB RAM | $5-10 |
| Small production app | 4 CPU, 8 GB RAM | $20-40 |
| Medium traffic | 8 CPU, 16 GB RAM | $40-80 |
| High traffic | 16+ CPU, 32+ GB RAM | $80-160+ |
These are rough guidelines. Actual needs depend on your application — a database-heavy app needs more RAM; a compute-heavy app needs more CPU.
Optimize before scaling
Optimization is free and often more effective than adding resources:
Database queries
The most common performance bottleneck. Check for:
- N+1 queries — Loading related data in a loop instead of a JOIN
- Missing indexes — Add indexes on columns used in WHERE, JOIN, and ORDER BY clauses
- Expensive queries — Use
EXPLAIN ANALYZEto find slow queries - Connection pooling — Set appropriate pool sizes (20-50 connections per container)
Caching
Add caching for data that does not change on every request:
import { createClient } from 'redis';
const redis = createClient({ url: process.env.REDIS_URL });
async function getUsers() {
const cached = await redis.get('users');
if (cached) return JSON.parse(cached);
const users = await db.query('SELECT * FROM users');
await redis.set('users', JSON.stringify(users), { EX: 300 }); // 5 min cache
return users;
}
Frontend performance
For static sites and SPAs deployed on Temps:
- Temps already applies gzip compression, cache headers, and ETags automatically
- Use code splitting to reduce initial bundle size
- Lazy load routes and heavy components
- Optimize images (use modern formats like WebP/AVIF)
Application profiling
Profile your application to find hotspots before guessing:
- Node.js:
node --inspectorclinic.js - Python:
cProfileorpy-spy - Go:
pprof - Rust:
flamegraph
Monitoring
Track these metrics to know when scaling is needed:
- Response time — Increasing trend means the server is under pressure
- CPU usage — Sustained >80% means CPU-bound; add replicas or upgrade
- Memory usage — Approaching the limit means containers may get killed
- Error rate — Spikes correlate with resource exhaustion
Temps includes built-in analytics and monitoring. For server-level metrics, use standard tools:
# Real-time resource usage
htop
# Container-level stats
docker stats
# Disk usage
df -h
For application-level monitoring, add OpenTelemetry tracing — Temps injects OTEL_* environment variables automatically. See the observability tutorial for setup instructions.
Troubleshooting
Container keeps restarting
The container is hitting its memory limit and being killed by Docker. Check the container logs for OOM (out of memory) errors. Increase the memory limit or fix memory leaks in your application.
Adding replicas does not help
Not all performance problems are solved by more replicas:
- Database bottleneck — All replicas share the same database. More replicas means more database connections but the same query performance.
- External API rate limits — More replicas make more requests to external APIs. You may hit rate limits faster.
- Single-threaded applications — Node.js runs on a single thread. Each replica uses one CPU core. If your server has 4 cores, 4 replicas fully utilize the CPU. Beyond that, you need a bigger server.
Deployment is slow
Build time is separate from runtime performance. If builds are slow:
- Use multi-stage Docker builds to cache dependency installation
- Use
.dockerignoreto reduce build context size - Upgrade the server for faster builds (builds use half of available CPU and memory)