Multi-Node Deployment

By default Temps runs everything on a single server. Multi-node mode lets you add worker nodes — additional servers that receive container deployments from the control plane — so you can distribute workloads across multiple machines.


Overview

Multi-node is useful when you:

  • Outgrow a single server — you need more CPU, memory, or disk than one machine provides
  • Want geographic distribution — deploy containers closer to your users by placing workers in different regions
  • Need workload isolation — run production and staging on separate physical machines

When no worker nodes are registered, Temps behaves exactly as before — all containers run locally on the control plane server.


Architecture

Control plane (your Temps server)

  • Runs the API, dashboard, proxy, and database
  • Accepts temps join registrations from workers
  • Schedules container deployments across nodes
  • Routes traffic to containers via private addresses
  • Watches every node's heartbeat and capacity, and self-reports its own

Worker nodes

  • Run the temps agent HTTPS server on port 3100
  • Accept container lifecycle commands from the control plane
  • Send heartbeats every 30 seconds with capacity metrics
  • Have Docker installed for running containers

Traffic from the internet still enters through the control plane's reverse proxy (Pingora). The proxy routes requests to the correct container — whether it runs locally or on a remote worker — using the node's private address.

The control plane also appears in the node list as a node of its own (id 0, role control-plane), so containers scheduled locally are visible alongside worker nodes, and the control plane's own CPU/memory/disk are monitored with the same alert thresholds.


Prerequisites

On each worker node:

  • Docker installed and running
  • The temps binary available (same version as the control plane)
  • Network connectivity to the control plane (see direct vs WireGuard modes below)
  • Port 3100 accessible from the control plane (or a WireGuard tunnel)

On the control plane:

  • Temps server running with temps serve
  • The /api/internal/nodes/register endpoint reachable from workers
  • An enrollment token minted on the control plane (see Securing node enrollment). Enrollment tokens are short-lived and single-use by default — they replace the old shared join token.
  • Mutual TLS certificate material created during enrollment. Fresh installations require mTLS by default; upgraded clusters can keep legacy workers on HTTP until you re-enroll them.

Adding a worker node (direct mode)

Add a worker node (direct mode)

  1. 1

    On the control plane, open Settings > Worker Nodes and create a one-hour, single-use enrollment token. Copy the returned token (it is shown once; the control plane only stores a SHA-256 hash).

    Checkpoint: Save the token before leaving -- you cannot retrieve the plaintext again. It expires in 1 hour and is consumed by a single successful join.

  2. 2

    On the worker machine, download install.sh, review it, then run the local copy.

  3. 3

    Join the cluster with the control plane URL and the enrollment token, passing --private-address <worker-ip> (e.g. 10.0.0.5) to enable direct mode so traffic routes straight to the worker's private address.

  4. 4

    Pass the returned trust pin as --ca-fingerprint <fp>, including for the first worker. Token creation initializes the cluster CA before returning, so no enrollment needs trust on first use.

  5. 5

    Optionally add --name worker-eu-1 and --labels region=eu-west,gpu=true to label the node (the default agent listen address is 127.0.0.1:3100; pass --agent-address 0.0.0.0:3100 to listen on all interfaces).

  6. 6

    After joining, the agent config is saved to ~/.temps/agent.json (and the node cert/key + cluster CA if mTLS is on). Start the agent with temps agent.

    Checkpoint: Open Settings > Worker Nodes and confirm the new node appears with status active and a recent heartbeat (within the last 90 seconds).

Direct mode is the simplest option when nodes can reach each other over a private network (e.g., same VPC, same data center, or a VPN).

1. Mint an enrollment token:

On the control plane, open Settings → Worker Nodes and create a short-lived, single-use token (see Securing node enrollment for API automation).

The response contains the plaintext token once — copy and save it. The control plane only stores a SHA-256 hash.

2. Install Temps on the worker machine:

curl -fsSL https://temps.sh/install.sh -o install.sh
less install.sh
bash install.sh

3. Join the cluster:

temps join <control-plane-url> <enrollment-token> --private-address <worker-ip>
  • Name
    control-plane-url
    Type
    string
    Description

    The URL of your Temps control plane, e.g. https://temps.example.com.

  • Name
    enrollment-token
    Type
    string
    Description

    The enrollment token minted in step 1 (or set TEMPS_JOIN_TOKEN). The token is hashed (SHA-256) before storage — the plaintext is never persisted on the control plane — and is consumed by a single successful join.

  • Name
    --private-address
    Type
    string
    Description

    The IP address the control plane should use to reach this worker's containers. Typically a private/internal IP like 10.0.0.5 or 192.168.1.50.

The --private-address flag enables direct mode — no WireGuard tunnel is created. The control plane routes traffic directly to the worker's private address.

Optional flags:

temps join <url> <token> --private-address 10.0.0.5 \
  --name worker-eu-1 \
  --labels region=eu-west,gpu=true \
  --agent-address 0.0.0.0:3100 \
  --ca-fingerprint <cluster-ca-fingerprint>
  • --name — A friendly name for this node (defaults to the hostname)
  • --labels — Key-value pairs stored on the node (e.g., region=us-east,tier=production)
  • --agent-address — The address the agent server listens on (default: 127.0.0.1:3100)
  • --ca-fingerprint — Pin the cluster CA when mutual TLS is enabled; the join aborts on mismatch

After joining, the node is registered and the agent config is saved to ~/.temps/agent.json. Start the agent separately:

temps agent

The agent reads its config from ~/.temps/agent.json and begins sending heartbeats to the control plane every 30 seconds.


Adding a worker node (WireGuard mode)

Add a worker node (WireGuard mode)

  1. 1

    Mint an enrollment token on the control plane (POST /api/settings/enrollment-tokens) and copy it.

  2. 2

    On the worker, run temps join <cluster-id> <enrollment-token> --relay-url <your-relay-url>; this generates a WireGuard keypair, contacts your relay for key exchange, brings up the wg0 interface on a 10.100.0.x address, and registers using that address. There is no public relay at the moment, so without --relay-url the command refuses to start and points you to direct mode. WireGuard runs in-process (boringtun) -- no wireguard-tools package needed, just root or CAP_NET_ADMIN to create the wg0 interface.

  3. 3

    Start the agent with temps agent. The config was saved to ~/.temps/agent.json during join.

    Checkpoint: Confirm the node shows status active in Settings > Worker Nodes with its 10.100.0.x WireGuard address.

WireGuard mode creates an encrypted tunnel between the worker and the control plane. Use this when nodes are on different networks without direct connectivity. It needs a relay that brokers the WireGuard key exchange, and you point the worker at it with --relay-url (or TEMPS_RELAY_URL).

There is no public relay at the moment. Without --relay-url, temps join refuses to start and tells you to choose direct mode or supply a relay you run. Direct mode is the supported path for most setups.

temps join <cluster-id> <enrollment-token> --relay-url https://relay.example.com

With --relay-url, the join command:

  1. Generates a WireGuard keypair
  2. Contacts your relay for key exchange
  3. Sets up a WireGuard tunnel (wg0 interface) with a 10.100.0.x address
  4. Registers with the control plane using the WireGuard address
  5. Saves the agent config to ~/.temps/agent.json

Then start the agent:

temps agent

WireGuard runs in-process via an embedded userspace implementation (boringtun) — no wireguard-tools package, wg, or wg-quick binaries are required on either machine. The only requirement is permission to create the wg0 interface: run temps join as root or grant the binary CAP_NET_ADMIN.


Running the agent

Run the worker agent

  1. 1

    Make sure temps join has already run on this machine so ~/.temps/agent.json exists with the node ID, control plane URL, and bearer token.

  2. 2

    Run temps agent in the foreground and confirm that the overlay, DNS resolver, and HTTPS agent listener start without errors.

  3. 3

    Install the agent as a systemd service using the same binary and data directory used during enrollment.

  4. 4

    The agent sends a heartbeat every 30 seconds; if heartbeats stop the control plane marks the node offline after 90 seconds and excludes it from scheduling.

    Checkpoint: In Settings > Worker Nodes confirm the node's last heartbeat updates and status stays active while the agent is running.

After temps join completes, run the agent server:

sudo temps agent

The agent loads its configuration from ~/.temps/agent.json (saved by temps join). CLI flags and environment variables override the saved config:

Variable / FlagDefaultDescription
TEMPS_AGENT_ADDRESS / --listen-address127.0.0.1:3100Address the agent listens on
TEMPS_AGENT_TOKEN / --token—Bearer token for authenticating control plane requests
TEMPS_NODE_NAME / --node-name—Name for this node
TEMPS_CONTROL_PLANE_URL / --control-plane-url—URL of the control plane
TEMPS_NODE_ID / --node-id—Node ID assigned during registration

If no ~/.temps/agent.json exists and required fields are missing, the agent exits with an error suggesting you run temps join first.

Run the agent after logout and reboot

Once foreground startup succeeds, run the agent as a systemd service. The unit below uses the default paths created by a root installation. Replace the binary or data-directory path when your installation uses a different location.

sudo tee /etc/systemd/system/temps-agent.service >/dev/null <<'EOF'
[Unit]
Description=Temps Worker Agent
Requires=docker.service
After=network-online.target docker.service
Wants=network-online.target

[Service]
Type=simple
Environment=TEMPS_DATA_DIR=/root/.temps
ExecStart=/usr/local/bin/temps agent
Restart=on-failure
RestartSec=5

[Install]
WantedBy=multi-user.target
EOF

sudo systemctl daemon-reload
sudo systemctl enable --now temps-agent

Do not run a foreground agent and the systemd service at the same time. Both processes try to bind the same listener and manage the same overlay state.

Confirm that the service remains active and inspect its startup logs:

sudo systemctl status temps-agent --no-pager
sudo journalctl -u temps-agent -n 100 --no-pager

Follow logs while you test deployments or service-to-service connectivity:

sudo journalctl -u temps-agent -f

After replacing the binary during an upgrade, restart the service:

sudo systemctl restart temps-agent

The agent sends a heartbeat to the control plane every 30 seconds. If heartbeats stop (e.g., the agent is stopped or the machine goes down), the control plane marks the node as offline after 90 seconds and excludes it from scheduling.

When mutual TLS is enabled and the node holds a signed certificate, the agent serves HTTPS with client-certificate verification instead of plaintext — the control plane must present its cluster-CA-signed identity on every call.


Securing node enrollment

Joining a node is gated by an enrollment token. Tokens are short-lived and single-use by default, and only their SHA-256 hash is ever stored on the control plane — the plaintext is shown once at mint time and never again.

Mint a token

The dashboard is the recommended interactive path. For automation, authenticate the request with a bearer token belonging to an operator with settings-write permission:

POST /api/settings/enrollment-tokens

export TEMPS_ACCESS_TOKEN='YOUR_OPERATOR_ACCESS_TOKEN'
curl -X POST https://temps.example.com/api/settings/enrollment-tokens \
  -H "Authorization: Bearer ${TEMPS_ACCESS_TOKEN}" \
  -H 'Content-Type: application/json' \
  -d '{"ttl_secs": 3600, "max_uses": 1, "bound_node_name": "worker-eu-1"}'
unset TEMPS_ACCESS_TOKEN
FieldDefaultNotes
ttl_secs3600 (1 h)Lifetime in seconds. Server-capped at 86400 (24 h).
max_uses1How many successful joins the token allows. Range 1–100.
bound_node_name—Optional. Restricts the token to a single node name (exact match).

The response returns the plaintext token once, its expires_at, and the cluster ca_fingerprint to pin on join. Minting the first token creates the per-cluster CA transactionally, so the first worker has the same trust pin as every later worker.

List and revoke

export TEMPS_ACCESS_TOKEN='YOUR_OPERATOR_ACCESS_TOKEN'

# List active (non-revoked, non-expired) tokens, newest first
curl https://temps.example.com/api/settings/enrollment-tokens \
  -H "Authorization: Bearer ${TEMPS_ACCESS_TOKEN}"

# Revoke a token by id
curl -X DELETE https://temps.example.com/api/settings/enrollment-tokens/<id> \
  -H "Authorization: Bearer ${TEMPS_ACCESS_TOKEN}"

unset TEMPS_ACCESS_TOKEN

The old shared join token still works during an upgrade — it is gated by multi_node.legacy_shared_token_enabled, which defaults to true on both existing and fresh clusters. Once every worker has re-enrolled with a per-node token, set it to false so only enrollment tokens are accepted. Expired, revoked, or exhausted enrollment tokens are rejected outright (no silent fallback to the legacy token), which prevents downgrade attacks.


Mutual TLS (mTLS)

Fresh Temps installations require mutual TLS when the control plane calls a worker agent on port 3100. The worker verifies the control plane's client certificate, the control plane verifies the worker's server certificate, and then the agent evaluates the bearer token. Existing clusters that predate this setting keep their serialized require_mtls = false value so an upgrade does not disconnect workers before you re-enroll them.

How it works

  • The control plane mints a per-cluster CA when the first enrollment token is created. The creation is transactional, so concurrent operators cannot publish different roots. The CA private key is encrypted at rest (AES-256-GCM) and never leaves the control plane.
  • Each joining node generates its own key pair and certificate signing request locally. The node's private key never leaves the worker. The control plane signs a per-node leaf certificate and replaces any requested SANs with the registered node IP and name.
  • On success the node serves mutual TLS, and the control plane presents its own cluster-CA-signed client identity on every call. The node's address is automatically switched to https://.
  • The bearer token remains a second authorization check. A valid cluster certificate alone does not authorize an agent API operation.

The worker's outbound heartbeat and network-sync requests go to the control plane's public HTTPS endpoint. Those requests verify the public server certificate and use the node bearer token; they do not currently present the node certificate as a client identity. The node certificate protects the agent API that accepts deployment, execution, log, and lifecycle commands from the control plane.

  • Name
    Fresh cluster
    Type
    mTLS enforced
    Description

    New installations default to multi_node.require_mtls = true. A worker must use a current temps join command that sends a CSR; a legacy CSR-less join is rejected.

  • Name
    Existing cluster
    Type
    migration mode
    Description

    Clusters upgraded from a version without node certificates keep multi_node.require_mtls = false while workers are migrated. A current worker that re-enrolls with a CSR still receives and uses mTLS even during this compatibility window.

Restarting an old agent is not a migration: it reloads the old plaintext configuration. Certificates are created only by a new temps join enrollment.

Fresh cluster mTLS setup

Use this path when no workers have joined the control plane yet.

  1. Install the same Temps version on the control plane and worker.
  2. Configure and verify the private network before enrollment.
  3. In Settings → Worker Nodes, mint a one-hour, single-use enrollment token bound to the worker's intended name. Copy both the plaintext token and the returned CA fingerprint.
  4. Join the worker with its stable private address and an agent listener that the control plane can reach:
export TEMPS_JOIN_TOKEN='YOUR_ENROLLMENT_TOKEN'
temps join https://your-temps-instance.com \
  --name worker-eu-1 \
  --private-address 10.0.0.5 \
  --agent-address 0.0.0.0:3100 \
  --ca-fingerprint 'YOUR_CLUSTER_CA_FINGERPRINT'
unset TEMPS_JOIN_TOKEN

Token creation initializes the CA before the token is returned. Verify the fingerprint through a second authenticated channel before joining, then compare it with the CA written on the worker:

# Operator machine
bunx @temps-sdk/cli login https://your-temps-instance.com --context production
export TEMPS_CONTEXT=production
bunx @temps-sdk/cli settings show --json \
  | jq -r '.multi_node.cluster_ca_fingerprint'
unset TEMPS_CONTEXT

# Worker machine
openssl x509 \
  -in ~/.temps/cluster-ca.pem \
  -noout -fingerprint -sha256

Normalize case and remove : separators when comparing the two SHA-256 values. Pass the trusted value to every temps join, including the first one.

Start the first agent in the foreground and verify its secure listener before putting it under a service supervisor:

temps agent

The log must contain Agent serving with mutual TLS. The saved configuration must have require_mtls: true, non-empty certificate paths, and a listener that is reachable only over the private network.

Existing cluster mTLS migration

Use this path when one or more workers already have an agent.json with require_mtls: null or false and no certificate paths. Migrate one worker at a time. Keep the cluster in compatibility mode until every active worker has been converted; this avoids taking all workers offline at once.

1. Record the worker's existing identity

On the worker, save the exact node name and control-plane URL. Re-enrollment uses the saved bearer token as proof of continuity, so using the same name and URL updates the existing node instead of creating a duplicate.

jq '{node_name,control_plane_url,node_id,listen_address,require_mtls}' \
  ~/.temps/agent.json

Also record the stable private address and underlay interface used by the worker. Do not delete agent.json before re-enrollment.

2. Mint a token for this worker

In Settings → Worker Nodes, mint a short-lived, single-use token and bind it to the exact node_name from the previous step. The token response includes the active CA fingerprint; verify it through a second trusted channel before continuing.

3. Re-enroll the existing identity

export TEMPS_JOIN_TOKEN='YOUR_ENROLLMENT_TOKEN'
temps join 'EXACT_EXISTING_CONTROL_PLANE_URL' \
  --name 'EXACT_EXISTING_NODE_NAME' \
  --private-address 10.0.0.5 \
  --agent-address 0.0.0.0:3100 \
  --ca-fingerprint 'YOUR_CLUSTER_CA_FINGERPRINT'
unset TEMPS_JOIN_TOKEN

Do not omit --ca-fingerprint. If a token response has no fingerprint, stop: the control plane is running an older or unhealthy build and cannot provide a pinned first enrollment.

4. Restart only this worker agent

Stop the old plaintext agent process and start temps agent again so it loads the new certificate, private key, CA, and HTTPS listener. Do not restart the control plane or the other workers as part of this step.

5. Verify before moving to the next worker

jq '{
  node_name,
  node_id,
  listen_address,
  require_mtls,
  tls_cert_path,
  tls_key_path,
  cluster_ca_path
}' ~/.temps/agent.json

Confirm all of the following:

  • node_id is unchanged and no duplicate worker appeared in the dashboard.
  • require_mtls is true.
  • The three certificate paths are populated and readable only by the worker owner; the private key is mode 0600.
  • The worker log contains Agent serving with mutual TLS.
  • The control plane records the worker address with an https:// scheme.
  • A test deployment, log stream, and container lifecycle operation succeed on this worker.

If verification fails, keep the other workers in compatibility mode, stop this agent, restore the saved agent.json and PEM files, and restart the previous agent. Investigate that single worker before continuing.

Repeat the procedure for every active worker. Only after every worker has a verified certificate-bearing configuration should the cluster reject legacy CSR-less registrations. Existing workers already using HTTPS+mTLS remain secure during the migration window; the enforcement setting controls whether an old client is still allowed to enroll, not whether a modern re-enrolled worker uses mTLS.

Certificate lifetime and rotation

Temps does not automatically renew node certificates yet. The current cluster CA and node leaf certificates use a long-lived validity period, so an unattended node does not stop when a short certificate window closes. This avoids expiry outages, but it is not a substitute for rotation after suspected key exposure or an ownership change.

Rotate a worker certificate by repeating the re-enrollment procedure above with a new single-use token. Re-enrollment generates a new private key locally, obtains a newly signed leaf certificate, replaces the PEM files in the Temps data directory, and updates the node's certificate fingerprint. Restart the agent after re-enrollment so it loads the new key and certificate.

Back up the control-plane encryption key with the control-plane database. The database stores the cluster CA private key encrypted with that key. If you restore only the database without its encryption key, the control plane cannot create its mTLS client identity or sign replacement worker certificates.

The cluster CA fingerprint is public and safe to verify through a separate trusted channel. Never copy node.key.pem or the encrypted cluster CA key between machines.

Recovering from a compromised cluster CA

A leaked worker private key does not require a cluster-wide root change. Revoke that worker's access, mint a node-bound single-use token, and re-enroll only that worker to replace its key and leaf certificate.

A leaked cluster CA private key is different: an attacker can mint trusted certificates. Treat it as an incident. The current trust model intentionally uses one active root, so emergency replacement is fail-closed and requires a maintenance window. There is no safe way to make existing workers trust a new root over a channel authenticated only by the compromised root.

1. Contain the incident

  • Stop worker agents or block control-plane-to-agent traffic on the private network. Running application containers can keep serving, but deployments and lifecycle operations pause.
  • Preserve the control-plane database, encryption key, audit logs, and the fingerprint you observed. Do not delete CA fields or edit settings directly.
  • Rotate any operator or node bearer tokens that may also have been exposed.

2. Read the current fingerprint through an authenticated channel

bunx @temps-sdk/cli login https://your-temps-instance.com --context production
export TEMPS_CONTEXT=production
bunx @temps-sdk/cli settings show --json \
  | jq -r '.multi_node.cluster_ca_fingerprint'
unset TEMPS_CONTEXT

3. Atomically replace the root

Sign in to the Temps console as a full Admin, open Settings → Worker Nodes → Cluster trust, and choose Rotate cluster CA. Rotation has its own cluster_ca:rotate permission; settings:write, Platform Admin, and project roles are not sufficient. Temps also requires an MFA-enrolled, persisted browser session and fresh MFA step-up. API keys, CLI tokens, and deployment tokens are rejected even if somebody explicitly assigns the rotation permission to them.

Enter the fingerprint read in the previous step and the exact destructive confirmation shown by the console. The operation replaces the certificate and encrypted key together and revokes every outstanding enrollment token. A stale fingerprint returns 409 Conflict instead of overwriting a newer root.

Record the new_fingerprint from the response through your incident channel. Old worker certificates stop authenticating immediately; this is the intended containment boundary.

4. Re-enroll every worker

For each worker, mint a fresh node-bound token, verify that its ca_fingerprint equals the recorded new value, and repeat the enrollment steps above. Re-enrollment creates a new worker private key locally. Restart that worker agent only after checking the saved CA fingerprint and certificate paths.

5. Prove recovery before reopening the cluster

  • Confirm every expected node ID is active and no unknown node appeared.
  • Run a deployment, log stream, and lifecycle operation on every worker.
  • Confirm old enrollment tokens are rejected and the audit log contains CLUSTER_CA_ROTATED with both fingerprints.
  • Remove temporary network blocks only after all workers use the new root.

Monitoring & alerts

The control plane watches every node automatically and delivers alerts through your existing notification providers (email, Slack, webhook) — there are no per-node rules to configure.

  • Node offline (critical). If a node misses heartbeats for 90 seconds, it is marked offline, its workloads are failed over to healthy nodes, and a critical alert fires once per outage.
  • Node recovered (info). When an offline node starts heart­beating again, an info notification is sent.
  • Resource pressure (warning). When a node's CPU, memory, or disk usage rises above a threshold, a warning alert fires. The control plane itself (node 0) is checked with the same thresholds.

Thresholds live under multi_node and are operator-configurable with PATCH /api/settings:

SettingDefaultNotes
multi_node.node_cpu_alert_percent90Set to null to disable CPU alerts
multi_node.node_memory_alert_percent90Set to null to disable memory alerts
multi_node.node_disk_alert_percent90Root mount (/); set to null to disable

A breach is strictly greater than the threshold — a node sitting at exactly 90% does not alert. The 90-second offline threshold is fixed. Notification delivery is best-effort: a provider failure is logged but never stalls health checks or failover.


Verifying nodes

From the dashboard:

Go to Settings > Worker Nodes to see all registered nodes, their status, and last heartbeat time.

From the API:

List all nodes

curl https://your-temps-instance/api/internal/nodes

Response

{
  "nodes": [
    {
      "id": 0,
      "name": "control-plane",
      "address": "local",
      "private_address": "127.0.0.1",
      "role": "control-plane",
      "status": "active",
      "capacity": { "cpu_percent": 5.5 },
      "last_heartbeat": "2026-03-04T12:00:00Z"
    },
    {
      "id": 1,
      "name": "worker-eu-1",
      "address": "https://10.0.0.5:3100",
      "private_address": "10.0.0.5",
      "role": "worker",
      "status": "active",
      "labels": { "region": "eu-west" },
      "last_heartbeat": "2026-03-04T12:00:00Z"
    }
  ],
  "total": 2
}

Get a specific node

curl https://your-temps-instance/api/internal/nodes/1

The control plane is always present as node 0 with role control-plane and self-reported capacity. Worker nodes start at id 1. A worker is considered active if it has sent a heartbeat within the last 90 seconds; workers that miss heartbeats are automatically marked offline and excluded from scheduling.


Node identity in containers

Every deployed container — on a worker or locally on the control plane — receives three environment variables so your app can report which node and replica is serving a request:

VariableExampleMeaning
TEMPS_NODE_NAMEworker-eu-1 (or control-plane)Name of the node running this container
TEMPS_NODE_ID1 (or 0 for the control plane)Numeric node ID, as a string
TEMPS_REPLICA1, 2, 3, …1-based replica index within the deployment
// Example: surface the serving node in a health endpoint
app.get('/whoami', (req, res) =>
  res.json({
    node: process.env.TEMPS_NODE_NAME,
    nodeId: process.env.TEMPS_NODE_ID,
    replica: process.env.TEMPS_REPLICA,
  })
)

Remote-node logs

Container logs from remote worker nodes are streamed back to the control plane and land in the same searchable history as local logs — both the live tail and the history viewer now include remote containers, each line tagged with the node it ran on.

A control-plane collector keeps a stream open to each remote container over the agent channel (encrypted with mTLS when enabled) and writes the lines into the shared log store. In the log viewer you can filter by Container and by Node (both default to "all", interleaved across every container) and a per-line source column shows which container and worker node each line came from.

Logs & debugging

How scheduling works

When you deploy an application with multiple replicas, the node scheduler decides which node each replica lands on:

  1. It starts from every node with status = "active" and a heartbeat within the last 90 seconds, plus the control plane itself (node 0). If no active worker nodes exist, all replicas run locally on the control plane — unless it was started with temps serve --profile control-plane (a control plane that runs no workloads), in which case the deployment fails with an explanation instead.
  2. Target nodes and label selectors, if set in the environment's or project's deployment settings, narrow the pool first. A label selector matches a node only if every key matches (a list of values matches any of them). The control plane has no labels, so a label selector never places replicas on it; target node 0 selects it explicitly. If these constraints leave no eligible node, the deployment fails with an error instead of ignoring them.
  3. Architecture compatibility narrows it further: a node is only eligible if this deployment has a built image for the platform its agent last reported (linux/amd64, linux/arm64, ...) — otherwise the container would pull the image, start, and immediately fail with exec format error. A node that hasn't reported its architecture yet (e.g. mid-upgrade) is kept in the pool with a warning instead of being dropped. Either way, the deploy log says so, e.g. Skipping node 'worker-2' runs linux/arm64, and this deployment only has images for [linux/amd64]. This filter needs to know which platforms the image exists for, either from a cross-architecture build or by inspecting the image on the control plane. When it can't tell (for example, the image isn't in the control plane's local image store, or the control plane has no local Docker daemon), the filter is skipped: every node stays eligible, and the deploy log doesn't mention it.
  4. From what's left, replicas are placed with least-loaded scheduling by default — each replica goes to the node with the lowest load score (60% CPU, 40% memory usage, from the node's last heartbeat), falling back to round-robin when no node reports capacity data. The control plane doesn't report capacity, so it (like any node without capacity data) counts as 50% loaded. Nodes at or above a 90% load score are skipped unless every remaining node is over that threshold, in which case the threshold is relaxed rather than failing the deployment.
  5. Anti-affinity (on by default) tries to give each replica its own node. When there are simply fewer nodes than replicas, replicas wrap around and share nodes. But if nodes were excluded for a lasting reason — for example, an architecture mismatch — and that leaves too few eligible nodes, the deployment fails with an error naming the excluded nodes instead of silently stacking replicas on the same node.
  6. Each container records which node it was assigned to in the database.

The deployment job creates a RemoteNodeDeployer for each remote assignment, which communicates with the worker's agent API to deploy, stop, and manage containers.


Current limitations

Multi-node is functional but still evolving. Known limitations:

  • No bin-packing — placement is least-loaded plus anti-affinity, not full bin-packing across resource dimensions.

What to explore next

Scaling strategies Resource allocation Logs & debugging

Last updated

Was this page helpful?