Cluster multi-process
VextJS uses one Master to manage multiple HTTP Workers. Each Worker has its own application instance and listens on the same service port. The Master starts and replaces Workers, checks heartbeats, and coordinates rolling reloads.
First start two Workers and verify requests; then tune the count and recovery policy. For development reloads see Hot Reload; for the production build prerequisite see Build.
Quick Start
Enable in configuration
Use an installed TypeScript API project, such as the CLI scaffold. Merge these settings into your existing production config while retaining business settings:
Add a diagnostic route:
Stop the same project's development server, then run from the project root:
Expect two ready Workers and a final workers=2/2 summary. Failure of the first Worker aborts startup; later failures may leave fewer Workers ready than requested. Check the actual ready count.
Startup and request verification
From another terminal, open separate connections:
The requests should return HTTP 200. data.pid and data.workerId identify the Worker serving each request. Confirm both ready Workers in the detailed startup logs; scheduling does not guarantee that two requests alternate between Workers.
Inspect .vext.pid in the project root. It records the Master PID, not an HTTP Worker PID. The status command has the limits described below. After verification, stop the foreground process with Ctrl+C; on Unix/macOS you may also run npx vextjs stop in another terminal. Verify the process, port, and PID file afterward. Keep or remove the diagnostic route as appropriate.
Enable through an environment variable
VEXT_CLUSTER=1 also enables Cluster. Configuration still controls the Worker count. Setting it to 0 does not disable an already configured cluster.enabled: true.
The Master loads configuration and checks the port before starting Workers. It passes the current bootstrap config provider patch to the Workers for reuse. Each Worker still initializes its own application; plugin side effects do not thereby run only once.
Architecture Overview
- Master process: does not process HTTP requests and is responsible for managing the life cycle of the Worker process
- Worker process: Each Worker runs a complete VextJS application instance and handles HTTP requests independently
- IPC communication: Messages are exchanged between Master and Worker through Node.js’ built-in inter-process communication (IPC)
Configuration options
Configure Cluster in the selected profile, such as src/config/production.ts. This example lists current defaults while enabling Cluster; usually you only need to override settings you intend to change:
Worker quantity strategy
Use a valid positive integer rather than relying on out-of-range fallback. CPU detection first tries os.availableParallelism(). Only when it is unavailable, throws, or returns an invalid value does the fallback start from os.cpus() and try Linux cgroup v1 quotas. A valid quota constrains the detected count. When files are missing, detection cannot read cgroup limits, contents are invalid, or no finite quota is set, the fallback may reflect the host CPU count. Automatic detection does not necessarily match every container CPU quota exactly.
In containers, first use this page's worker-info route to check the actual process count, then set an explicit value such as workers: 2 according to CPU, memory, and connection budgets. "auto-1" merely subtracts one from the detected count; it cannot repair failed quota detection.
Each Worker adds an application instance, database pool, cache, and heap. Choose a count based on request load, memory, and external connection limits, then verify under load; CPU multiples alone do not guarantee throughput.
State and multi-process boundaries
Workers do not share ordinary variables, Service instances, or in-memory Stores. Use shared storage for data that must be globally consistent. For example, an in-memory rate limit counts separately in each Worker and is not a global quota; see Rate Limiting. Sessions and caches likewise depend on their configured Stores.
Source IP affinity
Set cluster.sticky: "ip" when independent connections from the same TCP source IP must reach the same Worker. Master accepts and pauses each connection, chooses a stable logical slot using the source address, then hands the socket to its Worker over IPC. The Worker handles application traffic directly. All five built-in adapters support this through string configuration and official factories.
The default "none" keeps ordinary Cluster listening: Node SCHED_NONE on Linux, SCHED_RR on other platforms. "ip" directly implements affinity, replacing the previous round-robin placeholder; no new configuration or migration switch is required. It activates only during Cluster startup; dev/testing and ordinary single-process startup keep their existing paths.
Routing uses the TCP peer address, without inspecting X-Forwarded-For, Forwarded, cookies, or SIDs. Application trustProxy and req.ip parsing cannot change connection routing that has already happened. Proxies and NAT can merge many clients into one source IP and concentrate traffic on one Worker. External load balancers generally choose instances or separate upstreams, rather than Workers behind a shared port.
With stable slots and availability, new connections from one IP select the same slot. When a slot fails, only its traffic falls back; restoration can return it to its preferred slot. Rolling replacement preserves slot IDs, but does not preserve process memory. Old connections remain with the old Worker until completion or the shutdown budget expires. Changing sticky or workers requires a full Master restart; rolling reload rejects candidates whose policy differs from Master.
When a Worker enters application shutdown through SIGTERM/SIGINT or a fatal error, it notifies the Master to withdraw its routing eligibility. New connections can select remaining available slots, while in-flight connections retain the existing shutdown budget. Local exits still follow autoRestart and the restart budget to recover the same slot; Master-initiated replacements or shutdown do not create duplicate Workers.
Handoffs have bounded credits and deadlines: currently at most 1,024 pending globally and 64 per Worker, with a 5-second handoff deadline. These internal safeguards are separate from HTTP request execution timeouts. A saturated target slot rejects new connections rather than spilling them into another Worker. An unconfirmed, timed-out Worker is terminated and handled by the existing auto-restart policy. The framework does not replay requests; clients must consider idempotency when retrying. Runtime snapshot summary.connections records pending, committed, rejected, timedOut, sendFailed, and backpressure at Worker metrics intervals and lifecycle events. Committed means that the commit instruction was sent, rather than that an HTTP request succeeded.
IP affinity can serve applications with process-local state across multiple connections, but does not provide highly available sessions or automatically add Socket.IO, TLS, or HTTP/2. Shared storage remains preferable for consistency and recovery. Custom adapters must explicitly support socket handoff; see Adapters.
CLI commands
VextJS CLI provides complete Cluster management commands:
vext start — start
If cluster.enabled: true or VEXT_CLUSTER=1 is set in the configuration, vext start will automatically start in Cluster mode.
vext stop — stop
The command reads the PID file, sends SIGTERM, and waits up to 30 seconds for the Master to exit. Normal Master shutdown asks Workers to stop accepting requests, waits for work and cleanup, then removes the PID file on exit. A nonzero timeout does not prove the process has stopped.
On Windows, externally terminating a process is not equivalent to Unix signal-driven cleanup. Ctrl+C in a foreground vext start lets the CLI request shutdown through parent-child IPC; a separate vext stop or operating-system force termination cannot guarantee onClose runs. See Graceful shutdown for the timeout layers.
vext reload — rolling restart
The CLI returns after sending SIGHUP; successful delivery does not mean all Workers have been replaced. For each old Worker recorded at startup, the Master:
- Starts a new Worker and waits for ready.
- Asks the old Worker to shut down only after its replacement is ready.
- Waits for the old Worker to exit, forcing termination after timeout.
- Waits
workerDelay, then processes the next pair.
If a new Worker fails to start, the old Worker remains and failure is recorded while other replacements continue. Inspect replaced/total in the logs and make real requests; a complete message alone does not prove every replacement succeeded. Long connections, shutdown timeout, application errors, and insufficient resources can still interrupt traffic. Rolling replacement has no universal zero-downtime guarantee.
Reload does not recreate the Master. Worker count, Master heartbeat/backoff settings, and the provider patch captured at startup do not all refresh on a signal. Restart the complete service and shift traffic according to your deployment process when these settings change.
Windows does not support this reload signal operation; the command fails. Build valid TypeScript artifacts before deploying an update. Omitting cluster.reload does not disable rolling restart; the framework uses the default waits. Update code, config, and artifacts so old and new Workers can each read a consistent version.
vext status — View status
The command normally shows the Master PID and PID file, then probes http://<host>:<port>/health (default host 127.0.0.1, port 3000):
It appends PID, uptime, and memory details only when the health response contains those fields at the top level. It does not unpack the usual Vext response data, read the configured application port, or scan all Workers. The application must provide /health; a route at /api/health does not satisfy this fixed probe.
For example, if the target health route directly returns those top-level fields, one response could produce:
These numbers illustrate one health request, not all Workers. The default wrapped response in this page's demo does not automatically provide them. status may exit 0 even for not running, stale, or unreachable results; do not use its exit code as a deployment health gate. Pass the actual --host, --port, and --pid-file, and independently check business health.
Automatic failure recovery
Worker crashes and restarts
With autoRestart: true, an unintentional Worker exit usually triggers replacement. Candidate failures before ready are handled by their startup or replacement flow. The new Worker gets a new ID and PID; the old process is not revived in place. These log values are illustrative:
Worker memory, unfinished requests, and unpersisted state do not return automatically with the replacement.
Exponential backoff
When crashes occur continuously, the restart delay gradually increases (exponential backoff) to avoid frequent restarts consuming system resources:
Crash loop protection
restartWindow defaults to 60,000 ms and maxRestarts to 5. The whole Master shares this budget; abnormal exits from different Workers consume it together. The sixth attempted restart within the window is paused:
The end of the restart window does not itself schedule capacity restoration. When no Workers or pending replacement tasks remain, the standard bootstrap/CLI host records fatal capacity loss, removes its owned PID file, and exits with code 1 so an outer supervisor can restart it. With healthy Workers remaining, the Master stays up; check ready capacity and business requests. Direct users of the low-level ClusterMaster still receive all-workers-dead and choose their own shutdown policy.
Heartbeat detection
Workers send heartbeat messages spontaneously every 10 seconds by default. With healthCheck.enabled: true, the Master inspects ready Workers' last heartbeat every interval (default 15 seconds). After timeout (default 30 seconds), it forcibly terminates a timed-out Worker. Replacement still depends on autoRestart and the restart budget. This is not an HTTP /health request or an IPC health-check sent every 15 seconds. Interval-based inspection does not guarantee detection at exactly 30 seconds.
Memory threshold
Every 60 seconds, a Worker checks its heap used against cluster.memoryThreshold (default 1 GiB, in bytes). On crossing it, that Worker sends a single request-restart asking the Master to start a replacement first. It does not exit immediately, and this is not an RSS or container-memory hard limit. If replacement fails, do not assume memory was freed; inspect logs and live processes.
PID file
When started in Cluster mode, the Master process will write to the PID file (default .vext.pid), which is used for vext stop / vext reload / vext status commands to locate the process.
PID files are automatically managed at the following times:
- Create: when Master starts
- Delete: When Master exits normally
- Detection: Detect whether there is a running Cluster at startup
Relative paths resolve from the startup working directory. Stop, reload, and status do not automatically read this path from application config; pass the same --pid-file. Forced termination may leave a stale file. Check the process behind its PID rather than deleting the file as a substitute for stopping the service.
Add .vext.pid to .gitignore to avoid committing to version control.
Cooperation with graceful closing
Graceful shutdown process in Cluster mode:
There are separate timeout layers with different units:
When the Master's wait expires, it may SIGKILL a Worker, so application cleanup cannot be assumed to continue. Allow more time outside than inside the application, for example:
See Hooks for shutdown ordering. A forced termination or timeout is not proof that cleanup hooks finished.
Configure according to environment
It is recommended to use vext dev (hot reload mode) instead of Cluster mode for development environment. Cluster is mainly used for multi-core utilization and high availability in production environments.
Inter-process communication
Master and Worker communicate through an internal IPC protocol. The table helps explain runtime behavior; it is not a stable package-root API for application plugins. Applications must not manually send ready to bypass real initialization.
When launched through vext start, the Windows CLI sends its shutdown request to the Master over the parent-child IPC channel. The Master uses the same graceful shutdown sequence to notify Workers, wait for exit, and remove its PID file. Terminating a process directly through the operating system does not guarantee that close hooks run.
The message types below are the exact string literals of the IPC payload type field, without direction prefixes.
Worker → Master message
Master → Worker message
These messages are maintained by the framework. Placeholder request counts do not mean there was no traffic; production monitoring needs an actual request metrics provider.
When diagnosing internal process communication, message shapes include:
The PID and workerId are illustrative values; timeout is in milliseconds. Manage Workers through the CLI and lifecycle APIs on this page rather than sending internal messages manually from routes.
Deploying with Docker
Dockerfile example
This runtime image assumes a TypeScript API project has already built dist, all config needed by start is in the output, and no other runtime resources are required. See Build for a complete multi-stage build. JavaScript source mode must also carry src.
Suggestions
- Keep identity files inside
distand production dependencies. Carry custom frontend directories, workspace packages, and external resources separately. - Check Worker count against actual container CPUs, memory, and connection quotas; choose an explicit number when needed.
- The PID path must be writable and unique to each instance.
- The container stop grace period must exceed the Master wait and application cleanup time.
- With an outer process manager or multiple container replicas, count total Workers to avoid unintended double scaling.
Jobs and Cluster
HTTP Cluster Workers register scheduled timers after readiness and coordinate through config.jobs.redis. Active jobs without Redis fail at startup; empty directories or entirely disabled jobs do not require Redis. Namespace is generated automatically from package name, profile and runtime mode; matching replicas need no manual value. All replicas share the same Redis target and schedule definition. Each point is admitted once and same-job overlap is skipped; business side effects still require idempotency. See Jobs.
FAQ
What should we pay attention to when using WebSocket/SSE in Cluster mode?
An established long connection stays with the Worker holding it and can disconnect when that Worker exits. Design reconnection, cross-request state, broadcast, and shutdown timeouts explicitly. sticky: "ip" preserves source IP reconnection affinity while slots and availability are stable; replacement, failure, or address changes cannot preserve the original process or its memory. Protocol support also depends on the adapter and application.
What is the appropriate number of Workers?
Start with a controlled count, then adjust using CPU, memory, response latency, and total database connections. Every Worker initializes its own application and pools. "auto-1" only reduces process count; it does not reserve or pin a physical core for the Master.
How to monitor the status of each Worker?
Use vext status to view the Master PID, PID file state, and single health endpoint details when /health is reachable. The current command does not print a worker table or per-worker request counts; production environments should use Prometheus or another monitoring system for richer multi-worker metrics.
How is it different from PM2?
The built-in Master manages this framework's Workers. If an external process manager also supervises the app, decide whether it manages one Master or several independent application instances. Avoid two Cluster layers conflicting over Worker count, PID files, and shutdown. The outer supervisor still owns recovery when the Master itself exits.
Next step
- Learn about the detailed explanation of Cluster-related commands in CLI Commands
- View the complete configuration items of Cluster in Configuration
- Learn the relationship between hot reload and Cluster
- Explore Cluster-related testing methods in Testing