Graceful Shutdown in Node.js: Handling SIGTERM Without Dropping Requests
Every deploy and restart sends your Node.js app SIGTERM. How to shut down gracefully: stop accepting traffic, fail readiness, finish in-flight requests, close keep-alive connections, drain workers, close database pools — plus the Docker PID 1 and npm start signal traps.
Every deploy, scale-down, restart and container replacement ends the same way for your process: it receives a SIGTERM signal, and some seconds later — if it hasn't exited — a SIGKILL it can't catch. What your app does in between decides whether users see failed requests, whether jobs are half-done, and whether database connections are left dangling.
By default, Node.js does nothing special: on SIGTERM it simply exits, killing every in-flight request mid-response.
What a graceful shutdown does
- Stop receiving new traffic — tell the load balancer you're going away, and stop accepting new connections.
- Finish in-flight work — let active requests complete.
- Close idle connections, including HTTP keep-alive sockets.
- Stop background work — let the current job finish, don't take new ones.
- Release resources — close database pools, flush logs and metrics.
- Exit — before the platform's deadline, with a forced exit as a backstop.
A working implementation
import http from "node:http";
import { app } from "./app.js";
import { pool } from "./db.js";
const server = http.createServer(app);
server.listen(process.env.PORT ?? 3000);
let shuttingDown = false;
// Readiness endpoint used by the load balancer
app.get("/health/ready", (_req, res) => {
res.status(shuttingDown ? 503 : 200).end();
});
async function shutdown(signal) {
if (shuttingDown) return;
shuttingDown = true;
console.log(JSON.stringify({ msg: "shutdown started", signal }));
// Backstop: never outlive the platform's grace period
const forceExit = setTimeout(() => {
console.error("forced exit after timeout");
process.exit(1);
}, 25_000);
forceExit.unref();
// Give the load balancer time to notice /health/ready is failing
await new Promise((r) => setTimeout(r, 5_000));
// Stop accepting connections; resolves when all requests have finished
const closed = new Promise((resolve) => server.close(resolve));
server.closeIdleConnections(); // drop idle keep-alive sockets now
await closed;
await stopWorkers(); // finish current job, take no new ones
await pool.end(); // close database connections
clearTimeout(forceExit);
process.exit(0);
}
process.on("SIGTERM", () => shutdown("SIGTERM"));
process.on("SIGINT", () => shutdown("SIGINT"));
Why each step matters
- The readiness flip and short wait. Load balancers poll health checks every few seconds. Without the wait, they keep sending requests to a process that's stopped listening. Failing readiness first, then pausing briefly, lets traffic drain away. (Health check endpoints.)
server.close()stops accepting new connections and calls back once existing requests finish. But it won't close idle keep-alive connections on its own in all cases — clients holding an open keep-alive socket can keep the server "busy".server.closeIdleConnections()(Node 18.2+) closes the idle ones;server.closeAllConnections()is the blunt instrument for the final deadline.- The forced-exit timer guarantees you exit before the platform sends
SIGKILL, so you can log what happened. Keep it shorter than the grace period. - Closing the database pool releases connections cleanly instead of leaving the database to time them out — which matters when many instances restart at once. (Postgres connection pooling.)
Grace periods by platform
| Platform | Default time between SIGTERM and SIGKILL |
|---|---|
docker stop |
10 seconds (--time / stop_grace_period to change) |
| Kubernetes | 30 seconds (terminationGracePeriodSeconds) |
| systemd | 90 seconds (TimeoutStopSec) |
Tune your timings to fit — and raise the grace period if you have long requests or jobs. (PM2 vs systemd.)
The signal traps
Many "graceful shutdown doesn't work" reports aren't code bugs — the signal never reaches Node.
npm start doesn't forward signals reliably
CMD ["npm", "start"] # SIGTERM goes to npm, may never reach node
Run Node directly:
CMD ["node", "dist/server.js"]
Shell form in Dockerfiles
CMD node dist/server.js # runs under /bin/sh -c — sh receives SIGTERM, not node
CMD ["node", "dist/server.js"] # exec form — node is PID 1 and receives it
PID 1 has no default signal handling
In a container, your process is often PID 1, which the kernel treats specially: signals without a registered handler are ignored. With a SIGTERM handler registered (as above), Node handles it. To also reap zombie processes and get sane defaults, run with docker run --init (or init: true in Compose), which adds a tiny init process. (Writing a production Dockerfile.)
Background workers and jobs
For queue workers:
- On
SIGTERM, stop taking new jobs immediately. - Finish the current job if it can complete within the grace period.
- If it can't, make sure it's safe to be interrupted: jobs should be idempotent and the queue should hand unfinished jobs to another worker. With a Postgres-backed queue, an uncommitted job's lock is released when the connection closes, so it's picked up again. (Postgres as a job queue and idempotency keys.)
WebSockets and streams
Long-lived connections never "finish" on their own. On shutdown, tell clients to reconnect (a close frame with a "going away" code, or a final SSE event), then close them. Clients should reconnect with backoff — to a different instance. (Server-sent events vs WebSockets.)
Test it
node dist/server.js &
PID=$!
# start a slow request in another terminal, then:
kill -TERM $PID
The slow request should complete, new requests should be refused, and the process should exit cleanly with your shutdown logs. Do the same with docker stop on the real container, and run a load test through a deploy to confirm zero errors. (Load testing your app.)
The summary
- Deploys send SIGTERM, then SIGKILL after a grace period; Node exits immediately by default.
- Fail readiness, wait briefly,
server.close()pluscloseIdleConnections(), drain workers, close pools, exit — with a forced-exit backstop. - Run
nodedirectly in exec form; use--initin containers. - Make jobs idempotent so interrupted work is safely retried.
EasySpawn gives Claude Code a real server to test shutdown behaviour against — sending SIGTERM mid-request and confirming nothing is dropped — before it becomes part of your deploys. See how it works or join the waitlist.
Related: Zero-Downtime Deploys for a Small App · Finding Memory Leaks in Node.js · Reverse Proxies Explained · Do You Need Kubernetes? · PID 1 in Containers
Keep reading
Postgres Major Version Upgrades: pg_upgrade, Logical Replication, and Minimal Downtime
Major versions change the on-disk format, so upgrading PostgreSQL isn't a package update. Dump/restore vs pg_upgrade (copy, link, clone) vs logical replication cutover; extension and collation pitfalls; sequences and DDL gaps in logical replication; statistics after upgrade; and a rehearsed runbook.
What Is Nginx? The Web Server in Front of Half the Internet
Nginx ("engine-x") is a web server that serves files, forwards requests to your app, handles HTTPS and balances load. What it does, what a basic config looks like, where the files live, the commands you'll need, and whether you need it at all.