Blog
5 min read

Graceful Shutdown in Node.js: Handling SIGTERM Without Dropping Requests

Every deploy and restart sends your Node.js app SIGTERM. How to shut down gracefully: stop accepting traffic, fail readiness, finish in-flight requests, close keep-alive connections, drain workers, close database pools — plus the Docker PID 1 and npm start signal traps.

Every deploy, scale-down, restart and container replacement ends the same way for your process: it receives a SIGTERM signal, and some seconds later — if it hasn't exited — a SIGKILL it can't catch. What your app does in between decides whether users see failed requests, whether jobs are half-done, and whether database connections are left dangling.

By default, Node.js does nothing special: on SIGTERM it simply exits, killing every in-flight request mid-response.

What a graceful shutdown does

  1. Stop receiving new traffic — tell the load balancer you're going away, and stop accepting new connections.
  2. Finish in-flight work — let active requests complete.
  3. Close idle connections, including HTTP keep-alive sockets.
  4. Stop background work — let the current job finish, don't take new ones.
  5. Release resources — close database pools, flush logs and metrics.
  6. Exit — before the platform's deadline, with a forced exit as a backstop.

A working implementation

import http from "node:http";
import { app } from "./app.js";
import { pool } from "./db.js";

const server = http.createServer(app);
server.listen(process.env.PORT ?? 3000);

let shuttingDown = false;

// Readiness endpoint used by the load balancer
app.get("/health/ready", (_req, res) => {
  res.status(shuttingDown ? 503 : 200).end();
});

async function shutdown(signal) {
  if (shuttingDown) return;
  shuttingDown = true;
  console.log(JSON.stringify({ msg: "shutdown started", signal }));

  // Backstop: never outlive the platform's grace period
  const forceExit = setTimeout(() => {
    console.error("forced exit after timeout");
    process.exit(1);
  }, 25_000);
  forceExit.unref();

  // Give the load balancer time to notice /health/ready is failing
  await new Promise((r) => setTimeout(r, 5_000));

  // Stop accepting connections; resolves when all requests have finished
  const closed = new Promise((resolve) => server.close(resolve));
  server.closeIdleConnections(); // drop idle keep-alive sockets now
  await closed;

  await stopWorkers();  // finish current job, take no new ones
  await pool.end();     // close database connections

  clearTimeout(forceExit);
  process.exit(0);
}

process.on("SIGTERM", () => shutdown("SIGTERM"));
process.on("SIGINT", () => shutdown("SIGINT"));

Why each step matters

  • The readiness flip and short wait. Load balancers poll health checks every few seconds. Without the wait, they keep sending requests to a process that's stopped listening. Failing readiness first, then pausing briefly, lets traffic drain away. (Health check endpoints.)
  • server.close() stops accepting new connections and calls back once existing requests finish. But it won't close idle keep-alive connections on its own in all cases — clients holding an open keep-alive socket can keep the server "busy". server.closeIdleConnections() (Node 18.2+) closes the idle ones; server.closeAllConnections() is the blunt instrument for the final deadline.
  • The forced-exit timer guarantees you exit before the platform sends SIGKILL, so you can log what happened. Keep it shorter than the grace period.
  • Closing the database pool releases connections cleanly instead of leaving the database to time them out — which matters when many instances restart at once. (Postgres connection pooling.)

Grace periods by platform

Platform Default time between SIGTERM and SIGKILL
docker stop 10 seconds (--time / stop_grace_period to change)
Kubernetes 30 seconds (terminationGracePeriodSeconds)
systemd 90 seconds (TimeoutStopSec)

Tune your timings to fit — and raise the grace period if you have long requests or jobs. (PM2 vs systemd.)

The signal traps

Many "graceful shutdown doesn't work" reports aren't code bugs — the signal never reaches Node.

npm start doesn't forward signals reliably

CMD ["npm", "start"]   # SIGTERM goes to npm, may never reach node

Run Node directly:

CMD ["node", "dist/server.js"]

Shell form in Dockerfiles

CMD node dist/server.js        # runs under /bin/sh -c — sh receives SIGTERM, not node
CMD ["node", "dist/server.js"] # exec form — node is PID 1 and receives it

PID 1 has no default signal handling

In a container, your process is often PID 1, which the kernel treats specially: signals without a registered handler are ignored. With a SIGTERM handler registered (as above), Node handles it. To also reap zombie processes and get sane defaults, run with docker run --init (or init: true in Compose), which adds a tiny init process. (Writing a production Dockerfile.)

Background workers and jobs

For queue workers:

  • On SIGTERM, stop taking new jobs immediately.
  • Finish the current job if it can complete within the grace period.
  • If it can't, make sure it's safe to be interrupted: jobs should be idempotent and the queue should hand unfinished jobs to another worker. With a Postgres-backed queue, an uncommitted job's lock is released when the connection closes, so it's picked up again. (Postgres as a job queue and idempotency keys.)

WebSockets and streams

Long-lived connections never "finish" on their own. On shutdown, tell clients to reconnect (a close frame with a "going away" code, or a final SSE event), then close them. Clients should reconnect with backoff — to a different instance. (Server-sent events vs WebSockets.)

Test it

node dist/server.js &
PID=$!
# start a slow request in another terminal, then:
kill -TERM $PID

The slow request should complete, new requests should be refused, and the process should exit cleanly with your shutdown logs. Do the same with docker stop on the real container, and run a load test through a deploy to confirm zero errors. (Load testing your app.)

The summary

  • Deploys send SIGTERM, then SIGKILL after a grace period; Node exits immediately by default.
  • Fail readiness, wait briefly, server.close() plus closeIdleConnections(), drain workers, close pools, exit — with a forced-exit backstop.
  • Run node directly in exec form; use --init in containers.
  • Make jobs idempotent so interrupted work is safely retried.

EasySpawn gives Claude Code a real server to test shutdown behaviour against — sending SIGTERM mid-request and confirming nothing is dropped — before it becomes part of your deploys. See how it works or join the waitlist.

Related: Zero-Downtime Deploys for a Small App · Finding Memory Leaks in Node.js · Reverse Proxies Explained · Do You Need Kubernetes? · PID 1 in Containers

Keep reading