Blog
5 min read

PID 1 in Containers: Signals, Zombies, and Why Your Container Won't Stop

Inside a container your app is PID 1, and PID 1 is special: the kernel won't apply default signal handlers to it and it must reap orphaned children. Why docker stop takes 10 seconds, why shell-form CMD swallows SIGTERM, zombies, and the fixes: exec form, tini/--init, signal handling.

A container starts one process, and inside its PID namespace that process is PID 1 — the same position init/systemd holds on a normal Linux system. (Linux namespaces explained) PID 1 has two special responsibilities that ordinary programs were never written for. Getting them wrong causes three classic symptoms:

  • docker stop always takes 10 seconds, then the container is killed.
  • Deploys drop in-flight requests or corrupt work because the app never shut down gracefully.
  • Zombie processes pile up until the container can't create new ones.

Special rule 1: no default signal actions for PID 1

Normally, if a process hasn't installed a handler for SIGTERM or SIGINT, the kernel applies the default action: terminate. For PID 1 in a PID namespace, the kernel ignores signals that have no handler installed (except SIGKILL and SIGSTOP sent from the parent namespace).

So a program that relies on default behaviour — many do — simply doesn't react to SIGTERM when it's PID 1. docker stop sends SIGTERM, waits the grace period (10 s by default), then sends SIGKILL. Your app never got to finish requests, close database connections or flush logs.

Node.js is a common case: without a SIGTERM listener, a Node process running as PID 1 ignores it.

Special rule 2: PID 1 reaps orphans

When a process exits, it becomes a zombie until its parent calls wait() to collect its exit status. If the parent has already died, the orphan is re-parented to PID 1, which is expected to reap it.

Your app isn't an init system. If it spawns children that spawn grandchildren (a shell script, headless Chrome, git, a language server, an AI agent running commands), orphaned grandchildren get re-parented to your app, which never reaps them. Zombies hold PID slots; with a cgroup pids.max limit, the container eventually fails with fork: Resource temporarily unavailable. (Container CPU and memory limits)

ps -eo pid,ppid,stat,cmd | awk '$3 ~ /Z/'    # list zombies

Trap: shell-form CMD

CMD npm start                 # shell form

runs /bin/sh -c "npm start". Now sh is PID 1. sh doesn't forward SIGTERM to its child, so your app never hears it — even if it handles signals perfectly. And npm start adds another layer: npm as the parent of node.

Use the exec form, and run the runtime directly:

CMD ["node", "dist/server.js"]

Now node is PID 1 and receives signals directly. (Running via npm start is worth avoiding in production containers for this reason.) (Production Dockerfile for Node)

Trap: entrypoint scripts

Entrypoint scripts are useful for setup (waiting for a database, templating config). End them with exec so the app replaces the shell as PID 1:

#!/bin/sh
set -e
./bin/migrate-if-needed
exec node dist/server.js "$@"

Without exec, the shell stays PID 1 and swallows signals.

The fix for both rules: a tiny init

tini (and dumb-init) is a minimal init designed for containers. It runs as PID 1, forwards signals to your app, and reaps zombies.

Docker has it built in:

docker run --init myapp
# docker-compose.yml
services:
  app:
    image: myapp
    init: true

or bake it into the image:

RUN apt-get update && apt-get install -y --no-install-recommends tini && rm -rf /var/lib/apt/lists/*
ENTRYPOINT ["/usr/bin/tini", "--"]
CMD ["node", "dist/server.js"]

In Kubernetes, there's no --init flag; bake tini into the image, or enable shareProcessNamespace (the pause container then reaps zombies, at the cost of sharing the PID namespace between containers in the pod).

The app still has to shut down properly

An init forwards SIGTERM; your app must do something with it: stop accepting new connections, finish in-flight requests, drain job workers, close the database pool, then exit.

const server = app.listen(port)

function shutdown(signal) {
  console.log(`${signal} received, shutting down`)
  server.close(async () => {
    await db.end()
    process.exit(0)
  })
  setTimeout(() => process.exit(1), 25_000).unref()   // hard stop before the orchestrator's SIGKILL
}

process.on('SIGTERM', shutdown)
process.on('SIGINT', shutdown)

Make sure your internal timeout is shorter than the platform's grace period (docker stop -t, Compose stop_grace_period, Kubernetes terminationGracePeriodSeconds). Fail your readiness check as soon as shutdown begins so load balancers stop sending traffic. (Graceful shutdown in Node, health check endpoints, zero-downtime deploys)

STOPSIGNAL

Some programs expect a different signal for graceful shutdown (Nginx uses SIGQUIT for graceful stop; some apps use SIGINT). Set it in the image:

STOPSIGNAL SIGQUIT

Quick diagnosis

docker exec myapp ps -o pid,ppid,cmd       # who is PID 1? sh? npm? node?
time docker stop myapp                      # ~10 s means SIGTERM was ignored
docker inspect -f '{{.State.ExitCode}}' myapp   # 137 = killed by SIGKILL (or OOM)

The summary

  • PID 1 gets no default signal handling and must reap orphaned processes.
  • Shell-form CMD and entrypoint scripts without exec put a shell at PID 1 that swallows SIGTERM.
  • Use exec-form CMD, exec in scripts, and tini / --init / init: true.
  • Handle SIGTERM in the app: stop accepting, drain, close, exit before the grace period ends.
  • A 10-second docker stop and exit code 137 are the tell-tale signs.

EasySpawn runs your app and Claude Code's processes on a full VM with a real init system, so long-running agents, child processes and graceful restarts behave like they do on any Linux server. See how it works or join the waitlist.

Related: Graceful Shutdown in Node.js · Writing a Production Dockerfile · PM2 vs systemd · Docker Image Layers and OverlayFS

Keep reading