Blog
5 min read

Finding Memory Leaks in Node.js: Heap Snapshots, Retainers, and Common Culprits

Your Node.js server's memory climbs until it crashes. How to confirm a leak, read process.memoryUsage, capture heap snapshots in production safely, compare them in Chrome DevTools, follow retainer chains, and fix the usual causes: unbounded caches, listeners, timers and closures.

A Node.js service starts at 150 MB, sits at 600 MB by evening, and crashes overnight with JavaScript heap out of memory — or gets killed with exit code 137. Restarting "fixes" it until tomorrow. That's a memory leak: something keeps references to objects that should have been garbage-collected.

Here's a systematic way to find it.

First: confirm it's a leak

Memory that rises and then plateaus is usually fine — caches warming, the V8 heap growing to a comfortable size. A leak looks like a sawtooth that trends upward: garbage collection reclaims some memory each cycle, but the floor keeps rising with traffic.

Log memory periodically:

setInterval(() => {
  const m = process.memoryUsage();
  console.log(JSON.stringify({
    rss_mb: Math.round(m.rss / 1e6),
    heap_used_mb: Math.round(m.heapUsed / 1e6),
    heap_total_mb: Math.round(m.heapTotal / 1e6),
    external_mb: Math.round(m.external / 1e6),
    array_buffers_mb: Math.round(m.arrayBuffers / 1e6),
  }));
}, 60_000).unref();

What the numbers tell you:

  • heapUsed rising → a JavaScript object leak. Heap snapshots will find it.
  • rss rising but heapUsed flat → memory outside the JS heap: Buffers (external/arrayBuffers), native addons, or memory fragmentation. Heap snapshots won't show it directly; look at Buffer-heavy code (streams, file handling, image processing) and native modules.

Correlate growth with traffic: does memory rise per request, per WebSocket connection, per job?

Reproduce it locally if you can

Run the app with the inspector and generate load against the suspected endpoint. (Load testing your app.)

node --inspect dist/server.js

Open chrome://inspect in Chrome, click inspect under your process, and go to the Memory tab.

The three-snapshot technique

  1. Warm up the app (a few hundred requests), then take snapshot 1.
  2. Send a batch of requests — say 1,000 — then snapshot 2.
  3. Send another 1,000, then snapshot 3.

DevTools runs garbage collection before each snapshot, so what remains is genuinely retained.

In snapshot 3, switch the view to Comparison against snapshot 2 and sort by # Delta or Size Delta. Objects whose count grows by roughly the number of requests you sent are your leak candidates — for example, 1,000 new IncomingMessage objects, or 1,000 new closures from one function.

Follow the retainers

Select a leaking object and look at the Retainers panel at the bottom. It shows the chain of references keeping the object alive, from a GC root down:

(GC root) → global → requestCache (Map) → entry → { req, user, ... }

The first thing in that chain that shouldn't be holding on is your bug — here, a global Map used as a cache. Ignore entries in parentheses like (system) and focus on your own variable and function names. Naming functions (rather than anonymous arrows) makes this far easier to read.

Capturing snapshots in production

Sometimes the leak only happens with real traffic. Options:

  • On demand by signal: start Node with --heapsnapshot-signal=SIGUSR2, then kill -USR2 <pid> writes a snapshot to the working directory.
  • Programmatically: require("node:v8").writeHeapSnapshot() from an admin-only endpoint.
  • Automatically before a crash: --heapsnapshot-near-heap-limit=1 writes a snapshot when the heap approaches its limit.

Be careful:

  • Taking a snapshot pauses the process and can need memory comparable to the heap itself — do it on one instance taken out of the load balancer, or when traffic is low.
  • Snapshots contain everything in memory, including user data, tokens and secrets. Treat the files as sensitive and delete them afterwards.

The usual culprits

Unbounded caches

const cache = new Map();
app.get("/user/:id", async (req, res) => {
  if (!cache.has(req.params.id)) cache.set(req.params.id, await loadUser(req.params.id));
  res.json(cache.get(req.params.id));
});

Every distinct ID stays forever. Use an LRU cache with a maximum size and TTL, or an external cache. (What is caching?)

Event listeners added repeatedly

app.get("/stream", (req, res) => {
  emitter.on("update", (data) => res.write(data)); // never removed
});

Each request adds a listener that holds res forever. Remove it on close: req.on("close", () => emitter.off("update", handler)). Node's "MaxListenersExceededWarning" is often the first clue.

Timers that are never cleared

A setInterval created per request or per connection, never cleared, keeps its closure — and everything it references — alive.

Closures capturing large objects

A small callback stored long-term (in a queue, a map of pending operations) that closes over a large request or response object keeps the whole thing alive. Capture only what you need.

Per-request data in module-level state

Arrays used as logs, metrics buckets or "recent requests" that only ever grow.

Pending promises that never settle

Promises waiting on something that never happens (a lost callback, a request with no timeout) accumulate with their closures. Add timeouts to outbound calls.

Libraries and native modules

Sometimes the leak is in a dependency — an old version of a client library or a native addon. If retainers point into node_modules, check the library's issue tracker and upgrade.

Prevent regressions

  • Track heapUsed and rss as metrics with an alert on sustained growth.
  • Use bounded data structures by default.
  • Run a soak test (steady load for an hour) before major releases.
  • Set a sane --max-old-space-size and a memory limit with restart as a safety net — not a fix.

The summary

  • A leak shows as a rising floor in memory, not just high memory.
  • heapUsed rising → JS objects; rss rising alone → Buffers or native memory.
  • Take three heap snapshots around load, compare, and follow retainer chains.
  • Usual causes: unbounded caches, unremoved listeners, uncleared timers, closures, ever-growing arrays.
  • Snapshots pause the process and contain sensitive data — handle with care.

EasySpawn gives Claude Code your real server, where it can load-test your app, capture heap snapshots and compare them — and you can watch memory on server sizes up to 16 GB. See how it works or join the waitlist.

Related: How Container CPU and Memory Limits Actually Work · Graceful Shutdown in Node.js · Structured Logging · Error Monitoring for Beginners

Keep reading