Blog
4 min read

Health Check Endpoints: What /health Should (and Shouldn't) Check

A health check endpoint tells load balancers, orchestrators and monitors whether your app can serve traffic. Liveness vs readiness, what to check and what not to, response formats, timeouts, security, and examples for Express, Next.js and Docker.

A health check endpoint — usually /health, /healthz or /api/health — is a URL that answers one question for machines: "Is this app instance OK to send traffic to?" Load balancers, container platforms, deploy scripts and uptime monitors call it constantly and act on the answer: route traffic, restart a container, roll back a deploy, page you.

Because things act on it automatically, a badly designed health check can cause outages instead of preventing them.

Two different questions: liveness and readiness

Liveness: "Is the process stuck?"

If liveness fails, the platform restarts the instance. It should only fail when a restart would help — the process is deadlocked, out of memory, or broken beyond repair.

Keep it trivial: if the server can answer an HTTP request at all, it's alive.

app.get("/health/live", (_req, res) => res.status(200).send("ok"));

Readiness: "Can it serve requests right now?"

If readiness fails, the platform stops sending traffic to this instance (but doesn't kill it). Use it for:

  • startup — still warming up, running migrations, loading config,
  • shutdown — draining connections during a deploy (graceful shutdown),
  • hard dependencies — the database this instance needs is unreachable.
let shuttingDown = false;

app.get("/health/ready", async (_req, res) => {
  if (shuttingDown) return res.status(503).json({ status: "shutting_down" });
  try {
    await withTimeout(db.query("SELECT 1"), 1000);
    res.status(200).json({ status: "ok" });
  } catch {
    res.status(503).json({ status: "unavailable", reason: "database" });
  }
});

Small apps often use a single endpoint. That's fine — just decide consciously what it means, and remember that if a platform restarts instances on failure, checking the database in that endpoint means a database blip restarts every app instance at once.

The mistake that causes cascading failures

Don't make liveness depend on external services. If /health checks the database and the platform restarts on failure, then a 30-second database hiccup restarts all your app instances simultaneously — turning a brief slowdown into a full outage, followed by a thundering herd of reconnects.

Rules of thumb:

  • Liveness: no external dependencies.
  • Readiness: only hard dependencies this instance can't function without (usually the primary database).
  • Never check optional or third-party services (email provider, analytics, a payments API) in a health check. If Stripe is down, your app is degraded, not dead — and restarting it won't fix Stripe.

Make it fast and cheap

Health checks run every few seconds, from several places.

  • Use a trivial query (SELECT 1), never a real one.
  • Set a short timeout on dependency checks (around a second), so a slow database makes the check fail fast rather than hang.
  • Don't log every successful check — it floods your logs. (Structured logging.)
  • Exclude it from rate limiting and authentication, or monitors will be blocked.

Response format

Status codes are what machines read:

  • 200 — healthy.
  • 503 Service Unavailable — not healthy / not ready.

A small JSON body helps humans debugging:

{ "status": "ok", "version": "1.42.0", "uptime_s": 86400 }

Including the app version (or git commit) is handy: after a deploy, you can confirm which version each instance is running.

Security

Health endpoints are usually public, so don't leak information attackers can use:

  • no stack traces, connection strings, hostnames or dependency versions in public responses,
  • put detailed diagnostics on a separate endpoint that requires authentication or is only reachable internally.

Using it

  • Docker:

    HEALTHCHECK --interval=30s --timeout=3s --start-period=20s --retries=3 \
      CMD curl -fsS http://localhost:3000/health/live || exit 1
    

    (Needs curl or similar in the image — or use a tiny Node script.)

  • Docker Compose: a healthcheck: block, and depends_on: condition: service_healthy to start services in order. (Docker Compose for local development.)

  • Load balancers and reverse proxies remove unhealthy instances from rotation. (Horizontal vs vertical scaling.)

  • Deploy scripts wait for readiness before switching traffic. (Zero-downtime deploys.)

  • Uptime monitors check it from outside every minute. (How to know when your app is down.)

Next.js

A route handler works fine:

// app/api/health/route.ts
export const dynamic = "force-dynamic";

export function GET() {
  return Response.json({ status: "ok" });
}

force-dynamic makes sure the response isn't cached at build time — a cached "ok" is worse than no health check.

The summary

  • Liveness: "restart me if this fails" — keep it trivial, no external dependencies.
  • Readiness: "don't send me traffic if this fails" — startup, shutdown, hard dependencies only.
  • Never check optional third-party services; use short timeouts and cheap queries.
  • Return 200 or 503, include the version, leak nothing sensitive.

EasySpawn gives Claude Code your real running app and database on one server, so it can add a health check, then stop the database and confirm the endpoint and your app behave the way you intended. See how it works or join the waitlist.

Related: Do You Need Kubernetes? · PM2 vs systemd · HTTP Status Codes Explained · Load Testing Your App

Keep reading