Blog
4 min read

How to Deploy a FastAPI App to Production

Running uvicorn main:app --reload is for development. A production FastAPI deploy needs worker processes, a process manager or container, a reverse proxy with HTTPS, proper settings, migrations, and health checks. Step-by-step options for a VPS and Docker.

FastAPI apps are easy to start — uvicorn main:app --reload and you're running. That command is for development. In production you need more processes, automatic restarts, HTTPS, and settings that don't leak debug information.

Here's a production setup, first on a plain server and then in Docker. (Choosing a framework? Flask vs Django vs FastAPI.)

The production stack

Internet → Caddy/Nginx (HTTPS, port 443) → Uvicorn workers (127.0.0.1:8000) → FastAPI app → Postgres
  • Uvicorn is the ASGI server that runs your app.
  • Multiple workers let you use more than one CPU core (Python processes are effectively single-core for CPU work).
  • A process manager (systemd) or container runtime restarts it on crashes and boot.
  • A reverse proxy terminates HTTPS and forwards to Uvicorn. (Reverse proxies explained)

Step 1: Make the app production-ready

Settings from environment variables, validated at startup with pydantic-settings:

# settings.py
from pydantic_settings import BaseSettings

class Settings(BaseSettings):
    database_url: str
    secret_key: str
    environment: str = "production"
    cors_origins: list[str] = []

settings = Settings()   # fails fast if DATABASE_URL is missing

(Environment variables explained)

A health endpoint that checks the database:

@app.get("/healthz")
async def healthz():
    async with engine.connect() as conn:
        await conn.execute(text("SELECT 1"))
    return {"status": "ok"}

(Health check endpoints)

Turn off the interactive docs in production if your API is private: FastAPI(docs_url=None, redoc_url=None) — or protect them.

CORS: list exact origins; never allow_origins=["*"] with credentials. (CORS errors explained)

Pin dependencies with a lock file (uv.lock, poetry.lock, or a pinned requirements.txt). (Python virtual environments)

Step 2 (option A): A plain server with systemd

On the server, in a virtual environment:

cd /srv/api
uv sync --frozen            # or: python -m venv .venv && .venv/bin/pip install -r requirements.txt

Create /etc/systemd/system/api.service:

[Unit]
Description=FastAPI app
After=network.target

[Service]
User=deploy
WorkingDirectory=/srv/api
EnvironmentFile=/srv/api/.env
ExecStart=/srv/api/.venv/bin/uvicorn main:app --host 127.0.0.1 --port 8000 --workers 2 --proxy-headers
Restart=always

[Install]
WantedBy=multi-user.target
sudo systemctl enable --now api
curl http://127.0.0.1:8000/healthz

How many workers? Start with the number of CPU cores (for async apps, often 1–2 per core is plenty) and adjust by watching memory and latency. Each worker is a full copy of your app in memory. Some teams use Gunicorn as the process manager with Uvicorn workers (gunicorn -k uvicorn.workers.UvicornWorker); Uvicorn's own --workers is fine for most apps.

--proxy-headers makes FastAPI trust X-Forwarded-For/X-Forwarded-Proto from the proxy, so URLs and client IPs are correct. Only trust them from your proxy (--forwarded-allow-ips).

Step 2 (option B): Docker

FROM python:3.13-slim
WORKDIR /app

COPY pyproject.toml uv.lock ./
RUN pip install uv && uv sync --frozen --no-dev
COPY . .

RUN useradd --create-home app
USER app

EXPOSE 8000
CMD ["/app/.venv/bin/uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "2", "--proxy-headers"]

Inside a container, listen on 0.0.0.0 (the container's interface), but publish the port only to localhost on the host: -p 127.0.0.1:8000:8000. (Reduce Docker image size, production Dockerfiles — the principles carry over.)

Step 3: HTTPS with a reverse proxy

With Caddy, the whole config:

api.yourdomain.com {
    reverse_proxy 127.0.0.1:8000
}

Caddy obtains and renews the certificate. With Nginx, add certbot. (What is Nginx?) For streaming responses, disable proxy buffering on that route. (Streaming LLM responses)

Step 4: Database migrations

Run Alembic migrations as a deploy step, before the new code starts — not on app startup with multiple workers racing:

.venv/bin/alembic upgrade head
sudo systemctl restart api

(Database migrations explained)

Use a connection pool sized for workers × pool size staying under Postgres's connection limit. (Connection pooling)

Step 5: Background work

Don't run long tasks inside request handlers — BackgroundTasks is fine for small fire-and-forget work, but anything slow or important belongs in a proper queue with a separate worker process. (Background jobs)

Pre-launch checklist

  • Settings from env, validated; no secrets in code
  • Multiple workers under systemd or a container with restart policy
  • Uvicorn bound to localhost; HTTPS via proxy
  • /healthz checks the database
  • Docs disabled or protected; CORS restricted
  • Migrations as a deploy step
  • Structured logs and error monitoring (error monitoring)
  • Backups of the database

The summary

  • Production = Uvicorn workers + process manager/container + HTTPS proxy.
  • Validate settings at startup; add a health check; lock down docs and CORS.
  • Bind to localhost (or publish container ports to localhost only).
  • Run migrations as a separate deploy step; move slow work to a queue.

EasySpawn runs Python APIs on a dedicated server with Postgres alongside, automatic HTTPS on your domain and daily backups — and Claude Code can set up and deploy the FastAPI app for you. See how it works or join the waitlist.

Related: Flask vs Django vs FastAPI · Python Virtual Environments · How to Secure a New VPS · PM2 vs systemd

Keep reading