Postgres Read Replicas: Streaming Replication, Lag, and Read-Your-Writes
How Postgres physical streaming replication works, sync vs async, measuring replication lag, query conflicts and hot_standby_feedback, routing reads in your app without breaking read-your-writes, replication slots that fill disks, and when a replica is the wrong fix.
A read replica is a copy of your Postgres database that continuously receives changes from the primary and can serve read-only queries. Replicas are used for three different jobs, and it's worth being clear which you want:
- Read scaling — offload
SELECTtraffic from the primary. - Isolation — run heavy analytics or reports without hurting production.
- High availability — a warm standby to promote if the primary dies.
Before adding one for (1), make sure you actually need it: a missing index, an N+1 query, or no connection pooling is a far more common cause of an overloaded primary. (Database indexes, N+1 queries, connection pooling)
How physical streaming replication works
Every change in Postgres is first written to the write-ahead log (WAL). Streaming replication ships that WAL to standbys, which replay it — producing a byte-for-byte copy of the primary's data files.
Primary: write → WAL → (walsender) ──stream──▶ (walreceiver) Standby: write WAL → replay → readable
Characteristics:
- Whole cluster: every database and table; same major version; same architecture.
- Read-only standby (
hot_standby = on): queries allowed, writes rejected. - Asynchronous by default: the primary doesn't wait for the standby.
Setting one up (managed services do this with a click):
# on the new standby host, with the primary allowing a replication connection in pg_hba.conf
pg_basebackup -h primary -U replicator -D /var/lib/postgresql/18/main \
-X stream -R -S replica1 -C
-R writes the connection settings and standby.signal; -S replica1 -C creates a replication slot. Start Postgres and it begins streaming.
Replication slots: safety with a sharp edge
A slot makes the primary keep WAL until that standby has received it, so a standby that disconnects for a while can catch up instead of needing a full rebuild.
The edge: if a standby disappears and its slot remains, WAL accumulates on the primary indefinitely until the disk fills and the primary stops. (No space left on device) Monitor slots, and set max_slot_wal_keep_size to cap retention:
SELECT slot_name, active,
pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained
FROM pg_replication_slots;
Sync vs async
| Asynchronous (default) | Synchronous | |
|---|---|---|
| Commit waits for standby | No | Yes (per synchronous_commit level) |
| Commit latency | Unaffected | + network round trip |
| Data loss if primary dies | Possibly the last moments of transactions | None for confirmed commits |
| Risk | Lag | Primary stalls if sync standby is unavailable (unless using quorum) |
synchronous_commit = remote_apply also guarantees the change is visible on the standby before commit returns — useful for read-your-writes, expensive for latency. Most read-scaling setups stay async and handle lag in the application.
Measuring lag
On the primary:
SELECT application_name, state,
write_lag, flush_lag, replay_lag,
pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn)) AS bytes_behind
FROM pg_stat_replication;
On the standby:
SELECT now() - pg_last_xact_replay_timestamp() AS replay_delay;
(The standby measure looks large when the primary is idle — no new transactions to replay — so prefer the primary's view or a heartbeat table updated every few seconds.)
Lag spikes come from bulk writes (big migrations, mass updates), slow standby disks, network limits, and replay blocked by queries on the standby.
Query conflicts on the standby
Replay sometimes needs to remove row versions that a long-running standby query still needs (because VACUUM on the primary cleaned them). Postgres must either delay replay (lag grows) or cancel the query:
ERROR: canceling statement due to conflict with recovery
Knobs:
max_standby_streaming_delay— how long replay waits before cancelling queries.hot_standby_feedback = on— the standby tells the primary which rows it still needs, so VACUUM there keeps them. Fewer cancellations, but long standby queries now cause bloat on the primary. (Postgres MVCC explained)
A common split: a low-lag replica for app reads (short queries, default settings) and a separate analytics replica tolerant of lag.
Routing reads in the application
The hard part isn't replication; it's deciding which queries may go to a replica.
Read-your-writes is the classic bug: a user saves their profile (primary), the page reloads and reads from a lagging replica, and their change appears lost. Strategies:
- Pin after write: after a user writes, send their reads to the primary for a few seconds (store a timestamp in the session).
- LSN-based: after a write, record
pg_current_wal_lsn(); route reads to a replica only oncepg_last_wal_replay_lsn()has passed it — otherwise use the primary. - Route by use case, not per query: dashboards, search, exports and public pages → replica; anything in a request that just wrote, auth checks, and checkout → primary.
- Inside a transaction, stay on one node.
Many ORMs support read/write splitting (e.g. read-replica extensions or separate clients); poolers and proxies like PgBouncer (with separate pools), Pgpool-II or cloud proxies can route at the connection level, but only your application knows which reads must be fresh.
Logical replication is different
Logical replication sends row-level changes for chosen tables via publications/subscriptions. It works across major versions and allows a writable subscriber with different indexes — great for upgrades, migrations and feeding other systems — but it's not a full physical copy, doesn't replicate DDL or sequences automatically, and isn't the usual tool for read replicas. (Postgres major version upgrades)
Replicas are not backups
A replica faithfully replays your mistakes: DROP TABLE on the primary is a dropped table on the replica a moment later. You still need backups and point-in-time recovery. (Postgres PITR)
The summary
- Streaming replication ships WAL to standbys that replay it; async by default.
- Slots prevent standby gaps but can fill the primary's disk — cap and monitor them.
- Measure lag from
pg_stat_replication; expect spikes during bulk writes. - Standby query conflicts trade off against primary bloat (
hot_standby_feedback). - Route reads deliberately and protect read-your-writes; replicas aren't backups.
EasySpawn servers come with PostgreSQL on the same machine as your app — plenty for most small and medium apps — plus daily backups and snapshots for when something goes wrong. See pricing or join the waitlist.
Related: Postgres Connection Pooling · Postgres Point-in-Time Recovery · Horizontal vs Vertical Scaling · Postgres MVCC Explained
Keep reading
Postgres VACUUM and Table Bloat: How It Works and How to Keep It Under Control
Why Postgres tables bloat, what VACUUM and autovacuum actually do, tuning autovacuum for large tables, what blocks cleanup (long transactions, replication slots), transaction ID wraparound, and how to reclaim space without VACUUM FULL's exclusive lock.
The Transactional Outbox Pattern: Reliable Events Without Dual Writes
Writing to your database and publishing an event can't be made atomic, so one eventually happens without the other. How the transactional outbox fixes it: polling relays vs CDC, ordering, at-least-once delivery, idempotent consumers with an inbox, cleanup, and monitoring.