wal_retrieve_retry_interval
Fact — official short description: “Sets the time to wait before retrying to retrieve WAL after a failed attempt.”
Identity
Lifecycle
| Fact | Value |
|---|---|
| First observed | PG9.5 |
| Present in | PG9.5–19 Beta 3 |
| Removed in | No |
| Introduction commit | 5d2b45e3f78a — Add GUC to control the time to wait before retrieving WAL after failed attempt. |
| Commit date | 2015-02-23 |
| Discussion | — |
Default history
| Versions | Raw boot_val |
Unit | Human value |
|---|---|---|---|
| PG9.5–19 Beta 3 | 5000 |
ms |
5 s |
How it works
Sets the time to wait before retrying to retrieve WAL after a failed attempt. A configuration reload applies a new value; existing work already in flight is not retroactively changed.
After archive, local pg_wal, and streaming retrieval fail, recovery waits this long before another attempt. Small values reduce recovery/standby reaction time but can hammer a missing archive or noisy network with retries.
Monitor and change wal_retrieve_retry_interval together with primary_conninfo, primary_slot_name, restore_command. Validate on the relevant server role and real workload, then use its sighup context to choose session change, reload, or restart; a historical boot default is not the current effective value.
Tuning advice
Advice. These are workload-specific starting points and must be validated with measurements.
| Workload | Guidance |
|---|---|
| OLTP | Size wal_retrieve_retry_interval from topology, failover roles, slot/subscription count, and reconnect headroom. Test worst-case primary latency, standby replay, and disk retention before production. |
| OLAP | Read standbys and logical subscribers often see long queries or large transactions. Put explicit bounds on replay/apply and monitor lag, worker saturation, slot restart_lsn, and conflict cancellations. |
| Small nodes | Configure only replication capacity that is actually used. Even a small topology needs bounded timeouts and slot lifecycle; unlimited retention is not reliability. |
Pigsty
Values use the fixed 8-vCPU, 32-GiB, 100-GiB SSD fixture and render the current Pigsty templates for PG19 Beta 3; this does not assert current Pigsty support for that historical or beta release.
| Template | Effective value | Versus upstream boot | Source expression |
|---|---|---|---|
| OLTP | Unmodified | — | — |
| OLAP | Unmodified | — | — |
| CRIT | Unmodified | — | — |
| TINY | Unmodified | — | — |
Advice — pending human review. Fact from the current Pigsty template projection: OLTP: PG9.5–19 Beta 3 unmodified; OLAP: PG9.5–19 Beta 3 unmodified; CRIT: PG9.5–19 Beta 3 unmodified; TINY: PG9.5–19 Beta 3 unmodified. No Pigsty-specific rationale is inferred from an absent override.
Common pitfalls
- Changing it on the wrong primary, standby, sender, or subscriber role.
- Watching only configured bytes/time instead of actual lag, slot position, and worker state.
- Failing over to a node that lacks the old primary’s capacity or prerequisites.
- Using infinite waits or WAL retention to hide a failed consumer.
Related parameters
primary_conninfo · primary_slot_name · restore_command · wal_receiver_timeout · hot_standby · hot_standby_feedback
References
- PostgreSQL 19 Beta 3 — wal_retrieve_retry_interval
- PostgreSQL 19 release notes
- Machine-readable GUC export