Re: pg_resetwal with replication slot (17.11) - Mailing list pgsql-hackers

From Chao Li
Subject Re: pg_resetwal with replication slot (17.11)
Date
Msg-id C8035643-1DD9-4E00-ADB6-EE612DA94C84@gmail.com
Whole thread
In response to pg_resetwal with replication slot (17.11)  (Rıdvan Korkmaz <serkan.ridvan.korkmaz@gmail.com>)
List pgsql-hackers

> On Sep 21, 2026, at 20:21, Rıdvan Korkmaz <serkan.ridvan.korkmaz@gmail.com> wrote:
>
> Hi Dear Experts,
> I hit a case seems odd. I wonder if I do something unexpected, or something is here I can't see.
>
> First, these are all on test environment. The version is PostgreSQL 17.11, uses homebrew installation on MacOS.
>
> I have a master - replica setup, both are on the same host.
> master's
> PGDATA = m
> port = 15432
> rs = rep slot for streaming replication used by instance "r", created on master (m instance)
> max_wal_size = 4GB
> min_wal_size = 2GB
> wal_level = replica
>
>
> replica's
> PGDATA = r
> port = 25432
> primary_conninfo = created by pg_basebackup
> primary_slot_name = rs
>
>
> Case: I have 16MB WAL files on master instance (so on replica). I want to utilize 1GB WAL files.
> Here are the steps I take.
>
> 1. setup master - replica run on same host in respective directories and on ports
> 2. verify streaming replication works
> 3. verify "rs" (replication slot), master ("m" instance), and replica ("r" instance) have SAME WAL lsn
> 4. stop master, keep replica online (simulation for actual case) (pg_ctl-17 stop -D m)
> 5. run "pg_resetwal-17 -D m --wal-segsize=1024" on master instance.
> 6. start master instance, started, fine. Replica complains about "ERROR: requested WAL segment
000000010000000000000000has already been removed", no worry. 
> 7. stop master again, cool, done. (last log lines: "checkpoint complete", "database system is shut down")
> 8. start master -> bamm, could not start. (pg_ctl-17 start -D m -l m.log)
>
> There is no read, write between after step 3 (after verification of WAL lsns)
>
> Final failure log (step 8):
> 2026-09-21 14:47:45.525 +03 [22817] LOG: starting PostgreSQL 17.11 (Homebrew) on aarch64-apple-darwin25.6.0, compiled
byApple clang version 21.0.0 (clang-2100.1.1.101), 64-bit 
> 2026-09-21 14:47:45.525 +03 [22817] LOG: listening on IPv4 address "127.0.0.1", port 15432
> 2026-09-21 14:47:45.525 +03 [22817] LOG: listening on Unix socket "/tmp/.s.PGSQL.15432"
> 2026-09-21 14:47:45.528 +03 [22820] LOG: database system was shut down at 2026-09-21 14:46:42 +03
> 2026-09-21 14:47:45.528 +03 [22820] LOG: invalid checkpoint record
> 2026-09-21 14:47:45.528 +03 [22820] PANIC: could not locate a valid checkpoint record at 0/40000110
> 2026-09-21 14:47:45.528 +03 [22817] LOG: startup process (PID 22820) was terminated by signal 6: Abort trap: 6
> 2026-09-21 14:47:45.528 +03 [22817] LOG: terminating any other active server processes
> 2026-09-21 14:47:45.529 +03 [22817] LOG: shutting down due to startup process failure
> 2026-09-21 14:47:45.529 +03 [22817] LOG: database system is shut down
>
>
>
> After pg_resetwal, first start of master successful, but a second start fails.
> I guess this causes master to be lost.
>
> I'm able to spot the issue:
> The issue is replication slot. If I would have removed replication slot before second start (do it between 6 and 7),
itsucceeds. 
>
> Questions:
> 1. Is this behavior is expected?
> 2. Should replication slot case mentioned in PostgreSQL documents? (I checked yet could not see)
> 3. Am I doing something out of order, unexpected?
> 4. Once I understood the case, I dropped replication slot and able to start master. Now I want to copy
m/global/pg_controlto replica and m/pg_wal to replica as well and complete wal segment size change. I wonder if this
wayis documented or supported. I can say "it works" but does not mean "supported or documented at all". 
>
> Thank you in advance.
>
> Attachments:<1-master-replica-setup-info.txt><3-all-wal-lsn-same.txt>
>

To make the issue easier to reproduce, I created the attached repro_slot_wal_segsize.sh. It basically follows the
proceduredescribed by Rıdvan. I added "sleep 1" before the second server stop so that the log messages generated by
thatstop are easier to distinguish. I also added some temporary logging to show the problem more explicitly. See the
attachedtemp_log.diff for those changes. 

Here are the server logs from my reproduction:
```
2026-09-28 16:49:57.428 CST [76061] LOG:  EVAN checkpoint cleanup after decrement: segment 0
2026-09-28 16:49:57.428 CST [76061] LOG:  EVAN RemoveOldXlogFiles: remove through segment 0, boundary
000000000000000000000000,end segment 1, recycle through segment 128 
2026-09-28 16:49:58.786 CST [76148] LOG:  EVAN restoring replication slot "rs" from disk
2026-09-28 16:49:58.787 CST [76148] LOG:  EVAN computed replication slot minimum LSN 0/0151FA80
2026-09-28 16:49:58.882 CST [76155] LOG:  EVAN KeepLogSeg entry: end 0/400000E8, slot minimum 0/0151FA80, current
segment1, input segment 1 
2026-09-28 16:49:58.882 CST [76155] LOG:  EVAN KeepLogSeg mapped slot minimum to segment 0
2026-09-28 16:49:58.882 CST [76155] LOG:  EVAN KeepLogSeg exit: candidate segment 0, output segment 0
```

The following messages are generated by the second server stop, as shown by their timestamps:
```
2026-09-28 16:50:00.899 CST [76146] LOG:  EVAN checkpoint cleanup before KeepLogSeg: redo 0/400000E8, end 0/40000170,
segment1 
2026-09-28 16:50:00.899 CST [76146] LOG:  EVAN KeepLogSeg entry: end 0/40000170, slot minimum 0/0151FA80, current
segment1, input segment 1 
2026-09-28 16:50:00.899 CST [76146] LOG:  EVAN KeepLogSeg mapped slot minimum to segment 0
```

Here is where the problem begins. The slot’s restart_lsn is 0/0151FA80. After wal_segment_size has been changed to 1
GB,this call in KeepLogSeg(): 
```
XLByteToSeg(keep, segno, wal_segment_size);
```

maps that LSN to segment zero. The slot’s restart_lsn refers to the old WAL history and is no longer meaningful after
pg_resetwalhas replaced that history: 
```
2026-09-28 16:50:00.899 CST [76146] LOG:  EVAN KeepLogSeg exit: candidate segment 0, output segment 0
2026-09-28 16:50:00.899 CST [76146] LOG:  EVAN checkpoint cleanup after KeepLogSeg: segment 0
2026-09-28 16:50:00.899 CST [76146] LOG:  EVAN checkpoint cleanup before decrement: segment 0
2026-09-28 16:50:00.899 CST [76146] LOG:  EVAN checkpoint cleanup after decrement: segment 18446744073709551615
```

CreateCheckPoint() then makes the problem worse by unconditionally decrementing _logSegNo. This decrement is normally
requiredbecause RemoveOldXlogFiles() removes files whose segment numbers are less than or equal to the supplied
boundary.However, because _logSegNo is already 0 and XLogSegNo is unsigned, the decrement wraps to UINT64_MAX: 
```
2026-09-28 16:50:00.899 CST [76146] LOG:  EVAN RemoveOldXlogFiles: remove through segment 18446744073709551615,
boundary00000000FFFFFFFF00000003, end segment 1, recycle through segment 2 
2026-09-28 16:50:00.899 CST [76146] LOG:  EVAN RemoveOldXlogFiles: candidate 000000010000000000000001, next segment 1,
recyclethrough segment 2 
2026-09-28 16:50:00.899 CST [76146] LOG:  EVAN RemoveOldXlogFiles: removing WAL segment 000000010000000000000001
2026-09-28 16:50:00.899 CST [76146] LOG:  EVAN RemoveXlogFile: candidate 000000010000000000000001, next segment 1,
recyclethrough segment 2 
2026-09-28 16:50:00.899 CST [76146] LOG:  EVAN InstallXLogFileSegment: renaming pg_wal/000000010000000000000001 to
pg_wal/000000010000000000000002as segment 2 
```

Consequently, RemoveOldXlogFiles() incorrectly recycles 000000010000000000000001 as 000000010000000000000002. Segment 1
isno longer available under its expected name, even though it contains the checkpoint created after pg_resetwal. The
nextstartup therefore cannot locate the checkpoint referenced by pg_control. 

Running pg_resetwal again creates another checkpoint and allows the server to start again. However, this is only a
temporaryrecovery. If the stale slot remains, a later checkpoint cleanup can reproduce the same failure. 

Actually, the repro must be run with assertions disabled to reach the WAL recycling. I initially reproduced it that way
becauseI had disabled assertions in my sandbox while load testing another patch. With assertions enabled, the second
serverstop fails earlier: 
```
TRAP: failed Assert("!(possible_causes & RS_INVAL_WAL_REMOVED) || oldestSegno > 0"), File: "slot.c", Line: 2230, PID:
32858
0   postgres                            0x0000000104d538a0 ExceptionalCondition + 216
1   postgres                            0x0000000104a3b664 InvalidateObsoleteReplicationSlots + 148
2   postgres                            0x0000000104546fb0 CreateCheckPoint + 3656
3   postgres                            0x00000001045456a0 ShutdownXLOG + 424
4   postgres                            0x00000001049c0328 CheckpointerMain + 2276
5   postgres                            0x00000001049c5b98 postmaster_child_launch + 464
6   postgres                            0x00000001049cabd4 StartChildProcess + 308
7   postgres                            0x00000001049c9ae8 PostmasterMain + 6128
8   postgres                            0x000000010483c158 main + 924
9   dyld                                0x00000001827ac4e4 start + 6992
2026-09-28 14:05:39.589 CST [32855] LOG:  checkpointer process (PID 32858) was terminated by signal 6: Abort trap: 6
```

This shows that passing segment 0 to InvalidateObsoleteReplicationSlots() violates the function’s existing
precondition.

Changing wal_segment_size caused the stale restart_lsn in this repro to map to segment 0, which exposed the underflow.
However,the real problem is that pg_resetwal discards the previous WAL history while leaving replication slot state
unchanged.The stored restart_lsn values are no longer valid in the new WAL history. Even when the primary-startup
failuredoes not occur, an existing standby cannot resume replication from the discarded history. Its slot must be
recreated,and the standby must be rebuilt from a fresh base backup. 

How to fix? My first thought was to prevent changing wal_segment_size when replication slots exist. In that case, users
wouldhave to drop all replication slots before changing wal_segment_size, otherwise, pg_resetwal would fail with a
hint.

On second thought, I don't think we need to add this complexity to pg_resetwal. When pg_resetwal is used to recover a
clusterwith corrupted WAL or a corrupted control file, the doc already instructs users to immediately dump the data,
runinitdb, and restore into a new cluster. Replication slots do not play a meaningful role in that recovery procedure. 

The relevant special case is using --wal-segsize to change the WAL segment size of an otherwise sound cluster without
runninginitdb. In that case, existing replication slots retain restart_lsn values referring to the discarded WAL
history.As this repro demonstrates, increasing the segment size can make such a value map to incorrect segment numbers
andtrigger the failure. 

Since pg_resetwal is not frequently used, this problem is limited to this special use of --wal-segsize, and the server
canbe recovered by running pg_resetwal again and then immediately dropping the stale slots, I think enhancing the doc
issufficient. The doc for --wal-segsize should tell users to drop all replication slots before changing the segment
size.It should also explain that existing standbys cannot resume replication afterward and must be rebuilt from a new
basebackup. 

See the attached v1 patch for my proposed documentation change.

To answer Rıdvan's questions:

> Questions:
> 1. Is this behavior is expected?

No.

> 2. Should replication slot case mentioned in PostgreSQL documents? (I checked yet could not see)

I think so.

> 3. Am I doing something out of order, unexpected?

For a planned WAL segment-size change on a sound cluster, the standby should first be stopped and the old replication
slotshould be dropped before running pg_resetwal. 

If the slot was not dropped, it may be possible to start the primary once and drop the slot before the next checkpoint
orshutdown. The standby must be stopped first so that it does not reacquire the slot. If startup has already failed,
anotherpg_resetwal may allow one more startup, but the stale slot must then be removed immediately. 

> 4. Once I understood the case, I dropped replication slot and able to start master. Now I want to copy
m/global/pg_controlto replica and m/pg_wal to replica as well and complete wal segment size change. I wonder if this
wayis documented or supported. I can say "it works" but does not mean "supported or documented at all”. 

AFAIK, no. Copying only global/pg_control and pg_wal does not produce a consistent standby and is not a supported
procedure.After the primary’s WAL history has been reset, the standby should be recreated from a fresh base backup and
attachedusing a newly created replication slot. 

Best regards,
--
Chao Li (Evan)
HighGo Software Co., Ltd.
https://www.highgo.com/





Attachment

pgsql-hackers by date:

Previous
From: Zsolt Parragi
Date:
Subject: Re: injection_points: canceled or terminated waiters leak their wait slots
Next
From: Peter Smith
Date:
Subject: Re: [PATCH] Table sync race with REFRESH PUBLICATION