Hi PostgreSQL team,
We encountered a standby startup PANIC after a switchover on PostgreSQL 18.6. A standalone reproducer and recovery TAP test reproduce it on 18.6, 17.11, REL_18_STABLE, and master; the same test passes on 18.4 and 17.10.
Expected: a standby that has replayed a heap/visibility-map truncation restarts recovery successfully.
Actual on restart:
LOG: redo starts at 0/3059738
WARNING: page 0 of relation base/5/16396_vm does not exist
PANIC: WAL contains references to invalid pages
LOG: startup process (PID 55286) was terminated by signal 6: Abort trap: 6
Reproduction: Build PostgreSQL with --enable-tap-tests --enable-injection-points and install injection_points. Copy the attached t/058_vm_truncate_invalid_pages.pl into src/test/recovery/t/, then run:
make -C src/test/recovery check PROVE_TESTS=t/058_vm_truncate_invalid_pages.pl
The test holds a standby restartpoint, promotes that standby, deletes and vacuums an all-visible table (truncating its heap and VM forks), rejoins the old primary as a standby, and restarts it. On 18.6, 20/20 TAP runs failed with this PANIC; on 18.4, 20/20 passed. A shell reproducer, exact steps, WAL excerpts, version matrix, and two candidate patches are included in the attachment.
The affected 18.6 build is REL_18_6 (724edf9); 17.11 is REL_17_11 (083ac03). The runs were on macOS 15 / Darwin 24.6.0, arm64, clang 21.1.8, with debug and injection points enabled and without cassert. The original symptom was on Linux in a three-node streaming setup. The analysis in REPORT.md points to VM block reads without an FPI during redo after an already replayed VM truncation; that diagnosis and the candidate fixes are provided for review.
Regards,
Jacky Nguyen