Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation - Mailing list pgsql-bugs

From Kirill Reshke
Subject Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
Date
Msg-id CALdSSPjEM90Y_6RQGox7OFJC+oAurWy5YSuXaH8Mr7m-TMOyWA@mail.gmail.com
Whole thread
In response to PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation  (Jacky Nguyen <nktpro@gmail.com>)
Responses Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
List pgsql-bugs
On Sun, 27 Sept 2026 at 11:08, Jacky Nguyen <nktpro@gmail.com> wrote:
>
> Hi PostgreSQL team,
>
> We encountered a standby startup PANIC after a switchover on PostgreSQL 18.6. A standalone reproducer and recovery
TAPtest reproduce it on 18.6, 17.11, REL_18_STABLE, and master; the same test passes on 18.4 and 17.10. 
>
> Expected: a standby that has replayed a heap/visibility-map truncation restarts recovery successfully.
> Actual on restart:
>
> LOG:  redo starts at 0/3059738
> WARNING:  page 0 of relation base/5/16396_vm does not exist
> PANIC:  WAL contains references to invalid pages
> LOG:  startup process (PID 55286) was terminated by signal 6: Abort trap: 6
>
> Reproduction: Build PostgreSQL with --enable-tap-tests --enable-injection-points and install injection_points. Copy
theattached t/058_vm_truncate_invalid_pages.pl into src/test/recovery/t/, then run: 
>
> make -C src/test/recovery check PROVE_TESTS=t/058_vm_truncate_invalid_pages.pl
>
> The test holds a standby restartpoint, promotes that standby, deletes and vacuums an all-visible table (truncating
itsheap and VM forks), rejoins the old primary as a standby, and restarts it. On 18.6, 20/20 TAP runs failed with this
PANIC;on 18.4, 20/20 passed. A shell reproducer, exact steps, WAL excerpts, version matrix, and two candidate patches
areincluded in the attachment. 
>
> The affected 18.6 build is REL_18_6 (724edf9); 17.11 is REL_17_11 (083ac03). The runs were on macOS 15 / Darwin
24.6.0,arm64, clang 21.1.8, with debug and injection points enabled and without cassert. The original symptom was on
Linuxin a three-node streaming setup. The analysis in REPORT.md points to VM block reads without an FPI during redo
afteran already replayed VM truncation; that diagnosis and the candidate fixes are provided for review. 
>
> Regards,
> Jacky Nguyen

Hi!
Thanks for report. Looks like this bug existed since first VM commit,
which was already missing invalid page guard from [0].
So, case from [0] reintroduces one we started to register VM pages.

Also I prefer candidate fix b from your email

[0] https://github.com/postgres/postgres/commit/defe93463c69f8e0bb717294a34d67c34ac0b03f


--
Best regards,
Kirill Reshke



pgsql-bugs by date:

Previous
From: Jacky Nguyen
Date:
Subject: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
Next
From: Palak Chaturvedi
Date:
Subject: Re: BUG #19701: GIN trigram index loses rows at similarity_threshold 0