Hi Yuriy,
Thanks for v2. All my comments are addressed, and I'm marking the
entry Ready for Committer.
Testing (macOS arm64, meson debug build with assertions, master at
9510a826e4 plus v2):
- Both patches apply cleanly with git am, build without warnings, and
the regression tests pass.
- test_xlogreader passes. Without 0001 it fails the four malformed
cases and still accepts the three valid ones, so it does catch the
bug.
- Crash recovery with wal_compression = off, pglz, lz4 and zstd, each
replaying about 1,000 full-page images with holes, plus a pglz run
with wal_consistency_checking = all (about 3,500 images). The data
matched before and after the crash, amcheck found nothing, and the
new Asserts never fired.
- Records with the hole moved past the page, one uncompressed and one
pglz-compressed: pg_waldump reports the BKPIMAGE_HAS_HOLE error,
pg_waldump --save-fullpage exits with an error instead of crashing,
and crash recovery stops at the bad record.
One small suggestion for 0002: the uncompressed cases don't test the
boundary. hole_offset == bimg_len (accepted, the hole ends exactly at
BLCKSZ) and hole_offset == bimg_len + 1 (rejected) would match what
the compressed cases already do.
The same runs show the LSN problem again: for the corrupted record at
0/01E582A0, pg_waldump reported 0/01E58218, the previous record, and
recovery reported 0/01E581B8. I'll send that patch in a separate
thread.
The crafting script is attached. It finds the first full-page image
with a hole after a given LSN in a copy of pg_wal, moves the hole past
the end of the page, and recomputes the record CRC.
Regards,
Rahul Yadav