I ran into the same pg_current_wal_insert_lsn() page-boundary problem discussed in this thread, independently, outside the regression tests, while evaluating WAIT FOR for read-your-writes on a standby using a custom coded application to test the usability of this feature.
At this boundary, pg_current_wal_insert_lsn() returns the next insertion address, which is after the header gap. The test thus waited for a position beyond the existing WAL, requiring another WAL record to make progress. The unrelated progression comes from the primary's background writer who generated a RUNNING_XACTS record, advancing WAL to 0/0301A050, and then notified the waiters.
About 1 in 2000 waits never succeeded and ran to the full TIMEOUT
Since all my sessions were using WAIT FOR LSN, it was a complete stall, whenever it happens.
Interestingly, is this a pressure test? Can you please elaborate a bit more on how the application tests it? The stall in the TAP test 049 persisted for ~8s because of no WAL activity in that test env, and it was saved by a later RUNNING_XACTS record generated by bgwriter. In practice, I expected the problem to be harder to notice since WAL is more active.
A client side workaround similar to the following is a temporary solution for me for my tests:
CREATE FUNCTION wal_insert_end_lsn(l pg_lsn DEFAULT pg_current_wal_insert_lsn()) RETURNS pg_lsn LANGUAGE sql STABLE AS $$ SELECT CASE WHEN (l - '0/0'::pg_lsn) % pg_size_bytes(current_setting('wal_segment_size')) < current_setting('wal_block_size')::int -- first page of a segment THEN CASE WHEN (l - '0/0'::pg_lsn) % current_setting('wal_block_size')::int = 40 THEN l - 40 ELSE l END ELSE CASE WHEN (l - '0/0'::pg_lsn) % current_setting('wal_block_size')::int = 24 THEN l - 24 ELSE l END END $$;
A server-side fix would be better
Could WAIT FOR itself also tolerate such targets? When the target points exactly past a page header treat it as the page boundary.
I think we better fix the page/segment header issue on the server side.