remote_apply commit hangs when wal_receiver_status_interval = 0 - Mailing list pgsql-hackers

From Vaibhav Dalvi
Subject remote_apply commit hangs when wal_receiver_status_interval = 0
Date
Msg-id CA+vB=AHKUhc9GCDwoLr=SshTBCQG09kp8TEXew50UtD5rYqowg@mail.gmail.com
Whole thread
Responses Re: remote_apply commit hangs when wal_receiver_status_interval = 0
List pgsql-hackers
Hi hackers,

On master, a commit with synchronous_commit = remote_apply does not
return when the standby has wal_receiver_status_interval = 0. The
standby applies the commit, but it never sends the apply reply to the
primary, so the backend keeps waiting in SyncRep. With default
timeouts it waits about 30 seconds, until the walsender keepalive
forces a reply. With wal_sender_timeout = 0 and wal_receiver_timeout
= 0 it waits forever.

The docs for wal_receiver_status_interval say that updates are sent
while ignoring this parameter "when synchronous_commit is set to
remote_apply", so I think this is a regression.

Steps to reproduce:

     initdb -D pg20_primary -U postgres -A trust
    cat >> pg20_primary/postgresql.conf <<EOF
    port = 5521
    wal_sender_timeout = 0
    EOF
    pg_ctl -D pg20_primary -l primary.log -w st
    psql -p 5521 -U postgres -c "CREATE TABLE t (id int, note text);"

    pg_basebackup -h 127.0.0.1 -p 5521 -U postgres -D pg20_standby -R -d "application_name=standby_5522"
    cat >> pg20_standby/postgresql.conf <<EOF
    port = 5522
    wal_receiver_status_interval = 0
    wal_receiver_timeout = 0
    hot_standby_feedback = off
    EOF
    pg_ctl -D pg20_standby -l standby.log -w start

    psql -p 5521 -U postgres \
        -c "ALTER SYSTEM SET synchronous_standby_names = 'standby_5522';"
        -c "SELECT pg_reload_conf();"


The standby is streaming and shows as sync:

     application_name |   state   | sync_state
    ------------------+-----------+------------
     standby_5522     | streaming | sync


Now run a remote_apply commit on the primary:

     postgres=# SET synchronous_commit = remote_apply;
    postgres=# \timing on
    postgres=# INSERT INTO t VALUES (1, 'remote_apply');
    ^CCancel request sent
    WARNING:  canceling wait for synchronous replication due to user request
    DETAIL:  The transaction has already committed locally, but might
    not have been replicated to the standby.
    INSERT 0 1
    Time: 184888.783 ms (03:04.889)

While the INSERT was stuck, the backend was waiting in SyncRep:

       pid   | wait_event_type | wait_event |     waiting
    ---------+-----------------+------------+--
     2947177 | IPC             | SyncRep    | 00:00:51.399588


But the row was already visible on the standby:

     id |     note
    ----+--------------
      1 | remote_apply


The primary's replay_lsn in pg_stat_replication stays at the value
from connection time. So the standby has applied the commit, but the
primary never learns about it.

Root cause:

Commit 400a790a48e changed the apply notification call in the
walreceiver main loop from

    XLogWalRcvSendReply(true, false);

to

    XLogWalRcvSendReply(false, false, true);  

But XLogWalRcvSendReply() still has this early return at the top:

    if (!force && wal_receiver_status_interval <=0)
        return;


So with the interval set to 0, the apply reply is dropped before the
new checkApply logic is reached. The same happens for the reply
requested when the startup process has replayed all the WAL
received so far.

Fix:

The attached patch skips the early return when checkApply is set:

    -   if (!force && wal_receiver_status_interval <= 0)
   +   if (!force && !checkApply && wal_receiver_status_interval <= 0)

With interval 0 the reply wakeup is infinity, so the call goes into
the existing duplicate check, and the reply is sent only when the
apply position advances. This keeps the purpose of 400a790a48e,
which was to avoid sending duplicate replies.

The patch also adds a TAP test,
src/test/recovery/t/058_remote_apply_status_interval.pl. It fails on
master because the commit never completes, and passes with the fix.

Thoughts?

Regards,
Vaibhav Dalvi
EnterpriseDB
Attachment

pgsql-hackers by date:

Previous
From: Andrey Borodin
Date:
Subject: Re: Protocol Compression (fourth attempt)
Next
From: Vlad Lesin
Date:
Subject: Re: ReplicationSlotRelease() clobbers another backend's statusFlags entry