Re: aio: Async fsyncs for crash recovery and checkpointer - Mailing list pgsql-hackers

From Nazir Bilal Yavuz
Subject Re: aio: Async fsyncs for crash recovery and checkpointer
Date
Msg-id CAN55FZ2i-=87NWg8v66hbbm18o+kBmCjyLyX4fE7s49c9JbEBw@mail.gmail.com
Whole thread
In response to Re: aio: Async fsyncs for crash recovery and checkpointer  (Nitin Jadhav <nitinjadhavpostgres@gmail.com>)
List pgsql-hackers
Hi,

On Tue, 22 Sept 2026 at 18:21, Nitin Jadhav
<nitinjadhavpostgres@gmail.com> wrote:
>
> I found two further points, and the worker-side SLRU test-coverage
> question from v1 still appears applicable.
>
> An I/O worker can skip a required fsync after fsync is enabled. As
> Yuhang noted, there appears to be a race when reloading fsync from off
> to on. pgaio_io_start_fsync()  makes the dispatch decision using the
> issuing process's enableFsync.

You are right, that is fixed in v3 [1]. I used the issuing process's enableFsync


> Worker SMGR cleanup appears dependent on the worker becoming idle.
> Patch 0002 calls smgrdestroyall() only after the worker finds that no
> request is available. This means a continuously busy worker may never
> perform the cleanup. Worker-side relation reopening calls smgropen(),
> and those unpinned SMgrRelation objects remain in the worker's SMGR
> hash until smgrdestroyall() is called. With sustained I/O over many
> distinct relations, a worker whose queue never becomes empty could
> therefore retain an increasing number of SMGR entries, including
> entries for relations that have since been dropped. Would it be safer
> to check FirstCallSinceLastCheckpoint() at a safe point after
> completing each request, before consuming the next request, rather
> than only on the idle path? At that point any descriptor reopened for
> the completed operation has already been released.

We have created another thread for this. Now, we have a dedicated SMGR
entry limits, and the worker can't surpass this limit. For more
information you can look [2], but briefly, SMGR entries can't surpass
this limit when the worker is busy; they are cleaned when we reach
this limit. Also, SMGR entries are cleaned before the worker goes to
idle. These should solve the problem.


> The worker-side SLRU path does not appear to have targeted test
> coverage. The test_slru_page_sync() still registers the test SLRU with
> SYNC_HANDLER_NONE. Consequently, SlruSyncFileTag() selects
> PGAIO_TID_SYNC, which has no reopen callback and is executed
> synchronously in the submitting process under io_method=worker.

Yes, that is not tested. I will add that.

[1] https://postgr.es/m/CAN55FZ2NT6529QhB_dcrZ5LaSL20aayhP33ZisZVFSg2hOqLXg%40mail.gmail.com
[2] https://postgr.es/m/CAN55FZ2BesKUnajdgpw1fPSe3S6_CHOugryaUEtD7vdP%3DdRKEQ%40mail.gmail.com

--
Regards,
Nazir Bilal Yavuz
Microsoft



pgsql-hackers by date:

Previous
From: Nazir Bilal Yavuz
Date:
Subject: Re: aio: Async fsyncs for crash recovery and checkpointer
Next
From: Anthonin Bonnefoy
Date:
Subject: Re: Protocol Compression (fourth attempt)