Re: enhancing pg_basebackup speeds up to ~23Gbps (small fixes + io_uring/Direct I/O) - Mailing list pgsql-hackers

From Jakub Wartak
Subject Re: enhancing pg_basebackup speeds up to ~23Gbps (small fixes + io_uring/Direct I/O)
Date
Msg-id CAKZiRmxEha69zr_prZbbSi=uzzu9pd9o1kHBfyY872HcNGMuog@mail.gmail.com
Whole thread
In response to Re: enhancing pg_basebackup speeds up to ~23Gbps (small fixes + io_uring/Direct I/O)  (Gustavo William <gustavowilliam0805@gmail.com>)
List pgsql-hackers
On Tue, Sep 22, 2026 at 9:40 PM Gustavo William
<gustavowilliam0805@gmail.com> wrote:
>
> Hi Jakub,
>
> Thanks for reviewing it. I'm resending the chain of patches here, with 0005 already merged + your suggestions
applied.
[..]

Hi Gustavo,

TLDR; I've verified ~9.4x improvement with 0005-incremental-backup just for
the read path on the NVMe (by avoiding stalled pread()). Good job :)

Testing method:
  initdb # summarize_wal=on , maintenance_io_concurrency=2, ...
  pgbench -i -s 500 --fillfactor=100 # ~7.5GB base/
  pg_basebackup -D /tmp/full -c fast -v
  pgbench -N -c 8 -j 8 -t 10000 postgres # -N=updates
  psql -c "CHECKPOINT" postgres
  sudo /usr/local/bin/drop_fs_cache.sh
  time pg_basebackup -i /tmp/full/backup_manifest --target=server-blackhole \
     -Xnone -c fast --no-sync --no-verify-checksums

Avg of 3 runs (drop_caches + time pg_basebackup -i)
  Without 0005: 7.53s
  With 0005: 0.80s (note it's just reading due to blackhole, so no writes!)

I don't have access to SATA drive, but I suspect it would be way better
there.

I was kind of sceptical of the result at first, but explanation of the
phehomena seems to be like this: basebackup_incremnetal.c has
GetFileBackupMethod()->qsort() so it sorts the blocks to be read. Without 0005
patch, preads() of increasing block numbers (offsets) all cause hit page cache
misses, and apparently - at least here on kernel 6.17.x, but earlier for sure
too - this seem to kick off the readahead (because kernel's readahead seems to
be implemented to only kick in by by page-cache misses , it won't even start
if there are hits, kind of makes sense), so pure preads() tend to overread a
lot (like reading N-times more data from a a 1GB segment than we need).

I think the kernel's assumption is that reading sequentially is cheaper than
reading one by one, but depending on density of changes we often - at least
in this cenario, that is in incremetanl basebackup mode - we read much less
than the kernel reads with readahead (simply said: it overreads a lot!).

So using fadvise(FADV_WILLNEED) here caused kind of disarms such aggressive
readahead completley (as it results in 100% hit to pagecaches, so readahead is
not even activated), and this causes way faster incremental times, so fast
that I'm had to actually reserach why and how it happens :o

Issues with the patch
- it seems maintenance_io_concurrency=1 disables prefetch? (it's not queue
  depth, it's more of how many to prefetch, so "1" means IMHO pread() +
  one posix_fadvise() or am I wrong?)
- pgindent might be needed.

-J.



pgsql-hackers by date:

Previous
From: Cagri Biroglu
Date:
Subject: Re: Per-table resync for logical replication subscriptions
Next
From: Peter Smith
Date:
Subject: Re: [PATCH] Table sync race with REFRESH PUBLICATION