On Sat, Sep 26, 2026 at 12:25 PM Manu <manuelreyesbravo@gmail.com> wrote:
>
> Hi Nikhil,
>
> > v4 attached, rebased on e27f3b2cad7 and renumbered now that the first
> > patch is in:
>
> I ran some checks and numbers on v4, on master 3c5d9d914fa. Scripts
> and the full tables are attached.
>
> On-disk format. The cover letter says pglz and lz4 values keep their
> representation bit for bit. Writing the same data with master and
> with v4, the raw bytes of the inline compressed datums (pageinspect)
> and of the TOAST chunks hash the same, for pglz and lz4. pg_upgrade
> from master to v4 keeps every value, and verify_heapam(check_toast)
> finds nothing. It is also clean on v4 for a column holding pglz, lz4
> and zstd values at once, and after VACUUM FULL, UPDATE and DELETE on
> a zstd table.
>
Thanks a lot for checking both, and for the scripts. Those are the
two things the format patch must not break.
> The cost is on reads of small values. For 50,000 JSON values of
> 4 kB, reading them all takes 140 ms with zstd against 53 ms with lz4,
> and a 100-byte prefix costs the same as the whole value (140 ms,
> against 23 ms with lz4).
>
The prefix part comes from the format: a zstd block holds up to 128 kB
and can only be decoded whole, so for any smaller value a prefix is a
full decode. v7 has a comment saying so above the slice loop.
> One thing in 0004 that is easy to improve:
> zstd_decompress_datum_slice() creates and frees a ZSTD_DCtx for every
> value. The attached diff keeps one per backend, reset on each use,
> and uses it in both decompression paths. zstd, 4 rounds, ratio of
> the diff to v4:
>
I agree reusing the contexts is worth doing; zstd itself recommends
it, and ZSTD_compress() allocates one per call too, so the write side
would gain as well. For this series I'd rather stay with what the
tree already does: the WAL code uses the one-shot ZSTD_compress() and
ZSTD_decompress() calls, and 0003 matches that. A context kept for
the backend's lifetime also brings its own questions (how much memory
it may keep, and when to release it), which I'd rather settle
separately than attach to the format change.
So I'll propose it as a follow-up once this series is in, starting
from your diff and your numbers, if that works for you.
>
> Regards,
> Manu
--
Nikhil Veldanda