Re: Protocol Compression (fourth attempt) - Mailing list pgsql-hackers

From Anthonin Bonnefoy
Subject Re: Protocol Compression (fourth attempt)
Date
Msg-id CAO6_XqrhOnpon4pg+aceN_-Z+qOAW_Xp1DsmgDUQz6+wiPMwBA@mail.gmail.com
Whole thread
In response to Re: Protocol Compression (fourth attempt)  (Andrey Borodin <x4mmm@yandex-team.ru>)
List pgsql-hackers
On Wed, Sep 30, 2026 at 9:53 AM Andrey Borodin <x4mmm@yandex-team.ru> wrote:
> Would you mind if I prepared that revision of your patchset? It would
> help me understand your implementation better and give us a concrete
> starting point for combining our work. I'd focus on the simplifications
> we agree on.
>
> If you've already started on v2, let me know so we can split the work.

Sounds good to me, I haven't started the v2 yet.

> On enabling and disabling compression within a session, I agree that
> PQcommMethods makes the switch simple. But a USERSET GUC must handle
> SET LOCAL and rollback while output is buffered or a frame is open.
> Could we keep the choice at connection startup for v1? Ordinary messages
> would still be allowed on a compressed connection, so we could add
> session-level control later without changing the wire format. Is there a
> use case where choosing at startup would not be enough?

That's currently handled by flushing + pq_send_messages which closes
the frame and sends the buffered messages. I imagine that on some
workloads (table with mostly random data), compression would have
mostly negative effects and users may want to disable it for specific
queries. For a v1, that's fine to leave this out.

> On thresholds and batch sizes, I agree that the knobs are useful for
> experiments. Your first-packet timings show that the flush policy needs
> work. I'd first try to address that internally, for example by bounding
> the amount of uncompressed input processed before flushing output.
> Publishing these thresholds as GUCs would mean supporting their
> semantics as we change the buffering policy. Could we keep them in a
> benchmarking patch for now, and add public controls if measurements
> show a trade-off that users need to choose themselves?

Sounds reasonable to me. Using the input size to trigger flush was
something I had in mind, but that wouldn't be useful if the messages
are incomplete as the client won't start processing the message until
it is fully available. Maybe a combination of both (limited input size
with at least x full messages) could work? Adding it and benchmarking
it would definitely help to see the impact.

Regards,
Anthonin Bonnefoy



pgsql-hackers by date:

Previous
From: Nazir Bilal Yavuz
Date:
Subject: Re: aio: Async fsyncs for crash recovery and checkpointer
Next
From: Hannu Krosing
Date:
Subject: Re: Direct TOAST v2, faster, smaller and no migration needed