Hi,
On 2026-10-02 20:03:24 +0300, ahmed wrote:
> Thanks a lot for the review, and first, please ignore the previous email;
> it has bad formatting = Gmail tricked me :(.
>
> > Of course it'd be even better if would be to stop counting the same stuff
> in
> > pgBufferUsage.{shared,local}_blk_{read,write}_time and
> > pgStatBlock{Read,Write}Time. That's pretty darn silly.
> >
> > pgstat_update_dbstats() should just keep a pgBufferUsage snapshot from the
> > last report and add up the relevant pgBufferUsage fields. And vacuum &
> analyze
> > already diff, so they just would need to add the fields to get the same
> > results.
Just to be clear, I don't think that should be done in the same commit, that'd
be separate work imo.
> We tried to implement this and directly use `pgBufferUsage` in
> `pgstat_update_dbstats()` but we came across some issues that caused
> pg_stat_database to overcount, specifically when having parallel workers in
> the plan, each worker will acuumulate its own statistics in its own
> `pgBufferUsage` and will report them at the end and update the db stats
> entry, but the problem is after that the parent backend will also acuumulate
> all the statistics from its parallel workers and add them to its
> `pgBufferUsage` and update the db stats entry again hence causing each
> worker's I/O write/read times to be counted twice.
I really really dislike a lot of how the parallel worker stats stuff is done
:(. The whole idea of just mixing it into the client backends stats is just
bonkers IMO. We do it inconsistently, sometimes we then have to back out the
worker numbers again during explain. Just a bad idea overall.
I think Lukas had some WIP patches to change some of that? He might also have
some opinions independent of parallel workers, as he has pending work to
refactor things so that the explain stuff is not diff based anymore (the
diffing makes it rather expensive atm).
So maybe this is all moot - I'll let him chime in on that.
Greetings,
Andres Freund