Hi Andres,
Thanks a lot for the review, and first, please ignore the previous email; it has bad formatting = Gmail tricked me :(.
> Of course it'd be even better if would be to stop counting the same stuff in
> pgBufferUsage.{shared,local}_blk_{read,write}_time and
> pgStatBlock{Read,Write}Time. That's pretty darn silly.
>
> pgstat_update_dbstats() should just keep a pgBufferUsage snapshot from the
> last report and add up the relevant pgBufferUsage fields. And vacuum & analyze
> already diff, so they just would need to add the fields to get the same
> results.
We tried to implement this and directly use `pgBufferUsage` in `pgstat_update_dbstats()` but we came across some issues that caused pg_stat_database to overcount, specifically when having parallel workers in the plan, each worker will acuumulate its own statistics in its own `pgBufferUsage` and will report them at the end and update the db stats entry, but the problem is after that the parent backend will also acuumulate all the statistics from its parallel workers and add them to its `pgBufferUsage` and update the db stats entry again hence causing each worker's I/O write/read times to be counted twice.
Meanwhile using `pgStatBlock{Read,Write}Time` in `pg_stat_count_io_op_time()` will count every I/O timing a single time, without having to introduce changes in the parallel worker infrasturcture to handle the above case? Maybe thats why `pgStatBlock{Read,Write}Time` was introduced in the beginning?
Or else, would it be the right approach to look into changing the existing infastructure for I/O time counting in the context of parallel workers in your opinion?
What do you think? Are we missing something?
Best Regards,
Ahmed Gouda and Bernd Reiß