Hi,
At Databricks, we noticed a weird plan in one of our TPCC benchmarks
on pg18 (and reproduced on master), which was basically as follows:
Gather
NestedLoop
NestedLoop
Parallel Seq Scan (Filter: <primary key = constant>)
BitmapScan
BitmapScan
At face value that plan looks OK, but when you look at it in detail
it's quite a strange plan: The outermost plan is expected to only
produce one row. Joins can generally only produce tuples once rows are
produced on the outer side, so the paralellism is wasted: only a
single worker will find a row, and thus only a single worker will
execute the joins. Given that cap on parallelism, an index scan on
the primary key would've been much cheaper.
And that is indeed true, but due to the paralellism the inner nodes of
the nested loop joins were discounted by a factor of about n_workers.
This means that, in effect, we prefer parallel seqscans over unique
row lookups in indexscans, as long as the planner launches sufficient
workers to apply sufficient discount to the partial join plans.
This means that whilst all subsequent plan nodes inside the parallel
portion of the plan will be costed as if they're processing data in
parallel, only one of workers will ever produce tuples in the parallel
seqscan, and get to executing the other paralellized plan nodes. So,
the only part that actually benefits from parallelization here is the
seqscan itself (which is costed correctly) whilst the other plan nodes
are all significantly discounted.
Attached is a patch that introduces Path->effective_workers, which
attempts to account for this issue. I've attached a reproducer in
scratch_70.sql. It generates a ~large database (~19GB); I haven't
worked on minimizing its data size. The reproducer is derived from
HammerDB's TPCC schema & initializer script, removing the parts from
the schema that the query didn't look at, adding fillfactor padding
instead. On master, it scans the "district" table in parallel,
looking for the one PK-matching row; with the patch it uses a
non-parallel PK scan for the lookup in the district table.
Kind regards,
Matthias van de Meent
Databricks