On Wed, 30 Sept 2026 at 07:37, Haibo Yan <tristan.yim@gmail.com> wrote:
>
> Hi Matthias,
>
> I spent some time testing v1 and found two issues in the current
> `effective_workers` propagation.
>
> The first one looks like a correctness bug.
>
> `subpath_adjusted_effective_workers()` can see `subpath->parent->rows == 0`
> while building paths for `UPPERREL_PARTIAL_GROUP_AGG`. With leader
> participation enabled this can make `effective_workers` negative.
>
> For example, with a four-worker Partial Aggregate followed by Sort/Group,
> the calculation can end up with:
Thanks, I've adjusted this in the latest version.
> best_row_estimate = 0
> effective_workers = ceil(0) - 1 = -1
>
> and `get_parallel_divisor()` then returns 0.3. With leader participation
> disabled the corresponding value is 0, and the divisor becomes 0.
>
> That feeds into `compute_gather_rows()`, so I was able to get cases where
> a partial path producing 100 rows per participant was reconstructed as a
> Gather Merge producing only 30 rows with leader participation enabled,
> or 1 row with it disabled.
>
> The second issue is a discontinuity around the low-cardinality threshold.
>
> With four planned workers and leader participation enabled I get:
>
> estimated rows effective_workers divisor
> 2 1 1.7
> 3 2 2.4
> 4 3 3.1
> 5 4 4.0
>
> In a parameterized Nested Loop reproducer, that makes the estimated
> partial output decrease when the global row estimate increases:
>
> R=4: partial NL rows = 1290, total cost = 23659.35
> R=5: partial NL rows = 1250, total cost = 23559.37
>
> The R=5 case has 25% more qualifying outer rows and 25% more final output,
> but is costed lower.
Yes, that's presumably because the parallel portion of the plan is
spread across more workers, and the join portions are otherwise costed
equivalently (although in this case it expects fewer rows per worker
total).
And, note that each worker up to 3 adds 0.7 to the divisor, whilst the
4th worker adds .9, and every worker after that contributes 1.0. This
is just how the divisor is calculated, and changing that is not part
of what I'm planning to do.
--------------------
Attached is v2, which changes a few things:
* Changed types of the Path->*_workers fields to int16.
This is primarily to avoid size changes, and safe because we don't
support values larger than 1024.
* Updated the calculations of subpath_adjusted_effective_workers to
have better accounting in certain cases
This fixes the underflow issues mentioned by Haibo Yan.
* Added effective_worker to create_append_path.
This was needed to avoid otherwise unrelated plan changes in the
regression tests.
Kind regards,
Matthias van de Meent
Databricks (https://www.databricks.com)