Hi,
I did some additional testing of v4 and found three issues when parallel
initialization is combined with range partitioning:
1. --tablespace is not applied to the standalone child partitions,
so they are created in the database's default tablespace.
2. Running -I G after -I t fails because the t step does not create the child partitions.
3. Re-running -I g fails because createStandalonePartitions() attempts to
create child tables that already exist.
I reproduced issues 2 and 3 with the following commands, respectively:
```
pgbench -i -I t -j 2 --partitions=4 --partition-method=range postgres
pgbench -i -I G -j 2 --partitions=4 --partition-method=range postgres
```
```
pgbench -i -I t -j 2 --partitions=4 --partition-method=range postgres
pgbench -i -I g -j 2 --partitions=4 --partition-method=range postgres
pgbench -i -I g -j 2 --partitions=4 --partition-method=range postgres
```
One possible approach would be to use a scheme similar to v3: create the parent
and child partitions in the t step, then have each worker TRUNCATE its assigned
partition and run COPY ... WITH (FREEZE) on it in the same transaction during the g step.
In a simple test, TRUNCATE operations on different leaf partitions did not block
one another. TRUNCATE followed by COPY FREEZE also worked on an already attached partition.
Would this approach preserve both the independence of the t, g, and G steps
and the use of COPY FREEZE?
Koshi Shibagaki
FUJITSU LIMITED
https://www.fujitsu.com/
________________________________________
差出人: lakshmi <lakshmigcdac@gmail.com>
送信日時: 2026年5月9日 17:02
宛先: Mircea Cadariu
CC: Kuroda, Hayato/黒田 隼人; PostgreSQL Hackers; tomas@vondra.me; Heikki Linnakangas
件名: Re: parallel data loading for pgbench -i
On Fri, May 8, 2026 at 11:41 PM Mircea Cadariu <cadariu.mircea@gmail.com<mailto:cadariu.mircea@gmail.com>> wrote:
Hi Lakshmi and Hayato,
Thanks a lot for your feedback.
Attached for your consideration is v4, in which I address your remarks.
Hi Mircea, Hayato,
I tested the v4 patch on 19devel with a few different thread/partition combinations.
The updated API looks much better now. I verified that:
* parallel loading works correctly with -j
* uneven partition distribution (for example 5 partitions with 2 threads) also works fine
* serial mode with -j 1 works again as expected
The workers appear to run concurrently, and VACUUM time remains relatively small in my tests.
Overall, the new approach looks much cleaner and more flexible compared to the earlier versions.
Thanks again for the update.
Best regards,
Lakshmi