In-Reply-To: <CAP2G_gGpV=rKMLuET4RW-qO41GvBvi6Czsjsj16eT9hAwHxG6w@mail.gmail.com>
References: <CAP2G_gGpV=rKMLuET4RW-qO41GvBvi6Czsjsj16eT9hAwHxG6w@mail.gmail.com>
Hi,
Independent confirmation on a packaged release build, with a different
application query, plus a measured workaround.
Environment: PostgreSQL 18.6 (Debian package 18.6-1.pgdg13+2), x86_64.
JIT is not available on this build (llvmjit not installed).
Query shape: an outbox relay claims rows of a range-partitioned table
(5 partitions on created_at) in one statement:
WITH RECURSIVE keys AS (...),
heads AS (... CROSS JOIN LATERAL (... LIMIT $2) ...),
locked AS MATERIALIZED (
SELECT ... FROM outbox AS candidate JOIN heads ON <primary key>
WHERE <unpublished, lease expired>
ORDER BY ... LIMIT $1
FOR UPDATE OF candidate SKIP LOCKED),
candidates AS (SELECT ... FROM locked ...),
claimed AS (
UPDATE outbox SET status = 'PUBLISHING', lease_until = ...
FROM candidates
WHERE <primary key, including created_at>
RETURNING ...)
SELECT ... FROM claimed ...;
Several relay processes run this statement at the same time against the
same table.
Backtrace, identical in five core dumps (postgresql-18-dbgsym):
#0 ExecEvalExprSwitchContext (state=0x64, ...) executor.h:440
#1 partkey_datum_from_expr (stateidx=3) partprune.c:3823
#2 perform_pruning_base_step partprune.c:3503
#3 get_matching_partitions partprune.c:881
#4 find_matching_subplans_recurse (initial_prune=false) execPartition.c:2592
#5 ExecFindMatchingSubPlans execPartition.c:2535
#6 choose_next_subplan_locally nodeAppend.c:583
#7 ExecAppend nodeAppend.c:330
#9 ExecNestLoop nodeNestloop.c:159
#11 ExecModifyTable (CMD_UPDATE) nodeModifyTable.c:4280
#13 CteScanNext nodeCtescan.c:103
In two of the five dumps frame #0 is "??": a call through a freed function
pointer. es_epq_active is NULL at the crash. pprune->exec_context.planstate
and .exprstates point into memory allocated after the parent EState; the
"planstate" there has type 0xffffffff. This matches the analysis in the
thread: EvalPlanQualStart() shares es_part_prune_states, and
InitExecPartitionPruneContexts() re-points the shared exec_context at EPQ
memory that EvalPlanQualEnd() then frees.
Measured matrix. Each trial claims 1M rows; trial order interleaved; crash
means a backend terminated by signal 11:
enable_partition_pruning = on, 3-4 concurrent sessions: 4 of 4 crashed
(after 95-283 s)
enable_partition_pruning = on, 2 concurrent sessions: 0 of 2 crashed
enable_partition_pruning = off, 2-4 concurrent sessions: 0 of 6 crashed
One session alone never crashed.
Workaround in production use until a fixed release:
ALTER ROLE <relay_role> IN DATABASE <db> SET enable_partition_pruning = off;
With 5 partitions the claim throughput did not drop. This may help users
who wait for the fix, and may be worth a line in the release notes.
I did not test the attached patch.
Regards,
V.R.