Subject: Re: PG18: use-after-free in exec partition pruning after an EPQ recheck in LockRows - Mailing list pgsql-bugs

From Victor Rusaleev
Subject Subject: Re: PG18: use-after-free in exec partition pruning after an EPQ recheck in LockRows
Date
Msg-id CAA_xEx409dLbZh5DM6U5wa3S=ZaFTh+7=34xL+vt0H7_CdkRbA@mail.gmail.com
Whole thread
List pgsql-bugs
In-Reply-To: <CAP2G_gGpV=rKMLuET4RW-qO41GvBvi6Czsjsj16eT9hAwHxG6w@mail.gmail.com>
References: <CAP2G_gGpV=rKMLuET4RW-qO41GvBvi6Czsjsj16eT9hAwHxG6w@mail.gmail.com>

Hi,

Independent confirmation on a packaged release build, with a different
application query, plus a measured workaround.

Environment: PostgreSQL 18.6 (Debian package 18.6-1.pgdg13+2), x86_64.
JIT is not available on this build (llvmjit not installed).

Query shape: an outbox relay claims rows of a range-partitioned table
(5 partitions on created_at) in one statement:

  WITH RECURSIVE keys AS (...),
  heads AS (... CROSS JOIN LATERAL (... LIMIT $2) ...),
  locked AS MATERIALIZED (
      SELECT ... FROM outbox AS candidate JOIN heads ON <primary key>
      WHERE <unpublished, lease expired>
      ORDER BY ... LIMIT $1
      FOR UPDATE OF candidate SKIP LOCKED),
  candidates AS (SELECT ... FROM locked ...),
  claimed AS (
      UPDATE outbox SET status = 'PUBLISHING', lease_until = ...
      FROM candidates
      WHERE <primary key, including created_at>
      RETURNING ...)
  SELECT ... FROM claimed ...;

Several relay processes run this statement at the same time against the
same table.

Backtrace, identical in five core dumps (postgresql-18-dbgsym):

  #0  ExecEvalExprSwitchContext (state=0x64, ...)     executor.h:440
  #1  partkey_datum_from_expr (stateidx=3)            partprune.c:3823
  #2  perform_pruning_base_step                       partprune.c:3503
  #3  get_matching_partitions                         partprune.c:881
  #4  find_matching_subplans_recurse (initial_prune=false) execPartition.c:2592
  #5  ExecFindMatchingSubPlans                        execPartition.c:2535
  #6  choose_next_subplan_locally                     nodeAppend.c:583
  #7  ExecAppend                                      nodeAppend.c:330
  #9  ExecNestLoop                                    nodeNestloop.c:159
  #11 ExecModifyTable (CMD_UPDATE)                    nodeModifyTable.c:4280
  #13 CteScanNext                                     nodeCtescan.c:103

In two of the five dumps frame #0 is "??": a call through a freed function
pointer. es_epq_active is NULL at the crash. pprune->exec_context.planstate
and .exprstates point into memory allocated after the parent EState; the
"planstate" there has type 0xffffffff. This matches the analysis in the
thread: EvalPlanQualStart() shares es_part_prune_states, and
InitExecPartitionPruneContexts() re-points the shared exec_context at EPQ
memory that EvalPlanQualEnd() then frees.

Measured matrix. Each trial claims 1M rows; trial order interleaved; crash
means a backend terminated by signal 11:

  enable_partition_pruning = on,  3-4 concurrent sessions: 4 of 4 crashed
                                                          (after 95-283 s)
  enable_partition_pruning = on,  2 concurrent sessions:   0 of 2 crashed
  enable_partition_pruning = off, 2-4 concurrent sessions: 0 of 6 crashed

One session alone never crashed.

Workaround in production use until a fixed release:

  ALTER ROLE <relay_role> IN DATABASE <db> SET enable_partition_pruning = off;

With 5 partitions the claim throughput did not drop. This may help users
who wait for the fix, and may be worth a line in the release notes.

I did not test the attached patch.

Regards,
V.R.



pgsql-bugs by date:

Previous
From: shihao zhong
Date:
Subject: Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
Next
From: PG Bug reporting form
Date:
Subject: BUG #19742: `INTERSECT` under a `UNION ALL` with an empty arm fails with "could not find pathkey item t"