Re: [Patch]The Case For WAL-Logging pg_upgrade - Mailing list pgsql-hackers
| From | Bohyun Lee |
|---|---|
| Subject | Re: [Patch]The Case For WAL-Logging pg_upgrade |
| Date | |
| Msg-id | CAMPh8MrRP+C4HHvhgrqDX=8Q7NGfkjAiSuSEpt1qiP75Z-kKkw@mail.gmail.com Whole thread |
| In response to | Re: [Patch]The Case For WAL-Logging pg_upgrade (John Naylor <johncnaylorls@gmail.com>) |
| List | pgsql-hackers |
Hi all,
Thanks everyone for all the feedback. I thought about how to get --wal-upgrade
right over the past couple of months and prepared a second patch. The main
changes compared to the first patch:
- Automated the handoff during pg_upgrade, right before the old cluster primary
shuts down. Removed the separate handoff subcommand.
- Refactored the relink WAL record to capture and verify the correctness of the
pre-/post-upgrade storage state. Incorrect storage state fails the upgrade.
Standbys use it to replay the operations per relation/directory.
- Extended the patch to properly support cascaded standbys. Cascaded standbys
also ingest upgrade WALs like normal replication.
Github repo: https://github.com/LeeBohyun/postgres/blob/wal-upgrade-patch/
Thanks everyone for all the feedback. I thought about how to get --wal-upgrade
right over the past couple of months and prepared a second patch. The main
changes compared to the first patch:
- Automated the handoff during pg_upgrade, right before the old cluster primary
shuts down. Removed the separate handoff subcommand.
- Refactored the relink WAL record to capture and verify the correctness of the
pre-/post-upgrade storage state. Incorrect storage state fails the upgrade.
Standbys use it to replay the operations per relation/directory.
- Extended the patch to properly support cascaded standbys. Cascaded standbys
also ingest upgrade WALs like normal replication.
Github repo: https://github.com/LeeBohyun/postgres/blob/wal-upgrade-patch/
Extended high level doc: https://github.com/LeeBohyun/postgres/blob/wal-upgrade-patch/doc/high-level/wal-logging-pg-upgrade-v2.pdf
Since some feedback and concerns overlapped, I grouped and summarized them
Since some feedback and concerns overlapped, I grouped and summarized them
by high-level topic.
1. Comparing the proposed pg_upgrade --wal-upgrade vs. pg_upgrade_replica
I agree that pg_upgrade_replica is less disruptive (no redo/rmgr/new WAL records).
1. Comparing the proposed pg_upgrade --wal-upgrade vs. pg_upgrade_replica
I agree that pg_upgrade_replica is less disruptive (no redo/rmgr/new WAL records).
Why WAL-logging pg_upgrade may still be worth it:
- Backup-chain continuity. The upgrade becomes part of the WAL stream, so the
existing base backup plus archive recover straight across the upgrade boundary
with no new base backup (see Section 5 of the high-level doc).
pg_upgrade_replica only resyncs standbys and leaves this gap open.
- Less per-standby backup orchestration. With
pg_upgrade_replica, each standby is a backup pipeline the operator drives: point
the tool at the new primary, it pulls an incremental pg_basebackup and runs
pg_combinebackup to rebuild that standby from its old files. A cascade is handled
node by node, each node its own backup pipeline against the new primary.
--wal-upgrade instead lets standbys adopt the upgrade through the WAL replay they
already run. The operator connects a fresh new-version standby (over its retained
old datadir) to its parent, and it ingests and replays the upgrade WALs. Once
they are fully replayed, the standby automatically finalizes as a hot standby and
keeps following its parent; no coordination for that transition. The only
operator follow-up is reconfiguring it onto its named slot. No per-standby backup
pipeline, and in reference modes no bulk user data is pulled from the primary (user
files come from the standby's own retained old datadir).
- Seamless upgrade with cascaded standbys. The upgrade propagates down the cascade's
existing replication links: each standby replays the upgrade from its real
parent (primary -> standby -> standby), exactly as it already streams. Therefore,
bringing up a deep cascade is just starting each node in order, with no per-node
backup against the primary. The transitive handoff composes the caught-up
guarantee up the tree on its own, whereas pg_upgrade_replica must rebuild each
cascade level separately. The two-level cascade is exercised by
t/011_wal_upgrade_standby.pl and t/WalUpgradeCascade.pm.
- No WAL-summary window, and failures retry instead of re-cloning.
pg_upgrade_replica requires summarize_wal on from the new primary's first
post-upgrade startup, with summaries pruned after wal_summary_keep_time (10 days
by default). If that window is missed, the incremental backup cannot be built,
so the fallback is a full re-clone. WAL-logging has no such window: the migrated
slot retains the upgrade WALs, so a standby can replay it whenever it is brought
up. The upgrade is also atomic. Recovery never makes an incomplete upgrade bootable
or finalizes it as a valid cluster, so a failed attempt retries rather than
re-clones: a standby bring-up failure replays the retained window from a fresh
skeleton, and a primary upgrade failure reruns pg_upgrade (as long as the old
source is still intact).
To sum up, it seems to me that pg_upgrade_replica solves the narrower problem
(avoid re-cloning standbys) at lower backend risk, while --wal-upgrade can serve
multiple purposes. It solves that problem plus the backup-chain gap and
cascade adoption, with less per-standby backup orchestration.
--------------------------------------------------------------------------------
2. Handoff atomicity and shutdown ordering
Feedback: the old primary might continue processing transactions after the
handoff. The handoff command should shut down the primary at the handoff record.
The handoff could also be represented by the shutdown checkpoint itself.
How it is addressed:
This patch removes the separate handoff subcommand. The upgrade now owns the
entire handoff and shutdown sequence. It emits the handoff during the old
primary's final shutdown, immediately before the shutdown checkpoint, and flushes
the boundary before shutdown completes. The primary
cannot accept transactions after that boundary. Standbys associate their pause
with the shutdown checkpoint and restore it after a restart.
--------------------------------------------------------------------------------
3. Operational coordination and standby configuration
Feedback: standbys must specify the old and new cluster paths and require
explicit restarts. An alternative design would use operator-supplied standby
mapping, such as --tablespace-mapping, instead of embedding transfer behavior in
replayable WAL.
Feedback: WAL generation depends on observing every
relation file. If collection omits a relation, fork, or segment, replay could
diverge. RELINK also consults a standby GUC, so the response must explain which
results are fixed by WAL and which may differ between standbys.
Feedback: every replica must be configured with
pg_upgrade_standby_old_datadir. At fleet scale, manual configuration creates
operational effort and additional failure modes. A helper should validate every
replica before the primary is upgraded, including its retained directory and
tablespace layout.
Clarification and how this is addressed:
The operator supplies the retained old data directory and chooses the standby's
local transfer mode. The standby starts from a fresh new-version skeleton with
upgrade recovery enabled. WAL describes which relation/directory must be retained,
created, reset, or deleted, while the standby chooses how retained files are
placed on its filesystem.
Replay is deterministic: the relation set and the retain/create/reset/delete
operations are fixed by the WAL and identical on every standby. Only the local
placement method differs, chosen by the standby transfer-mode GUC. Capture
completeness is checked during emission, so an omitted relation, fork, or segment
fails the upgrade rather than letting replay diverge.
This still requires configuring the old and new paths and restarting each
standby. There is no operator-supplied per-tablespace mapping, and a fleet-wide
preflight that validates every replica's retained directory and tablespace layout
before the primary is upgraded is not implemented. Please refer to Sections 2 and 3 of
the doc for the primary-standby upgrade workflow.
--------------------------------------------------------------------------------
4. Rollback and --link behavior
Feedback: the old replica should remain read-only and bootable during the
upgrade. --link intentionally trades away that property to save disk space.
Feedback: the documentation described
--wal-upgrade-rollback and --wal-upgrade-delete-old even though neither command
was implemented. It also needed to list the literal values accepted by the
standby transfer-mode GUC.
How it is addressed:
Addressed with a mode-dependent rollback boundary.
- Copy-based modes keep the old cluster independent and bootable.
- Link disables the old cluster (pg_control is renamed) because the new cluster
shares its files by hardlink, giving up rollback like stock --link.
- Swap loses rollback when files are moved.
The standby transfer-mode GUC (pg_upgrade_standby_transfer_mode) accepts mirror
(the default, which reproduces the primary's mode), clone, copy, copy_file_range,
link, and swap.
The same boundary applies to a standby's retained old data directory. The
documentation no longer states unimplemented rollback or old-directory
deletion commands.
--------------------------------------------------------------------------------
5. WAL volume and scalability
Feedback: clarify how large XLOG_UPGRADE_RELFILE_DATA and
XLOG_UPGRADE_SLRU_DATA become with large catalogs and many databases. Compare
the size and performance with ordinary pg_upgrade --link, and test beyond the
basic TAP suite.
Feedback: WAL contains full-block images for every
changed catalog and SLRU block. How large does it become on a real cluster,
rather than only in the TAP tests?
Clarification:
In terms of upgrade scalability, both normal pg_upgrade --link and pg_upgrade
--wal-upgrade --link dump/restore the schema and hardlink user files (no bulk
copy). --wal-upgrade --link additionally emits the upgrade as WAL; catalog as
full-page images, SLRU/pg_control as RAWFILE, and a compact RELINK entry per
relation fork, but not user-relation data (except forks that must be rebuilt,
such as an unlogged relation's init fork). So its extra cost scales with catalog
+ SLRU + relation count, not user-data size, and emission runs one database at a
time so memory stays bounded.
However, that extra cost is a one-time primary-side expense that every standby
reuses: instead of a per-standby rsync resync (the stock --link procedure) or a
re-clone, each standby just replays the same upgrade WAL. Therefore, the cost
amortizes across the fleet.
Additionally, large-scale transfer-mode tests and an ordinary pg_upgrade --link
baseline have been run outside the TAP suite with bigger relations and SLRU files.
Though there is no production-cluster measurement yet.
- Backup-chain continuity. The upgrade becomes part of the WAL stream, so the
existing base backup plus archive recover straight across the upgrade boundary
with no new base backup (see Section 5 of the high-level doc).
pg_upgrade_replica only resyncs standbys and leaves this gap open.
- Less per-standby backup orchestration. With
pg_upgrade_replica, each standby is a backup pipeline the operator drives: point
the tool at the new primary, it pulls an incremental pg_basebackup and runs
pg_combinebackup to rebuild that standby from its old files. A cascade is handled
node by node, each node its own backup pipeline against the new primary.
--wal-upgrade instead lets standbys adopt the upgrade through the WAL replay they
already run. The operator connects a fresh new-version standby (over its retained
old datadir) to its parent, and it ingests and replays the upgrade WALs. Once
they are fully replayed, the standby automatically finalizes as a hot standby and
keeps following its parent; no coordination for that transition. The only
operator follow-up is reconfiguring it onto its named slot. No per-standby backup
pipeline, and in reference modes no bulk user data is pulled from the primary (user
files come from the standby's own retained old datadir).
- Seamless upgrade with cascaded standbys. The upgrade propagates down the cascade's
existing replication links: each standby replays the upgrade from its real
parent (primary -> standby -> standby), exactly as it already streams. Therefore,
bringing up a deep cascade is just starting each node in order, with no per-node
backup against the primary. The transitive handoff composes the caught-up
guarantee up the tree on its own, whereas pg_upgrade_replica must rebuild each
cascade level separately. The two-level cascade is exercised by
t/011_wal_upgrade_standby.pl and t/WalUpgradeCascade.pm.
- No WAL-summary window, and failures retry instead of re-cloning.
pg_upgrade_replica requires summarize_wal on from the new primary's first
post-upgrade startup, with summaries pruned after wal_summary_keep_time (10 days
by default). If that window is missed, the incremental backup cannot be built,
so the fallback is a full re-clone. WAL-logging has no such window: the migrated
slot retains the upgrade WALs, so a standby can replay it whenever it is brought
up. The upgrade is also atomic. Recovery never makes an incomplete upgrade bootable
or finalizes it as a valid cluster, so a failed attempt retries rather than
re-clones: a standby bring-up failure replays the retained window from a fresh
skeleton, and a primary upgrade failure reruns pg_upgrade (as long as the old
source is still intact).
To sum up, it seems to me that pg_upgrade_replica solves the narrower problem
(avoid re-cloning standbys) at lower backend risk, while --wal-upgrade can serve
multiple purposes. It solves that problem plus the backup-chain gap and
cascade adoption, with less per-standby backup orchestration.
--------------------------------------------------------------------------------
2. Handoff atomicity and shutdown ordering
Feedback: the old primary might continue processing transactions after the
handoff. The handoff command should shut down the primary at the handoff record.
The handoff could also be represented by the shutdown checkpoint itself.
How it is addressed:
This patch removes the separate handoff subcommand. The upgrade now owns the
entire handoff and shutdown sequence. It emits the handoff during the old
primary's final shutdown, immediately before the shutdown checkpoint, and flushes
the boundary before shutdown completes. The primary
cannot accept transactions after that boundary. Standbys associate their pause
with the shutdown checkpoint and restore it after a restart.
--------------------------------------------------------------------------------
3. Operational coordination and standby configuration
Feedback: standbys must specify the old and new cluster paths and require
explicit restarts. An alternative design would use operator-supplied standby
mapping, such as --tablespace-mapping, instead of embedding transfer behavior in
replayable WAL.
Feedback: WAL generation depends on observing every
relation file. If collection omits a relation, fork, or segment, replay could
diverge. RELINK also consults a standby GUC, so the response must explain which
results are fixed by WAL and which may differ between standbys.
Feedback: every replica must be configured with
pg_upgrade_standby_old_datadir. At fleet scale, manual configuration creates
operational effort and additional failure modes. A helper should validate every
replica before the primary is upgraded, including its retained directory and
tablespace layout.
Clarification and how this is addressed:
The operator supplies the retained old data directory and chooses the standby's
local transfer mode. The standby starts from a fresh new-version skeleton with
upgrade recovery enabled. WAL describes which relation/directory must be retained,
created, reset, or deleted, while the standby chooses how retained files are
placed on its filesystem.
Replay is deterministic: the relation set and the retain/create/reset/delete
operations are fixed by the WAL and identical on every standby. Only the local
placement method differs, chosen by the standby transfer-mode GUC. Capture
completeness is checked during emission, so an omitted relation, fork, or segment
fails the upgrade rather than letting replay diverge.
This still requires configuring the old and new paths and restarting each
standby. There is no operator-supplied per-tablespace mapping, and a fleet-wide
preflight that validates every replica's retained directory and tablespace layout
before the primary is upgraded is not implemented. Please refer to Sections 2 and 3 of
the doc for the primary-standby upgrade workflow.
--------------------------------------------------------------------------------
4. Rollback and --link behavior
Feedback: the old replica should remain read-only and bootable during the
upgrade. --link intentionally trades away that property to save disk space.
Feedback: the documentation described
--wal-upgrade-rollback and --wal-upgrade-delete-old even though neither command
was implemented. It also needed to list the literal values accepted by the
standby transfer-mode GUC.
How it is addressed:
Addressed with a mode-dependent rollback boundary.
- Copy-based modes keep the old cluster independent and bootable.
- Link disables the old cluster (pg_control is renamed) because the new cluster
shares its files by hardlink, giving up rollback like stock --link.
- Swap loses rollback when files are moved.
The standby transfer-mode GUC (pg_upgrade_standby_transfer_mode) accepts mirror
(the default, which reproduces the primary's mode), clone, copy, copy_file_range,
link, and swap.
The same boundary applies to a standby's retained old data directory. The
documentation no longer states unimplemented rollback or old-directory
deletion commands.
--------------------------------------------------------------------------------
5. WAL volume and scalability
Feedback: clarify how large XLOG_UPGRADE_RELFILE_DATA and
XLOG_UPGRADE_SLRU_DATA become with large catalogs and many databases. Compare
the size and performance with ordinary pg_upgrade --link, and test beyond the
basic TAP suite.
Feedback: WAL contains full-block images for every
changed catalog and SLRU block. How large does it become on a real cluster,
rather than only in the TAP tests?
Clarification:
In terms of upgrade scalability, both normal pg_upgrade --link and pg_upgrade
--wal-upgrade --link dump/restore the schema and hardlink user files (no bulk
copy). --wal-upgrade --link additionally emits the upgrade as WAL; catalog as
full-page images, SLRU/pg_control as RAWFILE, and a compact RELINK entry per
relation fork, but not user-relation data (except forks that must be rebuilt,
such as an unlogged relation's init fork). So its extra cost scales with catalog
+ SLRU + relation count, not user-data size, and emission runs one database at a
time so memory stays bounded.
However, that extra cost is a one-time primary-side expense that every standby
reuses: instead of a per-standby rsync resync (the stock --link procedure) or a
re-clone, each standby just replays the same upgrade WAL. Therefore, the cost
amortizes across the fleet.
Additionally, large-scale transfer-mode tests and an ordinary pg_upgrade --link
baseline have been run outside the TAP suite with bigger relations and SLRU files.
Though there is no production-cluster measurement yet.
Those tests are not included in the patch, but
please refer to the github link to see what was additionally run.
--------------------------------------------------------------------------------
6. Missing-file handling and corruption detection
Feedback: silently continuing on ENOENT during RELINK replay could hide an
incorrect path or a damaged old data directory. Unexpected missing files should
cause a clear failure.
Feedback: RELINK's ENOENT skip: is it reachable only for files that are
correctly absent (an unlogged relation's main fork), or could a real
missing file hit it too and leave a silent gap?
How it is addressed:
The patch validates every required relation file and its physical size while
building and replaying the RELINK records. The ENOENT skip is reachable for
exactly two forks, the free-space map and the visibility map.
check_inherited_file() returns false (skip) only when errno is ENOENT and the
fork is FSM or VM. For any other fork a missing file is ereport(PANIC) "could not
stat upgrade source", and a wrong size is ereport(PANIC) "inherited upgrade file
has the wrong size". So a real missing required file cannot pass through the skip.
An unlogged relation's main fork is never an inherited entry: relation_fork_mask
gives an unlogged relation only its init fork, so the main fork is not emitted or
looked for. It is reset (recreated from the init fork) by the normal
unlogged-relation reset during recovery, not silently skipped.
The missing-source-fork rejection is exercised by t/012_wal_upgrade_validation.pl.
See Section 6, Upgrade Storage Validation, of the high-level doc for details.
--------------------------------------------------------------------------------
7. Crash before finalization
Feedback: if a server crashes before the upgrade is
finalized, the operator need to rerun the upgrade. The behavior and recovery
procedure need to be defined and documented.
How it is addressed:
Finalization is atomic. The upgrade finalizes only when the completion checkpoint
runs after a committed COMPLETE, setting upgrade_finalized in pg_control. A crash
before that can leave a physically partial skeleton on disk, but it never makes
an incomplete upgrade bootable or finalizes it as a valid cluster: startup
rejects the incomplete window and requires a fresh skeleton.
Recovery is retry, not repair: rerun pg_upgrade on the primary, or discard the
skeleton and start a fresh one that replays the still-retained window. Retrying
this way relies on the retained old source still being intact, which holds for
copy-based modes. Link and swap may already have modified it.
This is exercised by the TAP suite: 010 truncates the upgrade WAL and 012
disconnects after COMPLETE before COMMIT, both asserting the unfinalized state is
rejected, and 011 asserts a partial skeleton is discarded for a fresh one.
I'd appreciate any feedbacks, please flag any potential concerns.
Best,
Bohyun
please refer to the github link to see what was additionally run.
--------------------------------------------------------------------------------
6. Missing-file handling and corruption detection
Feedback: silently continuing on ENOENT during RELINK replay could hide an
incorrect path or a damaged old data directory. Unexpected missing files should
cause a clear failure.
Feedback: RELINK's ENOENT skip: is it reachable only for files that are
correctly absent (an unlogged relation's main fork), or could a real
missing file hit it too and leave a silent gap?
How it is addressed:
The patch validates every required relation file and its physical size while
building and replaying the RELINK records. The ENOENT skip is reachable for
exactly two forks, the free-space map and the visibility map.
check_inherited_file() returns false (skip) only when errno is ENOENT and the
fork is FSM or VM. For any other fork a missing file is ereport(PANIC) "could not
stat upgrade source", and a wrong size is ereport(PANIC) "inherited upgrade file
has the wrong size". So a real missing required file cannot pass through the skip.
An unlogged relation's main fork is never an inherited entry: relation_fork_mask
gives an unlogged relation only its init fork, so the main fork is not emitted or
looked for. It is reset (recreated from the init fork) by the normal
unlogged-relation reset during recovery, not silently skipped.
The missing-source-fork rejection is exercised by t/012_wal_upgrade_validation.pl.
See Section 6, Upgrade Storage Validation, of the high-level doc for details.
--------------------------------------------------------------------------------
7. Crash before finalization
Feedback: if a server crashes before the upgrade is
finalized, the operator need to rerun the upgrade. The behavior and recovery
procedure need to be defined and documented.
How it is addressed:
Finalization is atomic. The upgrade finalizes only when the completion checkpoint
runs after a committed COMPLETE, setting upgrade_finalized in pg_control. A crash
before that can leave a physically partial skeleton on disk, but it never makes
an incomplete upgrade bootable or finalizes it as a valid cluster: startup
rejects the incomplete window and requires a fresh skeleton.
Recovery is retry, not repair: rerun pg_upgrade on the primary, or discard the
skeleton and start a fresh one that replays the still-retained window. Retrying
this way relies on the retained old source still being intact, which holds for
copy-based modes. Link and swap may already have modified it.
This is exercised by the TAP suite: 010 truncates the upgrade WAL and 012
disconnects after COMPLETE before COMMIT, both asserting the unfinalized state is
rejected, and 011 asserts a partial skeleton is discarded for a fresh one.
I'd appreciate any feedbacks, please flag any potential concerns.
Best,
Bohyun
On Tue, Sep 8, 2026 at 2:59 PM Marco Nenciarini <marco.nenciarini@enterprisedb.com> wrote:
Reposting this: my first reply (Aug 12) never threaded correctly and
didn't show up here. Sorry for the duplicate to those who already saw
it.
This thread solves the same problem as mine (pg_upgrade_replica [1]).
Different design, worth comparing.
> The transfer mode setting seems strange and not well motivated to me.
I have a similar choice: --tablespace-mapping. It's a command-line flag,
not a WAL record, so it never has to mean the same thing twice.
That's also why pg_upgrade_replica stays in src/bin/: it reuses
pg_upgrade's own manifest, a real backup_manifest, and pg_basebackup
--incremental's own WAL-summary code. No new WAL format, no core redo
change. Its only new core footprint is one small manifest file.
> For rollback, can't the operator just pause the rollback target at the
> handoff checkpoint while still on the old binary and promote if
> necessary?
Yes, in my design: --old-replica stays read-only and bootable on the
old binary. Unless --link is used, then the new standby's first write
corrupts it too, the same tradeoff pg_upgrade's own --link already
makes.
One thing your design fixes that mine doesn't: a primary that crashes
before the next backup or resync. WAL-logging makes recovery just
normal replay across the upgrade. My tool only runs after the upgrade,
against an already-running primary.
Hüseyin Demir's point about tablespace layout and cascading replicas
applies to me too. --tablespace-mapping doesn't assume the standard
pg_tblspc/ layout, but cascading isn't special-cased: each hop still
needs its own run.
Two questions for Bohyun:
- WAL size: full block images for every changed catalog and SLRU
block. How big does this get on a real cluster, not just the TAP
tests?
- RELINK's ENOENT skip: is it reachable only for files that are
correctly absent (an unlogged relation's main fork), or could a real
missing file hit it too and leave a silent gap?
Two valid answers to the same problem. Wanted the comparison on
record.
[1] https://www.postgresql.org/message-id/flat/CA%2BnrD2fqdeEJkGJrDt%2B-a7Uqr4OucXZHvSVxLCb9J0EkN%2BhLhw%40mail.gmail.com
Marco Nenciarini
EnterpriseDBOn Tue, Aug 11, 2026 04:30 PM, John Naylor <johncnaylorls@gmail.com> wrote:On Thu, Aug 6, 2026 at 2:43 PM Heikki Linnakangas <hlinnaka@iki.fi> wrote:
>
> On 06/08/2026 01:17, John Naylor wrote:
> > On Wed, Aug 5, 2026 at 9:36 PM Bohyun Lee
> > <bohyun.lee@databricks.com> wrote:
> >> The GUCs are introduced because we cannot assume the standby's
> >> storage layout exactly matches the primary's, neither where the
> >> retained pre-upgrade directory sits, nor how its files are
> >> physically placed. It also depends on how the cluster intends to
> >> use the standby. For instance, if the operator wants to keep the
> >> standby as a rollback target, it may be worthwhile to use a
> >> different transfer mode via the newly introduced
> >> pg_upgrade_standby_transfer_mode GUC, which I believe is useful.
> >> Nevertheless, the reconstructed cluster will be logically
> >> identical, even if the physical representation diverges.
> >
> > Given the above design concepts, I still think WAL is fundamentally
> > the wrong mechanism for this.
>
> Can you elaborate? Do you think the changes that pg_upgrade makes should
> be written somewhere else than WAL, or is this just about the transfer
> mode setting, or something else? How would you do it?
The transfer mode setting seems strange and not well motivated to me.
For rollback, can't the operator just pause the rollback target at the
handoff checkpoint while still on the old binary and promote if
necessary? Am I missing something?
--
John Naylor
Amazon Web Services
Attachment
pgsql-hackers by date: