Hi hackers,
pg_resetwal does not look for backup_label today. On a restored base
backup, pg_control says "in production", so pg_resetwal asks for -f, and
with -f it goes ahead. The server then fails with "could not locate
required checkpoint record" and a hint to remove backup_label. Removing it
at that point throws away what the backup needs to become consistent.
The attached patch makes pg_resetwal refuse when backup_label exists, even
with -f. This follows the postmaster.pid check, which -f does not override
either. It came up in the CF 4997 thread [1], and David preferred that -f
not bypass it [2]. On a pg_basebackup copy it now says:
pg_resetwal: error: backup label file "backup_label" exists
pg_resetwal: hint: If you are restoring from a backup, configure recovery instead. If you are not restoring from a backup, delete the backup label file and try again.
-n is refused too. A dry run should show what a real run would do, and the
real run refuses. pg_controldata still works for reading the control file.
Robert worried in [3] that people with a bad backup will just run
pg_resetwal, which is worse than starting from the wrong checkpoint. This
patch targets that step. A backup that still has its backup_label is the
case where the user has everything needed for a correct restore, and
pg_resetwal is the wrong tool. Now it says so before it does any damage. A
user can still delete the file and rerun, but that is a second deliberate
step, and the docs now say when that is safe.
0002 adds TAP tests and is optional. I think this is master only, since it
changes what an existing command accepts.
[1]
https://postgr.es/m/e2636c5d-c031-43c9-a5d6-5e5c7e4c5514@pgmasters.net[2]
https://postgr.es/m/556bef11-938d-45ca-94b3-7143a3cb532d@pgbackrest.org[3]
https://postgr.es/m/CA+TgmoZkdrWyd7KiPFHaJBg+tjM3UFrqOBK1EtG3NtVs97-7Xw@mail.gmail.comThanks,
Shihao