> On windows, it seems it can randomly happen, but it's very unlikely
I ended up debugging for some more on windows, and then looking into
the history of this issue.
The following modification works for testing Andrey's patch or also my
v2, and doesn't use a hard coded timeout, it reproduces the original
reported issue 100% in my tests on windows. It could be still timing
dependent on CI under load, but that just means that sometimes the
test would pass with the original output, not _1.
step terminate3 {
SELECT pg_terminate_backend(pid) FROM pg_stat_activity
WHERE wait_event = 'injection-points-wait';
+ SELECT pg_sleep(0.1);
}
It doesn't test the remainder issue fixed by Nikolay's 0003 or also by
v2-0001, but it is a simple change. Alternatively, if we do this in a
wait loop we can also reproduce the additional failure reported/fixed
by Nikolay without a hardcoded timeout.
I attached v3, which contains this additional change (the looped version).
I also looked into the history of this, and I think we might solve the
problem instead, at least on master:
6051857fc on 2021-12-02 added closesocket() in socket_close() under
#ifdef WIN32.
ed52c3707 on 2021-12-07 also shutdown(sock, SD_SEND)
these would fix the issue properly, making the alternative outputs and
the isolation tester fix unnecessary, however
75674c7ec on 2022-01-25 reverted in back branches because of issues
with walreciever
29992a6a5 on 2022-03-22 reverted also in master for the same reason
but also
a8458f508a7 on 2024-07-13 fixed the issue causing that instability, so
now it should be safe to revert the revert, at least I can't reproduce
the problem mentioned in the reverts with this fix in place, and I can
without it.
In theory a8458f508a7 was backported everywhere, but that doesn't mean
this change would be completely risk-free.