Re: doc: Document Linux cgroup memory limits - Mailing list pgsql-hackers

From Manu
Subject Re: doc: Document Linux cgroup memory limits
Date
Msg-id 179099690687.256144.11549763261274643957@gmail.com
Whole thread
In response to doc: Document Linux cgroup memory limits  (Joao Detomini <joao.detomini@enterprisedb.com>)
List pgsql-hackers
Hi João,

Thanks for writing this up.  I tried it on a cgroup v2 host (kernel
7.2, systemd 262), running master in a systemd scope with
MemoryMax=400M and no swap.

The main behaviour checks out.  With four sessions sorting with
work_mem = 300MB, one process was SIGKILLed, the postmaster
reinitialized, every session lost its connection, and nobody got an
out-of-memory error.  Reading a 638 MB table twice took usage up to
memory.max without any kill, since the page cache was reclaimed.

Two things I would change:

> Memory allocated from the huge page pool (see huge_pages) is by
> default not counted towards the memory usage of the cgroup.

That is the kernel default, but systemd 259 and later mount cgroup2
with memory_hugetlb_accounting [1], and then it is counted.  Here,
with shared_buffers = 256MB and huge_pages = on, memory.stat showed
290MB under hugetlb.  Perhaps say that whether it counts depends on
that mount option.

> If that is a server process other than the postmaster, the
> postmaster treats it as a crash

With memory.oom.group set to 1, the OOM killer kills every process in
the cgroup, the postmaster included.  There is no crash recovery then,
and nothing reaches the server log: the last line was "database system
is ready to accept connections".  Kubernetes sets memory.oom.group for
containers on cgroup v2 unless singleProcessOOMKill is enabled [2], so
this is probably the case most container users will see.  A sentence
about it would help.

One observation, maybe not for the docs: the first process killed was
not one of the sorting backends but an io worker, with about 90MB of
shared buffers mapped and almost no private memory.  The OOM killer
counts mapped shared memory, so the server log named the io worker
rather than the sessions that used the memory.

[1] https://github.com/systemd/systemd/blob/main/NEWS (CHANGES WITH 259)
[2] https://kubernetes.io/docs/reference/config-api/kubelet-config.v1beta1/
    (singleProcessOOMKill)

Regards,
Manu



pgsql-hackers by date:

Previous
From: Amit Langote
Date:
Subject: Re: Revert RI fast-path batching from REL_19_STABLE
Next
From: shihao zhong
Date:
Subject: Re: BUG #19686: Rolling back SET TABLESPACE