Re: is hibernation usable?

6 messages, 4 authors, 2020-02-21 · open the first message on its own page

Re: is hibernation usable?

From: Luigi Semenzato <hidden>
Date: 2020-02-20 17:38:20

I was forgetting: forcing swap by eating up memory is dangerous
because it can lead to unexpected OOM kills, but you can mitigate that
by giving the memory-eaters a higher OOM kill score.  Still, some way
of calling try_to_free_pages() directly from user-level would be
preferable.  I wonder if such API has been discussed.


On Thu, Feb 20, 2020 at 9:16 AM Luigi Semenzato [off-list ref] wrote:
I think this is the right group for the memory issues.

I suspect that the problem with failed allocations (ENOMEM) boils down
to the unreliability of the page allocator.  In my experience, under
pressure (i.e. pages must be swapped out to be reclaimed) allocations
can fail even when in theory they should succeed.  (I wish I were
wrong and that someone would convincingly correct me.)

I have a workaround in which I use memcgroups to free pages before
starting hibernation.  The cgroup request "echo $limit >
.../memory.limit_in_bytes"  blocks until memory usage in the chosen
cgroup is below $limit.  However, I have seen this request fail even
when there is extra available swap space.

The callback for the operation is mem_cgroup_resize_limit() (BTW I am
looking at kernel version 4.3.5) and that code has a loop where
try_to_free_pages() is called up to retry_count, which is at least 5.
Why 5?  One suspects that the writer of that code must have also
realized that the page freeing request is unreliable and it's worth
trying multiple times.

So you could try something similar.  I don't know if there are
interfaces to try_to_free_pages() other than those in cgroups.  If
not, and you aren't using cgroups, one way might be to start several
memory-eating processes (such as "dd if=/dev/zero bs=1G count=1 |
sleep infinity") and monitor allocation, then when they use more than
50% of RAM kill them and immediately hibernate before the freed pages
are reused.  If you can build your custom kernel, maybe it's worth
adding a sysfs entry to invoke try_to_free_pages().  You could also
change the hibernation code to do that, but having the user-level hook
may be more flexible.


On Wed, Feb 19, 2020 at 6:56 PM Chris Murphy [off-list ref] wrote:
quoted
Also, is this the correct list for hibernation/swap discussion? Or linux-pm@?

Thanks,

Chris Murphy

Re: is hibernation usable?

From: Michal Hocko <mhocko@kernel.org>
Date: 2020-02-21 08:49:14

On Thu 20-02-20 09:38:06, Luigi Semenzato wrote:
I was forgetting: forcing swap by eating up memory is dangerous
because it can lead to unexpected OOM kills
Could you be more specific what you have in mind? swapoff causing the
OOM killer?
, but you can mitigate that
by giving the memory-eaters a higher OOM kill score.  Still, some way
of calling try_to_free_pages() directly from user-level would be
preferable.  I wonder if such API has been discussed.
No, there is no API to trigger the global memory reclaim. You could
start the reclaim by increasing min_free_kbytes but I wouldn't really
recommend that unless you know exactly what you are doing and also I
fail to see the point. If s2disk fails due to insufficient swap space
then how can a pro-active reclaim help in the first place?
-- 
Michal Hocko
SUSE Labs

Re: is hibernation usable?

From: "Rafael J. Wysocki" <rafael@kernel.org>
Date: 2020-02-21 09:04:32

On Fri, Feb 21, 2020 at 9:49 AM Michal Hocko [off-list ref] wrote:
On Thu 20-02-20 09:38:06, Luigi Semenzato wrote:
quoted
I was forgetting: forcing swap by eating up memory is dangerous
because it can lead to unexpected OOM kills
Could you be more specific what you have in mind? swapoff causing the
OOM killer?
quoted
, but you can mitigate that
by giving the memory-eaters a higher OOM kill score.  Still, some way
of calling try_to_free_pages() directly from user-level would be
preferable.  I wonder if such API has been discussed.
No, there is no API to trigger the global memory reclaim. You could
start the reclaim by increasing min_free_kbytes but I wouldn't really
recommend that unless you know exactly what you are doing and also I
fail to see the point. If s2disk fails due to insufficient swap space
then how can a pro-active reclaim help in the first place?
My understanding of the problem is that the size of swap is
(theoretically) sufficient, but it is not used as expected during the
preallocation of image memory.

It was stated in one of the previous messages (not in this thread,
cannot find it now) that swap (of the same size as RAM) was activated
(swapon) right before hibernation, so theoretically that should be
sufficient AFAICS.

Re: is hibernation usable?

From: Michal Hocko <mhocko@kernel.org>
Date: 2020-02-21 09:36:41

On Fri 21-02-20 10:04:18, Rafael J. Wysocki wrote:
On Fri, Feb 21, 2020 at 9:49 AM Michal Hocko [off-list ref] wrote:
quoted
On Thu 20-02-20 09:38:06, Luigi Semenzato wrote:
quoted
I was forgetting: forcing swap by eating up memory is dangerous
because it can lead to unexpected OOM kills
Could you be more specific what you have in mind? swapoff causing the
OOM killer?
quoted
, but you can mitigate that
by giving the memory-eaters a higher OOM kill score.  Still, some way
of calling try_to_free_pages() directly from user-level would be
preferable.  I wonder if such API has been discussed.
No, there is no API to trigger the global memory reclaim. You could
start the reclaim by increasing min_free_kbytes but I wouldn't really
recommend that unless you know exactly what you are doing and also I
fail to see the point. If s2disk fails due to insufficient swap space
then how can a pro-active reclaim help in the first place?
My understanding of the problem is that the size of swap is
(theoretically) sufficient, but it is not used as expected during the
preallocation of image memory.

It was stated in one of the previous messages (not in this thread,
cannot find it now) that swap (of the same size as RAM) was activated
(swapon) right before hibernation, so theoretically that should be
sufficient AFAICS.
Hmm, this is interesting. Let me have a closer look...

pm_restrict_gfp_mask which would completely rule out any IO
happens after hibernate_preallocate_memory is done and my limited
understanding tells me that this is where all the reclaim happens
(via shrink_all_memory). It is quite possible that the MM decides to
not swap in that path - depending on the memory usage - and miss it's
target. More details would be needed. E.g. vmscan tracepoints could tell
us more.

-- 
Michal Hocko
SUSE Labs

Re: is hibernation usable?

From: Chris Murphy <hidden>
Date: 2020-02-21 09:47:10

On Fri, Feb 21, 2020 at 2:04 AM Rafael J. Wysocki [off-list ref] wrote:
My understanding of the problem is that the size of swap is
(theoretically) sufficient, but it is not used as expected during the
preallocation of image memory.
Right. I have no idea how locality of pages is determined in the swap
device. But if it's sufficiently fragmented such that contiguous free
space for a hibernation image is not sufficient, then hibernation
could fail.
It was stated in one of the previous messages (not in this thread,
cannot find it now) that swap (of the same size as RAM) was activated
(swapon) right before hibernation, so theoretically that should be
sufficient AFAICS.
I mentioned it as an idea floated by systemd developers. I'm not sure
if it's mentioned elsewhere. Some folks wonder if such functionality
could be prone to racing.
https://lore.kernel.org/linux-mm/CAJCQCtSx0FOX7q0p=9XgDLJ6O0+hF_vc-wU4KL=c9xoSGGkstA@mail.gmail.com/T/#m4d47d127da493f998b232d42d81621335358aee1

Another idea that's been suggested for a while is formally separating
hibernation and paging into separate files (or partitions).
a. Guarantees hibernation image has the necessary contiguous free space.
b. Might be easier to create (or even obviate) a sane interface for
hibernation images in swapfiles; that is, if it were a dedicated
hibernationfile rather than being inserted in a swapfile. Right now
that interface doesn't exist, so e.g. on Btrfs while it can support
swapfiles and hibernation images, the offset has to be figured out
manually so resume can succeed.
https://github.com/systemd/systemd/issues/11939#issuecomment-471684411





--
Chris Murphy

Re: is hibernation usable?

From: Luigi Semenzato <hidden>
Date: 2020-02-21 17:13:25

On Fri, Feb 21, 2020 at 1:36 AM Michal Hocko [off-list ref] wrote:
On Fri 21-02-20 10:04:18, Rafael J. Wysocki wrote:
quoted
On Fri, Feb 21, 2020 at 9:49 AM Michal Hocko [off-list ref] wrote:
quoted
On Thu 20-02-20 09:38:06, Luigi Semenzato wrote:
quoted
I was forgetting: forcing swap by eating up memory is dangerous
because it can lead to unexpected OOM kills
Could you be more specific what you have in mind? swapoff causing the
OOM killer?
No, not swapoff, just fast allocation.

Also, in some earlier experiments I tried gradually increasing
min_free_kbytes (precisely as suggested) and this would randomly
trigger OOM kills when swap space was still available.
quoted
quoted
quoted
, but you can mitigate that
by giving the memory-eaters a higher OOM kill score.  Still, some way
of calling try_to_free_pages() directly from user-level would be
preferable.  I wonder if such API has been discussed.
No, there is no API to trigger the global memory reclaim. You could
start the reclaim by increasing min_free_kbytes but I wouldn't really
recommend that unless you know exactly what you are doing and also I
fail to see the point. If s2disk fails due to insufficient swap space
then how can a pro-active reclaim help in the first place?
My understanding of the problem is that the size of swap is
(theoretically) sufficient, but it is not used as expected during the
preallocation of image memory.

It was stated in one of the previous messages (not in this thread,
cannot find it now) that swap (of the same size as RAM) was activated
(swapon) right before hibernation, so theoretically that should be
sufficient AFAICS.
Correct, those were my experiments.  Search the archives for
"semenzato", there are a couple of threads on the topic.

But really, why not have a user-level interface for reclaim?  I find
it very difficult to understand the behavior of the reclaim code, and
any attempt to reclaim from user level (memory-eating processes,
raising min_free_kbytes) can end in the OOM-kill path.  Using cgroups'
memory.limit_in_bytes doesn't have this problem, precisely because it
only calls try_to_free_pages(), which doesn't trigger OOM killing.  If
I could make that call from user level (without cgroups) it would
greatly simplify my current workaround, and would be useful in other
situations as well.

Something like

  echo $page_count > /proc/sys/vm/try_to_free_pages
  cat /proc/sys/vm/pages_freed   # the number of pages freed at the
latest request
Hmm, this is interesting. Let me have a closer look...

pm_restrict_gfp_mask which would completely rule out any IO
happens after hibernate_preallocate_memory is done and my limited
understanding tells me that this is where all the reclaim happens
(via shrink_all_memory). It is quite possible that the MM decides to
not swap in that path - depending on the memory usage - and miss it's
target. More details would be needed. E.g. vmscan tracepoints could tell
us more.

--
Michal Hocko
SUSE Labs

Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help