[BUG] atlantic: resume from S3 fails with -ENOMEM in aq_ring_alloc(), leaves unusable netdev
From: Stephan Hauser <hidden>
Date: 2026-09-26 19:49:10
Hi,
The atlantic driver intermittently fails to resume an AQC107 from suspend-to-RAM.
aq_nic_init() reallocates its per-ring software buffer arrays from the PM resume
path, where the PM core has restricted allocations to GFP_NOIO. With the default
ring sizes these are order-5 and order-6 requests (128 KiB and 256 KiB), which
cannot reliably be satisfied in a context that can neither reclaim, compact, nor
dip into reserves.
Resume then returns -ENOMEM and the device is left half-initialised: the netdev
still exists and is still marked IFF_UP, but has no rings and no vectors. It
passes no traffic, DHCP never completes, and bringing the link down hangs.
I have 3 occurrences in 60 suspend/resume cycles (5%) over 5.5 weeks of journal,
across kernels 7.1.4 and 7.2.2. What makes this worth reporting rather than
filing under "memory was tight": the three failures have three *different*
proximate causes in the allocator, detailed below. The allocation is fragile in
several independent ways at once, which is why no amount of tuning fixes it.
Two distinct problems:
1. A >PAGE_ALLOC_COSTLY_ORDER kmalloc() on the resume path, for memory that
never needs to be physically contiguous.
2. atl_resume_common() does not unwind when aq_nic_init() fails, so an
allocation failure is converted into a wedged interface.
Disclosure
----------
The analysis in this report, and the text of the report itself, were produced
with the assistance of Claude (Anthropic). Please weigh it accordingly.
What is directly evidenced: the hardware is real and in front of me, and every
log line, zone dump and free-page histogram quoted below is verbatim from my
journal. The failure counts come from scanning all boots in that journal.
What is inference rather than measurement, and where I would welcome a second
opinion:
- The 64-byte element size is derived from the two reported allocation
orders, not read out of the source.
- The TX-before-RX ordering in aq_vec_ring_alloc(), and the claim that
aq_nic_init() aborts on first failure, are inferred from the backtrace and
from only ever seeing a single warning per event.
- The ZONE_DMA32 lowmem_reserve arithmetic in case 3 is hand-computed from
the zone dump.
- The aq_nic_stop() attribution for the link-down hang is a guess; I say so
again where it appears.
- The suggested fixes were not compiled or tested, and were written without
the source tree to hand. Treat them as a direction, not a patch.
I have read the whole thing and stand behind reporting it, but I have not
personally verified the driver internals against the code.
System
------
Kernel: 7.1.4 and 7.2.2 (also running 7.2.7), x86_64, PREEMPT(lazy)
Hardware: Micro-Star International Co., Ltd. MEG X570 UNIFY (MS-7C35),
BIOS A.80 01/22/2021
NIC: Aquantia AQC107 NBase-T/IEEE 802.3an [Atlantic 10G] (rev 02)
PCI 0000:24:00.0, [1d6a:07b1], subsystem [1d6a:0001]
Driver: atlantic, firmware-version 3.1.100
Rings: rx 2048 / tx 4096 (driver defaults; maximum 8184), 8 vectors
Memory: 64 GB, no swap configured
Reproduction
------------
Not reproducible on demand. Correlates with uptime and page cache size:
1. Boot, use the machine until the page cache has filled most of RAM.
2. echo mem > /sys/power/state
3. Resume.
In one boot, a suspend/resume at t=11793s succeeded and one at t=52899s in the
same boot failed.
The allocation
--------------
From the backtrace this is a large kmalloc (__GFP_COMP|__GFP_ZERO, falling
through __kmalloc_large_node_noprof() straight to the page allocator), i.e. a
kcalloc() of the per-ring software buffer array — CPU-side per-descriptor
bookkeeping, not the DMA descriptor ring, which is allocated separately via
dma_alloc_coherent().
The two reported orders pin the element size at 64 bytes:
tx ring 4096 * 64 B = 262144 B = order 6
rx ring 2048 * 64 B = 131072 B = order 5
aq_vec_ring_alloc() allocates TX before RX and aq_nic_init() aborts on the first
failure, so an order-6 report means the very first ring failed, and an order-5
report means some TX rings succeeded first. With 8 vectors the driver asks for
8 order-6 plus 8 order-5 blocks on every single resume.
pm_restrict_gfp_mask() clears __GFP_IO/__GFP_FS for the duration of
suspend/resume, so the driver's GFP_KERNEL is silently downgraded to GFP_NOIO
underneath it. It cannot write back, and compaction is largely defeated.
The three failures
------------------
1) 2026-08-31, kernel 7.1.4, order:6 — fragmentation at exactly order 6
Node 0 Normal free:3363056kB boost:283840kB min:348116kB
Node 0 Normal: 46873*4kB (UME) 74929*8kB (UME) 63460*16kB (UME) 33869*32kB (UME) 5476*64kB (UME) 959*128kB (UM) 0*256kB 1*512kB (H) 1*1024kB (H) 0*2048kB 0*4096kB = 3360844kB
3.2 GB free, comfortably above the watermark, and 959 free order-5 blocks —
but zero at order 6. The only blocks above that are marked (H),
MIGRATE_HIGHATOMIC, and so unavailable here.
Note what this means: had the TX ring been order 5 rather than order 6, this
resume would have succeeded. The default tx ring size of 4096 sits exactly
one order above what the zone could serve.
2) 2026-09-25, kernel 7.2.2, order:6 — below the min watermark
Node 0 Normal free:342528kB boost:283840kB min:348116kB
Node 0 Normal: 204*4kB (ME) 293*8kB (UME) 5378*16kB (UME) 1673*32kB (UME) 1175*64kB (UMH) 601*128kB (UMH) 173*256kB (M) 2*512kB (M) 0*1024kB 0*2048kB 0*4096kB = 340184kB
Here there were 173 free order-6 blocks and 601 at order 5 — fragmentation
was not the problem. The zone was simply 5.5 MB below its min watermark, so
the watermark check rejects the request at any order. GFP_NOIO cannot reclaim
its way back above the watermark and cannot touch reserves.
No ring size would have helped this one.
3) 2026-09-26, kernel 7.2.2, order:5 — high-order lists drained mid-loop
Node 0 Normal free:803236kB boost:283840kB min:348116kB
Node 0 Normal: 37466*4kB (UME) 19631*8kB (UME) 10676*16kB (UME) 8307*32kB (UME) 892*64kB (UME) 0*128kB 0*256kB 0*512kB 0*1024kB 0*2048kB 0*4096kB = 800640kB
800 MB free, above the watermark, but nothing at order 5 or above; 46 GB of
page cache (the bulk of it readahead) had shredded the zone. The earlier
order-6 TX allocations in the same loop succeeded and consumed the last of
the high-order blocks; the following order-5 RX allocation then found the
order-5 and order-6 lists both empty.
ZONE_DMA32 had order-5 blocks free but appears to be excluded by
lowmem_reserve rather than by lack of suitable pages:
Node 0 DMA32 free:247268kB min:3288kB
lowmem_reserve[]: 0 0 61020 61020 61020
free 61817 pages - (min 822 + reserve 61020) = -25 pages
About 25 pages short of passing the check, so the fallback zone was closed
too. (This last bit is my arithmetic off the zone dump, not instrumented.)
Note that ZONE_NORMAL is in watermark boost (boost:283840) in all three cases,
i.e. kswapd was already working on recovering high-order blocks and had not
caught up by the time resume ran.
Log (case 3)
------------
[52902.114499] kworker/u97:42: page allocation failure: order:5, mode:0x40d00(GFP_NOIO|__GFP_ZERO|__GFP_COMP), nodemask=(null),cpuset=/,mems_allowed=0
[52902.114508] CPU: 17 UID: 0 PID: 363518 Comm: kworker/u97:42 Not tainted 7.2.2 #1-NixOS PREEMPT(lazy)
[52902.114511] Hardware name: Micro-Star International Co., Ltd. MS-7C35/MEG X570 UNIFY (MS-7C35), BIOS A.80 01/22/2021
[52902.114512] Workqueue: async async_run_entry_fn
[52902.114517] Call Trace:
[52902.114518] <TASK>
[52902.114521] dump_stack_lvl+0x5d/0x80
[52902.114526] warn_alloc+0x158/0x180
[52902.114535] __alloc_frozen_pages_noprof+0x77a/0x1710
[52902.114544] alloc_pages_mpol+0xb6/0x170
[52902.114548] ___kmalloc_large_node+0xb3/0xd0
[52902.114552] __kmalloc_large_node_noprof+0x1d/0xc0
[52902.114555] __kmalloc_noprof+0x424/0x590
[52902.114573] aq_ring_alloc+0x4a/0xd0 [atlantic]
[52902.114581] aq_vec_ring_alloc+0x9d/0x1d0 [atlantic]
[52902.114589] aq_nic_init+0x116/0x1e0 [atlantic]
[52902.114597] atl_resume_common+0x43/0xd0 [atlantic]
[52902.114607] dpm_run_callback+0x4e/0x160
[52902.114613] device_resume+0x15c/0x260
[52902.114616] async_resume+0x21/0x30
[52902.114618] async_run_entry_fn+0x34/0x150
[52902.114620] process_one_work+0x199/0x370
[52902.114624] worker_thread+0x177/0x2e0
[52902.114628] kthread+0xe2/0x110
[52902.114633] ret_from_fork+0x251/0x330
[52902.114637] ret_from_fork_asm+0x1a/0x30
[52902.114642] </TASK>
[52902.119423] atlantic 0000:24:00.0: PM: dpm_run_callback(): pci_pm_resume returns -12
[52902.119427] atlantic 0000:24:00.0: PM: failed to resume async: error -12
Cases 1 and 2 are identical apart from the order and the aq_nic_init()/
aq_ring_alloc() offsets. Full logs available on request.
Consequences
------------
pci_pm_resume() returning -12 does not tear the device down. The netdev remains
registered and IFF_UP, but aq_nic_init() bailed partway through
aq_vec_ring_alloc(), so the vectors and rings behind it were never allocated:
3: enp36s0: <NO-CARRIER,BROADCAST,MULTICAST,UP> mtu 1500 qdisc mq state DOWN
No packets are received, DHCP never completes, and `ip link set enp36s0 down`
hangs indefinitely — presumably in the aq_nic_stop() path operating on NAPI
contexts and rings that do not exist, though I do not have a blocked-task trace
to confirm that. I will send one if I can capture `echo w > /proc/sysrq-trigger`
the next time it reproduces.
Recovering without a reboot requires re-probing the device; the netdev is beyond
help at that point:
echo 1 > /sys/bus/pci/devices/0000:24:00.0/remove
echo 1 > /sys/bus/pci/rescan
Suggested fixes
---------------
1. Use kvcalloc()/kvfree() for the ring buffer arrays in aq_ring_alloc() and
friends. The array is never handed to the device, so it has no need to be
physically contiguous, and a vmalloc fallback removes the high-order
dependency entirely. This addresses cases 1 and 3 outright, and should make
case 2 recoverable as well: order-0 requests can be satisfied from clean
unmapped page cache, which is reclaimable without __GFP_FS, whereas nothing
the allocator can do under GFP_NOIO produces a fresh order-6 block.
This looks like the minimal correct fix.
2. Make atl_resume_common() unwind on error. Whatever the allocator does, a
failed resume should not leave a registered, IFF_UP netdev with no rings
behind it. At minimum the error path should undo the partial aq_nic_init()
and leave the interface administratively down, so a subsequent `ip link set
down` / `up` can recover the device rather than hanging.
3. Worth considering separately: reallocating the rings on resume at all. They
are freed at suspend and immediately reallocated at identical sizes on
resume, in the one context where allocation is most constrained. Keeping the
ring memory across the suspend/resume cycle and only resetting the hardware
would avoid the problem structurally, at the cost of holding the memory while
suspended.
I am happy to test patches on this hardware.
Workaround in use
-----------------
For anyone hitting this before it is fixed: unbind the device before sleep and
rebind after, which moves the allocation out of the GFP_NOIO window into normal
process context where reclaim works.
# before suspend
echo 0000:24:00.0 > /sys/bus/pci/drivers/atlantic/unbind
# after resume
echo 0000:24:00.0 > /sys/bus/pci/drivers/atlantic/bind
Shrinking the rings (ethtool -G enp36s0 rx 1024 tx 1024, dropping both to
order 4) helps but is not sufficient on its own — it would have prevented cases
1 and 3 but not case 2. Note that atlantic does not implement ethtool -L, so
reducing the number of queues is not available as a lever.
Thanks,
Stephan