Thread (2 messages) 2 messages, 2 authors, 2d ago

[BUG] atlantic: resume from S3 fails with -ENOMEM in aq_ring_alloc(), leaves unusable netdev

From: Stephan Hauser <hidden>
Date: 2026-09-26 19:49:10

Hi,

The atlantic driver intermittently fails to resume an AQC107 from suspend-to-RAM.
aq_nic_init() reallocates its per-ring software buffer arrays from the PM resume
path, where the PM core has restricted allocations to GFP_NOIO. With the default
ring sizes these are order-5 and order-6 requests (128 KiB and 256 KiB), which
cannot reliably be satisfied in a context that can neither reclaim, compact, nor
dip into reserves.

Resume then returns -ENOMEM and the device is left half-initialised: the netdev
still exists and is still marked IFF_UP, but has no rings and no vectors. It
passes no traffic, DHCP never completes, and bringing the link down hangs.

I have 3 occurrences in 60 suspend/resume cycles (5%) over 5.5 weeks of journal,
across kernels 7.1.4 and 7.2.2. What makes this worth reporting rather than
filing under "memory was tight": the three failures have three *different*
proximate causes in the allocator, detailed below. The allocation is fragile in
several independent ways at once, which is why no amount of tuning fixes it.

Two distinct problems:

  1. A >PAGE_ALLOC_COSTLY_ORDER kmalloc() on the resume path, for memory that
     never needs to be physically contiguous.
  2. atl_resume_common() does not unwind when aq_nic_init() fails, so an
     allocation failure is converted into a wedged interface.


Disclosure
----------

The analysis in this report, and the text of the report itself, were produced
with the assistance of Claude (Anthropic). Please weigh it accordingly.

What is directly evidenced: the hardware is real and in front of me, and every
log line, zone dump and free-page histogram quoted below is verbatim from my
journal. The failure counts come from scanning all boots in that journal.

What is inference rather than measurement, and where I would welcome a second
opinion:

  - The 64-byte element size is derived from the two reported allocation
    orders, not read out of the source.
  - The TX-before-RX ordering in aq_vec_ring_alloc(), and the claim that
    aq_nic_init() aborts on first failure, are inferred from the backtrace and
    from only ever seeing a single warning per event.
  - The ZONE_DMA32 lowmem_reserve arithmetic in case 3 is hand-computed from
    the zone dump.
  - The aq_nic_stop() attribution for the link-down hang is a guess; I say so
    again where it appears.
  - The suggested fixes were not compiled or tested, and were written without
    the source tree to hand. Treat them as a direction, not a patch.

I have read the whole thing and stand behind reporting it, but I have not
personally verified the driver internals against the code.


System
------

  Kernel:    7.1.4 and 7.2.2 (also running 7.2.7), x86_64, PREEMPT(lazy)
  Hardware:  Micro-Star International Co., Ltd. MEG X570 UNIFY (MS-7C35),
             BIOS A.80 01/22/2021
  NIC:       Aquantia AQC107 NBase-T/IEEE 802.3an [Atlantic 10G] (rev 02)
             PCI 0000:24:00.0, [1d6a:07b1], subsystem [1d6a:0001]
  Driver:    atlantic, firmware-version 3.1.100
  Rings:     rx 2048 / tx 4096 (driver defaults; maximum 8184), 8 vectors
  Memory:    64 GB, no swap configured


Reproduction
------------

Not reproducible on demand. Correlates with uptime and page cache size:

  1. Boot, use the machine until the page cache has filled most of RAM.
  2. echo mem > /sys/power/state
  3. Resume.

In one boot, a suspend/resume at t=11793s succeeded and one at t=52899s in the
same boot failed.


The allocation
--------------

From the backtrace this is a large kmalloc (__GFP_COMP|__GFP_ZERO, falling
through __kmalloc_large_node_noprof() straight to the page allocator), i.e. a
kcalloc() of the per-ring software buffer array — CPU-side per-descriptor
bookkeeping, not the DMA descriptor ring, which is allocated separately via
dma_alloc_coherent().

The two reported orders pin the element size at 64 bytes:

  tx ring  4096 * 64 B = 262144 B = order 6
  rx ring  2048 * 64 B = 131072 B = order 5

aq_vec_ring_alloc() allocates TX before RX and aq_nic_init() aborts on the first
failure, so an order-6 report means the very first ring failed, and an order-5
report means some TX rings succeeded first. With 8 vectors the driver asks for
8 order-6 plus 8 order-5 blocks on every single resume.

pm_restrict_gfp_mask() clears __GFP_IO/__GFP_FS for the duration of
suspend/resume, so the driver's GFP_KERNEL is silently downgraded to GFP_NOIO
underneath it. It cannot write back, and compaction is largely defeated.


The three failures
------------------

1) 2026-08-31, kernel 7.1.4, order:6 — fragmentation at exactly order 6

   Node 0 Normal free:3363056kB boost:283840kB min:348116kB
   Node 0 Normal: 46873*4kB (UME) 74929*8kB (UME) 63460*16kB (UME) 33869*32kB (UME) 5476*64kB (UME) 959*128kB (UM) 0*256kB 1*512kB (H) 1*1024kB (H) 0*2048kB 0*4096kB = 3360844kB

   3.2 GB free, comfortably above the watermark, and 959 free order-5 blocks —
   but zero at order 6. The only blocks above that are marked (H),
   MIGRATE_HIGHATOMIC, and so unavailable here.

   Note what this means: had the TX ring been order 5 rather than order 6, this
   resume would have succeeded. The default tx ring size of 4096 sits exactly
   one order above what the zone could serve.

2) 2026-09-25, kernel 7.2.2, order:6 — below the min watermark

   Node 0 Normal free:342528kB boost:283840kB min:348116kB
   Node 0 Normal: 204*4kB (ME) 293*8kB (UME) 5378*16kB (UME) 1673*32kB (UME) 1175*64kB (UMH) 601*128kB (UMH) 173*256kB (M) 2*512kB (M) 0*1024kB 0*2048kB 0*4096kB = 340184kB

   Here there were 173 free order-6 blocks and 601 at order 5 — fragmentation
   was not the problem. The zone was simply 5.5 MB below its min watermark, so
   the watermark check rejects the request at any order. GFP_NOIO cannot reclaim
   its way back above the watermark and cannot touch reserves.

   No ring size would have helped this one.

3) 2026-09-26, kernel 7.2.2, order:5 — high-order lists drained mid-loop

   Node 0 Normal free:803236kB boost:283840kB min:348116kB
   Node 0 Normal: 37466*4kB (UME) 19631*8kB (UME) 10676*16kB (UME) 8307*32kB (UME) 892*64kB (UME) 0*128kB 0*256kB 0*512kB 0*1024kB 0*2048kB 0*4096kB = 800640kB

   800 MB free, above the watermark, but nothing at order 5 or above; 46 GB of
   page cache (the bulk of it readahead) had shredded the zone. The earlier
   order-6 TX allocations in the same loop succeeded and consumed the last of
   the high-order blocks; the following order-5 RX allocation then found the
   order-5 and order-6 lists both empty.

   ZONE_DMA32 had order-5 blocks free but appears to be excluded by
   lowmem_reserve rather than by lack of suitable pages:

     Node 0 DMA32 free:247268kB min:3288kB
     lowmem_reserve[]: 0 0 61020 61020 61020

     free 61817 pages - (min 822 + reserve 61020) = -25 pages

   About 25 pages short of passing the check, so the fallback zone was closed
   too. (This last bit is my arithmetic off the zone dump, not instrumented.)

Note that ZONE_NORMAL is in watermark boost (boost:283840) in all three cases,
i.e. kswapd was already working on recovering high-order blocks and had not
caught up by the time resume ran.


Log (case 3)
------------

[52902.114499] kworker/u97:42: page allocation failure: order:5, mode:0x40d00(GFP_NOIO|__GFP_ZERO|__GFP_COMP), nodemask=(null),cpuset=/,mems_allowed=0
[52902.114508] CPU: 17 UID: 0 PID: 363518 Comm: kworker/u97:42 Not tainted 7.2.2 #1-NixOS PREEMPT(lazy)
[52902.114511] Hardware name: Micro-Star International Co., Ltd. MS-7C35/MEG X570 UNIFY (MS-7C35), BIOS A.80 01/22/2021
[52902.114512] Workqueue: async async_run_entry_fn
[52902.114517] Call Trace:
[52902.114518]  <TASK>
[52902.114521]  dump_stack_lvl+0x5d/0x80
[52902.114526]  warn_alloc+0x158/0x180
[52902.114535]  __alloc_frozen_pages_noprof+0x77a/0x1710
[52902.114544]  alloc_pages_mpol+0xb6/0x170
[52902.114548]  ___kmalloc_large_node+0xb3/0xd0
[52902.114552]  __kmalloc_large_node_noprof+0x1d/0xc0
[52902.114555]  __kmalloc_noprof+0x424/0x590
[52902.114573]  aq_ring_alloc+0x4a/0xd0 [atlantic]
[52902.114581]  aq_vec_ring_alloc+0x9d/0x1d0 [atlantic]
[52902.114589]  aq_nic_init+0x116/0x1e0 [atlantic]
[52902.114597]  atl_resume_common+0x43/0xd0 [atlantic]
[52902.114607]  dpm_run_callback+0x4e/0x160
[52902.114613]  device_resume+0x15c/0x260
[52902.114616]  async_resume+0x21/0x30
[52902.114618]  async_run_entry_fn+0x34/0x150
[52902.114620]  process_one_work+0x199/0x370
[52902.114624]  worker_thread+0x177/0x2e0
[52902.114628]  kthread+0xe2/0x110
[52902.114633]  ret_from_fork+0x251/0x330
[52902.114637]  ret_from_fork_asm+0x1a/0x30
[52902.114642]  </TASK>

[52902.119423] atlantic 0000:24:00.0: PM: dpm_run_callback(): pci_pm_resume returns -12
[52902.119427] atlantic 0000:24:00.0: PM: failed to resume async: error -12

Cases 1 and 2 are identical apart from the order and the aq_nic_init()/
aq_ring_alloc() offsets. Full logs available on request.


Consequences
------------

pci_pm_resume() returning -12 does not tear the device down. The netdev remains
registered and IFF_UP, but aq_nic_init() bailed partway through
aq_vec_ring_alloc(), so the vectors and rings behind it were never allocated:

  3: enp36s0: <NO-CARRIER,BROADCAST,MULTICAST,UP> mtu 1500 qdisc mq state DOWN

No packets are received, DHCP never completes, and `ip link set enp36s0 down`
hangs indefinitely — presumably in the aq_nic_stop() path operating on NAPI
contexts and rings that do not exist, though I do not have a blocked-task trace
to confirm that. I will send one if I can capture `echo w > /proc/sysrq-trigger`
the next time it reproduces.

Recovering without a reboot requires re-probing the device; the netdev is beyond
help at that point:

  echo 1 > /sys/bus/pci/devices/0000:24:00.0/remove
  echo 1 > /sys/bus/pci/rescan


Suggested fixes
---------------

1. Use kvcalloc()/kvfree() for the ring buffer arrays in aq_ring_alloc() and
   friends. The array is never handed to the device, so it has no need to be
   physically contiguous, and a vmalloc fallback removes the high-order
   dependency entirely. This addresses cases 1 and 3 outright, and should make
   case 2 recoverable as well: order-0 requests can be satisfied from clean
   unmapped page cache, which is reclaimable without __GFP_FS, whereas nothing
   the allocator can do under GFP_NOIO produces a fresh order-6 block.
   This looks like the minimal correct fix.

2. Make atl_resume_common() unwind on error. Whatever the allocator does, a
   failed resume should not leave a registered, IFF_UP netdev with no rings
   behind it. At minimum the error path should undo the partial aq_nic_init()
   and leave the interface administratively down, so a subsequent `ip link set
   down` / `up` can recover the device rather than hanging.

3. Worth considering separately: reallocating the rings on resume at all. They
   are freed at suspend and immediately reallocated at identical sizes on
   resume, in the one context where allocation is most constrained. Keeping the
   ring memory across the suspend/resume cycle and only resetting the hardware
   would avoid the problem structurally, at the cost of holding the memory while
   suspended.

I am happy to test patches on this hardware.


Workaround in use
-----------------

For anyone hitting this before it is fixed: unbind the device before sleep and
rebind after, which moves the allocation out of the GFP_NOIO window into normal
process context where reclaim works.

  # before suspend
  echo 0000:24:00.0 > /sys/bus/pci/drivers/atlantic/unbind
  # after resume
  echo 0000:24:00.0 > /sys/bus/pci/drivers/atlantic/bind

Shrinking the rings (ethtool -G enp36s0 rx 1024 tx 1024, dropping both to
order 4) helps but is not sufficient on its own — it would have prevented cases
1 and 3 but not case 2. Note that atlantic does not implement ethtool -L, so
reducing the number of queues is not available as a lever.

Thanks,
Stephan
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help