Thread (7 messages) 7 messages, 3 authors, 2d ago

Re: [PATCH net v3 0/3] gve: various XDP fixes

flat view

From: Joshua Washington <joshwash@google.com>
Date: 2026-09-30 22:54:27
Also in: bpf, lkml

On Wed, Sep 30, 2026 at 2:54 PM [off-list ref] wrote:
Hi!

This is an automated message. This series looks like a fix, but its
commit messages seem to be missing some information:

 - How the issue was discovered, e.g. hit in production, hit during
   development, syzbot report, manual code inspection, LLM or static
   analysis tool scan.
(1/3) gve: fix XSK buffer leak when rings are stopped

This issue was hit in production.

(2/3) gve: fix XSK buffer leak on error descriptor

This issue was caught by LLM

(3/3)  gve: fix napi_disable deadlock when attempting to disable XSK pools

This issue was found in production.

 - Whether the issue was actually triggered, or is only theoretical
   (e.g. found by code inspection). If it was triggered please include
   the symptoms, like the stack trace or error messages.
(1/3) gve: fix XSK buffer leak when rings are stopped

This issue was actually triggered; symptoms included a buffer leak.
Reproduced by continually performing ip link up/down on an interface
while XSK traffic was flowing. Eventually, no more packets can pass
because all of the XSK buffers from the UMEM pool have been leaked.

(2/3) gve: fix XSK buffer leak on error descriptor

This issue is theoretical, as RX error packets are extremely rare in
my personal experience. But it is plain to see by static analysis that
the buffer will be leaked if an RX error is returned, due to the early
return in the driver. All other XSK-releated paths free the XSK buffer
in gve_rx_xsk_dqo().

(3/3)  gve: fix napi_disable deadlock when attempting to disable XSK pools

This one can be very easily reproduced by enabling an AF_XDP zero-copy
socket, and disabling it. The deadlock becomes more apparent when
attempting to enable a second AF_XDP zero-copy socket, as that
operation will stall waiting to get the netdev instance lock.

Snipped stacktrace from
https://github.com/GoogleCloudPlatform/compute-virtual-ethernet-linux/pull/96:

Workqueue: events xp_release_deferred
napi_disable+0x1d/0x50
gve_xsk_pool_disable+0xed/0x1d0 [gve]
gve_xdp+0x14a/0x1c0 [gve]
xp_disable_drv_zc+0x89/0xe0
xp_clear_dev+0x59/0xf0
xp_release_deferred+0x20/0x90
 - What hardware the change was tested on. For driver fixes please
   mention the device (and if relevant firmware version) used for
   testing, or say that the change was not tested on real hardware.
All of these changes were tested on the GVE driver using the DQO RDA
queue format.
Please do not repost the series just to address the above. Instead,
reply to this email with the missing information, so that reviewers
can take it into account. If the series needs another revision for
other reasons, please include the information in the commit messages
then.

The evaluation is done by an LLM so it may be wrong, if you think
that is the case please reply and explain.


-- 

Joshua Washington | Software Engineer | joshwash@google.com | (414) 366-4423
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help