Re: [PATCH net v3 0/3] gve: various XDP fixes
flat view
From: Joshua Washington <joshwash@google.com>
Date: 2026-09-30 22:54:27
Also in:
bpf, lkml
On Wed, Sep 30, 2026 at 2:54 PM [off-list ref] wrote:
Hi! This is an automated message. This series looks like a fix, but its commit messages seem to be missing some information: - How the issue was discovered, e.g. hit in production, hit during development, syzbot report, manual code inspection, LLM or static analysis tool scan.
(1/3) gve: fix XSK buffer leak when rings are stopped This issue was hit in production. (2/3) gve: fix XSK buffer leak on error descriptor This issue was caught by LLM (3/3) gve: fix napi_disable deadlock when attempting to disable XSK pools This issue was found in production.
- Whether the issue was actually triggered, or is only theoretical (e.g. found by code inspection). If it was triggered please include the symptoms, like the stack trace or error messages.
(1/3) gve: fix XSK buffer leak when rings are stopped This issue was actually triggered; symptoms included a buffer leak. Reproduced by continually performing ip link up/down on an interface while XSK traffic was flowing. Eventually, no more packets can pass because all of the XSK buffers from the UMEM pool have been leaked. (2/3) gve: fix XSK buffer leak on error descriptor This issue is theoretical, as RX error packets are extremely rare in my personal experience. But it is plain to see by static analysis that the buffer will be leaked if an RX error is returned, due to the early return in the driver. All other XSK-releated paths free the XSK buffer in gve_rx_xsk_dqo(). (3/3) gve: fix napi_disable deadlock when attempting to disable XSK pools This one can be very easily reproduced by enabling an AF_XDP zero-copy socket, and disabling it. The deadlock becomes more apparent when attempting to enable a second AF_XDP zero-copy socket, as that operation will stall waiting to get the netdev instance lock. Snipped stacktrace from https://github.com/GoogleCloudPlatform/compute-virtual-ethernet-linux/pull/96: Workqueue: events xp_release_deferred napi_disable+0x1d/0x50 gve_xsk_pool_disable+0xed/0x1d0 [gve] gve_xdp+0x14a/0x1c0 [gve] xp_disable_drv_zc+0x89/0xe0 xp_clear_dev+0x59/0xf0 xp_release_deferred+0x20/0x90
- What hardware the change was tested on. For driver fixes please mention the device (and if relevant firmware version) used for testing, or say that the change was not tested on real hardware.
All of these changes were tested on the GVE driver using the DQO RDA queue format.
Please do not repost the series just to address the above. Instead, reply to this email with the missing information, so that reviewers can take it into account. If the series needs another revision for other reasons, please include the information in the commit messages then. The evaluation is done by an LLM so it may be wrong, if you think that is the case please reply and explain.
-- Joshua Washington | Software Engineer | joshwash@google.com | (414) 366-4423