Thread (1 message) 1 message, 1 author, 8d ago

Re: Fw: [PATCH net 6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit

From: David Butler <hidden>
Date: 2026-10-02 21:10:09
Also in: stable

[Severity: Medium]
Should netdev->max_mtu be clamped to match?
Yes, the review feedback is valid.  The bug outlined by the review is
where the user sets an MTU
in the 9728 to ~16356 range and frames above 9728 are silently dropped
instead of the MTU set failing.

We should still move forward with this patch but reword the commit
message.  It will likely be at
least several weeks before I will be able to acquire and setup a lab
environment for reproduction
and testing for a new patch revision.  This patch as-is does not
introduce any code regressions.
In every case the behavior is unchanged or strictly better.
Practically speaking, most users would
not be impacted by the frame drop.  Standard jumbo (≤9728, e.g. 9000)
is entirely unaffected.

Since 5fa4caff59f2, E810 + ESXi icen + SR-IOV has been unusable: the
VF never finishes
 CONFIG_VSI_QUEUES, regardless of MTU.  The effects are not contained
to the guest
that triggers it. Under pass-through the mis-programmed queue can
raise an IOMMU fault that
 wedges the PF/VF on the ESXi host.  This usually takes out that
interface for subsequent
guests too (even if they have this patch).  Recovery requires a host reboot.

A follow-up to adopt the reviewer's fix is warranted.  I can provide
the patch if someone else
thinks they can test it sooner than I can.

A revised commit message follows:

iavf: cap advertised max_pkt_size at the single-buffer HW limit

Since commit 5fa4caff59f2 ("iavf: switch to Page Pool")
iavf_configure_queues() advertises max_pkt_size to the PF as:

        max_frame = LIBIE_MAX_RX_FRM_LEN(adapter->rx_rings->pp->p.offset);
        max_frame = min_not_zero(adapter->vf_res->max_mtu, max_frame);

LIBIE_MAX_RX_FRM_LEN (16382) is the multi-descriptor scatter/gather frame
ceiling, not a single-queue value, and it exceeds the E810 MAC frame size
maximum of 9728. Per the E810 datasheet (613875-009 section 13.2.2.17.1)
the Tx frame-size register PRTDCB_TDPUC.MAX_TXFRAME has a maximum of
0x2600 (9728); larger frames are discarded. The in-tree ice driver encodes
the same value as ICE_AQ_SET_MAC_FRAME_SIZE_MAX (== LIBIE_MAX_RX_BUF_LEN ==
9728), and the VF clamped max_frame to IAVF_MAX_RXBUFFER (9728) before this
commit.

When the PF advertises vf_res->max_mtu as 0, min_not_zero() leaves
max_frame at 16382. The Linux ice PF advertises max_mtu = port MAC frame
size (<= 9728), so a VF behind ice never sends more than that. The ESXi
"icen" PF on E810 advertises max_mtu as 0, so the VF sends
max_pkt_size = 16382, which icen rejects while programming the queue
context for VIRTCHNL_OP_CONFIG_VSI_QUEUES (opcode 6):

        icen_ConfigureTxQueue: VSI 8: Failed to set LAN Tx queue context for
                               absolute Tx queue 64, Error: ICE_ERR_PARAM
        indrv_SendMsgToVf: VF 0: Failed opcode 6, Error -5

        iavf 0000:03:00.0: PF returned error -5 (IAVF_ERR_PARAM) to
our request 6
        iavf 0000:03:00.0 ethX: NETDEV WATCHDOG: transmit queue N timed out

The VF's queues never come up; under SR-IOV passthrough the mis-programmed
queue can also trigger a fatal IOMMU fault in the guest. Forcing only
max_pkt_size back to 9728 (and leaving the Page Pool rx_buf_len/
databuffer_size untouched) makes the VF come up; databuffer_size is not
involved. This was confirmed on two E810 NVM versions (3.00 and 4.51) under
icen, so the trigger is the icen PF, not the firmware revision. Reported by
several users on E810 + ESXi icen with v6.10+ guests:

Link: https://community.intel.com/t5/Ethernet-Products/E810-C-iavf-driver-issue-on-Linux-6-12/m-p/1737490
Link: https://access.redhat.com/solutions/6973766
Link: https://knowledge.broadcom.com/external/article/404315/sriov-enabled-vms-network-adaptor-goes-d.html

Cap max_frame at the single-buffer HW limit (LIBIE_MAX_RX_BUF_LEN, 9728) so
the VF advertises a max_pkt_size the PF can program into the Rx queue
context. This is the value the VF used before 5fa4caff59f2 (IAVF_MAX_RXBUFFER).

Note this only caps the advertised max_pkt_size. 5fa4caff59f2 also raised
netdev->max_mtu to LIBIE_MAX_MTU (derived from the same 16382 ceiling); when
the PF advertises max_mtu == 0 that leaves max_mtu above 9728, so an MTU in
the 9728 < mtu <= LIBIE_MAX_MTU range can still be set while frames above
9728 are dropped by the queue context. Clamping netdev->max_mtu to match is
left to a follow-up.

Fixes: 5fa4caff59f2 ("iavf: switch to Page Pool")
Signed-off-by: Dave Butler <redacted>
Cc: stable@vger.kernel.org # v6.10+
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help