Thread (31 messages) flat view 31 messages, 2 authors, 3d ago

Re: [PATCH net-next v6 14/14] misc: lan966x-pci: dts: add fdma interrupt to overlay

From: Daniel Machon <daniel.machon@microchip.com>
Date: 2026-09-15 10:52:28
Also in: bpf, linux-arm-kernel, lkml

Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider:

Critical: 0 · High: 0 · Medium: 1 · Low: 0

- [Medium] The FDMA interrupt is hooked up (and thereby unmasked in the
  OIC) long before any FDMA state exists.

--

misc: lan966x-pci: dts: add fdma interrupt to overlay

This adds OIC interrupt 14 as the "fdma" named interrupt to the lan966x
PCI device tree overlay, so FDMA-based injection/extraction can be used
when the switch is reached over PCIe.
quoted
diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
index 7bb726550caf..5bb12dbc0843 100644
--- a/drivers/misc/lan966x_pci.dtso
+++ b/drivers/misc/lan966x_pci.dtso
@@ -141,8 +141,9 @@ switch: switch@e0000000 {

                                      interrupt-parent = <&oic>;
                                      interrupts = <12 IRQ_TYPE_LEVEL_HIGH>,
+                                                  <14 IRQ_TYPE_LEVEL_HIGH>,
                                                   <9 IRQ_TYPE_LEVEL_HIGH>;
-                                     interrupt-names = "xtr", "ana";
+                                     interrupt-names = "xtr", "fdma", "ana";
[Severity: Medium]
With source 14 now named "fdma", platform_get_irq_byname(pdev, "fdma")
starts succeeding on the PCIe path, and lan966x_probe() requests the
handler right there:

drivers/net/ethernet/microchip/lan966x/lan966x_main.c:lan966x_probe() {
        lan966x->fdma_irq = platform_get_irq_byname(pdev, "fdma");
        if (lan966x->fdma_irq > 0) {
                err = devm_request_irq(&pdev->dev, lan966x->fdma_irq,
                                       lan966x_fdma_irq_handler, 0,
                                       "fdma irq", lan966x);
                ...
                lan966x->fdma = true;
        }
}

Is this ordering safe?  The OIC maps and unmasks source 14 at
devm_request_irq() time, but at that point no FDMA state exists yet:

- netif_napi_add()/napi_enable() only run later, from
  lan966x_port_init() -> lan966x_fdma_netdev_init() in the per-port
  probe loop, so lan966x->napi is still the zeroed devm_kzalloc()
  memory and napi.poll is NULL.

- the DCB rings, the LLP/ATU windows and the FDMA_CTRL NRESET pulse are
  all done even later, from lan966x->ops->fdma_init() ->
  lan966x_fdma_pci_init().

If the OIC delivers an FDMA interrupt in that window, the handler
schedules NAPI unconditionally:

drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c:lan966x_fdma_irq_handler() {
        db = lan_rd(lan966x, FDMA_INTR_DB);
        err = lan_rd(lan966x, FDMA_INTR_ERR);

        if (db) {
                lan_wr(0, lan966x, FDMA_INTR_DB_ENA);
                lan_wr(db, lan966x, FDMA_INTR_DB);

                napi_schedule(&lan966x->napi);
        }
}

Can this then oops in __napi_poll(), either calling n->poll() while it
is still NULL, or running lan966x_fdma_pci_napi_poll() against
rx->fdma.dcbs == NULL and n_dcbs == 0, where fdma_has_frames() ->
fdma_db_get() dereferences the dcbs array?

How reachable is that window in practice?  The series itself describes
the card arriving with FDMA state from a previous boot.  The commit
"net: lan966x: clear FDMA interrupt stickies after switch reset" says
the stickies "sit in FDMA_INTR_ERR until the FDMA IRQ is requested later
in probe, at which point the handler fires immediately and WARNs", and
"net: lan966x: add shutdown callback to stop FDMA on reboot" says that
on the endpoint "channels stay active and interrupt sources stay armed
across the reset" and that "FDMA_INTR_ENA persists on the card across a
warm reboot".

The sticky clear in lan966x_reset_switch() clears the latched bits at
one instant, but does not mask FDMA_INTR_ENA/FDMA_INTR_DB_ENA and does
not stop the channels.  After a reboot that skips device_shutdown()
(panic reboot, kexec/kdump, watchdog warm reset with the card still
powered), could a still-active channel latch FDMA_INTR_DB again after
the clear and have it delivered as soon as source 14 is unmasked?

Would it be more robust to reset and mask the FDMA block (the
FDMA_CTRL NRESET pulse that already exists in lan966x_fdma_pci_init(),
plus clearing FDMA_INTR_ENA/FDMA_INTR_DB_ENA) before
devm_request_irq(), or to request the FDMA IRQ only after fdma_init()
has built the rings and added the NAPI instance?
I think this is pre-existing. It can oops, both variants, and yes - a channel the
previous kernel left active can re-latch FDMA_INTR_DB before devm_request_irq()
unmasks the source. Your second suggestion is the right fix; the reset half of the
first is what fdma_init()'s NRESET pulse already does.

None of this originates here: the request has sat before netif_napi_add() and
ops->fdma_init() since c8349639324a ("net: lan966x: Add FDMA functionality"), and
the SoC device tree has named "fdma" since 2952d32c0611 ("ARM: dts: lan966x: add
switch node") - this patch adds one entry to an interrupt-names list. Nothing to
change in this patch; it goes to net with the same reorder as the MTU-change problem
earlier in the series.
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260909-lan966x-pci-fdma-v6-0-6f48dab9d671%40microchip.com
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help