Thread (8 messages) flat view 8 messages, 2 authors, 13d ago

Re: [PATCH net v2 3/5] net: phy: restore device state after probe failure

From: Jakub Kicinski <kuba@kernel.org>
Date: 2026-08-18 16:09:33
Also in: lkml

This is an AI-generated review of your patch. The human sending this
email has considered the AI review valid, or at least plausible.
---
net: phy: restore device state after probe failure

phy_probe() sets phydev->drv up front and PHY_READY before calling
of_phy_leds(), but never undoes either if a later step fails. The patch
clears phydev->drv on the common error path and resets the state to
PHY_DOWN when LED setup fails.
quoted hunk ↗ jump to hunk
diff --git a/drivers/net/phy/phy_device.c b/drivers/net/phy/phy_device.c
index 9c7ed9c61e6d5..c9e75bd3b81a2 100644
--- a/drivers/net/phy/phy_device.c
+++ b/drivers/net/phy/phy_device.c
@@ -3821,6 +3821,8 @@ static int phy_probe(struct device *dev)
 	if (!phydev->is_on_sfp_module)
 		phy_led_triggers_unregister(phydev);
 
+	phydev->state = PHY_DOWN;
+
[Severity: Medium]
This isn't a bug introduced by this patch, but since the subject says the
device state is restored after probe failure, is the set of fields being
restored complete?

Earlier in phy_probe() a few fields that describe the hardware are
overwritten based on the matched driver:

	if (!phy_drv_supports_irq(phydrv) && phy_interrupt_is_valid(phydev))
		phydev->irq = PHY_POLL;

	if (phydrv->flags & PHY_IS_INTERNAL)
		phydev->is_internal = true;

and later:

	if (!linkmode_test_bit(ETHTOOL_LINK_MODE_Autoneg_BIT,
			       phydev->supported))
		phydev->autoneg = 0;

None of these are undone here. phydev->irq is only ever derived from
bus->irq[addr] in phy_device_create(), and phy_attach_direct() can only
force it to PHY_POLL again:

	if (!phy_drv_supports_irq(phydev->drv) && phy_interrupt_is_valid(phydev))
		phydev->irq = PHY_POLL;

So if a driver's probe fails after those assignments and the same phydev is
later bound again (genphy fallback in phy_attach_direct(), or another module
matching the same phy_id), does the PHY stay stuck in polling mode, keep the
is_internal flag from the failed driver, and keep autoneg disabled?
quoted hunk ↗ jump to hunk
 out_ports:
 	phy_sfp_release(phydev);
 	phy_cleanup_ports(phydev);
@@ -3828,6 +3830,7 @@ static int phy_probe(struct device *dev)
 out_reset:
 	/* Re-assert the reset signal on error */
 	phy_device_reset(phydev, 1);
+	phydev->drv = NULL;
[Severity: Medium]
Can this store race with a concurrent phy_attach_direct() and produce a
NULL dereference?

At this point phydev->mdio.dev.driver is still set: really_probe() only
clears dev->driver later in device_unbind_cleanup(), after the
BUS_NOTIFY_DRIVER_NOT_BOUND notifiers, device_links_no_driver() and
devres_release_all(). Meanwhile phydev->drv is already NULL.

phy_attach_direct() decides on the genphy fallback using the other field:

	if (!d->driver) {
		...
		phydev->is_genphy_driven = 1;
	}

so in that window it skips the fallback probe (nothing repopulates
phydev->drv) and then does:

	if (!phy_drv_supports_irq(phydev->drv) && phy_interrupt_is_valid(phydev))
		phydev->irq = PHY_POLL;

with phy_drv_supports_irq() dereferencing unconditionally:

	static bool phy_drv_supports_irq(const struct phy_driver *phydrv)
	{
		return phydrv->config_intr && phydrv->handle_interrupt;
	}

phy_attach_direct() holds only RTNL, not device_lock() and not phydev->lock,
and phy_probe() takes no phylib lock either. Before this patch the same race
read a stale but non-NULL pointer, so it did not oops.

Related: devres-registered interfaces created by the PHY driver's own probe
(hwmon, PTP) stay live until devres_release_all(), i.e. after this store,
and the paged register helpers dereference phydev->drv unconditionally:

	drivers/net/phy/phy-core.c:__phy_read_page() {
		if (WARN_ONCE(!phydev->drv->read_page,
			      "read_page callback not available, PHY driver not loaded?\n"))
			return -EOPNOTSUPP;

		return phydev->drv->read_page(phydev);
	}

That ordering already exists in phy_remove(), but does this patch newly
extend it to the probe-failure path?

[Severity: Medium]
This is a pre-existing issue and phy_attach_direct() is not touched by this
patch, but while looking at error-path symmetry in phy_probe(): is there any
unwind for a successful phy_probe() followed by a failed
device_bind_driver()?

	drivers/net/phy/phy_device.c:phy_attach_direct() {
		if (phydev->is_genphy_driven) {
			err = d->driver->probe(d);
			if (err >= 0)
				err = device_bind_driver(d);

			if (err)
				goto error_module_put;
		}
	...
	error_module_put:
		module_put(d->driver->owner);
		phydev->is_genphy_driven = 0;
		d->driver = NULL;
	error_put_device:
		put_device(d);
	}

Since the device was never bound, the driver core never calls phy_remove(),
and this path only does module_put()/put_device(). Does that leak the
kmalloc'ed phy->led_link_trigger and phy->phy_led_triggers[] objects that
phy_led_triggers_register() published into the global LED trigger list, plus
the phy_port objects from phy_setup_ports() and any SFP upstream
registration, while drv->remove() is never called?

phydev->drv and PHY_READY are also left set for an unbound device, and a
later attach re-runs phy_probe() which overwrites phy->phy_led_triggers and
phy->led_link_trigger. Would that orphan the earlier registrations in the
global trigger list?
 
 	return err;
 }
-- 
pw-bot: cr
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help