From: Clemens Gruber <hidden> Date: 2016-02-12 18:09:12
Commit 113c74d83eef ("net: phy: turn carrier off on phy attach") breaks
the eth0 link coming up on all my i.MX6Q boards with a Marvell 88E1510.
If I then do a ifconfig eth0 down/up cycle I first get a MDIO read
timeout but then the link becomes ready and everything is back to
normal.
Without this step however, the link stays down forever, an unusually
high amount of phy interrupts occur (about 10000/second) and kworker/0:2
is constantly using over 60% of the CPU.
Reverting it fixes the problems with the link not coming up at boot as
well as the high amount of phy interrupts and kworker load in that
state.
This reverts commit 113c74d83eef870e43a0d9279044e9d5435f0d07.
Signed-off-by: Clemens Gruber <redacted>
---
drivers/net/phy/phy_device.c | 5 -----
1 file changed, 5 deletions(-)
@@ -901,11 +901,6 @@ int phy_attach_direct(struct net_device *dev, struct phy_device *phydev,phydev->state=PHY_READY;-/* Initial carrier state is off as the phy is about to be-*(re)initialized.-*/-netif_carrier_off(phydev->attached_dev);-/* Do initial configuration here, now that*wehavecertainkeyparameters*(dev_flagsandinterface)
Commit 113c74d83eef ("net: phy: turn carrier off on phy attach") breaks
the eth0 link coming up on all my i.MX6Q boards with a Marvell 88E1510.
If I then do a ifconfig eth0 down/up cycle I first get a MDIO read
timeout but then the link becomes ready and everything is back to
normal.
Without this step however, the link stays down forever, an unusually
high amount of phy interrupts occur (about 10000/second) and kworker/0:2
is constantly using over 60% of the CPU.
Reverting it fixes the problems with the link not coming up at boot as
well as the high amount of phy interrupts and kworker load in that
state.
You are seeing this with the FEC driver right? We probably want to
carefully audit the driver and understand what could be going wrong, the
initial change is correct, so there must be something else going on here.
quoted hunk
This reverts commit 113c74d83eef870e43a0d9279044e9d5435f0d07.
Signed-off-by: Clemens Gruber <redacted>
---
drivers/net/phy/phy_device.c | 5 -----
1 file changed, 5 deletions(-)
@@ -901,11 +901,6 @@ int phy_attach_direct(struct net_device *dev, struct phy_device *phydev,phydev->state=PHY_READY;-/* Initial carrier state is off as the phy is about to be-*(re)initialized.-*/-netif_carrier_off(phydev->attached_dev);-/* Do initial configuration here, now that*wehavecertainkeyparameters*(dev_flagsandinterface)
From: Clemens Gruber <hidden> Date: 2016-02-12 21:07:40
Hi Florian,
On Fri, Feb 12, 2016 at 10:56:04AM -0800, Florian Fainelli wrote:
On 12/02/16 10:01, Clemens Gruber wrote:
quoted
Commit 113c74d83eef ("net: phy: turn carrier off on phy attach") breaks
the eth0 link coming up on all my i.MX6Q boards with a Marvell 88E1510.
If I then do a ifconfig eth0 down/up cycle I first get a MDIO read
timeout but then the link becomes ready and everything is back to
normal.
Without this step however, the link stays down forever, an unusually
high amount of phy interrupts occur (about 10000/second) and kworker/0:2
is constantly using over 60% of the CPU.
Reverting it fixes the problems with the link not coming up at boot as
well as the high amount of phy interrupts and kworker load in that
state.
You are seeing this with the FEC driver right? We probably want to
carefully audit the driver and understand what could be going wrong, the
initial change is correct, so there must be something else going on here.
Yes, this occurs with the fec driver.
In the fec_probe function at line 3471 of
net/ethernet/freescale/fec_main.c, netif_carrier_off is called and a
comment states "Carrier starts down, phylib will bring it up".
Could this be the source of the problem, both fec and phy expecting the
other one to turn on the carrier?
If you have an idea about what could be going wrong in the fec driver,
please let me know.
Thanks,
Clemens
From: Clemens Gruber <hidden> Date: 2016-02-14 18:25:36
On Fri, Feb 12, 2016 at 10:56:04AM -0800, Florian Fainelli wrote:
On 12/02/16 10:01, Clemens Gruber wrote:
quoted
Commit 113c74d83eef ("net: phy: turn carrier off on phy attach") breaks
the eth0 link coming up on all my i.MX6Q boards with a Marvell 88E1510.
If I then do a ifconfig eth0 down/up cycle I first get a MDIO read
timeout but then the link becomes ready and everything is back to
normal.
Without this step however, the link stays down forever, an unusually
high amount of phy interrupts occur (about 10000/second) and kworker/0:2
is constantly using over 60% of the CPU.
Reverting it fixes the problems with the link not coming up at boot as
well as the high amount of phy interrupts and kworker load in that
state.
You are seeing this with the FEC driver right? We probably want to
carefully audit the driver and understand what could be going wrong, the
initial change is correct, so there must be something else going on here.
I think I found the underlying problem!
It was not the fec driver but the marvell phy driver, more specifically
the marvell_of_reg_init call being made too late, which lead to the
observed problem at half the boot ups where the link never came up.
With the marvell,reg-init device tree parameter, a flag needs to be set
to tell the Marvell 88E1510 that it should enable the interrupt output.
(At a specific pin, in my case LED[2])
If this is not set (or set too late), the phydev->state is set to UP in
phy_start (called from fec_enet_open) but then, the auto-negotiation
never starts.
In comparison, now, after I called marvell_of_reg_init not in
m88e1510_config_aneg but in marvell_probe, everything works again :)
About a second after the fec_enet_open/phy_start calls, the
auto-negotiation starts (m88e1510_config_aneg) and the phy state changes
from UP to AN, to CHANGELINK and finally to RUNNING.
I will send a patch shortly, calling marvell_of_reg_init from a new
m88e1510_probe function instead of the m88e1510_config_aneg function.
Thanks.
Clemens
On February 14, 2016 10:25:30 AM PST, Clemens Gruber [off-list ref] wrote:
On Fri, Feb 12, 2016 at 10:56:04AM -0800, Florian Fainelli wrote:
quoted
On 12/02/16 10:01, Clemens Gruber wrote:
quoted
Commit 113c74d83eef ("net: phy: turn carrier off on phy attach")
breaks
quoted
quoted
the eth0 link coming up on all my i.MX6Q boards with a Marvell
88E1510.
quoted
quoted
If I then do a ifconfig eth0 down/up cycle I first get a MDIO read
timeout but then the link becomes ready and everything is back to
normal.
Without this step however, the link stays down forever, an
unusually
quoted
quoted
high amount of phy interrupts occur (about 10000/second) and
kworker/0:2
quoted
quoted
is constantly using over 60% of the CPU.
Reverting it fixes the problems with the link not coming up at boot
as
quoted
quoted
well as the high amount of phy interrupts and kworker load in that
state.
You are seeing this with the FEC driver right? We probably want to
carefully audit the driver and understand what could be going wrong,
the
quoted
initial change is correct, so there must be something else going on
here.
I think I found the underlying problem!
It was not the fec driver but the marvell phy driver, more specifically
the marvell_of_reg_init call being made too late, which lead to the
observed problem at half the boot ups where the link never came up.
With the marvell,reg-init device tree parameter, a flag needs to be set
to tell the Marvell 88E1510 that it should enable the interrupt output.
(At a specific pin, in my case LED[2])
If this is not set (or set too late), the phydev->state is set to UP in
phy_start (called from fec_enet_open) but then, the auto-negotiation
never starts.
In comparison, now, after I called marvell_of_reg_init not in
m88e1510_config_aneg but in marvell_probe, everything works again :)
About a second after the fec_enet_open/phy_start calls, the
auto-negotiation starts (m88e1510_config_aneg) and the phy state
changes
from UP to AN, to CHANGELINK and finally to RUNNING.
I will send a patch shortly, calling marvell_of_reg_init from a new
m88e1510_probe function instead of the m88e1510_config_aneg function.
config_init is more appropriate here since this call back will be called even if there is a software reset (e.g: from phy_init_hw). config_aneg is definitely too late, thanks for finding this!
--
Florian