From: Stefan Wahren <hidden> Date: 2016-06-11 01:01:31
Hi,
[add Rafael, J?rg and J?rgen to CC]
Stefan Wahren [off-list ref] hat am 1. Juni 2016 um 22:51
geschrieben:
Hi,
i'm currently working on standby support for MXS platform. If i trigger the
standby per sysfs everything works fine, but mostly on resume i get a bunch of
warnings (see below for full output):
[ 67.040000] Interrupts enabled before system core resume.
My current working branch [1] based on Linux 4.5 ( cmdline has
no_console_suspend=1 ). The test hardware is a i.MX23 board (
iMX233-OLinuXino-MAXI ).
How can i narrow down this issue to the relevant driver / interrupt ?
What is the right way to fix this issue?
I have the suspicion it's triggered by clocksource/mxs-timer.c because the
vendor kernel 2.6.35 had a function call [2] to suspend the timers.
[1] - https://github.com/lategoodbye/linux-mxs-power/tree/rebase-4.5
[2] -
http://git.freescale.com/git/cgit.cgi/imx/linux-2.6-imx.git/tree/arch/arm/mach-mx23/pm.c?h=imx_2.6.35_maintain#n284
Full output in error case:
# echo standby > /sys/power/state
[ 65.820000] PM: Syncing filesystems ... done.
[ 66.890000] Freezing user space processes ... (elapsed 0.018 seconds) done.
[ 66.920000] Freezing remaining freezable tasks ... (elapsed 0.004 seconds)
done.
[ 66.960000] smsc95xx 1-1.1:1.0 eth0: entering SUSPEND2 mode
[ 66.990000] PM: suspend of devices complete after 56.500 msecs
[ 67.010000] PM: late suspend of devices complete after 11.906 msecs
[ 67.030000] PM: noirq suspend of devices complete after 12.250 msecs
[ 67.040000] ------------[ cut here ]------------
[ 67.040000] WARNING: CPU: 0 PID: 271 at drivers/base/syscore.c:99
syscore_resume+0x14c/0x1bc()
[ 67.040000] Interrupts enabled before system core resume.
[ 67.040000] Modules linked in: mxs_lradc(C) industrialio_triggered_buffer
[ 67.040000] CPU: 0 PID: 271 Comm: bash Tainted: G C
4.5.0-g8b79a7d-dirty #86
[ 67.040000] Hardware name: Freescale MXS (Device Tree)
[ 67.040000] [<c000fdc4>] (unwind_backtrace) from [<c000de80>]
(show_stack+0x10/0x14)
[ 67.040000] [<c000de80>] (show_stack) from [<c001d094>]
(warn_slowpath_common+0x7c/0xb4)
[ 67.040000] [<c001d094>] (warn_slowpath_common) from [<c001d160>]
(warn_slowpath_fmt+0x30/0x40)
[ 67.040000] [<c001d160>] (warn_slowpath_fmt) from [<c035c6b8>]
(syscore_resume+0x14c/0x1bc)
[ 67.040000] [<c035c6b8>] (syscore_resume) from [<c005d37c>]
(suspend_devices_and_enter+0x480/0x770)
[ 67.040000] [<c005d37c>] (suspend_devices_and_enter) from [<c005dad4>]
(pm_suspend+0x468/0x4e4)
[ 67.040000] [<c005dad4>] (pm_suspend) from [<c005c324>]
(state_store+0x80/0xcc)
[ 67.040000] [<c005c324>] (state_store) from [<c02dc0d8>]
(kobj_attr_store+0x18/0x1c)
[ 67.040000] [<c02dc0d8>] (kobj_attr_store) from [<c0181988>]
(sysfs_kf_write+0x48/0x4c)
[ 67.040000] [<c0181988>] (sysfs_kf_write) from [<c0180b34>]
(kernfs_fop_write+0xe0/0x1a0)
[ 67.040000] [<c0180b34>] (kernfs_fop_write) from [<c0114b08>]
(__vfs_write+0x2c/0xe8)
[ 67.040000] [<c0114b08>] (__vfs_write) from [<c0114c6c>]
(vfs_write+0xa8/0x198)
[ 67.040000] [<c0114c6c>] (vfs_write) from [<c0114e30>]
(SyS_write+0x44/0x88)
[ 67.040000] [<c0114e30>] (SyS_write) from [<c000a280>]
(ret_fast_syscall+0x0/0x1c)
[ 67.040000] ---[ end trace 22f59f91777b0ede ]---
any hints or ideas to narrow down this issue are very welcome.
Regards Stefan
From: Rafael J. Wysocki <hidden> Date: 2016-06-11 01:42:01
On Saturday, June 11, 2016 03:01:31 AM Stefan Wahren wrote:
Hi,
[add Rafael, J?rg and J?rgen to CC]
quoted
Stefan Wahren [off-list ref] hat am 1. Juni 2016 um 22:51
geschrieben:
Hi,
i'm currently working on standby support for MXS platform. If i trigger the
standby per sysfs everything works fine, but mostly on resume i get a bunch of
warnings (see below for full output):
[ 67.040000] Interrupts enabled before system core resume.
My current working branch [1] based on Linux 4.5 ( cmdline has
no_console_suspend=1 ). The test hardware is a i.MX23 board (
iMX233-OLinuXino-MAXI ).
How can i narrow down this issue to the relevant driver / interrupt ?
What is the right way to fix this issue?
I have the suspicion it's triggered by clocksource/mxs-timer.c because the
vendor kernel 2.6.35 had a function call [2] to suspend the timers.
[1] - https://github.com/lategoodbye/linux-mxs-power/tree/rebase-4.5
[2] -
http://git.freescale.com/git/cgit.cgi/imx/linux-2.6-imx.git/tree/arch/arm/mach-mx23/pm.c?h=imx_2.6.35_maintain#n284
Full output in error case:
# echo standby > /sys/power/state
[ 65.820000] PM: Syncing filesystems ... done.
[ 66.890000] Freezing user space processes ... (elapsed 0.018 seconds) done.
[ 66.920000] Freezing remaining freezable tasks ... (elapsed 0.004 seconds)
done.
[ 66.960000] smsc95xx 1-1.1:1.0 eth0: entering SUSPEND2 mode
[ 66.990000] PM: suspend of devices complete after 56.500 msecs
[ 67.010000] PM: late suspend of devices complete after 11.906 msecs
[ 67.030000] PM: noirq suspend of devices complete after 12.250 msecs
[ 67.040000] ------------[ cut here ]------------
[ 67.040000] WARNING: CPU: 0 PID: 271 at drivers/base/syscore.c:99
syscore_resume+0x14c/0x1bc()
[ 67.040000] Interrupts enabled before system core resume.
[ 67.040000] Modules linked in: mxs_lradc(C) industrialio_triggered_buffer
[ 67.040000] CPU: 0 PID: 271 Comm: bash Tainted: G C
4.5.0-g8b79a7d-dirty #86
[ 67.040000] Hardware name: Freescale MXS (Device Tree)
[ 67.040000] [<c000fdc4>] (unwind_backtrace) from [<c000de80>]
(show_stack+0x10/0x14)
[ 67.040000] [<c000de80>] (show_stack) from [<c001d094>]
(warn_slowpath_common+0x7c/0xb4)
[ 67.040000] [<c001d094>] (warn_slowpath_common) from [<c001d160>]
(warn_slowpath_fmt+0x30/0x40)
[ 67.040000] [<c001d160>] (warn_slowpath_fmt) from [<c035c6b8>]
(syscore_resume+0x14c/0x1bc)
[ 67.040000] [<c035c6b8>] (syscore_resume) from [<c005d37c>]
(suspend_devices_and_enter+0x480/0x770)
[ 67.040000] [<c005d37c>] (suspend_devices_and_enter) from [<c005dad4>]
(pm_suspend+0x468/0x4e4)
[ 67.040000] [<c005dad4>] (pm_suspend) from [<c005c324>]
(state_store+0x80/0xcc)
[ 67.040000] [<c005c324>] (state_store) from [<c02dc0d8>]
(kobj_attr_store+0x18/0x1c)
[ 67.040000] [<c02dc0d8>] (kobj_attr_store) from [<c0181988>]
(sysfs_kf_write+0x48/0x4c)
[ 67.040000] [<c0181988>] (sysfs_kf_write) from [<c0180b34>]
(kernfs_fop_write+0xe0/0x1a0)
[ 67.040000] [<c0180b34>] (kernfs_fop_write) from [<c0114b08>]
(__vfs_write+0x2c/0xe8)
[ 67.040000] [<c0114b08>] (__vfs_write) from [<c0114c6c>]
(vfs_write+0xa8/0x198)
[ 67.040000] [<c0114c6c>] (vfs_write) from [<c0114e30>]
(SyS_write+0x44/0x88)
[ 67.040000] [<c0114e30>] (SyS_write) from [<c000a280>]
(ret_fast_syscall+0x0/0x1c)
[ 67.040000] ---[ end trace 22f59f91777b0ede ]---
any hints or ideas to narrow down this issue are very welcome.
Interrupts should not be enabled before syscore_resume() is called, but they
are, apparently by the platform code.
Thanks,
Rafael
From: Stefan Wahren <hidden> Date: 2016-06-12 10:27:02
Hi Rafael,
"Rafael J. Wysocki" [off-list ref] hat am 11. Juni 2016 um 03:42
geschrieben:
On Saturday, June 11, 2016 03:01:31 AM Stefan Wahren wrote:
quoted
Hi,
[add Rafael, J?rg and J?rgen to CC]
quoted
Stefan Wahren [off-list ref] hat am 1. Juni 2016 um 22:51
geschrieben:
Hi,
i'm currently working on standby support for MXS platform. If i trigger
the
standby per sysfs everything works fine, but mostly on resume i get a
bunch of
warnings (see below for full output):
[ 67.040000] Interrupts enabled before system core resume.
My current working branch [1] based on Linux 4.5 ( cmdline has
no_console_suspend=1 ). The test hardware is a i.MX23 board (
iMX233-OLinuXino-MAXI ).
How can i narrow down this issue to the relevant driver / interrupt ?
What is the right way to fix this issue?
any hints or ideas to narrow down this issue are very welcome.
Interrupts should not be enabled before syscore_resume() is called, but they
are, apparently by the platform code.
sure, but it isn't that simple. I can't reproduce the warning everytime. At
least i need 5 tries and more to reproduce it.
I place the following code in the suspend code, before the low level suspend
code in SRAM gets executed:
if (!irqs_disabled())
pr_info("IRQs not disabled before suspend\n");
The info message above only get printed if the warning "Interrupts enabled
before system core resume." appears.
After that i dumped all enabled IRQs (at the same code place) from the interrupt
collector (irqchip) and compared them with cat /proc/interrupts. I noticed that
the dump contains all interrupt from /proc/interrupts plus the IRQs from the 3
GPIO banks, which were not listed under PROC FS.
So i take a look at the GPIO driver gpio-mxs.c and noticed that it seems to miss
the flag IRQCHIP_MASK_ON_SUSPEND compared to a driver gpio-mxc.c for very
similiar hardware. Unfortunately adding the flag in the GPIO driver doesn't fix
the issue.
So any hints about narrowing down the issue would be helpful.
Stefan
From: linux@armlinux.org.uk (Russell King - ARM Linux) Date: 2016-06-12 13:54:28
On Sun, Jun 12, 2016 at 12:27:02PM +0200, Stefan Wahren wrote:
Hi Rafael,
quoted
"Rafael J. Wysocki" [off-list ref] hat am 11. Juni 2016 um 03:42
geschrieben:
Interrupts should not be enabled before syscore_resume() is called, but they
are, apparently by the platform code.
sure, but it isn't that simple. I can't reproduce the warning everytime. At
least i need 5 tries and more to reproduce it.
I place the following code in the suspend code, before the low level suspend
code in SRAM gets executed:
if (!irqs_disabled())
pr_info("IRQs not disabled before suspend\n");
The info message above only get printed if the warning "Interrupts enabled
before system core resume." appears.
After that i dumped all enabled IRQs (at the same code place) from the interrupt
collector (irqchip) and compared them with cat /proc/interrupts.
This isn't about the state of the interrupt controller. It's about
the state of the IRQ mask bit in the CPUs CPSR.
irq_disabled() returns true when the I bit in CPSR is set, false
otherwise.
What it's pointing towards is some driver being unreasonable, and
clearing the CPUs CPSR I bit after the core PM code has set it. Causes
can be using spin_lock_irq()/spin_unlock_irq()/local_irq_enable() etc
inappropriately, rather than using the irqsave/irqrestore versions.
suspend_enter() does this:
arch_suspend_disable_irqs();
BUG_ON(!irqs_disabled());
error = syscore_suspend();
So, we can be sure that if we reach syscore_suspend(), then IRQs were
disabled.
Now, syscore_suspend() calls a set of suspend callbacks, and verifies
that interrupts are not re-enabled after each one. So, you should be
getting a complaint from the kernel if one of those does trigger, and
these are about the last things that happen before the ->enter
callback is called. IRQs should be disabled when mxs_suspend_enter()
is entered.
I'm confused by your statement about "the low level suspend code in SRAM
gets executed" - from what I can see in arch/arm/mach-mxs (you said MXS
in the subject), there is no SRAM code that gets executed for S2RAM.
Since you also said MX23, that ties up with mach-mxs, so I can only
conclude that you're using patches on top of mainline which change the
platform suspend code.
--
RMK's Patch system: http://www.armlinux.org.uk/developer/patches/
FTTC broadband for 0.8mile line: currently at 9.6Mbps down 400kbps up
according to speedtest.net.
From: Stefan Wahren <hidden> Date: 2016-06-12 14:46:18
Hi Russell,
Russell King - ARM Linux [off-list ref] hat am 12. Juni 2016 um 15:54
geschrieben:
On Sun, Jun 12, 2016 at 12:27:02PM +0200, Stefan Wahren wrote:
quoted
Hi Rafael,
quoted
"Rafael J. Wysocki" [off-list ref] hat am 11. Juni 2016 um 03:42
geschrieben:
Interrupts should not be enabled before syscore_resume() is called, but
they
are, apparently by the platform code.
sure, but it isn't that simple. I can't reproduce the warning everytime. At
least i need 5 tries and more to reproduce it.
I place the following code in the suspend code, before the low level suspend
code in SRAM gets executed:
if (!irqs_disabled())
pr_info("IRQs not disabled before suspend\n");
The info message above only get printed if the warning "Interrupts enabled
before system core resume." appears.
After that i dumped all enabled IRQs (at the same code place) from the
interrupt
collector (irqchip) and compared them with cat /proc/interrupts.
This isn't about the state of the interrupt controller. It's about
the state of the IRQ mask bit in the CPUs CPSR.
irq_disabled() returns true when the I bit in CPSR is set, false
otherwise.
What it's pointing towards is some driver being unreasonable, and
clearing the CPUs CPSR I bit after the core PM code has set it. Causes
can be using spin_lock_irq()/spin_unlock_irq()/local_irq_enable() etc
inappropriately, rather than using the irqsave/irqrestore versions.
suspend_enter() does this:
arch_suspend_disable_irqs();
BUG_ON(!irqs_disabled());
error = syscore_suspend();
So, we can be sure that if we reach syscore_suspend(), then IRQs were
disabled.
Now, syscore_suspend() calls a set of suspend callbacks, and verifies
that interrupts are not re-enabled after each one. So, you should be
getting a complaint from the kernel if one of those does trigger, and
these are about the last things that happen before the ->enter
callback is called. IRQs should be disabled when mxs_suspend_enter()
is entered.
thanks for this explanation.
I'm confused by your statement about "the low level suspend code in SRAM
gets executed" - from what I can see in arch/arm/mach-mxs (you said MXS
in the subject), there is no SRAM code that gets executed for S2RAM.
Since you also said MX23, that ties up with mach-mxs, so I can only
conclude that you're using patches on top of mainline which change the
platform suspend code.
Yes, as i wrote in my first email i'm working on this feature and i hope to
contribute it to mainline in the near future. My changes based on the old
Freescale BSP and the mainline i.MX5x S2RAM code [1].
Stefan
[1] - https://github.com/lategoodbye/linux-mxs-power/tree/rebase-4.5
From: Stefan Wahren <hidden> Date: 2016-06-12 21:45:50
Hi,
Russell King - ARM Linux [off-list ref] hat am 12. Juni 2016 um 15:54
geschrieben:
On Sun, Jun 12, 2016 at 12:27:02PM +0200, Stefan Wahren wrote:
quoted
Hi Rafael,
quoted
"Rafael J. Wysocki" [off-list ref] hat am 11. Juni 2016 um 03:42
geschrieben:
Interrupts should not be enabled before syscore_resume() is called, but
they
are, apparently by the platform code.
sure, but it isn't that simple. I can't reproduce the warning everytime. At
least i need 5 tries and more to reproduce it.
I place the following code in the suspend code, before the low level suspend
code in SRAM gets executed:
if (!irqs_disabled())
pr_info("IRQs not disabled before suspend\n");
The info message above only get printed if the warning "Interrupts enabled
before system core resume." appears.
After that i dumped all enabled IRQs (at the same code place) from the
interrupt
collector (irqchip) and compared them with cat /proc/interrupts.
This isn't about the state of the interrupt controller. It's about
the state of the IRQ mask bit in the CPUs CPSR.
irq_disabled() returns true when the I bit in CPSR is set, false
otherwise.
What it's pointing towards is some driver being unreasonable, and
clearing the CPUs CPSR I bit after the core PM code has set it. Causes
can be using spin_lock_irq()/spin_unlock_irq()/local_irq_enable() etc
inappropriately, rather than using the irqsave/irqrestore versions.
i finally found the reason for the warning: clk_get_sys was called inside of
mxs_suspend_enter.
After moving it into the init function the issue disappear.
From: linux@armlinux.org.uk (Russell King - ARM Linux) Date: 2016-06-12 22:00:26
On Sun, Jun 12, 2016 at 11:45:50PM +0200, Stefan Wahren wrote:
Hi,
quoted
Russell King - ARM Linux [off-list ref] hat am 12. Juni 2016 um 15:54
geschrieben:
This isn't about the state of the interrupt controller. It's about
the state of the IRQ mask bit in the CPUs CPSR.
irq_disabled() returns true when the I bit in CPSR is set, false
otherwise.
What it's pointing towards is some driver being unreasonable, and
clearing the CPUs CPSR I bit after the core PM code has set it. Causes
can be using spin_lock_irq()/spin_unlock_irq()/local_irq_enable() etc
inappropriately, rather than using the irqsave/irqrestore versions.
i finally found the reason for the warning: clk_get_sys was called inside of
mxs_suspend_enter.
After moving it into the init function the issue disappear.
Hmm. I suspect a faster way to have found that issue would be to build
your development kernels with more debugging options enabled - eg,
check that you have DEBUG_ATOMIC_SLEEP enabled.
That should have caught the attempt to take a mutex in an IRQs-disabled
region.
--
RMK's Patch system: http://www.armlinux.org.uk/developer/patches/
FTTC broadband for 0.8mile line: currently at 9.6Mbps down 400kbps up
according to speedtest.net.