[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

Subsystems: arm port, omap2+ support, the rest

STALE3491d

14 messages, 2 authors, 2017-02-17 · open the first message on its own page

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: tony@atomide.com (Tony Lindgren)
Date: 2017-02-13 21:50:13

Commit 3251885285e1 ("ARM: OMAP4+: Reset CPU1 properly for kexec") started
resetting cpu1 because of a kexec boot issue I was seeing earlier in 2016
on omap4 when doing kexec boot between two different kernel versions. The
booted kernel ended up trying to use the old kernel start-up address unless
cpu1 was reset before configuring the cpu1 start-up address.

It seems the reset part was not correct but probably working around some
other issue. I have not been able to reproduce this issue any longer despite
testing with backported patches back to v4.6 kernel. So it is possible this
issue was caused by other work in progress kexec patches I had applied. Or
it is possible some other fixes have made the issue go way.

The unconditional reset of cpu1 can cause issues booting some devices. For
example, bootloader configured secure OS running on cpu1 will fail as the
configuration is not preserved as reported by Andrew F. Davis [off-list ref].

Let's fix the issue by reverting the cpu1 reset parts. If it turns out we
still need to reset cpu1 in some cases, we can add it back and do it
conditionally.

Fixes: 3251885285e1 ("ARM: OMAP4+: Reset CPU1 properly for kexec")
Cc: Keerthy <j-keerthy@ti.com>
Cc: Tero Kristo <redacted>
Reported-by: Andrew F. Davis <redacted>
Signed-off-by: Tony Lindgren <tony@atomide.com>
---
 arch/arm/mach-omap2/omap-smp.c | 10 ----------
 1 file changed, 10 deletions(-)
diff --git a/arch/arm/mach-omap2/omap-smp.c b/arch/arm/mach-omap2/omap-smp.c
--- a/arch/arm/mach-omap2/omap-smp.c
+++ b/arch/arm/mach-omap2/omap-smp.c
@@ -300,16 +300,6 @@ static void __init omap4_smp_prepare_cpus(unsigned int max_cpus)
 		scu_enable(cfg.scu_base);
 
 	/*
-	 * Reset CPU1 before configuring, otherwise kexec will
-	 * end up trying to use old kernel startup address.
-	 */
-	if (cfg.cpu1_rstctrl_va) {
-		writel_relaxed(1, cfg.cpu1_rstctrl_va);
-		readl_relaxed(cfg.cpu1_rstctrl_va);
-		writel_relaxed(0, cfg.cpu1_rstctrl_va);
-	}
-
-	/*
 	 * Write the address of secondary startup routine into the
 	 * AuxCoreBoot1 where ROM code will jump and start executing
 	 * on secondary core once out of WFE
-- 
2.11.1

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: tony@atomide.com (Tony Lindgren)
Date: 2017-02-14 19:36:45

* Tony Lindgren [off-list ref] [170213 13:51]:
Commit 3251885285e1 ("ARM: OMAP4+: Reset CPU1 properly for kexec") started
resetting cpu1 because of a kexec boot issue I was seeing earlier in 2016
on omap4 when doing kexec boot between two different kernel versions. The
booted kernel ended up trying to use the old kernel start-up address unless
cpu1 was reset before configuring the cpu1 start-up address.

It seems the reset part was not correct but probably working around some
other issue. I have not been able to reproduce this issue any longer despite
testing with backported patches back to v4.6 kernel. So it is possible this
issue was caused by other work in progress kexec patches I had applied. Or
it is possible some other fixes have made the issue go way.

The unconditional reset of cpu1 can cause issues booting some devices. For
example, bootloader configured secure OS running on cpu1 will fail as the
configuration is not preserved as reported by Andrew F. Davis [off-list ref].

Let's fix the issue by reverting the cpu1 reset parts. If it turns out we
still need to reset cpu1 in some cases, we can add it back and do it
conditionally.
Actually with this I'm now seeing cpu1 not come up after a suspend/resume
cycle on duovero:

[  118.257415] CPU1: shutdown
[  118.294616] Error taking CPU1 up: -2
[  118.299072] PM: noirq resume of devices complete after 3.723 msecs
[  118.303802] PM: early resume of devices complete after 3.723 msecs

So this issue needs to be investigated more.

Regards,

Tony
quoted hunk
Fixes: 3251885285e1 ("ARM: OMAP4+: Reset CPU1 properly for kexec")
Cc: Keerthy <j-keerthy@ti.com>
Cc: Tero Kristo <redacted>
Reported-by: Andrew F. Davis <redacted>
Signed-off-by: Tony Lindgren <tony@atomide.com>
---
 arch/arm/mach-omap2/omap-smp.c | 10 ----------
 1 file changed, 10 deletions(-)
diff --git a/arch/arm/mach-omap2/omap-smp.c b/arch/arm/mach-omap2/omap-smp.c
--- a/arch/arm/mach-omap2/omap-smp.c
+++ b/arch/arm/mach-omap2/omap-smp.c
@@ -300,16 +300,6 @@ static void __init omap4_smp_prepare_cpus(unsigned int max_cpus)
 		scu_enable(cfg.scu_base);
 
 	/*
-	 * Reset CPU1 before configuring, otherwise kexec will
-	 * end up trying to use old kernel startup address.
-	 */
-	if (cfg.cpu1_rstctrl_va) {
-		writel_relaxed(1, cfg.cpu1_rstctrl_va);
-		readl_relaxed(cfg.cpu1_rstctrl_va);
-		writel_relaxed(0, cfg.cpu1_rstctrl_va);
-	}
-
-	/*
 	 * Write the address of secondary startup routine into the
 	 * AuxCoreBoot1 where ROM code will jump and start executing
 	 * on secondary core once out of WFE
-- 
2.11.1
--
To unsubscribe from this list: send the line "unsubscribe linux-omap" in
the body of a message to majordomo at vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: tony@atomide.com (Tony Lindgren)
Date: 2017-02-15 18:39:16

* Tony Lindgren [off-list ref] [170214 11:39]:
* Tony Lindgren [off-list ref] [170213 13:51]:
quoted
Commit 3251885285e1 ("ARM: OMAP4+: Reset CPU1 properly for kexec") started
resetting cpu1 because of a kexec boot issue I was seeing earlier in 2016
on omap4 when doing kexec boot between two different kernel versions. The
booted kernel ended up trying to use the old kernel start-up address unless
cpu1 was reset before configuring the cpu1 start-up address.

It seems the reset part was not correct but probably working around some
other issue. I have not been able to reproduce this issue any longer despite
testing with backported patches back to v4.6 kernel. So it is possible this
issue was caused by other work in progress kexec patches I had applied. Or
it is possible some other fixes have made the issue go way.

The unconditional reset of cpu1 can cause issues booting some devices. For
example, bootloader configured secure OS running on cpu1 will fail as the
configuration is not preserved as reported by Andrew F. Davis [off-list ref].

Let's fix the issue by reverting the cpu1 reset parts. If it turns out we
still need to reset cpu1 in some cases, we can add it back and do it
conditionally.
Actually with this I'm now seeing cpu1 not come up after a suspend/resume
cycle on duovero:

[  118.257415] CPU1: shutdown
[  118.294616] Error taking CPU1 up: -2
[  118.299072] PM: noirq resume of devices complete after 3.723 msecs
[  118.303802] PM: early resume of devices complete after 3.723 msecs

So this issue needs to be investigated more.
And then today the omap4 suspend/resume issue is no longer reproducable..
Go figure.

But then doing more testing I noticed that also omap5 needs the reset.
Without it we get the following on omap5-uevm doing a kexec boot. So clearly
the reset cannot be just removed at least for omap4 and omap5.

Regards,

Tony

8< ---------------------
[    0.156796] CPU0: thread -1, cpu 0, socket 0, mpidr 80000000
[    0.163396] Setting up static identity map for 0x80100000 - 0x80100070
[    0.172246] smp: Bringing up secondary CPUs ...
[    0.178970] Unable to handle kernel NULL pointer dereference at virtual address 00000000
[    0.178974] pgd = c0004000
[    0.178977] [00000000] *pgd=00000000
[    0.178990] Internal error: Oops: 80000005 [#1] SMP ARM
[    0.178995] Modules linked in:
[    0.179005] CPU: 1 PID: 0 Comm: swapper/1 Not tainted 4.10.0-rc8-next-20170215+ #120
[    0.179008] Hardware name: Generic OMAP5 (Flattened Device Tree)
[    0.179013] task: ee0c8ec0 task.stack: ee0ca000
[    0.179018] PC is at 0x0
[    0.179029] LR is@omap4_cpu_die+0x58/0x98
[    0.179034] pc : [<00000000>]    lr : [<c01243dc>]    psr: 60000093
[    0.179034] sp : ee0cbfb8  ip : 00000000  fp : 00000000
[    0.179038] r10: 00000000  r9 : c0d50569  r8 : 00000000
[    0.179042] r7 : c0c76448  r6 : c0d0792c  r5 : 00000001  r4 : c0b08054
[    0.179046] r3 : 00000001  r2 : f0880000  r1 : 00000003  r0 : 00000001
[    0.179051] Flags: nZCv  IRQs off  FIQs on  Mode SVC_32  ISA ARM  Segment none
[    0.179055] Control: 10c5387d  Table: 8000406a  DAC: 00000051
[    0.179059] Process swapper/1 (pid: 0, stack limit = 0xee0ca218)
[    0.179063] Stack: (0xee0cbfb8 to 0xee0cc000)
[    0.179068] bfa0:                                                       00000000 00000000
[    0.179075] bfc0: 00000000 00000000 00000000 00000000 00000000 00000000 00000000 00000000
[    0.179082] bfe0: 00000000 00000000 00000000 00000000 00000013 00000000 681b0041 cf3e4021
[    0.179092] [<c01243dc>] (omap4_cpu_die) from [<00000000>] (  (null))
[    0.179098] Code: bad PC value
[    0.179115] ---[ end trace e14406c260ce69db ]---
[    0.179121] Kernel panic - not syncing: Attempted to kill the idle task!
[    0.179135] CPU0: stopping
[    0.179141] ---[ end Kernel panic - not syncing: Attempted to kill the idle task!
[    0.339715] CPU: 0 PID: 0 Comm: swapper/0 Tainted: G      D         4.10.0-rc8-next-20170215+ #120
[    0.348927] Hardware name: Generic OMAP5 (Flattened Device Tree)
[    0.355112] [<c0110228>] (unwind_backtrace) from [<c010c224>] (show_stack+0x10/0x14)
[    0.363083] [<c010c224>] (show_stack) from [<c04ca860>] (dump_stack+0xac/0xe0)
[    0.370513] [<c04ca860>] (dump_stack) from [<c010e72c>] (handle_IPI+0x358/0x3f8)
[    0.378120] [<c010e72c>] (handle_IPI) from [<c01015a4>] (gic_handle_irq+0x9c/0xb8)
[    0.385909] [<c01015a4>] (gic_handle_irq) from [<c083b270>] (__irq_svc+0x70/0x98)
[    0.393602] Exception stack(0xc0d01f38 to 0xc0d01f80)
[    0.398794] 1f20:                                                       c0108284 00000000
[    0.407205] 1f40: 00000000 00000000 c0d00000 c0d07994 c0d0792c c0c76448 c0d08560 c0d50569
[    0.415616] 1f60: 00000000 00000000 00000000 c0d01f88 c0108284 c0108288 60000013 ffffffff
[    0.424032] [<c083b270>] (__irq_svc) from [<c0108288>] (arch_cpu_idle+0x20/0x3c)
[    0.431643] [<c0108288>] (arch_cpu_idle) from [<c0190bc4>] (do_idle+0x164/0x218)
[    0.439251] [<c0190bc4>] (do_idle) from [<c0190ffc>] (cpu_startup_entry+0x18/0x1c)
[    0.447040] [<c0190ffc>] (cpu_startup_entry) from [<c0c00c40>] (start_kernel+0x35c/0x3d4)
[    0.455451] [<c0c00c40>] (start_kernel) from [<8000807c>] (0x8000807c)

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: tony@atomide.com (Tony Lindgren)
Date: 2017-02-15 19:12:42

* Tony Lindgren [off-list ref] [170215 10:40]:
* Tony Lindgren [off-list ref] [170214 11:39]:
quoted
* Tony Lindgren [off-list ref] [170213 13:51]:
quoted
Commit 3251885285e1 ("ARM: OMAP4+: Reset CPU1 properly for kexec") started
resetting cpu1 because of a kexec boot issue I was seeing earlier in 2016
on omap4 when doing kexec boot between two different kernel versions. The
booted kernel ended up trying to use the old kernel start-up address unless
cpu1 was reset before configuring the cpu1 start-up address.

It seems the reset part was not correct but probably working around some
other issue. I have not been able to reproduce this issue any longer despite
testing with backported patches back to v4.6 kernel. So it is possible this
issue was caused by other work in progress kexec patches I had applied. Or
it is possible some other fixes have made the issue go way.

The unconditional reset of cpu1 can cause issues booting some devices. For
example, bootloader configured secure OS running on cpu1 will fail as the
configuration is not preserved as reported by Andrew F. Davis [off-list ref].

Let's fix the issue by reverting the cpu1 reset parts. If it turns out we
still need to reset cpu1 in some cases, we can add it back and do it
conditionally.
Actually with this I'm now seeing cpu1 not come up after a suspend/resume
cycle on duovero:

[  118.257415] CPU1: shutdown
[  118.294616] Error taking CPU1 up: -2
[  118.299072] PM: noirq resume of devices complete after 3.723 msecs
[  118.303802] PM: early resume of devices complete after 3.723 msecs

So this issue needs to be investigated more.
And then today the omap4 suspend/resume issue is no longer reproducable..
Go figure.

But then doing more testing I noticed that also omap5 needs the reset.
Without it we get the following on omap5-uevm doing a kexec boot. So clearly
the reset cannot be just removed at least for omap4 and omap5.
And also the same issue happens doing kexec on beagle-x15 naturally if
the cpu1 reset is removed.

Regards,

Tony
8< ---------------------
[    0.156796] CPU0: thread -1, cpu 0, socket 0, mpidr 80000000
[    0.163396] Setting up static identity map for 0x80100000 - 0x80100070
[    0.172246] smp: Bringing up secondary CPUs ...
[    0.178970] Unable to handle kernel NULL pointer dereference at virtual address 00000000
[    0.178974] pgd = c0004000
[    0.178977] [00000000] *pgd=00000000
[    0.178990] Internal error: Oops: 80000005 [#1] SMP ARM
[    0.178995] Modules linked in:
[    0.179005] CPU: 1 PID: 0 Comm: swapper/1 Not tainted 4.10.0-rc8-next-20170215+ #120
[    0.179008] Hardware name: Generic OMAP5 (Flattened Device Tree)
[    0.179013] task: ee0c8ec0 task.stack: ee0ca000
[    0.179018] PC is at 0x0
[    0.179029] LR is at omap4_cpu_die+0x58/0x98
[    0.179034] pc : [<00000000>]    lr : [<c01243dc>]    psr: 60000093
[    0.179034] sp : ee0cbfb8  ip : 00000000  fp : 00000000
[    0.179038] r10: 00000000  r9 : c0d50569  r8 : 00000000
[    0.179042] r7 : c0c76448  r6 : c0d0792c  r5 : 00000001  r4 : c0b08054
[    0.179046] r3 : 00000001  r2 : f0880000  r1 : 00000003  r0 : 00000001
[    0.179051] Flags: nZCv  IRQs off  FIQs on  Mode SVC_32  ISA ARM  Segment none
[    0.179055] Control: 10c5387d  Table: 8000406a  DAC: 00000051
[    0.179059] Process swapper/1 (pid: 0, stack limit = 0xee0ca218)
[    0.179063] Stack: (0xee0cbfb8 to 0xee0cc000)
[    0.179068] bfa0:                                                       00000000 00000000
[    0.179075] bfc0: 00000000 00000000 00000000 00000000 00000000 00000000 00000000 00000000
[    0.179082] bfe0: 00000000 00000000 00000000 00000000 00000013 00000000 681b0041 cf3e4021
[    0.179092] [<c01243dc>] (omap4_cpu_die) from [<00000000>] (  (null))
[    0.179098] Code: bad PC value
[    0.179115] ---[ end trace e14406c260ce69db ]---
[    0.179121] Kernel panic - not syncing: Attempted to kill the idle task!
[    0.179135] CPU0: stopping
[    0.179141] ---[ end Kernel panic - not syncing: Attempted to kill the idle task!
[    0.339715] CPU: 0 PID: 0 Comm: swapper/0 Tainted: G      D         4.10.0-rc8-next-20170215+ #120
[    0.348927] Hardware name: Generic OMAP5 (Flattened Device Tree)
[    0.355112] [<c0110228>] (unwind_backtrace) from [<c010c224>] (show_stack+0x10/0x14)
[    0.363083] [<c010c224>] (show_stack) from [<c04ca860>] (dump_stack+0xac/0xe0)
[    0.370513] [<c04ca860>] (dump_stack) from [<c010e72c>] (handle_IPI+0x358/0x3f8)
[    0.378120] [<c010e72c>] (handle_IPI) from [<c01015a4>] (gic_handle_irq+0x9c/0xb8)
[    0.385909] [<c01015a4>] (gic_handle_irq) from [<c083b270>] (__irq_svc+0x70/0x98)
[    0.393602] Exception stack(0xc0d01f38 to 0xc0d01f80)
[    0.398794] 1f20:                                                       c0108284 00000000
[    0.407205] 1f40: 00000000 00000000 c0d00000 c0d07994 c0d0792c c0c76448 c0d08560 c0d50569
[    0.415616] 1f60: 00000000 00000000 00000000 c0d01f88 c0108284 c0108288 60000013 ffffffff
[    0.424032] [<c083b270>] (__irq_svc) from [<c0108288>] (arch_cpu_idle+0x20/0x3c)
[    0.431643] [<c0108288>] (arch_cpu_idle) from [<c0190bc4>] (do_idle+0x164/0x218)
[    0.439251] [<c0190bc4>] (do_idle) from [<c0190ffc>] (cpu_startup_entry+0x18/0x1c)
[    0.447040] [<c0190ffc>] (cpu_startup_entry) from [<c0c00c40>] (start_kernel+0x35c/0x3d4)
[    0.455451] [<c0c00c40>] (start_kernel) from [<8000807c>] (0x8000807c)
--
To unsubscribe from this list: send the line "unsubscribe linux-omap" in
the body of a message to majordomo at vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: Andrew F. Davis <hidden>
Date: 2017-02-15 22:13:19

On 02/15/2017 01:12 PM, Tony Lindgren wrote:
* Tony Lindgren [off-list ref] [170215 10:40]:
quoted
* Tony Lindgren [off-list ref] [170214 11:39]:
quoted
* Tony Lindgren [off-list ref] [170213 13:51]:
quoted
Commit 3251885285e1 ("ARM: OMAP4+: Reset CPU1 properly for kexec") started
resetting cpu1 because of a kexec boot issue I was seeing earlier in 2016
on omap4 when doing kexec boot between two different kernel versions. The
booted kernel ended up trying to use the old kernel start-up address unless
cpu1 was reset before configuring the cpu1 start-up address.

It seems the reset part was not correct but probably working around some
other issue. I have not been able to reproduce this issue any longer despite
testing with backported patches back to v4.6 kernel. So it is possible this
issue was caused by other work in progress kexec patches I had applied. Or
it is possible some other fixes have made the issue go way.

The unconditional reset of cpu1 can cause issues booting some devices. For
example, bootloader configured secure OS running on cpu1 will fail as the
configuration is not preserved as reported by Andrew F. Davis [off-list ref].

Let's fix the issue by reverting the cpu1 reset parts. If it turns out we
still need to reset cpu1 in some cases, we can add it back and do it
conditionally.
Actually with this I'm now seeing cpu1 not come up after a suspend/resume
cycle on duovero:

[  118.257415] CPU1: shutdown
[  118.294616] Error taking CPU1 up: -2
[  118.299072] PM: noirq resume of devices complete after 3.723 msecs
[  118.303802] PM: early resume of devices complete after 3.723 msecs

So this issue needs to be investigated more.
And then today the omap4 suspend/resume issue is no longer reproducable..
Go figure.

But then doing more testing I noticed that also omap5 needs the reset.
Without it we get the following on omap5-uevm doing a kexec boot. So clearly
the reset cannot be just removed at least for omap4 and omap5.
And also the same issue happens doing kexec on beagle-x15 naturally if
the cpu1 reset is removed.
When a core actually powers up it idles in ROM code waiting for
OMAP_AUX_CORE_BOOT_0 to be set. When we shutdown a core it is not really
powered off, we just let it spin in omap4_cpu_die() or
omap4_secondary_startup() waiting on OMAP_AUX_CORE_BOOT_0, just like if
it were still trapped in ROM after a reset.

The issue with this fake startup idle loop is that, unlike the ROM based
startup idle loop, these do *not* jump to the address we stored in
OMAP_AUX_CORE_BOOT_1, they just make the assumption that they can safely
jump to the kernel startup function.

So when we tell this core to boot, and it is not in the real ROM startup
loop, it breaks stuff as it jumps to the old kernel's
secondary_startup() even though we gave it the correct address in
OMAP_AUX_CORE_BOOT_1.

Reseting the core to put it back in the real ROM idle loop is wrong, the
two idle loop functions above should be fixed to respect the address in
OMAP_AUX_CORE_BOOT_1 and not to make assumptions, this should take care
of the kexec failure in a sane way.

Andrew
Regards,

Tony
quoted
8< ---------------------
[    0.156796] CPU0: thread -1, cpu 0, socket 0, mpidr 80000000
[    0.163396] Setting up static identity map for 0x80100000 - 0x80100070
[    0.172246] smp: Bringing up secondary CPUs ...
[    0.178970] Unable to handle kernel NULL pointer dereference at virtual address 00000000
[    0.178974] pgd = c0004000
[    0.178977] [00000000] *pgd=00000000
[    0.178990] Internal error: Oops: 80000005 [#1] SMP ARM
[    0.178995] Modules linked in:
[    0.179005] CPU: 1 PID: 0 Comm: swapper/1 Not tainted 4.10.0-rc8-next-20170215+ #120
[    0.179008] Hardware name: Generic OMAP5 (Flattened Device Tree)
[    0.179013] task: ee0c8ec0 task.stack: ee0ca000
[    0.179018] PC is at 0x0
[    0.179029] LR is at omap4_cpu_die+0x58/0x98
[    0.179034] pc : [<00000000>]    lr : [<c01243dc>]    psr: 60000093
[    0.179034] sp : ee0cbfb8  ip : 00000000  fp : 00000000
[    0.179038] r10: 00000000  r9 : c0d50569  r8 : 00000000
[    0.179042] r7 : c0c76448  r6 : c0d0792c  r5 : 00000001  r4 : c0b08054
[    0.179046] r3 : 00000001  r2 : f0880000  r1 : 00000003  r0 : 00000001
[    0.179051] Flags: nZCv  IRQs off  FIQs on  Mode SVC_32  ISA ARM  Segment none
[    0.179055] Control: 10c5387d  Table: 8000406a  DAC: 00000051
[    0.179059] Process swapper/1 (pid: 0, stack limit = 0xee0ca218)
[    0.179063] Stack: (0xee0cbfb8 to 0xee0cc000)
[    0.179068] bfa0:                                                       00000000 00000000
[    0.179075] bfc0: 00000000 00000000 00000000 00000000 00000000 00000000 00000000 00000000
[    0.179082] bfe0: 00000000 00000000 00000000 00000000 00000013 00000000 681b0041 cf3e4021
[    0.179092] [<c01243dc>] (omap4_cpu_die) from [<00000000>] (  (null))
[    0.179098] Code: bad PC value
[    0.179115] ---[ end trace e14406c260ce69db ]---
[    0.179121] Kernel panic - not syncing: Attempted to kill the idle task!
[    0.179135] CPU0: stopping
[    0.179141] ---[ end Kernel panic - not syncing: Attempted to kill the idle task!
[    0.339715] CPU: 0 PID: 0 Comm: swapper/0 Tainted: G      D         4.10.0-rc8-next-20170215+ #120
[    0.348927] Hardware name: Generic OMAP5 (Flattened Device Tree)
[    0.355112] [<c0110228>] (unwind_backtrace) from [<c010c224>] (show_stack+0x10/0x14)
[    0.363083] [<c010c224>] (show_stack) from [<c04ca860>] (dump_stack+0xac/0xe0)
[    0.370513] [<c04ca860>] (dump_stack) from [<c010e72c>] (handle_IPI+0x358/0x3f8)
[    0.378120] [<c010e72c>] (handle_IPI) from [<c01015a4>] (gic_handle_irq+0x9c/0xb8)
[    0.385909] [<c01015a4>] (gic_handle_irq) from [<c083b270>] (__irq_svc+0x70/0x98)
[    0.393602] Exception stack(0xc0d01f38 to 0xc0d01f80)
[    0.398794] 1f20:                                                       c0108284 00000000
[    0.407205] 1f40: 00000000 00000000 c0d00000 c0d07994 c0d0792c c0c76448 c0d08560 c0d50569
[    0.415616] 1f60: 00000000 00000000 00000000 c0d01f88 c0108284 c0108288 60000013 ffffffff
[    0.424032] [<c083b270>] (__irq_svc) from [<c0108288>] (arch_cpu_idle+0x20/0x3c)
[    0.431643] [<c0108288>] (arch_cpu_idle) from [<c0190bc4>] (do_idle+0x164/0x218)
[    0.439251] [<c0190bc4>] (do_idle) from [<c0190ffc>] (cpu_startup_entry+0x18/0x1c)
[    0.447040] [<c0190ffc>] (cpu_startup_entry) from [<c0c00c40>] (start_kernel+0x35c/0x3d4)
[    0.455451] [<c0c00c40>] (start_kernel) from [<8000807c>] (0x8000807c)
--
To unsubscribe from this list: send the line "unsubscribe linux-omap" in
the body of a message to majordomo at vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: tony@atomide.com (Tony Lindgren)
Date: 2017-02-15 22:27:11

* Andrew F. Davis [off-list ref] [170215 14:14]:
On 02/15/2017 01:12 PM, Tony Lindgren wrote:
quoted
* Tony Lindgren [off-list ref] [170215 10:40]:
quoted
* Tony Lindgren [off-list ref] [170214 11:39]:
quoted
* Tony Lindgren [off-list ref] [170213 13:51]:
quoted
Commit 3251885285e1 ("ARM: OMAP4+: Reset CPU1 properly for kexec") started
resetting cpu1 because of a kexec boot issue I was seeing earlier in 2016
on omap4 when doing kexec boot between two different kernel versions. The
booted kernel ended up trying to use the old kernel start-up address unless
cpu1 was reset before configuring the cpu1 start-up address.

It seems the reset part was not correct but probably working around some
other issue. I have not been able to reproduce this issue any longer despite
testing with backported patches back to v4.6 kernel. So it is possible this
issue was caused by other work in progress kexec patches I had applied. Or
it is possible some other fixes have made the issue go way.

The unconditional reset of cpu1 can cause issues booting some devices. For
example, bootloader configured secure OS running on cpu1 will fail as the
configuration is not preserved as reported by Andrew F. Davis [off-list ref].

Let's fix the issue by reverting the cpu1 reset parts. If it turns out we
still need to reset cpu1 in some cases, we can add it back and do it
conditionally.
Actually with this I'm now seeing cpu1 not come up after a suspend/resume
cycle on duovero:

[  118.257415] CPU1: shutdown
[  118.294616] Error taking CPU1 up: -2
[  118.299072] PM: noirq resume of devices complete after 3.723 msecs
[  118.303802] PM: early resume of devices complete after 3.723 msecs

So this issue needs to be investigated more.
And then today the omap4 suspend/resume issue is no longer reproducable..
Go figure.

But then doing more testing I noticed that also omap5 needs the reset.
Without it we get the following on omap5-uevm doing a kexec boot. So clearly
the reset cannot be just removed at least for omap4 and omap5.
And also the same issue happens doing kexec on beagle-x15 naturally if
the cpu1 reset is removed.
When a core actually powers up it idles in ROM code waiting for
OMAP_AUX_CORE_BOOT_0 to be set. When we shutdown a core it is not really
powered off, we just let it spin in omap4_cpu_die() or
omap4_secondary_startup() waiting on OMAP_AUX_CORE_BOOT_0, just like if
it were still trapped in ROM after a reset.

The issue with this fake startup idle loop is that, unlike the ROM based
startup idle loop, these do *not* jump to the address we stored in
OMAP_AUX_CORE_BOOT_1, they just make the assumption that they can safely
jump to the kernel startup function.

So when we tell this core to boot, and it is not in the real ROM startup
loop, it breaks stuff as it jumps to the old kernel's
secondary_startup() even though we gave it the correct address in
OMAP_AUX_CORE_BOOT_1.
Yes this is probably what's going on here. Note that the error I pasted
was booting the same kernel where that address should be correct though.
So there might be something else to it also.
Reseting the core to put it back in the real ROM idle loop is wrong, the
two idle loop functions above should be fixed to respect the address in
OMAP_AUX_CORE_BOOT_1 and not to make assumptions, this should take care
of the kexec failure in a sane way.
OK care to try to patch it as now you also have a reproducable test
case for kexec too?

Regards,

Tony

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: tony@atomide.com (Tony Lindgren)
Date: 2017-02-16 16:10:10

* Tony Lindgren [off-list ref] [170215 14:28]:
* Andrew F. Davis [off-list ref] [170215 14:14]:
quoted
On 02/15/2017 01:12 PM, Tony Lindgren wrote:
quoted
And also the same issue happens doing kexec on beagle-x15 naturally if
the cpu1 reset is removed.
When a core actually powers up it idles in ROM code waiting for
OMAP_AUX_CORE_BOOT_0 to be set. When we shutdown a core it is not really
powered off, we just let it spin in omap4_cpu_die() or
omap4_secondary_startup() waiting on OMAP_AUX_CORE_BOOT_0, just like if
it were still trapped in ROM after a reset.
OK so I debugged this a bit more. We have CPU1 in omap_do_wfi()
as we don't currently have omap5_secondary_startup() or any deeper
idle mode support beyond retention for omap5 or dra7 in the mainline
kernel.
quoted
The issue with this fake startup idle loop is that, unlike the ROM based
startup idle loop, these do *not* jump to the address we stored in
OMAP_AUX_CORE_BOOT_1, they just make the assumption that they can safely
jump to the kernel startup function.
This does not seem to be the case here.
quoted
So when we tell this core to boot, and it is not in the real ROM startup
loop, it breaks stuff as it jumps to the old kernel's
secondary_startup() even though we gave it the correct address in
OMAP_AUX_CORE_BOOT_1.
And this is not happening. I think this is what I was seeing earlier,
but it's not the omap5/dra7 issue.

What we have is cpu1 returning from previous kernel's omap_do_wfi()
in the kexec booted kernel's code and that's when things go wrong.

So if cpu1 was configured for idle for any reason, it will never gets
to omap5_secondary_startup without the reset currently.

The reason kexec and suspend/resume mostly works for omap4 without
cpu1 reset is that we usually enter off mode for cpu1 and the context
is lost and then we properly go through omap4_secondary_startup. Or
that's my theory so far for the occasional flakeyness I've been seeing :)

Any ideas what we should try to check to see if cpu1 is in idle
mode so we can do the reset if needed?

Regards,

Tony

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: tony@atomide.com (Tony Lindgren)
Date: 2017-02-16 16:21:42

* Tony Lindgren [off-list ref] [170216 08:11]:
* Tony Lindgren [off-list ref] [170215 14:28]:
quoted
* Andrew F. Davis [off-list ref] [170215 14:14]:
quoted
On 02/15/2017 01:12 PM, Tony Lindgren wrote:
quoted
And also the same issue happens doing kexec on beagle-x15 naturally if
the cpu1 reset is removed.
When a core actually powers up it idles in ROM code waiting for
OMAP_AUX_CORE_BOOT_0 to be set. When we shutdown a core it is not really
powered off, we just let it spin in omap4_cpu_die() or
omap4_secondary_startup() waiting on OMAP_AUX_CORE_BOOT_0, just like if
it were still trapped in ROM after a reset.
OK so I debugged this a bit more. We have CPU1 in omap_do_wfi()
as we don't currently have omap5_secondary_startup() or any deeper
idle mode support beyond retention for omap5 or dra7 in the mainline
kernel.
quoted
quoted
The issue with this fake startup idle loop is that, unlike the ROM based
startup idle loop, these do *not* jump to the address we stored in
OMAP_AUX_CORE_BOOT_1, they just make the assumption that they can safely
jump to the kernel startup function.
This does not seem to be the case here.
quoted
quoted
So when we tell this core to boot, and it is not in the real ROM startup
loop, it breaks stuff as it jumps to the old kernel's
secondary_startup() even though we gave it the correct address in
OMAP_AUX_CORE_BOOT_1.
And this is not happening. I think this is what I was seeing earlier,
but it's not the omap5/dra7 issue.

What we have is cpu1 returning from previous kernel's omap_do_wfi()
in the kexec booted kernel's code and that's when things go wrong.

So if cpu1 was configured for idle for any reason, it will never gets
to omap5_secondary_startup without the reset currently.

The reason kexec and suspend/resume mostly works for omap4 without
cpu1 reset is that we usually enter off mode for cpu1 and the context
is lost and then we properly go through omap4_secondary_startup. Or
that's my theory so far for the occasional flakeyness I've been seeing :)

Any ideas what we should try to check to see if cpu1 is in idle
mode so we can do the reset if needed?
Maybe we should do the reset if OMAP5_CPU1_WAKEUP_NS_PA_ADDR_OFFSET
is not 0?

Regards,

Tony

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: Andrew F. Davis <hidden>
Date: 2017-02-16 16:29:16

On 02/16/2017 10:10 AM, Tony Lindgren wrote:
* Tony Lindgren [off-list ref] [170215 14:28]:
quoted
* Andrew F. Davis [off-list ref] [170215 14:14]:
quoted
On 02/15/2017 01:12 PM, Tony Lindgren wrote:
quoted
And also the same issue happens doing kexec on beagle-x15 naturally if
the cpu1 reset is removed.
When a core actually powers up it idles in ROM code waiting for
OMAP_AUX_CORE_BOOT_0 to be set. When we shutdown a core it is not really
powered off, we just let it spin in omap4_cpu_die() or
omap4_secondary_startup() waiting on OMAP_AUX_CORE_BOOT_0, just like if
it were still trapped in ROM after a reset.
OK so I debugged this a bit more. We have CPU1 in omap_do_wfi()
as we don't currently have omap5_secondary_startup() or any deeper
idle mode support beyond retention for omap5 or dra7 in the mainline
kernel.
quoted
quoted
The issue with this fake startup idle loop is that, unlike the ROM based
startup idle loop, these do *not* jump to the address we stored in
OMAP_AUX_CORE_BOOT_1, they just make the assumption that they can safely
jump to the kernel startup function.
This does not seem to be the case here.
Well this is what I am seeing every time, this code only works when it
is the same kernel we kexec, any changed addresses here will not work.
quoted
quoted
So when we tell this core to boot, and it is not in the real ROM startup
loop, it breaks stuff as it jumps to the old kernel's
secondary_startup() even though we gave it the correct address in
OMAP_AUX_CORE_BOOT_1.
And this is not happening. I think this is what I was seeing earlier,
but it's not the omap5/dra7 issue.

What we have is cpu1 returning from previous kernel's omap_do_wfi()
in the kexec booted kernel's code and that's when things go wrong.
We are the ones sending it to omap_do_wfi(), in omap4_cpu_die() it gets
idled in a loop, it shouldn't be idled after it is shut off, it should
get parked, we should do this like we do in omap5_secondary_startup().
So if cpu1 was configured for idle for any reason, it will never gets
to omap5_secondary_startup without the reset currently.

The reason kexec and suspend/resume mostly works for omap4 without
cpu1 reset is that we usually enter off mode for cpu1 and the context
is lost and then we properly go through omap4_secondary_startup. Or
that's my theory so far for the occasional flakeyness I've been seeing :)

Any ideas what we should try to check to see if cpu1 is in idle
mode so we can do the reset if needed?
You can never reset the core, resetting the core is not allowed on HS
devices and so it really doesn't matter what the core is doing. In no
case is reseting the core a valid work-around for not correctly parking
it. We need to fix the omap4_cpu_die() to not let the core go idle if
the return from idle path is the problem.

Andrew
Regards,

Tony

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: tony@atomide.com (Tony Lindgren)
Date: 2017-02-16 16:54:09

* Andrew F. Davis [off-list ref] [170216 08:30]:
On 02/16/2017 10:10 AM, Tony Lindgren wrote:
quoted
* Tony Lindgren [off-list ref] [170215 14:28]:
quoted
* Andrew F. Davis [off-list ref] [170215 14:14]:
quoted
On 02/15/2017 01:12 PM, Tony Lindgren wrote:
quoted
And also the same issue happens doing kexec on beagle-x15 naturally if
the cpu1 reset is removed.
When a core actually powers up it idles in ROM code waiting for
OMAP_AUX_CORE_BOOT_0 to be set. When we shutdown a core it is not really
powered off, we just let it spin in omap4_cpu_die() or
omap4_secondary_startup() waiting on OMAP_AUX_CORE_BOOT_0, just like if
it were still trapped in ROM after a reset.
OK so I debugged this a bit more. We have CPU1 in omap_do_wfi()
as we don't currently have omap5_secondary_startup() or any deeper
idle mode support beyond retention for omap5 or dra7 in the mainline
kernel.
quoted
quoted
The issue with this fake startup idle loop is that, unlike the ROM based
startup idle loop, these do *not* jump to the address we stored in
OMAP_AUX_CORE_BOOT_1, they just make the assumption that they can safely
jump to the kernel startup function.
This does not seem to be the case here.
Well this is what I am seeing every time, this code only works when it
is the same kernel we kexec, any changed addresses here will not work.
Hmm let's talk the mainline kernel here. Currently things do work in
the mainline kernel because of the cpu1 reset. And without cpu1 reset
things will currently go wrong in the mainline kernel both for kexec
and suspend/resume.
quoted
quoted
quoted
So when we tell this core to boot, and it is not in the real ROM startup
loop, it breaks stuff as it jumps to the old kernel's
secondary_startup() even though we gave it the correct address in
OMAP_AUX_CORE_BOOT_1.
And this is not happening. I think this is what I was seeing earlier,
but it's not the omap5/dra7 issue.

What we have is cpu1 returning from previous kernel's omap_do_wfi()
in the kexec booted kernel's code and that's when things go wrong.
We are the ones sending it to omap_do_wfi(), in omap4_cpu_die() it gets
idled in a loop, it shouldn't be idled after it is shut off, it should
get parked, we should do this like we do in omap5_secondary_startup().
Yup agreed. We need to figure out if it's just normal cpuidle hot-unplug
event vs shut down and park for kexec. Probably cpu_kill() is the place
to park it, need to check.
quoted
So if cpu1 was configured for idle for any reason, it will never gets
to omap5_secondary_startup without the reset currently.

The reason kexec and suspend/resume mostly works for omap4 without
cpu1 reset is that we usually enter off mode for cpu1 and the context
is lost and then we properly go through omap4_secondary_startup. Or
that's my theory so far for the occasional flakeyness I've been seeing :)

Any ideas what we should try to check to see if cpu1 is in idle
mode so we can do the reset if needed?
You can never reset the core, resetting the core is not allowed on HS
devices and so it really doesn't matter what the core is doing. In no
case is reseting the core a valid work-around for not correctly parking
it. We need to fix the omap4_cpu_die() to not let the core go idle if
the return from idle path is the problem.
Yeah well from Linux point of view, what we're interested in is that
cpu1 comes up reliably in all cases no matter what it takes. I agree
doing a reset on it should be only done if nothing else helps. And I
can see some HS implementations not allowing cpu1 reset. And I can see
some product specific bootloaders idle cpu1 and that's where things
break again.

For your use case, probably all we need is runtime checks for HS in
addition to parking cpu1 for kexec. If that's not enough, then maybe
a device specific DT property for never-reset-no-matter-what.

Regards,

Tony

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: tony@atomide.com (Tony Lindgren)
Date: 2017-02-16 19:07:01

* Tony Lindgren [off-list ref] [170216 08:55]:
* Andrew F. Davis [off-list ref] [170216 08:30]:
quoted
On 02/16/2017 10:10 AM, Tony Lindgren wrote:
quoted
* Tony Lindgren [off-list ref] [170215 14:28]:
quoted
* Andrew F. Davis [off-list ref] [170215 14:14]:
quoted
On 02/15/2017 01:12 PM, Tony Lindgren wrote:
quoted
And also the same issue happens doing kexec on beagle-x15 naturally if
the cpu1 reset is removed.
When a core actually powers up it idles in ROM code waiting for
OMAP_AUX_CORE_BOOT_0 to be set. When we shutdown a core it is not really
powered off, we just let it spin in omap4_cpu_die() or
omap4_secondary_startup() waiting on OMAP_AUX_CORE_BOOT_0, just like if
it were still trapped in ROM after a reset.
OK so I debugged this a bit more. We have CPU1 in omap_do_wfi()
as we don't currently have omap5_secondary_startup() or any deeper
idle mode support beyond retention for omap5 or dra7 in the mainline
kernel.
quoted
quoted
The issue with this fake startup idle loop is that, unlike the ROM based
startup idle loop, these do *not* jump to the address we stored in
OMAP_AUX_CORE_BOOT_1, they just make the assumption that they can safely
jump to the kernel startup function.
This does not seem to be the case here.
Well this is what I am seeing every time, this code only works when it
is the same kernel we kexec, any changed addresses here will not work.
Hmm let's talk the mainline kernel here. Currently things do work in
the mainline kernel because of the cpu1 reset. And without cpu1 reset
things will currently go wrong in the mainline kernel both for kexec
and suspend/resume.
quoted
quoted
quoted
quoted
So when we tell this core to boot, and it is not in the real ROM startup
loop, it breaks stuff as it jumps to the old kernel's
secondary_startup() even though we gave it the correct address in
OMAP_AUX_CORE_BOOT_1.
And this is not happening. I think this is what I was seeing earlier,
but it's not the omap5/dra7 issue.

What we have is cpu1 returning from previous kernel's omap_do_wfi()
in the kexec booted kernel's code and that's when things go wrong.
We are the ones sending it to omap_do_wfi(), in omap4_cpu_die() it gets
idled in a loop, it shouldn't be idled after it is shut off, it should
get parked, we should do this like we do in omap5_secondary_startup().
Yup agreed. We need to figure out if it's just normal cpuidle hot-unplug
event vs shut down and park for kexec. Probably cpu_kill() is the place
to park it, need to check.
quoted
quoted
So if cpu1 was configured for idle for any reason, it will never gets
to omap5_secondary_startup without the reset currently.

The reason kexec and suspend/resume mostly works for omap4 without
cpu1 reset is that we usually enter off mode for cpu1 and the context
is lost and then we properly go through omap4_secondary_startup. Or
that's my theory so far for the occasional flakeyness I've been seeing :)

Any ideas what we should try to check to see if cpu1 is in idle
mode so we can do the reset if needed?
You can never reset the core, resetting the core is not allowed on HS
devices and so it really doesn't matter what the core is doing. In no
case is reseting the core a valid work-around for not correctly parking
it. We need to fix the omap4_cpu_die() to not let the core go idle if
the return from idle path is the problem.
Yeah well from Linux point of view, what we're interested in is that
cpu1 comes up reliably in all cases no matter what it takes. I agree
doing a reset on it should be only done if nothing else helps. And I
can see some HS implementations not allowing cpu1 reset. And I can see
some product specific bootloaders idle cpu1 and that's where things
break again.

For your use case, probably all we need is runtime checks for HS in
addition to parking cpu1 for kexec. If that's not enough, then maybe
a device specific DT property for never-reset-no-matter-what.
Below is a first take on the last resort cpu1 reset done based on
configured CPU1_WAKEUP_NS_PA_ADDR_OFFSET for omap4 and 5. Note that
we can't merge it yet as it will break kexec boot for dra7. To fix that,
we need to first do what you're suggesting and properly park cpu1 for
kexec.

Regards,

Tony

8< -----------------------------
diff --git a/arch/arm/mach-omap2/common.h b/arch/arm/mach-omap2/common.h
--- a/arch/arm/mach-omap2/common.h
+++ b/arch/arm/mach-omap2/common.h
@@ -270,6 +270,7 @@ extern const struct smp_operations omap4_smp_ops;
 extern int omap4_mpuss_init(void);
 extern int omap4_enter_lowpower(unsigned int cpu, unsigned int power_state);
 extern int omap4_hotplug_cpu(unsigned int cpu, unsigned int power_state);
+extern bool omap4_cpu1_may_need_reset(void);
 #else
 static inline int omap4_enter_lowpower(unsigned int cpu,
 					unsigned int power_state)
diff --git a/arch/arm/mach-omap2/omap-mpuss-lowpower.c b/arch/arm/mach-omap2/omap-mpuss-lowpower.c
--- a/arch/arm/mach-omap2/omap-mpuss-lowpower.c
+++ b/arch/arm/mach-omap2/omap-mpuss-lowpower.c
@@ -64,6 +64,7 @@
 #include "prm-regbits-44xx.h"
 
 static void __iomem *sar_base;
+static u32 old_cpu1_ns_pa_addr;
 
 #if defined(CONFIG_PM) && defined(CONFIG_SMP)
 
@@ -213,6 +214,35 @@ static void __init save_l2x0_context(void)
 #endif
 
 /**
+ * omap4_cpu1_may_need_reset: Check if cpu1 needs to be reset on boot
+ *
+ * If cpu1 is configured for idle in the bootloader or in the previous
+ * kernel after kexec, it will wake-up from idle state to the configured
+ * restore address which is unsafe for booting kernel. In that case all
+ * we can do is reset cpu1 to get it to start_secondary. In at least
+ * omap5 case, the ROM code properly parks the bootloader at a higher
+ * kernel address, so for those cases no reset is needed. For anything
+ * configured for the first 1GiB@0x80000000 let's assume we need a
+ * reset.
+ */
+bool omap4_cpu1_may_need_reset(void)
+{
+	bool unsafe_cpu1_ns_pa_addr;
+
+	if (!old_cpu1_ns_pa_addr)
+		return false;
+
+	unsafe_cpu1_ns_pa_addr = ((old_cpu1_ns_pa_addr >> 24) == 0x80);
+	if (!unsafe_cpu1_ns_pa_addr)
+		return false;
+
+	pr_info("smp: omap has configured ns_pa_addr: 0x%08x\n",
+		old_cpu1_ns_pa_addr);
+
+	return true;
+}
+
+/**
  * omap4_enter_lowpower: OMAP4 MPUSS Low Power Entry Function
  * The purpose of this function is to manage low power programming
  * of OMAP4 MPUSS subsystem
@@ -371,7 +401,7 @@ int __init omap4_mpuss_init(void)
 	pm_info = &per_cpu(omap4_pm_info, 0x0);
 	if (sar_base) {
 		pm_info->scu_sar_addr = sar_base + SCU_OFFSET0;
-		if (cpu_is_omap44xx())
+		if (soc_is_omap44xx())
 			pm_info->wkup_sar_addr = sar_base +
 				CPU0_WAKEUP_NS_PA_ADDR_OFFSET;
 		else
@@ -395,7 +425,7 @@ int __init omap4_mpuss_init(void)
 	pm_info = &per_cpu(omap4_pm_info, 0x1);
 	if (sar_base) {
 		pm_info->scu_sar_addr = sar_base + SCU_OFFSET1;
-		if (cpu_is_omap44xx())
+		if (soc_is_omap44xx())
 			pm_info->wkup_sar_addr = sar_base +
 				CPU1_WAKEUP_NS_PA_ADDR_OFFSET;
 		else
@@ -432,7 +462,7 @@ int __init omap4_mpuss_init(void)
 		save_l2x0_context();
 	}
 
-	if (cpu_is_omap44xx()) {
+	if (soc_is_omap44xx()) {
 		omap_pm_ops.finish_suspend = omap4_finish_suspend;
 		omap_pm_ops.resume = omap4_cpu_resume;
 		omap_pm_ops.scu_prepare = scu_pwrst_prepare;
@@ -443,7 +473,7 @@ int __init omap4_mpuss_init(void)
 		enable_mercury_retention_mode();
 	}
 
-	if (cpu_is_omap446x())
+	if (soc_is_omap446x())
 		omap_pm_ops.hotplug_restart = omap4460_secondary_startup;
 
 	return 0;
@@ -460,22 +490,30 @@ int __init omap4_mpuss_init(void)
 void __init omap4_mpuss_early_init(void)
 {
 	unsigned long startup_pa;
+	void __iomem *ns_pa_addr;
 
-	if (!(cpu_is_omap44xx() || soc_is_omap54xx()))
+	if (!(soc_is_omap44xx() || soc_is_omap54xx()))
 		return;
 
 	sar_base = omap4_get_sar_ram_base();
 
-	if (cpu_is_omap443x())
+	/* Restore old NS_PA_ADDR for validity checks later on */
+	if (soc_is_omap44xx())
+		ns_pa_addr = sar_base + CPU1_WAKEUP_NS_PA_ADDR_OFFSET;
+	else
+		ns_pa_addr = sar_base + OMAP5_CPU1_WAKEUP_NS_PA_ADDR_OFFSET;
+	old_cpu1_ns_pa_addr = readl_relaxed(ns_pa_addr);
+
+	if (soc_is_omap443x())
 		startup_pa = __pa_symbol(omap4_secondary_startup);
-	else if (cpu_is_omap446x())
+	else if (soc_is_omap446x())
 		startup_pa = __pa_symbol(omap4460_secondary_startup);
 	else if ((__boot_cpu_mode & MODE_MASK) == HYP_MODE)
 		startup_pa = __pa_symbol(omap5_secondary_hyp_startup);
 	else
 		startup_pa = __pa_symbol(omap5_secondary_startup);
 
-	if (cpu_is_omap44xx())
+	if (soc_is_omap44xx())
 		writel_relaxed(startup_pa, sar_base +
 			       CPU1_WAKEUP_NS_PA_ADDR_OFFSET);
 	else
diff --git a/arch/arm/mach-omap2/omap-smp.c b/arch/arm/mach-omap2/omap-smp.c
--- a/arch/arm/mach-omap2/omap-smp.c
+++ b/arch/arm/mach-omap2/omap-smp.c
@@ -300,10 +300,13 @@ static void __init omap4_smp_prepare_cpus(unsigned int max_cpus)
 		scu_enable(cfg.scu_base);
 
 	/*
-	 * Reset CPU1 before configuring, otherwise kexec will
-	 * end up trying to use old kernel startup address.
+	 * Reset CPU1 before configuring, otherwise kexec can
+	 * end up trying to use old kernel startup address or
+	 * suspen/resume will fail bring up CPU1. Seen only on
+	 * 4430..
 	 */
-	if (cfg.cpu1_rstctrl_va) {
+	if (omap4_cpu1_may_need_reset() && cfg.cpu1_rstctrl_va) {
+		pr_info("smp: omap cpu1 idle configured, needs reset\n");
 		writel_relaxed(1, cfg.cpu1_rstctrl_va);
 		readl_relaxed(cfg.cpu1_rstctrl_va);
 		writel_relaxed(0, cfg.cpu1_rstctrl_va);
-- 
2.11.1

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: tony@atomide.com (Tony Lindgren)
Date: 2017-02-17 15:55:02

* Tony Lindgren [off-list ref] [170216 11:08]:
* Tony Lindgren [off-list ref] [170216 08:55]:
quoted
For your use case, probably all we need is runtime checks for HS in
addition to parking cpu1 for kexec. If that's not enough, then maybe
a device specific DT property for never-reset-no-matter-what.
Below is a first take on the last resort cpu1 reset done based on
configured CPU1_WAKEUP_NS_PA_ADDR_OFFSET for omap4 and 5. Note that
we can't merge it yet as it will break kexec boot for dra7. To fix that,
we need to first do what you're suggesting and properly park cpu1 for
kexec.
Something like the patch below might be doable for the first fix
before we have something to park cpu1 for kexec. I've added more checks
to attemp to detect cpu1 being in WFI. At least things keep on working
for kexec on omap4/5 and GP dra7.

Andrew, care to give it a try and see if can add some HS dra7 checks
there too to have kexec work on it? Basically we want to attempt to
detect if there's a chance cpu1 is in WFI, and if I think the only
optino is to reset it before bringing it up as otherwise outcome will
be unpredictable.

Regards,

Tony

8< -----------------------
diff --git a/arch/arm/mach-omap2/common.h b/arch/arm/mach-omap2/common.h
--- a/arch/arm/mach-omap2/common.h
+++ b/arch/arm/mach-omap2/common.h
@@ -270,6 +270,7 @@ extern const struct smp_operations omap4_smp_ops;
 extern int omap4_mpuss_init(void);
 extern int omap4_enter_lowpower(unsigned int cpu, unsigned int power_state);
 extern int omap4_hotplug_cpu(unsigned int cpu, unsigned int power_state);
+extern u32 omap4_get_cpu1_ns_pa_addr(void);
 #else
 static inline int omap4_enter_lowpower(unsigned int cpu,
 					unsigned int power_state)
diff --git a/arch/arm/mach-omap2/omap-mpuss-lowpower.c b/arch/arm/mach-omap2/omap-mpuss-lowpower.c
--- a/arch/arm/mach-omap2/omap-mpuss-lowpower.c
+++ b/arch/arm/mach-omap2/omap-mpuss-lowpower.c
@@ -64,6 +64,7 @@
 #include "prm-regbits-44xx.h"
 
 static void __iomem *sar_base;
+static u32 old_cpu1_ns_pa_addr;
 
 #if defined(CONFIG_PM) && defined(CONFIG_SMP)
 
@@ -212,6 +213,11 @@ static void __init save_l2x0_context(void)
 {}
 #endif
 
+u32 omap4_get_cpu1_ns_pa_addr(void)
+{
+	return old_cpu1_ns_pa_addr;
+}
+
 /**
  * omap4_enter_lowpower: OMAP4 MPUSS Low Power Entry Function
  * The purpose of this function is to manage low power programming
@@ -371,7 +377,7 @@ int __init omap4_mpuss_init(void)
 	pm_info = &per_cpu(omap4_pm_info, 0x0);
 	if (sar_base) {
 		pm_info->scu_sar_addr = sar_base + SCU_OFFSET0;
-		if (cpu_is_omap44xx())
+		if (soc_is_omap44xx())
 			pm_info->wkup_sar_addr = sar_base +
 				CPU0_WAKEUP_NS_PA_ADDR_OFFSET;
 		else
@@ -395,7 +401,7 @@ int __init omap4_mpuss_init(void)
 	pm_info = &per_cpu(omap4_pm_info, 0x1);
 	if (sar_base) {
 		pm_info->scu_sar_addr = sar_base + SCU_OFFSET1;
-		if (cpu_is_omap44xx())
+		if (soc_is_omap44xx())
 			pm_info->wkup_sar_addr = sar_base +
 				CPU1_WAKEUP_NS_PA_ADDR_OFFSET;
 		else
@@ -432,7 +438,7 @@ int __init omap4_mpuss_init(void)
 		save_l2x0_context();
 	}
 
-	if (cpu_is_omap44xx()) {
+	if (soc_is_omap44xx()) {
 		omap_pm_ops.finish_suspend = omap4_finish_suspend;
 		omap_pm_ops.resume = omap4_cpu_resume;
 		omap_pm_ops.scu_prepare = scu_pwrst_prepare;
@@ -443,7 +449,7 @@ int __init omap4_mpuss_init(void)
 		enable_mercury_retention_mode();
 	}
 
-	if (cpu_is_omap446x())
+	if (soc_is_omap446x())
 		omap_pm_ops.hotplug_restart = omap4460_secondary_startup;
 
 	return 0;
@@ -460,22 +466,30 @@ int __init omap4_mpuss_init(void)
 void __init omap4_mpuss_early_init(void)
 {
 	unsigned long startup_pa;
+	void __iomem *ns_pa_addr;
 
-	if (!(cpu_is_omap44xx() || soc_is_omap54xx()))
+	if (!(soc_is_omap44xx() || soc_is_omap54xx()))
 		return;
 
 	sar_base = omap4_get_sar_ram_base();
 
-	if (cpu_is_omap443x())
+	/* Restore old NS_PA_ADDR for validity checks later on */
+	if (soc_is_omap44xx())
+		ns_pa_addr = sar_base + CPU1_WAKEUP_NS_PA_ADDR_OFFSET;
+	else
+		ns_pa_addr = sar_base + OMAP5_CPU1_WAKEUP_NS_PA_ADDR_OFFSET;
+	old_cpu1_ns_pa_addr = readl_relaxed(ns_pa_addr);
+
+	if (soc_is_omap443x())
 		startup_pa = __pa_symbol(omap4_secondary_startup);
-	else if (cpu_is_omap446x())
+	else if (soc_is_omap446x())
 		startup_pa = __pa_symbol(omap4460_secondary_startup);
 	else if ((__boot_cpu_mode & MODE_MASK) == HYP_MODE)
 		startup_pa = __pa_symbol(omap5_secondary_hyp_startup);
 	else
 		startup_pa = __pa_symbol(omap5_secondary_startup);
 
-	if (cpu_is_omap44xx())
+	if (soc_is_omap44xx())
 		writel_relaxed(startup_pa, sar_base +
 			       CPU1_WAKEUP_NS_PA_ADDR_OFFSET);
 	else
diff --git a/arch/arm/mach-omap2/omap-smp.c b/arch/arm/mach-omap2/omap-smp.c
--- a/arch/arm/mach-omap2/omap-smp.c
+++ b/arch/arm/mach-omap2/omap-smp.c
@@ -44,6 +44,7 @@ struct omap_smp_config {
 	unsigned long cpu1_rstctrl_pa;
 	void __iomem *cpu1_rstctrl_va;
 	void __iomem *scu_base;
+	void __iomem *wakeupgen_base;
 	void *startup_addr;
 };
 
@@ -140,7 +141,6 @@ static int omap4_boot_secondary(unsigned int cpu, struct task_struct *idle)
 	static struct clockdomain *cpu1_clkdm;
 	static bool booted;
 	static struct powerdomain *cpu1_pwrdm;
-	void __iomem *base = omap_get_wakeupgen_base();
 
 	/*
 	 * Set synchronisation state between this boot processor
@@ -157,7 +157,7 @@ static int omap4_boot_secondary(unsigned int cpu, struct task_struct *idle)
 	if (omap_secure_apis_support())
 		omap_modify_auxcoreboot0(0x200, 0xfffffdff);
 	else
-		writel_relaxed(0x20, base + OMAP_AUX_CORE_BOOT_0);
+		writel_relaxed(0x20, cfg.wakeupgen_base + OMAP_AUX_CORE_BOOT_0);
 
 	if (!cpu1_clkdm && !cpu1_pwrdm) {
 		cpu1_clkdm = clkdm_lookup("mpu1_clkdm");
@@ -261,9 +261,62 @@ static void __init omap4_smp_init_cpus(void)
 		set_cpu_possible(i, true);
 }
 
+/*
+ * For now, just make sure the start-up address is not within
+ * the first 1GB as that most likely means that CPU1 is configured
+ * by either the bootloader or previous kernel in kexec boot to
+ * something that will most likely fail without a reset.
+ */
+static bool __init omap4_smp_cpu1_startup_valid(unsigned long addr)
+{
+	if ((addr >> 24) == 0x80)
+		return false;
+
+	return true;
+}
+
+/*
+ * We may need to reset CPU1 before configuring, otherwise kexec can end up
+ * trying to use old kernel startup address or suspend-resume will occasionally
+ * fail to bring up CPU1 on 4430 if CPU1 fails to enter deeper idle states.
+ */
+static void __init omap4_smp_maybe_reset_cpu1(struct omap_smp_config *c)
+{
+	unsigned long cpu1_startup_pa, cpu1_ns_pa_addr;
+	bool needs_reset = false;
+
+	cpu1_startup_pa = readl_relaxed(cfg.wakeupgen_base +
+					OMAP_AUX_CORE_BOOT_1);
+	cpu1_ns_pa_addr = omap4_get_cpu1_ns_pa_addr();
+
+	/* REVISIT: Anything to check for HS dra7? */
+	if (soc_is_dra74x() && omap_secure_apis_support())
+		return;
+
+	/* If dra7 has AUX_CORE_BOOT_1 within first 1GB, CPU1 may be in WFI */
+	if (soc_is_dra74x() && !omap_secure_apis_support() &&
+	    !omap4_smp_cpu1_startup_valid(cpu1_startup_pa))
+		needs_reset = true;
+
+	/* If omap4 or 5 has NS_PA_ADDR within first 1GB, CPU1 may be in WFI */
+	if ((soc_is_omap44xx() || soc_is_omap54xx()) &&
+	    !omap4_smp_cpu1_startup_valid(cpu1_ns_pa_addr) &&
+	    c->cpu1_rstctrl_va)
+		needs_reset = true;
+
+	if (!needs_reset || !c->cpu1_rstctrl_va)
+		return;
+
+	pr_info("smp: Already configured omap cpu1, needs reset (0x%lx 0x%lx)\n",
+		cpu1_startup_pa, cpu1_ns_pa_addr);
+
+	writel_relaxed(1, c->cpu1_rstctrl_va);
+	readl_relaxed(c->cpu1_rstctrl_va);
+	writel_relaxed(0, c->cpu1_rstctrl_va);
+}
+
 static void __init omap4_smp_prepare_cpus(unsigned int max_cpus)
 {
-	void __iomem *base = omap_get_wakeupgen_base();
 	const struct omap_smp_config *c = NULL;
 
 	if (soc_is_omap443x())
@@ -281,6 +334,7 @@ static void __init omap4_smp_prepare_cpus(unsigned int max_cpus)
 	/* Must preserve cfg.scu_base set earlier */
 	cfg.cpu1_rstctrl_pa = c->cpu1_rstctrl_pa;
 	cfg.startup_addr = c->startup_addr;
+	cfg.wakeupgen_base = omap_get_wakeupgen_base();
 
 	if (soc_is_dra74x() || soc_is_omap54xx()) {
 		if ((__boot_cpu_mode & MODE_MASK) == HYP_MODE)
@@ -299,15 +353,7 @@ static void __init omap4_smp_prepare_cpus(unsigned int max_cpus)
 	if (cfg.scu_base)
 		scu_enable(cfg.scu_base);
 
-	/*
-	 * Reset CPU1 before configuring, otherwise kexec will
-	 * end up trying to use old kernel startup address.
-	 */
-	if (cfg.cpu1_rstctrl_va) {
-		writel_relaxed(1, cfg.cpu1_rstctrl_va);
-		readl_relaxed(cfg.cpu1_rstctrl_va);
-		writel_relaxed(0, cfg.cpu1_rstctrl_va);
-	}
+	omap4_smp_maybe_reset_cpu1(&cfg);
 
 	/*
 	 * Write the address of secondary startup routine into the
@@ -319,7 +365,7 @@ static void __init omap4_smp_prepare_cpus(unsigned int max_cpus)
 		omap_auxcoreboot_addr(__pa_symbol(cfg.startup_addr));
 	else
 		writel_relaxed(__pa_symbol(cfg.startup_addr),
-			       base + OMAP_AUX_CORE_BOOT_1);
+			       cfg.wakeupgen_base + OMAP_AUX_CORE_BOOT_1);
 }
 
 const struct smp_operations omap4_smp_ops __initconst = {
-- 
2.11.1

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: Andrew F. Davis <hidden>
Date: 2017-02-17 20:27:07

On 02/17/2017 09:55 AM, Tony Lindgren wrote:
* Tony Lindgren [off-list ref] [170216 11:08]:
quoted
* Tony Lindgren [off-list ref] [170216 08:55]:
quoted
For your use case, probably all we need is runtime checks for HS in
addition to parking cpu1 for kexec. If that's not enough, then maybe
a device specific DT property for never-reset-no-matter-what.
Below is a first take on the last resort cpu1 reset done based on
configured CPU1_WAKEUP_NS_PA_ADDR_OFFSET for omap4 and 5. Note that
we can't merge it yet as it will break kexec boot for dra7. To fix that,
we need to first do what you're suggesting and properly park cpu1 for
kexec.
Something like the patch below might be doable for the first fix
before we have something to park cpu1 for kexec. I've added more checks
to attemp to detect cpu1 being in WFI. At least things keep on working
for kexec on omap4/5 and GP dra7.

Andrew, care to give it a try and see if can add some HS dra7 checks
there too to have kexec work on it? Basically we want to attempt to
detect if there's a chance cpu1 is in WFI, and if I think the only
optino is to reset it before bringing it up as otherwise outcome will
be unpredictable.
This patch seems stop the reset issue on my boards, so as long as it is
just a temporary workaround until we can fix the real issue (a core that
should be off going into omap_do_wfi() instead of parking), then this
patch works for me.

[...]
+
+	/* REVISIT: Anything to check for HS dra7? */
+	if (soc_is_dra74x() && omap_secure_apis_support())
+		return;
omap_secure_apis_support() only returns true when cpu_is_omap44xx() is
true, so this will never be true. Someday we may add the secure APIs to
secure DRA7xx devices, but I don't know of a good way to detect this
from kernel right now.

Andrew

[PATCH] ARM: omap2+: Revert omap-smp.c changes resetting cpu1 during boot

From: tony@atomide.com (Tony Lindgren)
Date: 2017-02-17 21:09:40

* Andrew F. Davis [off-list ref] [170217 12:28]:
On 02/17/2017 09:55 AM, Tony Lindgren wrote:
quoted
* Tony Lindgren [off-list ref] [170216 11:08]:
quoted
* Tony Lindgren [off-list ref] [170216 08:55]:
quoted
For your use case, probably all we need is runtime checks for HS in
addition to parking cpu1 for kexec. If that's not enough, then maybe
a device specific DT property for never-reset-no-matter-what.
Below is a first take on the last resort cpu1 reset done based on
configured CPU1_WAKEUP_NS_PA_ADDR_OFFSET for omap4 and 5. Note that
we can't merge it yet as it will break kexec boot for dra7. To fix that,
we need to first do what you're suggesting and properly park cpu1 for
kexec.
Something like the patch below might be doable for the first fix
before we have something to park cpu1 for kexec. I've added more checks
to attemp to detect cpu1 being in WFI. At least things keep on working
for kexec on omap4/5 and GP dra7.

Andrew, care to give it a try and see if can add some HS dra7 checks
there too to have kexec work on it? Basically we want to attempt to
detect if there's a chance cpu1 is in WFI, and if I think the only
optino is to reset it before bringing it up as otherwise outcome will
be unpredictable.
This patch seems stop the reset issue on my boards, so as long as it is
just a temporary workaround until we can fix the real issue (a core that
should be off going into omap_do_wfi() instead of parking), then this
patch works for me.
OK
quoted
+
+	/* REVISIT: Anything to check for HS dra7? */
+	if (soc_is_dra74x() && omap_secure_apis_support())
+		return;
omap_secure_apis_support() only returns true when cpu_is_omap44xx() is
true, so this will never be true. Someday we may add the secure APIs to
secure DRA7xx devices, but I don't know of a good way to detect this
from kernel right now.
OK so that can be removed.

Can you check what happens with suspend/resume cycle? Based on what
I've seen omap5 and dra7 currently always hit the "smp_ops.cpu_die()
returned, trying to resuscitate" on resume meaning it will get sent to
secondary_start_kernel on resume.

Regards,

Tony
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help