Re: Machine Check in P2010(e500v2)
From: Joakim Tjernlund <hidden>
Date: 2017-09-14 16:55:49
On Sat, 2017-09-09 at 14:59 +0200, Joakim Tjernlund wrote:
On Sat, 2017-09-09 at 14:45 +0200, Joakim Tjernlund wrote:quoted
On Fri, 2017-09-08 at 22:27 +0000, Leo Li wrote:quoted
quoted
-----Original Message----- From: Joakim Tjernlund [mailto:Joakim.Tjernlund@infinera.com] Sent: Friday, September 08, 2017 7:51 AM To: linuxppc-dev@lists.ozlabs.org; Leo Li <redacted>; Yor=
k Sun
quoted
quoted
quoted
[off-list ref] Subject: Re: Machine Check in P2010(e500v2) =20 On Fri, 2017-09-08 at 11:54 +0200, Joakim Tjernlund wrote:quoted
On Thu, 2017-09-07 at 18:54 +0000, Leo Li wrote:quoted
quoted
-----Original Message----- From: Joakim Tjernlund [mailto:Joakim.Tjernlund@infinera.com] Sent: Thursday, September 07, 2017 3:41 AM To: linuxppc-dev@lists.ozlabs.org; Leo Li <leoyang.li@nxp.com=;quoted
quoted
quoted
quoted
quoted
quoted
York Sun [off-list ref] Subject: Re: Machine Check in P2010(e500v2) =20 On Thu, 2017-09-07 at 00:50 +0200, Joakim Tjernlund wrote:quoted
On Wed, 2017-09-06 at 21:13 +0000, Leo Li wrote:quoted
quoted
-----Original Message----- From: Joakim Tjernlund [mailto:Joakim.Tjernlund@infinera.com] Sent: Wednesday, September 06, 2017 3:54 PM To: linuxppc-dev@lists.ozlabs.org; Leo Li [off-list ref]; York Sun [off-list ref] Subject: Re: Machine Check in P2010(e500v2) =20 On Wed, 2017-09-06 at 20:28 +0000, Leo Li wrote:quoted
quoted
-----Original Message----- From: Joakim Tjernlund [mailto:Joakim.Tjernlund@infinera.com] Sent: Wednesday, September 06, 2017 3:17 PM To: linuxppc-dev@lists.ozlabs.org; Leo Li [off-list ref]; York Sun [off-list ref] Subject: Re: Machine Check in P2010(e500v2) =20 On Wed, 2017-09-06 at 19:31 +0000, Leo Li wrote:quoted
quoted
-----Original Message----- From: York Sun Sent: Wednesday, September 06, 2017 10:38 AM To: Joakim Tjernlund [off-list ref]; linuxppc- dev@lists.ozlabs.org; Leo Li [off-list ref] Subject: Re: Machine Check in P2010(e500v2) =20 Scott is no longer with Freescale/NXP. Adding L=
eo.
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
=20 On 09/05/2017 01:40 AM, Joakim Tjernlund wrote:quoted
So after some debugging I found this bug:@@ -996,7 +998,7 @@ intfsl_pci_mcheck_exception(struct pt_regs=20 *regs)quoted
quoted
quoted
quoted
quoted
if (is_in_pci_mem_space(addr)) { if (user_mode(regs)) { pagefault_disable(); - ret =3D get_user(regs=
->nip, &inst);
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
+ ret =3D get_user(inst=
,
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
+ (__u32 __user *)regs->nip); pagefault_enable(); } else { ret =3D probe_kernel_address(regs->nip, inst); =20 However, the kernel still locked up after fix=
ing that.
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
Now I wonder why this fixup is there in the f=
irst place?
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
The routine will not really fixup the insn, j=
ust
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
return 0xffffffff for the failing read and th=
en advance the
quoted
quoted
quoted
=20 process NIP.quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
=20 You are right. The code here only gives 0xffffff=
ff to
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
the load instructions and=20 continue with the next instruction when the load instruction is causing the machine check. This wil=
l
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
prevent a system lockup when reading from PCI/Rapid=
IO device
quoted
quoted
quoted
=20 which is link down.quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
=20 I don't know what is actual problem in your case. Maybe it is a write=20 instruction instead of read? Or the code is in a =
infinite loop
quoted
quoted
quoted
=20 waiting forquoted
quoted
quoted
=20 aquoted
quoted
quoted
=20 validquoted
quoted
read result? Are you able to do some further debug=
ging
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
with the NIP correctly printed?quoted
=20=20 According to the MC it is a Read and the NIP also l=
eads
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
to a read in the=20 program.quoted
quoted
ATM, I have disabled the fixup but I will enable th=
at again.
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
Question, is it safe add a small printk when this M=
C
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
happens(after fixing up)? I need to see that it has happened as the error is somewhat=20 random.quoted
=20 I think it is safe to add printk as the current machi=
ne
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
check handlers are also=20 using printk. =20 I hope so, but if the fixup fires there is no printk at=
all so I was a bit
quoted
quoted
quoted
=20 unsure.quoted
quoted
quoted
quoted
quoted
quoted
Don't like this fixup though, is there not a better way=
than
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
faking a read to user space(or kernel for that matter) =
?
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
=20 I don't have a better idea. Without the fixup, the offen=
ding
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
load instruction=20 will never finish if there is anything wrong with the backing device and freeze the whole system. Do you have any suggesti=
on in mind?
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
=20=20 But it never finishes the load, it just fakes a load of 0xfffffffff, for user space I rather have it signal a SIGBU=
S but
quoted
quoted
quoted
quoted
quoted
quoted
quoted
that does not seem to work either, at least not for us but =
that
quoted
quoted
quoted
quoted
quoted
quoted
quoted
could be a bug in general MC code=20 maybe.quoted
This fixup might be valid for kernel only as it has never w=
orked
quoted
quoted
quoted
quoted
quoted
quoted
quoted
for user space=20 due to the bug I found.quoted
=20 Where can I read about this errata ?=20 I have look high and low an cannot find an errata which maps =
to this fixup.
quoted
quoted
quoted
quoted
quoted
quoted
The closest I get is A-005125 which seems to have another workaround, I cannot find any evidence that this workaround h=
as been
quoted
quoted
quoted
=20 applied in Linux, can you?quoted
quoted
=20 This is not A-005125. There was an erratum for this issue with=
older silicons
quoted
quoted
quoted
=20 (e.g. erratum PCI-ex 3 for MPC8572).quoted
quoted
" When its link goes down, the PCI Express controller clears al=
l
quoted
quoted
quoted
quoted
quoted
outstanding transactions with an error indicator and sends a li=
nk
quoted
quoted
quoted
quoted
quoted
down exception to the interrupt controller if PEX_PME_MES_DISR[=
LDDD]
quoted
quoted
quoted
quoted
quoted
=3D 0. If, however, any transactions are sent to the controller=
after
quoted
quoted
quoted
quoted
quoted
the link down event, they are accepted by the controller and wa=
it
quoted
quoted
quoted
quoted
quoted
for the link to come back up before starting any timeout counte=
rs (for
quoted
quoted
quoted
=20 example, completion timeout). There is no mechanism to cancel the n=
ew
quoted
quoted
quoted
transactions short of a device HRESET. "quoted
quoted
=20 But it was removed in newer silicon like P2020/P2010 probably b=
ecause a
quoted
quoted
quoted
=20 Machine Check will be triggered in this situation to deal with the =
stalled
quoted
quoted
quoted
instruction and no longer considered it as a hardware issue.quoted
quoted
=20=20 Maybe this fixup should be configurable then?=20 No. My point is that the problem was no longer considered a hardware=
issue because of the machine check mechanism is in place to handle it. If= there is no handling of this special case, we would still experience a sys= tem hang if this situation really occurs.
quoted
quoted
=20quoted
quoted
=20quoted
The A-005125 is dealt with in u-boot.=20 https://emea01.safelinks.protection.outlook.com/?url=3Dhttps%3A%2F%=
2Flists.de
quoted
quoted
quoted
nx.de%2Fpipermail%2Fu-boot%2F2013- August%2F161185.html&data=3D01%7C01%7Cleoyang.li%40nxp.com%7Ccb8a93=
e
quoted
quoted
quoted
0090e48eb53a008d4f6b84235%7C686ea1d3bc2b4c6fa92cd99c5c301635%7C0& sdata=3D8sR4yoXA4adqMHz6TY%2BvmYpfCBTcYEZHjPuANjz%2F1EQ%3D&reserve d=3D0quoted
=20 Yes, I found it eventually :) =20 However, I cannot return to normal execution. I can follow the co=
de to
quoted
quoted
quoted
quoted
returning from machine_check_exception() and moving into ASM handler for returni=
ng
quoted
quoted
quoted
quoted
from a ME but then I am a bit lost. It does not seem to be any pr=
oblem
quoted
quoted
quoted
quoted
executing, it feels more like a SW bug dealing with machine check=
s. Don't
quoted
quoted
quoted
=20 known how to diagnose this further and could use some pointers.=20 Is the execution returned to the user application? I doubt the syste=
m hang is caused by the machine check handling.
quoted
quoted
You can try to comment out the machine check handling code and check =
if there is any improvement and see if
quoted
quoted
this is related to the machine check handling.=20 It tries to return to user app but I cannot see what happens as the sys=
tem lock up when the
quoted
MC returns. How do you mean comment out MC handling? The simplest path is the PCI f=
ixup which will
quoted
just do regs->nip +=3D 4; and then return to user space. That still doe=
s not work as
quoted
as soon MC handling returns, the system is locked up. =20quoted
=20 Machine check is a serious situation and not always possible to be re=
covered from.=20
quoted
=20 This one should at least not kill the whole system. It is a simple bus =
error in user space and
quoted
the app should get SIGBUS and the the system should carry on.=20 =20quoted
I would focus more on debugging why the machine check is triggered by=
the user space application.
quoted
quoted
Can you locate what code is causing this machine check from user spac=
e? =20
quoted
quoted
Is it accessing some hardware related space which is not ready?=20 Or is it accessing address that it shouldn't have accessed?=20 of course, this is ongoing and getting closer a solution. The MC lookin=
g the machine completely
quoted
does not make this any easier though. These are 2 separate things, fixing the cause and not having a simple b=
us error lock up the machine.
quoted
I am focusing on fixing the lockup. =20 I have been following the execution in the kernel and I always end up i=
n the ASM returning
quoted
from the MC. The other day we got a similar PCI MC(bus error) on T1042 CPU(e5500/e50=
0mc) and there
quoted
the system survived. The one thing I see different there is that MSR RI=
is set
quoted
when entering MC, why is that?=20 Before you ask, I have tried to add MSR_RI to both msr and mcsrr1. Didn't=
help. I managed to provoke another Machine Check, much earlier this time: [ 15.047108] Machine check in kernel mode. [ 15.051120] Caused by (from MCSR=3D10008): Bus - Read Data Bus Error [ 15.057302] Oops: Machine check, sig: 7 [#1] [ 15.061567] P1010 RDB [ 15.063832] Modules linked in: linux_bcm_knet(PO) linux_user_bde(PO) lin= ux_kernel_bde(PO) [ 15.072022] CPU: 0 PID: 472 Comm: emxp2_hw_bl Tainted: P O = 4.1.43+ #52 [ 15.079680] task: db1a7990 ti: df18c000 task.ti: df18c000 [ 15.085075] NIP: 00000000 LR: 109e7648 CTR: 00000000 [ 15.090036] REGS: df18df10 TRAP: 0204 Tainted: P O (4.1.= 43+) [ 15.097082] MSR: 0002d000 <CE,EE,PR,ME> CR: 280004e8 XER: 20000000 [ 15.103448] DEAR: b6e44140 ESR: 00000000=20 GPR00: 10ac1160 bfa44010 b79734a0 136eb4a0 bfa44030 01010101 bfa44038 00000= 020=20 GPR08: 00000000 b6e13000 063e521e 0f9ed9c4 22000422 11db7334 00000000 00000= 000=20 GPR16: 10f8b054 10f895e5 10f8a8bf 00031150 136eb4d0 00030000 00031140 00031= 140=20 GPR24: 00000000 00000000 136f10a0 00000000 00000000 00000000 00031140 136eb= 4a0=20 [ 15.135690] NIP [00000000] (null) [ 15.139174] LR [109e7648] 0x109e7648 [ 15.142743] Call Trace: [ 15.145184] ---[ end trace c00af6117685cb6e ]--- The fun part is that now the OS did NOT lock up! Looking that the faulting process, emxp2_hw_bl, I see it is in Zombie state= (cd /proc/472): cat status=20 Name: emxp2_hw_bl State: Z (zombie) Tgid: 472 Ngid: 0 Pid: 472 PPid: 468 TracerPid: 0 Uid: 0 0 0 0 Gid: 0 0 0 0 FDSize: 0 Groups:=09 Threads: 8 SigQ: 0/3462 SigPnd: 0000000000000000 ShdPnd: 0000000000000000 SigBlk: 0000000000000000 SigIgn: 0000000000001000 SigCgt: 00000001c0000628 CapInh: 0000000000000000 CapPrm: 0000003fffffffff CapEff: 0000003fffffffff CapBnd: 0000003fffffffff Cpus_allowed: 1 Cpus_allowed_list: 0 voluntary_ctxt_switches: 1126 nonvoluntary_ctxt_switches: 376 This even after parent process has called waitid(2) for emxp2_hw_bl If I now do a kill -s SIGBUS/TERM <pid of emxp2_hw_bl> this signal is propagated to the parent and emxp2_hw_bl goes away. Stack: cat stack=20 [<c0071c04>] do_futex+0x150/0x874 [<c0027670>] do_exit+0x4e8/0x7d0 [<c000a164>] die+0x178/0x1d8 [<c000a7c8>] machine_check_exception+0xcc/0x17c [<c000dd94>] ret_from_mcheck_exc+0x0/0x144 So emxp2_hw_bl is stuck somewhere in down in machine_check_exception(). This all looks like Linux bugs when asked to kill a user process from Machine Check. I don't think I will get any further without some pointers now. Jocke