Re: Machine Check in P2010(e500v2)
From: Joakim Tjernlund <hidden>
Date: 2017-09-06 22:50:13
On Wed, 2017-09-06 at 21:13 +0000, Leo Li wrote:
quoted
-----Original Message----- From: Joakim Tjernlund [mailto:Joakim.Tjernlund@infinera.com] Sent: Wednesday, September 06, 2017 3:54 PM To: linuxppc-dev@lists.ozlabs.org; Leo Li <redacted>; York Su=
n
quoted
[off-list ref] Subject: Re: Machine Check in P2010(e500v2) =20 On Wed, 2017-09-06 at 20:28 +0000, Leo Li wrote:quoted
quoted
-----Original Message----- From: Joakim Tjernlund [mailto:Joakim.Tjernlund@infinera.com] Sent: Wednesday, September 06, 2017 3:17 PM To: linuxppc-dev@lists.ozlabs.org; Leo Li <redacted>; Yor=
k
quoted
quoted
quoted
Sun [off-list ref] Subject: Re: Machine Check in P2010(e500v2) =20 On Wed, 2017-09-06 at 19:31 +0000, Leo Li wrote:quoted
quoted
-----Original Message----- From: York Sun Sent: Wednesday, September 06, 2017 10:38 AM To: Joakim Tjernlund <redacted>; linuxppc- dev@lists.ozlabs.org; Leo Li [off-list ref] Subject: Re: Machine Check in P2010(e500v2) =20 Scott is no longer with Freescale/NXP. Adding Leo. =20 On 09/05/2017 01:40 AM, Joakim Tjernlund wrote:quoted
So after some debugging I found this bug:@@ -996,7 +998,7 @@ int fsl_pci_mcheck_exception(struct pt_re=
gs
quoted
=20 *regs)quoted
quoted
quoted
quoted
quoted
if (is_in_pci_mem_space(addr)) { if (user_mode(regs)) { pagefault_disable(); - ret =3D get_user(regs->nip, &inst); + ret =3D get_user(inst, (__u32 __user + *)regs->nip); pagefault_enable(); } else { ret =3D probe_kernel_address(regs->n=
ip,
quoted
quoted
quoted
quoted
quoted
quoted
inst); =20 However, the kernel still locked up after fixing that. Now I wonder why this fixup is there in the first place? The routine will not really fixup the insn, just return 0xfffffff=
f
quoted
quoted
quoted
quoted
quoted
quoted
for the failing read and then advance the process NIP.=20 You are right. The code here only gives 0xffffffff to the load instructions and=20 continue with the next instruction when the load instruction is causing the machine check. This will prevent a system lockup when reading from PCI/RapidIO device which is link down.quoted
=20 I don't know what is actual problem in your case. Maybe it is a write=20 instruction instead of read? Or the code is in a infinite loop wa=
iting for a
quoted
=20 validquoted
quoted
read result? Are you able to do some further debugging with the NI=
P
quoted
quoted
quoted
correctly printed?quoted
=20=20 According to the MC it is a Read and the NIP also leads to a read i=
n the
quoted
=20 program.quoted
quoted
ATM, I have disabled the fixup but I will enable that again. Question, is it safe add a small printk when this MC happens(after fixing up)? I need to see that it has happened as the error is some=
what
quoted
=20 random.quoted
=20 I think it is safe to add printk as the current machine check handler=
s are also
quoted
=20 using printk. =20 I hope so, but if the fixup fires there is no printk at all so I was a =
bit unsure.
quoted
Don't like this fixup though, is there not a better way than faking a r=
ead to user
quoted
space(or kernel for that matter) ?=20 I don't have a better idea. Without the fixup, the offending load instru=
ction will never finish if there is anything wrong with the backing device = and freeze the whole system. Do you have any suggestion in mind?
=20
But it never finishes the load, it just fakes a load of 0xfffffffff, for us= er space I rather have it signal a SIGBUS but that does not seem to work either, at least not for us but tha= t could be a bug in general MC code maybe. This fixup might be valid for kernel only as it has never worked for user s= pace due to the bug I found. Where can I read about this errata ? Jocke