Thread (20 messages) flat view 20 messages, 5 authors, 2017-09-21

Re: Machine Check in P2010(e500v2)

From: Joakim Tjernlund <hidden>
Date: 2017-09-06 20:53:52

On Wed, 2017-09-06 at 20:28 +0000, Leo Li wrote:
quoted
-----Original Message-----
From: Joakim Tjernlund [mailto:Joakim.Tjernlund@infinera.com]
Sent: Wednesday, September 06, 2017 3:17 PM
To: linuxppc-dev@lists.ozlabs.org; Leo Li <redacted>; York Su=
n
quoted
[off-list ref]
Subject: Re: Machine Check in P2010(e500v2)
=20
On Wed, 2017-09-06 at 19:31 +0000, Leo Li wrote:
quoted
quoted
-----Original Message-----
From: York Sun
Sent: Wednesday, September 06, 2017 10:38 AM
To: Joakim Tjernlund <redacted>; linuxppc-
dev@lists.ozlabs.org; Leo Li [off-list ref]
Subject: Re: Machine Check in P2010(e500v2)
=20
Scott is no longer with Freescale/NXP. Adding Leo.
=20
On 09/05/2017 01:40 AM, Joakim Tjernlund wrote:
quoted
So after some debugging I found this bug:
@@ -996,7 +998,7 @@ int fsl_pci_mcheck_exception(struct pt_regs *=
regs)
quoted
quoted
quoted
quoted
         if (is_in_pci_mem_space(addr)) {
                 if (user_mode(regs)) {
                         pagefault_disable();
-                       ret =3D get_user(regs->nip, &inst);
+                       ret =3D get_user(inst, (__u32 __user
+ *)regs->nip);
                         pagefault_enable();
                 } else {
                         ret =3D probe_kernel_address(regs->nip,
inst);
=20
However, the kernel still locked up after fixing that.
Now I wonder why this fixup is there in the first place? The
routine will not really fixup the insn, just return 0xffffffff fo=
r
quoted
quoted
quoted
quoted
the failing read and then advance the process NIP.
=20
You are right.  The code here only gives 0xffffffff to the load instr=
uctions and
quoted
=20
continue with the next instruction when the load instruction is causing=
 the
quoted
machine check.  This will prevent a system lockup when reading from
PCI/RapidIO device which is link down.
quoted
=20
I don't know what is actual problem in your case.  Maybe it is a writ=
e
quoted
=20
instruction instead of read?   Or the code is in a infinite loop waitin=
g for a valid
quoted
read result?  Are you able to do some further debugging with the NIP co=
rrectly
quoted
printed?
quoted
=20
=20
According to the MC it is a Read and the NIP also leads to a read in th=
e program.
quoted
ATM, I have disabled the fixup but I will enable that again.
Question, is it safe add a small printk when this MC happens(after fixi=
ng up)? I
quoted
need to see that it has happened as the error is somewhat random.
=20
I think it is safe to add printk as the current machine check handlers ar=
e also using printk.

I hope so, but if the fixup fires there is no printk at all so I was a bit =
unsure.
Don't like this fixup though, is there not a better way than faking a read
to user space(or kernel for that matter) ?

 Jocke=
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help