Thread (20 messages) flat view 20 messages, 5 authors, 2017-09-21

Re: Machine Check in P2010(e500v2)

From: Joakim Tjernlund <hidden>
Date: 2017-09-09 12:45:51

On Fri, 2017-09-08 at 22:27 +0000, Leo Li wrote:
quoted
-----Original Message-----
From: Joakim Tjernlund [mailto:Joakim.Tjernlund@infinera.com]
Sent: Friday, September 08, 2017 7:51 AM
To: linuxppc-dev@lists.ozlabs.org; Leo Li <redacted>; York Su=
n
quoted
[off-list ref]
Subject: Re: Machine Check in P2010(e500v2)
=20
On Fri, 2017-09-08 at 11:54 +0200, Joakim Tjernlund wrote:
quoted
On Thu, 2017-09-07 at 18:54 +0000, Leo Li wrote:
quoted
quoted
-----Original Message-----
From: Joakim Tjernlund [mailto:Joakim.Tjernlund@infinera.com]
Sent: Thursday, September 07, 2017 3:41 AM
To: linuxppc-dev@lists.ozlabs.org; Leo Li <redacted>;
York Sun [off-list ref]
Subject: Re: Machine Check in P2010(e500v2)
=20
On Thu, 2017-09-07 at 00:50 +0200, Joakim Tjernlund wrote:
quoted
On Wed, 2017-09-06 at 21:13 +0000, Leo Li wrote:
quoted
quoted
-----Original Message-----
From: Joakim Tjernlund
[mailto:Joakim.Tjernlund@infinera.com]
Sent: Wednesday, September 06, 2017 3:54 PM
To: linuxppc-dev@lists.ozlabs.org; Leo Li
[off-list ref]; York Sun [off-list ref]
Subject: Re: Machine Check in P2010(e500v2)
=20
On Wed, 2017-09-06 at 20:28 +0000, Leo Li wrote:
quoted
quoted
-----Original Message-----
From: Joakim Tjernlund
[mailto:Joakim.Tjernlund@infinera.com]
Sent: Wednesday, September 06, 2017 3:17 PM
To: linuxppc-dev@lists.ozlabs.org; Leo Li
[off-list ref]; York Sun [off-list ref]
Subject: Re: Machine Check in P2010(e500v2)
=20
On Wed, 2017-09-06 at 19:31 +0000, Leo Li wrote:
quoted
quoted
-----Original Message-----
From: York Sun
Sent: Wednesday, September 06, 2017 10:38 AM
To: Joakim Tjernlund
[off-list ref];
linuxppc- dev@lists.ozlabs.org; Leo Li
[off-list ref]
Subject: Re: Machine Check in P2010(e500v2)
=20
Scott is no longer with Freescale/NXP. Adding Leo.
=20
On 09/05/2017 01:40 AM, Joakim Tjernlund wrote:
quoted
So after some debugging I found this bug:
@@ -996,7 +998,7 @@ int
fsl_pci_mcheck_exception(struct pt_regs
=20
*regs)
quoted
quoted
quoted
quoted
quoted
         if (is_in_pci_mem_space(addr)) {
                 if (user_mode(regs)) {
                         pagefault_disable();
-                       ret =3D get_user(regs->ni=
p, &inst);
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
+                       ret =3D get_user(inst,
+ (__u32 __user *)regs->nip);
                         pagefault_enable();
                 } else {
                         ret =3D
probe_kernel_address(regs->nip, inst);
=20
However, the kernel still locked up after fixing =
that.
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
Now I wonder why this fixup is there in the first=
 place?
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
The routine will not really fixup the insn, just
return 0xffffffff for the failing read and then a=
dvance the
quoted
=20
process NIP.
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
=20
You are right.  The code here only gives 0xffffffff t=
o
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
the load instructions and
=20
continue with the next instruction when the load
instruction is causing the machine check.  This will
prevent a system lockup when reading from PCI/RapidIO d=
evice
quoted
=20
which is link down.
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
=20
I don't know what is actual problem in your case.
Maybe it is a write
=20
instruction instead of read?   Or the code is in a infi=
nite loop
quoted
=20
waiting for
quoted
quoted
quoted
=20
a
quoted
quoted
quoted
=20
valid
quoted
quoted
read result?  Are you able to do some further debugging
with the NIP correctly printed?
quoted
=20
=20
According to the MC it is a Read and the NIP also leads
to a read in the
=20
program.
quoted
quoted
ATM, I have disabled the fixup but I will enable that a=
gain.
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
Question, is it safe add a small printk when this MC
happens(after fixing up)? I need to see that it has
happened as the error is somewhat
=20
random.
quoted
=20
I think it is safe to add printk as the current machine
check handlers are also
=20
using printk.
=20
I hope so, but if the fixup fires there is no printk at all=
 so I was a bit
quoted
=20
unsure.
quoted
quoted
quoted
quoted
quoted
quoted
Don't like this fixup though, is there not a better way tha=
n
quoted
quoted
quoted
quoted
quoted
quoted
quoted
faking a read to user space(or kernel for that matter) ?
=20
I don't have a better idea.  Without the fixup, the offending
load instruction
=20
will never finish if there is anything wrong with the backing
device and freeze the whole system.  Do you have any suggestion i=
n mind?
quoted
quoted
quoted
quoted
quoted
quoted
=20
=20
But it never finishes the load, it just fakes a load of
0xfffffffff, for user space I rather have it signal a SIGBUS bu=
t
quoted
quoted
quoted
quoted
quoted
that does not seem to work either, at least not for us but that
could be a bug in general MC code
=20
maybe.
quoted
This fixup might be valid for kernel only as it has never worke=
d
quoted
quoted
quoted
quoted
quoted
for user space
=20
due to the bug I found.
quoted
=20
Where can I read about this errata ?
=20
I have look high and low an cannot find an errata which maps to t=
his fixup.
quoted
quoted
quoted
quoted
The closest I get is A-005125 which seems to have another
workaround, I cannot find any evidence that this workaround has b=
een
quoted
=20
applied in Linux, can you?
quoted
quoted
=20
This is not A-005125.  There was an erratum for this issue with old=
er silicons
quoted
=20
(e.g. erratum PCI-ex 3 for MPC8572).
quoted
quoted
" When its link goes down, the PCI Express controller clears all
outstanding transactions with an error indicator and sends a link
down exception to the interrupt controller if PEX_PME_MES_DISR[LDDD=
]
quoted
quoted
quoted
=3D 0. If, however, any transactions are sent to the controller aft=
er
quoted
quoted
quoted
the link down event, they are accepted by the controller and wait
for the link to come back up before starting any timeout counters (=
for
quoted
=20
example, completion timeout). There is no mechanism to cancel the new
transactions short of a device HRESET. "
quoted
quoted
=20
But it was removed in newer silicon like P2020/P2010 probably becau=
se a
quoted
=20
Machine Check will be triggered in this situation to deal with the stal=
led
quoted
instruction and no longer considered it as a hardware issue.
quoted
quoted
=20
=20
Maybe this fixup should be configurable then?
=20
No.  My point is that the problem was no longer considered a hardware iss=
ue because of the machine check mechanism is in place to handle it.  If the=
re is no handling of this special case, we would still experience a system =
hang if this situation really occurs.
=20
quoted
quoted
=20
quoted
The A-005125 is dealt with in u-boot.
=20
https://emea01.safelinks.protection.outlook.com/?url=3Dhttps%3A%2F%2Fli=
sts.de
quoted
nx.de%2Fpipermail%2Fu-boot%2F2013-
August%2F161185.html&data=3D01%7C01%7Cleoyang.li%40nxp.com%7Ccb8a93e
0090e48eb53a008d4f6b84235%7C686ea1d3bc2b4c6fa92cd99c5c301635%7C0&
sdata=3D8sR4yoXA4adqMHz6TY%2BvmYpfCBTcYEZHjPuANjz%2F1EQ%3D&reserve
d=3D0
quoted
=20
Yes, I found it eventually :)
=20
However, I cannot return to normal execution. I can follow the code t=
o
quoted
quoted
returning from
machine_check_exception() and moving into ASM handler for returning
from a ME but then I am a bit lost. It does not seem to be any proble=
m
quoted
quoted
executing, it feels more like a SW bug dealing with machine checks. D=
on't
quoted
=20
known how to diagnose this further and could use some pointers.
=20
Is the execution returned to the user application?  I doubt the system ha=
ng is caused by the machine check handling.
You can try to comment out the machine check handling code and check if t=
here is any improvement and see if
this is related to the machine check handling.
It tries to return to user app but I cannot see what happens as the system =
lock up when the
MC returns.
How do you mean comment out MC handling? The simplest path is the PCI fixup=
 which will
just do regs->nip +=3D 4; and then return to user space. That still does no=
t work as
as soon MC handling returns, the system is locked up.
=20
Machine check is a serious situation and not always possible to be recove=
red from.=20

This one should at least not kill the whole system. It is a simple bus erro=
r in user space and
the app should get SIGBUS and the the system should carry on.=20
I would focus more on debugging why the machine check is triggered by the=
 user space application.
Can you locate what code is causing this machine check from user space? =
=20
Is it accessing some hardware related space which is not ready?=20
Or is it accessing address that it shouldn't have accessed?
of course, this is ongoing and getting closer a solution. The MC looking th=
e machine completely
does not make this any easier though.
These are 2 separate things, fixing the cause and not having a simple bus e=
rror lock up the machine.
I am focusing on fixing the lockup.

I have been following the execution in the kernel and I always end up in th=
e ASM returning
from the MC.
The other day we got a similar PCI MC(bus error) on T1042 CPU(e5500/e500mc)=
 and there
the system survived. The one thing I see different there is that MSR RI is =
set
when entering MC, why is that?

 Jocke=
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help