Thread (20 messages) flat view 20 messages, 5 authors, 2017-09-21

Re: Machine Check in P2010(e500v2)

From: Joakim Tjernlund <hidden>
Date: 2017-09-08 09:54:39

On Thu, 2017-09-07 at 18:54 +0000, Leo Li wrote:
quoted
-----Original Message-----
From: Joakim Tjernlund [mailto:Joakim.Tjernlund@infinera.com]
Sent: Thursday, September 07, 2017 3:41 AM
To: linuxppc-dev@lists.ozlabs.org; Leo Li <redacted>; York Su=
n
quoted
[off-list ref]
Subject: Re: Machine Check in P2010(e500v2)
=20
On Thu, 2017-09-07 at 00:50 +0200, Joakim Tjernlund wrote:
quoted
On Wed, 2017-09-06 at 21:13 +0000, Leo Li wrote:
quoted
quoted
-----Original Message-----
From: Joakim Tjernlund [mailto:Joakim.Tjernlund@infinera.com]
Sent: Wednesday, September 06, 2017 3:54 PM
To: linuxppc-dev@lists.ozlabs.org; Leo Li <redacted>;
York Sun [off-list ref]
Subject: Re: Machine Check in P2010(e500v2)
=20
On Wed, 2017-09-06 at 20:28 +0000, Leo Li wrote:
quoted
quoted
-----Original Message-----
From: Joakim Tjernlund [mailto:Joakim.Tjernlund@infinera.com]
Sent: Wednesday, September 06, 2017 3:17 PM
To: linuxppc-dev@lists.ozlabs.org; Leo Li
[off-list ref]; York Sun [off-list ref]
Subject: Re: Machine Check in P2010(e500v2)
=20
On Wed, 2017-09-06 at 19:31 +0000, Leo Li wrote:
quoted
quoted
-----Original Message-----
From: York Sun
Sent: Wednesday, September 06, 2017 10:38 AM
To: Joakim Tjernlund <redacted>;
linuxppc- dev@lists.ozlabs.org; Leo Li
[off-list ref]
Subject: Re: Machine Check in P2010(e500v2)
=20
Scott is no longer with Freescale/NXP. Adding Leo.
=20
On 09/05/2017 01:40 AM, Joakim Tjernlund wrote:
quoted
So after some debugging I found this bug:
@@ -996,7 +998,7 @@ int fsl_pci_mcheck_exception(struct
pt_regs
=20
*regs)
quoted
quoted
quoted
quoted
quoted
         if (is_in_pci_mem_space(addr)) {
                 if (user_mode(regs)) {
                         pagefault_disable();
-                       ret =3D get_user(regs->nip, &in=
st);
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
+                       ret =3D get_user(inst, (__u32
+ __user *)regs->nip);
                         pagefault_enable();
                 } else {
                         ret =3D
probe_kernel_address(regs->nip, inst);
=20
However, the kernel still locked up after fixing that.
Now I wonder why this fixup is there in the first place=
?
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
The routine will not really fixup the insn, just return
0xffffffff for the failing read and then advance the pr=
ocess NIP.
quoted
quoted
quoted
quoted
quoted
quoted
quoted
=20
You are right.  The code here only gives 0xffffffff to the
load instructions and
=20
continue with the next instruction when the load instruction
is causing the machine check.  This will prevent a system
lockup when reading from PCI/RapidIO device which is link dow=
n.
quoted
quoted
quoted
quoted
quoted
quoted
quoted
=20
I don't know what is actual problem in your case.  Maybe it
is a write
=20
instruction instead of read?   Or the code is in a infinite l=
oop waiting for
quoted
=20
a
quoted
quoted
quoted
=20
valid
quoted
quoted
read result?  Are you able to do some further debugging with
the NIP correctly printed?
quoted
=20
=20
According to the MC it is a Read and the NIP also leads to a
read in the
=20
program.
quoted
quoted
ATM, I have disabled the fixup but I will enable that again.
Question, is it safe add a small printk when this MC
happens(after fixing up)? I need to see that it has happened
as the error is somewhat
=20
random.
quoted
=20
I think it is safe to add printk as the current machine check
handlers are also
=20
using printk.
=20
I hope so, but if the fixup fires there is no printk at all so I =
was a bit unsure.
quoted
quoted
quoted
quoted
Don't like this fixup though, is there not a better way than
faking a read to user space(or kernel for that matter) ?
=20
I don't have a better idea.  Without the fixup, the offending load =
instruction
quoted
=20
will never finish if there is anything wrong with the backing device an=
d freeze the
quoted
whole system.  Do you have any suggestion in mind?
quoted
quoted
=20
=20
But it never finishes the load, it just fakes a load of 0xfffffffff,
for user space I rather have it signal a SIGBUS but that does not see=
m
quoted
quoted
to work either, at least not for us but that could be a bug in genera=
l MC code
quoted
=20
maybe.
quoted
This fixup might be valid for kernel only as it has never worked for =
user space
quoted
=20
due to the bug I found.
quoted
=20
Where can I read about this errata ?
=20
I have look high and low an cannot find an errata which maps to this fi=
xup.
quoted
The closest I get is A-005125 which seems to have another workaround, I=
 cannot
quoted
find any evidence that this workaround has been applied in Linux, can y=
ou?
=20
This is not A-005125.  There was an erratum for this issue with older sil=
icons (e.g. erratum PCI-ex 3 for MPC8572). =20
" When its link goes down, the PCI Express controller clears all outstand=
ing transactions with an
error indicator and sends a link down exception to the interrupt controll=
er if
PEX_PME_MES_DISR[LDDD] =3D 0. If, however, any transactions are sent to t=
he controller after
the link down event, they are accepted by the controller and wait for the=
 link to come back up
before starting any timeout counters (for example, completion timeout). T=
here is no mechanism to
cancel the new transactions short of a device HRESET. "

But it was removed in newer silicon like P2020/P2010 probably because a M=
achine Check will be triggered in this situation to deal with the stalled i=
nstruction and no longer considered it as a hardware issue.
=20
Maybe this fixup should be configurable then?
The A-005125 is dealt with in u-boot.   https://lists.denx.de/pipermail/u=
-boot/2013-August/161185.html

Yes, I found it eventually :)

However, I cannot return to normal execution. I can follow the code to retu=
rning from
machine_check_exception() and moving into ASM handler for returning from a =
ME but then I
am a bit lost. It does not seem to be any problem executing, it feels more =
like a SW bug
dealing with machine checks. Don't known how to diagnose this further and c=
ould use some pointers.

 Jocke=
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help