From: Mahesh Salgaonkar <redacted>
There are chances that multiple CPUs can call crash_fadump() simultaneously
and would start duplicating same info to vmcoreinfo ELF note section. This
causes makedumpfile to fail during kdump capture. One example is,
triggering dumprestart from HMC which sends system reset to all the CPUs at
once.
makedumpfile --dump-dmesg /proc/vmcore
read_vmcoreinfo_basic_info: Invalid data in /tmp/vmcoreinfoyjgxlL: CRASHTIME=1475605971CRASHTIME=1475605971CRASHTIME=1475605971CRASHTIME=1475605971CRASHTIME=1475605971CRASHTIME=1475605971CRASHTIME=1475605971CRASHTIME=1475605971
makedumpfile Failed.
Running makedumpfile --dump-dmesg /proc/vmcore failed (1).
makedumpfile -d 31 -l /proc/vmcore
read_vmcoreinfo_basic_info: Invalid data in /tmp/vmcoreinfo1mmVdO: CRASHTIME=1475605971CRASHTIME=1475605971CRASHTIME=1475605971CRASHTIME=1475605971CRASHTIME=1475605971CRASHTIME=1475605971CRASHTIME=1475605971CRASHTIME=1475605971
makedumpfile Failed.
Running makedumpfile -d 31 -l /proc/vmcore failed (1).
Signed-off-by: Mahesh Salgaonkar <redacted>
---
arch/powerpc/kernel/fadump.c | 11 ++++++++++-
1 file changed, 10 insertions(+), 1 deletion(-)
From: Michael Ellerman <mpe@ellerman.id.au> Date: 2016-10-10 10:52:11
Mahesh J Salgaonkar [off-list ref] writes:
From: Mahesh Salgaonkar <redacted>
There are chances that multiple CPUs can call crash_fadump() simultaneously
and would start duplicating same info to vmcoreinfo ELF note section. This
causes makedumpfile to fail during kdump capture. One example is,
triggering dumprestart from HMC which sends system reset to all the CPUs at
once.
From: Mahesh Salgaonkar <redacted>
There are chances that multiple CPUs can call crash_fadump() simultaneously
and would start duplicating same info to vmcoreinfo ELF note section. This
causes makedumpfile to fail during kdump capture. One example is,
triggering dumprestart from HMC which sends system reset to all the CPUs at
once.
What happens when a crashing CPU can't get the mutex and goes to sleep?
Got your point. I think I should use mutex_trylock() here. There is only
two reason crashing CPU can't get mutex, 1) Another CPU also crashing
that got the mutex and on its way to trigger fadump. OR 2) We are in
middle of fadump register/un-register, in which case we can just return
and go to normal panic.
Thanks,
-Mahesh.
From: Mahesh Salgaonkar <redacted>
There are chances that multiple CPUs can call crash_fadump() simultaneously
and would start duplicating same info to vmcoreinfo ELF note section. This
causes makedumpfile to fail during kdump capture. One example is,
triggering dumprestart from HMC which sends system reset to all the CPUs at
once.
What happens when a crashing CPU can't get the mutex and goes to sleep?
Got your point. I think I should use mutex_trylock() here. There is only
two reason crashing CPU can't get mutex, 1) Another CPU also crashing
that got the mutex and on its way to trigger fadump. OR 2) We are in
middle of fadump register/un-register, in which case we can just return
and go to normal panic.
mutex_trylock() is probably OK, though with DEBUG options on it might
still run a bunch of code, you don't want it to go and call into lockdep
for example.
Normally the advice is not to write your own locks, but in this case the
safest option might be just to do a cmpxchg() or something similarly
simple.
cheers
On 13/10/16 04:48, Mahesh Jagannath Salgaonkar wrote:
On 10/10/2016 04:22 PM, Michael Ellerman wrote:
quoted
Mahesh J Salgaonkar [off-list ref] writes:
quoted
From: Mahesh Salgaonkar <redacted>
There are chances that multiple CPUs can call crash_fadump() simultaneously
and would start duplicating same info to vmcoreinfo ELF note section. This
causes makedumpfile to fail during kdump capture. One example is,
triggering dumprestart from HMC which sends system reset to all the CPUs at
once.
What happens when a crashing CPU can't get the mutex and goes to sleep?
Got your point. I think I should use mutex_trylock() here. There is only
two reason crashing CPU can't get mutex, 1) Another CPU also crashing
that got the mutex and on its way to trigger fadump. OR 2) We are in
middle of fadump register/un-register, in which case we can just return
and go to normal panic.
I think trylock is a good idea, but having said that not getting the lock
and having the CPU active will still lead to the same issue.
I don't quite know the source of failure in makedumpfile but
should we fix makedumpfile to deal better with these issues?
Another option is to check to see if anyone started writing at the ELF
note section and have others bail out if they get there after the try
lock
Balbir Singh.