This patch series adds kdump support on arm64.
To load a crash-dump kernel to the systems, a series of patches to
kexec-tools, which have not yet been merged upstream, are needed.
Please use my kdump patches [1].
To examine vmcore (/proc/vmcore) on a crash-dump kernel, you can use
- crash utility (coming v7.1.6 or later) [2]
(Necessary patches have already been queued in the master.)
[1] T.B.D.
[2] https://github.com/crash-utility/crash.git
Changes for v24 (Aug 9, 2016):
o Rebase to Linux-4.8-rc1
o Update descriptions about newly added DT proerties
Changes for v23 (July 26, 2016):
o Move memblock_reserve() to a single place in reserve_crashkernel()
o Use cpu_park_loop() in ipi_cpu_crash_stop()
o Always enforce ARCH_LOW_ADDRESS_LIMIT to the memory range of crash kernel
o Re-implement fdt_enforce_memory_region() to remove non-reserve regions
(for ACPI) from usable memory at crash kernel
Changes for v22 (July 12, 2016):
o Export "crashkernel-base" and "crashkernel-size" via device-tree,
and add some descriptions about them in chosen.txt
o Rename "usable-memory" to "usable-memory-range" to avoid inconsistency
with powerpc's "usable-memory"
o Make cosmetic changes regarding "ifdef" usage
o Correct some wordings in kdump.txt
Changes for v21 (July 6, 2016):
o Remove kexec patches.
o Rebase to arm64's for-next/core (Linux-4.7-rc4 based).
o Clarify the description about kvm in kdump.txt.
See the following link [3] for older changes:
[3] http://lists.infradead.org/pipermail/linux-arm-kernel/2016-June/438780.html
AKASHI Takahiro (8):
arm64: kdump: reserve memory for crash dump kernel
memblock: add memblock_cap_memory_range()
arm64: limit memory regions based on DT property, usable-memory-range
arm64: kdump: implement machine_crash_shutdown()
arm64: kdump: add kdump support
arm64: kdump: add VMCOREINFO's for user-space coredump tools
arm64: kdump: enable kdump in the arm64 defconfig
arm64: kdump: update a kernel doc
James Morse (1):
Documentation: dt: chosen properties for arm64 kdump
Documentation/devicetree/bindings/chosen.txt | 45 ++++++
Documentation/kdump/kdump.txt | 16 ++-
arch/arm64/Kconfig | 11 ++
arch/arm64/configs/defconfig | 1 +
arch/arm64/include/asm/hardirq.h | 2 +-
arch/arm64/include/asm/kexec.h | 41 +++++-
arch/arm64/include/asm/smp.h | 2 +
arch/arm64/kernel/Makefile | 1 +
arch/arm64/kernel/crash_dump.c | 71 ++++++++++
arch/arm64/kernel/machine_kexec.c | 67 ++++++++-
arch/arm64/kernel/setup.c | 7 +-
arch/arm64/kernel/smp.c | 63 +++++++++
arch/arm64/mm/init.c | 202 +++++++++++++++++++++++++++
include/linux/memblock.h | 1 +
mm/memblock.c | 28 ++++
15 files changed, 551 insertions(+), 7 deletions(-)
create mode 100644 arch/arm64/kernel/crash_dump.c
--
2.9.0
AKASHI Takahiro (8):
arm64: kdump: reserve memory for crash dump kernel
memblock: add memblock_cap_memory_range()
arm64: limit memory regions based on DT property, usable-memory-range
arm64: kdump: implement machine_crash_shutdown()
arm64: kdump: add kdump support
arm64: kdump: add VMCOREINFO's for user-space coredump tools
arm64: kdump: enable kdump in the arm64 defconfig
arm64: kdump: update a kernel doc
James Morse (1):
Documentation: dt: chosen properties for arm64 kdump
Documentation/devicetree/bindings/chosen.txt | 50 +++++++
Documentation/kdump/kdump.txt | 16 ++-
arch/arm64/Kconfig | 11 ++
arch/arm64/configs/defconfig | 1 +
arch/arm64/include/asm/hardirq.h | 2 +-
arch/arm64/include/asm/kexec.h | 41 +++++-
arch/arm64/include/asm/smp.h | 2 +
arch/arm64/kernel/Makefile | 1 +
arch/arm64/kernel/crash_dump.c | 71 ++++++++++
arch/arm64/kernel/machine_kexec.c | 67 ++++++++-
arch/arm64/kernel/setup.c | 7 +-
arch/arm64/kernel/smp.c | 63 +++++++++
arch/arm64/mm/init.c | 202 +++++++++++++++++++++++++++
include/linux/memblock.h | 1 +
mm/memblock.c | 28 ++++
15 files changed, 556 insertions(+), 7 deletions(-)
create mode 100644 arch/arm64/kernel/crash_dump.c
--
2.9.0
Crash dump kernel uses only a limited range of memory as System RAM.
On arm64 implementation, a new device tree property,
"linux,usable-memory-range," is used to notify crash dump kernel of
this range.[1]
But simply excluding all the other regions, whatever their memory types
are, doesn't work, especially, on the systems with ACPI. Since some of
such regions will be later mapped as "device memory" by ioremap()/
acpi_os_ioremap(), it can cause errors like unalignment accesses.[2]
This issue is akin to the one reported in [3].
So this patch follows Chen's approach, and implements a new function,
memblock_cap_memory_range(), which will exclude only the memory regions
that are not marked "NOMAP" from memblock.memory.
[1] http://lists.infradead.org/pipermail/linux-arm-kernel/2016-July/442817.html
[2] http://lists.infradead.org/pipermail/linux-arm-kernel/2016-July/444165.html
[3] http://lists.infradead.org/pipermail/linux-arm-kernel/2016-July/443356.html
Signed-off-by: AKASHI Takahiro <redacted>
---
include/linux/memblock.h | 1 +
mm/memblock.c | 28 ++++++++++++++++++++++++++++
2 files changed, 29 insertions(+)
@@ -1539,6 +1539,34 @@ void __init memblock_mem_limit_remove_map(phys_addr_t limit)(phys_addr_t)ULLONG_MAX);}+void__initmemblock_cap_memory_range(phys_addr_tbase,phys_addr_tsize)+{+intstart_rgn,end_rgn;+inti,ret;++if(!size)+return;++ret=memblock_isolate_range(&memblock.memory,base,size,+&start_rgn,&end_rgn);+if(ret)+return;++/* remove all the MAP regions */+for(i=memblock.memory.cnt-1;i>=end_rgn;i--)+if(!memblock_is_nomap(&memblock.memory.regions[i]))+memblock_remove_region(&memblock.memory,i);++for(i=start_rgn-1;i>=0;i--)+if(!memblock_is_nomap(&memblock.memory.regions[i]))+memblock_remove_region(&memblock.memory,i);++/* truncate the reserved regions */+memblock_remove_range(&memblock.reserved,0,base);+memblock_remove_range(&memblock.reserved,+base+size,(phys_addr_t)ULLONG_MAX);+}+staticint__init_memblockmemblock_search(structmemblock_type*type,phys_addr_taddr){unsignedintleft=0,right=type->cnt;
On the startup of primary kernel, the memory region used by crash dump
kernel must be specified by "crashkernel=" kernel parameter.
reserve_crashkernel() will allocate and reserve the region for later use.
User space tools, like kexec-tools, will be able to find that region as
- "Crash kernel" in /proc/iomem, or
- "linux,crashkernel-base" and "linux,crashkernel-size" under
/sys/firmware/devicetree/base/chosen
Signed-off-by: AKASHI Takahiro <redacted>
Signed-off-by: Mark Salter <redacted>
Signed-off-by: Pratyush Anand <redacted>
Reviewed-by: James Morse <james.morse@arm.com>
---
arch/arm64/kernel/setup.c | 7 ++-
arch/arm64/mm/init.c | 113 ++++++++++++++++++++++++++++++++++++++++++++++
2 files changed, 119 insertions(+), 1 deletion(-)
@@ -220,6 +219,12 @@ static void __init request_standard_resources(void)kernel_data.end<=res->end)request_resource(res,&kernel_data);}++#ifdef CONFIG_KEXEC_CORE+/* User space tools will find "Crash kernel" region in /proc/iomem. */+if(crashk_res.end)+insert_resource(&iomem_resource,&crashk_res);+#endif}u64__cpu_logical_map[NR_CPUS]={[0...NR_CPUS-1]=INVALID_HWID};
@@ -76,6 +78,114 @@ static int __init early_initrd(char *p)early_param("initrd",early_initrd);#endif+#ifdef CONFIG_KEXEC_CORE+staticunsignedlonglongcrash_size,crash_base;+staticstructpropertycrash_base_prop={+.name="linux,crashkernel-base",+.length=sizeof(u64),+.value=&crash_base+};+staticstructpropertycrash_size_prop={+.name="linux,crashkernel-size",+.length=sizeof(u64),+.value=&crash_size,+};++staticint__initexport_crashkernel(void)+{+structdevice_node*node;+intret;++if(!crashk_res.end)+return0;++crash_base=cpu_to_be64(crashk_res.start);+crash_size=cpu_to_be64(crashk_res.end-crashk_res.start+1);++/* Add /chosen/linux,crashkernel-* properties */+node=of_find_node_by_path("/chosen");+if(!node)+return-ENOENT;++/*+*Theremightbeexistingcrashkernelproperties,butwecan't+*besurewhat'sinthem,soremovethem.+*/+of_remove_property(node,of_find_property(node,+"linux,crashkernel-base",NULL));+of_remove_property(node,of_find_property(node,+"linux,crashkernel-size",NULL));++ret=of_add_property(node,&crash_base_prop);+if(ret)+gotoret_err;++ret=of_add_property(node,&crash_size_prop);+if(ret)+gotoret_err;++return0;++ret_err:+pr_warn("Exporting crashkernel region to device tree failed\n");+returnret;+}+late_initcall(export_crashkernel);++/*+*reserve_crashkernel()-reservesmemoryforcrashkernel+*+*Thisfunctionreservesmemoryareagivenin"crashkernel="kernelcommand+*lineparameter.Thememoryreservedisusedbydumpcapturekernelwhen+*primarykerneliscrashing.+*/+staticvoid__initreserve_crashkernel(void)+{+intret;++ret=parse_crashkernel(boot_command_line,memblock_phys_mem_size(),+&crash_size,&crash_base);+/* no crashkernel= or invalid value specified */+if(ret||!crash_size)+return;++if(crash_base==0){+/* Current arm64 boot protocol requires 2MB alignment */+crash_base=memblock_find_in_range(0,ARCH_LOW_ADDRESS_LIMIT,+crash_size,SZ_2M);+if(crash_base==0){+pr_warn("Unable to allocate crashkernel (size:%llx)\n",+crash_size);+return;+}+}else{+/* User specifies base address explicitly. */+if(!memblock_is_region_memory(crash_base,crash_size)||+memblock_is_region_reserved(crash_base,crash_size)){+pr_warn("crashkernel has wrong address or size\n");+return;+}++if(!IS_ALIGNED(crash_base,SZ_2M)){+pr_warn("crashkernel base address is not 2MB aligned\n");+return;+}+}+memblock_reserve(crash_base,crash_size);++pr_info("Reserving %lldMB of memory at %lldMB for crashkernel\n",+crash_size>>20,crash_base>>20);++crashk_res.start=crash_base;+crashk_res.end=crash_base+crash_size-1;+}+#else+staticvoid__initreserve_crashkernel(void)+{+;+}+#endif /* CONFIG_KEXEC_CORE */+/**ReturnthemaximumphysicaladdressforZONE_DMA(DMA_BIT_MASK(32)).It*currentlyassumesthatformemorystartingabove4G,32-bitdeviceswill
Crash dump kernel will be run with a limited range of memory as System
RAM.
On arm64, we will use a device-tree property under /chosen,
linux,usable-memory-range = <BASE SIZE>
in order for primary kernel either on uefi or non-uefi (device tree only)
system to hand over the information about usable memory region to crash
dump kernel. This property will supercede entries in uefi memory map table
and "memory" nodes in a device tree.
Signed-off-by: AKASHI Takahiro <redacted>
Reviewed-by: Geoff Levand <geoff@infradead.org>
---
arch/arm64/mm/init.c | 35 +++++++++++++++++++++++++++++++++++
1 file changed, 35 insertions(+)
Primary kernel calls machine_crash_shutdown() to shut down non-boot cpus
and save registers' status in per-cpu ELF notes before starting crash
dump kernel. See kernel_kexec().
Even if not all secondary cpus have shut down, we do kdump anyway.
As we don't have to make non-boot(crashed) cpus offline (to preserve
correct status of cpus at crash dump) before shutting down, this patch
also adds a variant of smp_send_stop().
Signed-off-by: AKASHI Takahiro <redacted>
---
arch/arm64/include/asm/hardirq.h | 2 +-
arch/arm64/include/asm/kexec.h | 41 ++++++++++++++++++++++++-
arch/arm64/include/asm/smp.h | 2 ++
arch/arm64/kernel/machine_kexec.c | 56 ++++++++++++++++++++++++++++++++--
arch/arm64/kernel/smp.c | 63 +++++++++++++++++++++++++++++++++++++++
5 files changed, 159 insertions(+), 5 deletions(-)
@@ -808,6 +811,29 @@ static void ipi_cpu_stop(unsigned int cpu)cpu_relax();}+#ifdef CONFIG_KEXEC_CORE+staticatomic_twaiting_for_crash_ipi;+#endif++staticvoidipi_cpu_crash_stop(unsignedintcpu,structpt_regs*regs)+{+#ifdef CONFIG_KEXEC_CORE+crash_save_cpu(regs,cpu);++atomic_dec(&waiting_for_crash_ipi);++local_irq_disable();++#ifdef CONFIG_HOTPLUG_CPU+if(cpu_ops[cpu]->cpu_die)+cpu_ops[cpu]->cpu_die(cpu);+#endif++/* just in case */+cpu_park_loop();+#endif+}+/**Mainhandlerforinter-processorinterrupts*/
@@ -910,6 +945,34 @@ void smp_send_stop(void)cpumask_pr_args(cpu_online_mask));}+#ifdef CONFIG_KEXEC_CORE+voidsmp_send_crash_stop(void)+{+cpumask_tmask;+unsignedlongtimeout;++if(num_online_cpus()==1)+return;++cpumask_copy(&mask,cpu_online_mask);+cpumask_clear_cpu(smp_processor_id(),&mask);++atomic_set(&waiting_for_crash_ipi,num_online_cpus()-1);++pr_crit("SMP: stopping secondary CPUs\n");+smp_cross_call(&mask,IPI_CPU_CRASH_STOP);++/* Wait up to one second for other CPUs to stop */+timeout=USEC_PER_SEC;+while((atomic_read(&waiting_for_crash_ipi)>0)&&timeout--)+udelay(1);++if(atomic_read(&waiting_for_crash_ipi)>0)+pr_warning("SMP: failed to stop secondary CPUs %*pbl\n",+cpumask_pr_args(cpu_online_mask));+}+#endif+/**notsupportedhere*/
On crash dump kernel, all the information about primary kernel's system
memory (core image) is available in elf core header.
The primary kernel will set aside this header with reserve_elfcorehdr()
at boot time and inform crash dump kernel of its location via a new
device-tree property, "linux,elfcorehdr".
Please note that all other architectures use traditional "elfcorehdr="
kernel parameter for this purpose.
Then crash dump kernel will access the primary kernel's memory with
copy_oldmem_page(), which reads one page by ioremap'ing it since it does
not reside in linear mapping on crash dump kernel.
We also need our own elfcorehdr_read() here since the header is placed
within crash dump kernel's usable memory.
Signed-off-by: AKASHI Takahiro <redacted>
---
arch/arm64/Kconfig | 11 +++++++
arch/arm64/kernel/Makefile | 1 +
arch/arm64/kernel/crash_dump.c | 71 ++++++++++++++++++++++++++++++++++++++++++
arch/arm64/mm/init.c | 54 ++++++++++++++++++++++++++++++++
4 files changed, 137 insertions(+)
create mode 100644 arch/arm64/kernel/crash_dump.c
For the current crash utility, we need to know, at least,
- kimage_voffset
- PHYS_OFFSET
to handle the contents of core dump file (/proc/vmcore) correctly due to
the introduction of KASLR (CONFIG_RANDOMIZE_BASE) in v4.6.
This patch puts them as VMCOREINFO's into the file.
- VA_BITS
is also added for makedumpfile command.
More VMCOREINFO's may be added later.
Signed-off-by: AKASHI Takahiro <redacted>
---
arch/arm64/kernel/machine_kexec.c | 11 +++++++++++
1 file changed, 11 insertions(+)
@@ -73,6 +73,7 @@ CONFIG_TRANSPARENT_HUGEPAGE=y CONFIG_CMA=y CONFIG_XEN=y CONFIG_KEXEC=y+CONFIG_CRASH_DUMP=y # CONFIG_CORE_DUMP_DEFAULT_ELF_HEADERS is not set CONFIG_COMPAT=y CONFIG_CPU_IDLE=y
This patch adds arch specific descriptions about kdump usage on arm64
to kdump.txt.
Signed-off-by: AKASHI Takahiro <redacted>
Reviewed-by: Baoquan He <redacted>
Acked-by: Dave Young <redacted>
---
Documentation/kdump/kdump.txt | 16 +++++++++++++++-
1 file changed, 15 insertions(+), 1 deletion(-)
@@ -18,7 +18,7 @@ memory image to a dump file on the local disk, or across the network to a remote system. Kdump and kexec are currently supported on the x86, x86_64, ppc64, ia64,-s390x and arm architectures.+s390x, arm and arm64 architectures. When the system kernel boots, it reserves a small section of memory for the dump-capture kernel. This ensures that ongoing Direct Memory Access
@@ -249,6 +249,13 @@ Dump-capture kernel config options (Arch Dependent, arm) AUTO_ZRELADDR=y+Dump-capture kernel config options (Arch Dependent, arm64)+----------------------------------------------------------++- Please note that kvm of the dump-capture kernel will not be enabled+ on non-VHE systems even if it is configured. This is because the CPU+ cannot be reset to EL2 on panic.+ Extended crashkernel syntax ===========================
@@ -305,6 +312,8 @@ Boot into System Kernel kernel will automatically locate the crash kernel image within the first 512MB of RAM if X is not given.+ On arm64, use "crashkernel=Y[@X]". Note that the start address of+ the kernel, X if explicitly specified, must be aligned to 2MiB (0x200000). Load the Dump-capture Kernel ============================
@@ -327,6 +336,8 @@ For s390x: - Use image or bzImage For arm: - Use zImage+For arm64:+ - Use vmlinux or Image If you are using a uncompressed vmlinux image then use following command to load dump-capture kernel.
@@ -370,6 +381,9 @@ For s390x: For arm: "1 maxcpus=1 reset_devices"+For arm64:+ "1 maxcpus=1 reset_devices"+ Notes on loading the dump-capture kernel: * By default, the ELF headers are stored in ELF64 format to support
On Tue, Aug 09, 2016 at 10:52:47AM +0900, AKASHI Takahiro wrote:
This patch series adds kdump support on arm64.
To load a crash-dump kernel to the systems, a series of patches to
kexec-tools, which have not yet been merged upstream, are needed.
Please use my kdump patches [1].
To examine vmcore (/proc/vmcore) on a crash-dump kernel, you can use
- crash utility (coming v7.1.6 or later) [2]
(Necessary patches have already been queued in the master.)
[1] T.B.D.
From: James Morse <james.morse@arm.com> Date: 2016-08-10 16:26:49
Hi Akashi,
On 09/08/16 02:55, AKASHI Takahiro wrote:
Crash dump kernel uses only a limited range of memory as System RAM.
On arm64 implementation, a new device tree property,
"linux,usable-memory-range," is used to notify crash dump kernel of
this range.[1]
But simply excluding all the other regions, whatever their memory types
are, doesn't work, especially, on the systems with ACPI. Since some of
such regions will be later mapped as "device memory" by ioremap()/
acpi_os_ioremap(), it can cause errors like unalignment accesses.[2]
This issue is akin to the one reported in [3].
So this patch follows Chen's approach, and implements a new function,
memblock_cap_memory_range(), which will exclude only the memory regions
that are not marked "NOMAP" from memblock.memory.
This (and the next patch) fixes the acpi related unaligned access problem I had.
I've tested it on a Juno r1 and Seattle B0.
Tested-by: James Morse <james.morse@arm.com>
Thanks,
James
Hi Akashi,
On 09/08/16 02:56, AKASHI Takahiro wrote:
On crash dump kernel, all the information about primary kernel's system
memory (core image) is available in elf core header.
The primary kernel will set aside this header with reserve_elfcorehdr()
at boot time and inform crash dump kernel of its location via a new
device-tree property, "linux,elfcorehdr".
Please note that all other architectures use traditional "elfcorehdr="
kernel parameter for this purpose.
Then crash dump kernel will access the primary kernel's memory with
copy_oldmem_page(), which reads one page by ioremap'ing it since it does
not reside in linear mapping on crash dump kernel.
We also need our own elfcorehdr_read() here since the header is placed
within crash dump kernel's usable memory.
On Seattle when I panic and boot the kdump kernel, I am unable to read the
/proc/vmcore file. Instead I get:
nanook at frikadeller:~$ sudo cp /proc/vmcore /
[ 174.393875] Unhandled fault: synchronous external abort (0x96000210) at
0xffffff80096b6000
[ 174.402158] Internal error: : 96000210 [#1] PREEMPT SMP
[ 174.407370] Modules linked in:
[ 174.410417] CPU: 6 PID: 2059 Comm: cp Tainted: G S W I 4.8.0-rc1+ #4708
[ 174.417799] Hardware name: AMD Overdrive/Supercharger/Default string, BIOS
ROD1002C 04/08/2016
[ 174.426396] task: ffffffc0fdec5780 task.stack: ffffffc0f34bc000
[ 174.432313] PC is at __arch_copy_to_user+0x180/0x280
[ 174.437274] LR is at copy_oldmem_page+0xac/0xf0
[ 174.441791] pc : [<ffffff800835e080>] lr : [<ffffff8008095b9c>] pstate: 20000145
[ 174.449173] sp : ffffffc0f34bfc90
[ 174.452474] x29: ffffffc0f34bfc90 x28: 0000000000000000
[ 174.457776] x27: 0000000008000000 x26: 000000000000d000
[ 174.463077] x25: 0000000000000001 x24: ffffff8008eb5000
[ 174.468378] x23: 0000000000000000 x22: ffffff80096b6000
[ 174.473679] x21: 0000000000000001 x20: 0000000030127000
[ 174.478979] x19: 0000000000001000 x18: 0000007ff7085d60
[ 174.484279] x17: 0000000000429358 x16: ffffff80081d9e88
[ 174.489579] x15: 0000007fae377590 x14: 0000000000000000
[ 174.494880] x13: 0000000000000000 x12: ffffff8008dd1000
[ 174.500180] x11: ffffff80096b6fff x10: ffffff80096b6fff
[ 174.505480] x9 : 0000000040000000 x8 : ffffff8008db6000
[ 174.510781] x7 : ffffff80096b7000 x6 : 0000000030127000
[ 174.516082] x5 : 0000000030128000 x4 : 0000000000000000
[ 174.521382] x3 : 00e8000000000713 x2 : 0000000000000f80
[ 174.526682] x1 : ffffff80096b6000 x0 : 0000000030127000
[ 174.531982]
[ 174.533461] Process cp (pid: 2059, stack limit = 0xffffffc0f34bc020)
[ 174.848448] [<ffffff800835e080>] __arch_copy_to_user+0x180/0x280
[ 174.854448] [<ffffff8008245f34>] read_from_oldmem.part.4+0xb4/0xf4
[ 174.860615] [<ffffff8008246074>] read_vmcore+0x100/0x22c
[ 174.865919] [<ffffff8008239378>] proc_reg_read+0x64/0x90
[ 174.871223] [<ffffff80081d7da8>] __vfs_read+0x28/0x108
[ 174.876348] [<ffffff80081d8ae4>] vfs_read+0x84/0x144
[ 174.881301] [<ffffff80081d9ecc>] SyS_read+0x44/0xa0
[ 174.886167] [<ffffff8008082ef0>] el0_svc_naked+0x24/0x28
[ 174.891466] Code: 00000000 00000000 00000000 00000000 (a8c12027)
[ 174.897562] ---[ end trace 00801b2e35b0cd1f ]---
The offending call is:
This page is 'Runtime Data', and marked as nomap by both the original and kdump
kernels, but copy_oldmem_page() doesn't know this.
In this case because we have already parsed the efi memory map again in the
kdump kernel and re-marked these regions as nomap, the below hunk fixes the
problem for me:
=========================%<=========================
@@ -37,6 +37,11 @@ ssize_t copy_oldmem_page(unsigned long pfn, char *buf,if(!csize)return0;+if(memblock_is_memory(pfn<<PAGE_SHIFT)&&+!memblock_is_map_memory(pfn<<PAGE_SHIFT))+/* skip this nomap memory region, reserved by firmware */+return0;+vaddr=ioremap_cache(__pfn_to_phys(pfn),PAGE_SIZE);if(!vaddr)return-ENOMEM;
=========================%<=========================
With this I can copy the vmcore file, and feed it to crash to read dmesg, task
list etc...
This could be a deeper/wider issue, but I can't see any other users of
memblock_mark_nomap().
Do you think depending on this this 're-learning' is robust enough, or should
the nomap ranges be described in the vmcoreinfo elf notes?
Thanks,
James
Hi Akashi,
On 09/08/16 02:56, AKASHI Takahiro wrote:
quoted
On crash dump kernel, all the information about primary kernel's system
memory (core image) is available in elf core header.
The primary kernel will set aside this header with reserve_elfcorehdr()
at boot time and inform crash dump kernel of its location via a new
device-tree property, "linux,elfcorehdr".
Please note that all other architectures use traditional "elfcorehdr="
kernel parameter for this purpose.
Then crash dump kernel will access the primary kernel's memory with
copy_oldmem_page(), which reads one page by ioremap'ing it since it does
not reside in linear mapping on crash dump kernel.
We also need our own elfcorehdr_read() here since the header is placed
within crash dump kernel's usable memory.
On Seattle when I panic and boot the kdump kernel, I am unable to read the
/proc/vmcore file. Instead I get:
nanook at frikadeller:~$ sudo cp /proc/vmcore /
[ 174.393875] Unhandled fault: synchronous external abort (0x96000210) at
0xffffff80096b6000
Yes, I see the same while executing vmcore-dmesg or copying vmcore.
This page is 'Runtime Data', and marked as nomap by both the original and kdump
kernels, but copy_oldmem_page() doesn't know this.
In this case because we have already parsed the efi memory map again in the
kdump kernel and re-marked these regions as nomap, the below hunk fixes the
problem for me:
=========================%<=========================
In any case kernel must not panic, so I think we must have above hunk. However,
we also need to look into kexec-tools that why it is asking kernel to copy those
unneeded chunks.
I will test tomorrow with above hunk.
~Pratyush
With this I can copy the vmcore file, and feed it to crash to read dmesg, task
list etc...
This could be a deeper/wider issue, but I can't see any other users of
memblock_mark_nomap().
Do you think depending on this this 're-learning' is robust enough, or should
the nomap ranges be described in the vmcoreinfo elf notes?
@@ -37,6 +37,11 @@ ssize_t copy_oldmem_page(unsigned long pfn, char *buf,if(!csize)return0;+if(memblock_is_memory(pfn<<PAGE_SHIFT)&&+!memblock_is_map_memory(pfn<<PAGE_SHIFT))+/* skip this nomap memory region, reserved by firmware */+return0;
This should return 0 or -EINVAL? because, its caller does not care properly
about 0 return value (when csize is non-zero). So either we need to return
-EINVAL or we need to fix it's caller so that pread() would know that required
number of data were not read.
quoted
+
vaddr = ioremap_cache(__pfn_to_phys(pfn), PAGE_SIZE);
if (!vaddr)
return -ENOMEM;
=========================%<=========================
In any case kernel must not panic, so I think we must have above hunk. However,
we also need to look into kexec-tools that why it is asking kernel to copy those
unneeded chunks.
I will test tomorrow with above hunk.
After that hunk it did not crash but vmcore-dmesg fails with following message:
"No program header covering vaddr 0x401ff0found kexec bug?"
It happened because vmcore-dmesg is sending wrong offset to the pread(), and so
it did not crash after the above kernel hunk but it still read garbage wrong
log_buf virtual address pointer.
vmcore-dmesg is sending wrong offset because page_offset(vp_offset) calculation
is not perfect for my case, explained here [1].
So, if I correct page_offset(vp_offset) (as arm64_mem.page_offset = ehdr.e_entry
- "kernel Code Start PA" + phys_offset), then vmcore-dmesg and vmcore copy
worked fine, however if I use makedumpfile to copy(compressed) data from
/proc/vmcore then it still generates "synchronous external abort". I think, it
generated because it would have found garbage data in EFI memory region. My
/proc/iomem shows following:
8000000000-8001e7ffff : System RAM
8001e80000-83ff17ffff : System RAM
8002080000-8002b3ffff : Kernel code
8002c40000-800348ffff : Kernel data
807fe00000-80ffdfffff : Crash kernel
83ff180000-83ff1cffff : System RAM
83ff1d0000-83ff21ffff : System RAM
83ff220000-83ffe4ffff : System RAM
83ffe50000-83ffffffff : System RAM
If I clip all the region before "kernel code" and provide that clipped
input to kexec-tools then everything works fine.
~Pratyush
[1] http://lists.infradead.org/pipermail/kexec/2016-August/016834.html
@@ -37,6 +37,11 @@ ssize_t copy_oldmem_page(unsigned long pfn, char *buf,if(!csize)return0;+if(memblock_is_memory(pfn<<PAGE_SHIFT)&&+!memblock_is_map_memory(pfn<<PAGE_SHIFT))+/* skip this nomap memory region, reserved by firmware */+return0;
This should return 0 or -EINVAL? because, its caller does not care properly
about 0 return value (when csize is non-zero). So either we need to return
-EINVAL or we need to fix it's caller so that pread() would know that required
number of data were not read.
I blindly followed 'number of bytes copied' -> 0. It worked for me, but may not
be correct.
remap_oldmem_pfn_checked() looks like it substitutes the zero page in this (or
at least a similar) case, maybe we should do the same for nomap pages.
quoted
quoted
+
vaddr = ioremap_cache(__pfn_to_phys(pfn), PAGE_SIZE);
if (!vaddr)
return -ENOMEM;
=========================%<=========================
In any case kernel must not panic, so I think we must have above hunk. However,
we also need to look into kexec-tools that why it is asking kernel to copy those
unneeded chunks.
I will test tomorrow with above hunk.
After that hunk it did not crash but vmcore-dmesg fails with following message:
"No program header covering vaddr 0x401ff0found kexec bug?"
It happened because vmcore-dmesg is sending wrong offset to the pread(), and so
it did not crash after the above kernel hunk but it still read garbage wrong
log_buf virtual address pointer.
vmcore-dmesg is sending wrong offset because page_offset(vp_offset) calculation
is not perfect for my case, explained here [1].
So, if I correct page_offset(vp_offset) (as arm64_mem.page_offset = ehdr.e_entry
- "kernel Code Start PA" + phys_offset), then vmcore-dmesg and vmcore copy
worked fine, however if I use makedumpfile to copy(compressed) data from
/proc/vmcore then it still generates "synchronous external abort". I think, it
At a guess makedumpfile is mmap()ing /proc/vmcore so it can use multiple
threads to read (then compress) the data. This bypasses the check added to
copy_oldmem_page(). We probably need to provide a remap_oldmem_pfn_range() that
checks whether the range contains nomap pages.
I will try and send a fixup patch to do this later this week, (unless someone
beats me to it!)
generated because it would have found garbage data in EFI memory region.
If it was marked as belonging to efi in the efi memory map, the kernel shouldn't
be touching it. If you add 'efi=debug' to your kernel cmdline you get a table of
the addresses and properties.
Thanks,
James
copy_oldmem_page() and mmap_vmcore() provide two ways for userspace to read
from /proc/vmcore. Neither of these check with memblock to see if the page
they are accessing is nomap. On Seattle this causes:
[ 174.393875] Unhandled fault: synchronous external abort (0x96000210) at
0xffffff80096b6000
[ 174.402158] Internal error: : 96000210 [#1] PREEMPT SMP
[ 174.407370] Modules linked in:
[ 174.410417] CPU: 6 PID: 2059 Comm: cp Tainted: G S W I 4.8.0-rc1+ #4708
[ 174.417799] Hardware name: AMD Overdrive/Supercharger/Default string, BIOS
ROD1002C 04/08/2016
[ 174.426396] task: ffffffc0fdec5780 task.stack: ffffffc0f34bc000
[ 174.432313] PC is at __arch_copy_to_user+0x180/0x280
[ 174.437274] LR is at copy_oldmem_page+0xac/0xf0
[ 174.441791] pc : [<ffffff800835e080>] lr : [<ffffff8008095b9c>] pstate: 20000145
[ 174.449173] sp : ffffffc0f34bfc90
[ 174.452474] x29: ffffffc0f34bfc90 x28: 0000000000000000
[ 174.457776] x27: 0000000008000000 x26: 000000000000d000
[ 174.463077] x25: 0000000000000001 x24: ffffff8008eb5000
[ 174.468378] x23: 0000000000000000 x22: ffffff80096b6000
[ 174.473679] x21: 0000000000000001 x20: 0000000030127000
[ 174.478979] x19: 0000000000001000 x18: 0000007ff7085d60
[ 174.484279] x17: 0000000000429358 x16: ffffff80081d9e88
[ 174.489579] x15: 0000007fae377590 x14: 0000000000000000
[ 174.494880] x13: 0000000000000000 x12: ffffff8008dd1000
[ 174.500180] x11: ffffff80096b6fff x10: ffffff80096b6fff
[ 174.505480] x9 : 0000000040000000 x8 : ffffff8008db6000
[ 174.510781] x7 : ffffff80096b7000 x6 : 0000000030127000
[ 174.516082] x5 : 0000000030128000 x4 : 0000000000000000
[ 174.521382] x3 : 00e8000000000713 x2 : 0000000000000f80
[ 174.526682] x1 : ffffff80096b6000 x0 : 0000000030127000
[ 174.531982]
[ 174.533461] Process cp (pid: 2059, stack limit = 0xffffffc0f34bc020)
[ 174.848448] [<ffffff800835e080>] __arch_copy_to_user+0x180/0x280
[ 174.854448] [<ffffff8008245f34>] read_from_oldmem.part.4+0xb4/0xf4
[ 174.860615] [<ffffff8008246074>] read_vmcore+0x100/0x22c
[ 174.865919] [<ffffff8008239378>] proc_reg_read+0x64/0x90
[ 174.871223] [<ffffff80081d7da8>] __vfs_read+0x28/0x108
[ 174.876348] [<ffffff80081d8ae4>] vfs_read+0x84/0x144
[ 174.881301] [<ffffff80081d9ecc>] SyS_read+0x44/0xa0
[ 174.886167] [<ffffff8008082ef0>] el0_svc_naked+0x24/0x28
[ 174.891466] Code: 00000000 00000000 00000000 00000000 (a8c12027)
[ 174.897562] ---[ end trace 00801b2e35b0cd1f ]---
When reading /proc/vmcore with cat/cp or or mmap()ing it with makedumpfile.
The fs/proc/vmcore.c code provides a hook to indicate whether oldmem pages
are ram or not. Use this to look for our earlier handiwork in memblock.
Signed-off-by: James Morse <james.morse@arm.com>
---
Hi Pratyush,
I couldn't get makedumpfile to build, or rather it depends on elfutils which
wouldn't build for autotools reasons. Does implementing this hook solve your
makedumpfile issue?
With this patch I can extract a usable vmcore file using read or mmap,
avoiding the earlier splat.
Akashi, if you agree this is the right thing to do, please consider folding
this into patch 5. (no need to keep the commit mesage or anything).
Thanks,
James
arch/arm64/kernel/crash_dump.c | 28 ++++++++++++++++++++++++++++
1 file changed, 28 insertions(+)
Hi James, Pratyush,
Thank you for your testing and reporting an issue.
I've been on vacation until yesterday.
On Wed, Aug 10, 2016 at 05:38:05PM +0100, James Morse wrote:
quoted hunk
Hi Akashi,
On 09/08/16 02:56, AKASHI Takahiro wrote:
quoted
On crash dump kernel, all the information about primary kernel's system
memory (core image) is available in elf core header.
The primary kernel will set aside this header with reserve_elfcorehdr()
at boot time and inform crash dump kernel of its location via a new
device-tree property, "linux,elfcorehdr".
Please note that all other architectures use traditional "elfcorehdr="
kernel parameter for this purpose.
Then crash dump kernel will access the primary kernel's memory with
copy_oldmem_page(), which reads one page by ioremap'ing it since it does
not reside in linear mapping on crash dump kernel.
We also need our own elfcorehdr_read() here since the header is placed
within crash dump kernel's usable memory.
On Seattle when I panic and boot the kdump kernel, I am unable to read the
/proc/vmcore file. Instead I get:
nanook at frikadeller:~$ sudo cp /proc/vmcore /
[ 174.393875] Unhandled fault: synchronous external abort (0x96000210) at
0xffffff80096b6000
[ 174.402158] Internal error: : 96000210 [#1] PREEMPT SMP
[ 174.407370] Modules linked in:
[ 174.410417] CPU: 6 PID: 2059 Comm: cp Tainted: G S W I 4.8.0-rc1+ #4708
[ 174.417799] Hardware name: AMD Overdrive/Supercharger/Default string, BIOS
ROD1002C 04/08/2016
[ 174.426396] task: ffffffc0fdec5780 task.stack: ffffffc0f34bc000
[ 174.432313] PC is at __arch_copy_to_user+0x180/0x280
[ 174.437274] LR is at copy_oldmem_page+0xac/0xf0
[ 174.441791] pc : [<ffffff800835e080>] lr : [<ffffff8008095b9c>] pstate: 20000145
[ 174.449173] sp : ffffffc0f34bfc90
[ 174.452474] x29: ffffffc0f34bfc90 x28: 0000000000000000
[ 174.457776] x27: 0000000008000000 x26: 000000000000d000
[ 174.463077] x25: 0000000000000001 x24: ffffff8008eb5000
[ 174.468378] x23: 0000000000000000 x22: ffffff80096b6000
[ 174.473679] x21: 0000000000000001 x20: 0000000030127000
[ 174.478979] x19: 0000000000001000 x18: 0000007ff7085d60
[ 174.484279] x17: 0000000000429358 x16: ffffff80081d9e88
[ 174.489579] x15: 0000007fae377590 x14: 0000000000000000
[ 174.494880] x13: 0000000000000000 x12: ffffff8008dd1000
[ 174.500180] x11: ffffff80096b6fff x10: ffffff80096b6fff
[ 174.505480] x9 : 0000000040000000 x8 : ffffff8008db6000
[ 174.510781] x7 : ffffff80096b7000 x6 : 0000000030127000
[ 174.516082] x5 : 0000000030128000 x4 : 0000000000000000
[ 174.521382] x3 : 00e8000000000713 x2 : 0000000000000f80
[ 174.526682] x1 : ffffff80096b6000 x0 : 0000000030127000
[ 174.531982]
[ 174.533461] Process cp (pid: 2059, stack limit = 0xffffffc0f34bc020)
[ 174.848448] [<ffffff800835e080>] __arch_copy_to_user+0x180/0x280
[ 174.854448] [<ffffff8008245f34>] read_from_oldmem.part.4+0xb4/0xf4
[ 174.860615] [<ffffff8008246074>] read_vmcore+0x100/0x22c
[ 174.865919] [<ffffff8008239378>] proc_reg_read+0x64/0x90
[ 174.871223] [<ffffff80081d7da8>] __vfs_read+0x28/0x108
[ 174.876348] [<ffffff80081d8ae4>] vfs_read+0x84/0x144
[ 174.881301] [<ffffff80081d9ecc>] SyS_read+0x44/0xa0
[ 174.886167] [<ffffff8008082ef0>] el0_svc_naked+0x24/0x28
[ 174.891466] Code: 00000000 00000000 00000000 00000000 (a8c12027)
[ 174.897562] ---[ end trace 00801b2e35b0cd1f ]---
The offending call is:
This page is 'Runtime Data', and marked as nomap by both the original and kdump
kernels, but copy_oldmem_page() doesn't know this.
In this case because we have already parsed the efi memory map again in the
kdump kernel and re-marked these regions as nomap, the below hunk fixes the
problem for me:
=========================%<=========================
@@ -37,6 +37,11 @@ ssize_t copy_oldmem_page(unsigned long pfn, char *buf,if(!csize)return0;+if(memblock_is_memory(pfn<<PAGE_SHIFT)&&+!memblock_is_map_memory(pfn<<PAGE_SHIFT))+/* skip this nomap memory region, reserved by firmware */+return0;+vaddr=ioremap_cache(__pfn_to_phys(pfn),PAGE_SIZE);
Here I'm wandering why my original code doesn't work.
If !memblock_is_map_memory(), ioremap_cache() would call __ioremap_caller()
and return a valid virtual address mapped in vmalloc area.
if (!vaddr)
return -ENOMEM;
=========================%<=========================
With this I can copy the vmcore file, and feed it to crash to read dmesg, task
list etc...
This could be a deeper/wider issue, but I can't see any other users of
memblock_mark_nomap().
Do you think depending on this this 're-learning' is robust enough, or should
the nomap ranges be described in the vmcoreinfo elf notes?
The current kexec-tools identifies all the memory regions from
/proc/iomem and there is no way for user space tools to distinguish
"EFI runtime data," or any other nomap memory, from normal "System RAM"
because all those resources are currently marked as "System RAM."
So I think that such regions should be marked as, say, "reserved,"
so that we can exclude those memories from a crush dump file.
(I don't know whether this change may have a backward-compatibility
problem.)
-Takahiro AKASHI
From: Dave Young <hidden> Date: 2016-08-18 07:19:54
On 08/18/16 at 04:15pm, AKASHI Takahiro wrote:
Hi James, Pratyush,
Thank you for your testing and reporting an issue.
I've been on vacation until yesterday.
On Wed, Aug 10, 2016 at 05:38:05PM +0100, James Morse wrote:
quoted
Hi Akashi,
On 09/08/16 02:56, AKASHI Takahiro wrote:
quoted
On crash dump kernel, all the information about primary kernel's system
memory (core image) is available in elf core header.
The primary kernel will set aside this header with reserve_elfcorehdr()
at boot time and inform crash dump kernel of its location via a new
device-tree property, "linux,elfcorehdr".
Please note that all other architectures use traditional "elfcorehdr="
kernel parameter for this purpose.
Then crash dump kernel will access the primary kernel's memory with
copy_oldmem_page(), which reads one page by ioremap'ing it since it does
not reside in linear mapping on crash dump kernel.
We also need our own elfcorehdr_read() here since the header is placed
within crash dump kernel's usable memory.
On Seattle when I panic and boot the kdump kernel, I am unable to read the
/proc/vmcore file. Instead I get:
nanook at frikadeller:~$ sudo cp /proc/vmcore /
[ 174.393875] Unhandled fault: synchronous external abort (0x96000210) at
0xffffff80096b6000
[ 174.402158] Internal error: : 96000210 [#1] PREEMPT SMP
[ 174.407370] Modules linked in:
[ 174.410417] CPU: 6 PID: 2059 Comm: cp Tainted: G S W I 4.8.0-rc1+ #4708
[ 174.417799] Hardware name: AMD Overdrive/Supercharger/Default string, BIOS
ROD1002C 04/08/2016
[ 174.426396] task: ffffffc0fdec5780 task.stack: ffffffc0f34bc000
[ 174.432313] PC is at __arch_copy_to_user+0x180/0x280
[ 174.437274] LR is at copy_oldmem_page+0xac/0xf0
[ 174.441791] pc : [<ffffff800835e080>] lr : [<ffffff8008095b9c>] pstate: 20000145
[ 174.449173] sp : ffffffc0f34bfc90
[ 174.452474] x29: ffffffc0f34bfc90 x28: 0000000000000000
[ 174.457776] x27: 0000000008000000 x26: 000000000000d000
[ 174.463077] x25: 0000000000000001 x24: ffffff8008eb5000
[ 174.468378] x23: 0000000000000000 x22: ffffff80096b6000
[ 174.473679] x21: 0000000000000001 x20: 0000000030127000
[ 174.478979] x19: 0000000000001000 x18: 0000007ff7085d60
[ 174.484279] x17: 0000000000429358 x16: ffffff80081d9e88
[ 174.489579] x15: 0000007fae377590 x14: 0000000000000000
[ 174.494880] x13: 0000000000000000 x12: ffffff8008dd1000
[ 174.500180] x11: ffffff80096b6fff x10: ffffff80096b6fff
[ 174.505480] x9 : 0000000040000000 x8 : ffffff8008db6000
[ 174.510781] x7 : ffffff80096b7000 x6 : 0000000030127000
[ 174.516082] x5 : 0000000030128000 x4 : 0000000000000000
[ 174.521382] x3 : 00e8000000000713 x2 : 0000000000000f80
[ 174.526682] x1 : ffffff80096b6000 x0 : 0000000030127000
[ 174.531982]
[ 174.533461] Process cp (pid: 2059, stack limit = 0xffffffc0f34bc020)
[ 174.848448] [<ffffff800835e080>] __arch_copy_to_user+0x180/0x280
[ 174.854448] [<ffffff8008245f34>] read_from_oldmem.part.4+0xb4/0xf4
[ 174.860615] [<ffffff8008246074>] read_vmcore+0x100/0x22c
[ 174.865919] [<ffffff8008239378>] proc_reg_read+0x64/0x90
[ 174.871223] [<ffffff80081d7da8>] __vfs_read+0x28/0x108
[ 174.876348] [<ffffff80081d8ae4>] vfs_read+0x84/0x144
[ 174.881301] [<ffffff80081d9ecc>] SyS_read+0x44/0xa0
[ 174.886167] [<ffffff8008082ef0>] el0_svc_naked+0x24/0x28
[ 174.891466] Code: 00000000 00000000 00000000 00000000 (a8c12027)
[ 174.897562] ---[ end trace 00801b2e35b0cd1f ]---
The offending call is:
This page is 'Runtime Data', and marked as nomap by both the original and kdump
kernels, but copy_oldmem_page() doesn't know this.
In this case because we have already parsed the efi memory map again in the
kdump kernel and re-marked these regions as nomap, the below hunk fixes the
problem for me:
=========================%<=========================
@@ -37,6 +37,11 @@ ssize_t copy_oldmem_page(unsigned long pfn, char *buf,if(!csize)return0;+if(memblock_is_memory(pfn<<PAGE_SHIFT)&&+!memblock_is_map_memory(pfn<<PAGE_SHIFT))+/* skip this nomap memory region, reserved by firmware */+return0;+vaddr=ioremap_cache(__pfn_to_phys(pfn),PAGE_SIZE);
Here I'm wandering why my original code doesn't work.
If !memblock_is_map_memory(), ioremap_cache() would call __ioremap_caller()
and return a valid virtual address mapped in vmalloc area.
quoted
if (!vaddr)
return -ENOMEM;
=========================%<=========================
With this I can copy the vmcore file, and feed it to crash to read dmesg, task
list etc...
This could be a deeper/wider issue, but I can't see any other users of
memblock_mark_nomap().
Do you think depending on this this 're-learning' is robust enough, or should
the nomap ranges be described in the vmcoreinfo elf notes?
The current kexec-tools identifies all the memory regions from
/proc/iomem and there is no way for user space tools to distinguish
"EFI runtime data," or any other nomap memory, from normal "System RAM"
because all those resources are currently marked as "System RAM."
So I think that such regions should be marked as, say, "reserved,"
so that we can exclude those memories from a crush dump file.
Agreed.
EFI runtime memory is not system ram, in X86 they are "Reserved" ranges,
it sounds a better way to mark them ask reserved as well in arm64.
(I don't know whether this change may have a backward-compatibility
problem.)
-Takahiro AKASHI
On Wed, Aug 17, 2016 at 04:33:31PM +0100, James Morse wrote:
copy_oldmem_page() and mmap_vmcore() provide two ways for userspace to read
from /proc/vmcore. Neither of these check with memblock to see if the page
they are accessing is nomap. On Seattle this causes:
[ 174.393875] Unhandled fault: synchronous external abort (0x96000210) at
0xffffff80096b6000
[ 174.402158] Internal error: : 96000210 [#1] PREEMPT SMP
[ 174.407370] Modules linked in:
[ 174.410417] CPU: 6 PID: 2059 Comm: cp Tainted: G S W I 4.8.0-rc1+ #4708
[ 174.417799] Hardware name: AMD Overdrive/Supercharger/Default string, BIOS
ROD1002C 04/08/2016
[ 174.426396] task: ffffffc0fdec5780 task.stack: ffffffc0f34bc000
[ 174.432313] PC is at __arch_copy_to_user+0x180/0x280
[ 174.437274] LR is at copy_oldmem_page+0xac/0xf0
[ 174.441791] pc : [<ffffff800835e080>] lr : [<ffffff8008095b9c>] pstate: 20000145
[ 174.449173] sp : ffffffc0f34bfc90
[ 174.452474] x29: ffffffc0f34bfc90 x28: 0000000000000000
[ 174.457776] x27: 0000000008000000 x26: 000000000000d000
[ 174.463077] x25: 0000000000000001 x24: ffffff8008eb5000
[ 174.468378] x23: 0000000000000000 x22: ffffff80096b6000
[ 174.473679] x21: 0000000000000001 x20: 0000000030127000
[ 174.478979] x19: 0000000000001000 x18: 0000007ff7085d60
[ 174.484279] x17: 0000000000429358 x16: ffffff80081d9e88
[ 174.489579] x15: 0000007fae377590 x14: 0000000000000000
[ 174.494880] x13: 0000000000000000 x12: ffffff8008dd1000
[ 174.500180] x11: ffffff80096b6fff x10: ffffff80096b6fff
[ 174.505480] x9 : 0000000040000000 x8 : ffffff8008db6000
[ 174.510781] x7 : ffffff80096b7000 x6 : 0000000030127000
[ 174.516082] x5 : 0000000030128000 x4 : 0000000000000000
[ 174.521382] x3 : 00e8000000000713 x2 : 0000000000000f80
[ 174.526682] x1 : ffffff80096b6000 x0 : 0000000030127000
[ 174.531982]
[ 174.533461] Process cp (pid: 2059, stack limit = 0xffffffc0f34bc020)
[ 174.848448] [<ffffff800835e080>] __arch_copy_to_user+0x180/0x280
[ 174.854448] [<ffffff8008245f34>] read_from_oldmem.part.4+0xb4/0xf4
[ 174.860615] [<ffffff8008246074>] read_vmcore+0x100/0x22c
[ 174.865919] [<ffffff8008239378>] proc_reg_read+0x64/0x90
[ 174.871223] [<ffffff80081d7da8>] __vfs_read+0x28/0x108
[ 174.876348] [<ffffff80081d8ae4>] vfs_read+0x84/0x144
[ 174.881301] [<ffffff80081d9ecc>] SyS_read+0x44/0xa0
[ 174.886167] [<ffffff8008082ef0>] el0_svc_naked+0x24/0x28
[ 174.891466] Code: 00000000 00000000 00000000 00000000 (a8c12027)
[ 174.897562] ---[ end trace 00801b2e35b0cd1f ]---
When reading /proc/vmcore with cat/cp or or mmap()ing it with makedumpfile.
The fs/proc/vmcore.c code provides a hook to indicate whether oldmem pages
are ram or not. Use this to look for our earlier handiwork in memblock.
I'm not quite sure about the background that oldmem_pfn_is_ram() was
originally introduced on x86, but I think that this feature be deserved
for fixing an issue on Xen.
See:
commit 997c136
Author: Olaf Hering [off-list ref]
Date: Thu May 26 16:25:54 2011 -0700
fs/proc/vmcore.c: add hook to read_from_oldmem() to check for non-ram pages
Thanks,
-Takahiro AKASHI
quoted hunk
Signed-off-by: James Morse <james.morse@arm.com>
---
Hi Pratyush,
I couldn't get makedumpfile to build, or rather it depends on elfutils which
wouldn't build for autotools reasons. Does implementing this hook solve your
makedumpfile issue?
With this patch I can extract a usable vmcore file using read or mmap,
avoiding the earlier splat.
Akashi, if you agree this is the right thing to do, please consider folding
this into patch 5. (no need to keep the commit mesage or anything).
Thanks,
James
arch/arm64/kernel/crash_dump.c | 28 ++++++++++++++++++++++++++++
1 file changed, 28 insertions(+)
James,
On Thu, Aug 18, 2016 at 04:15:48PM +0900, AKASHI Takahiro wrote:
Hi James, Pratyush,
Thank you for your testing and reporting an issue.
I've been on vacation until yesterday.
On Wed, Aug 10, 2016 at 05:38:05PM +0100, James Morse wrote:
quoted
Hi Akashi,
On 09/08/16 02:56, AKASHI Takahiro wrote:
quoted
On crash dump kernel, all the information about primary kernel's system
memory (core image) is available in elf core header.
The primary kernel will set aside this header with reserve_elfcorehdr()
at boot time and inform crash dump kernel of its location via a new
device-tree property, "linux,elfcorehdr".
Please note that all other architectures use traditional "elfcorehdr="
kernel parameter for this purpose.
Then crash dump kernel will access the primary kernel's memory with
copy_oldmem_page(), which reads one page by ioremap'ing it since it does
not reside in linear mapping on crash dump kernel.
We also need our own elfcorehdr_read() here since the header is placed
within crash dump kernel's usable memory.
On Seattle when I panic and boot the kdump kernel, I am unable to read the
/proc/vmcore file. Instead I get:
nanook at frikadeller:~$ sudo cp /proc/vmcore /
[ 174.393875] Unhandled fault: synchronous external abort (0x96000210) at
0xffffff80096b6000
[ 174.402158] Internal error: : 96000210 [#1] PREEMPT SMP
[ 174.407370] Modules linked in:
[ 174.410417] CPU: 6 PID: 2059 Comm: cp Tainted: G S W I 4.8.0-rc1+ #4708
[ 174.417799] Hardware name: AMD Overdrive/Supercharger/Default string, BIOS
ROD1002C 04/08/2016
[ 174.426396] task: ffffffc0fdec5780 task.stack: ffffffc0f34bc000
[ 174.432313] PC is at __arch_copy_to_user+0x180/0x280
[ 174.437274] LR is at copy_oldmem_page+0xac/0xf0
[ 174.441791] pc : [<ffffff800835e080>] lr : [<ffffff8008095b9c>] pstate: 20000145
[ 174.449173] sp : ffffffc0f34bfc90
[ 174.452474] x29: ffffffc0f34bfc90 x28: 0000000000000000
[ 174.457776] x27: 0000000008000000 x26: 000000000000d000
[ 174.463077] x25: 0000000000000001 x24: ffffff8008eb5000
[ 174.468378] x23: 0000000000000000 x22: ffffff80096b6000
[ 174.473679] x21: 0000000000000001 x20: 0000000030127000
[ 174.478979] x19: 0000000000001000 x18: 0000007ff7085d60
[ 174.484279] x17: 0000000000429358 x16: ffffff80081d9e88
[ 174.489579] x15: 0000007fae377590 x14: 0000000000000000
[ 174.494880] x13: 0000000000000000 x12: ffffff8008dd1000
[ 174.500180] x11: ffffff80096b6fff x10: ffffff80096b6fff
[ 174.505480] x9 : 0000000040000000 x8 : ffffff8008db6000
[ 174.510781] x7 : ffffff80096b7000 x6 : 0000000030127000
[ 174.516082] x5 : 0000000030128000 x4 : 0000000000000000
[ 174.521382] x3 : 00e8000000000713 x2 : 0000000000000f80
[ 174.526682] x1 : ffffff80096b6000 x0 : 0000000030127000
[ 174.531982]
[ 174.533461] Process cp (pid: 2059, stack limit = 0xffffffc0f34bc020)
[ 174.848448] [<ffffff800835e080>] __arch_copy_to_user+0x180/0x280
[ 174.854448] [<ffffff8008245f34>] read_from_oldmem.part.4+0xb4/0xf4
[ 174.860615] [<ffffff8008246074>] read_vmcore+0x100/0x22c
[ 174.865919] [<ffffff8008239378>] proc_reg_read+0x64/0x90
[ 174.871223] [<ffffff80081d7da8>] __vfs_read+0x28/0x108
[ 174.876348] [<ffffff80081d8ae4>] vfs_read+0x84/0x144
[ 174.881301] [<ffffff80081d9ecc>] SyS_read+0x44/0xa0
[ 174.886167] [<ffffff8008082ef0>] el0_svc_naked+0x24/0x28
[ 174.891466] Code: 00000000 00000000 00000000 00000000 (a8c12027)
[ 174.897562] ---[ end trace 00801b2e35b0cd1f ]---
The offending call is:
This page is 'Runtime Data', and marked as nomap by both the original and kdump
kernels, but copy_oldmem_page() doesn't know this.
In this case because we have already parsed the efi memory map again in the
kdump kernel and re-marked these regions as nomap, the below hunk fixes the
problem for me:
=========================%<=========================
@@ -37,6 +37,11 @@ ssize_t copy_oldmem_page(unsigned long pfn, char *buf,if(!csize)return0;+if(memblock_is_memory(pfn<<PAGE_SHIFT)&&+!memblock_is_map_memory(pfn<<PAGE_SHIFT))+/* skip this nomap memory region, reserved by firmware */+return0;+vaddr=ioremap_cache(__pfn_to_phys(pfn),PAGE_SIZE);
Here I'm wandering why my original code doesn't work.
If !memblock_is_map_memory(), ioremap_cache() would call __ioremap_caller()
and return a valid virtual address mapped in vmalloc area.
quoted
if (!vaddr)
return -ENOMEM;
=========================%<=========================
With this I can copy the vmcore file, and feed it to crash to read dmesg, task
list etc...
This could be a deeper/wider issue, but I can't see any other users of
memblock_mark_nomap().
Do you think depending on this this 're-learning' is robust enough, or should
the nomap ranges be described in the vmcoreinfo elf notes?
The current kexec-tools identifies all the memory regions from
/proc/iomem and there is no way for user space tools to distinguish
"EFI runtime data," or any other nomap memory, from normal "System RAM"
because all those resources are currently marked as "System RAM."
So I think that such regions should be marked as, say, "reserved,"
so that we can exclude those memories from a crush dump file.
Can you try the following change?
If it fixes your problem, I will post it as a patch.
Thanks,
-Takahiro AKASHI
===8<===
From 740563e4a437f0d6ecf6e421c91433f9b8f19041 Mon Sep 17 00:00:00 2001
Hi James,
On 17/08/2016:04:33:31 PM, James Morse wrote:
copy_oldmem_page() and mmap_vmcore() provide two ways for userspace to read
from /proc/vmcore. Neither of these check with memblock to see if the page
they are accessing is nomap. On Seattle this causes:
Thanks for the patch.It did resolve the kernel crash issue with makedumpfile,
however neither there was any data in vmcore-dmesg nor crash utility was able to
work the saved vmcore.
This happened, because we still do not have correct page_offset (or vp_offset as
per new patches) calculation in kexec-tools. I still need following fixup in
kexec-tools.
https://github.com/pratyushanand/kexec-tools/commit/2358de3ec614d8282a565b8d031a1a91ebc55475
~Pratyush
quoted hunk
[ 174.393875] Unhandled fault: synchronous external abort (0x96000210) at
0xffffff80096b6000
[ 174.402158] Internal error: : 96000210 [#1] PREEMPT SMP
[ 174.407370] Modules linked in:
[ 174.410417] CPU: 6 PID: 2059 Comm: cp Tainted: G S W I 4.8.0-rc1+ #4708
[ 174.417799] Hardware name: AMD Overdrive/Supercharger/Default string, BIOS
ROD1002C 04/08/2016
[ 174.426396] task: ffffffc0fdec5780 task.stack: ffffffc0f34bc000
[ 174.432313] PC is at __arch_copy_to_user+0x180/0x280
[ 174.437274] LR is at copy_oldmem_page+0xac/0xf0
[ 174.441791] pc : [<ffffff800835e080>] lr : [<ffffff8008095b9c>] pstate: 20000145
[ 174.449173] sp : ffffffc0f34bfc90
[ 174.452474] x29: ffffffc0f34bfc90 x28: 0000000000000000
[ 174.457776] x27: 0000000008000000 x26: 000000000000d000
[ 174.463077] x25: 0000000000000001 x24: ffffff8008eb5000
[ 174.468378] x23: 0000000000000000 x22: ffffff80096b6000
[ 174.473679] x21: 0000000000000001 x20: 0000000030127000
[ 174.478979] x19: 0000000000001000 x18: 0000007ff7085d60
[ 174.484279] x17: 0000000000429358 x16: ffffff80081d9e88
[ 174.489579] x15: 0000007fae377590 x14: 0000000000000000
[ 174.494880] x13: 0000000000000000 x12: ffffff8008dd1000
[ 174.500180] x11: ffffff80096b6fff x10: ffffff80096b6fff
[ 174.505480] x9 : 0000000040000000 x8 : ffffff8008db6000
[ 174.510781] x7 : ffffff80096b7000 x6 : 0000000030127000
[ 174.516082] x5 : 0000000030128000 x4 : 0000000000000000
[ 174.521382] x3 : 00e8000000000713 x2 : 0000000000000f80
[ 174.526682] x1 : ffffff80096b6000 x0 : 0000000030127000
[ 174.531982]
[ 174.533461] Process cp (pid: 2059, stack limit = 0xffffffc0f34bc020)
[ 174.848448] [<ffffff800835e080>] __arch_copy_to_user+0x180/0x280
[ 174.854448] [<ffffff8008245f34>] read_from_oldmem.part.4+0xb4/0xf4
[ 174.860615] [<ffffff8008246074>] read_vmcore+0x100/0x22c
[ 174.865919] [<ffffff8008239378>] proc_reg_read+0x64/0x90
[ 174.871223] [<ffffff80081d7da8>] __vfs_read+0x28/0x108
[ 174.876348] [<ffffff80081d8ae4>] vfs_read+0x84/0x144
[ 174.881301] [<ffffff80081d9ecc>] SyS_read+0x44/0xa0
[ 174.886167] [<ffffff8008082ef0>] el0_svc_naked+0x24/0x28
[ 174.891466] Code: 00000000 00000000 00000000 00000000 (a8c12027)
[ 174.897562] ---[ end trace 00801b2e35b0cd1f ]---
When reading /proc/vmcore with cat/cp or or mmap()ing it with makedumpfile.
The fs/proc/vmcore.c code provides a hook to indicate whether oldmem pages
are ram or not. Use this to look for our earlier handiwork in memblock.
Signed-off-by: James Morse <james.morse@arm.com>
---
Hi Pratyush,
I couldn't get makedumpfile to build, or rather it depends on elfutils which
wouldn't build for autotools reasons. Does implementing this hook solve your
makedumpfile issue?
With this patch I can extract a usable vmcore file using read or mmap,
avoiding the earlier splat.
Akashi, if you agree this is the right thing to do, please consider folding
this into patch 5. (no need to keep the commit mesage or anything).
Thanks,
James
arch/arm64/kernel/crash_dump.c | 28 ++++++++++++++++++++++++++++
1 file changed, 28 insertions(+)
It will help kexec-tools to prevent copying of any unnecessary data. I
think, then you also need to change phys_offset calculation in kexec-tools. That
should be start of either of first "reserved" or "System RAM" block.
~Pratyush
Can you try the following change?
If it fixes your problem, I will post it as a patch.
Almost! This causes booting with acpi=on to fail for the familiar
alignment-fault reasons[2],
details and a suggested fix below.
I think we should have this change as it matches x86's use of acpi, and means we
don't rely on re-parsing the efi memory map to learn which areas of memory
shouldn't be in the vmcore.
quoted hunk
===8<===
From 740563e4a437f0d6ecf6e421c91433f9b8f19041 Mon Sep 17 00:00:00 2001
From: AKASHI Takahiro <redacted>
Date: Fri, 19 Aug 2016 09:57:52 +0900
Subject: [PATCH] arm64: mark reserved memblock regions explicitly
---
arch/arm64/kernel/setup.c | 9 +++++++--
1 file changed, 7 insertions(+), 2 deletions(-)
This causes acpica to choke. arch/arm64/include/asm/acpi.h:acpi_os_ioremap()
calls page_is_ram(), which expects IORESOURCE_SYSTEM_RAM. From kernel/resource.c:
/*
* This generic page_is_ram() returns true if specified address is
* registered as System RAM in iomem_resource list.
*/
int __weak page_is_ram(unsigned long pfn)
We are trying to infer information about the EFI memory map by looking through
iomem_resource list generated from memblock.
drivers/firmware/efi/arm-init.c:reserve_regions() adds memory with the WB
attribute to memblock via early_init_dt_add_memory_arch(), so changing
page_is_ram() for memblock_is_memory() is one step closer to checking the
attributes in the efi memory map (which turns out to tricky).
With your v24 on v4.8-rc1 and 'mark reserved memblock regions explicitly', and
this extra hack [0], I can boot, kdump, extract the vmcore (with read() and
mmap()), and pull things out of it with crash.
Thanks,
James
copy_oldmem_page() and mmap_vmcore() provide two ways for userspace to read
from /proc/vmcore. Neither of these check with memblock to see if the page
they are accessing is nomap. On Seattle this causes:
Thanks for the patch.It did resolve the kernel crash issue with makedumpfile,
however neither there was any data in vmcore-dmesg nor crash utility was able to
work the saved vmcore.
vmcore-dmesg doesn't work for me, but crash did once I'd rebuilt it from the
most recent source.
The most recent commit I have is:
b349598bb755 ("Fix for the ARM64 "bt -R <symbol>" option if the only reference")
Are you using an older version?
Thanks,
James
[0] http://lists.infradead.org/pipermail/linux-arm-kernel/2016-August/447597.html
copy_oldmem_page() and mmap_vmcore() provide two ways for userspace to read
from /proc/vmcore. Neither of these check with memblock to see if the page
they are accessing is nomap. On Seattle this causes:
Thanks for the patch.It did resolve the kernel crash issue with makedumpfile,
however neither there was any data in vmcore-dmesg nor crash utility was able to
work the saved vmcore.
vmcore-dmesg doesn't work for me, but crash did once I'd rebuilt it from the
most recent source.
Yes, saved vmcore worked with latest crash. However, we will need to correct
phys_offset and page_offset in kexec-tools to get meaningful output from vmcore-dmesg.
~Pratyush
James,
On Fri, Aug 19, 2016 at 02:28:06PM +0100, James Morse wrote:
On 19/08/16 02:26, AKASHI Takahiro wrote:
quoted
Can you try the following change?
If it fixes your problem, I will post it as a patch.
Almost! This causes booting with acpi=on to fail for the familiar
alignment-fault reasons[2],
details and a suggested fix below.
Thank you for the fix.
I will merge your hunk to the patch.
-Takahiro AKASHI
quoted hunk
I think we should have this change as it matches x86's use of acpi, and means we
don't rely on re-parsing the efi memory map to learn which areas of memory
shouldn't be in the vmcore.
quoted
===8<===
From 740563e4a437f0d6ecf6e421c91433f9b8f19041 Mon Sep 17 00:00:00 2001
From: AKASHI Takahiro <redacted>
Date: Fri, 19 Aug 2016 09:57:52 +0900
Subject: [PATCH] arm64: mark reserved memblock regions explicitly
---
arch/arm64/kernel/setup.c | 9 +++++++--
1 file changed, 7 insertions(+), 2 deletions(-)
This causes acpica to choke. arch/arm64/include/asm/acpi.h:acpi_os_ioremap()
calls page_is_ram(), which expects IORESOURCE_SYSTEM_RAM. From kernel/resource.c:
quoted
/*
* This generic page_is_ram() returns true if specified address is
* registered as System RAM in iomem_resource list.
*/
int __weak page_is_ram(unsigned long pfn)
We are trying to infer information about the EFI memory map by looking through
iomem_resource list generated from memblock.
drivers/firmware/efi/arm-init.c:reserve_regions() adds memory with the WB
attribute to memblock via early_init_dt_add_memory_arch(), so changing
page_is_ram() for memblock_is_memory() is one step closer to checking the
attributes in the efi memory map (which turns out to tricky).
With your v24 on v4.8-rc1 and 'mark reserved memblock regions explicitly', and
this extra hack [0], I can boot, kdump, extract the vmcore (with read() and
mmap()), and pull things out of it with crash.
Thanks,
James
It will help kexec-tools to prevent copying of any unnecessary data. I
think, then you also need to change phys_offset calculation in kexec-tools. That
should be start of either of first "reserved" or "System RAM" block.
Good point, but I'm not sure this is always true.
Is there any system whose ACPI memory is *not* part of DRAM
(so not part of linear mapping)?
Thanks,
-Takahiro AKASHI
It will help kexec-tools to prevent copying of any unnecessary data. I
think, then you also need to change phys_offset calculation in kexec-tools. That
should be start of either of first "reserved" or "System RAM" block.
Good point, but I'm not sure this is always true.
Is there any system whose ACPI memory is *not* part of DRAM
(so not part of linear mapping)?
Looking into kernel/resource.c:reserve_setup(), it seems that there could be
some none-DRAM area as well, which could be marked as "reserved". So, I think if
we mark nomap region as "reserved" then applications like kexec-tools may not
always identify start of DRAM correctly. Probably, we should give an unique name
to reserved system ram area.
~Pratyush
On Fri, Aug 19, 2016 at 04:52:17PM +0530, Pratyush Anand wrote:
quoted
It will help kexec-tools to prevent copying of any unnecessary data. I
think, then you also need to change phys_offset calculation in kexec-tools. That
should be start of either of first "reserved" or "System RAM" block.
Good point, but I'm not sure this is always true.
Is there any system whose ACPI memory is *not* part of DRAM
From the spec, it looks like this is allowed.
What do you mean by 'DRAM'? Any ACPI region will be in the UEFI memory map, so
the question is what is its type and memory attributes?
The UEFI spec[0] says ACPI regions can have a type of EfiACPIReclaimMemory or
EfiACPIMemoryNVS, the memory attributes aren't specified, so are chosen by the
firmware.
It is possible these regions have to be mapped non-cacheable, page 40 has a
couple of:
If no information about the table location exists in the UEFI memory map or
ACPI memory
descriptors, the table is assumed to be non-cached.
reserve_regions() in drivers/firmware/efi/arm-init.c will add any entry in the
memory map that has a 'WB' attribute to the memblock.memory list (via
early_init_dt_add_memory_arch()), it will also mark as no-map regions that have
this attribute and aren't in the is_reserve_region() list.
If these ACPI regions have the 'WB' attribute, we add them as memory and mark
them nomap. These show up as either a hole, or 'reserved' in /proc/iomem.
If they don't have the 'WB' attribute, then then they are left out of memblock
and aren't part of DRAM, I don't think these will show up in /proc/iomem at all.
Thanks,
James
[0] '2.3.6 AArch64 Platforms' of version 2.6 of the UEFI spec at
http://uefi.org/specifications
On Mon, Aug 22, 2016 at 02:47:30PM +0100, James Morse wrote:
On 22/08/16 02:29, AKASHI Takahiro wrote:
quoted
On Fri, Aug 19, 2016 at 04:52:17PM +0530, Pratyush Anand wrote:
quoted
It will help kexec-tools to prevent copying of any unnecessary data. I
think, then you also need to change phys_offset calculation in kexec-tools. That
should be start of either of first "reserved" or "System RAM" block.
Good point, but I'm not sure this is always true.
quoted
Is there any system whose ACPI memory is *not* part of DRAM
From the spec, it looks like this is allowed.
What do you mean by 'DRAM'? Any ACPI region will be in the UEFI memory map, so
the question is what is its type and memory attributes?
Yes.
The UEFI spec[0] says ACPI regions can have a type of EfiACPIReclaimMemory or
EfiACPIMemoryNVS, the memory attributes aren't specified, so are chosen by the
firmware.
It is possible these regions have to be mapped non-cacheable, page 40 has a
couple of:
quoted
If no information about the table location exists in the UEFI memory map or
ACPI memory
quoted
descriptors, the table is assumed to be non-cached.
reserve_regions() in drivers/firmware/efi/arm-init.c will add any entry in the
memory map that has a 'WB' attribute to the memblock.memory list (via
early_init_dt_add_memory_arch()), it will also mark as no-map regions that have
this attribute and aren't in the is_reserve_region() list.
If these ACPI regions have the 'WB' attribute, we add them as memory and mark
them nomap. These show up as either a hole, or 'reserved' in /proc/iomem.
If they don't have the 'WB' attribute, then then they are left out of memblock
and aren't part of DRAM, I don't think these will show up in /proc/iomem at all.
Let's say,
0x1000-0x1fff: reserved (SRAM for UEFI, WB)
0x80000000-0xffffffff: System RAM (DRAM)
If, as Pratyush suggested, "reserved" resources are added to phys_offset
calculation, the kernel linear mapping area starts at PAGE_OFFSET, but
there is no actual mapping around PAGE_OFFSET.
It won't hurt anything, but looks funny.
So we'd better not include "reserved" in phys_offset calculation anyway.
-> Pratyush
Thanks,
-Takahiro AKASHI
On Mon, Aug 22, 2016 at 02:47:30PM +0100, James Morse wrote:
quoted
On 22/08/16 02:29, AKASHI Takahiro wrote:
quoted
On Fri, Aug 19, 2016 at 04:52:17PM +0530, Pratyush Anand wrote:
quoted
It will help kexec-tools to prevent copying of any unnecessary data. I
think, then you also need to change phys_offset calculation in kexec-tools. That
should be start of either of first "reserved" or "System RAM" block.
Good point, but I'm not sure this is always true.
quoted
Is there any system whose ACPI memory is *not* part of DRAM
From the spec, it looks like this is allowed.
What do you mean by 'DRAM'? Any ACPI region will be in the UEFI memory map, so
the question is what is its type and memory attributes?
Yes.
quoted
The UEFI spec[0] says ACPI regions can have a type of EfiACPIReclaimMemory or
EfiACPIMemoryNVS, the memory attributes aren't specified, so are chosen by the
firmware.
It is possible these regions have to be mapped non-cacheable, page 40 has a
couple of:
quoted
If no information about the table location exists in the UEFI memory map or
ACPI memory
quoted
descriptors, the table is assumed to be non-cached.
reserve_regions() in drivers/firmware/efi/arm-init.c will add any entry in the
memory map that has a 'WB' attribute to the memblock.memory list (via
early_init_dt_add_memory_arch()), it will also mark as no-map regions that have
this attribute and aren't in the is_reserve_region() list.
If these ACPI regions have the 'WB' attribute, we add them as memory and mark
them nomap. These show up as either a hole, or 'reserved' in /proc/iomem.
If they don't have the 'WB' attribute, then then they are left out of memblock
and aren't part of DRAM, I don't think these will show up in /proc/iomem at all.
Let's say,
0x1000-0x1fff: reserved (SRAM for UEFI, WB)
0x80000000-0xffffffff: System RAM (DRAM)
May be slightly more complicated:
0x80000000-0x80001fff: System RAM (DRAM) for UEFI, WB
0x80002000-0xffffffff: System RAM (DRAM)
Kernel will have phys_offset 0x80000000, however kexec-tools will calculate it
as 0x80002000.
If, as Pratyush suggested, "reserved" resources are added to phys_offset
calculation, the kernel linear mapping area starts at PAGE_OFFSET, but
there is no actual mapping around PAGE_OFFSET.
It won't hurt anything, but looks funny.
So we'd better not include "reserved" in phys_offset calculation anyway.
-> Pratyush
My only concern is that, then we will have different values of phys_offset in
kernel and kexec-tools, which might lead to further confusion.
~Pratyush
From: Dave Young <hidden> Date: 2016-08-24 08:04:43
Ccing uefi people.
On 08/23/16 at 04:53pm, Pratyush Anand wrote:
On 23/08/2016:09:38:16 AM, AKASHI Takahiro wrote:
quoted
On Mon, Aug 22, 2016 at 02:47:30PM +0100, James Morse wrote:
quoted
On 22/08/16 02:29, AKASHI Takahiro wrote:
quoted
On Fri, Aug 19, 2016 at 04:52:17PM +0530, Pratyush Anand wrote:
quoted
It will help kexec-tools to prevent copying of any unnecessary data. I
think, then you also need to change phys_offset calculation in kexec-tools. That
should be start of either of first "reserved" or "System RAM" block.
Good point, but I'm not sure this is always true.
quoted
Is there any system whose ACPI memory is *not* part of DRAM
From the spec, it looks like this is allowed.
What do you mean by 'DRAM'? Any ACPI region will be in the UEFI memory map, so
the question is what is its type and memory attributes?
Yes.
quoted
The UEFI spec[0] says ACPI regions can have a type of EfiACPIReclaimMemory or
EfiACPIMemoryNVS, the memory attributes aren't specified, so are chosen by the
firmware.
It is possible these regions have to be mapped non-cacheable, page 40 has a
couple of:
quoted
If no information about the table location exists in the UEFI memory map or
ACPI memory
quoted
descriptors, the table is assumed to be non-cached.
reserve_regions() in drivers/firmware/efi/arm-init.c will add any entry in the
memory map that has a 'WB' attribute to the memblock.memory list (via
early_init_dt_add_memory_arch()), it will also mark as no-map regions that have
this attribute and aren't in the is_reserve_region() list.
Looking the arm-init.c, I suspect it missed the some efi ranges as
reserved ranges like runtime code and runtime data etc. But I might be
wrong.
Below is the is_reserve_region, it will regard any regions which is not
in the EFI_* below as normal memory, it does not include the runtime
ranges and other types.
static __init int is_reserve_region(efi_memory_desc_t *md)
{
switch (md->type) {
case EFI_LOADER_CODE:
case EFI_LOADER_DATA:
case EFI_BOOT_SERVICES_CODE:
case EFI_BOOT_SERVICES_DATA:
case EFI_CONVENTIONAL_MEMORY:
case EFI_PERSISTENT_MEMORY:
return 0;
default:
break;
}
return is_normal_ram(md);
}
Let's see the x86 do_add_efi_mem_map, the default case set all other
types as reserved. Shouldn't this be same in all arches though there's
no e820 in arm(64)?
static void __init do_add_efi_memmap(void)
{
[snip]
switch (md->type) {
case EFI_LOADER_CODE:
case EFI_LOADER_DATA:
case EFI_BOOT_SERVICES_CODE:
case EFI_BOOT_SERVICES_DATA:
case EFI_CONVENTIONAL_MEMORY:
if (md->attribute & EFI_MEMORY_WB)
e820_type = E820_RAM;
else
e820_type = E820_RESERVED;
break;
[snip]
default:
/*
* EFI_RESERVED_TYPE EFI_RUNTIME_SERVICES_CODE
* EFI_RUNTIME_SERVICES_DATA
* EFI_MEMORY_MAPPED_IO
* EFI_MEMORY_MAPPED_IO_PORT_SPACE EFI_PAL_CODE
*/
e820_type = E820_RESERVED;
break;
}
[snip]
}
quoted
quoted
If these ACPI regions have the 'WB' attribute, we add them as memory and mark
them nomap. These show up as either a hole, or 'reserved' in /proc/iomem.
If they don't have the 'WB' attribute, then then they are left out of memblock
and aren't part of DRAM, I don't think these will show up in /proc/iomem at all.
Let's say,
0x1000-0x1fff: reserved (SRAM for UEFI, WB)
0x80000000-0xffffffff: System RAM (DRAM)
May be slightly more complicated:
0x80000000-0x80001fff: System RAM (DRAM) for UEFI, WB
0x80002000-0xffffffff: System RAM (DRAM)
Kernel will have phys_offset 0x80000000, however kexec-tools will calculate it
as 0x80002000.
quoted
If, as Pratyush suggested, "reserved" resources are added to phys_offset
calculation, the kernel linear mapping area starts at PAGE_OFFSET, but
there is no actual mapping around PAGE_OFFSET.
It won't hurt anything, but looks funny.
So we'd better not include "reserved" in phys_offset calculation anyway.
-> Pratyush
My only concern is that, then we will have different values of phys_offset in
kernel and kexec-tools, which might lead to further confusion.
~Pratyush
Looking the arm-init.c, I suspect it missed the some efi ranges as
reserved ranges like runtime code and runtime data etc. But I might be
wrong.
This had me confused for too... I think I get it, my understanding is:
static __init int is_reserve_region(efi_memory_desc_t *md)
{
switch (md->type) {
case EFI_LOADER_CODE:
case EFI_LOADER_DATA:
case EFI_BOOT_SERVICES_CODE:
case EFI_BOOT_SERVICES_DATA:
case EFI_CONVENTIONAL_MEMORY:
case EFI_PERSISTENT_MEMORY:
return 0;
return false - this is the list of region-types to never reserve, regardless of
memory attributes.
default:
break;
}
return is_normal_ram(md);
If its not in the 'never reserve' list above, then we check if the region is
'normal' ram. If it is then it will end up in memblock.memory so we return true,
causing it to be marked nomap too.
reserve_regions() in that same file calls is_normal_ram() directly before adding
all regions with the WB attribute to memblock.memory via
early_init_dt_add_memory_arch().
A runtime region with the WB attribute will be caught by is_reserve_region(),
and is_normal_ram(), so it ends up in memblock.memory and memblock.nomap.
}
Let's see the x86 do_add_efi_mem_map, the default case set all other
types as reserved. Shouldn't this be same in all arches though there's
no e820 in arm(64)?
static void __init do_add_efi_memmap(void)
{
[snip]
switch (md->type) {
case EFI_LOADER_CODE:
case EFI_LOADER_DATA:
case EFI_BOOT_SERVICES_CODE:
case EFI_BOOT_SERVICES_DATA:
case EFI_CONVENTIONAL_MEMORY:
if (md->attribute & EFI_MEMORY_WB)
e820_type = E820_RAM;
In this case reserve_regions() will add the memory to memblock.memory because it
has the WB attribute, and not reserve it.
else
e820_type = E820_RESERVED;
Without the WB attribute, these regions are in neither memblock.memory nor
memblock.nomap.
On 9 August 2016 at 03:56, AKASHI Takahiro [off-list ref] wrote:
quoted hunk
On crash dump kernel, all the information about primary kernel's system
memory (core image) is available in elf core header.
The primary kernel will set aside this header with reserve_elfcorehdr()
at boot time and inform crash dump kernel of its location via a new
device-tree property, "linux,elfcorehdr".
Please note that all other architectures use traditional "elfcorehdr="
kernel parameter for this purpose.
Then crash dump kernel will access the primary kernel's memory with
copy_oldmem_page(), which reads one page by ioremap'ing it since it does
not reside in linear mapping on crash dump kernel.
We also need our own elfcorehdr_read() here since the header is placed
within crash dump kernel's usable memory.
Signed-off-by: AKASHI Takahiro <redacted>
---
arch/arm64/Kconfig | 11 +++++++
arch/arm64/kernel/Makefile | 1 +
arch/arm64/kernel/crash_dump.c | 71 ++++++++++++++++++++++++++++++++++++++++++
arch/arm64/mm/init.c | 54 ++++++++++++++++++++++++++++++++
4 files changed, 137 insertions(+)
create mode 100644 arch/arm64/kernel/crash_dump.c
... and memunmap here?
ioremap_cache() is not very well defined, and memremap() has been
introduced specifically to replace it, so I think we should use it in
new code.
Thanks,
Ard.
quoted hunk
+
+ return csize;
+}
+
+/**
+ * elfcorehdr_read - read from ELF core header
+ * @buf: buffer where the data is placed
+ * @csize: number of bytes to read
+ * @ppos: address in the memory
+ *
+ * This function reads @count bytes from elf core header which exists
+ * on crash dump kernel's memory.
+ */
+ssize_t elfcorehdr_read(char *buf, size_t count, u64 *ppos)
+{
+ memcpy(buf, phys_to_virt((phys_addr_t)*ppos), count);
+ return count;
+}
From: Dave Young <hidden> Date: 2016-08-25 01:04:26
On 08/24/16 at 11:25am, James Morse wrote:
Hi Dave,
On 24/08/16 09:04, Dave Young wrote:
quoted
Looking the arm-init.c, I suspect it missed the some efi ranges as
reserved ranges like runtime code and runtime data etc. But I might be
wrong.
This had me confused for too... I think I get it, my understanding is:
James, thanks for your clarification.
quoted
static __init int is_reserve_region(efi_memory_desc_t *md)
{
switch (md->type) {
case EFI_LOADER_CODE:
case EFI_LOADER_DATA:
case EFI_BOOT_SERVICES_CODE:
case EFI_BOOT_SERVICES_DATA:
case EFI_CONVENTIONAL_MEMORY:
case EFI_PERSISTENT_MEMORY:
return 0;
return false - this is the list of region-types to never reserve, regardless of
memory attributes.
quoted
default:
break;
}
return is_normal_ram(md);
If its not in the 'never reserve' list above, then we check if the region is
'normal' ram. If it is then it will end up in memblock.memory so we return true,
causing it to be marked nomap too.
reserve_regions() in that same file calls is_normal_ram() directly before adding
all regions with the WB attribute to memblock.memory via
early_init_dt_add_memory_arch().
A runtime region with the WB attribute will be caught by is_reserve_region(),
and is_normal_ram(), so it ends up in memblock.memory and memblock.nomap.
Hmm, It is not straitforward like the do_add_efi_memmap. I got it.
BTW, I believe there is same problem in arm as well as arm64, it also
need mark the runtime ranges as "reserved" /proc/iomem.
quoted
}
Let's see the x86 do_add_efi_mem_map, the default case set all other
types as reserved. Shouldn't this be same in all arches though there's
no e820 in arm(64)?
quoted
static void __init do_add_efi_memmap(void)
{
[snip]
switch (md->type) {
case EFI_LOADER_CODE:
case EFI_LOADER_DATA:
case EFI_BOOT_SERVICES_CODE:
case EFI_BOOT_SERVICES_DATA:
case EFI_CONVENTIONAL_MEMORY:
if (md->attribute & EFI_MEMORY_WB)
e820_type = E820_RAM;
In this case reserve_regions() will add the memory to memblock.memory because it
has the WB attribute, and not reserve it.
quoted
else
e820_type = E820_RESERVED;
Without the WB attribute, these regions are in neither memblock.memory nor
memblock.nomap.
On Wed, Aug 24, 2016 at 04:44:09PM +0200, Ard Biesheuvel wrote:
On 9 August 2016 at 03:56, AKASHI Takahiro [off-list ref] wrote:
quoted
On crash dump kernel, all the information about primary kernel's system
memory (core image) is available in elf core header.
The primary kernel will set aside this header with reserve_elfcorehdr()
at boot time and inform crash dump kernel of its location via a new
device-tree property, "linux,elfcorehdr".
Please note that all other architectures use traditional "elfcorehdr="
kernel parameter for this purpose.
Then crash dump kernel will access the primary kernel's memory with
copy_oldmem_page(), which reads one page by ioremap'ing it since it does
not reside in linear mapping on crash dump kernel.
We also need our own elfcorehdr_read() here since the header is placed
within crash dump kernel's usable memory.
Signed-off-by: AKASHI Takahiro <redacted>
---
arch/arm64/Kconfig | 11 +++++++
arch/arm64/kernel/Makefile | 1 +
arch/arm64/kernel/crash_dump.c | 71 ++++++++++++++++++++++++++++++++++++++++++
arch/arm64/mm/init.c | 54 ++++++++++++++++++++++++++++++++
4 files changed, 137 insertions(+)
create mode 100644 arch/arm64/kernel/crash_dump.c
... and memunmap here?
ioremap_cache() is not very well defined, and memremap() has been
introduced specifically to replace it, so I think we should use it in
new code.
Sure. I will use memremap(MEMREMAP_WB) instead.
Thanks,
-Takahiro AKASHI
Thanks,
Ard.
quoted
+
+ return csize;
+}
+
+/**
+ * elfcorehdr_read - read from ELF core header
+ * @buf: buffer where the data is placed
+ * @csize: number of bytes to read
+ * @ppos: address in the memory
+ *
+ * This function reads @count bytes from elf core header which exists
+ * on crash dump kernel's memory.
+ */
+ssize_t elfcorehdr_read(char *buf, size_t count, u64 *ppos)
+{
+ memcpy(buf, phys_to_virt((phys_addr_t)*ppos), count);
+ return count;
+}
Hi Akashi,
On 08/09/2016 07:22 AM, AKASHI Takahiro wrote:
This patch series adds kdump support on arm64.
To load a crash-dump kernel to the systems, a series of patches to
kexec-tools, which have not yet been merged upstream, are needed.
Please use my kdump patches [1].
To examine vmcore (/proc/vmcore) on a crash-dump kernel, you can use
- crash utility (coming v7.1.6 or later) [2]
(Necessary patches have already been queued in the master.)
[1] T.B.D.
[2] https://github.com/crash-utility/crash.git
Changes for v24 (Aug 9, 2016):
o Rebase to Linux-4.8-rc1
o Update descriptions about newly added DT proerties
Changes for v23 (July 26, 2016):
o Move memblock_reserve() to a single place in reserve_crashkernel()
o Use cpu_park_loop() in ipi_cpu_crash_stop()
o Always enforce ARCH_LOW_ADDRESS_LIMIT to the memory range of crash kernel
o Re-implement fdt_enforce_memory_region() to remove non-reserve regions
(for ACPI) from usable memory at crash kernel
Changes for v22 (July 12, 2016):
o Export "crashkernel-base" and "crashkernel-size" via device-tree,
and add some descriptions about them in chosen.txt
o Rename "usable-memory" to "usable-memory-range" to avoid inconsistency
with powerpc's "usable-memory"
o Make cosmetic changes regarding "ifdef" usage
o Correct some wordings in kdump.txt
Changes for v21 (July 6, 2016):
o Remove kexec patches.
o Rebase to arm64's for-next/core (Linux-4.7-rc4 based).
o Clarify the description about kvm in kdump.txt.
See the following link [3] for older changes:
[3] http://lists.infradead.org/pipermail/linux-arm-kernel/2016-June/438780.html
AKASHI Takahiro (8):
arm64: kdump: reserve memory for crash dump kernel
memblock: add memblock_cap_memory_range()
arm64: limit memory regions based on DT property, usable-memory-range
arm64: kdump: implement machine_crash_shutdown()
arm64: kdump: add kdump support
arm64: kdump: add VMCOREINFO's for user-space coredump tools
arm64: kdump: enable kdump in the arm64 defconfig
arm64: kdump: update a kernel doc
James Morse (1):
Documentation: dt: chosen properties for arm64 kdump
Documentation/devicetree/bindings/chosen.txt | 45 ++++++
Documentation/kdump/kdump.txt | 16 ++-
arch/arm64/Kconfig | 11 ++
arch/arm64/configs/defconfig | 1 +
arch/arm64/include/asm/hardirq.h | 2 +-
arch/arm64/include/asm/kexec.h | 41 +++++-
arch/arm64/include/asm/smp.h | 2 +
arch/arm64/kernel/Makefile | 1 +
arch/arm64/kernel/crash_dump.c | 71 ++++++++++
arch/arm64/kernel/machine_kexec.c | 67 ++++++++-
arch/arm64/kernel/setup.c | 7 +-
arch/arm64/kernel/smp.c | 63 +++++++++
arch/arm64/mm/init.c | 202 +++++++++++++++++++++++++++
include/linux/memblock.h | 1 +
mm/memblock.c | 28 ++++
15 files changed, 551 insertions(+), 7 deletions(-)
create mode 100644 arch/arm64/kernel/crash_dump.c
Couple of points
a) Just a note, while testing, the crashkernel reserved memory should be less than ARCH_LOW_ADDRESS_LIMIT (=arm64_dma_phys_limit).
b) Has anyone tested this on a SoC with Gicv3 ITS ?
Should the GICD/R be reset prior to switching to crash kernel ?
I am seeing lot of GICv3: RWP timeout, gone fishing while crash kernel boots.
Thanks,
Manish
Manish,
Thank you for testing my kdump and reporting issues.
On Wed, Aug 31, 2016 at 09:11:52AM +0530, Manish Jaggi wrote:
Hi Akashi,
On 08/09/2016 07:22 AM, AKASHI Takahiro wrote:
quoted
This patch series adds kdump support on arm64.
To load a crash-dump kernel to the systems, a series of patches to
kexec-tools, which have not yet been merged upstream, are needed.
Please use my kdump patches [1].
To examine vmcore (/proc/vmcore) on a crash-dump kernel, you can use
- crash utility (coming v7.1.6 or later) [2]
(Necessary patches have already been queued in the master.)
[1] T.B.D.
[2] https://github.com/crash-utility/crash.git
Changes for v24 (Aug 9, 2016):
o Rebase to Linux-4.8-rc1
o Update descriptions about newly added DT proerties
Changes for v23 (July 26, 2016):
o Move memblock_reserve() to a single place in reserve_crashkernel()
o Use cpu_park_loop() in ipi_cpu_crash_stop()
o Always enforce ARCH_LOW_ADDRESS_LIMIT to the memory range of crash kernel
o Re-implement fdt_enforce_memory_region() to remove non-reserve regions
(for ACPI) from usable memory at crash kernel
Changes for v22 (July 12, 2016):
o Export "crashkernel-base" and "crashkernel-size" via device-tree,
and add some descriptions about them in chosen.txt
o Rename "usable-memory" to "usable-memory-range" to avoid inconsistency
with powerpc's "usable-memory"
o Make cosmetic changes regarding "ifdef" usage
o Correct some wordings in kdump.txt
Changes for v21 (July 6, 2016):
o Remove kexec patches.
o Rebase to arm64's for-next/core (Linux-4.7-rc4 based).
o Clarify the description about kvm in kdump.txt.
See the following link [3] for older changes:
[3] http://lists.infradead.org/pipermail/linux-arm-kernel/2016-June/438780.html
AKASHI Takahiro (8):
arm64: kdump: reserve memory for crash dump kernel
memblock: add memblock_cap_memory_range()
arm64: limit memory regions based on DT property, usable-memory-range
arm64: kdump: implement machine_crash_shutdown()
arm64: kdump: add kdump support
arm64: kdump: add VMCOREINFO's for user-space coredump tools
arm64: kdump: enable kdump in the arm64 defconfig
arm64: kdump: update a kernel doc
James Morse (1):
Documentation: dt: chosen properties for arm64 kdump
Documentation/devicetree/bindings/chosen.txt | 45 ++++++
Documentation/kdump/kdump.txt | 16 ++-
arch/arm64/Kconfig | 11 ++
arch/arm64/configs/defconfig | 1 +
arch/arm64/include/asm/hardirq.h | 2 +-
arch/arm64/include/asm/kexec.h | 41 +++++-
arch/arm64/include/asm/smp.h | 2 +
arch/arm64/kernel/Makefile | 1 +
arch/arm64/kernel/crash_dump.c | 71 ++++++++++
arch/arm64/kernel/machine_kexec.c | 67 ++++++++-
arch/arm64/kernel/setup.c | 7 +-
arch/arm64/kernel/smp.c | 63 +++++++++
arch/arm64/mm/init.c | 202 +++++++++++++++++++++++++++
include/linux/memblock.h | 1 +
mm/memblock.c | 28 ++++
15 files changed, 551 insertions(+), 7 deletions(-)
create mode 100644 arch/arm64/kernel/crash_dump.c
Couple of points
a) Just a note, while testing, the crashkernel reserved memory should be less than ARCH_LOW_ADDRESS_LIMIT (=arm64_dma_phys_limit).
I think that this is a common mistake not only for kdump, but also
for general kernels.
Since request_standard_resources() calls alloc_bootmem_low(),
the kernel will panic if any of usable "System RAM" is located
above ARCH_LOW_ADDRESS_LIMIT.
For kdump, using "crashkernel=SS" notation is a convenient way
to avoid this issue.
b) Has anyone tested this on a SoC with Gicv3 ITS ?
Should the GICD/R be reset prior to switching to crash kernel ?
I am seeing lot of GICv3: RWP timeout, gone fishing while crash kernel boots.
I've never seen this kind of messages.
I usually do my testing on a fast model.
"compatible" of interrupt-controller is "arm,gic-v3."
Thanks,
-Takahiro AKASHI
Manish,
Thank you for testing my kdump and reporting issues.
On Wed, Aug 31, 2016 at 09:11:52AM +0530, Manish Jaggi wrote:
quoted
Hi Akashi,
On 08/09/2016 07:22 AM, AKASHI Takahiro wrote:
quoted
This patch series adds kdump support on arm64.
To load a crash-dump kernel to the systems, a series of patches to
kexec-tools, which have not yet been merged upstream, are needed.
Please use my kdump patches [1].
To examine vmcore (/proc/vmcore) on a crash-dump kernel, you can use
- crash utility (coming v7.1.6 or later) [2]
(Necessary patches have already been queued in the master.)
[1] T.B.D.
[2] https://github.com/crash-utility/crash.git
Changes for v24 (Aug 9, 2016):
o Rebase to Linux-4.8-rc1
o Update descriptions about newly added DT proerties
Changes for v23 (July 26, 2016):
o Move memblock_reserve() to a single place in reserve_crashkernel()
o Use cpu_park_loop() in ipi_cpu_crash_stop()
o Always enforce ARCH_LOW_ADDRESS_LIMIT to the memory range of crash kernel
o Re-implement fdt_enforce_memory_region() to remove non-reserve regions
(for ACPI) from usable memory at crash kernel
Changes for v22 (July 12, 2016):
o Export "crashkernel-base" and "crashkernel-size" via device-tree,
and add some descriptions about them in chosen.txt
o Rename "usable-memory" to "usable-memory-range" to avoid inconsistency
with powerpc's "usable-memory"
o Make cosmetic changes regarding "ifdef" usage
o Correct some wordings in kdump.txt
Changes for v21 (July 6, 2016):
o Remove kexec patches.
o Rebase to arm64's for-next/core (Linux-4.7-rc4 based).
o Clarify the description about kvm in kdump.txt.
See the following link [3] for older changes:
[3] http://lists.infradead.org/pipermail/linux-arm-kernel/2016-June/438780.html
AKASHI Takahiro (8):
arm64: kdump: reserve memory for crash dump kernel
memblock: add memblock_cap_memory_range()
arm64: limit memory regions based on DT property, usable-memory-range
arm64: kdump: implement machine_crash_shutdown()
arm64: kdump: add kdump support
arm64: kdump: add VMCOREINFO's for user-space coredump tools
arm64: kdump: enable kdump in the arm64 defconfig
arm64: kdump: update a kernel doc
James Morse (1):
Documentation: dt: chosen properties for arm64 kdump
Documentation/devicetree/bindings/chosen.txt | 45 ++++++
Documentation/kdump/kdump.txt | 16 ++-
arch/arm64/Kconfig | 11 ++
arch/arm64/configs/defconfig | 1 +
arch/arm64/include/asm/hardirq.h | 2 +-
arch/arm64/include/asm/kexec.h | 41 +++++-
arch/arm64/include/asm/smp.h | 2 +
arch/arm64/kernel/Makefile | 1 +
arch/arm64/kernel/crash_dump.c | 71 ++++++++++
arch/arm64/kernel/machine_kexec.c | 67 ++++++++-
arch/arm64/kernel/setup.c | 7 +-
arch/arm64/kernel/smp.c | 63 +++++++++
arch/arm64/mm/init.c | 202 +++++++++++++++++++++++++++
include/linux/memblock.h | 1 +
mm/memblock.c | 28 ++++
15 files changed, 551 insertions(+), 7 deletions(-)
create mode 100644 arch/arm64/kernel/crash_dump.c
Couple of points
a) Just a note, while testing, the crashkernel reserved memory should be less than ARCH_LOW_ADDRESS_LIMIT (=arm64_dma_phys_limit).
I think that this is a common mistake not only for kdump, but also
for general kernels.
Since request_standard_resources() calls alloc_bootmem_low(),
the kernel will panic if any of usable "System RAM" is located
above ARCH_LOW_ADDRESS_LIMIT.
For kdump, using "crashkernel=SS" notation is a convenient way
to avoid this issue.
quoted
b) Has anyone tested this on a SoC with Gicv3 ITS ?
Should the GICD/R be reset prior to switching to crash kernel ?
I am seeing lot of GICv3: RWP timeout, gone fishing while crash kernel boots.
I've never seen this kind of messages.
I usually do my testing on a fast model.
"compatible" of interrupt-controller is "arm,gic-v3."
I suspect gic_cpu_pm_notifier is not being called on any of the cores prior to start of crash kernel.
We might have to call it explicitly.
[Cc: Marc]
On Fri, Sep 02, 2016 at 06:23:25PM +0530, Manish Jaggi wrote:
On 08/31/2016 11:01 AM, AKASHI Takahiro wrote:
quoted
Manish,
Thank you for testing my kdump and reporting issues.
On Wed, Aug 31, 2016 at 09:11:52AM +0530, Manish Jaggi wrote:
quoted
Hi Akashi,
On 08/09/2016 07:22 AM, AKASHI Takahiro wrote:
quoted
This patch series adds kdump support on arm64.
To load a crash-dump kernel to the systems, a series of patches to
kexec-tools, which have not yet been merged upstream, are needed.
Please use my kdump patches [1].
To examine vmcore (/proc/vmcore) on a crash-dump kernel, you can use
- crash utility (coming v7.1.6 or later) [2]
(Necessary patches have already been queued in the master.)
[1] T.B.D.
[2] https://github.com/crash-utility/crash.git
Changes for v24 (Aug 9, 2016):
o Rebase to Linux-4.8-rc1
o Update descriptions about newly added DT proerties
Changes for v23 (July 26, 2016):
o Move memblock_reserve() to a single place in reserve_crashkernel()
o Use cpu_park_loop() in ipi_cpu_crash_stop()
o Always enforce ARCH_LOW_ADDRESS_LIMIT to the memory range of crash kernel
o Re-implement fdt_enforce_memory_region() to remove non-reserve regions
(for ACPI) from usable memory at crash kernel
Changes for v22 (July 12, 2016):
o Export "crashkernel-base" and "crashkernel-size" via device-tree,
and add some descriptions about them in chosen.txt
o Rename "usable-memory" to "usable-memory-range" to avoid inconsistency
with powerpc's "usable-memory"
o Make cosmetic changes regarding "ifdef" usage
o Correct some wordings in kdump.txt
Changes for v21 (July 6, 2016):
o Remove kexec patches.
o Rebase to arm64's for-next/core (Linux-4.7-rc4 based).
o Clarify the description about kvm in kdump.txt.
See the following link [3] for older changes:
[3] http://lists.infradead.org/pipermail/linux-arm-kernel/2016-June/438780.html
AKASHI Takahiro (8):
arm64: kdump: reserve memory for crash dump kernel
memblock: add memblock_cap_memory_range()
arm64: limit memory regions based on DT property, usable-memory-range
arm64: kdump: implement machine_crash_shutdown()
arm64: kdump: add kdump support
arm64: kdump: add VMCOREINFO's for user-space coredump tools
arm64: kdump: enable kdump in the arm64 defconfig
arm64: kdump: update a kernel doc
James Morse (1):
Documentation: dt: chosen properties for arm64 kdump
Documentation/devicetree/bindings/chosen.txt | 45 ++++++
Documentation/kdump/kdump.txt | 16 ++-
arch/arm64/Kconfig | 11 ++
arch/arm64/configs/defconfig | 1 +
arch/arm64/include/asm/hardirq.h | 2 +-
arch/arm64/include/asm/kexec.h | 41 +++++-
arch/arm64/include/asm/smp.h | 2 +
arch/arm64/kernel/Makefile | 1 +
arch/arm64/kernel/crash_dump.c | 71 ++++++++++
arch/arm64/kernel/machine_kexec.c | 67 ++++++++-
arch/arm64/kernel/setup.c | 7 +-
arch/arm64/kernel/smp.c | 63 +++++++++
arch/arm64/mm/init.c | 202 +++++++++++++++++++++++++++
include/linux/memblock.h | 1 +
mm/memblock.c | 28 ++++
15 files changed, 551 insertions(+), 7 deletions(-)
create mode 100644 arch/arm64/kernel/crash_dump.c
Couple of points
a) Just a note, while testing, the crashkernel reserved memory should be less than ARCH_LOW_ADDRESS_LIMIT (=arm64_dma_phys_limit).
I think that this is a common mistake not only for kdump, but also
for general kernels.
Since request_standard_resources() calls alloc_bootmem_low(),
the kernel will panic if any of usable "System RAM" is located
above ARCH_LOW_ADDRESS_LIMIT.
For kdump, using "crashkernel=SS" notation is a convenient way
to avoid this issue.
quoted
b) Has anyone tested this on a SoC with Gicv3 ITS ?
Should the GICD/R be reset prior to switching to crash kernel ?
I am seeing lot of GICv3: RWP timeout, gone fishing while crash kernel boots.
I've never seen this kind of messages.
I usually do my testing on a fast model.
"compatible" of interrupt-controller is "arm,gic-v3."
I suspect gic_cpu_pm_notifier is not being called on any of the cores prior to start of crash kernel.
We might have to call it explicitly.
I'm not sure that it is the cause, but anyway none of any cpu_pm_notifier's
will be called at panic. That is the reason why "maxcpus=1" should be
specified (for kdump on arm64).
-Takahiro AKASHI
[Cc: Marc]
On Fri, Sep 02, 2016 at 06:23:25PM +0530, Manish Jaggi wrote:
quoted
On 08/31/2016 11:01 AM, AKASHI Takahiro wrote:
quoted
Manish,
Thank you for testing my kdump and reporting issues.
On Wed, Aug 31, 2016 at 09:11:52AM +0530, Manish Jaggi wrote:
quoted
Hi Akashi,
On 08/09/2016 07:22 AM, AKASHI Takahiro wrote:
quoted
This patch series adds kdump support on arm64.
To load a crash-dump kernel to the systems, a series of patches to
kexec-tools, which have not yet been merged upstream, are needed.
Please use my kdump patches [1].
To examine vmcore (/proc/vmcore) on a crash-dump kernel, you can use
- crash utility (coming v7.1.6 or later) [2]
(Necessary patches have already been queued in the master.)
[1] T.B.D.
[2] https://github.com/crash-utility/crash.git
Changes for v24 (Aug 9, 2016):
o Rebase to Linux-4.8-rc1
o Update descriptions about newly added DT proerties
Changes for v23 (July 26, 2016):
o Move memblock_reserve() to a single place in reserve_crashkernel()
o Use cpu_park_loop() in ipi_cpu_crash_stop()
o Always enforce ARCH_LOW_ADDRESS_LIMIT to the memory range of crash kernel
o Re-implement fdt_enforce_memory_region() to remove non-reserve regions
(for ACPI) from usable memory at crash kernel
Changes for v22 (July 12, 2016):
o Export "crashkernel-base" and "crashkernel-size" via device-tree,
and add some descriptions about them in chosen.txt
o Rename "usable-memory" to "usable-memory-range" to avoid inconsistency
with powerpc's "usable-memory"
o Make cosmetic changes regarding "ifdef" usage
o Correct some wordings in kdump.txt
Changes for v21 (July 6, 2016):
o Remove kexec patches.
o Rebase to arm64's for-next/core (Linux-4.7-rc4 based).
o Clarify the description about kvm in kdump.txt.
See the following link [3] for older changes:
[3] http://lists.infradead.org/pipermail/linux-arm-kernel/2016-June/438780.html
AKASHI Takahiro (8):
arm64: kdump: reserve memory for crash dump kernel
memblock: add memblock_cap_memory_range()
arm64: limit memory regions based on DT property, usable-memory-range
arm64: kdump: implement machine_crash_shutdown()
arm64: kdump: add kdump support
arm64: kdump: add VMCOREINFO's for user-space coredump tools
arm64: kdump: enable kdump in the arm64 defconfig
arm64: kdump: update a kernel doc
James Morse (1):
Documentation: dt: chosen properties for arm64 kdump
Documentation/devicetree/bindings/chosen.txt | 45 ++++++
Documentation/kdump/kdump.txt | 16 ++-
arch/arm64/Kconfig | 11 ++
arch/arm64/configs/defconfig | 1 +
arch/arm64/include/asm/hardirq.h | 2 +-
arch/arm64/include/asm/kexec.h | 41 +++++-
arch/arm64/include/asm/smp.h | 2 +
arch/arm64/kernel/Makefile | 1 +
arch/arm64/kernel/crash_dump.c | 71 ++++++++++
arch/arm64/kernel/machine_kexec.c | 67 ++++++++-
arch/arm64/kernel/setup.c | 7 +-
arch/arm64/kernel/smp.c | 63 +++++++++
arch/arm64/mm/init.c | 202 +++++++++++++++++++++++++++
include/linux/memblock.h | 1 +
mm/memblock.c | 28 ++++
15 files changed, 551 insertions(+), 7 deletions(-)
create mode 100644 arch/arm64/kernel/crash_dump.c
Couple of points
a) Just a note, while testing, the crashkernel reserved memory should be less than ARCH_LOW_ADDRESS_LIMIT (=arm64_dma_phys_limit).
I think that this is a common mistake not only for kdump, but also
for general kernels.
Since request_standard_resources() calls alloc_bootmem_low(),
the kernel will panic if any of usable "System RAM" is located
above ARCH_LOW_ADDRESS_LIMIT.
For kdump, using "crashkernel=SS" notation is a convenient way
to avoid this issue.
quoted
b) Has anyone tested this on a SoC with Gicv3 ITS ?
Should the GICD/R be reset prior to switching to crash kernel ?
I am seeing lot of GICv3: RWP timeout, gone fishing while crash kernel boots.
I've never seen this kind of messages.
I usually do my testing on a fast model.
"compatible" of interrupt-controller is "arm,gic-v3."
I suspect gic_cpu_pm_notifier is not being called on any of the cores prior to start of crash kernel.
We might have to call it explicitly.
I'm not sure that it is the cause, but anyway none of any cpu_pm_notifier's
will be called at panic. That is the reason why "maxcpus=1" should be
specified (for kdump on arm64).
What I meant was that since cpu_pm_notifier is not called before crash kernel is started,
GIC Distributor/re-distributor/ITS is not set in quiescent state.
In my setup the GICD_CTRL[RWP] bit is not cleared in the crashkernels' distributor init function.
Marc what do you think?
From: Marc Zyngier <hidden> Date: 2016-09-06 15:33:57
On 05/09/16 13:42, Manish Jaggi wrote:
On 09/05/2016 01:45 PM, AKASHI Takahiro wrote:
quoted
[Cc: Marc]
On Fri, Sep 02, 2016 at 06:23:25PM +0530, Manish Jaggi wrote:
quoted
On 08/31/2016 11:01 AM, AKASHI Takahiro wrote:
quoted
Manish,
Thank you for testing my kdump and reporting issues.
On Wed, Aug 31, 2016 at 09:11:52AM +0530, Manish Jaggi wrote:
quoted
Hi Akashi,
On 08/09/2016 07:22 AM, AKASHI Takahiro wrote:
quoted
This patch series adds kdump support on arm64.
To load a crash-dump kernel to the systems, a series of patches to
kexec-tools, which have not yet been merged upstream, are needed.
Please use my kdump patches [1].
To examine vmcore (/proc/vmcore) on a crash-dump kernel, you can use
- crash utility (coming v7.1.6 or later) [2]
(Necessary patches have already been queued in the master.)
[1] T.B.D.
[2] https://github.com/crash-utility/crash.git
Changes for v24 (Aug 9, 2016):
o Rebase to Linux-4.8-rc1
o Update descriptions about newly added DT proerties
Changes for v23 (July 26, 2016):
o Move memblock_reserve() to a single place in reserve_crashkernel()
o Use cpu_park_loop() in ipi_cpu_crash_stop()
o Always enforce ARCH_LOW_ADDRESS_LIMIT to the memory range of crash kernel
o Re-implement fdt_enforce_memory_region() to remove non-reserve regions
(for ACPI) from usable memory at crash kernel
Changes for v22 (July 12, 2016):
o Export "crashkernel-base" and "crashkernel-size" via device-tree,
and add some descriptions about them in chosen.txt
o Rename "usable-memory" to "usable-memory-range" to avoid inconsistency
with powerpc's "usable-memory"
o Make cosmetic changes regarding "ifdef" usage
o Correct some wordings in kdump.txt
Changes for v21 (July 6, 2016):
o Remove kexec patches.
o Rebase to arm64's for-next/core (Linux-4.7-rc4 based).
o Clarify the description about kvm in kdump.txt.
See the following link [3] for older changes:
[3] http://lists.infradead.org/pipermail/linux-arm-kernel/2016-June/438780.html
AKASHI Takahiro (8):
arm64: kdump: reserve memory for crash dump kernel
memblock: add memblock_cap_memory_range()
arm64: limit memory regions based on DT property, usable-memory-range
arm64: kdump: implement machine_crash_shutdown()
arm64: kdump: add kdump support
arm64: kdump: add VMCOREINFO's for user-space coredump tools
arm64: kdump: enable kdump in the arm64 defconfig
arm64: kdump: update a kernel doc
James Morse (1):
Documentation: dt: chosen properties for arm64 kdump
Documentation/devicetree/bindings/chosen.txt | 45 ++++++
Documentation/kdump/kdump.txt | 16 ++-
arch/arm64/Kconfig | 11 ++
arch/arm64/configs/defconfig | 1 +
arch/arm64/include/asm/hardirq.h | 2 +-
arch/arm64/include/asm/kexec.h | 41 +++++-
arch/arm64/include/asm/smp.h | 2 +
arch/arm64/kernel/Makefile | 1 +
arch/arm64/kernel/crash_dump.c | 71 ++++++++++
arch/arm64/kernel/machine_kexec.c | 67 ++++++++-
arch/arm64/kernel/setup.c | 7 +-
arch/arm64/kernel/smp.c | 63 +++++++++
arch/arm64/mm/init.c | 202 +++++++++++++++++++++++++++
include/linux/memblock.h | 1 +
mm/memblock.c | 28 ++++
15 files changed, 551 insertions(+), 7 deletions(-)
create mode 100644 arch/arm64/kernel/crash_dump.c
Couple of points
a) Just a note, while testing, the crashkernel reserved memory should be less than ARCH_LOW_ADDRESS_LIMIT (=arm64_dma_phys_limit).
I think that this is a common mistake not only for kdump, but also
for general kernels.
Since request_standard_resources() calls alloc_bootmem_low(),
the kernel will panic if any of usable "System RAM" is located
above ARCH_LOW_ADDRESS_LIMIT.
For kdump, using "crashkernel=SS" notation is a convenient way
to avoid this issue.
quoted
b) Has anyone tested this on a SoC with Gicv3 ITS ?
Should the GICD/R be reset prior to switching to crash kernel ?
I am seeing lot of GICv3: RWP timeout, gone fishing while crash kernel boots.
I've never seen this kind of messages.
I usually do my testing on a fast model.
"compatible" of interrupt-controller is "arm,gic-v3."
I suspect gic_cpu_pm_notifier is not being called on any of the cores prior to start of crash kernel.
We might have to call it explicitly.
I'm not sure that it is the cause, but anyway none of any cpu_pm_notifier's
will be called at panic. That is the reason why "maxcpus=1" should be
specified (for kdump on arm64).
What I meant was that since cpu_pm_notifier is not called before
crash kernel is started, GIC Distributor/re-distributor/ITS is not
set in quiescent state.
Which is fine, they are not expected to be in a sane state anyway
(that's what a crash is about...). The ITS now has provision to be put
in a disabled state before being reinitialized. As for GICD, it is
disabled before being reprogrammed, which should be enough.
In my setup the GICD_CTRL[RWP] bit is not cleared in the
crashkernels' distributor init function.
Which instance is failing? The initial one (just after the initial
disable)? Or the one called from gic_dist_config()?
Thanks,
M.
--
Jazz is not dead. It just smells funny...
[Cc: Marc]
On Fri, Sep 02, 2016 at 06:23:25PM +0530, Manish Jaggi wrote:
quoted
On 08/31/2016 11:01 AM, AKASHI Takahiro wrote:
quoted
Manish,
Thank you for testing my kdump and reporting issues.
On Wed, Aug 31, 2016 at 09:11:52AM +0530, Manish Jaggi wrote:
quoted
Hi Akashi,
On 08/09/2016 07:22 AM, AKASHI Takahiro wrote:
quoted
This patch series adds kdump support on arm64.
To load a crash-dump kernel to the systems, a series of patches to
kexec-tools, which have not yet been merged upstream, are needed.
Please use my kdump patches [1].
To examine vmcore (/proc/vmcore) on a crash-dump kernel, you can use
- crash utility (coming v7.1.6 or later) [2]
(Necessary patches have already been queued in the master.)
[1] T.B.D.
[2] https://github.com/crash-utility/crash.git
Changes for v24 (Aug 9, 2016):
o Rebase to Linux-4.8-rc1
o Update descriptions about newly added DT proerties
Changes for v23 (July 26, 2016):
o Move memblock_reserve() to a single place in reserve_crashkernel()
o Use cpu_park_loop() in ipi_cpu_crash_stop()
o Always enforce ARCH_LOW_ADDRESS_LIMIT to the memory range of crash kernel
o Re-implement fdt_enforce_memory_region() to remove non-reserve regions
(for ACPI) from usable memory at crash kernel
Changes for v22 (July 12, 2016):
o Export "crashkernel-base" and "crashkernel-size" via device-tree,
and add some descriptions about them in chosen.txt
o Rename "usable-memory" to "usable-memory-range" to avoid inconsistency
with powerpc's "usable-memory"
o Make cosmetic changes regarding "ifdef" usage
o Correct some wordings in kdump.txt
Changes for v21 (July 6, 2016):
o Remove kexec patches.
o Rebase to arm64's for-next/core (Linux-4.7-rc4 based).
o Clarify the description about kvm in kdump.txt.
See the following link [3] for older changes:
[3] http://lists.infradead.org/pipermail/linux-arm-kernel/2016-June/438780.html
AKASHI Takahiro (8):
arm64: kdump: reserve memory for crash dump kernel
memblock: add memblock_cap_memory_range()
arm64: limit memory regions based on DT property, usable-memory-range
arm64: kdump: implement machine_crash_shutdown()
arm64: kdump: add kdump support
arm64: kdump: add VMCOREINFO's for user-space coredump tools
arm64: kdump: enable kdump in the arm64 defconfig
arm64: kdump: update a kernel doc
James Morse (1):
Documentation: dt: chosen properties for arm64 kdump
Documentation/devicetree/bindings/chosen.txt | 45 ++++++
Documentation/kdump/kdump.txt | 16 ++-
arch/arm64/Kconfig | 11 ++
arch/arm64/configs/defconfig | 1 +
arch/arm64/include/asm/hardirq.h | 2 +-
arch/arm64/include/asm/kexec.h | 41 +++++-
arch/arm64/include/asm/smp.h | 2 +
arch/arm64/kernel/Makefile | 1 +
arch/arm64/kernel/crash_dump.c | 71 ++++++++++
arch/arm64/kernel/machine_kexec.c | 67 ++++++++-
arch/arm64/kernel/setup.c | 7 +-
arch/arm64/kernel/smp.c | 63 +++++++++
arch/arm64/mm/init.c | 202 +++++++++++++++++++++++++++
include/linux/memblock.h | 1 +
mm/memblock.c | 28 ++++
15 files changed, 551 insertions(+), 7 deletions(-)
create mode 100644 arch/arm64/kernel/crash_dump.c
Couple of points
a) Just a note, while testing, the crashkernel reserved memory should be less than ARCH_LOW_ADDRESS_LIMIT (=arm64_dma_phys_limit).
I think that this is a common mistake not only for kdump, but also
for general kernels.
Since request_standard_resources() calls alloc_bootmem_low(),
the kernel will panic if any of usable "System RAM" is located
above ARCH_LOW_ADDRESS_LIMIT.
For kdump, using "crashkernel=SS" notation is a convenient way
to avoid this issue.
quoted
b) Has anyone tested this on a SoC with Gicv3 ITS ?
Should the GICD/R be reset prior to switching to crash kernel ?
I am seeing lot of GICv3: RWP timeout, gone fishing while crash kernel boots.
I've never seen this kind of messages.
I usually do my testing on a fast model.
"compatible" of interrupt-controller is "arm,gic-v3."
I suspect gic_cpu_pm_notifier is not being called on any of the cores prior to start of crash kernel.
We might have to call it explicitly.
I'm not sure that it is the cause, but anyway none of any cpu_pm_notifier's
will be called at panic. That is the reason why "maxcpus=1" should be
specified (for kdump on arm64).
What I meant was that since cpu_pm_notifier is not called before
crash kernel is started, GIC Distributor/re-distributor/ITS is not
set in quiescent state.
Which is fine, they are not expected to be in a sane state anyway
(that's what a crash is about...). The ITS now has provision to be put
in a disabled state before being reinitialized. As for GICD, it is
disabled before being reprogrammed, which should be enough.
quoted
In my setup the GICD_CTRL[RWP] bit is not cleared in the
crashkernels' distributor init function.
Which instance is failing? The initial one (just after the initial
disable)? Or the one called from gic_dist_config()?
In crash kernel, when the GICD_CTRL is set to 0x0, RWP is not getting clear.
And is never cleared for any subsequent writes.
From: Marc Zyngier <hidden> Date: 2016-09-06 16:42:00
On 06/09/16 17:15, Manish Jaggi wrote:
quoted
quoted
In my setup the GICD_CTRL[RWP] bit is not cleared in the
crashkernels' distributor init function.
Which instance is failing? The initial one (just after the initial
disable)? Or the one called from gic_dist_config()?
In crash kernel, when the GICD_CTRL is set to 0x0, RWP is not getting clear.
And is never cleared for any subsequent writes.
That's weird. It means writes are still pending, and never drained. What
happens if you put a dsb(sy) in the wait_for_rwp() loop?
Thanks,
M.
--
Jazz is not dead. It just smells funny...