The crash memory allocation, and the exclude of crashk_res, crashk_low_res
and crashk_cma memory are almost identical across different architectures,
This patch set handle them in crash core in a general way, which eliminate
a lot of duplication code.
And add support for crashkernel CMA reservation for arm64 and riscv.
Rebased on v7.1-rc1.
Basic second kernel boot test were performed on QEMU platforms for x86,
ARM64 and RISC-V architectures with the following parameters:
"cma=256M crashkernel=4G crashkernel=64M,cma"
For first kernel, there will be such log:
# dmesg | grep crash
[ 0.000000] crashkernel low memory reserved: 0xe8000000 - 0xf0000000 (128 MB)
[ 0.000000] crashkernel reserved: 0x000000023e600000 - 0x000000033e600000 (4096 MB)
[ 0.000000] crashkernel CMA reserved: 64 MB in 1 ranges
# dmesg | grep cma
[ 0.000000] cma: Reserved 256 MiB at 0x00000000f0000000
[ 0.000000] cma: Reserved 64 MiB at 0x0000000100000000
For second kernel, there will be such log:
[ 0.000000] OF: fdt: Looking for usable-memory-range property...
[ 0.000000] OF: fdt: cap_mem_regions[0]: base=0x000000023e600000, size=0x0000000100000000
[ 0.000000] OF: fdt: cap_mem_regions[1]: base=0x00000000e8000000, size=0x0000000008000000
[ 0.000000] OF: fdt: cap_mem_regions[2]: base=0x0000000100000000, size=0x0000000004000000
Changes in v13:
- Rebased on v7.1-rc1.
- Update the commit message.
- Add Reviewed-by.
- Link to v12: https://lore.kernel.org/all/20260402072701.628293-1-ruanjinjie@huawei.com/
Changes in v12:
- Remove the unused "nr_mem_ranges" for x86.
- Add "Fix crashk_low_res not exclude bug" test log.
- Provide a separate patch for each architecture for using
crash_prepare_headers(), which will make the review more convenient.
- Add Reviewed-by and Tested-by.
- Link to v11: https://lore.kernel.org/all/20260328074013.3589544-1-ruanjinjie@huawei.com/
Changes in v11:
- Avoid silently drop crash memory if the crash kernel is built without
CONFIG_CMA.
- Remove unnecessary "cmem->nr_ranges = 0" for arch_crash_populate_cmem()
as we use kvzalloc().
- Provide a separate patch for each architecture to fix the existing
buffer overflow issue.
- Add Acked-bys for arm64.
Changes in v10:
- Fix crashk_low_res not excluded bug in the existing
RISC-V code.
- Fix an existing memory leak issue in the existing PowerPC code.
- Fix the ordering issue of adding CMA ranges to
"linux,usable-memory-range".
- Fix an existing concurrency issue. A Concurrent memory hotplug may occur
between reading memblock and attempting to fill cmem during kexec_load()
for almost all existing architectures.
- Link to v9: https://lore.kernel.org/all/20260323072745.2481719-1-ruanjinjie@huawei.com/
Changes in v9:
- Collect Reviewed-by and Acked-by, and prepare for Sashiko AI review.
- Link to v8: https://lore.kernel.org/all/20260302035315.3892241-1-ruanjinjie@huawei.com/
Changes in v8:
- Fix the build issues reported by kernel test robot and Sourabh.
- Link to v7: https://lore.kernel.org/all/20260226130437.1867658-1-ruanjinjie@huawei.com/
Changes in v7:
- Correct the inclusion of CMA-reserved ranges for kdump kernel in of/kexec
for arm64 and riscv.
- Add Acked-by.
- Link to v6: https://lore.kernel.org/all/20260224085342.387996-1-ruanjinjie@huawei.com/
Changes in v6:
- Update the crash core exclude code as Mike suggested.
- Rebased on v7.0-rc1.
- Add acked-by.
- Link to v5: https://lore.kernel.org/all/20260212101001.343158-1-ruanjinjie@huawei.com/
Jinjie Ruan (14):
riscv: kexec_file: Fix crashk_low_res not exclude bug
powerpc/crash: Fix possible memory leak in update_crash_elfcorehdr()
x86/kexec: Fix potential buffer overflow in prepare_elf_headers()
arm64: kexec_file: Fix potential buffer overflow in
prepare_elf_headers()
riscv: kexec_file: Fix potential buffer overflow in
prepare_elf_headers()
LoongArch: kexec: Fix potential buffer overflow in
prepare_elf_headers()
crash: Add crash_prepare_headers() to exclude crash kernel memory
arm64: kexec_file: Use crash_prepare_headers() helper to simplify code
x86/kexec: Use crash_prepare_headers() helper to simplify code
riscv: kexec_file: Use crash_prepare_headers() helper to simplify code
LoongArch: kexec: Use crash_prepare_headers() helper to simplify code
crash: Use crash_exclude_core_ranges() on powerpc
arm64: kexec: Add support for crashkernel CMA reservation
riscv: kexec: Add support for crashkernel CMA reservation
Sourabh Jain (1):
powerpc/crash: sort crash memory ranges before preparing elfcorehdr
.../admin-guide/kernel-parameters.txt | 16 +--
arch/arm64/kernel/machine_kexec_file.c | 43 +++-----
arch/arm64/mm/init.c | 5 +-
arch/loongarch/kernel/machine_kexec_file.c | 43 +++-----
arch/powerpc/include/asm/kexec_ranges.h | 1 -
arch/powerpc/kexec/crash.c | 7 +-
arch/powerpc/kexec/ranges.c | 101 +-----------------
arch/riscv/kernel/machine_kexec_file.c | 42 +++-----
arch/riscv/mm/init.c | 5 +-
arch/x86/kernel/crash.c | 92 +++-------------
drivers/of/fdt.c | 9 +-
drivers/of/kexec.c | 9 ++
include/linux/crash_core.h | 9 ++
include/linux/crash_reserve.h | 4 +-
kernel/crash_core.c | 98 ++++++++++++++++-
15 files changed, 202 insertions(+), 282 deletions(-)
--
2.34.1
As done in commit 944a45abfabc ("arm64: kdump: Reimplement crashkernel=X")
and commit 4831be702b95 ("arm64/kexec: Fix missing extra range for
crashkres_low.") for arm64, while implementing crashkernel=X,[high,low],
riscv should have excluded the "crashk_low_res" reserved ranges from
the crash kernel memory to prevent them from being exported through
/proc/vmcore, and the exclusion would need an extra crash_mem range.
Just simply tested on qemu with crashkernel=4G with kexec in [1] mentioned
in [2]. And the second kernel can be started normally.
# dmesg | grep crash
[ 0.000000] crashkernel low memory reserved: 0xf8000000 - 0x100000000 (128 MB)
[ 0.000000] crashkernel reserved: 0x000000017fe00000 - 0x000000027fe00000 (4096 MB)
Cc: Guo Ren <guoren@kernel.org>
Cc: Baoquan He <redacted>
[1]: https://github.com/chenjh005/kexec-tools/tree/build-test-riscv-v2
[2]: https://lore.kernel.org/all/20230726175000.2536220-1-chenjiahao16@huawei.com/
Fixes: 5882e5acf18d ("riscv: kdump: Implement crashkernel=X,[high,low]")
Reviewed-by: Guo Ren <guoren@kernel.org>
Signed-off-by: Jinjie Ruan <redacted>
---
arch/riscv/kernel/machine_kexec_file.c | 14 +++++++++++---
1 file changed, 11 insertions(+), 3 deletions(-)
@@ -61,7 +61,7 @@ static int prepare_elf_headers(void **addr, unsigned long *sz)unsignedintnr_ranges;intret;-nr_ranges=1;/* For exclusion of crashkernel region */+nr_ranges=2;/* For exclusion of crashkernel region */walk_system_ram_res(0,-1,&nr_ranges,get_nr_ram_ranges_callback);cmem=kmalloc_flex(*cmem,ranges,nr_ranges);
@@ -76,8 +76,16 @@ static int prepare_elf_headers(void **addr, unsigned long *sz)/* Exclude crashkernel region */ret=crash_exclude_mem_range(cmem,crashk_res.start,crashk_res.end);-if(!ret)-ret=crash_prepare_elf64_headers(cmem,true,addr,sz);+if(ret)+gotoout;++if(crashk_low_res.end){+ret=crash_exclude_mem_range(cmem,crashk_low_res.start,crashk_low_res.end);+if(ret)+gotoout;+}++ret=crash_prepare_elf64_headers(cmem,true,addr,sz);out:kfree(cmem);
In get_crash_memory_ranges(), if crash_exclude_mem_range() failed
after realloc_mem_ranges() has successfully allocated the cmem
memory, it just returns an error but leaves cmem pointing to
the allocated memory, nor is it freed in the caller
update_crash_elfcorehdr(), which cause a memory leak, goto out
to free the cmem.
Cc: Sourabh Jain <redacted>
Cc: Hari Bathini <hbathini@linux.ibm.com>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Fixes: 849599b702ef ("powerpc/crash: add crash memory hotplug support")
Reviewed-by: Sourabh Jain <redacted>
Signed-off-by: Jinjie Ruan <redacted>
---
arch/powerpc/kexec/crash.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
There is a race condition between the kexec_load() system call
(crash kernel loading path) and memory hotplug operations that can lead
to buffer overflow and potential kernel crash.
During prepare_elf_headers(), the following steps occur:
1. get_nr_ram_ranges_callback() queries current System RAM memory ranges
2. Allocates buffer based on queried count
3. prepare_elf64_ram_headers_callback() populates ranges from memblock
If memory hotplug occurs between step 1 and step 3, the number of ranges
can increase, causing out-of-bounds write when populating cmem->ranges[].
This happens because kexec_load() uses kexec_trylock (atomic_t) while
memory hotplug uses device_hotplug_lock (mutex), so they don't serialize
with each other.
Since x86 supports crash hotplug, any data inconsistency caused by
a race during the initial load will be corrected by the subsequent
hotplug update. However, we must prevent a buffer overflow if the
number of memory regions increases between the two passes.
Add a boundary checking in prepare_elf64_ram_headers_callback() to ensure
that the number of populated ranges does not exceed
the allocated maximum.
Cc: Thomas Gleixner <tglx@kernel.org>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Borislav Petkov <bp@alien8.de>
Cc: "H. Peter Anvin" <hpa@zytor.com>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Baoquan He <redacted>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: stable@vger.kernel.org
Fixes: 8d5f894a3108 ("x86: kexec_file: lift CRASH_MAX_RANGES limit on crash_mem buffer")
Signed-off-by: Jinjie Ruan <redacted>
---
arch/x86/kernel/crash.c | 3 +++
1 file changed, 3 insertions(+)
There is a race condition between the kexec_load() system call
(crash kernel loading path) and memory hotplug operations that can
lead to buffer overflow and potential kernel crash.
During prepare_elf_headers(), the following steps occur:
1. The first for_each_mem_range() queries current System RAM memory ranges
2. Allocates buffer based on queried count
3. The 2st for_each_mem_range() populates ranges from memblock
If memory hotplug occurs between step 1 and step 3, the number of ranges
can increase, causing out-of-bounds write when populating cmem->ranges[].
This happens because kexec_load() uses kexec_trylock (atomic_t) while
memory hotplug uses device_hotplug_lock (mutex), so they don't serialize
with each other.
Add the explicit bounds checking to prevent out-of-bounds access.
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Will Deacon <redacted>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Baoquan He <redacted>
Cc: Breno Leitao <leitao@debian.org>
Cc: stable@vger.kernel.org
Fixes: 3751e728cef2 ("arm64: kexec_file: add crash dump support")
Closes: https://sashiko.dev/#/patchset/20260323072745.2481719-1-ruanjinjie%40huawei.com
Signed-off-by: Jinjie Ruan <redacted>
---
arch/arm64/kernel/machine_kexec_file.c | 5 +++++
1 file changed, 5 insertions(+)
There is a race condition between the kexec_load() system call
(crash kernel loading path) and memory hotplug operations that can lead
to buffer overflow and potential kernel crash.
During prepare_elf_headers(), the following steps occur:
1. get_nr_ram_ranges_callback() queries current System RAM memory ranges
2. Allocates buffer based on queried count
3. prepare_elf64_ram_headers_callback() populates ranges from memblock
If memory hotplug occurs between step 1 and step 3, the number of ranges
can increase, causing out-of-bounds write when populating cmem->ranges[].
This happens because kexec_load() uses kexec_trylock (atomic_t) while
memory hotplug uses device_hotplug_lock (mutex), so they don't serialize
with each other.
While this works today because RISC-V server hardware with hotplug
support is still rare and most deployments use fixed memory configurations
(e.g., QEMU virt machine), it is technically fragile. So add bounds
checking in prepare_elf64_ram_headers_callback() to prevent
out-of-bounds (OOB) access.
No functional change for current RISC-V deployments, but makes
the code robust against future hotplug-capable platforms.
Cc: Paul Walmsley <pjw@kernel.org>
Cc: Palmer Dabbelt <palmer@dabbelt.com>
Cc: Albert Ou <aou@eecs.berkeley.edu>
Cc: Alexandre Ghiti <alex@ghiti.fr>
Cc: songshuaishuai@tinylab.org
Cc: bjorn@rivosinc.com
Cc: leitao@debian.org
Fixes: 8acea455fafa ("RISC-V: Support for kexec_file on panic")
Reviewed-by: Guo Ren <guoren@kernel.org>
Signed-off-by: Jinjie Ruan <redacted>
---
arch/riscv/kernel/machine_kexec_file.c | 3 +++
1 file changed, 3 insertions(+)
The crash memory alloc, and the exclude of crashk_res, crashk_low_res
and crashk_cma memory are almost identical across different architectures,
handling them in the crash core would eliminate a lot of duplication, so
add crash_prepare_headers() helper to handle them in the common code.
To achieve the above goal, three architecture-specific functions are
introduced:
- arch_get_system_nr_ranges(). Pre-counts the max number of memory ranges.
- arch_crash_populate_cmem(). Collects the memory ranges and fills them
into cmem.
- arch_crash_exclude_ranges(). Architecture's additional crash memory
ranges exclusion, defaulting to empty.
Reviewed-by: Sourabh Jain <redacted>
Acked-by: Baoquan He <redacted>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
Signed-off-by: Jinjie Ruan <redacted>
---
include/linux/crash_core.h | 5 +++
kernel/crash_core.c | 91 ++++++++++++++++++++++++++++++++++++--
2 files changed, 93 insertions(+), 3 deletions(-)
@@ -272,6 +270,93 @@ int crash_prepare_elf64_headers(struct crash_mem *mem, int need_kernel_map,return0;}+staticstructcrash_mem*alloc_cmem(unsignedintnr_ranges)+{+structcrash_mem*cmem;++cmem=kvzalloc_flex(*cmem,ranges,nr_ranges);+if(!cmem)+returnNULL;++cmem->max_nr_ranges=nr_ranges;+returncmem;+}++unsignedint__weakarch_get_system_nr_ranges(void){return0;}+int__weakarch_crash_populate_cmem(structcrash_mem*cmem){return-1;}+int__weakarch_crash_exclude_ranges(structcrash_mem*cmem){return0;}++staticintcrash_exclude_core_ranges(structcrash_mem*cmem)+{+intret,i;++/* Exclude crashkernel region */+ret=crash_exclude_mem_range(cmem,crashk_res.start,crashk_res.end);+if(ret)+returnret;++if(crashk_low_res.end){+ret=crash_exclude_mem_range(cmem,crashk_low_res.start,crashk_low_res.end);+if(ret)+returnret;+}++for(i=0;i<crashk_cma_cnt;++i){+ret=crash_exclude_mem_range(cmem,crashk_cma_ranges[i].start,+crashk_cma_ranges[i].end);+if(ret)+returnret;+}++return0;+}++intcrash_prepare_headers(intneed_kernel_map,void**addr,unsignedlong*sz,+unsignedlong*nr_mem_ranges)+{+unsignedintmax_nr_ranges;+structcrash_mem*cmem;+intret;++get_online_mems();+max_nr_ranges=arch_get_system_nr_ranges();+if(!max_nr_ranges){+put_online_mems();+return-ENOMEM;+}++cmem=alloc_cmem(max_nr_ranges);+if(!cmem){+put_online_mems();+return-ENOMEM;+}++ret=arch_crash_populate_cmem(cmem);+if(ret){+put_online_mems();+gotoout;+}++put_online_mems();+ret=crash_exclude_core_ranges(cmem);+if(ret)+gotoout;++ret=arch_crash_exclude_ranges(cmem);+if(ret)+gotoout;++/* Return the computed number of memory ranges, for hotplug usage */+if(nr_mem_ranges)+*nr_mem_ranges=cmem->nr_ranges;++ret=crash_prepare_elf64_headers(cmem,need_kernel_map,addr,sz);++out:+kvfree(cmem);+returnret;+}+/***crash_exclude_mem_range-excludeamemrangeforexistingranges*@mem:mem->rangecontainsanarrayofrangessortedinascendingorder
There is a race condition between the kexec_load() system call
(crash kernel loading path) and memory hotplug operations that can lead
to buffer overflow and potential kernel crash.
During prepare_elf_headers(), the following steps occur:
1. The first for_each_mem_range() queries current System RAM memory ranges
2. Allocates buffer based on queried count
3. The 2st for_each_mem_range() populates ranges from memblock
If memory hotplug occurs between step 1 and step 3, the number of ranges
can increase, causing out-of-bounds write when populating cmem->ranges[].
This happens because kexec_load() uses kexec_trylock (atomic_t) while
memory hotplug uses device_hotplug_lock (mutex), so they don't serialize
with each other.
Just add bounds checking to prevent out-of-bounds access.
Cc: Youling Tang <redacted>
Cc: Huacai Chen <redacted>
Cc: WANG Xuerui <kernel@xen0n.name>
Cc: stable@vger.kernel.org
Fixes: 1bcca8620a91 ("LoongArch: Add crash dump support for kexec_file")
Signed-off-by: Jinjie Ruan <redacted>
---
arch/loongarch/kernel/machine_kexec_file.c | 5 +++++
1 file changed, 5 insertions(+)
Use the newly introduced crash_prepare_headers() function to replace
the existing prepare_elf_headers(), allocate cmem and exclude crash kernel
memory in the crash core, which reduce code duplication.
Only the following three architecture functions need to be implemented:
- arch_get_system_nr_ranges(). Call get_nr_ram_ranges_callback()
to pre-count the max number of memory ranges.
- arch_crash_populate_cmem(). Use prepare_elf64_ram_headers_callback()
to collect the memory ranges and fills them into cmem.
- arch_crash_exclude_ranges(). Exclude the low 1M for x86.
By the way, remove the unused "nr_mem_ranges" in
arch_crash_handle_hotplug_event().
Cc: Thomas Gleixner <tglx@kernel.org>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Borislav Petkov <bp@alien8.de>
Cc: Dave Hansen <dave.hansen@linux.intel.com>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Vivek Goyal <vgoyal@redhat.com>
Reviewed-by: Sourabh Jain <redacted>
Acked-by: Baoquan He <redacted>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
Signed-off-by: Jinjie Ruan <redacted>
---
arch/x86/kernel/crash.c | 89 +++++------------------------------------
1 file changed, 11 insertions(+), 78 deletions(-)
@@ -153,16 +153,8 @@ static int get_nr_ram_ranges_callback(struct resource *res, void *arg)return0;}-/* Gather all the required information to prepare elf headers for ram regions */-staticstructcrash_mem*fill_up_crash_elf_data(void)+unsignedintarch_get_system_nr_ranges(void){-unsignedintnr_ranges=0;-structcrash_mem*cmem;--walk_system_ram_res(0,-1,&nr_ranges,get_nr_ram_ranges_callback);-if(!nr_ranges)-returnNULL;-/**Exclusionofcrashregion,crashk_low_resand/orcrashk_cma_ranges*maycauserangesplits.Soaddextraslotshere.
@@ -177,49 +169,16 @@ static struct crash_mem *fill_up_crash_elf_data(void)*Butinordertolestthelow1Mcouldbechangedinthefuture,*(e.g.[start,1M]),addaextraslot.*/-nr_ranges+=3+crashk_cma_cnt;-cmem=vzalloc(struct_size(cmem,ranges,nr_ranges));-if(!cmem)-returnNULL;--cmem->max_nr_ranges=nr_ranges;+unsignedintnr_ranges=3+crashk_cma_cnt;-returncmem;+walk_system_ram_res(0,-1,&nr_ranges,get_nr_ram_ranges_callback);+returnnr_ranges;}-/*-*Lookforanyunwantedrangesbetweenmstart,mendandremovethem.This-*mightleadtosplitandsplitrangesareputincmem->ranges[]array-*/-staticintelf_header_exclude_ranges(structcrash_mem*cmem)+intarch_crash_exclude_ranges(structcrash_mem*cmem){-intret=0;-inti;-/* Exclude the low 1M because it is always reserved */-ret=crash_exclude_mem_range(cmem,0,SZ_1M-1);-if(ret)-returnret;--/* Exclude crashkernel region */-ret=crash_exclude_mem_range(cmem,crashk_res.start,crashk_res.end);-if(ret)-returnret;--if(crashk_low_res.end)-ret=crash_exclude_mem_range(cmem,crashk_low_res.start,-crashk_low_res.end);-if(ret)-returnret;--for(i=0;i<crashk_cma_cnt;++i){-ret=crash_exclude_mem_range(cmem,crashk_cma_ranges[i].start,-crashk_cma_ranges[i].end);-if(ret)-returnret;-}--return0;+returncrash_exclude_mem_range(cmem,0,SZ_1M-1);}staticintprepare_elf64_ram_headers_callback(structresource*res,void*arg)
@@ -236,35 +195,9 @@ static int prepare_elf64_ram_headers_callback(struct resource *res, void *arg)return0;}-/* Prepare elf headers. Return addr and size */-staticintprepare_elf_headers(void**addr,unsignedlong*sz,-unsignedlong*nr_mem_ranges)+intarch_crash_populate_cmem(structcrash_mem*cmem){-structcrash_mem*cmem;-intret;--cmem=fill_up_crash_elf_data();-if(!cmem)-return-ENOMEM;--ret=walk_system_ram_res(0,-1,cmem,prepare_elf64_ram_headers_callback);-if(ret)-gotoout;--/* Exclude unwanted mem ranges */-ret=elf_header_exclude_ranges(cmem);-if(ret)-gotoout;--/* Return the computed number of memory ranges, for hotplug usage */-*nr_mem_ranges=cmem->nr_ranges;--/* By default prepare 64bit headers */-ret=crash_prepare_elf64_headers(cmem,IS_ENABLED(CONFIG_X86_64),addr,sz);--out:-vfree(cmem);-returnret;+returnwalk_system_ram_res(0,-1,cmem,prepare_elf64_ram_headers_callback);}#endif
@@ -422,7 +355,8 @@ int crash_load_segments(struct kimage *image).buf_max=ULONG_MAX,.top_down=false};/* Prepare elf headers and add a segment */-ret=prepare_elf_headers(&kbuf.buffer,&kbuf.bufsz,&pnum);+ret=crash_prepare_headers(IS_ENABLED(CONFIG_X86_64),&kbuf.buffer,+&kbuf.bufsz,&pnum);if(ret)returnret;
@@ -515,7 +449,6 @@ unsigned int arch_crash_get_elfcorehdr_size(void)voidarch_crash_handle_hotplug_event(structkimage*image,void*arg){void*elfbuf=NULL,*old_elfcorehdr;-unsignedlongnr_mem_ranges;unsignedlongmem,memsz;unsignedlongelfsz=0;
@@ -533,7 +466,7 @@ void arch_crash_handle_hotplug_event(struct kimage *image, void *arg)*CreatethenewelfcorehdrreflectingthechangestoCPUand/or*memoryresources.*/-if(prepare_elf_headers(&elfbuf,&elfsz,&nr_mem_ranges)){+if(crash_prepare_headers(IS_ENABLED(CONFIG_X86_64),&elfbuf,&elfsz,NULL)){pr_err("unable to create new elfcorehdr");gotoout;}
From: Sourabh Jain <redacted>
During a memory hot-remove event, the elfcorehdr is rebuilt to exclude
the removed memory. While updating the crash memory ranges for this
operation, the crash memory ranges array can become unsorted. This
happens because remove_mem_range() may split a memory range into two
parts and append the higher-address part as a separate range at the end
of the array.
So far, no issues have been observed due to the unsorted crash memory
ranges. However, this could lead to problems once crash memory range
removal is handled by generic code, as introduced in the upcoming
patches in this series.
Currently, powerpc uses a platform-specific function,
remove_mem_range(), to exclude hot-removed memory from the crash memory
ranges. This function performs the same task as the generic
crash_exclude_mem_range() in crash_core.c. The generic helper also
ensures that the crash memory ranges remain sorted. So remove the
redundant powerpc-specific implementation and instead call
crash_exclude_mem_range_guarded() (which internally calls
crash_exclude_mem_range()) to exclude the hot-removed memory ranges.
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Baoquan he <redacted>
Cc: Jinjie Ruan <redacted>
Cc: Hari Bathini <hbathini@linux.ibm.com>
Cc: Madhavan Srinivasan <maddy@linux.ibm.com>
Cc: Mahesh Salgaonkar <mahesh@linux.ibm.com>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Ritesh Harjani (IBM) <ritesh.list@gmail.com>
Cc: Shivang Upadhyay <redacted>
Cc: linux-kernel@vger.kernel.org
Acked-by: Baoquan He <redacted>
Reviewed-by: Ritesh Harjani (IBM) <ritesh.list@gmail.com>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
Signed-off-by: Sourabh Jain <redacted>
Signed-off-by: Jinjie Ruan <redacted>
---
arch/powerpc/include/asm/kexec_ranges.h | 4 +-
arch/powerpc/kexec/crash.c | 5 +-
arch/powerpc/kexec/ranges.c | 87 +------------------------
3 files changed, 7 insertions(+), 89 deletions(-)
Use the newly introduced crash_prepare_headers() function to replace
the existing prepare_elf_headers(), allocate cmem and exclude crash kernel
memory in the crash core, which reduce code duplication.
Only the following two architecture functions need to be implemented:
- arch_get_system_nr_ranges(). Call get_nr_ram_ranges_callback()
to pre-counts the max number of memory ranges.
- arch_crash_populate_cmem(). Use prepare_elf64_ram_headers_callback()
to collects the memory ranges and fills them into cmem.
Cc: Paul Walmsley <pjw@kernel.org>
Cc: Palmer Dabbelt <palmer@dabbelt.com>
Cc: Albert Ou <aou@eecs.berkeley.edu>
Cc: Alexandre Ghiti <alex@ghiti.fr>
Cc: Guo Ren <guoren@kernel.org>
Reviewed-by: Sourabh Jain <redacted>
Acked-by: Baoquan He <redacted>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
Signed-off-by: Jinjie Ruan <redacted>
---
arch/riscv/kernel/machine_kexec_file.c | 47 +++++++-------------------
1 file changed, 12 insertions(+), 35 deletions(-)
@@ -44,6 +44,15 @@ static int get_nr_ram_ranges_callback(struct resource *res, void *arg)return0;}+unsignedintarch_get_system_nr_ranges(void)+{+unsignedintnr_ranges=2;/* For exclusion of crashkernel region */++walk_system_ram_res(0,-1,&nr_ranges,get_nr_ram_ranges_callback);++returnnr_ranges;+}+staticintprepare_elf64_ram_headers_callback(structresource*res,void*arg){structcrash_mem*cmem=arg;
@@ -58,41 +67,9 @@ static int prepare_elf64_ram_headers_callback(struct resource *res, void *arg)return0;}-staticintprepare_elf_headers(void**addr,unsignedlong*sz)+intarch_crash_populate_cmem(structcrash_mem*cmem){-structcrash_mem*cmem;-unsignedintnr_ranges;-intret;--nr_ranges=2;/* For exclusion of crashkernel region */-walk_system_ram_res(0,-1,&nr_ranges,get_nr_ram_ranges_callback);--cmem=kmalloc_flex(*cmem,ranges,nr_ranges);-if(!cmem)-return-ENOMEM;--cmem->max_nr_ranges=nr_ranges;-cmem->nr_ranges=0;-ret=walk_system_ram_res(0,-1,cmem,prepare_elf64_ram_headers_callback);-if(ret)-gotoout;--/* Exclude crashkernel region */-ret=crash_exclude_mem_range(cmem,crashk_res.start,crashk_res.end);-if(ret)-gotoout;--if(crashk_low_res.end){-ret=crash_exclude_mem_range(cmem,crashk_low_res.start,crashk_low_res.end);-if(ret)-gotoout;-}--ret=crash_prepare_elf64_headers(cmem,true,addr,sz);--out:-kfree(cmem);-returnret;+returnwalk_system_ram_res(0,-1,cmem,prepare_elf64_ram_headers_callback);}staticchar*setup_kdump_cmdline(structkimage*image,char*cmdline,
@@ -284,7 +261,7 @@ int load_extra_segments(struct kimage *image, unsigned long kernel_start,if(image->type==KEXEC_TYPE_CRASH){void*headers;unsignedlongheaders_sz;-ret=prepare_elf_headers(&headers,&headers_sz);+ret=crash_prepare_headers(true,&headers,&headers_sz,NULL);if(ret){pr_err("Preparing elf core header failed\n");gotoout;
Use the newly introduced crash_prepare_headers() function to replace
the existing prepare_elf_headers(), allocate cmem and exclude crash
kernel memory in the crash core, which reduce code duplication.
Only the following two architecture functions need to be implemented:
- arch_get_system_nr_ranges(). Use for_each_mem_range() to traverse
and pre-count the max number of memory ranges.
- arch_crash_populate_cmem(). Use for_each_mem_range to traverse
and collect the memory ranges and fills them into cmem.
Acked-by: Catalin Marinas <catalin.marinas@arm.com>
Reviewed-by: Sourabh Jain <redacted>
Acked-by: Baoquan He <redacted>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
Signed-off-by: Jinjie Ruan <redacted>
---
arch/arm64/kernel/machine_kexec_file.c | 46 ++++++++------------------
1 file changed, 14 insertions(+), 32 deletions(-)
@@ -40,51 +40,33 @@ int arch_kimage_file_post_load_cleanup(struct kimage *image)}#ifdef CONFIG_CRASH_DUMP-staticintprepare_elf_headers(void**addr,unsignedlong*sz)+unsignedintarch_get_system_nr_ranges(void){-structcrash_mem*cmem;-unsignedintnr_ranges;-intret;-u64i;+unsignedintnr_ranges=2;/* for exclusion of crashkernel region */phys_addr_tstart,end;+u64i;-nr_ranges=2;/* for exclusion of crashkernel region */for_each_mem_range(i,&start,&end)nr_ranges++;-cmem=kmalloc_flex(*cmem,ranges,nr_ranges);-if(!cmem)-return-ENOMEM;+returnnr_ranges;+}++intarch_crash_populate_cmem(structcrash_mem*cmem)+{+phys_addr_tstart,end;+u64i;-cmem->max_nr_ranges=nr_ranges;-cmem->nr_ranges=0;for_each_mem_range(i,&start,&end){-if(cmem->nr_ranges>=cmem->max_nr_ranges){-ret=-ENOMEM;-gotoout;-}+if(cmem->nr_ranges>=cmem->max_nr_ranges)+return-ENOMEM;cmem->ranges[cmem->nr_ranges].start=start;cmem->ranges[cmem->nr_ranges].end=end-1;cmem->nr_ranges++;}-/* Exclude crashkernel region */-ret=crash_exclude_mem_range(cmem,crashk_res.start,crashk_res.end);-if(ret)-gotoout;--if(crashk_low_res.end){-ret=crash_exclude_mem_range(cmem,crashk_low_res.start,crashk_low_res.end);-if(ret)-gotoout;-}--ret=crash_prepare_elf64_headers(cmem,true,addr,sz);--out:-kfree(cmem);-returnret;+return0;}#endif
Use the newly introduced crash_prepare_headers() function to replace
the existing prepare_elf_headers(), allocate cmem and exclude crash kernel
memory in the crash core, which reduce code duplication.
Only the following two architecture functions need to be implemented:
- arch_get_system_nr_ranges(). Use for_each_mem_range to traverse
and pre-count the max number of memory ranges.
- arch_crash_populate_cmem(). Use for_each_mem_range to traverse
and collect the memory ranges and fills them into cmem.
Cc: Huacai Chen <chenhuacai@kernel.org>
Cc: WANG Xuerui <kernel@xen0n.name>
Cc: Youling Tang <redacted>
Cc: Baoquan He <redacted>
Reviewed-by: Sourabh Jain <redacted>
Acked-by: Baoquan He <redacted>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
Signed-off-by: Jinjie Ruan <redacted>
---
arch/loongarch/kernel/machine_kexec_file.c | 46 +++++++---------------
1 file changed, 14 insertions(+), 32 deletions(-)
@@ -56,51 +56,33 @@ static void cmdline_add_initrd(struct kimage *image, unsigned long *cmdline_tmpl}#ifdef CONFIG_CRASH_DUMP--staticintprepare_elf_headers(void**addr,unsignedlong*sz)+unsignedintarch_get_system_nr_ranges(void){-intret,nr_ranges;-uint64_ti;+intnr_ranges=2;/* for exclusion of crashkernel region */phys_addr_tstart,end;-structcrash_mem*cmem;+uint64_ti;-nr_ranges=2;/* for exclusion of crashkernel region */for_each_mem_range(i,&start,&end)nr_ranges++;-cmem=kmalloc_flex(*cmem,ranges,nr_ranges);-if(!cmem)-return-ENOMEM;+returnnr_ranges;+}++intarch_crash_populate_cmem(structcrash_mem*cmem)+{+phys_addr_tstart,end;+uint64_ti;-cmem->max_nr_ranges=nr_ranges;-cmem->nr_ranges=0;for_each_mem_range(i,&start,&end){-if(cmem->nr_ranges>=cmem->max_nr_ranges){-ret=-ENOMEM;-gotoout;-}+if(cmem->nr_ranges>=cmem->max_nr_ranges)+return-ENOMEM;cmem->ranges[cmem->nr_ranges].start=start;cmem->ranges[cmem->nr_ranges].end=end-1;cmem->nr_ranges++;}-/* Exclude crashkernel region */-ret=crash_exclude_mem_range(cmem,crashk_res.start,crashk_res.end);-if(ret<0)-gotoout;--if(crashk_low_res.end){-ret=crash_exclude_mem_range(cmem,crashk_low_res.start,crashk_low_res.end);-if(ret<0)-gotoout;-}--ret=crash_prepare_elf64_headers(cmem,true,addr,sz);--out:-kfree(cmem);-returnret;+return0;}/*
The crash memory exclude of crashk_res and crashk_cma memory on powerpc
are almost identical to the generic crash_exclude_core_ranges().
By introducing the architecture-specific arch_crash_exclude_mem_range()
function with a default implementation of crash_exclude_mem_range(),
and using crash_exclude_mem_range_guarded as powerpc's separate
implementation, the generic crash_exclude_core_ranges() helper function
can be reused.
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Hari Bathini <hbathini@linux.ibm.com>
Cc: Madhavan Srinivasan <maddy@linux.ibm.com>
Cc: Mahesh Salgaonkar <mahesh@linux.ibm.com>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Ritesh Harjani (IBM) <ritesh.list@gmail.com>
Cc: Shivang Upadhyay <redacted>
Acked-by: Baoquan He <redacted>
Reviewed-by: Sourabh Jain <redacted>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
Signed-off-by: Jinjie Ruan <redacted>
---
arch/powerpc/include/asm/kexec_ranges.h | 3 ---
arch/powerpc/kexec/crash.c | 2 +-
arch/powerpc/kexec/ranges.c | 16 ++++------------
include/linux/crash_core.h | 4 ++++
kernel/crash_core.c | 19 +++++++++++++------
5 files changed, 22 insertions(+), 22 deletions(-)
Commit 35c18f2933c5 ("Add a new optional ",cma" suffix to the
crashkernel= command line option") and commit ab475510e042 ("kdump:
implement reserve_crashkernel_cma") added CMA support for kdump
crashkernel reservation. This allows the kernel to dynamically allocate
contiguous memory for crash dumping when needed, rather than permanently
reserving a fixed region at boot time.
So extend crashkernel CMA reservation support to riscv. The following
changes are made to enable CMA reservation:
- Parse and obtain the CMA reservation size along with other crashkernel
parameters.
- Call reserve_crashkernel_cma() to allocate the CMA region for kdump.
- Include the CMA-reserved ranges for kdump kernel to use, which was
already done in of_kexec_alloc_and_setup_fdt().
- Exclude the CMA-reserved ranges from the crash kernel memory to
prevent them from being exported through /proc/vmcore, which was
already done in the crash core.
Update kernel-parameters.txt to document CMA support for crashkernel on
riscv architecture.
Cc: Paul Walmsley <pjw@kernel.org>
Cc: Palmer Dabbelt <palmer@dabbelt.com>
Cc: Albert Ou <aou@eecs.berkeley.edu>
Cc: Alexandre Ghiti <alex@ghiti.fr>
Acked-by: Baoquan He <redacted>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
Acked-by: Paul Walmsley <pjw@kernel.org> # arch/riscv
Signed-off-by: Jinjie Ruan <redacted>
---
Documentation/admin-guide/kernel-parameters.txt | 16 ++++++++--------
arch/riscv/kernel/machine_kexec_file.c | 2 +-
arch/riscv/mm/init.c | 5 +++--
3 files changed, 12 insertions(+), 11 deletions(-)
@@ -1119,14 +1119,14 @@ Kernel parameters It will be ignored when crashkernel=X,high is not used or memory reserved is below 4G. crashkernel=size[KMG],cma- [KNL, X86, ARM64, PPC] Reserve additional crash kernel memory from- CMA. This reservation is usable by the first system's- userspace memory and kernel movable allocations (memory- balloon, zswap). Pages allocated from this memory range- will not be included in the vmcore so this should not- be used if dumping of userspace memory is intended and- it has to be expected that some movable kernel pages- may be missing from the dump.+ [KNL, X86, ARM64, RISCV, PPC] Reserve additional crash+ kernel memory from CMA. This reservation is usable by+ the first system's userspace memory and kernel movable+ allocations (memory balloon, zswap). Pages allocated+ from this memory range will not be included in the vmcore+ so this should not be used if dumping of userspace memory+ is intended and it has to be expected that some movable+ kernel pages may be missing from the dump. A standard crashkernel reservation, as described above, is still needed to hold the crash kernel and initrd.
@@ -46,7 +46,7 @@ static int get_nr_ram_ranges_callback(struct resource *res, void *arg)unsignedintarch_get_system_nr_ranges(void){-unsignedintnr_ranges=2;/* For exclusion of crashkernel region */+unsignedintnr_ranges=2+crashk_cma_cnt;/* For exclusion of crashkernel region */walk_system_ram_res(0,-1,&nr_ranges,get_nr_ram_ranges_callback);
Commit 35c18f2933c5 ("Add a new optional ",cma" suffix to the
crashkernel= command line option") and commit ab475510e042 ("kdump:
implement reserve_crashkernel_cma") added CMA support for kdump
crashkernel reservation.
Crash kernel memory reservation wastes production resources if too
large, risks kdump failure if too small, and faces allocation difficulties
on fragmented systems due to contiguous block constraints. The new
CMA-based crashkernel reservation scheme splits the "large fixed
reservation" into a "small fixed region + large CMA dynamic region": the
CMA memory is available to userspace during normal operation to avoid
waste, and is reclaimed for kdump upon crash—saving memory while
improving reliability.
So extend crashkernel CMA reservation support to arm64. The following
changes are made to enable CMA reservation:
- Parse and obtain the CMA reservation size along with other crashkernel
parameters.
- Call reserve_crashkernel_cma() to allocate the CMA region for kdump.
- Include the CMA-reserved ranges for kdump kernel to use.
- Exclude the CMA-reserved ranges from the crash kernel memory to
prevent them from being exported through /proc/vmcore, which is already
done in the crash core.
Update kernel-parameters.txt to document CMA support for crashkernel on
arm64 architecture.
Tested-by: Breno Leitao <leitao@debian.org>
Acked-by: Catalin Marinas <catalin.marinas@arm.com>
Acked-by: Rob Herring (Arm) <robh@kernel.org>
Acked-by: Baoquan He <redacted>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
Acked-by: Ard Biesheuvel <ardb@kernel.org>
Signed-off-by: Jinjie Ruan <redacted>
---
v7:
- Correct the inclusion of CMA-reserved ranges for kdump
kernel in of/kexec.
v3:
- Add Acked-by.
v2:
- Free cmem in prepare_elf_headers()
- Add the mtivation.
---
Documentation/admin-guide/kernel-parameters.txt | 2 +-
arch/arm64/kernel/machine_kexec_file.c | 2 +-
arch/arm64/mm/init.c | 5 +++--
drivers/of/fdt.c | 9 +++++----
drivers/of/kexec.c | 9 +++++++++
include/linux/crash_reserve.h | 4 +++-
6 files changed, 22 insertions(+), 9 deletions(-)
@@ -1119,7 +1119,7 @@ Kernel parameters It will be ignored when crashkernel=X,high is not used or memory reserved is below 4G. crashkernel=size[KMG],cma- [KNL, X86, ppc] Reserve additional crash kernel memory from+ [KNL, X86, ARM64, PPC] Reserve additional crash kernel memory from CMA. This reservation is usable by the first system's userspace memory and kernel movable allocations (memory balloon, zswap). Pages allocated from this memory range
@@ -42,7 +42,7 @@ int arch_kimage_file_post_load_cleanup(struct kimage *image)#ifdef CONFIG_CRASH_DUMPunsignedintarch_get_system_nr_ranges(void){-unsignedintnr_ranges=2;/* for exclusion of crashkernel region */+unsignedintnr_ranges=2+crashk_cma_cnt;/* for exclusion of crashkernel region */phys_addr_tstart,end;u64i;
On Mon, May 11, 2026 at 11:04:43AM +0800, Jinjie Ruan wrote:
There is a race condition between the kexec_load() system call
(crash kernel loading path) and memory hotplug operations that can
lead to buffer overflow and potential kernel crash.
During prepare_elf_headers(), the following steps occur:
1. The first for_each_mem_range() queries current System RAM memory ranges
2. Allocates buffer based on queried count
3. The 2st for_each_mem_range() populates ranges from memblock
If memory hotplug occurs between step 1 and step 3, the number of ranges
can increase, causing out-of-bounds write when populating cmem->ranges[].
This happens because kexec_load() uses kexec_trylock (atomic_t) while
memory hotplug uses device_hotplug_lock (mutex), so they don't serialize
with each other.
Add the explicit bounds checking to prevent out-of-bounds access.
It seems you have a TOCTOU type of issue, and this seems to be shrinking
the window, but not fully solving it?
On Mon, May 11, 2026 at 11:04:43AM +0800, Jinjie Ruan wrote:
quoted
There is a race condition between the kexec_load() system call
(crash kernel loading path) and memory hotplug operations that can
lead to buffer overflow and potential kernel crash.
During prepare_elf_headers(), the following steps occur:
1. The first for_each_mem_range() queries current System RAM memory ranges
2. Allocates buffer based on queried count
3. The 2st for_each_mem_range() populates ranges from memblock
If memory hotplug occurs between step 1 and step 3, the number of ranges
can increase, causing out-of-bounds write when populating cmem->ranges[].
This happens because kexec_load() uses kexec_trylock (atomic_t) while
memory hotplug uses device_hotplug_lock (mutex), so they don't serialize
with each other.
Add the explicit bounds checking to prevent out-of-bounds access.
It seems you have a TOCTOU type of issue, and this seems to be shrinking
the window, but not fully solving it?
Hi Breno,
Thanks for your comments regarding the TOCTOU issue.
You are correct that the current bounds checking only "shrinks the
window" and prevents a kernel crash, but doesn't fully guarantee header
consistency if a race occurs.
In my local environment, this race is extremely difficult to reproduce,
but it is theoretically possible.
To address this properly for arm64, I am considering two steps:
- For this patch: I will change the return value to -EAGAIN and keep the
bounds check. This ensures that even if a race happens, the kernel
remains safe (no OOB access), and user-space is notified to retry.
- Long-term solution: A better way to solve this is to implement ARM64
CRASH_HOTPLUG support (similar to x86). With crash hotplug, the kernel
will automatically re-generate the crash headers whenever a memory
hotplug event occurs. This makes the TOCTOU during the initial
kexec_load less critical, as any transient inconsistency will be
immediately corrected by the subsequent hotplug handler.
Does it make sense to you to use this patch as a safety guard first, and
then I (or someone else) follow up with the full CRASH_HOTPLUG support
for arm64 as [1]?
[1]:
https://lore.kernel.org/all/20260402081459.635022-1-ruanjinjie@huawei.com/
Best regards,
Jinjie
On Mon, May 11, 2026 at 07:30:44PM +0800, Jinjie Ruan wrote:
On 5/11/2026 5:46 PM, Breno Leitao wrote:
quoted
On Mon, May 11, 2026 at 11:04:43AM +0800, Jinjie Ruan wrote:
quoted
There is a race condition between the kexec_load() system call
(crash kernel loading path) and memory hotplug operations that can
lead to buffer overflow and potential kernel crash.
During prepare_elf_headers(), the following steps occur:
1. The first for_each_mem_range() queries current System RAM memory ranges
2. Allocates buffer based on queried count
3. The 2st for_each_mem_range() populates ranges from memblock
If memory hotplug occurs between step 1 and step 3, the number of ranges
can increase, causing out-of-bounds write when populating cmem->ranges[].
This happens because kexec_load() uses kexec_trylock (atomic_t) while
memory hotplug uses device_hotplug_lock (mutex), so they don't serialize
with each other.
Add the explicit bounds checking to prevent out-of-bounds access.
It seems you have a TOCTOU type of issue, and this seems to be shrinking
the window, but not fully solving it?
Hi Breno,
Thanks for your comments regarding the TOCTOU issue.
You are correct that the current bounds checking only "shrinks the
window" and prevents a kernel crash, but doesn't fully guarantee header
consistency if a race occurs.
In my local environment, this race is extremely difficult to reproduce,
but it is theoretically possible.
To address this properly for arm64, I am considering two steps:
- For this patch: I will change the return value to -EAGAIN and keep the
bounds check. This ensures that even if a race happens, the kernel
remains safe (no OOB access), and user-space is notified to retry.
- Long-term solution: A better way to solve this is to implement ARM64
CRASH_HOTPLUG support (similar to x86). With crash hotplug, the kernel
will automatically re-generate the crash headers whenever a memory
hotplug event occurs. This makes the TOCTOU during the initial
kexec_load less critical, as any transient inconsistency will be
immediately corrected by the subsequent hotplug handler.
Does it make sense to you to use this patch as a safety guard first, and
then I (or someone else) follow up with the full CRASH_HOTPLUG support
for arm64 as [1]?
It would be OK for me, but, make it explict that there is a TOCTOU
issue, that depends on CRASH_HOTPLUG.
On Mon, May 11, 2026 at 11:04:43AM +0800, Jinjie Ruan wrote:
quoted
There is a race condition between the kexec_load() system call
(crash kernel loading path) and memory hotplug operations that can
lead to buffer overflow and potential kernel crash.
During prepare_elf_headers(), the following steps occur:
1. The first for_each_mem_range() queries current System RAM memory ranges
2. Allocates buffer based on queried count
3. The 2st for_each_mem_range() populates ranges from memblock
If memory hotplug occurs between step 1 and step 3, the number of ranges
can increase, causing out-of-bounds write when populating cmem->ranges[].
This happens because kexec_load() uses kexec_trylock (atomic_t) while
memory hotplug uses device_hotplug_lock (mutex), so they don't serialize
with each other.
Add the explicit bounds checking to prevent out-of-bounds access.
It seems you have a TOCTOU type of issue, and this seems to be shrinking
the window, but not fully solving it?
I plan to fix this issue as follows, and would appreciate your feedback
on whether this is reasonable.
Sashiko AI code review pointed out there is a TOCTOU (Time-of-Check to
Time-of-Use) race condition in prepare_elf_headers() between the initial
pass that counts System RAM ranges and the second pass that populates them.
If a memory hotplug event occurs between these two steps, the number of
memory regions may increase, causing an out-of-bounds write to
the cmem->ranges[] array.
To resolve this and ensure data consistency, this patch:
1. Wraps the counting and population passes with get_online_mems() and
crash_hotplug_lock(). This serializes the kexec_file_load() path
with concurrent memory hotplug operations, ensuring the memory
map remains consistent throughout the header preparation.
2. Adds an explicit boundary check in prepare_elf64_ram_headers_callback().
If the number of ranges exceeds the allocated maximum, it now returns
-EAGAIN, which indicates a transient race, signaling userspace
kexec-tools to retry the syscall instead of leaving the system
without a loaded crash kernel.
index daf81a873bbd..546be6261177 100644