When CPU or memory hotplug events occur, the elfcorehdr in the kdump
image becomes stale, potentially leading to incomplete crash dumps.
Currently, userspace udev rules reload the entire kdump image upon such
events, which is inefficient and leaves kdump inactive for a long time.
Commit 247262756121 ("crash: add generic infrastructure for crash hotplug
support") introduced a kernel mechanism to update only the elfcorehdr.
This patch set implements crash hotplug support for arm64.
As Baoquan and Catalin suggested, it also addresses and fixes several
pre-existing code issues found by Sashiko AI [1][2].
The major improvements and fixes included in this series are:
- Fix several memory leaks for arm64, and similar issues on LoongArch.
- Fix TOCTOU race in crash memory range collection
- Simplify arm64 load_other_segments().
- Implement infrastructure for arm64 crash memory hotplug support.
This patch set is rebased on v7.3-rc1.
TESTING
=======
Only kexec_file_load path has been tested; kexec_load is expected to
work via KEXEC_CRASH_HOTPLUG_SUPPORT flag but not yet verified.
Tested on an arm64 guest using KVM-QEMU[3] with the following
configuration:
-M virt,acpi=on,highmem=on
-m 4G,slots=256,maxmem=16G
-smp cpus=4,maxcpus=8,cores=4,threads=2,sockets=1
1. Memory Hot-Add Test
[Step 1] Load kexec first:
./kexec --kexec-file-syscall ...
[Step 2] Hotplug and online 128M memory and trigger crash:
(qemu) object_add memory-backend-ram,id=mem1,size=128M
(qemu) device_add pc-dimm,id=dimm3,memdev=mem1,addr=0x160000000
echo 1 > /sys/devices/system/memory/memory44/online
echo c > /proc/sysrq-trigger
[Step 3] Verify vmcore layout in the secondary kernel:
readelf -l /proc/vmcore
The newly added 128M memory segment (0x160000000) is successfully
recognized and populated as a LOAD segment:
LOAD 0x... 0x0000000160000000 0x08000000 0x08000000 RWE 0x0
2. Memory Hot-Remove Test
[Step 1] Add memory device and online 128M memory first:
(qemu) device_add pc-dimm,id=dimm3,memdev=mem1,addr=0x160000000
echo 1 > /sys/devices/system/memory/memory44/online
[Step 2] Load kexec, offline memory, and crash:
./kexec --kexec-file-syscall ...
echo 0 > /sys/devices/system/memory/memory44/online
echo c > /proc/sysrq-trigger
[Step 3] Verify vmcore layout:
readelf -l /proc/vmcore
Result: The 0x160000000 segment is cleanly excluded from the vmcore
program headers, and the dump completes without any hang.
3. CPU Hot-Add Test
[Step 1] Load kexec first:
./kexec --kexec-file-syscall ...
[Step 2] hotplug and online one CPU, then crash:
(qemu) device_add driver=host-arm-cpu,core-id=2,thread-id=0,id=cpu4
echo 1 > /sys/devices/system/cpu/cpu4/online
echo c > /proc/sysrq-trigger
[Step 3] Verify notes count:
readelf -n /proc/vmcore | grep -w CORE | wc -l
5
Result: Crash hotplug responds correctly; the newly plugged CPU4 is
tracked, and 5 NT_PRSTATUS notes are generated.
4. CPU Hot-Remove Test
[Step 1] Add CPU device and online it first:
(qemu) device_add driver=host-arm-cpu,core-id=2,thread-id=0,id=cpu4
echo 1 > /sys/devices/system/cpu/cpu4/online
[Step 2] Load kexec, remove CPU, and crash:
./kexec --kexec-file-syscall ...
(qemu) device_del cpu4
echo c > /proc/sysrq-trigger
[Step 3] Verify notes count:
readelf -n /proc/vmcore | grep -w CORE | wc -l
4
Result: Crash hotplug automatically updates the headers upon CPU
eviction; only 4 online CPUs are registered in the vmcore.
[1]: https://lore.kernel.org/all/20260601094805.2928614-1-ruanjinjie@huawei.com/
[2]: https://sashiko.dev/#/patchset/20260729031235.2840255-1-ruanjinjie%40huawei.com
[3]: https://github.com/salil-mehta/qemu.git virt-cpuhp-armv8/rfc-v2
Changes in v4:
- Rebased on v7.3-rc1.
- Update the kexec_core code as Mike suggested.
- Update the LoongArch subject as Huacai suggested.
- Drop crash_dump_dm_crypt patch which will be fixed by Coiby in [4] as
Sourabh suggested.
- Drop x86 related patches because of branch conflict, which will
be done later.
- Drop the incorrect CRASH_MAX_MEMORY_RANGES patch.
- Handle elfcorehdr_index in arm64 arch code.
- Link to v3: https://lore.kernel.org/all/20260826092541.3905933-1-ruanjinjie@huawei.com/
[4] https://lore.kernel.org/all/20260828084900.1496839-2-coiby.xu@gmail.com/
Changes in v3:
- Handle "KEXEC_CRASH_HP_REMOVE_MEMORY" action.
- Fix several pre-existing code issues reported by Sashiko AI review [3].
- Introduce crash_extra_elfcorehdr_size() and elf64_phdr_size() helper.
- Rework related crash and arch code.
- Add test method.
- v2: https://lore.kernel.org/all/20260729031235.2840255-1-ruanjinjie@huawei.com/
Changes in v2:
- Split out Powerpc bugfix patch as Mike suggested.
- Use phys_to_virt() instead of __va() in update_crash_elfcorehdr().
- Convert pnum_hdr_sz() to a function.
- Only assign elfcorehdr_index after kexec_add_buffer succeeds, considering
crash_handle_hotplug_event() already performs validity check on
elfcorehdr_index:
- We can safely remove the check for CPU hotplug
in arch_crash_handle_hotplug_event().
- The elfcorehdr_index's segment mem will be valid in
update_crash_elfcorehdr(), so we can safely remove the NULL check.
- Simplify the commit message.
- v1: https://lore.kernel.org/all/20260723131242.1537633-1-ruanjinjie@huawei.com/#t
Jinjie Ruan (12):
kexec: Record allocated CMA pages to fix release size mismatch
kexec: Extract kexec_free_segment_cma() from kimage_free_cma()
arm64: kexec_file: Fix CMA page leaks in segment placement retry loops
arm64: kexec_file: Fix elf_headers memory leak in retry loop
LoongArch: kexec_file: Fix CMA page leaks in segment placement retry
loops
LoongArch: kexec_file: Fix elf_headers memory leak in retry loop
crash: Extract crash_get_memory_ranges() helper
crash: Fix TOCTOU race in crash memory range collection
elf: Introduce elf64_phdr_size() helper
crash: Introduce crash_extra_elfcorehdr_size() helper
arm64: kexec_file: Simplify load_other_segments()
arm64: crash: Add crash hotplug support
arch/arm64/Kconfig | 3 +
arch/arm64/include/asm/kexec.h | 11 ++
arch/arm64/kernel/Makefile | 2 +-
arch/arm64/kernel/crash.c | 167 +++++++++++++++++++++
arch/arm64/kernel/kexec_image.c | 1 +
arch/arm64/kernel/machine_kexec_file.c | 63 +++-----
arch/loongarch/kernel/kexec_efi.c | 1 +
arch/loongarch/kernel/machine_kexec_file.c | 10 +-
arch/powerpc/kexec/crash.c | 2 +-
arch/powerpc/kexec/file_load_64.c | 19 +--
arch/powerpc/platforms/powernv/opal-core.c | 3 +-
arch/x86/kernel/crash.c | 12 +-
fs/proc/vmcore.c | 6 +-
include/linux/crash_core.h | 21 +++
include/linux/elf.h | 4 +
include/linux/kexec.h | 3 +
kernel/crash_core.c | 50 +++++-
kernel/kexec_core.c | 26 ++--
kernel/kexec_file.c | 13 +-
19 files changed, 326 insertions(+), 91 deletions(-)
create mode 100644 arch/arm64/kernel/crash.c
--
2.34.1
The CMA pages allocated for a kexec segment are released using the
segment's memsz to calculate the number of pages. However, some
architecture loaders modify the segment's memsz after allocation
(e.g. arm64 subtracts text_offset), causing the release function to
free fewer pages than were originally allocated, leaking the remaining
CMA pages.
Add a per-segment `segment_cma_pages` array to store the number of
pages actually allocated from CMA. Populate it during
kexec_add_buffer() using the aligned memsz, and use it in
kimage_free_cma() to accurately release all allocated pages.
This avoids relying on the potentially modified segment->memsz and
prevents silent CMA memory leaks.
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Baoquan He <baoquan.he@linux.dev>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Pasha Tatashin <pasha.tatashin@soleen.com>
Cc: Pratyush Yadav <pratyush@kernel.org>
Cc: Brian Mak <redacted>
Cc: Pingfan Liu <redacted>
Cc: Sourabh Jain <redacted>
Cc: Justinien Bouron <redacted>
Cc: Li Chen <redacted>
Cc: stable@vger.kernel.org
Link: https://sashiko.dev/#/patchset/20260729031235.2840255-1-ruanjinjie%40huawei.com
Fixes: 07d24902977e ("kexec: enable CMA based contiguous allocation")
Signed-off-by: Jinjie Ruan <redacted>
---
include/linux/kexec.h | 1 +
kernel/kexec_core.c | 7 ++++---
kernel/kexec_file.c | 13 +++++++++----
3 files changed, 14 insertions(+), 7 deletions(-)
@@ -670,7 +670,7 @@ static int kexec_walk_resources(struct kexec_buf *kbuf,staticintkexec_alloc_contig(structkexec_buf*kbuf){-size_tnr_pages=kbuf->memsz>>PAGE_SHIFT;+size_tnr_pages=PFN_DOWN(kbuf->memsz);unsignedlongmem;structpage*p;
@@ -756,14 +756,16 @@ int kexec_locate_mem_hole(struct kexec_buf *kbuf)*/intkexec_add_buffer(structkexec_buf*kbuf){+unsignedlongnr_segments=kbuf->image->nr_segments;structkexec_segment*ksegment;+size_tnr_cma_pages=0;intret;/* Currently adding segment this way is allowed only in file mode */if(!kbuf->image->file_mode)return-EINVAL;-if(kbuf->image->nr_segments>=KEXEC_SEGMENT_MAX)+if(nr_segments>=KEXEC_SEGMENT_MAX)return-EINVAL;/*
@@ -789,12 +791,15 @@ int kexec_add_buffer(struct kexec_buf *kbuf)returnret;/* Found a suitable memory range */-ksegment=&kbuf->image->segment[kbuf->image->nr_segments];+ksegment=&kbuf->image->segment[nr_segments];ksegment->kbuf=kbuf->buffer;ksegment->bufsz=kbuf->bufsz;ksegment->mem=kbuf->mem;ksegment->memsz=kbuf->memsz;-kbuf->image->segment_cma[kbuf->image->nr_segments]=kbuf->cma;+kbuf->image->segment_cma[nr_segments]=kbuf->cma;+if(kbuf->cma)+nr_cma_pages=PFN_DOWN(kbuf->memsz);+kbuf->image->segment_cma_pages[nr_segments]=nr_cma_pages;kbuf->image->nr_segments++;return0;}
kimage_free_cma() relies on image->nr_segments to iterate over segments.
When an architecture loader (e.g., arm64) truncates nr_segments on a
mid-way failure, CMA pages allocated beyond the new boundary become
unreachable, causing silent memory leaks.
Extract the per-segment freeing logic into the exported helper
kexec_free_segment_cma(), so that architecture loaders can release
individual segments before nr_segments is truncated. Refactor
kimage_free_cma() to loop over the new helper, preserving existing
behavior.
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Baoquan He <baoquan.he@linux.dev>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Pasha Tatashin <pasha.tatashin@soleen.com>
Cc: Pratyush Yadav <pratyush@kernel.org>
Signed-off-by: Jinjie Ruan <redacted>
---
include/linux/kexec.h | 2 ++
kernel/kexec_core.c | 27 +++++++++++++++------------
2 files changed, 17 insertions(+), 12 deletions(-)
Extract the elfcorehdr extra space calculation from powerpc into a
generic helper crash_extra_elfcorehdr_size() for use by other
architectures like arm64.
Strengthen the original powerpc check: instead of only checking
the loose CONFIG_CRASH_MAX_MEMORY_RANGES, the new helper enforces
a strict compile-time BUILD_BUG_ON() to guarantee that the absolute
maximum theoretical number of ELF Program Headers will never exceed
the ELF physical limit of PN_XNUM. This ensures absolute safety across
all architectures with zero runtime overhead.
The helper also provides a zero-size stub when crash memory hotplug
is disabled.
Cc: Madhavan Srinivasan <maddy@linux.ibm.com>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Nicholas Piggin <npiggin@gmail.com>
Cc: "Christophe Leroy (CS GROUP)" <chleroy@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Baoquan He <baoquan.he@linux.dev>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Pasha Tatashin <pasha.tatashin@soleen.com>
Cc: Pratyush Yadav <pratyush@kernel.org>
Cc: Dave Young <ruirui.yang@linux.dev>
Cc: Sourabh Jain <redacted>
Signed-off-by: Jinjie Ruan <redacted>
---
arch/powerpc/kexec/file_load_64.c | 19 +------------------
include/linux/crash_core.h | 20 ++++++++++++++++++++
2 files changed, 21 insertions(+), 18 deletions(-)
@@ -374,23 +374,6 @@ static int load_backup_segment(struct kimage *image, struct kexec_buf *kbuf)return0;}-staticunsignedintkdump_extra_elfcorehdr_size(structcrash_mem*cmem)-{-#if defined(CONFIG_CRASH_HOTPLUG) && defined(CONFIG_MEMORY_HOTPLUG)-unsignedintextra_sz=0;--if(CONFIG_CRASH_MAX_MEMORY_RANGES>(unsignedint)PN_XNUM)-pr_warn("Number of Phdrs %u exceeds max\n",CONFIG_CRASH_MAX_MEMORY_RANGES);-elseif(cmem->nr_ranges>=CONFIG_CRASH_MAX_MEMORY_RANGES)-pr_warn("Configured crash mem ranges may not be enough\n");-else-extra_sz=(CONFIG_CRASH_MAX_MEMORY_RANGES-cmem->nr_ranges)*sizeof(Elf64_Phdr);--returnextra_sz;-#endif-return0;-}-/***load_elfcorehdr_segment-Setupcrashmemoryrangesandinitializeelfcorehdr*segmentneededtoloadkdumpkernel.
During kexec image placement retry loops, any midway failure causes
the loader to truncate `image->nr_segments` back to its initial state
to purge the failed segments.
However, this truncation introduces a memory leak. The CMA pages
allocated via kexec_add_buffer() during the failed attempt are tracked
in the `image->segment_cma` array. Because the subsequent cleanup paths
only iterate up to the truncated `nr_segments` boundary, these allocated
CMA pages outside the new boundary are permanently leaked.
Fix this by explicitly releasing the associated CMA buffers in
the failure paths before `image->nr_segments` is reduced.
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Will Deacon <will@kernel.org>
Cc: Breno Leitao <leitao@debian.org>
Cc: Pratyush Yadav <pratyush@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Yeoreum Yun <redacted>
Cc: Baoquan He <redacted>
Cc: stable@vger.kernel.org
Fixes: 07d24902977e4 ("kexec: enable CMA based contiguous allocation")
Signed-off-by: Jinjie Ruan <redacted>
---
arch/arm64/kernel/kexec_image.c | 1 +
arch/arm64/kernel/machine_kexec_file.c | 5 ++++-
2 files changed, 5 insertions(+), 1 deletion(-)
Add a common helper to compute the total size of an ELF64 header
(Ehdr + program headers) from the number of program headers.
Replace open-coded calculations in powerpc, x86, vmcore,
and crash_core.
On ppc64, struct elfhdr maps to elf64_hdr, so the powerpc change
is a pure cleanup.
No functional change intended.
Cc: Madhavan Srinivasan <maddy@linux.ibm.com>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Nicholas Piggin <npiggin@gmail.com>
Cc: "Christophe Leroy (CS GROUP)" <chleroy@kernel.org>
Cc: Thomas Gleixner <tglx@kernel.org>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Borislav Petkov <bp@alien8.de>
Cc: Dave Hansen <dave.hansen@linux.intel.com>
Cc: "H. Peter Anvin" <hpa@zytor.com>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Baoquan He <baoquan.he@linux.dev>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Pasha Tatashin <pasha.tatashin@soleen.com>
Cc: Pratyush Yadav <pratyush@kernel.org>
Cc: Dave Young <ruirui.yang@linux.dev>
Cc: Kees Cook <kees@kernel.org>
Cc: Sourabh Jain <redacted>
Signed-off-by: Jinjie Ruan <redacted>
---
arch/powerpc/kexec/crash.c | 2 +-
arch/powerpc/platforms/powernv/opal-core.c | 3 +--
arch/x86/kernel/crash.c | 3 +--
fs/proc/vmcore.c | 6 ++----
include/linux/elf.h | 4 ++++
kernel/crash_core.c | 2 +-
6 files changed, 10 insertions(+), 10 deletions(-)
@@ -309,8 +309,7 @@ static int __init create_opalcore(void)char*bufp;/* Get size of header & CPU notes for OPAL core */-hdr_size=(sizeof(Elf64_Ehdr)+-((oc_conf->ptload_cnt+1)*sizeof(Elf64_Phdr)));+hdr_size=elf64_phdr_size(oc_conf->ptload_cnt+1);cpu_notes_size=((oc_conf->num_cpus*(CRASH_CORE_NOTE_HEAD_BYTES+CRASH_CORE_NOTE_NAME_BYTES+CRASH_CORE_NOTE_DESC_BYTES))+
@@ -1238,8 +1238,7 @@ static int __init parse_crash_elf64_headers(void)}/* Read in all elf headers. */-elfcorebuf_sz_orig=sizeof(Elf64_Ehdr)+-ehdr.e_phnum*sizeof(Elf64_Phdr);+elfcorebuf_sz_orig=elf64_phdr_size(ehdr.e_phnum);elfcorebuf_sz=elfcorebuf_sz_orig;elfcorebuf=(void*)__get_free_pages(GFP_KERNEL|__GFP_ZERO,get_order(elfcorebuf_sz_orig));
@@ -1605,8 +1604,7 @@ static int vmcore_add_device_ram_elf64(struct list_head *list, size_t count)}/* elfcorebuf_sz must always cover full pages. */-new_size=sizeof(Elf64_Ehdr)+-(ehdr->e_phnum+count)*sizeof(Elf64_Phdr);+new_size=elf64_phdr_size(ehdr->e_phnum+count);new_size=roundup(new_size,PAGE_SIZE);/*
@@ -193,7 +193,7 @@ int crash_prepare_elf64_headers(struct crash_mem *mem, int need_kernel_map,*/nr_phdr++;-elf_sz=sizeof(Elf64_Ehdr)+nr_phdr*sizeof(Elf64_Phdr);+elf_sz=elf64_phdr_size(nr_phdr);elf_sz=ALIGN(elf_sz,ELF_CORE_HEADER_ALIGN);buf=vzalloc(elf_sz);
When CPU or memory hotplug events occur, the elfcorehdr in the kdump
image becomes stale, potentially leading to incomplete crash dumps.
Currently, userspace udev rules reload the entire kdump image upon such
events, which is inefficient and leaves kdump inactive for a long time.
Commit 247262756121 ("crash: add generic infrastructure for crash hotplug
support") introduced a kernel mechanism to update only the elfcorehdr.
This patch enables that support for arm64.
On arm64, only memory hotplug events require elfcorehdr updates:
- Physical CPU hotplug is not supported.
- For ACPI based vCPU hotplug [1], the elfcorehdr is built using
for_each_possible_cpu(), so no update is needed.
The patch:
- Adds CONFIG_ARCH_SUPPORTS_CRASH_HOTPLUG (default y).
- Implements following arch functions to handle memory hotplug:
1. arch_crash_hotplug_support()
2. arch_crash_get_elfcorehdr_size()
3. arch_crash_handle_hotplug_event()
- Moves arch_get_system_nr_ranges() and arch_crash_populate_cmem()
from machine_kexec_file.c to crash.c for crash hotplug reuse.
Follows the approach of x86 commit ea53ad9cf73b ("x86/crash: add x86 crash
hotplug support") and powerpc commit b741092d5976 ("powerpc/crash: add
crash CPU hotplug support").
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Will Deacon <will@kernel.org>
Cc: Baoquan He <redacted>
Cc: "Mike Rapoport (Microsoft)" <rppt@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Breno Leitao <leitao@debian.org>
Cc: Sourabh Jain <redacted>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Ard Biesheuvel <ardb@kernel.org>
Cc: Thomas Huth <redacted>
[1]: https://lore.kernel.org/all/20240529133446.28446-1-Jonathan.Cameron@huawei.com/
Signed-off-by: Jinjie Ruan <redacted>
---
arch/arm64/Kconfig | 3 +
arch/arm64/include/asm/kexec.h | 11 ++
arch/arm64/kernel/Makefile | 2 +-
arch/arm64/kernel/crash.c | 167 +++++++++++++++++++++++++
arch/arm64/kernel/machine_kexec_file.c | 42 ++-----
5 files changed, 192 insertions(+), 33 deletions(-)
create mode 100644 arch/arm64/kernel/crash.c
@@ -0,0 +1,167 @@+// SPDX-License-Identifier: GPL-2.0-only+/*+*Architecturespecificfunctionsforkexecbasedcrashdumps.+*/++#define pr_fmt(fmt) "crash hp: " fmt++#include<linux/cacheflush.h>+#include<linux/elf.h>+#include<linux/kexec.h>+#include<linux/memblock.h>+#include<linux/memory.h>+#include<linux/vmalloc.h>++#include<asm/kexec.h>++#if defined(CONFIG_KEXEC_FILE) || defined(CONFIG_CRASH_HOTPLUG)+unsignedintarch_get_system_nr_ranges(void)+{+unsignedintnr_ranges=2+crashk_cma_cnt;/* for exclusion of crashkernel region */+phys_addr_tstart,end;+u64i;++for_each_mem_range(i,&start,&end)+nr_ranges++;++returnnr_ranges;+}++intarch_crash_populate_cmem(structcrash_mem*cmem)+{+phys_addr_tstart,end;+u64i;++for_each_mem_range(i,&start,&end){+cmem->ranges[cmem->nr_ranges].start=start;+cmem->ranges[cmem->nr_ranges].end=end-1;+cmem->nr_ranges++;+}++return0;+}+#endif++#ifdef CONFIG_CRASH_HOTPLUG+intarch_crash_hotplug_support(structkimage*image,unsignedlongkexec_flags)+{+#ifdef CONFIG_KEXEC_FILE+if(image->file_mode)+return1;+#endif+/*+*Forkexec_loadsyscall,crashhotplugsupportrequires+*KEXEC_CRASH_HOTPLUG_SUPPORTflagtobepassedbyuserspace.+*/+returnkexec_flags&KEXEC_CRASH_HOTPLUG_SUPPORT;+}++unsignedintarch_crash_get_elfcorehdr_size(void)+{+unsignedlongphdr_cnt;++/* A program header for possible CPUs, vmcoreinfo and kernel_map */+phdr_cnt=2+num_possible_cpus();+if(IS_ENABLED(CONFIG_MEMORY_HOTPLUG))+phdr_cnt+=CONFIG_CRASH_MAX_MEMORY_RANGES;++returnelf64_phdr_size(phdr_cnt);+}++/**+*update_crash_elfcorehdr()-Recreatetheelfcorehdrandreplaceitwithold+*elfcorehdrinthekexecsegmentarray.+*@image:theactivestructkimage+*@mn:structmemory_notifydatahandler+*/+staticvoidupdate_crash_elfcorehdr(structkimage*image,structmemory_notify*mn)+{+void*elfbuf=NULL,*old_elfcorehdr;+unsignedlongmem,memsz,elfsz=0;+structcrash_mem*cmem=NULL;+u64start,end;+intret;++ret=crash_get_memory_ranges_nolock(&cmem);+if(ret){+pr_err("Failed to get crash memory ranges.\n");+gotoout;+}++/*+*Thehotunpluggedmemoryispartofcrashmemoryranges,+*removeithere.+*/+if(image->hp_action==KEXEC_CRASH_HP_REMOVE_MEMORY){+start=PFN_PHYS(mn->start_pfn);+end=start+PFN_PHYS(mn->nr_pages)-1;++ret=crash_exclude_mem_range(cmem,start,end);+if(ret){+pr_err("Failed to remove hot-unplugged memory from crash memory ranges.\n");+gotoout;+}+}++/*+*CreatethenewelfcorehdrreflectingthechangestoCPUand/or+*memoryresources.+*/+ret=crash_prepare_elf64_headers(cmem,true,&elfbuf,&elfsz);+if(ret){+pr_err("Failed to create new elfcorehdr");+gotoout;+}++/*+*Obtainaddressandsizeoftheelfcorehdrsegment,and+*checkitagainstthenewelfcorehdrbuffer.+*/+mem=image->segment[image->elfcorehdr_index].mem;+memsz=image->segment[image->elfcorehdr_index].memsz;+if(elfsz>memsz){+pr_err("update elfcorehdr elfsz %lu > memsz %lu",+elfsz,memsz);+gotoout;+}++/* Copy new elfcorehdr over the old elfcorehdr at destination. */+old_elfcorehdr=phys_to_virt(mem);++/*+*Temporarilyinvalidatethecrashimagewhilethe+*elfcorehdrisupdated.+*/+xchg(&kexec_crash_image,NULL);+memcpy(old_elfcorehdr,elfbuf,elfsz);+dcache_clean_inval_poc((unsignedlong)old_elfcorehdr,+(unsignedlong)(old_elfcorehdr+elfsz));+xchg(&kexec_crash_image,image);+pr_debug("updated elfcorehdr\n");++out:+kvfree(cmem);+vfree(elfbuf);+}++/**+*arch_crash_handle_hotplug_event()-Handlehotplugelfcorehdrchanges+*@image:apointertokexec_crash_image+*@arg:structmemory_notifyhandlerformemoryhotplugcaseand+*NULLforCPUhotplugcase.+*+*Updatethekdumpimagebasedonthetypeofhotplugevent:+*-CPUaddandremove:Noactionisneeded.+*-Memoryadd/remove:Updatetheelfcorehdrtoreflectthecurrentmemorylayout.+*+*Preparethenewelfcorehdrandreplacetheexistingelfcorehdr.+*/+voidarch_crash_handle_hotplug_event(structkimage*image,void*arg)+{+if(image->hp_action==KEXEC_CRASH_HP_ADD_CPU||+image->hp_action==KEXEC_CRASH_HP_REMOVE_CPU)+return;++update_crash_elfcorehdr(image,(structmemory_notify*)arg);+}+#endif /* CONFIG_CRASH_HOTPLUG */
@@ -39,34 +38,6 @@ int arch_kimage_file_post_load_cleanup(struct kimage *image)returnkexec_image_post_load_cleanup_default(image);}-#ifdef CONFIG_CRASH_DUMP-unsignedintarch_get_system_nr_ranges(void)-{-unsignedintnr_ranges=2+crashk_cma_cnt;/* for exclusion of crashkernel region */-phys_addr_tstart,end;-u64i;--for_each_mem_range(i,&start,&end)-nr_ranges++;--returnnr_ranges;-}--intarch_crash_populate_cmem(structcrash_mem*cmem)-{-phys_addr_tstart,end;-u64i;--for_each_mem_range(i,&start,&end){-cmem->ranges[cmem->nr_ranges].start=start;-cmem->ranges[cmem->nr_ranges].end=end-1;-cmem->nr_ranges++;-}--return0;-}-#endif-/**TriestoaddtheinitrdandDTBtotheimage.Ifitisnotpossibletofind*validlocations,thisfunctionwillundochangestotheimageandreturnnon
Factor out the crash memory range collection logic from
crash_prepare_headers() into a separate function. This allows
the memory hotplug path to obtain and modify the range list
(e.g. remove offlined memory) before generating the elfcorehdr.
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Baoquan He <baoquan.he@linux.dev>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Pasha Tatashin <pasha.tatashin@soleen.com>
Cc: Pratyush Yadav <pratyush@kernel.org>
Cc: Dave Young <ruirui.yang@linux.dev>
Signed-off-by: Jinjie Ruan <redacted>
---
include/linux/crash_core.h | 1 +
kernel/crash_core.c | 22 +++++++++++++++++++---
2 files changed, 20 insertions(+), 3 deletions(-)
@@ -317,8 +317,7 @@ int crash_exclude_core_ranges(struct crash_mem **cmem)return0;}-intcrash_prepare_headers(intneed_kernel_map,void**addr,unsignedlong*sz,-unsignedlong*nr_mem_ranges)+intcrash_get_memory_ranges(structcrash_mem**mem_ranges){unsignedintmax_nr_ranges;structcrash_mem*cmem;
@@ -344,13 +343,30 @@ int crash_prepare_headers(int need_kernel_map, void **addr, unsigned long *sz,if(ret)gotoout;+*mem_ranges=cmem;+return0;++out:+kvfree(cmem);+returnret;+}++intcrash_prepare_headers(intneed_kernel_map,void**addr,unsignedlong*sz,+unsignedlong*nr_mem_ranges)+{+structcrash_mem*cmem=NULL;+intret;++ret=crash_get_memory_ranges(&cmem);+if(ret)+returnret;+/* Return the computed number of memory ranges, for hotplug usage */if(nr_mem_ranges)*nr_mem_ranges=cmem->nr_ranges;ret=crash_prepare_elf64_headers(cmem,need_kernel_map,addr,sz);-out:kvfree(cmem);returnret;}
Use `kbuf` fields directly in crash_prepare_headers() to eliminate
the local variables "headers" and "headers_sz"..
Advance the assignment to image->elf_headers before
calling kexec_add_buffer(). If kexec_add_buffer() fails, the explicit
vfree() in the error path can be removed, as the global infrastructure
in arch_kimage_file_post_load_cleanup() will handle the cleanup.
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Will Deacon <will@kernel.org>
Cc: Baoquan He <redacted>
Cc: Breno Leitao <leitao@debian.org>
Signed-off-by: Jinjie Ruan <redacted>
---
arch/arm64/kernel/machine_kexec_file.c | 24 +++++++++---------------
1 file changed, 9 insertions(+), 15 deletions(-)
If load_other_segments() fails after image->elf_headers is assigned,
the memory lifecycle is safely managed by the global kimage object
and will be freed in arch_kimage_file_post_load_cleanup().
However, during a retry loop in image_load(), a subsequent iteration
will allocate a new buffer and overwrite image->elf_headers. This
permanently leaks the stale memory from the previous iteration before
the global cleanup can track it.
Fix this by explicitly freeing the stale `image->elf_headers` buffer
before assigning the newly allocated headers.
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Will Deacon <will@kernel.org>
Cc: Thomas Huth <redacted>
Cc: Breno Leitao <leitao@debian.org>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Yeoreum Yun <redacted>
Cc: Baoquan He <redacted>
Cc: stable@vger.kernel.org
Fixes: 108aa503657e ("arm64: kexec_file: try more regions if loading segments fails")
Signed-off-by: Jinjie Ruan <redacted>
---
arch/arm64/kernel/machine_kexec_file.c | 4 ++++
1 file changed, 4 insertions(+)
During kexec image placement retry loops, any midway failure causes
the loader to truncate `image->nr_segments` back to its initial state
to purge the failed segments.
However, this truncation introduces a memory leak. The CMA pages
allocated via kexec_add_buffer() during the failed attempt are tracked
in the `image->segment_cma` array. Because the subsequent cleanup paths
only iterate up to the truncated `nr_segments` boundary, these allocated
CMA pages outside the new boundary are permanently leaked.
Fix this by explicitly releasing the associated CMA buffers in
the failure paths before `image->nr_segments` is reduced.
Cc: Huacai Chen <chenhuacai@kernel.org>
Cc: WANG Xuerui <kernel@xen0n.name>
Cc: Youling Tang <redacted>
Cc: "Mike Rapoport (Microsoft)" <rppt@kernel.org>
Cc: Sourabh Jain <redacted>
Cc: Kees Cook <kees@kernel.org>
Cc: stable@vger.kernel.org
Link: https://sashiko.dev/#/patchset/20260729031235.2840255-1-ruanjinjie%40huawei.com
Fixes: 55d990f0084c ("LoongArch: Add EFI binary support for kexec_file")
Signed-off-by: Jinjie Ruan <redacted>
---
arch/loongarch/kernel/kexec_efi.c | 1 +
arch/loongarch/kernel/machine_kexec_file.c | 6 +++++-
2 files changed, 6 insertions(+), 1 deletion(-)
The crash kernel ELF core header construction counts system memory
ranges via `arch_get_system_nr_ranges()`, allocates the crash_mem
buffer, and then populates it via `arch_crash_populate_cmem()`.
This sequence has a time-of-check-to-time-of-use (TOCTOU) race with
memory hotplug: a concurrent hotplug event between the count
and populate steps can increase the number of ranges beyond the allocated
capacity, causing an out-of-bounds write. If the event triggers
memblock_double_array(), the memblock array can be freed and reallocated
during iteration, leading to a use-after-free.
Protect the entire range collection with device_hotplug_lock. Since
the hotplug notification path already holds that lock, add a lockless
helper, crash_get_memory_ranges_nolock(), for use there. The regular
crash_get_memory_ranges() acquires the lock and calls the helper.
Cc: stable@vger.kernel.org
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Baoquan He <baoquan.he@linux.dev>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Pasha Tatashin <pasha.tatashin@soleen.com>
Cc: Pratyush Yadav <pratyush@kernel.org>
Cc: Dave Young <ruirui.yang@linux.dev>
Cc: AKASHI Takahiro <redacted>
Cc: Will Deacon <will@kernel.org>
Cc: James Morse <james.morse@arm.com>
Cc: Palmer Dabbelt <redacted>
Cc: Youling Tang <redacted>
Cc: Huacai Chen <chenhuacai@kernel.org>
Fixes: 8d5f894a3108 ("x86: kexec_file: lift CRASH_MAX_RANGES limit on crash_mem buffer")
Fixes: 3751e728cef2 ("arm64: kexec_file: add crash dump support")
Fixes: 8acea455fafa ("RISC-V: Support for kexec_file on panic")
Fixes: 1bcca8620a91 ("LoongArch: Add crash dump support for kexec_file")
Link: https://sashiko.dev/#/patchset/20260729031235.2840255-1-ruanjinjie%40huawei.com
Signed-off-by: Jinjie Ruan <redacted>
---
arch/x86/kernel/crash.c | 9 ++++++++-
include/linux/crash_core.h | 2 +-
kernel/crash_core.c | 28 +++++++++++++++++++++++++++-
3 files changed, 36 insertions(+), 3 deletions(-)
@@ -448,6 +448,7 @@ unsigned int arch_crash_get_elfcorehdr_size(void)voidarch_crash_handle_hotplug_event(structkimage*image,void*arg){void*elfbuf=NULL,*old_elfcorehdr;+structcrash_mem*cmem=NULL;unsignedlongmem,memsz;unsignedlongelfsz=0;
@@ -461,11 +462,16 @@ void arch_crash_handle_hotplug_event(struct kimage *image, void *arg)(image->hp_action==KEXEC_CRASH_HP_REMOVE_CPU)))return;+if(crash_get_memory_ranges_nolock(&cmem)){+pr_err("Failed to get crash mem range\n");+gotoout;+}+/**CreatethenewelfcorehdrreflectingthechangestoCPUand/or*memoryresources.*/-if(crash_prepare_headers(IS_ENABLED(CONFIG_X86_64),&elfbuf,&elfsz,NULL)){+if(crash_prepare_elf64_headers(cmem,IS_ENABLED(CONFIG_X86_64),&elfbuf,&elfsz)){pr_err("unable to create new elfcorehdr");gotoout;}
If load_other_segments() fails after image->elf_headers is assigned,
the memory lifecycle is safely managed by the global kimage object
and will be freed in arch_kimage_file_post_load_cleanup().
However, during a retry loop in efi_kexec_load(), a subsequent iteration
will allocate a new buffer and overwrite image->elf_headers. This
permanently leaks the stale memory from the previous iteration before
the global cleanup can track it.
Fix this by explicitly freeing the stale `image->elf_headers` buffer
before assigning the newly allocated headers.
Cc: Huacai Chen <chenhuacai@kernel.org>
Cc: WANG Xuerui <kernel@xen0n.name>
Cc: Youling Tang <redacted>
Cc: "Mike Rapoport (Microsoft)" <rppt@kernel.org>
Cc: Sourabh Jain <redacted>
Cc: Kees Cook <kees@kernel.org>
Cc: stable@vger.kernel.org
Link: https://sashiko.dev/#/patchset/20260729031235.2840255-1-ruanjinjie%40huawei.com
Fixes: 55d990f0084c ("LoongArch: Add EFI binary support for kexec_file")
Signed-off-by: Jinjie Ruan <redacted>
---
arch/loongarch/kernel/machine_kexec_file.c | 4 ++++
1 file changed, 4 insertions(+)
Hi, Jinjie,
On Mon, Sep 7, 2026 at 8:53 PM Jinjie Ruan [off-list ref] wrote:
During kexec image placement retry loops, any midway failure causes
the loader to truncate `image->nr_segments` back to its initial state
to purge the failed segments.
However, this truncation introduces a memory leak. The CMA pages
allocated via kexec_add_buffer() during the failed attempt are tracked
in the `image->segment_cma` array. Because the subsequent cleanup paths
only iterate up to the truncated `nr_segments` boundary, these allocated
CMA pages outside the new boundary are permanently leaked.
Fix this by explicitly releasing the associated CMA buffers in
the failure paths before `image->nr_segments` is reduced.
I'm not sure but does the elf version have similar problems as the efi version?
Huacai
On Mon, Sep 7, 2026 at 8:53 PM Jinjie Ruan [off-list ref] wrote:
quoted
During kexec image placement retry loops, any midway failure causes
the loader to truncate `image->nr_segments` back to its initial state
to purge the failed segments.
However, this truncation introduces a memory leak. The CMA pages
allocated via kexec_add_buffer() during the failed attempt are tracked
in the `image->segment_cma` array. Because the subsequent cleanup paths
only iterate up to the truncated `nr_segments` boundary, these allocated
CMA pages outside the new boundary are permanently leaked.
Fix this by explicitly releasing the associated CMA buffers in
the failure paths before `image->nr_segments` is reduced.
I'm not sure but does the elf version have similar problems as the efi version?
This issue does not exist in the ELF version, as the retry loop for
load_other_segments() does not exit here.
Best regards,
Jinjie