Hi all,
I am not really sure whose bug this is, as it only appears when three
seemingly independent patch series are applied together, so I have added
the patch authors and their committers (along with the tracing
maintainers) to this thread. Feel free to expand or reduce that list as
necessary.
Our continuous integration has noticed a crash when booting
ppc64_guest_defconfig in QEMU on the past few -next versions.
https://github.com/ClangBuiltLinux/continuous-integration2/actions/runs/23311154492/job/67811527112
This does not appear to be clang related, as it can be reproduced with
GCC 15.2.0 as well. Through multiple bisects, I was able to land on
applying:
mm: improve RSS counter approximation accuracy for proc interfaces [1]
vdso/datastore: Allocate data pages dynamically [2]
kho: fix deferred init of kho scratch [3]
and their dependent changes on top of 7.0-rc4 is enough to reproduce
this (at least on two of my machines with the same commands). I have
attached the diff from the result of the following 'git apply' commands
below, done in a linux-next checkout.
$ git checkout v7.0-rc4
HEAD is now at f338e7738378 Linux 7.0-rc4
# [1]
$ git diff 60ddf3eed4999bae440d1cf9e5868ccb3f308b64^..087dd6d2cc12c82945ab859194c32e8e977daae3 | git apply -3v
...
# [2]
# Fix trivial conflict in init/main.c around headers
$ git diff dc432ab7130bb39f5a351281a02d4bc61e85a14a^..05988dba11791ccbb458254484826b32f17f4ad2 | git apply -3v
...
# [3]
# Fix conflict in kernel/liveupdate/kexec_handover.c due to lack of kho_mem_retrieve(), just add pfn_is_kho_scratch()
$ git show 4a78467ffb537463486968232daef1e8a2f105e3 | git apply -3v
...
$ make -skj"$(nproc)" ARCH=powerpc CROSS_COMPILE=powerpc64-linux- mrproper ppc64_guest_defconfig vmlinux
$ curl -LSs https://github.com/ClangBuiltLinux/boot-utils/releases/download/20241120-044434/ppc64-rootfs.cpio.zst | zstd -d >rootfs.cpio
$ qemu-system-ppc64 \
-display none \
-nodefaults \
-cpu power8 \
-machine pseries \
-vga none \
-kernel vmlinux \
-initrd rootfs.cpio \
-m 1G \
-serial mon:stdio
...
[ 0.000000][ T0] Linux version 7.0.0-rc4-dirty (nathan@framework-amd-ryzen-maxplus-395) (powerpc64-linux-gcc (GCC) 15.2.0, GNU ld (GNU Binutils) 2.45) #1 SMP PREEMPT Thu Mar 19 15:45:53 MST 2026
...
[ 0.216764][ T1] vgaarb: loaded
[ 0.217590][ T1] clocksource: Switched to clocksource timebase
[ 0.221007][ T12] BUG: Kernel NULL pointer dereference at 0x00000010
[ 0.221049][ T12] Faulting instruction address: 0xc00000000044947c
[ 0.221237][ T12] Oops: Kernel access of bad area, sig: 11 [#1]
[ 0.221276][ T12] BE PAGE_SIZE=64K MMU=Hash SMP NR_CPUS=2048 NUMA pSeries
[ 0.221359][ T12] Modules linked in:
[ 0.221556][ T12] CPU: 0 UID: 0 PID: 12 Comm: kworker/u4:0 Not tainted 7.0.0-rc4-dirty #1 PREEMPTLAZY
[ 0.221631][ T12] Hardware name: IBM pSeries (emulated by qemu) POWER8 (architected) 0x4d0200 0xf000004 of:SLOF,HEAD pSeries
[ 0.221765][ T12] Workqueue: trace_init_wq tracer_init_tracefs_work_func
[ 0.222065][ T12] NIP: c00000000044947c LR: c00000000041a584 CTR: c00000000053aa90
[ 0.222084][ T12] REGS: c000000003bc7960 TRAP: 0380 Not tainted (7.0.0-rc4-dirty)
[ 0.222111][ T12] MSR: 8000000000009032 <SF,EE,ME,IR,DR,RI> CR: 44000204 XER: 00000000
[ 0.222287][ T12] CFAR: c000000000449420 IRQMASK: 0
[ 0.222287][ T12] GPR00: c00000000041a584 c000000003bc7c00 c000000001c08100 c000000002892f20
[ 0.222287][ T12] GPR04: c0000000019cfa68 c0000000019cfa60 0000000000000001 0000000000000064
[ 0.222287][ T12] GPR08: 0000000000000002 0000000000000000 c000000003bba000 0000000000000010
[ 0.222287][ T12] GPR12: c00000000053aa90 c000000002c50000 c000000001ab25f8 c000000001626690
[ 0.222287][ T12] GPR16: 0000000000000000 0000000000000000 0000000000000000 0000000000000000
[ 0.222287][ T12] GPR20: c000000001624868 c000000001ab2708 c0000000019cfa08 c000000001a00d18
[ 0.222287][ T12] GPR24: c0000000019cfa18 fffffffffffffef7 c000000003051205 c0000000019cfa68
[ 0.222287][ T12] GPR28: 0000000000000000 c0000000019cfa60 c000000002894e90 0000000000000000
[ 0.222526][ T12] NIP [c00000000044947c] __find_event_file+0x9c/0x110
[ 0.222572][ T12] LR [c00000000041a584] init_tracer_tracefs+0x274/0xcc0
[ 0.222643][ T12] Call Trace:
[ 0.222690][ T12] [c000000003bc7c00] [c000000000b943b0] tracefs_create_file+0x1a0/0x2b0 (unreliable)
[ 0.222766][ T12] [c000000003bc7c50] [c00000000041a584] init_tracer_tracefs+0x274/0xcc0
[ 0.222791][ T12] [c000000003bc7dc0] [c000000002046f1c] tracer_init_tracefs_work_func+0x50/0x320
[ 0.222809][ T12] [c000000003bc7e50] [c000000000276958] process_one_work+0x1b8/0x530
[ 0.222828][ T12] [c000000003bc7f10] [c00000000027778c] worker_thread+0x1dc/0x3d0
[ 0.222883][ T12] [c000000003bc7f90] [c000000000284c44] kthread+0x194/0x1b0
[ 0.222900][ T12] [c000000003bc7fe0] [c00000000000cf30] start_kernel_thread+0x14/0x18
[ 0.222961][ T12] Code: 7c691b78 7f63db78 2c090000 40820018 e89c0000 49107f21 60000000 2c030000 41820048 ebff0000 7c3ff040 41820038 <e93f0010> 7fa3eb78 81490058 e8890018
[ 0.223190][ T12] ---[ end trace 0000000000000000 ]---
...
Interestingly, turning on CONFIG_KASAN appears to hide this, maybe
pointing to some sort of memory corruption (or something timing
related)? If there is any other information I can provide, I am more
than happy to do so.
[1]: https://lore.kernel.org/20260227153730.1556542-4-mathieu.desnoyers@efficios.com/
[2]: https://lore.kernel.org/20260304-vdso-sparc64-generic-2-v6-3-d8eb3b0e1410@linutronix.de/
[3]: https://lore.kernel.org/20260311125539.4123672-2-mclapinski@google.com/
Cheers,
Nathan
From: Harry Yoo <hidden> Date: 2026-03-20 04:18:44
On Thu, Mar 19, 2026 at 04:37:45PM -0700, Nathan Chancellor wrote:
Hi all,
I am not really sure whose bug this is, as it only appears when three
seemingly independent patch series are applied together, so I have added
the patch authors and their committers (along with the tracing
maintainers) to this thread. Feel free to expand or reduce that list as
necessary.
Our continuous integration has noticed a crash when booting
ppc64_guest_defconfig in QEMU on the past few -next versions.
https://github.com/ClangBuiltLinux/continuous-integration2/actions/runs/23311154492/job/67811527112
This does not appear to be clang related, as it can be reproduced with
GCC 15.2.0 as well. Through multiple bisects, I was able to land on
applying:
mm: improve RSS counter approximation accuracy for proc interfaces [1]
vdso/datastore: Allocate data pages dynamically [2]
kho: fix deferred init of kho scratch [3]
and their dependent changes on top of 7.0-rc4 is enough to reproduce
this (at least on two of my machines with the same commands). I have
attached the diff from the result of the following 'git apply' commands
below, done in a linux-next checkout.
$ git checkout v7.0-rc4
HEAD is now at f338e7738378 Linux 7.0-rc4
# [1]
$ git diff 60ddf3eed4999bae440d1cf9e5868ccb3f308b64^..087dd6d2cc12c82945ab859194c32e8e977daae3 | git apply -3v
...
# [2]
# Fix trivial conflict in init/main.c around headers
$ git diff dc432ab7130bb39f5a351281a02d4bc61e85a14a^..05988dba11791ccbb458254484826b32f17f4ad2 | git apply -3v
...
# [3]
# Fix conflict in kernel/liveupdate/kexec_handover.c due to lack of kho_mem_retrieve(), just add pfn_is_kho_scratch()
$ git show 4a78467ffb537463486968232daef1e8a2f105e3 | git apply -3v
...
$ make -skj"$(nproc)" ARCH=powerpc CROSS_COMPILE=powerpc64-linux- mrproper ppc64_guest_defconfig vmlinux
$ curl -LSs https://github.com/ClangBuiltLinux/boot-utils/releases/download/20241120-044434/ppc64-rootfs.cpio.zst | zstd -d >rootfs.cpio
$ qemu-system-ppc64 \
-display none \
-nodefaults \
-cpu power8 \
-machine pseries \
-vga none \
-kernel vmlinux \
-initrd rootfs.cpio \
-m 1G \
-serial mon:stdio
Thanks, such a detailed steps to reproduce!
Interestingly, the combination of my compiler (GCC 13.3.0) and
QEMU (8.2.2) don't trigger this bug.
[ 0.000000][ T0] Linux version 7.0.0-rc4-dirty (nathan@framework-amd-ryzen-maxplus-395) (powerpc64-linux-gcc (GCC) 15.2.0, GNU ld (GNU Binutils) 2.45) #1 SMP PREEMPT Thu Mar 19 15:45:53 MST 2026
...
[ 0.216764][ T1] vgaarb: loaded
[ 0.217590][ T1] clocksource: Switched to clocksource timebase
[ 0.221007][ T12] BUG: Kernel NULL pointer dereference at 0x00000010
[ 0.221049][ T12] Faulting instruction address: 0xc00000000044947c
[ 0.221237][ T12] Oops: Kernel access of bad area, sig: 11 [#1]
[ 0.221276][ T12] BE PAGE_SIZE=64K MMU=Hash SMP NR_CPUS=2048 NUMA pSeries
[ 0.221359][ T12] Modules linked in:
[ 0.221556][ T12] CPU: 0 UID: 0 PID: 12 Comm: kworker/u4:0 Not tainted 7.0.0-rc4-dirty #1 PREEMPTLAZY
[ 0.221631][ T12] Hardware name: IBM pSeries (emulated by qemu) POWER8 (architected) 0x4d0200 0xf000004 of:SLOF,HEAD pSeries
[ 0.221765][ T12] Workqueue: trace_init_wq tracer_init_tracefs_work_func
[ 0.222065][ T12] NIP: c00000000044947c LR: c00000000041a584 CTR: c00000000053aa90
[ 0.222084][ T12] REGS: c000000003bc7960 TRAP: 0380 Not tainted (7.0.0-rc4-dirty)
[ 0.222111][ T12] MSR: 8000000000009032 <SF,EE,ME,IR,DR,RI> CR: 44000204 XER: 00000000
[ 0.222287][ T12] CFAR: c000000000449420 IRQMASK: 0
[ 0.222287][ T12] GPR00: c00000000041a584 c000000003bc7c00 c000000001c08100 c000000002892f20
[ 0.222287][ T12] GPR04: c0000000019cfa68 c0000000019cfa60 0000000000000001 0000000000000064
[ 0.222287][ T12] GPR08: 0000000000000002 0000000000000000 c000000003bba000 0000000000000010
[ 0.222287][ T12] GPR12: c00000000053aa90 c000000002c50000 c000000001ab25f8 c000000001626690
[ 0.222287][ T12] GPR16: 0000000000000000 0000000000000000 0000000000000000 0000000000000000
[ 0.222287][ T12] GPR20: c000000001624868 c000000001ab2708 c0000000019cfa08 c000000001a00d18
[ 0.222287][ T12] GPR24: c0000000019cfa18 fffffffffffffef7 c000000003051205 c0000000019cfa68
[ 0.222287][ T12] GPR28: 0000000000000000 c0000000019cfa60 c000000002894e90 0000000000000000
[ 0.222526][ T12] NIP [c00000000044947c] __find_event_file+0x9c/0x110
[ 0.222572][ T12] LR [c00000000041a584] init_tracer_tracefs+0x274/0xcc0
[ 0.222643][ T12] Call Trace:
[ 0.222690][ T12] [c000000003bc7c00] [c000000000b943b0] tracefs_create_file+0x1a0/0x2b0 (unreliable)
[ 0.222766][ T12] [c000000003bc7c50] [c00000000041a584] init_tracer_tracefs+0x274/0xcc0
[ 0.222791][ T12] [c000000003bc7dc0] [c000000002046f1c] tracer_init_tracefs_work_func+0x50/0x320
[ 0.222809][ T12] [c000000003bc7e50] [c000000000276958] process_one_work+0x1b8/0x530
[ 0.222828][ T12] [c000000003bc7f10] [c00000000027778c] worker_thread+0x1dc/0x3d0
[ 0.222883][ T12] [c000000003bc7f90] [c000000000284c44] kthread+0x194/0x1b0
[ 0.222900][ T12] [c000000003bc7fe0] [c00000000000cf30] start_kernel_thread+0x14/0x18
[ 0.222961][ T12] Code: 7c691b78 7f63db78 2c090000 40820018 e89c0000 49107f21 60000000 2c030000 41820048 ebff0000 7c3ff040 41820038 <e93f0010> 7fa3eb78 81490058 e8890018
[ 0.223190][ T12] ---[ end trace 0000000000000000 ]---
...
Interestingly, turning on CONFIG_KASAN appears to hide this, maybe
pointing to some sort of memory corruption (or something timing
related)? If there is any other information I can provide, I am more
than happy to do so.
I don't have much idea on how things end up causing
NULL-pointer-deref... but let's point out suspicious things.
Changes since v7:
- Explicitly initialize the subsystem from start_kernel() right
after mm_core_init() so it is up and running before the creation of
the first mm at boot.
@@ -2078,9 +2082,11 @@ deferred_init_memmap_chunk(unsigned long start_pfn, unsigned long end_pfn,unsignedlongmo_pfn=ALIGN(spfn+1,MAX_ORDER_NR_PAGES);unsignedlongchunk_end=min(mo_pfn,epfn);-nr_pages+=deferred_init_pages(zone,spfn,chunk_end);
Previously, deferred_init_pages() returned nr of pages to add, which is
(end_pfn (= chunk_end) - spfn).
quoted hunk
- deferred_free_pages(spfn, chunk_end - spfn);+ // KHO scratch is MAX_ORDER_NR_PAGES aligned.+ if (!pfn_is_kho_scratch(spfn))+ deferred_init_pages(zone, spfn, chunk_end);
But since the function is not always called with the change,
the calculation is moved to...
quoted hunk
+ deferred_free_pages(spfn, chunk_end - spfn); spfn = chunk_end; if (can_resched)
@@ -2088,6 +2094,7 @@ deferred_init_memmap_chunk(unsigned long start_pfn, unsigned long end_pfn, else touch_nmi_watchdog(); }+ nr_pages += epfn - spfn;
Here.
But this is incorrect, because here we have:
static unsigned long __init
deferred_init_memmap_chunk(unsigned long start_pfn, unsigned long end_pfn,
struct zone *zone, bool can_resched)
{
int nid = zone_to_nid(zone);
unsigned long nr_pages = 0;
phys_addr_t start, end;
u64 i = 0;
for_each_free_mem_range(i, nid, 0, &start, &end, NULL) {
unsigned long spfn = PFN_UP(start);
unsigned long epfn = PFN_DOWN(end);
if (spfn >= end_pfn)
break;
spfn = max(spfn, start_pfn);
epfn = min(epfn, end_pfn);
while (spfn < epfn) {
The loop condition is (spfn < epfn), and by the time the loop terminates...
unsigned long mo_pfn = ALIGN(spfn + 1, MAX_ORDER_NR_PAGES);
unsigned long chunk_end = min(mo_pfn, epfn);
// KHO scratch is MAX_ORDER_NR_PAGES aligned.
if (!pfn_is_kho_scratch(spfn))
deferred_init_pages(zone, spfn, chunk_end);
deferred_free_pages(spfn, chunk_end - spfn);
spfn = chunk_end;
if (can_resched)
cond_resched();
else
touch_nmi_watchdog();
}
nr_pages += epfn - spfn;
epfn - spfn <= 0.
So the number of pages returned by deferred_init_memmap_chunk() becomes
incorrect.
The equivalent translation of what's there before would be doing
`nr_pages += chunk_end - spfn;` within the loop.
--
Cheers,
Harry / Hyeonggon
From: Michał Cłapiński <hidden> Date: 2026-03-20 12:23:27
On Fri, Mar 20, 2026 at 5:18 AM Harry Yoo [off-list ref] wrote:
On Thu, Mar 19, 2026 at 04:37:45PM -0700, Nathan Chancellor wrote:
quoted
Hi all,
I am not really sure whose bug this is, as it only appears when three
seemingly independent patch series are applied together, so I have added
the patch authors and their committers (along with the tracing
maintainers) to this thread. Feel free to expand or reduce that list as
necessary.
Our continuous integration has noticed a crash when booting
ppc64_guest_defconfig in QEMU on the past few -next versions.
https://github.com/ClangBuiltLinux/continuous-integration2/actions/runs/23311154492/job/67811527112
This does not appear to be clang related, as it can be reproduced with
GCC 15.2.0 as well. Through multiple bisects, I was able to land on
applying:
mm: improve RSS counter approximation accuracy for proc interfaces [1]
vdso/datastore: Allocate data pages dynamically [2]
kho: fix deferred init of kho scratch [3]
and their dependent changes on top of 7.0-rc4 is enough to reproduce
this (at least on two of my machines with the same commands). I have
attached the diff from the result of the following 'git apply' commands
below, done in a linux-next checkout.
$ git checkout v7.0-rc4
HEAD is now at f338e7738378 Linux 7.0-rc4
# [1]
$ git diff 60ddf3eed4999bae440d1cf9e5868ccb3f308b64^..087dd6d2cc12c82945ab859194c32e8e977daae3 | git apply -3v
...
# [2]
# Fix trivial conflict in init/main.c around headers
$ git diff dc432ab7130bb39f5a351281a02d4bc61e85a14a^..05988dba11791ccbb458254484826b32f17f4ad2 | git apply -3v
...
# [3]
# Fix conflict in kernel/liveupdate/kexec_handover.c due to lack of kho_mem_retrieve(), just add pfn_is_kho_scratch()
$ git show 4a78467ffb537463486968232daef1e8a2f105e3 | git apply -3v
...
$ make -skj"$(nproc)" ARCH=powerpc CROSS_COMPILE=powerpc64-linux- mrproper ppc64_guest_defconfig vmlinux
$ curl -LSs https://github.com/ClangBuiltLinux/boot-utils/releases/download/20241120-044434/ppc64-rootfs.cpio.zst | zstd -d >rootfs.cpio
$ qemu-system-ppc64 \
-display none \
-nodefaults \
-cpu power8 \
-machine pseries \
-vga none \
-kernel vmlinux \
-initrd rootfs.cpio \
-m 1G \
-serial mon:stdio
Thanks, such a detailed steps to reproduce!
Interestingly, the combination of my compiler (GCC 13.3.0) and
QEMU (8.2.2) don't trigger this bug.
quoted
[ 0.000000][ T0] Linux version 7.0.0-rc4-dirty (nathan@framework-amd-ryzen-maxplus-395) (powerpc64-linux-gcc (GCC) 15.2.0, GNU ld (GNU Binutils) 2.45) #1 SMP PREEMPT Thu Mar 19 15:45:53 MST 2026
...
[ 0.216764][ T1] vgaarb: loaded
[ 0.217590][ T1] clocksource: Switched to clocksource timebase
[ 0.221007][ T12] BUG: Kernel NULL pointer dereference at 0x00000010
[ 0.221049][ T12] Faulting instruction address: 0xc00000000044947c
[ 0.221237][ T12] Oops: Kernel access of bad area, sig: 11 [#1]
[ 0.221276][ T12] BE PAGE_SIZE=64K MMU=Hash SMP NR_CPUS=2048 NUMA pSeries
[ 0.221359][ T12] Modules linked in:
[ 0.221556][ T12] CPU: 0 UID: 0 PID: 12 Comm: kworker/u4:0 Not tainted 7.0.0-rc4-dirty #1 PREEMPTLAZY
[ 0.221631][ T12] Hardware name: IBM pSeries (emulated by qemu) POWER8 (architected) 0x4d0200 0xf000004 of:SLOF,HEAD pSeries
[ 0.221765][ T12] Workqueue: trace_init_wq tracer_init_tracefs_work_func
[ 0.222065][ T12] NIP: c00000000044947c LR: c00000000041a584 CTR: c00000000053aa90
[ 0.222084][ T12] REGS: c000000003bc7960 TRAP: 0380 Not tainted (7.0.0-rc4-dirty)
[ 0.222111][ T12] MSR: 8000000000009032 <SF,EE,ME,IR,DR,RI> CR: 44000204 XER: 00000000
[ 0.222287][ T12] CFAR: c000000000449420 IRQMASK: 0
[ 0.222287][ T12] GPR00: c00000000041a584 c000000003bc7c00 c000000001c08100 c000000002892f20
[ 0.222287][ T12] GPR04: c0000000019cfa68 c0000000019cfa60 0000000000000001 0000000000000064
[ 0.222287][ T12] GPR08: 0000000000000002 0000000000000000 c000000003bba000 0000000000000010
[ 0.222287][ T12] GPR12: c00000000053aa90 c000000002c50000 c000000001ab25f8 c000000001626690
[ 0.222287][ T12] GPR16: 0000000000000000 0000000000000000 0000000000000000 0000000000000000
[ 0.222287][ T12] GPR20: c000000001624868 c000000001ab2708 c0000000019cfa08 c000000001a00d18
[ 0.222287][ T12] GPR24: c0000000019cfa18 fffffffffffffef7 c000000003051205 c0000000019cfa68
[ 0.222287][ T12] GPR28: 0000000000000000 c0000000019cfa60 c000000002894e90 0000000000000000
[ 0.222526][ T12] NIP [c00000000044947c] __find_event_file+0x9c/0x110
[ 0.222572][ T12] LR [c00000000041a584] init_tracer_tracefs+0x274/0xcc0
[ 0.222643][ T12] Call Trace:
[ 0.222690][ T12] [c000000003bc7c00] [c000000000b943b0] tracefs_create_file+0x1a0/0x2b0 (unreliable)
[ 0.222766][ T12] [c000000003bc7c50] [c00000000041a584] init_tracer_tracefs+0x274/0xcc0
[ 0.222791][ T12] [c000000003bc7dc0] [c000000002046f1c] tracer_init_tracefs_work_func+0x50/0x320
[ 0.222809][ T12] [c000000003bc7e50] [c000000000276958] process_one_work+0x1b8/0x530
[ 0.222828][ T12] [c000000003bc7f10] [c00000000027778c] worker_thread+0x1dc/0x3d0
[ 0.222883][ T12] [c000000003bc7f90] [c000000000284c44] kthread+0x194/0x1b0
[ 0.222900][ T12] [c000000003bc7fe0] [c00000000000cf30] start_kernel_thread+0x14/0x18
[ 0.222961][ T12] Code: 7c691b78 7f63db78 2c090000 40820018 e89c0000 49107f21 60000000 2c030000 41820048 ebff0000 7c3ff040 41820038 <e93f0010> 7fa3eb78 81490058 e8890018
[ 0.223190][ T12] ---[ end trace 0000000000000000 ]---
...
Interestingly, turning on CONFIG_KASAN appears to hide this, maybe
pointing to some sort of memory corruption (or something timing
related)? If there is any other information I can provide, I am more
than happy to do so.
I don't have much idea on how things end up causing
NULL-pointer-deref... but let's point out suspicious things.
Changes since v7:
- Explicitly initialize the subsystem from start_kernel() right
after mm_core_init() so it is up and running before the creation of
the first mm at boot.
@@ -2078,9 +2082,11 @@ deferred_init_memmap_chunk(unsigned long start_pfn, unsigned long end_pfn,unsignedlongmo_pfn=ALIGN(spfn+1,MAX_ORDER_NR_PAGES);unsignedlongchunk_end=min(mo_pfn,epfn);-nr_pages+=deferred_init_pages(zone,spfn,chunk_end);
Previously, deferred_init_pages() returned nr of pages to add, which is
(end_pfn (= chunk_end) - spfn).
quoted
- deferred_free_pages(spfn, chunk_end - spfn);+ // KHO scratch is MAX_ORDER_NR_PAGES aligned.+ if (!pfn_is_kho_scratch(spfn))+ deferred_init_pages(zone, spfn, chunk_end);
But since the function is not always called with the change,
the calculation is moved to...
quoted
+ deferred_free_pages(spfn, chunk_end - spfn); spfn = chunk_end; if (can_resched)
@@ -2088,6 +2094,7 @@ deferred_init_memmap_chunk(unsigned long start_pfn, unsigned long end_pfn, else touch_nmi_watchdog(); }+ nr_pages += epfn - spfn;
Here.
But this is incorrect, because here we have:
quoted
static unsigned long __init
deferred_init_memmap_chunk(unsigned long start_pfn, unsigned long end_pfn,
struct zone *zone, bool can_resched)
{
int nid = zone_to_nid(zone);
unsigned long nr_pages = 0;
phys_addr_t start, end;
u64 i = 0;
for_each_free_mem_range(i, nid, 0, &start, &end, NULL) {
unsigned long spfn = PFN_UP(start);
unsigned long epfn = PFN_DOWN(end);
if (spfn >= end_pfn)
break;
spfn = max(spfn, start_pfn);
epfn = min(epfn, end_pfn);
while (spfn < epfn) {
The loop condition is (spfn < epfn), and by the time the loop terminates...
quoted
unsigned long mo_pfn = ALIGN(spfn + 1, MAX_ORDER_NR_PAGES);
unsigned long chunk_end = min(mo_pfn, epfn);
// KHO scratch is MAX_ORDER_NR_PAGES aligned.
if (!pfn_is_kho_scratch(spfn))
deferred_init_pages(zone, spfn, chunk_end);
deferred_free_pages(spfn, chunk_end - spfn);
spfn = chunk_end;
if (can_resched)
cond_resched();
else
touch_nmi_watchdog();
}
nr_pages += epfn - spfn;
epfn - spfn <= 0.
So the number of pages returned by deferred_init_memmap_chunk() becomes
incorrect.
The equivalent translation of what's there before would be doing
`nr_pages += chunk_end - spfn;` within the loop.
Good point, thank you. This patch has already been removed from mm-new.
Changes since v7:
- Explicitly initialize the subsystem from start_kernel() right
after mm_core_init() so it is up and running before the creation of
the first mm at boot.
But how does this work when someone calls mm_cpumask() on init_mm early?
Looks like it will behave incorrectly because get_rss_stat_items_size()
returns zero?
It doesn't work as expected at all. I missed that all users of mm_cpumask()
end up relying on get_rss_stat_items_size(), which now calls
percpu_counter_tree_items_size(), which depends on initialization from
percpu_counter_tree_subsystem_init().
If you add a call to percpu_counter_tree_subsystem_init in
arch/powerpc/kernel/setup_arch() just before:
VM_WARN_ON(cpumask_test_cpu(smp_processor_id(), mm_cpumask(&init_mm)));
cpumask_set_cpu(smp_processor_id(), mm_cpumask(&init_mm));
Does the warning go away ?
Alternatively, would could use a lazy initialization invoking
percpu_counter_tree_subsystem_init from percpu_counter_tree_items_size
when the initialization is not already done.
Any preference ?
Mathieu
--
Mathieu Desnoyers
EfficiOS Inc.
https://www.efficios.com
Changes since v7:
- Explicitly initialize the subsystem from start_kernel() right
after mm_core_init() so it is up and running before the creation of
the first mm at boot.
But how does this work when someone calls mm_cpumask() on init_mm early?
Looks like it will behave incorrectly because get_rss_stat_items_size()
returns zero?
It doesn't work as expected at all. I missed that all users of mm_cpumask()
end up relying on get_rss_stat_items_size(), which now calls
percpu_counter_tree_items_size(), which depends on initialization from
percpu_counter_tree_subsystem_init().
If you add a call to percpu_counter_tree_subsystem_init in
arch/powerpc/kernel/setup_arch() just before:
VM_WARN_ON(cpumask_test_cpu(smp_processor_id(), mm_cpumask(&init_mm)));
cpumask_set_cpu(smp_processor_id(), mm_cpumask(&init_mm));
Does the warning go away ?
Hmm it goes away, but I'm not sure if it is it okay to use nr_cpu_ids
before setup_nr_cpu_ids() is called?
Alternatively, would could use a lazy initialization invoking
percpu_counter_tree_subsystem_init from percpu_counter_tree_items_size
when the initialization is not already done.
So this probably isn't a way to go?
Hmm perhaps we should treat init_mm as a special case in
mm_cpus_allowed() and mm_cpumask().
--
Cheers,
Harry / Hyeonggon
Changes since v7:
- Explicitly initialize the subsystem from start_kernel() right
after mm_core_init() so it is up and running before the creation of
the first mm at boot.
But how does this work when someone calls mm_cpumask() on init_mm early?
Looks like it will behave incorrectly because get_rss_stat_items_size()
returns zero?
It doesn't work as expected at all. I missed that all users of mm_cpumask()
end up relying on get_rss_stat_items_size(), which now calls
percpu_counter_tree_items_size(), which depends on initialization from
percpu_counter_tree_subsystem_init().
If you add a call to percpu_counter_tree_subsystem_init in
arch/powerpc/kernel/setup_arch() just before:
VM_WARN_ON(cpumask_test_cpu(smp_processor_id(), mm_cpumask(&init_mm)));
cpumask_set_cpu(smp_processor_id(), mm_cpumask(&init_mm));
Does the warning go away ?
Hmm it goes away, but I'm not sure if it is it okay to use nr_cpu_ids
before setup_nr_cpu_ids() is called?
AFAIU on powerpc setup_nr_cpu_ids() is called near the end of
smp_setup_cpu_maps(), which is called early in setup_arch,
at least before the two lines which use mm_cpumask.
quoted
Alternatively, would could use a lazy initialization invoking
percpu_counter_tree_subsystem_init from percpu_counter_tree_items_size
when the initialization is not already done.
So this probably isn't a way to go?
I'd favor explicit initialization, so the inter-dependencies are clear.
Hmm perhaps we should treat init_mm as a special case in
mm_cpus_allowed() and mm_cpumask().
I'd prefer not to go there if boot sequence permits and keep things
simple.
I think we're in a situation very similar to tree RCU, here is what
is done in rcu_init_geometry:
static bool initialized;
if (initialized) {
/*
* Warn if setup_nr_cpu_ids() had not yet been invoked,
* unless nr_cpus_ids == NR_CPUS, in which case who cares?
*/
WARN_ON_ONCE(old_nr_cpu_ids != nr_cpu_ids);
return;
}
old_nr_cpu_ids = nr_cpu_ids;
initialized = true;
Thanks,
Mathieu
--
Mathieu Desnoyers
EfficiOS Inc.
https://www.efficios.com
Changes since v7:
- Explicitly initialize the subsystem from start_kernel() right
after mm_core_init() so it is up and running before the creation of
the first mm at boot.
But how does this work when someone calls mm_cpumask() on init_mm early?
Looks like it will behave incorrectly because get_rss_stat_items_size()
returns zero?
It doesn't work as expected at all. I missed that all users of mm_cpumask()
end up relying on get_rss_stat_items_size(), which now calls
percpu_counter_tree_items_size(), which depends on initialization from
percpu_counter_tree_subsystem_init().
If you add a call to percpu_counter_tree_subsystem_init in
arch/powerpc/kernel/setup_arch() just before:
[...]
One thing we could do to catch this kind of init sequence issue
is to add a WARN_ON_ONCE in percpu_counter_tree_items_size:
size_t percpu_counter_tree_items_size(void)
{
if (WARN_ON_ONCE(!nr_cpus_order))
return 0;
return counter_config->nr_items * sizeof(struct percpu_counter_tree_level_item);
}
Thanks,
Mathieu
--
Mathieu Desnoyers
EfficiOS Inc.
https://www.efficios.com
Changes since v7:
- Explicitly initialize the subsystem from start_kernel() right
after mm_core_init() so it is up and running before the
creation of
the first mm at boot.
But how does this work when someone calls mm_cpumask() on init_mm
early?
Looks like it will behave incorrectly because get_rss_stat_items_size()
returns zero?
It doesn't work as expected at all. I missed that all users of
mm_cpumask()
end up relying on get_rss_stat_items_size(), which now calls
percpu_counter_tree_items_size(), which depends on initialization from
percpu_counter_tree_subsystem_init().
If you add a call to percpu_counter_tree_subsystem_init in
arch/powerpc/kernel/setup_arch() just before:
Even though powerpc is showing the warning because of VM_WARN_ON_ONCE(),
but this looks more of a generic problem, where use of mm_cpumask()
before and after percpu_counter_tree_items_size() could lead to
different results (as you also pointed above).
Looks like this is causing regressions in linux-next with warnings
similar to what Harry also pointed out. Do we have any solution for
this, or are we planning to hold on to this patch[1] and maybe even
remove it temporarily from linux-next, until this is fixed?
[1]: https://lore.kernel.org/all/20260227153730.1556542-1-mathieu.desnoyers@efficios.com/
[ 0.000000] WARNING: arch/powerpc/mm/mmu_context.c:106 at switch_mm_irqs_off+0x1a0/0x1d0, CPU#2: swapper/0
[ 0.000000] Modules linked in:
[ 0.000000] CPU: 2 UID: 0 PID: 0 Comm: swapper Not tainted 7.0.0-rc4-next-20260317-00008-g5585e414f073 #4 PREEMPTLAZY
[ 0.000000] Hardware name: IBM PowerNV (emulated by qemu) POWER10 0x801200 opal:v7.1 PowerNV
[ 0.000000] NIP: c00000000008f3b0 LR: c00000000008f330 CTR: c000000000090e20
[ 0.000000] REGS: c000000003cb79b0 TRAP: 0700 Not tainted (7.0.0-rc4-next-20260317-00008-g5585e414f073)
[ 0.000000] MSR: 9000000002021033 <SF,HV,VEC,ME,IR,DR,RI,LE> CR:24022224 XER: 00000000
<...>
[ 0.000000] NIP [c00000000008f3b0] switch_mm_irqs_off+0x1a0/0x1d0
[ 0.000000] LR [c00000000008f330] switch_mm_irqs_off+0x120/0x1d0
[ 0.000000] Call Trace:
[ 0.000000] [c000000003cb7c50] [0500210400000080] 0x500210400000080 (unreliable)
[ 0.000000] [c000000003cb7cb0] [c0000000000ad850] start_using_temp_mm+0x34/0xb0
[ 0.000000] [c000000003cb7cf0] [c0000000000ae8b8] patch_mem+0x110/0x530
[ 0.000000] [c000000003cb7d70] [c000000000077f30] ftrace_modify_code+0x114/0x154
[ 0.000000] [c000000003cb7dd0] [c00000000036a690] ftrace_process_locs+0x408/0x810
[ 0.000000] [c000000003cb7ec0] [c0000000030584ec] ftrace_init+0x68/0x1c4
[ 0.000000] [c000000003cb7f30] [c00000000300d3b8] start_kernel+0x680/0xc44
[ 0.000000] [c000000003cb7fe0] [c00000000000e99c] start_here_common+0x1c/0x20
-ritesh
From: Andrew Morton <akpm@linux-foundation.org> Date: 2026-03-21 02:21:55
On Sat, 21 Mar 2026 06:42:41 +0530 Ritesh Harjani (IBM) [off-list ref] wrote:
Looks like this is causing regressions in linux-next with warnings
similar to what Harry also pointed out. Do we have any solution for
this, or are we planning to hold on to this patch[1] and maybe even
remove it temporarily from linux-next, until this is fixed?
Changes since v7:
- Explicitly initialize the subsystem from start_kernel() right
after mm_core_init() so it is up and running before the creation of
the first mm at boot.
But how does this work when someone calls mm_cpumask() on init_mm early?
Looks like it will behave incorrectly because get_rss_stat_items_size()
returns zero?
It doesn't work as expected at all. I missed that all users of mm_cpumask()
end up relying on get_rss_stat_items_size(), which now calls
percpu_counter_tree_items_size(), which depends on initialization from
percpu_counter_tree_subsystem_init().
If you add a call to percpu_counter_tree_subsystem_init in
arch/powerpc/kernel/setup_arch() just before:
VM_WARN_ON(cpumask_test_cpu(smp_processor_id(), mm_cpumask(&init_mm)));
cpumask_set_cpu(smp_processor_id(), mm_cpumask(&init_mm));
Does the warning go away ?
Hmm it goes away, but I'm not sure if it is it okay to use nr_cpu_ids
before setup_nr_cpu_ids() is called?
AFAIU on powerpc setup_nr_cpu_ids() is called near the end of
smp_setup_cpu_maps(), which is called early in setup_arch,
at least before the two lines which use mm_cpumask.
Right.
quoted
quoted
Alternatively, would could use a lazy initialization invoking
percpu_counter_tree_subsystem_init from percpu_counter_tree_items_size
when the initialization is not already done.
So this probably isn't a way to go?
I'd favor explicit initialization, so the inter-dependencies are clear.
Ack.
quoted
Hmm perhaps we should treat init_mm as a special case in
mm_cpus_allowed() and mm_cpumask().
I'd prefer not to go there if boot sequence permits and keep things
simple.
I think we're in a situation very similar to tree RCU, here is what
is done in rcu_init_geometry:
static bool initialized;
if (initialized) {
/*
* Warn if setup_nr_cpu_ids() had not yet been invoked,
* unless nr_cpus_ids == NR_CPUS, in which case who cares?
*/
WARN_ON_ONCE(old_nr_cpu_ids != nr_cpu_ids);
return;
}
old_nr_cpu_ids = nr_cpu_ids;
initialized = true;
Yeah, as long as nr_cpus_order doesn't change after init,
that will work for HPCC. powerpc seems to be a special case that calls
mm_cpumask() very early in the boot process, so explicitly calling the
init function seems to be fair.
By the way, thinking about it differently - it would probably be simpler
to just eliminate mm_cpumask's dependency on HPCC init dependency by
placing those cpumasks before percpu counter tree items... (but yeah,
that would make mm_struct a bit larger due to alignment requirements)
--
Cheers,
Harry / Hyeonggon
Changes since v7:
- Explicitly initialize the subsystem from start_kernel() right
after mm_core_init() so it is up and running before
the creation of
the first mm at boot.
But how does this work when someone calls mm_cpumask() on
init_mm early?
Looks like it will behave incorrectly because get_rss_stat_items_size()
returns zero?
It doesn't work as expected at all. I missed that all users of
mm_cpumask()
end up relying on get_rss_stat_items_size(), which now calls
percpu_counter_tree_items_size(), which depends on initialization from
percpu_counter_tree_subsystem_init().
If you add a call to percpu_counter_tree_subsystem_init in
arch/powerpc/kernel/setup_arch() just before:
[...]
One thing we could do to catch this kind of init sequence issue
is to add a WARN_ON_ONCE in percpu_counter_tree_items_size:
size_t percpu_counter_tree_items_size(void)
{
if (WARN_ON_ONCE(!nr_cpus_order))
return 0;
return counter_config->nr_items * sizeof(struct percpu_counter_tree_level_item);
On Sat, 21 Mar 2026 06:42:41 +0530 Ritesh Harjani (IBM) [off-list ref] wrote:
quoted
Looks like this is causing regressions in linux-next with warnings
similar to what Harry also pointed out. Do we have any solution for
this, or are we planning to hold on to this patch[1] and maybe even
remove it temporarily from linux-next, until this is fixed?
Yes, I'll disable this patchset.
Hi Andrew,
I have prepared fixes for this issue. On which branch should I rebase
them ? Do you still have the HPCC series in your branch or should I
send it anew ?
Thanks,
Mathieu
--
Mathieu Desnoyers
EfficiOS Inc.
https://www.efficios.com
On Sat, 21 Mar 2026 06:42:41 +0530 Ritesh Harjani (IBM) [off-list ref] wrote:
quoted
Looks like this is causing regressions in linux-next with warnings
similar to what Harry also pointed out. Do we have any solution for
this, or are we planning to hold on to this patch[1] and maybe even
remove it temporarily from linux-next, until this is fixed?
Yes, I'll disable this patchset.
Hi Andrew,
I have prepared fixes for this issue. On which branch should I rebase
them ? Do you still have the HPCC series in your branch or should I
send it anew ?
Cool thanks.
It's best to do a full resend after -rc1 please, presumably against
mainline. Show reviewers the latest version, refresh memories, etc.