Re: early boot crash in init_tracer_tracefs
flat view
From: Amit Machhiwal <hidden>
Date: 2026-10-09 16:43:48
On 2026/09/30 12:46 PM, Michal Suchánek wrote:
On Wed, Sep 30, 2026 at 01:19:25AM +0530, Amit Machhiwal wrote:quoted
Hi Michal, Thanks for sharing the host information and the full guest boot log. On 2026/09/29 02:18 PM, Michal Suchánek wrote:quoted
On Tue, Sep 29, 2026 at 04:39:05PM +0530, Amit Machhiwal wrote:quoted
On 2026/09/29 01:04 PM, Michal Suchánek wrote:quoted
On Tue, Sep 29, 2026 at 09:36:05AM +0530, Amit Machhiwal wrote:quoted
On 2026/09/28 11:43 AM, Amit Machhiwal wrote:quoted
Hi Michal, Thanks for the report. We have been looking at the crash and have some findings and follow-up questions. On 2026/09/25 10:10 AM, Michal Suchánek wrote:quoted
Hello, There appears to be a regression between< snip >quoted
I have been trying to recreate the issue with this kernel and config but I haven't been able to. The L2 KVM guest boots fine everytime. In addition to the requested information, could you please also share your qemu cmdline/guest xml you used?The XML of the VM is below. What other information do you require?Thanks for sharing the guest XML. I had requested some more info [1]. Could you please share that? [1] https://lore.kernel.org/all/20260928111234.faec09f7-8d-amachhiw@linux.ibm.com/ (local)Hello, full log below. host: numactl --hardware available: 2 nodes (0-1) node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 node 0 size: 122139 MB node 0 free: 103635 MB node 1 cpus: 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 node 1 size: 139516 MB node 1 free: 128483 MB node distances: node 0 1 0: 10 20 1: 20 10 The patch does indeed make it possible to boot when the buffer allocation fails.Glad the first fix works. That fixes the crash (the NULL deref in __find_event_file()), but the ring buffer allocation failure itself is a separate bug worth fixing too. My suspicion is that with CONFIG_DEFERRED_STRUCT_PAGE_INIT=y, si_mem_available() returns falsely negative during early boot because NR_FREE_PAGES only reflects the non-deferred memory pool at that point — the bulk of RAM hasn't been handed to the buddy allocator yet. The check in __rb_allocate_pages() rejects the allocation prematurely. [ 0.149312][ T743] node 0 deferred pages initialised in 20ms Since the issue is not recreating on my environment, could you please give this patch a try and see if the actual crash goes away?diff --git a/kernel/trace/ring_buffer.c b/kernel/trace/ring_buffer.c index 04bb94c29f58..224cc0e5c066 100644 --- a/kernel/trace/ring_buffer.c +++ b/kernel/trace/ring_buffer.c@@ -2454,7 +2454,7 @@ static int __rb_allocate_pages(struct ring_buffer_per_cpu *cpu_buffer, * not going to succeed. */ i = si_mem_available(); - if (i < nr_pages) + if (system_state != SYSTEM_BOOTING && i < nr_pages) return -ENOMEM; /*Hello, this does seem to mitigate the problem with allocating the trace buffer. Related to reproducing the problem it seems that it is much more likely to happen on reboot than first boot, and the typical solution for kernel test automation produces one-off VMs that never reboot. With the VM XML as posted earlier using a full production OS image testing is not very efficient. Still with something like n=1 ; while true ; do echo $n ; n=$(expr $n + 1) ; ssh 192.168.0.2 reboot ; sleep 10 ; ssh 192.168.0.2 dmesg | { grep ERROR: && exit 1 ; } ; done the VM can be rebooted a number of times without additional user interaction.
Hi Michal, Apologies for the delayed response, as I was caught up with some higher-priority work items. Thanks for testing and confirming both fixes, and for sharing your reboot loop reproduction snippet. I have submitted the formal 2-patch series fixing both the tracefs workqueue NULL dereference crash and the early boot ring buffer allocation failure: https://lore.kernel.org/all/20261009163557.88467-1-amachhiw@linux.ibm.com/ (local) Thanks again for reporting and helping verify the fixes! Thanks, Amit