Thread (20 messages) flat view 20 messages, 7 authors, 16d ago

Re: [PATCH 2/2] perf/x86: Disable precise sampling for PERF_SAMPLE_STACK_USER

From: Mi, Dapeng <hidden>
Date: 2026-09-09 00:59:46
Also in: lkml

On 9/9/2026 4:56 AM, Ian Rogers wrote:
On Tue, Sep 8, 2026 at 8:10 AM Andi Kleen [off-list ref] wrote:
quoted
quoted
That's what we already do, no? I have distinct memories of making the
stack unwind use the NMI regs rather then the PEBS regs.
quoted
In my opinion, it could even make the thing worse. User
requires to get precise samplings, but perf silently returns imprecise
records, this would mislead user.
Mostly just the unwind might be off a little, the rest is accurate. This
has been the case 'forever'. Performance analysis isn't for silly
people, if they can't deal with a little fuzz then perhaps they're in
the wrong business.
Is the main problem that the stack doesn't agree? Perhaps there
could be a check for regs->rsp == pebs->user rsp (if in user space)
to detect problematic samples.

The question is how to report it and who should do the checking.

It may need new fields in the ABI either to communicate the extra PEBS RSP
or a bit to indicate that there might be a mismatch.

I guess checking in the kernel and reporting an error might be simpler
and maybe cleaner, but it would likely limit more advanced recovery
possibilities.

Are there other mismatches that break the unwinding? Perhaps the same
for RBP?
For DWARF unwinding any register may be the source of a frame pointer
(e.g. the OpenSSL library would use R11 rather than RBP).

There is redundancy on x86 you can sample the PERF_REG_X86_IP register
in the user register and there is PERF_SAMPLE_IP in the sample event
itself.

My understanding is that IBS can only sample IP and so for precise
samples we can use PERF_SAMPLE_IP as the precise location and the user
register PERF_REG_X86_IP as the interrupt IP - this would match the
other register values in the interrupt.

In DWARF unwinding, we initialize the register state using the sampled
user registers:
https://web.git.kernel.org/pub/scm/linux/kernel/git/perf/perf-tools-next.git/tree/tools/perf/util/unwind-libdw.c?h=perf-tools-next#n270
and on x86 we sample all registers for DWARF unwinding:
https://web.git.kernel.org/pub/scm/linux/kernel/git/perf/perf-tools-next.git/tree/tools/perf/arch/x86/include/perf_regs.h?h=perf-tools-next#n20
https://web.git.kernel.org/pub/scm/linux/kernel/git/perf/perf-tools-next.git/tree/tools/perf/util/perf-regs-arch/perf_regs_x86.c?h=perf-tools-next#n238

Having PERF_SAMPLE_IP be precise and the user registers from the
interrupt I believe works for AMD IBS and ARM SPE, but for Intel PEBS
there is the ability to use the PEBS register samples for other
non-redundant registers. In the x86 driver could we disable PEBS
sampling for these registers when doing user stack sampling, so that
the sampled user registers match the stack sample? We can keep the
PERF_SAMPLE_IP precise, and make all the registers precise when there
is no stack sampling.
That sounds the best way to fix this issue by decoupling PERF_SAMPLE_IP
with PERF_REG_X86_IP. Then we can keep the precise SAMPLE_IP and the PMI
context user register snapshot simultaneously.

I would post v2 patch with this fix. Thanks.

Thanks,
Ian
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help