Jeremy reported a bug[1] while executing a BPF program containing
timed_may_goto instructions with private stacks.
timed_may_goto[2] is a runtime safety mechanism that allows BPF
programs to execute longer loops[3]. The BPF verifier replaces each
may_goto instruction with a loop counter initialized to 0xffff and a
timestamp check[4] that terminates the loop after 250 ms.
Private stacks[5] allow BPF programs to use per-CPU memory instead of
consuming more of the native kernel stack when BPF programs are deeply
nested.
To make timed may_goto work, the BPF program reserves 16 bytes of stack
space. The first 8 bytes store the loop counter and the next 8 bytes
store the timestamp.
After the loop counter is exhausted, arch_bpf_timed_may_goto() is
called. On x86, it adds the counter's stack offset to RBP to obtain a
pointer to the counter and timestamp[6]. This works when the BPF
program uses the normal stack because RBP is also the BPF frame pointer.
When a private stack is used, the x86 JIT uses R9 as the BPF frame pointer.
The verifier-generated loads and stores therefore access the counter and
timestamp through R9. However, arch_bpf_timed_may_goto() still adds the
offset to RBP and accesses an unrelated location in the native JIT stack
frame. This is the mismatch Jeremy reported.
Fix the mismatch by resolving the address in the generated BPF
instructions:
BPF_REG_AX = BPF_REG_FP
BPF_REG_AX += stack_offset
The JIT can then select the correct BPF frame pointer before calling
arch_bpf_timed_may_goto(). The function receives the resolved pointer
instead of reconstructing it from RBP.
The LoongArch timed may_goto implementation is currently queued through
the loongarch-next tree[7], while its selftests were merged separately
through the bpf-next tree[8]. This series is based on bpf-next and
therefore does not include the LoongArch trampoline update.
[1] https://lore.kernel.org/all/20260824213158.3755932-2-Jeremy.Jean@oss.cyber.gouv.fr/
[2] https://lore.kernel.org/all/20250304003239.2390751-1-memxor@gmail.com/
[3] https://elixir.bootlin.com/linux/v7.2.2/source/tools/testing/selftests/bpf/libarena/include/bpf_may_goto.h#L8
[4] https://elixir.bootlin.com/linux/v7.2.2/source/kernel/bpf/core.c#L3407
[5] https://lore.kernel.org/bpf/20260417034658.2625353-1-yonghong.song@linux.dev/
[6] https://elixir.bootlin.com/linux/v7.2.2/source/arch/x86/net/bpf_timed_may_goto.S#L18
[7] https://lore.kernel.org/loongarch/20260804153938.16129-3-dongtai.guo@linux.dev/
[8] https://lore.kernel.org/bpf/20260813070906.5164-1-yangtiezhu@loongson.cn/
Siddharth Chintamaneni (7):
bpf: Fix timed may_goto stack pointer for private stacks
bpf, x86: Use resolved pointer for timed may_goto
bpf, arm64: Use resolved pointer for timed may_goto
bpf, powerpc64: Use resolved pointer for timed may_goto
bpf, riscv: Use resolved pointer for timed may_goto
bpf, s390: Use resolved pointer for timed may_goto
selftests/bpf: Test timed may_goto with private stacks
arch/arm64/net/bpf_timed_may_goto.S | 12 ++------
arch/powerpc/net/bpf_timed_may_goto.S | 8 ++---
arch/riscv/net/bpf_timed_may_goto.S | 13 ++++----
arch/s390/net/bpf_jit_comp.c | 6 ++--
arch/s390/net/bpf_timed_may_goto.S | 8 ++---
arch/x86/net/bpf_timed_may_goto.S | 6 ----
kernel/bpf/fixups.c | 19 ++++++------
.../bpf/progs/verifier_bpf_fastcall.c | 30 ++++++++++---------
.../selftests/bpf/progs/verifier_may_goto_1.c | 17 ++++++-----
.../bpf/progs/verifier_private_stack.c | 19 ++++++++++++
10 files changed, 75 insertions(+), 63 deletions(-)
base-commit: d761934c9483ecde93fe99d8705282f716dfee50
--
2.43.0
The timed may_goto fixup now passes the resolved counter pointer through
BPF_REG_AX instead of a stack offset and BPF frame pointer pair.
Copy the pointer directly into the first argument register and update the
special calling convention documentation.
Fixes: b8efa810c1db ("s390/bpf: Add s390 JIT support for timed may_goto")
Reported-by: Jeremy Jean <redacted>
Link: https://lore.kernel.org/all/20260824213158.3755932-2-Jeremy.Jean@oss.cyber.gouv.fr/
Assisted-by: Copilot:gpt-5.6-sol
Signed-off-by: Siddharth Chintamaneni <redacted>
---
arch/s390/net/bpf_jit_comp.c | 6 +++---
arch/s390/net/bpf_timed_may_goto.S | 8 ++++----
2 files changed, 7 insertions(+), 7 deletions(-)
@@ -21,9 +21,9 @@SYM_FUNC_START(arch_bpf_timed_may_goto)/*-*ThisfunctionhasaspecialABI:theparametersarein%r12and-*%r13; the return value is in %r12;allGPRsexcept%r0, %r1,and-*%r12 are callee-saved; and the return address is in %r0.+*ThisfunctionhasaspecialABI:theparameterandreturnvalueare+*in%r12; all GPRs except %r0,%r1, and %r12arecallee-saved;and+*thereturnaddressisin%r0.*/stmg%r2,%r5,FRAME_OFF+R2_OFF(%r15)stg%r14,FRAME_OFF+R14_OFF(%r15)
The timed may_goto fixup now passes the resolved counter pointer through
BPF_REG_AX instead of a stack offset.
Copy the pointer directly into the first argument register rather than
adding it to the BPF frame pointer again.
Fixes: 16175375da36 ("bpf, arm64: Add JIT support for timed may_goto")
Reported-by: Jeremy Jean <redacted>
Link: https://lore.kernel.org/all/20260824213158.3755932-2-Jeremy.Jean@oss.cyber.gouv.fr/
Assisted-by: Copilot:gpt-5.6-sol
Signed-off-by: Siddharth Chintamaneni <redacted>
---
arch/arm64/net/bpf_timed_may_goto.S | 12 +++---------
1 file changed, 3 insertions(+), 9 deletions(-)
Add private_stack_timed_may_goto() with enough stack usage to select a
private stack. Check that the JIT uses the private-stack frame pointer
when constructing the pointer passed to the architecture trampoline.
Update the translated instruction expectations in may_goto_batch_2(),
may_goto_interaction_x86_64(), and may_goto_interaction() for the
additional pointer construction instruction and adjusted branch offsets.
Signed-off-by: Siddharth Chintamaneni <redacted>
---
.../bpf/progs/verifier_bpf_fastcall.c | 30 ++++++++++---------
.../selftests/bpf/progs/verifier_may_goto_1.c | 17 ++++++-----
.../bpf/progs/verifier_private_stack.c | 19 ++++++++++++
3 files changed, 44 insertions(+), 22 deletions(-)
The timed may_goto fixup now passes the resolved counter pointer through
BPF_REG_AX instead of a stack offset.
Copy the pointer directly into the first argument register rather than
adding it to the BPF frame pointer again.
Fixes: 6ef8ff20c30b ("bpf, riscv: Add support for timed may_goto")
Reported-by: Jeremy Jean <redacted>
Link: https://lore.kernel.org/all/20260824213158.3755932-2-Jeremy.Jean@oss.cyber.gouv.fr/
Assisted-by: Copilot:gpt-5.6-sol
Signed-off-by: Siddharth Chintamaneni <redacted>
---
arch/riscv/net/bpf_timed_may_goto.S | 13 ++++++++-----
1 file changed, 8 insertions(+), 5 deletions(-)
The timed may_goto fixup now passes the resolved counter pointer through
BPF_REG_AX instead of a stack offset.
Copy the pointer directly into the first argument register rather than
adding it to the BPF frame pointer again.
Fixes: b55b6b9ad76c ("powerpc64/bpf: Add powerpc64 JIT support for timed may_goto")
Reported-by: Jeremy Jean <redacted>
Link: https://lore.kernel.org/all/20260824213158.3755932-2-Jeremy.Jean@oss.cyber.gouv.fr/
Assisted-by: Copilot:gpt-5.6-sol
Signed-off-by: Siddharth Chintamaneni <redacted>
---
arch/powerpc/net/bpf_timed_may_goto.S | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
timed may_goto passes a stack offset to the architecture trampoline,
which reconstructs the counter pointer from its BPF frame pointer. This
breaks when the JIT uses a private stack with a different frame pointer.
Resolve the counter pointer in the fixup using BPF_REG_FP and pass the
pointer through BPF_REG_AX. Account for the extra instruction in the
internal branch offsets.
Fixes: e723608bf428 ("bpf: Add verifier support for timed may_goto")
Reported-by: Jeremy Jean <redacted>
Link: https://lore.kernel.org/all/20260824213158.3755932-2-Jeremy.Jean@oss.cyber.gouv.fr/
Signed-off-by: Siddharth Chintamaneni <redacted>
---
kernel/bpf/fixups.c | 19 +++++++++----------
1 file changed, 9 insertions(+), 10 deletions(-)
The timed may_goto fixup now passes the resolved counter pointer through
BPF_REG_AX instead of a stack offset.
Use the pointer directly rather than adding it to RBP. This preserves the
private-stack address selected by the JIT through R9.
Fixes: 2fb761823ead ("bpf, x86: Add x86 JIT support for timed may_goto")
Reported-by: Jeremy Jean <redacted>
Link: https://lore.kernel.org/all/20260824213158.3755932-2-Jeremy.Jean@oss.cyber.gouv.fr/
Signed-off-by: Siddharth Chintamaneni <redacted>
---
arch/x86/net/bpf_timed_may_goto.S | 6 ------
1 file changed, 6 deletions(-)
bpf, riscv: Use resolved pointer for timed may_goto
The timed may_goto fixup now passes the resolved counter pointer through
BPF_REG_AX instead of a stack offset.
Copy the pointer directly into the first argument register rather than
adding it to the BPF frame pointer again.
Fixes: 6ef8ff20c30b ("bpf, riscv: Add support for timed may_goto")
Should the Fixes: tag point to d8319a04dafc ("bpf: Fix timed may_goto
stack pointer for private stacks") instead? That commit changed the
calling convention from passing a stack offset to passing a resolved
pointer through BPF_REG_AX, which is what broke the riscv trampoline.
The original riscv implementation (6ef8ff20c30b) worked correctly with
the original calling convention. riscv64 cannot hit the private-stack bug
that d8319a04dafc targeted: it does not implement
bpf_jit_supports_private_stack(), so the __weak default in
kernel/bpf/core.c returns false, and private stacks are gated on it in
kernel/bpf/verifier.c. On riscv, regmap[BPF_REG_FP] = RV_REG_S5 is the
one and only BPF frame pointer, and the pre-patch 'add a0, t0, s5'
computed exactly the right address.
What this patch actually does on riscv is adapt to the new BPF_REG_AX
calling convention introduced by d8319a04dafc, which is patch 1 of the
same series. If stable/AUTOSEL tooling picks this commit up on the
strength of its Fixes: tag without also taking d8319a04dafc, BPF_REG_AX
still holds the raw immediate from the old fixup (BPF_MOV64_IMM(BPF_REG_AX,
stack_off_cnt), i.e. -stack_depth-16, up to -528):
mv a0, t0 /* a0 = -528, not a pointer */
call bpf_check_timed_may_goto
bpf_check_timed_may_goto() then dereferences p->timestamp
(kernel/bpf/core.c) at 0xfffffffffffffdf0, oopsing on every timed
may_goto loop.
Should the changelog also declare that d8319a04dafc is a prerequisite
for this patch?
Because d8319a04dafc lands first and each architecture is converted in a
later commit, riscv64 (and arm64, powerpc64, s390) BPF timed may_goto is
broken at every intermediate commit of the series. Does the series need to
be structured differently to remain bisectable for those architectures?
A subsystem pattern flags this as potentially concerning: This commit
changes the x86 trampoline to the new 'pointer in BPF_REG_AX' ABI, but the
producer side of that ABI was changed one commit earlier, in d8319a04dafc
("bpf: Fix timed may_goto stack pointer for private stacks"), which
rewrote kernel/bpf/fixups.c to emit:
insn_buf[4] = BPF_MOV64_REG(BPF_REG_AX, BPF_REG_FP);
insn_buf[5] = BPF_ALU64_IMM(BPF_ADD, BPF_REG_AX, stack_off_cnt);
At commit d8319a04dafc (i.e. HEAD~1), does x86 still execute
'leaq (%rbp, %r10, 1), %r10' on a register that already holds a resolved
pointer, producing r10 = rbp + (frame_ptr + stack_off_cnt) -- roughly
2*rbp, a non-canonical address?
bpf_check_timed_may_goto() then reads and writes p->timestamp / p->count
through it (kernel/bpf/core.c:3410). Would any BPF program containing
may_goto oops on x86_64 at that commit?
The same window exists for arm64, riscv, powerpc64 and s390, whose
trampolines are only converted in the four later commits d35642770d98,
7053db7d3c0d, 6e03227c5677 and 20356163d51e. Is the tree bisectable across
the series?
The two halves carry different Fixes: tags (this one 2fb761823ead, the
fixups.c one e723608bf428), which invites a partial stable backport that
reintroduces exactly this wild-pointer write. Would squashing the fixups.c
ABI change with the arch trampoline updates, or ordering all arch updates
before the generic change, address the bisectability concern?
/* Setup frame. */
pushq %rbp
movq %rsp, %rbp
The x86-only __xlated expectations for the timed may_goto expansion in
tools/testing/selftests/bpf/progs/verifier_bpf_fastcall.c were never
updated for the new 7-insn -> 8-insn sequence. The test expects:
__xlated("7: if r12 == 0x0 goto pc+6")
__xlated("8: r12 -= 1")
__xlated("9: if r12 != 0x0 goto pc+2")
__xlated("10: r12 = -24")
__xlated("11: call unknown")
__xlated("12: *(u64 *)(r10 -24) = r12")
Does this test fail on x86_64 for the whole series? Every line from index
7 onward is now wrong: the first branch becomes 'goto pc+7' (insn->off + 6
instead of insn->off + 5), the second becomes 'goto pc+3' instead of
'pc+2', 'r12 = -24' is replaced by the pair 'r12 = r10' / 'r12 += -24',
and the trailing indices shift by one ('call unknown' at 12, the store at
13, then 14 and 15).
git log 28bd77eb835b..8dc5c7d5340d shows the only selftest commit in the
series is 8dc5c7d5340d ("selftests/bpf: Test timed may_goto with private
stacks"), which adds new coverage but does not touch
verifier_bpf_fastcall.c. Note the original x86 JIT enablement commit
2fb761823ead -- the commit this patch lists in its Fixes: tag -- adjusted
this very test for the same reason, so the update belongs with the ABI
change.
commit 28bd77eb835bc1a95190315e93afbc6bb186b9ff
Author: Siddharth Chintamaneni [off-list ref]
bpf, x86: Use resolved pointer for timed may_goto
The timed may_goto fixup now passes the resolved counter pointer through
BPF_REG_AX instead of a stack offset.
Use the pointer directly rather than adding it to RBP. This preserves the
private-stack address selected by the JIT through R9.
Fixes: 2fb761823ead ("bpf, x86: Add x86 JIT support for timed may_goto")
^^^^^^^^^^^
The existing Fixes: tag points to commit 2fb761823ead, which is not an
ancestor of this fix and is on a different branch (bpf-next/res_spin_lock).
The bug was actually introduced by commit 47ab38f3f4d9 in this branch's
history, which added the incorrect code 'leaq (%rbp, %r10, 1), %r10' that
overwrites r10 instead of preserving it.
Should this use:
Fixes: 47ab38f3f4d9 ("adding ci files")
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/33914557238
commit 20356163d51e7b8551f96780e7e969a26767c5c4
Author: Siddharth Chintamaneni [off-list ref]
AuthorDate: Fri Sep 4 19:51:31 2026 +0000
Commit: Kernel Patches Daemon [off-list ref]
CommitDate: Fri Sep 4 13:03:02 2026 -0700
bpf, s390: Use resolved pointer for timed may_goto
The timed may_goto fixup now passes the resolved counter pointer through
BPF_REG_AX instead of a stack offset and BPF frame pointer pair.
Copy the pointer directly into the first argument register and update the
special calling convention documentation.
Fixes: b8efa810c1db ("s390/bpf: Add s390 JIT support for timed may_goto")
Reported-by: Jeremy Jean [off-list ref]
Link: https://lore.kernel.org/all/20260824213158.3755932-2-Jeremy.Jean@oss.cyber.gouv.fr/
Assisted-by: Copilot:gpt-5.6-sol
Signed-off-by: Siddharth Chintamaneni [off-list ref]
Should the Fixes: tag point at b8efa810c1db?
At that commit the s390 trampoline was correct for the ABI in force at the
time. kernel/bpf/fixups.c emitted `insn_buf[4] = BPF_MOV64_IMM(BPF_REG_AX,
stack_off_cnt)`, so %r12 held a signed byte offset and `lay
%r2,0(%r12,%r13)` correctly resolved it against the s390 BPF frame pointer
%r13. Nothing was wrong with b8efa810c1db.
What makes the old code wrong is the parent commit in this series,
d8319a04dafc ("bpf: Fix timed may_goto stack pointer for private stacks"),
which replaced that with:
insn_buf[4] = BPF_MOV64_REG(BPF_REG_AX, BPF_REG_FP);
insn_buf[5] = BPF_ALU64_IMM(BPF_ADD, BPF_REG_AX, stack_off_cnt)
so that %r12 now carries an already-resolved pointer. Only after that does
`lay %r2,0(%r12,%r13)` double-add the frame pointer.
If a stable maintainer or AUTOSEL picks 20356163d51e on the strength of
its Fixes: tag without also taking d8319a04dafc, s390 is left with `lgr
%r2,%r12` while the fixup still puts a raw offset (stack_off_cnt =
-stack_depth - 16, e.g. -528) in %r12:
bpf_check_timed_may_goto(struct bpf_timed_may_goto *p)
then does `p->count` / `p->timestamp` loads and stores through a small
negative integer treated as a kernel pointer, which would trigger an oops
or memory corruption on every timed may_goto timeout.
The mirror case is equally bad: d8319a04dafc carries `Fixes:
e723608bf428` and will be selected for stable on its own, leaving s390
(and arm64/ppc64/riscv) double-adding the frame pointer.
The generic ABI change and its five arch consumers are not independently
correct, so either they belong in one commit, or the arch patches should
carry the same `Fixes: e723608bf428 ("bpf: Add verifier support for timed
may_goto")` and a `Cc: [off-list ref] # <sha>: <generic
commit>` prerequisite note. No Cc: stable and no prerequisite note is
present today.
Note: the same pattern is used by all four sibling commits (28bd77eb835b
x86, d35642770d98 arm64, 6e03227c5677 ppc64, 7053db7d3c0d riscv), so this
is a series-wide tagging decision rather than an s390-only slip.
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/33914557238
A subsystem pattern flags this as potentially concerning: This commit
changes the x86 trampoline to the new 'pointer in BPF_REG_AX' ABI, but the
producer side of that ABI was changed one commit earlier, in d8319a04dafc
("bpf: Fix timed may_goto stack pointer for private stacks"), which
rewrote kernel/bpf/fixups.c to emit:
insn_buf[4] = BPF_MOV64_REG(BPF_REG_AX, BPF_REG_FP);
insn_buf[5] = BPF_ALU64_IMM(BPF_ADD, BPF_REG_AX, stack_off_cnt);
At commit d8319a04dafc (i.e. HEAD~1), does x86 still execute
'leaq (%rbp, %r10, 1), %r10' on a register that already holds a resolved
pointer, producing r10 = rbp + (frame_ptr + stack_off_cnt) -- roughly
2*rbp, a non-canonical address?
bpf_check_timed_may_goto() then reads and writes p->timestamp / p->count
through it (kernel/bpf/core.c:3410). Would any BPF program containing
may_goto oops on x86_64 at that commit?
The same window exists for arm64, riscv, powerpc64 and s390, whose
trampolines are only converted in the four later commits d35642770d98,
7053db7d3c0d, 6e03227c5677 and 20356163d51e. Is the tree bisectable across
the series?
The two halves carry different Fixes: tags (this one 2fb761823ead, the
fixups.c one e723608bf428), which invites a partial stable backport that
reintroduces exactly this wild-pointer write. Would squashing the fixups.c
ABI change with the arch trampoline updates, or ordering all arch updates
before the generic change, address the bisectability concern?
quoted
/* Setup frame. */
pushq %rbp
movq %rsp, %rbp
I tried to follow the same pattern used when the patches were
introduced, but this is a valid concern.
Should I just squash all the patches to one then?
The x86-only __xlated expectations for the timed may_goto expansion in
tools/testing/selftests/bpf/progs/verifier_bpf_fastcall.c were never
updated for the new 7-insn -> 8-insn sequence. The test expects:
__xlated("7: if r12 == 0x0 goto pc+6")
__xlated("8: r12 -= 1")
__xlated("9: if r12 != 0x0 goto pc+2")
__xlated("10: r12 = -24")
__xlated("11: call unknown")
__xlated("12: *(u64 *)(r10 -24) = r12")
I've fixed this in the selftests patch.
Does this test fail on x86_64 for the whole series? Every line from index
7 onward is now wrong: the first branch becomes 'goto pc+7' (insn->off + 6
instead of insn->off + 5), the second becomes 'goto pc+3' instead of
'pc+2', 'r12 = -24' is replaced by the pair 'r12 = r10' / 'r12 += -24',
and the trailing indices shift by one ('call unknown' at 12, the store at
13, then 14 and 15).
git log 28bd77eb835b..8dc5c7d5340d shows the only selftest commit in the
series is 8dc5c7d5340d ("selftests/bpf: Test timed may_goto with private
stacks"), which adds new coverage but does not touch
verifier_bpf_fastcall.c. Note the original x86 JIT enablement commit
2fb761823ead -- the commit this patch lists in its Fixes: tag -- adjusted
this very test for the same reason, so the update belongs with the ABI
change.
quoted
commit 28bd77eb835bc1a95190315e93afbc6bb186b9ff
Author: Siddharth Chintamaneni [off-list ref]
bpf, x86: Use resolved pointer for timed may_goto
The timed may_goto fixup now passes the resolved counter pointer through
BPF_REG_AX instead of a stack offset.
Use the pointer directly rather than adding it to RBP. This preserves the
private-stack address selected by the JIT through R9.
Fixes: 2fb761823ead ("bpf, x86: Add x86 JIT support for timed may_goto")
^^^^^^^^^^^
The existing Fixes: tag points to commit 2fb761823ead, which is not an
ancestor of this fix and is on a different branch (bpf-next/res_spin_lock).
I think we can keep this fix local to x86. Would be less churn in the end.
The problem is that x86 switches the BPF frame pointer from RBP to R9 for
private stacks, but the trampoline still uses RBP. arm64 and powerpc64 keep
using X25 and R31 respectively for both stack modes, so their trampolines
already use the right frame pointer.
Instead of changing the generic fixup, can we resolve the pointer in the x86 JIT
when emitting the call to arch_bpf_timed_may_goto()? We already know whether
priv_frame_ptr is set there. R10 contains the offset passed through BPF_REG_AX,
so we can emit:
/* Private stack */
leaq (%r9, %r10), %r10
/* Normal stack */
leaq (%rbp, %r10), %r10
Then remove the LEA from the x86 trampoline, as patch 2 already does.
This gives bpf_check_timed_may_goto() the same address used by the generated
loads and stores: the BPF frame pointer plus the counter’s stack offset. For
normal stacks, we just move the existing calculation into the JIT. For private
stacks, we use R9 instead of RBP, which fixes the mismatch.
The existing save/restore of R9 around the call should stay. We also leave RBP
alone, since it is still needed for the native frame chain and unwinding.
I haven’t tested this approach, so it still needs checking with both normal and
private stacks.
pw-bot: cr
Siddharth Chintamaneni (7):
bpf: Fix timed may_goto stack pointer for private stacks
bpf, x86: Use resolved pointer for timed may_goto
bpf, arm64: Use resolved pointer for timed may_goto
bpf, powerpc64: Use resolved pointer for timed may_goto
bpf, riscv: Use resolved pointer for timed may_goto
bpf, s390: Use resolved pointer for timed may_goto
selftests/bpf: Test timed may_goto with private stacks
arch/arm64/net/bpf_timed_may_goto.S | 12 ++------
arch/powerpc/net/bpf_timed_may_goto.S | 8 ++---
arch/riscv/net/bpf_timed_may_goto.S | 13 ++++----
arch/s390/net/bpf_jit_comp.c | 6 ++--
arch/s390/net/bpf_timed_may_goto.S | 8 ++---
arch/x86/net/bpf_timed_may_goto.S | 6 ----
kernel/bpf/fixups.c | 19 ++++++------
.../bpf/progs/verifier_bpf_fastcall.c | 30 ++++++++++---------
.../selftests/bpf/progs/verifier_may_goto_1.c | 17 ++++++-----
.../bpf/progs/verifier_private_stack.c | 19 ++++++++++++
10 files changed, 75 insertions(+), 63 deletions(-)
base-commit: d761934c9483ecde93fe99d8705282f716dfee50
I think we can keep this fix local to x86. Would be less churn in the end.
The problem is that x86 switches the BPF frame pointer from RBP to R9 for
private stacks, but the trampoline still uses RBP. arm64 and powerpc64 keep
using X25 and R31 respectively for both stack modes, so their trampolines
already use the right frame pointer.
Instead of changing the generic fixup, can we resolve the pointer in the x86 JIT
when emitting the call to arch_bpf_timed_may_goto()? We already know whether
priv_frame_ptr is set there. R10 contains the offset passed through BPF_REG_AX,
so we can emit:
/* Private stack */
leaq (%r9, %r10), %r10
/* Normal stack */
leaq (%rbp, %r10), %r10
Then remove the LEA from the x86 trampoline, as patch 2 already does.
make sense! I thought that the remaining arch's have a similar private
stacks approach to x86. I'll give it a try and re-spin the patch.
This gives bpf_check_timed_may_goto() the same address used by the generated
loads and stores: the BPF frame pointer plus the counter’s stack offset. For
normal stacks, we just move the existing calculation into the JIT. For private
stacks, we use R9 instead of RBP, which fixes the mismatch.
The existing save/restore of R9 around the call should stay. We also leave RBP
alone, since it is still needed for the native frame chain and unwinding.
I haven’t tested this approach, so it still needs checking with both normal and
private stacks.
pw-bot: cr
quoted
Siddharth Chintamaneni (7):
bpf: Fix timed may_goto stack pointer for private stacks
bpf, x86: Use resolved pointer for timed may_goto
bpf, arm64: Use resolved pointer for timed may_goto
bpf, powerpc64: Use resolved pointer for timed may_goto
bpf, riscv: Use resolved pointer for timed may_goto
bpf, s390: Use resolved pointer for timed may_goto
selftests/bpf: Test timed may_goto with private stacks
arch/arm64/net/bpf_timed_may_goto.S | 12 ++------
arch/powerpc/net/bpf_timed_may_goto.S | 8 ++---
arch/riscv/net/bpf_timed_may_goto.S | 13 ++++----
arch/s390/net/bpf_jit_comp.c | 6 ++--
arch/s390/net/bpf_timed_may_goto.S | 8 ++---
arch/x86/net/bpf_timed_may_goto.S | 6 ----
kernel/bpf/fixups.c | 19 ++++++------
.../bpf/progs/verifier_bpf_fastcall.c | 30 ++++++++++---------
.../selftests/bpf/progs/verifier_may_goto_1.c | 17 ++++++-----
.../bpf/progs/verifier_private_stack.c | 19 ++++++++++++
10 files changed, 75 insertions(+), 63 deletions(-)
base-commit: d761934c9483ecde93fe99d8705282f716dfee50