powerpc unconditionally selects GENERIC_ENTRY. The GENERIC_ENTRY
infrastructure relies on the compiler emitting __asan_mem*() calls at
instrumented mem*() sites rather than plain memset/memcpy/memmove, so
that entry/exit paths calling those functions are not instrumented.
This assumption is encoded in two places:
mm/kasan/shadow.c:
#if !defined(CONFIG_CC_HAS_KASAN_MEMINTRINSIC_PREFIX) &&
!defined(CONFIG_GENERIC_ENTRY)
include/linux/fortify-string.h:
#if !defined(CONFIG_CC_HAS_KASAN_MEMINTRINSIC_PREFIX) &&
!defined(CONFIG_GENERIC_ENTRY)
When GENERIC_ENTRY is set, both guards suppress the C wrappers for
memset/memcpy/memmove and the __underlying_mem*() redirections. This
is only safe when the compiler supports the prefixed __asan_mem*()
intrinsics. On older toolchains (e.g. GCC 9) that lack this support,
plain mem*() calls from instrumented code fall through to the raw
assembly implementations in mem_64.S / copy_32.S, completely bypassing
the KASAN shadow check.
Other arches with GENERIC_ENTRY (x86, s390, loongarch, riscv) do not
hit this because their CI toolchains are always new enough to support
the prefix flag.
Background: the !GENERIC_ENTRY guard was introduced by commit 69d4c0d32186
("entry, kasan, x86: Disallow overriding mem*() functions", Peter Zijlstra,
Jan 2023). The root problem is that the KASAN C wrappers override the
linker symbol memset/memcpy/memmove globally, so any call from noinstr or
__no_sanitize_address code (e.g. irqentry_enter/irqentry_exit) would still
reach the KASAN shadow-check wrapper -- at a point where KASAN invariants
may not hold. The compiler prefix approach (Marco Elver, Feb 2023,
commit 51287dcb00cc) solves this by having the compiler emit __asan_memset
at instrumented call sites and bare memset inside __no_sanitize_address
functions, splitting the decision at code-generation time rather than at
link time.
A manual C-level override cannot replicate this split: a single linker
symbol cannot be made to resolve differently depending on the caller.
x86 also placed its raw memset/memcpy/memmove implementations in
.noinstr.text (same commit, 69d4c0d32186), which is the other half of
the fix: noinstr callers hit the raw assembly directly, safely bypassing
KASAN. PowerPC has not done this. Placing mem_64.S / memcpy_64.S /
copy_32.S implementations in .noinstr.text would be the complementary
long-term fix that could re-enable KASAN on older toolchains, but it
requires care around linker stub overflow on large PPC64 kernels (the
same reason powerpc uses NOKPROBE_SYMBOL rather than noinstr for its
interrupt handlers -- see the comment in asm/interrupt.h). That work
is left as a follow-up.
For now, introduce PPC_CC_HAS_KASAN_MEMINTRINSIC_PREFIX, an arch-local
compiler probe that mirrors the same check as CC_HAS_KASAN_MEMINTRINSIC_PREFIX
in lib/Kconfig.kasan but lives outside the 'if KASAN' block to avoid a
recursive dependency (CC_HAS_KASAN_MEMINTRINSIC_PREFIX depends on KASAN
which depends on HAVE_ARCH_KASAN). Gate the three HAVE_ARCH_KASAN selects
on this new symbol so that KASAN is not offered as a config option on
toolchains that cannot support it correctly with GENERIC_ENTRY.
Since KASAN on powerpc now unconditionally implies
CC_HAS_KASAN_MEMINTRINSIC_PREFIX, the old !CC_HAS_KASAN_MEMINTRINSIC_PREFIX
code paths in asm/kasan.h and asm/string.h are dead. Clean them up:
- asm/kasan.h: remove the dual-entry-point variant of _GLOBAL_KASAN /
_GLOBAL_TOC_KASAN / EXPORT_SYMBOL_KASAN that emitted both memset and
__memset as entry points to the same assembly. These aliases were only
needed so the C KASAN wrappers in shadow.c could call __memset() to
reach raw memory ops; with the compiler prefix approach the wrappers are
not used at all for mem* on powerpc.
- asm/string.h: remove the separate __memset/__memcpy/__memmove symbol
declarations and the memset/memcpy/memmove macro redirections for
uninstrumented files that were needed on old toolchains. Simplify the
CONFIG_KASAN block to just the three #define aliases (which are still
used by shadow.c as raw backends).
- cputable.c, prom_init.c: update stale comments that said GCC replaces
memcpy() with __memcpy() under KASAN; with the prefix flag it emits
__asan_memcpy() instead.
Reported-by: Venkat Rao Bagalkote <redacted>
Closes: https://lore.kernel.org/all/96dae110-f79b-4e54-8a9b-514369019b5d@linux.ibm.com
Signed-off-by: Mukesh Kumar Chaurasiya (IBM) <redacted>
---
arch/powerpc/Kconfig | 10 +++++++---
arch/powerpc/include/asm/kasan.h | 19 ++++++++-----------
arch/powerpc/include/asm/string.h | 25 +++++--------------------
arch/powerpc/kernel/cputable.c | 6 +++---
arch/powerpc/kernel/prom_init.c | 4 ++--
5 files changed, 25 insertions(+), 39 deletions(-)
@@ -7,6 +7,10 @@ config CC_HAS_ELFV2configCC_HAS_PREFIXEDdef_boolPPC64&&$(cc-option,-mcpu=power10-mprefixed)+configPPC_CC_HAS_KASAN_MEMINTRINSIC_PREFIX+def_bool(CC_IS_CLANG&&$(cc-option,-fsanitize=kernel-address-mllvm-asan-kernel-mem-intrinsic-prefix=1))||\+(CC_IS_GCC&&$(cc-option,-fsanitize=kernel-address--paramasan-kernel-mem-intrinsic-prefix=1))+configCC_HAS_PCREL# Clang has a bug (https://github.com/llvm/llvm-project/issues/62372)# where pcrel code is not generated if -msoft-float, -mno-altivec, or
On 08/09/26 12:19 pm, Mukesh Kumar Chaurasiya (IBM) wrote:
powerpc unconditionally selects GENERIC_ENTRY. The GENERIC_ENTRY
infrastructure relies on the compiler emitting __asan_mem*() calls at
instrumented mem*() sites rather than plain memset/memcpy/memmove, so
that entry/exit paths calling those functions are not instrumented.
This assumption is encoded in two places:
mm/kasan/shadow.c:
#if !defined(CONFIG_CC_HAS_KASAN_MEMINTRINSIC_PREFIX) &&
!defined(CONFIG_GENERIC_ENTRY)
include/linux/fortify-string.h:
#if !defined(CONFIG_CC_HAS_KASAN_MEMINTRINSIC_PREFIX) &&
!defined(CONFIG_GENERIC_ENTRY)
When GENERIC_ENTRY is set, both guards suppress the C wrappers for
memset/memcpy/memmove and the __underlying_mem*() redirections. This
is only safe when the compiler supports the prefixed __asan_mem*()
intrinsics. On older toolchains (e.g. GCC 9) that lack this support,
plain mem*() calls from instrumented code fall through to the raw
assembly implementations in mem_64.S / copy_32.S, completely bypassing
the KASAN shadow check.
Other arches with GENERIC_ENTRY (x86, s390, loongarch, riscv) do not
hit this because their CI toolchains are always new enough to support
the prefix flag.
Background: the !GENERIC_ENTRY guard was introduced by commit 69d4c0d32186
("entry, kasan, x86: Disallow overriding mem*() functions", Peter Zijlstra,
Jan 2023). The root problem is that the KASAN C wrappers override the
linker symbol memset/memcpy/memmove globally, so any call from noinstr or
__no_sanitize_address code (e.g. irqentry_enter/irqentry_exit) would still
reach the KASAN shadow-check wrapper -- at a point where KASAN invariants
may not hold. The compiler prefix approach (Marco Elver, Feb 2023,
commit 51287dcb00cc) solves this by having the compiler emit __asan_memset
at instrumented call sites and bare memset inside __no_sanitize_address
functions, splitting the decision at code-generation time rather than at
link time.
A manual C-level override cannot replicate this split: a single linker
symbol cannot be made to resolve differently depending on the caller.
x86 also placed its raw memset/memcpy/memmove implementations in
.noinstr.text (same commit, 69d4c0d32186), which is the other half of
the fix: noinstr callers hit the raw assembly directly, safely bypassing
KASAN. PowerPC has not done this. Placing mem_64.S / memcpy_64.S /
copy_32.S implementations in .noinstr.text would be the complementary
long-term fix that could re-enable KASAN on older toolchains, but it
requires care around linker stub overflow on large PPC64 kernels (the
same reason powerpc uses NOKPROBE_SYMBOL rather than noinstr for its
interrupt handlers -- see the comment in asm/interrupt.h). That work
is left as a follow-up.
For now, introduce PPC_CC_HAS_KASAN_MEMINTRINSIC_PREFIX, an arch-local
compiler probe that mirrors the same check as CC_HAS_KASAN_MEMINTRINSIC_PREFIX
in lib/Kconfig.kasan but lives outside the 'if KASAN' block to avoid a
recursive dependency (CC_HAS_KASAN_MEMINTRINSIC_PREFIX depends on KASAN
which depends on HAVE_ARCH_KASAN). Gate the three HAVE_ARCH_KASAN selects
on this new symbol so that KASAN is not offered as a config option on
toolchains that cannot support it correctly with GENERIC_ENTRY.
Since KASAN on powerpc now unconditionally implies
CC_HAS_KASAN_MEMINTRINSIC_PREFIX, the old !CC_HAS_KASAN_MEMINTRINSIC_PREFIX
code paths in asm/kasan.h and asm/string.h are dead. Clean them up:
- asm/kasan.h: remove the dual-entry-point variant of _GLOBAL_KASAN /
_GLOBAL_TOC_KASAN / EXPORT_SYMBOL_KASAN that emitted both memset and
__memset as entry points to the same assembly. These aliases were only
needed so the C KASAN wrappers in shadow.c could call __memset() to
reach raw memory ops; with the compiler prefix approach the wrappers are
not used at all for mem* on powerpc.
- asm/string.h: remove the separate __memset/__memcpy/__memmove symbol
declarations and the memset/memcpy/memmove macro redirections for
uninstrumented files that were needed on old toolchains. Simplify the
CONFIG_KASAN block to just the three #define aliases (which are still
used by shadow.c as raw backends).
- cputable.c, prom_init.c: update stale comments that said GCC replaces
memcpy() with __memcpy() under KASAN; with the prefix flag it emits
__asan_memcpy() instead.
Reported-by: Venkat Rao Bagalkote <redacted>
Closes: https://lore.kernel.org/all/96dae110-f79b-4e54-8a9b-514369019b5d@linux.ibm.com
Signed-off-by: Mukesh Kumar Chaurasiya (IBM) <redacted>
---
@@ -7,6 +7,10 @@ config CC_HAS_ELFV2configCC_HAS_PREFIXEDdef_boolPPC64&&$(cc-option,-mcpu=power10-mprefixed)+configPPC_CC_HAS_KASAN_MEMINTRINSIC_PREFIX+def_bool(CC_IS_CLANG&&$(cc-option,-fsanitize=kernel-address-mllvm-asan-kernel-mem-intrinsic-prefix=1))||\+(CC_IS_GCC&&$(cc-option,-fsanitize=kernel-address--paramasan-kernel-mem-intrinsic-prefix=1))+configCC_HAS_PCREL# Clang has a bug (https://github.com/llvm/llvm-project/issues/62372)# where pcrel code is not generated if -msoft-float, -mno-altivec, or
Le 08/09/2026 à 08:49, Mukesh Kumar Chaurasiya (IBM) a écrit :
quoted hunk
powerpc unconditionally selects GENERIC_ENTRY. The GENERIC_ENTRY
infrastructure relies on the compiler emitting __asan_mem*() calls at
instrumented mem*() sites rather than plain memset/memcpy/memmove, so
that entry/exit paths calling those functions are not instrumented.
This assumption is encoded in two places:
mm/kasan/shadow.c:
#if !defined(CONFIG_CC_HAS_KASAN_MEMINTRINSIC_PREFIX) &&
!defined(CONFIG_GENERIC_ENTRY)
include/linux/fortify-string.h:
#if !defined(CONFIG_CC_HAS_KASAN_MEMINTRINSIC_PREFIX) &&
!defined(CONFIG_GENERIC_ENTRY)
When GENERIC_ENTRY is set, both guards suppress the C wrappers for
memset/memcpy/memmove and the __underlying_mem*() redirections. This
is only safe when the compiler supports the prefixed __asan_mem*()
intrinsics. On older toolchains (e.g. GCC 9) that lack this support,
plain mem*() calls from instrumented code fall through to the raw
assembly implementations in mem_64.S / copy_32.S, completely bypassing
the KASAN shadow check.
Other arches with GENERIC_ENTRY (x86, s390, loongarch, riscv) do not
hit this because their CI toolchains are always new enough to support
the prefix flag.
Background: the !GENERIC_ENTRY guard was introduced by commit 69d4c0d32186
("entry, kasan, x86: Disallow overriding mem*() functions", Peter Zijlstra,
Jan 2023). The root problem is that the KASAN C wrappers override the
linker symbol memset/memcpy/memmove globally, so any call from noinstr or
__no_sanitize_address code (e.g. irqentry_enter/irqentry_exit) would still
reach the KASAN shadow-check wrapper -- at a point where KASAN invariants
may not hold. The compiler prefix approach (Marco Elver, Feb 2023,
commit 51287dcb00cc) solves this by having the compiler emit __asan_memset
at instrumented call sites and bare memset inside __no_sanitize_address
functions, splitting the decision at code-generation time rather than at
link time.
A manual C-level override cannot replicate this split: a single linker
symbol cannot be made to resolve differently depending on the caller.
x86 also placed its raw memset/memcpy/memmove implementations in
.noinstr.text (same commit, 69d4c0d32186), which is the other half of
the fix: noinstr callers hit the raw assembly directly, safely bypassing
KASAN. PowerPC has not done this. Placing mem_64.S / memcpy_64.S /
copy_32.S implementations in .noinstr.text would be the complementary
long-term fix that could re-enable KASAN on older toolchains, but it
requires care around linker stub overflow on large PPC64 kernels (the
same reason powerpc uses NOKPROBE_SYMBOL rather than noinstr for its
interrupt handlers -- see the comment in asm/interrupt.h). That work
is left as a follow-up.
For now, introduce PPC_CC_HAS_KASAN_MEMINTRINSIC_PREFIX, an arch-local
compiler probe that mirrors the same check as CC_HAS_KASAN_MEMINTRINSIC_PREFIX
in lib/Kconfig.kasan but lives outside the 'if KASAN' block to avoid a
recursive dependency (CC_HAS_KASAN_MEMINTRINSIC_PREFIX depends on KASAN
which depends on HAVE_ARCH_KASAN). Gate the three HAVE_ARCH_KASAN selects
on this new symbol so that KASAN is not offered as a config option on
toolchains that cannot support it correctly with GENERIC_ENTRY.
Since KASAN on powerpc now unconditionally implies
CC_HAS_KASAN_MEMINTRINSIC_PREFIX, the old !CC_HAS_KASAN_MEMINTRINSIC_PREFIX
code paths in asm/kasan.h and asm/string.h are dead. Clean them up:
- asm/kasan.h: remove the dual-entry-point variant of _GLOBAL_KASAN /
_GLOBAL_TOC_KASAN / EXPORT_SYMBOL_KASAN that emitted both memset and
__memset as entry points to the same assembly. These aliases were only
needed so the C KASAN wrappers in shadow.c could call __memset() to
reach raw memory ops; with the compiler prefix approach the wrappers are
not used at all for mem* on powerpc.
- asm/string.h: remove the separate __memset/__memcpy/__memmove symbol
declarations and the memset/memcpy/memmove macro redirections for
uninstrumented files that were needed on old toolchains. Simplify the
CONFIG_KASAN block to just the three #define aliases (which are still
used by shadow.c as raw backends).
- cputable.c, prom_init.c: update stale comments that said GCC replaces
memcpy() with __memcpy() under KASAN; with the prefix flag it emits
__asan_memcpy() instead.
Reported-by: Venkat Rao Bagalkote <redacted>
Closes: https://lore.kernel.org/all/96dae110-f79b-4e54-8a9b-514369019b5d@linux.ibm.com
Signed-off-by: Mukesh Kumar Chaurasiya (IBM) <redacted>
---
arch/powerpc/Kconfig | 10 +++++++---
arch/powerpc/include/asm/kasan.h | 19 ++++++++-----------
arch/powerpc/include/asm/string.h | 25 +++++--------------------
arch/powerpc/kernel/cputable.c | 6 +++---
arch/powerpc/kernel/prom_init.c | 4 ++--
5 files changed, 25 insertions(+), 39 deletions(-)
@@ -7,6 +7,10 @@ config CC_HAS_ELFV2configCC_HAS_PREFIXEDdef_boolPPC64&&$(cc-option,-mcpu=power10-mprefixed)+configPPC_CC_HAS_KASAN_MEMINTRINSIC_PREFIX+def_bool(CC_IS_CLANG&&$(cc-option,-fsanitize=kernel-address-mllvm-asan-kernel-mem-intrinsic-prefix=1))||\+(CC_IS_GCC&&$(cc-option,-fsanitize=kernel-address--paramasan-kernel-mem-intrinsic-prefix=1))+configCC_HAS_PCREL# Clang has a bug (https://github.com/llvm/llvm-project/issues/62372)# where pcrel code is not generated if -msoft-float, -mno-altivec, or
Does the initial problem still exist with the new __asan_memcpy() approach ? If not the comment should be removed.
quoted hunk
memcpy(t, s, sizeof(*t));
@@ -55,7 +55,7 @@ static struct cpu_spec * __init setup_cpu_spec(unsigned long offset, /* * Copy everything, then do fixups. Use memcpy() instead of *t = *s- * so that GCC replaces it by __memcpy() when KASAN is active+ * so that the compiler replaces it by __asan_memcpy() when KASAN is active
Does the initial problem still exist with the new __asan_memcpy() approach ? If not the comment should be removed.
Does the initial problem still exist with the new __asan_memcpy() approach ?
If not the comment should be removed.
Hey Christophe,
Thanks for pointing it out, i took a deeper look into this, here's my
understanding on it.
On PowerPC during very early boot the kernel is loaded by the
bootloader/firmware at some physical address, but the kernel was linked
expecting it to run at KERNELBASE(virtual address like
0xc000000000000000). The MMU mapping that makes that virtual address
valid hasn't been set up yet. So far for a window of early boot, code is
executing at the physical load address while all symbol addresses in the
binary refer to the virtual linked address. reloc_offset() computes the
gap between these two and PTRRELOC applies it to any pointer.
So PTRRELOC(&the_cpu_spec) gives the physical address where the struct
actually lives in memory right now, not where the linker thinks it lives.
Why *t = *s would be wrong?
In set_cur_cpu_spec:
struct cpu_spec *t = &the_cpu_spec; // linked (virtual) address
t = PTRRELOC(t); // physical address — where it actually is
memcpy(t, s, sizeof(*t)); // copy into the right place
If you wrote *t = *s instead, the compiler generates a struct assignment.
For a large struct like cpu_spec, GCC is free to implement that however
it likes — including emitting a call to memcpy(). But crucially, a
compiler-generated memcpy call resolves through the GOT/PLT or direct
symbol — which points to the linked virtual address of memcpy, not the
physical address. At this point in boot, calling through the wrong
address would jump to garbage or an unmapped page.
memcpy(t, s, sizeof(*t)) written explicitly is different: t is already
the corrected physical address, s points into the cpu_specs table which
has also been PTRRELOC'd. The explicit call goes through the normal
early-boot call mechanism which is safe.
The original comment said:
"use memcpy() instead of *t = *s so that GCC replaces it by __memcpy()
when KASAN is active"
This was added because under the old KASAN scheme
(!CC_HAS_KASAN_MEMINTRINSIC_PREFIX), KASAN overrode the memset/memcpy
linker symbols globally with C wrappers that called kasan_check_range().
If the compiler turned *t = *s into an implicit memcpy(), that would hit
the KASAN wrapper — calling kasan_check_range() at a point in early boot
where the KASAN shadow isn't mapped yet, causing a crash.
Writing memcpy(t, s, sizeof(*t)) explicitly made GCC emit __memcpy()
(the raw assembly alias exposed by _GLOBAL_KASAN) instead of the
KASAN-wrapped memcpy(), bypassing the shadow check.
That was the secondary reason. The primary reason that t is a
PTRRELOC-adjusted physical pointer and the copy must go through it
correctly was never stated.
So the KASAN comment is not required but i think we still need to state
why memcpy is required. For PTRRELOC adjustment, comment should reflect
that.
I'll update the comment and commit message and send out a new version.
Regards,
Mukesh
Does the initial problem still exist with the new __asan_memcpy() approach ?
If not the comment should be removed.
Hey Christophe,
Thanks for pointing it out, i took a deeper look into this, here's my
understanding on it.
On PowerPC during very early boot the kernel is loaded by the
bootloader/firmware at some physical address, but the kernel was linked
expecting it to run at KERNELBASE(virtual address like
0xc000000000000000). The MMU mapping that makes that virtual address
valid hasn't been set up yet. So far for a window of early boot, code is
executing at the physical load address while all symbol addresses in the
binary refer to the virtual linked address. reloc_offset() computes the
gap between these two and PTRRELOC applies it to any pointer.
So PTRRELOC(&the_cpu_spec) gives the physical address where the struct
actually lives in memory right now, not where the linker thinks it lives.
Why *t = *s would be wrong?
In set_cur_cpu_spec:
struct cpu_spec *t = &the_cpu_spec; // linked (virtual) address
t = PTRRELOC(t); // physical address — where it actually is
memcpy(t, s, sizeof(*t)); // copy into the right place
If you wrote *t = *s instead, the compiler generates a struct assignment.
For a large struct like cpu_spec, GCC is free to implement that however
it likes — including emitting a call to memcpy(). But crucially, a
compiler-generated memcpy call resolves through the GOT/PLT or direct
symbol — which points to the linked virtual address of memcpy, not the
physical address. At this point in boot, calling through the wrong
address would jump to garbage or an unmapped page.
memcpy(t, s, sizeof(*t)) written explicitly is different: t is already
the corrected physical address, s points into the cpu_specs table which
has also been PTRRELOC'd. The explicit call goes through the normal
early-boot call mechanism which is safe.
The original comment said:
"use memcpy() instead of *t = *s so that GCC replaces it by __memcpy()
when KASAN is active"
This was added because under the old KASAN scheme
(!CC_HAS_KASAN_MEMINTRINSIC_PREFIX), KASAN overrode the memset/memcpy
linker symbols globally with C wrappers that called kasan_check_range().
If the compiler turned *t = *s into an implicit memcpy(), that would hit
the KASAN wrapper — calling kasan_check_range() at a point in early boot
where the KASAN shadow isn't mapped yet, causing a crash.
Writing memcpy(t, s, sizeof(*t)) explicitly made GCC emit __memcpy()
(the raw assembly alias exposed by _GLOBAL_KASAN) instead of the
KASAN-wrapped memcpy(), bypassing the shadow check.
That was the secondary reason. The primary reason that t is a
PTRRELOC-adjusted physical pointer and the copy must go through it
correctly was never stated.
So the KASAN comment is not required but i think we still need to state
why memcpy is required. For PTRRELOC adjustment, comment should reflect
that.
I'll update the comment and commit message and send out a new version.
Explanation based on kernel v5.10
The problem was not linked to PTRRELOC, the t = PTRRELOC(t) followed by *t = *s works well in term of adressing, regardless of whether CONFIG_KASAN is enabled or not.
The problem is that with *t = *s, gcc emits a call to memcpy(). When CONFIG_KASAN is enabled, memcpy() is instrumented. But we don't want cputable.o instrumented as we have KASAN_SANITIZE_cputable.o := n in Makefile.
In asm/string.h we have:
#if defined(CONFIG_KASAN) && !defined(__SANITIZE_ADDRESS__)
/*
* For files that are not instrumented (e.g. mm/slub.c) we
* should use not instrumented version of mem* functions.
*/
#define memcpy(dst, src, len) __memcpy(dst, src, len)
#define memmove(dst, src, len) __memmove(dst, src, len)
#define memset(s, c, n) __memset(s, c, n)
Because in non-instrumented files like cputable.o we want memcpy() to be replaced at buildtime by __memcpy() to skip KASAN instrumentation. But this is resolved by pre-processing, and pre-processor doesn't know that the compiler will emit a call to memcpy().
By replacing *t = *s by the memcpy(), the pre-processor replaces memcpy() by __memcpy() when CONFIG_KASAN is enabled.
See the difference:
This is v5.10
00000000 <set_cur_cpu_spec>:
0: 94 21 ff e0 stwu r1,-32(r1)
4: 7c 69 1b 78 mr r9,r3
8: bf c1 00 18 stmw r30,24(r1)
c: 3f e0 00 00 lis r31,0
e: R_PPC_ADDR16_HA .data..read_mostly
10: 3b ff 00 00 addi r31,r31,0
12: R_PPC_ADDR16_LO .data..read_mostly
14: 7c 08 02 a6 mflr r0
18: 7d 3e 4b 78 mr r30,r9
1c: 7f e3 fb 78 mr r3,r31
20: 90 01 00 24 stw r0,36(r1)
24: 48 00 00 01 bl 24 <set_cur_cpu_spec+0x24>
24: R_PPC_REL24 add_reloc_offset
28: 7f c4 f3 78 mr r4,r30
2c: 38 a0 00 58 li r5,88
30: 48 00 00 01 bl 30 <set_cur_cpu_spec+0x30>
30: R_PPC_REL24 __memcpy
34: 38 7f 00 58 addi r3,r31,88
38: 48 00 00 01 bl 38 <set_cur_cpu_spec+0x38>
38: R_PPC_REL24 add_reloc_offset
3c: 93 e3 00 00 stw r31,0(r3)
40: 80 01 00 24 lwz r0,36(r1)
44: 83 c1 00 18 lwz r30,24(r1)
48: 83 e1 00 1c lwz r31,28(r1)
4c: 7c 08 03 a6 mtlr r0
50: 38 21 00 20 addi r1,r1,32
54: 4e 80 00 20 blr
This is v5.10 with commit adcf59187e270 reverted:
00000000 <set_cur_cpu_spec>:
0: 94 21 ff e0 stwu r1,-32(r1)
4: 7c 69 1b 78 mr r9,r3
8: bf c1 00 18 stmw r30,24(r1)
c: 3f e0 00 00 lis r31,0
e: R_PPC_ADDR16_HA .data..read_mostly
10: 3b ff 00 00 addi r31,r31,0
12: R_PPC_ADDR16_LO .data..read_mostly
14: 7c 08 02 a6 mflr r0
18: 7d 3e 4b 78 mr r30,r9
1c: 7f e3 fb 78 mr r3,r31
20: 90 01 00 24 stw r0,36(r1)
24: 48 00 00 01 bl 24 <set_cur_cpu_spec+0x24>
24: R_PPC_REL24 add_reloc_offset
28: 7f c4 f3 78 mr r4,r30
2c: 38 a0 00 58 li r5,88
30: 48 00 00 01 bl 30 <set_cur_cpu_spec+0x30>
30: R_PPC_REL24 memcpy
34: 38 7f 00 58 addi r3,r31,88
38: 48 00 00 01 bl 38 <set_cur_cpu_spec+0x38>
38: R_PPC_REL24 add_reloc_offset
3c: 93 e3 00 00 stw r31,0(r3)
40: 80 01 00 24 lwz r0,36(r1)
44: 83 c1 00 18 lwz r30,24(r1)
48: 83 e1 00 1c lwz r31,28(r1)
4c: 7c 08 03 a6 mtlr r0
50: 38 21 00 20 addi r1,r1,32
54: 4e 80 00 20 blr
So my question is ? Do we still have this issue nowadays ?
Christophe
t = PTRRELOC(t);
/*
- * use memcpy() instead of *t = *s so that GCC replaces it
- * by __memcpy() when KASAN is active
+ * use memcpy() instead of *t = *s so that the compiler replaces it
+ * by __asan_memcpy() when KASAN is active
*/
Does the initial problem still exist with the new __asan_memcpy() approach ?
If not the comment should be removed.
Hey Christophe,
Thanks for pointing it out, i took a deeper look into this, here's my
understanding on it.
On PowerPC during very early boot the kernel is loaded by the
bootloader/firmware at some physical address, but the kernel was linked
expecting it to run at KERNELBASE(virtual address like
0xc000000000000000). The MMU mapping that makes that virtual address
valid hasn't been set up yet. So far for a window of early boot, code is
executing at the physical load address while all symbol addresses in the
binary refer to the virtual linked address. reloc_offset() computes the
gap between these two and PTRRELOC applies it to any pointer.
So PTRRELOC(&the_cpu_spec) gives the physical address where the struct
actually lives in memory right now, not where the linker thinks it lives.
Why *t = *s would be wrong?
In set_cur_cpu_spec:
struct cpu_spec *t = &the_cpu_spec; // linked (virtual) address
t = PTRRELOC(t); // physical address — where it actually is
memcpy(t, s, sizeof(*t)); // copy into the right place
If you wrote *t = *s instead, the compiler generates a struct assignment.
For a large struct like cpu_spec, GCC is free to implement that however
it likes — including emitting a call to memcpy(). But crucially, a
compiler-generated memcpy call resolves through the GOT/PLT or direct
symbol — which points to the linked virtual address of memcpy, not the
physical address. At this point in boot, calling through the wrong
address would jump to garbage or an unmapped page.
memcpy(t, s, sizeof(*t)) written explicitly is different: t is already
the corrected physical address, s points into the cpu_specs table which
has also been PTRRELOC'd. The explicit call goes through the normal
early-boot call mechanism which is safe.
The original comment said:
"use memcpy() instead of *t = *s so that GCC replaces it by __memcpy()
when KASAN is active"
This was added because under the old KASAN scheme
(!CC_HAS_KASAN_MEMINTRINSIC_PREFIX), KASAN overrode the memset/memcpy
linker symbols globally with C wrappers that called kasan_check_range().
If the compiler turned *t = *s into an implicit memcpy(), that would hit
the KASAN wrapper — calling kasan_check_range() at a point in early boot
where the KASAN shadow isn't mapped yet, causing a crash.
Writing memcpy(t, s, sizeof(*t)) explicitly made GCC emit __memcpy()
(the raw assembly alias exposed by _GLOBAL_KASAN) instead of the
KASAN-wrapped memcpy(), bypassing the shadow check.
That was the secondary reason. The primary reason that t is a
PTRRELOC-adjusted physical pointer and the copy must go through it
correctly was never stated.
So the KASAN comment is not required but i think we still need to state
why memcpy is required. For PTRRELOC adjustment, comment should reflect
that.
I'll update the comment and commit message and send out a new version.
Explanation based on kernel v5.10
The problem was not linked to PTRRELOC, the t = PTRRELOC(t) followed by *t = *s works well in term of adressing, regardless of whether CONFIG_KASAN is enabled or not.
The problem is that with *t = *s, gcc emits a call to memcpy(). When CONFIG_KASAN is enabled, memcpy() is instrumented. But we don't want cputable.o instrumented as we have KASAN_SANITIZE_cputable.o := n in Makefile.
In asm/string.h we have:
#if defined(CONFIG_KASAN) && !defined(__SANITIZE_ADDRESS__)
/*
* For files that are not instrumented (e.g. mm/slub.c) we
* should use not instrumented version of mem* functions.
*/
#define memcpy(dst, src, len) __memcpy(dst, src, len)
#define memmove(dst, src, len) __memmove(dst, src, len)
#define memset(s, c, n) __memset(s, c, n)
Because in non-instrumented files like cputable.o we want memcpy() to be replaced at buildtime by __memcpy() to skip KASAN instrumentation. But this is resolved by pre-processing, and pre-processor doesn't know that the compiler will emit a call to memcpy().
By replacing *t = *s by the memcpy(), the pre-processor replaces memcpy() by __memcpy() when CONFIG_KASAN is enabled.
See the difference:
This is v5.10
00000000 <set_cur_cpu_spec>:
0: 94 21 ff e0 stwu r1,-32(r1)
4: 7c 69 1b 78 mr r9,r3
8: bf c1 00 18 stmw r30,24(r1)
c: 3f e0 00 00 lis r31,0
e: R_PPC_ADDR16_HA .data..read_mostly
10: 3b ff 00 00 addi r31,r31,0
12: R_PPC_ADDR16_LO .data..read_mostly
14: 7c 08 02 a6 mflr r0
18: 7d 3e 4b 78 mr r30,r9
1c: 7f e3 fb 78 mr r3,r31
20: 90 01 00 24 stw r0,36(r1)
24: 48 00 00 01 bl 24 <set_cur_cpu_spec+0x24>
24: R_PPC_REL24 add_reloc_offset
28: 7f c4 f3 78 mr r4,r30
2c: 38 a0 00 58 li r5,88
30: 48 00 00 01 bl 30 <set_cur_cpu_spec+0x30>
30: R_PPC_REL24 __memcpy
34: 38 7f 00 58 addi r3,r31,88
38: 48 00 00 01 bl 38 <set_cur_cpu_spec+0x38>
38: R_PPC_REL24 add_reloc_offset
3c: 93 e3 00 00 stw r31,0(r3)
40: 80 01 00 24 lwz r0,36(r1)
44: 83 c1 00 18 lwz r30,24(r1)
48: 83 e1 00 1c lwz r31,28(r1)
4c: 7c 08 03 a6 mtlr r0
50: 38 21 00 20 addi r1,r1,32
54: 4e 80 00 20 blr
This is v5.10 with commit adcf59187e270 reverted:
00000000 <set_cur_cpu_spec>:
0: 94 21 ff e0 stwu r1,-32(r1)
4: 7c 69 1b 78 mr r9,r3
8: bf c1 00 18 stmw r30,24(r1)
c: 3f e0 00 00 lis r31,0
e: R_PPC_ADDR16_HA .data..read_mostly
10: 3b ff 00 00 addi r31,r31,0
12: R_PPC_ADDR16_LO .data..read_mostly
14: 7c 08 02 a6 mflr r0
18: 7d 3e 4b 78 mr r30,r9
1c: 7f e3 fb 78 mr r3,r31
20: 90 01 00 24 stw r0,36(r1)
24: 48 00 00 01 bl 24 <set_cur_cpu_spec+0x24>
24: R_PPC_REL24 add_reloc_offset
28: 7f c4 f3 78 mr r4,r30
2c: 38 a0 00 58 li r5,88
30: 48 00 00 01 bl 30 <set_cur_cpu_spec+0x30>
30: R_PPC_REL24 memcpy
34: 38 7f 00 58 addi r3,r31,88
38: 48 00 00 01 bl 38 <set_cur_cpu_spec+0x38>
38: R_PPC_REL24 add_reloc_offset
3c: 93 e3 00 00 stw r31,0(r3)
40: 80 01 00 24 lwz r0,36(r1)
44: 83 c1 00 18 lwz r30,24(r1)
48: 83 e1 00 1c lwz r31,28(r1)
4c: 7c 08 03 a6 mtlr r0
50: 38 21 00 20 addi r1,r1,32
54: 4e 80 00 20 blr
So my question is ? Do we still have this issue nowadays ?
I now did the same test with v7.2 without and with adcf59187e270 reverted. I both cases I get memcpy().
Christophe
t = PTRRELOC(t);
/*
- * use memcpy() instead of *t = *s so that GCC replaces it
- * by __memcpy() when KASAN is active
+ * use memcpy() instead of *t = *s so that the compiler
replaces it
+ * by __asan_memcpy() when KASAN is active
*/
Does the initial problem still exist with the new
__asan_memcpy() approach ?
If not the comment should be removed.
Hey Christophe,
Thanks for pointing it out, i took a deeper look into this, here's my
understanding on it.
On PowerPC during very early boot the kernel is loaded by the
bootloader/firmware at some physical address, but the kernel was linked
expecting it to run at KERNELBASE(virtual address like
0xc000000000000000). The MMU mapping that makes that virtual address
valid hasn't been set up yet. So far for a window of early boot, code is
executing at the physical load address while all symbol addresses in the
binary refer to the virtual linked address. reloc_offset() computes the
gap between these two and PTRRELOC applies it to any pointer.
So PTRRELOC(&the_cpu_spec) gives the physical address where the struct
actually lives in memory right now, not where the linker thinks it lives.
Why *t = *s would be wrong?
In set_cur_cpu_spec:
struct cpu_spec *t = &the_cpu_spec; // linked (virtual) address
t = PTRRELOC(t); // physical address — where it
actually is
memcpy(t, s, sizeof(*t)); // copy into the right place
If you wrote *t = *s instead, the compiler generates a struct assignment.
For a large struct like cpu_spec, GCC is free to implement that however
it likes — including emitting a call to memcpy(). But crucially, a
compiler-generated memcpy call resolves through the GOT/PLT or direct
symbol — which points to the linked virtual address of memcpy, not the
physical address. At this point in boot, calling through the wrong
address would jump to garbage or an unmapped page.
memcpy(t, s, sizeof(*t)) written explicitly is different: t is already
the corrected physical address, s points into the cpu_specs table which
has also been PTRRELOC'd. The explicit call goes through the normal
early-boot call mechanism which is safe.
The original comment said:
"use memcpy() instead of *t = *s so that GCC replaces it by __memcpy()
when KASAN is active"
This was added because under the old KASAN scheme
(!CC_HAS_KASAN_MEMINTRINSIC_PREFIX), KASAN overrode the memset/memcpy
linker symbols globally with C wrappers that called kasan_check_range().
If the compiler turned *t = *s into an implicit memcpy(), that would hit
the KASAN wrapper — calling kasan_check_range() at a point in early boot
where the KASAN shadow isn't mapped yet, causing a crash.
Writing memcpy(t, s, sizeof(*t)) explicitly made GCC emit __memcpy()
(the raw assembly alias exposed by _GLOBAL_KASAN) instead of the
KASAN-wrapped memcpy(), bypassing the shadow check.
That was the secondary reason. The primary reason that t is a
PTRRELOC-adjusted physical pointer and the copy must go through it
correctly was never stated.
So the KASAN comment is not required but i think we still need to state
why memcpy is required. For PTRRELOC adjustment, comment should reflect
that.
I'll update the comment and commit message and send out a new version.
Explanation based on kernel v5.10
The problem was not linked to PTRRELOC, the t = PTRRELOC(t) followed by
*t = *s works well in term of adressing, regardless of whether
CONFIG_KASAN is enabled or not.
The problem is that with *t = *s, gcc emits a call to memcpy(). When
CONFIG_KASAN is enabled, memcpy() is instrumented. But we don't want
cputable.o instrumented as we have KASAN_SANITIZE_cputable.o := n in
Makefile.
In asm/string.h we have:
#if defined(CONFIG_KASAN) && !defined(__SANITIZE_ADDRESS__)
/*
* For files that are not instrumented (e.g. mm/slub.c) we
* should use not instrumented version of mem* functions.
*/
#define memcpy(dst, src, len) __memcpy(dst, src, len)
#define memmove(dst, src, len) __memmove(dst, src, len)
#define memset(s, c, n) __memset(s, c, n)
Because in non-instrumented files like cputable.o we want memcpy() to be
replaced at buildtime by __memcpy() to skip KASAN instrumentation. But
this is resolved by pre-processing, and pre-processor doesn't know that
the compiler will emit a call to memcpy().
By replacing *t = *s by the memcpy(), the pre-processor replaces
memcpy() by __memcpy() when CONFIG_KASAN is enabled.
See the difference:
This is v5.10
00000000 <set_cur_cpu_spec>:
0: 94 21 ff e0 stwu r1,-32(r1)
4: 7c 69 1b 78 mr r9,r3
8: bf c1 00 18 stmw r30,24(r1)
c: 3f e0 00 00 lis r31,0
e: R_PPC_ADDR16_HA .data..read_mostly
10: 3b ff 00 00 addi r31,r31,0
12: R_PPC_ADDR16_LO .data..read_mostly
14: 7c 08 02 a6 mflr r0
18: 7d 3e 4b 78 mr r30,r9
1c: 7f e3 fb 78 mr r3,r31
20: 90 01 00 24 stw r0,36(r1)
24: 48 00 00 01 bl 24 <set_cur_cpu_spec+0x24>
24: R_PPC_REL24 add_reloc_offset
28: 7f c4 f3 78 mr r4,r30
2c: 38 a0 00 58 li r5,88
30: 48 00 00 01 bl 30 <set_cur_cpu_spec+0x30>
30: R_PPC_REL24 __memcpy
34: 38 7f 00 58 addi r3,r31,88
38: 48 00 00 01 bl 38 <set_cur_cpu_spec+0x38>
38: R_PPC_REL24 add_reloc_offset
3c: 93 e3 00 00 stw r31,0(r3)
40: 80 01 00 24 lwz r0,36(r1)
44: 83 c1 00 18 lwz r30,24(r1)
48: 83 e1 00 1c lwz r31,28(r1)
4c: 7c 08 03 a6 mtlr r0
50: 38 21 00 20 addi r1,r1,32
54: 4e 80 00 20 blr
This is v5.10 with commit adcf59187e270 reverted:
00000000 <set_cur_cpu_spec>:
0: 94 21 ff e0 stwu r1,-32(r1)
4: 7c 69 1b 78 mr r9,r3
8: bf c1 00 18 stmw r30,24(r1)
c: 3f e0 00 00 lis r31,0
e: R_PPC_ADDR16_HA .data..read_mostly
10: 3b ff 00 00 addi r31,r31,0
12: R_PPC_ADDR16_LO .data..read_mostly
14: 7c 08 02 a6 mflr r0
18: 7d 3e 4b 78 mr r30,r9
1c: 7f e3 fb 78 mr r3,r31
20: 90 01 00 24 stw r0,36(r1)
24: 48 00 00 01 bl 24 <set_cur_cpu_spec+0x24>
24: R_PPC_REL24 add_reloc_offset
28: 7f c4 f3 78 mr r4,r30
2c: 38 a0 00 58 li r5,88
30: 48 00 00 01 bl 30 <set_cur_cpu_spec+0x30>
30: R_PPC_REL24 memcpy
34: 38 7f 00 58 addi r3,r31,88
38: 48 00 00 01 bl 38 <set_cur_cpu_spec+0x38>
38: R_PPC_REL24 add_reloc_offset
3c: 93 e3 00 00 stw r31,0(r3)
40: 80 01 00 24 lwz r0,36(r1)
44: 83 c1 00 18 lwz r30,24(r1)
48: 83 e1 00 1c lwz r31,28(r1)
4c: 7c 08 03 a6 mtlr r0
50: 38 21 00 20 addi r1,r1,32
54: 4e 80 00 20 blr
So my question is ? Do we still have this issue nowadays ?
I now did the same test with v7.2 without and with adcf59187e270 reverted. I
both cases I get memcpy().
Christophe
Oh I got it now. Thanks for the detailed explanation. Yeah the comment
is not required anymore.
I'll send out a V3.
Thanks and Regards,
Mukesh