From: Benjamin Gray <hidden> Date: 2022-10-21 05:25:19
This is a revision of Chris and Jordan's series to introduces a per cpu temporary
mm to be used for patching with strict rwx on radix mmus.
It is just rebased on powerpc/next. I am aware there are several code patching
patches on the list and can rebase when necessary. For now I figure this'll get
changes requested for a v9 either way.
v8: * Merge the temp mm 'introduction' and usage into one patch.
x86 split it because their temp MMU swap mechanism may be
used for other purposes, but ours cannot (it is local to
code-patching.c).
* Shuffle v7,3/5 cpu_patching_addr usage to the end (v8,4/4)
after cpu_patching_addr is actually introduced.
* Clearer formatting of the cpuhp_setup_state arguments
* Only allocate patching resources as CPU comes online. Free
them when CPU goes offline or if an error occurs during allocation.
* Refactored the random address calculation to make the page
alignment more obvious.
* Manually perform the allocation page walk to avoid taking locks
(which, given they are not necessary to take, is misleading) and
prevent memory leaks if page tree allocation fails.
* Cache the pte pointer.
* Stop using the patching mm first, then clear the patching PTE & TLB.
* Only clear the VA with the writable mapping from the TLB. Leaving
the other TLB entries helps performance, especially when patching
many times in a row (e.g., ftrace activation).
* Instruction patch verification moved to it's own patch onto shared
path with existing mechanism.
* Detect missing patching_mm and return an error for the caller to
decide what to do.
* Comment the purposes of each synchronisation, and why it is safe to
omit some at certain points.
Previous versions:
v7: https://lore.kernel.org/all/20211110003717.1150965-1-jniethe5@gmail.com/
v6: https://lore.kernel.org/all/20210911022904.30962-1-cmr@bluescreens.de/
v5: https://lore.kernel.org/all/20210713053113.4632-1-cmr@linux.ibm.com/
v4: https://lore.kernel.org/all/20210429072057.8870-1-cmr@bluescreens.de/
v3: https://lore.kernel.org/all/20200827052659.24922-1-cmr@codefail.de/
v2: https://lore.kernel.org/all/20200709040316.12789-1-cmr@informatik.wtf/
v1: https://lore.kernel.org/all/20200603051912.23296-1-cmr@informatik.wtf/
RFC: https://lore.kernel.org/all/20200323045205.20314-1-cmr@informatik.wtf/
x86: https://lore.kernel.org/kernel-hardening/20190426232303.28381-1-nadav.amit@gmail.com/
Benjamin Gray (5):
powerpc/code-patching: Use WARN_ON and fix check in poking_init
powerpc/code-patching: Verify instruction patch succeeded
powerpc/tlb: Add local flush for page given mm_struct and psize
powerpc/code-patching: Use temporary mm for Radix MMU
powerpc/code-patching: Use CPU local patch address directly
Jordan Niethe (1):
powerpc: Allow clearing and restoring registers independent of saved
breakpoint state
arch/powerpc/include/asm/book3s/32/tlbflush.h | 9 +
.../include/asm/book3s/64/tlbflush-hash.h | 5 +
arch/powerpc/include/asm/book3s/64/tlbflush.h | 8 +
arch/powerpc/include/asm/debug.h | 2 +
arch/powerpc/include/asm/nohash/tlbflush.h | 1 +
arch/powerpc/kernel/process.c | 36 ++-
arch/powerpc/lib/code-patching.c | 236 +++++++++++++++++-
7 files changed, 284 insertions(+), 13 deletions(-)
base-commit: 8636df94ec917019c4cb744ba0a1f94cf9057790
prerequisite-patch-id: b8387303be6478fdf94264d485d5e08994f305c7
prerequisite-patch-id: 06e54849e6c9e45a9b24668fa12cc0ece3f831a7
prerequisite-patch-id: f4be9e7d613761fba33fb2f7a81839cef36fe0fe
prerequisite-patch-id: 4ea0e36de5c393f9f6ae6243cb21a0ddb364c263
prerequisite-patch-id: 47a1294f0a5d5531ec5c32a761269cb5a1158515
prerequisite-patch-id: d72e371d3d820fdf529f03d2544c7f7f8bb6327a
prerequisite-patch-id: 3024e700433cb6a20dc1e1c6476ea1e98409d8b7
prerequisite-patch-id: f136637f7a8fe92dc4f60b908e2e7aa24aac3f43
--
2.37.3
From: Benjamin Gray <hidden> Date: 2022-10-21 05:26:12
From: Jordan Niethe <redacted>
For the coming temporary mm used for instruction patching, the
breakpoint registers need to be cleared to prevent them from
accidentally being triggered. As soon as the patching is done, the
breakpoints will be restored. The breakpoint state is stored in the per
cpu variable current_brk[]. Add a pause_breakpoints() function which will
clear the breakpoint registers without touching the state in
current_bkr[]. Add a pair function unpause_breakpoints() which will move
the state in current_brk[] back to the registers.
Signed-off-by: Jordan Niethe <redacted>
Signed-off-by: Benjamin Gray <redacted>
---
arch/powerpc/include/asm/debug.h | 2 ++
arch/powerpc/kernel/process.c | 36 +++++++++++++++++++++++++++++---
2 files changed, 35 insertions(+), 3 deletions(-)
@@ -862,10 +863,8 @@ static inline int set_breakpoint_8xx(struct arch_hw_breakpoint *brk)return0;}-void__set_breakpoint(intnr,structarch_hw_breakpoint*brk)+staticvoid____set_breakpoint(intnr,structarch_hw_breakpoint*brk){-memcpy(this_cpu_ptr(¤t_brk[nr]),brk,sizeof(*brk));-if(dawr_enabled())// Power8 or laterset_dawr(nr,brk);
@@ -879,6 +878,12 @@ void __set_breakpoint(int nr, struct arch_hw_breakpoint *brk)WARN_ON_ONCE(1);}+void__set_breakpoint(intnr,structarch_hw_breakpoint*brk)+{+memcpy(this_cpu_ptr(¤t_brk[nr]),brk,sizeof(*brk));+____set_breakpoint(nr,brk);+}+/* Check if we have DAWR or DABR hardware */boolppc_breakpoint_available(void){
@@ -891,6 +896,31 @@ bool ppc_breakpoint_available(void)}EXPORT_SYMBOL_GPL(ppc_breakpoint_available);+/* Disable the breakpoint in hardware without touching current_brk[] */+voidpause_breakpoints(void)+{+structarch_hw_breakpointbrk={0};+inti;++if(!ppc_breakpoint_available())+return;++for(i=0;i<nr_wp_slots();i++)+____set_breakpoint(i,&brk);+}++/* Renable the breakpoint in hardware from current_brk[] */+voidunpause_breakpoints(void)+{+inti;++if(!ppc_breakpoint_available())+return;++for(i=0;i<nr_wp_slots();i++)+____set_breakpoint(i,this_cpu_ptr(¤t_brk[i]));+}+#ifdef CONFIG_PPC_TRANSACTIONAL_MEMstaticinlinebooltm_enabled(structtask_struct*tsk)
From: Benjamin Gray <hidden> Date: 2022-10-21 05:27:06
From: "Christopher M. Riedl" <redacted>
The latest kernel docs list BUG_ON() as 'deprecated' and that they
should be replaced with WARN_ON() (or pr_warn()) when possible. The
BUG_ON() in poking_init() warrants a WARN_ON() rather than a pr_warn()
since the error condition is deemed "unreachable".
Also take this opportunity to fix the failure check in the WARN_ON():
cpuhp_setup_state(CPUHP_AP_ONLINE_DYN, ...) returns a positive integer
on success and a negative integer on failure.
Signed-off-by: Benjamin Gray <redacted>
---
arch/powerpc/lib/code-patching.c | 13 +++++--------
1 file changed, 5 insertions(+), 8 deletions(-)
@@ -81,16 +81,13 @@ static int text_area_cpu_down(unsigned int cpu)static__ro_after_initDEFINE_STATIC_KEY_FALSE(poking_init_done);-/*-*AlthoughBUG_ON()isrude,inthiscaseitshouldonlyhappenifENOMEM,and-*wejudgeitasbeingpreferabletoakernelthatwillcrashlaterwhen-*someonetriestousepatch_instruction().-*/void__initpoking_init(void){-BUG_ON(!cpuhp_setup_state(CPUHP_AP_ONLINE_DYN,-"powerpc/text_poke:online",text_area_cpu_up,-text_area_cpu_down));+WARN_ON(cpuhp_setup_state(CPUHP_AP_ONLINE_DYN,+"powerpc/text_poke:online",+text_area_cpu_up,+text_area_cpu_down)<0);+static_branch_enable(&poking_init_done);}
From: Benjamin Gray <hidden> Date: 2022-10-21 05:28:08
Verifies that if the instruction patching did not return an error then
the value stored at the given address to patch is now equal to the
instruction we patched it to.
Signed-off-by: Benjamin Gray <redacted>
---
arch/powerpc/lib/code-patching.c | 2 ++
1 file changed, 2 insertions(+)
From: Benjamin Gray <hidden> Date: 2022-10-21 05:29:02
Adds a local TLB flush operation that works given an mm_struct, VA to
flush, and page size representation.
This removes the need to create a vm_area_struct, which the temporary
patching mm work does not need.
Signed-off-by: Benjamin Gray <redacted>
---
arch/powerpc/include/asm/book3s/32/tlbflush.h | 9 +++++++++
arch/powerpc/include/asm/book3s/64/tlbflush-hash.h | 5 +++++
arch/powerpc/include/asm/book3s/64/tlbflush.h | 8 ++++++++
arch/powerpc/include/asm/nohash/tlbflush.h | 1 +
4 files changed, 23 insertions(+)
From: Benjamin Gray <hidden> Date: 2022-10-21 05:30:03
From: "Christopher M. Riedl" <redacted>
x86 supports the notion of a temporary mm which restricts access to
temporary PTEs to a single CPU. A temporary mm is useful for situations
where a CPU needs to perform sensitive operations (such as patching a
STRICT_KERNEL_RWX kernel) requiring temporary mappings without exposing
said mappings to other CPUs. Another benefit is that other CPU TLBs do
not need to be flushed when the temporary mm is torn down.
Mappings in the temporary mm can be set in the userspace portion of the
address-space.
Interrupts must be disabled while the temporary mm is in use. HW
breakpoints, which may have been set by userspace as watchpoints on
addresses now within the temporary mm, are saved and disabled when
loading the temporary mm. The HW breakpoints are restored when unloading
the temporary mm. All HW breakpoints are indiscriminately disabled while
the temporary mm is in use - this may include breakpoints set by perf.
Use the `poking_init` init hook to prepare a temporary mm and patching
address. Initialize the temporary mm by copying the init mm. Choose a
randomized patching address inside the temporary mm userspace address
space. The patching address is randomized between PAGE_SIZE and
DEFAULT_MAP_WINDOW-PAGE_SIZE.
Bits of entropy with 64K page size on BOOK3S_64:
bits of entropy = log2(DEFAULT_MAP_WINDOW_USER64 / PAGE_SIZE)
PAGE_SIZE=64K, DEFAULT_MAP_WINDOW_USER64=128TB
bits of entropy = log2(128TB / 64K)
bits of entropy = 31
The upper limit is DEFAULT_MAP_WINDOW due to how the Book3s64 Hash MMU
operates - by default the space above DEFAULT_MAP_WINDOW is not
available. Currently the Hash MMU does not use a temporary mm so
technically this upper limit isn't necessary; however, a larger
randomization range does not further "harden" this overall approach and
future work may introduce patching with a temporary mm on Hash as well.
Randomization occurs only once during initialization for each CPU as it
comes online.
The patching page is mapped with PAGE_KERNEL to set EAA[0] for the PTE
which ignores the AMR (so no need to unlock/lock KUAP) according to
PowerISA v3.0b Figure 35 on Radix.
Based on x86 implementation:
commit 4fc19708b165
("x86/alternatives: Initialize temporary mm for patching")
and:
commit b3fd8e83ada0
("x86/alternatives: Use temporary mm for text poking")
---
Synchronisation is done according to Book 3 Chapter 13 "Synchronization
Requirements for Context Alterations". Switching the mm is a change to
the PID, which requires a context synchronising instruction before and
after the change, and a hwsync between the last instruction that
performs address translation for an associated storage access.
Instruction fetch is an associated storage access, but the instruction
address mappings are not being changed, so it should not matter which
context they use. We must still perform a hwsync to guard arbitrary
prior code that may have access a userspace address.
TLB invalidation is local and VA specific. Local because only this core
used the patching mm, and VA specific because we only care that the
writable mapping is purged. Leaving the other mappings intact is more
efficient, especially when performing many code patches in a row (e.g.,
as ftrace would).
Signed-off-by: Benjamin Gray <redacted>
---
arch/powerpc/lib/code-patching.c | 226 ++++++++++++++++++++++++++++++-
1 file changed, 221 insertions(+), 5 deletions(-)
@@ -42,11 +47,59 @@ int raw_patch_instruction(u32 *addr, ppc_inst_t instr)}#ifdef CONFIG_STRICT_KERNEL_RWX+staticDEFINE_PER_CPU(structvm_struct*,text_poke_area);+staticDEFINE_PER_CPU(structmm_struct*,cpu_patching_mm);+staticDEFINE_PER_CPU(unsignedlong,cpu_patching_addr);+staticDEFINE_PER_CPU(pte_t*,cpu_patching_pte);staticintmap_patch_area(void*addr,unsignedlongtext_poke_addr);staticvoidunmap_patch_area(unsignedlongaddr);+structtemp_mm_state{+structmm_struct*mm;+};++staticboolmm_patch_enabled(void)+{+returnIS_ENABLED(CONFIG_SMP)&&radix_enabled();+}++/*+*ThefollowingappliesforRadixMMU.HashMMUhasdifferentrequirements,+*andsoisnotsupported.+*+*Changingmmrequirescontextsynchronisinginstructionsonbothsidesof+*thecontextswitch,aswellasahwsyncbetweenthelastinstructionfor+*whichtheaddressofanassociatedstorageaccesswastranslatedusing+*thecurrentcontext.+*+*switch_mm_irqs_offperformsanisyncafterthecontextswitch.Itis+*theresponsibilityofthecallertoperformtheCSIandhwsyncbefore+*starting/stoppingthetempmm.+*/+staticstructtemp_mm_statestart_using_temp_mm(structmm_struct*mm)+{+structtemp_mm_statetemp_state;++lockdep_assert_irqs_disabled();+temp_state.mm=current->active_mm;+switch_mm_irqs_off(temp_state.mm,mm,current);++WARN_ON(!mm_is_thread_local(mm));++pause_breakpoints();+returntemp_state;+}++staticvoidstop_using_temp_mm(structmm_struct*temp_mm,+structtemp_mm_stateprev_state)+{+lockdep_assert_irqs_disabled();+switch_mm_irqs_off(temp_mm,prev_state.mm,current);+unpause_breakpoints();+}+staticinttext_area_cpu_up(unsignedintcpu){structvm_struct*area;
@@ -79,14 +132,127 @@ static int text_area_cpu_down(unsigned int cpu)return0;}+staticinttext_area_cpu_up_mm(unsignedintcpu)+{+structmm_struct*mm;+unsignedlongaddr;+pgd_t*pgdp;+p4d_t*p4dp;+pud_t*pudp;+pmd_t*pmdp;+pte_t*ptep;++mm=copy_init_mm();+if(WARN_ON(!mm))+gotofail_no_mm;++/*+*Choosearandompage-alignedaddressfromtheinterval+*[PAGE_SIZE..DEFAULT_MAP_WINDOW-PAGE_SIZE].+*TheloweraddressboundisPAGE_SIZEtoavoidthezero-page.+*/+addr=(1+(get_random_long()%(DEFAULT_MAP_WINDOW/PAGE_SIZE-2)))<<PAGE_SHIFT;++/*+*PTEallocationusesGFP_KERNELwhichmeansweneedto+*pre-allocatethePTEherebecausewecannotdothe+*allocationduringpatchingwhenIRQsaredisabled.+*/+pgdp=pgd_offset(mm,addr);++p4dp=p4d_alloc(mm,pgdp,addr);+if(WARN_ON(!p4dp))+gotofail_no_p4d;++pudp=pud_alloc(mm,p4dp,addr);+if(WARN_ON(!pudp))+gotofail_no_pud;++pmdp=pmd_alloc(mm,pudp,addr);+if(WARN_ON(!pmdp))+gotofail_no_pmd;++ptep=pte_alloc_map(mm,pmdp,addr);+if(WARN_ON(!ptep))+gotofail_no_pte;++this_cpu_write(cpu_patching_mm,mm);+this_cpu_write(cpu_patching_addr,addr);+this_cpu_write(cpu_patching_pte,ptep);++return0;++fail_no_pte:+pmd_free(mm,pmdp);+mm_dec_nr_pmds(mm);+fail_no_pmd:+pud_free(mm,pudp);+mm_dec_nr_puds(mm);+fail_no_pud:+p4d_free(patching_mm,p4dp);+fail_no_p4d:+mmput(mm);+fail_no_mm:+return-ENOMEM;+}++staticinttext_area_cpu_down_mm(unsignedintcpu)+{+structmm_struct*mm;+unsignedlongaddr;+pte_t*ptep;+pmd_t*pmdp;+pud_t*pudp;+p4d_t*p4dp;+pgd_t*pgdp;++mm=this_cpu_read(cpu_patching_mm);+addr=this_cpu_read(cpu_patching_addr);++pgdp=pgd_offset(mm,addr);+p4dp=p4d_offset(pgdp,addr);+pudp=pud_offset(p4dp,addr);+pmdp=pmd_offset(pudp,addr);+ptep=pte_offset_map(pmdp,addr);++pte_free(mm,ptep);+pmd_free(mm,pmdp);+pud_free(mm,pudp);+p4d_free(mm,p4dp);+/* pgd is dropped in mmput */++mm_dec_nr_ptes(mm);+mm_dec_nr_pmds(mm);+mm_dec_nr_puds(mm);++mmput(mm);++this_cpu_write(cpu_patching_mm,NULL);+this_cpu_write(cpu_patching_addr,0);+this_cpu_write(cpu_patching_pte,NULL);++return0;+}+static__ro_after_initDEFINE_STATIC_KEY_FALSE(poking_init_done);void__initpoking_init(void){-WARN_ON(cpuhp_setup_state(CPUHP_AP_ONLINE_DYN,-"powerpc/text_poke:online",-text_area_cpu_up,-text_area_cpu_down)<0);+intret;++if(mm_patch_enabled())+ret=cpuhp_setup_state(CPUHP_AP_ONLINE_DYN,+"powerpc/text_poke_mm:online",+text_area_cpu_up_mm,+text_area_cpu_down_mm);+else+ret=cpuhp_setup_state(CPUHP_AP_ONLINE_DYN,+"powerpc/text_poke:online",+text_area_cpu_up,+text_area_cpu_down);++/* cpuhp_setup_state returns >= 0 on success */+WARN_ON(ret<0);static_branch_enable(&poking_init_done);}
@@ -144,6 +310,53 @@ static void unmap_patch_area(unsigned long addr)flush_tlb_kernel_range(addr,addr+PAGE_SIZE);}+staticint__do_patch_instruction_mm(u32*addr,ppc_inst_tinstr)+{+interr;+u32*patch_addr;+unsignedlongtext_poke_addr;+pte_t*pte;+unsignedlongpfn=get_patch_pfn(addr);+structmm_struct*patching_mm;+structtemp_mm_stateprev;++patching_mm=__this_cpu_read(cpu_patching_mm);+pte=__this_cpu_read(cpu_patching_pte);+text_poke_addr=__this_cpu_read(cpu_patching_addr);+patch_addr=(u32*)(text_poke_addr+offset_in_page(addr));++if(unlikely(!patching_mm))+return-ENOMEM;++set_pte_at(patching_mm,text_poke_addr,pte,pfn_pte(pfn,PAGE_KERNEL));++/* order PTE update before use, also serves as the hwsync */+asmvolatile("ptesync":::"memory");++/* order context switch after arbitrary prior code */+isync();++prev=start_using_temp_mm(patching_mm);++err=__patch_instruction(addr,instr,patch_addr);++/* hwsync performed by __patch_instruction (sync) if successful */+if(err)+mb();/* sync */++/* context synchronisation performed by __patch_instruction (isync or exception) */+stop_using_temp_mm(patching_mm,prev);++pte_clear(patching_mm,text_poke_addr,pte);+/*+*ptesynctoorderPTEupdatebeforeTLBinvalidationdone+*byradix__local_flush_tlb_page_psize(in_tlbiel_va)+*/+local_flush_tlb_page_psize(patching_mm,text_poke_addr,mmu_virtual_psize);++returnerr;+}+staticint__do_patch_instruction(u32*addr,ppc_inst_tinstr){interr;
@@ -183,7 +396,10 @@ static int do_patch_instruction(u32 *addr, ppc_inst_t instr)returnraw_patch_instruction(addr,instr);local_irq_save(flags);-err=__do_patch_instruction(addr,instr);+if(mm_patch_enabled())+err=__do_patch_instruction_mm(addr,instr);+else+err=__do_patch_instruction(addr,instr);local_irq_restore(flags);WARN_ON(!err&&!ppc_inst_equal(instr,ppc_inst_read(addr)));
From: Benjamin Gray <hidden> Date: 2022-10-21 05:30:56
With the isolated mm context support, there is a CPU local variable that
can hold the patch address. Use it instead of adding a level of
indirection through the text_poke_area vm_struct.
Signed-off-by: Benjamin Gray <redacted>
---
arch/powerpc/lib/code-patching.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
@@ -122,6 +122,7 @@ static int text_area_cpu_up(unsigned int cpu)unmap_patch_area(addr);this_cpu_write(text_poke_area,area);+this_cpu_write(cpu_patching_addr,addr);return0;}
@@ -365,7 +366,7 @@ static int __do_patch_instruction(u32 *addr, ppc_inst_t instr)pte_t*pte;unsignedlongpfn=get_patch_pfn(addr);-text_poke_addr=(unsignedlong)__this_cpu_read(text_poke_area)->addr&PAGE_MASK;+text_poke_addr=(unsignedlong)__this_cpu_read(cpu_patching_addr)&PAGE_MASK;patch_addr=(u32*)(text_poke_addr+offset_in_page(addr));pte=virt_to_kpte(text_poke_addr);
From: Russell Currey <hidden> Date: 2022-10-24 03:07:53
On Fri, 2022-10-21 at 16:22 +1100, Benjamin Gray wrote:
From: Jordan Niethe <redacted>
Hi Ben,
For the coming temporary mm used for instruction patching, the
breakpoint registers need to be cleared to prevent them from
accidentally being triggered. As soon as the patching is done, the
breakpoints will be restored. The breakpoint state is stored in the
per
cpu variable current_brk[]. Add a pause_breakpoints() function which
will
clear the breakpoint registers without touching the state in
current_bkr[]. Add a pair function unpause_breakpoints() which will
typo here ^
quoted hunk
move
the state in current_brk[] back to the registers.
Signed-off-by: Jordan Niethe <redacted>
Signed-off-by: Benjamin Gray <redacted>
---
arch/powerpc/include/asm/debug.h | 2 ++
arch/powerpc/kernel/process.c | 36 +++++++++++++++++++++++++++++-
--
2 files changed, 35 insertions(+), 3 deletions(-)
diff --git a/arch/powerpc/include/asm/debug.h
b/arch/powerpc/include/asm/debug.h
index 86a14736c76c..83f2dc3785e8 100644
From: Russell Currey <hidden> Date: 2022-10-24 03:09:28
On Fri, 2022-10-21 at 16:22 +1100, Benjamin Gray wrote:
From: "Christopher M. Riedl" <redacted>
The latest kernel docs list BUG_ON() as 'deprecated' and that they
should be replaced with WARN_ON() (or pr_warn()) when possible. The
BUG_ON() in poking_init() warrants a WARN_ON() rather than a
pr_warn()
since the error condition is deemed "unreachable".
Also take this opportunity to fix the failure check in the WARN_ON():
cpuhp_setup_state(CPUHP_AP_ONLINE_DYN, ...) returns a positive
integer
on success and a negative integer on failure.
Signed-off-by: Benjamin Gray <redacted>
From: Russell Currey <hidden> Date: 2022-10-24 03:21:13
On Fri, 2022-10-21 at 16:22 +1100, Benjamin Gray wrote:
quoted hunk
Verifies that if the instruction patching did not return an error
then
the value stored at the given address to patch is now equal to the
instruction we patched it to.
Signed-off-by: Benjamin Gray <redacted>
---
arch/powerpc/lib/code-patching.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/arch/powerpc/lib/code-patching.c
b/arch/powerpc/lib/code-patching.c
index 34fc7ac34d91..9b9eba574d7e 100644
As a side note, I had a look at test-code-patching.c and it doesn't
look like we don't have a test for ppc_inst_equal() with prefixed
instructions. We should fix that.
From: Russell Currey <hidden> Date: 2022-10-24 03:31:09
On Fri, 2022-10-21 at 16:22 +1100, Benjamin Gray wrote:
quoted hunk
Adds a local TLB flush operation that works given an mm_struct, VA to
flush, and page size representation.
This removes the need to create a vm_area_struct, which the temporary
patching mm work does not need.
Signed-off-by: Benjamin Gray <redacted>
---
arch/powerpc/include/asm/book3s/32/tlbflush.h | 9 +++++++++
arch/powerpc/include/asm/book3s/64/tlbflush-hash.h | 5 +++++
arch/powerpc/include/asm/book3s/64/tlbflush.h | 8 ++++++++
arch/powerpc/include/asm/nohash/tlbflush.h | 1 +
4 files changed, 23 insertions(+)
long start, unsigned long end
extern void flush_tlb_kernel_range(unsigned long start, unsigned
long end);
extern void local_flush_tlb_mm(struct mm_struct *mm);
extern void local_flush_tlb_page(struct vm_area_struct *vma,
unsigned long vmaddr);
+extern void local_flush_tlb_page_psize(struct mm_struct *mm,
unsigned long vmaddr, int psize);
extern void __local_flush_tlb_page(struct mm_struct *mm, unsigned
long vmaddr,
int tsize, int ind);
From: Russell Currey <hidden> Date: 2022-10-24 03:46:33
On Fri, 2022-10-21 at 16:22 +1100, Benjamin Gray wrote:
From: "Christopher M. Riedl" <redacted>
x86 supports the notion of a temporary mm which restricts access to
temporary PTEs to a single CPU. A temporary mm is useful for
situations
where a CPU needs to perform sensitive operations (such as patching a
STRICT_KERNEL_RWX kernel) requiring temporary mappings without
exposing
said mappings to other CPUs. Another benefit is that other CPU TLBs
do
not need to be flushed when the temporary mm is torn down.
Mappings in the temporary mm can be set in the userspace portion of
the
address-space.
Interrupts must be disabled while the temporary mm is in use. HW
breakpoints, which may have been set by userspace as watchpoints on
addresses now within the temporary mm, are saved and disabled when
loading the temporary mm. The HW breakpoints are restored when
unloading
the temporary mm. All HW breakpoints are indiscriminately disabled
while
the temporary mm is in use - this may include breakpoints set by
perf.
Use the `poking_init` init hook to prepare a temporary mm and
patching
address. Initialize the temporary mm by copying the init mm. Choose a
randomized patching address inside the temporary mm userspace address
space. The patching address is randomized between PAGE_SIZE and
DEFAULT_MAP_WINDOW-PAGE_SIZE.
Bits of entropy with 64K page size on BOOK3S_64:
bits of entropy = log2(DEFAULT_MAP_WINDOW_USER64 / PAGE_SIZE)
PAGE_SIZE=64K, DEFAULT_MAP_WINDOW_USER64=128TB
bits of entropy = log2(128TB / 64K)
bits of entropy = 31
The upper limit is DEFAULT_MAP_WINDOW due to how the Book3s64 Hash
MMU
operates - by default the space above DEFAULT_MAP_WINDOW is not
available. Currently the Hash MMU does not use a temporary mm so
technically this upper limit isn't necessary; however, a larger
randomization range does not further "harden" this overall approach
and
future work may introduce patching with a temporary mm on Hash as
well.
Randomization occurs only once during initialization for each CPU as
it
comes online.
The patching page is mapped with PAGE_KERNEL to set EAA[0] for the
PTE
which ignores the AMR (so no need to unlock/lock KUAP) according to
PowerISA v3.0b Figure 35 on Radix.
Based on x86 implementation:
commit 4fc19708b165
("x86/alternatives: Initialize temporary mm for patching")
and:
commit b3fd8e83ada0
("x86/alternatives: Use temporary mm for text poking")
---
Is the section following the --- your addendum to Chris' patch? That
cuts it off from git, including your signoff. It'd be better to have
it together as one commit message and note the bits you contributed
below the --- after your signoff.
Commits where you're modifying someone else's previous work should
include their signoff above yours, as well.
Synchronisation is done according to Book 3 Chapter 13
might want to mention the ISA version alongside this, since chapter
numbering can change
quoted hunk
"Synchronization
Requirements for Context Alterations". Switching the mm is a change
to
the PID, which requires a context synchronising instruction before
and
after the change, and a hwsync between the last instruction that
performs address translation for an associated storage access.
Instruction fetch is an associated storage access, but the
instruction
address mappings are not being changed, so it should not matter which
context they use. We must still perform a hwsync to guard arbitrary
prior code that may have access a userspace address.
TLB invalidation is local and VA specific. Local because only this
core
used the patching mm, and VA specific because we only care that the
writable mapping is purged. Leaving the other mappings intact is more
efficient, especially when performing many code patches in a row
(e.g.,
as ftrace would).
Signed-off-by: Benjamin Gray <redacted>
---
arch/powerpc/lib/code-patching.c | 226
++++++++++++++++++++++++++++++-
1 file changed, 221 insertions(+), 5 deletions(-)
diff --git a/arch/powerpc/lib/code-patching.c
b/arch/powerpc/lib/code-patching.c
index 9b9eba574d7e..eabdd74a26c0 100644
Is this a useful abstraction? This looks like a struct that used to
have more in it but is no longer necessary.
+
+static bool mm_patch_enabled(void)
+{
+ return IS_ENABLED(CONFIG_SMP) && radix_enabled();
+}
+
+/*
+ * The following applies for Radix MMU. Hash MMU has different
requirements,
+ * and so is not supported.
+ *
+ * Changing mm requires context synchronising instructions on both
sides of
+ * the context switch, as well as a hwsync between the last
instruction for
+ * which the address of an associated storage access was translated
using
+ * the current context.
+ *
+ * switch_mm_irqs_off performs an isync after the context switch. It
is
I'd prefer having parens here (switch_mm_irqs_off()) but I dunno if
that's actually a style guideline.
quoted hunk
+ * the responsibility of the caller to perform the CSI and hwsync
before
+ * starting/stopping the temp mm.
+ */
+static struct temp_mm_state start_using_temp_mm(struct mm_struct
*mm)
+{
+ struct temp_mm_state temp_state;
+
+ lockdep_assert_irqs_disabled();
+ temp_state.mm = current->active_mm;
+ switch_mm_irqs_off(temp_state.mm, mm, current);
+
+ WARN_ON(!mm_is_thread_local(mm));
+
+ pause_breakpoints();
+ return temp_state;
+}
+
+static void stop_using_temp_mm(struct mm_struct *temp_mm,
+ struct temp_mm_state prev_state)
+{
+ lockdep_assert_irqs_disabled();
+ switch_mm_irqs_off(temp_mm, prev_state.mm, current);
+ unpause_breakpoints();
+}
+
static int text_area_cpu_up(unsigned int cpu)
{
struct vm_struct *area;
@@ -79,14 +132,127 @@ static int text_area_cpu_down(unsigned int cpu)
return 0;
}
+static int text_area_cpu_up_mm(unsigned int cpu)
+{
+ struct mm_struct *mm;
+ unsigned long addr;
+ pgd_t *pgdp;
+ p4d_t *p4dp;
+ pud_t *pudp;
+ pmd_t *pmdp;
+ pte_t *ptep;
+
+ mm = copy_init_mm();
+ if (WARN_ON(!mm))
+ goto fail_no_mm;
+
+ /*
+ * Choose a random page-aligned address from the interval
+ * [PAGE_SIZE .. DEFAULT_MAP_WINDOW - PAGE_SIZE].
+ * The lower address bound is PAGE_SIZE to avoid the zero-
page.
+ */
+ addr = (1 + (get_random_long() % (DEFAULT_MAP_WINDOW /
PAGE_SIZE - 2))) << PAGE_SHIFT;
+
+ /*
+ * PTE allocation uses GFP_KERNEL which means we need to
+ * pre-allocate the PTE here because we cannot do the
+ * allocation during patching when IRQs are disabled.
+ */
+ pgdp = pgd_offset(mm, addr);
+
+ p4dp = p4d_alloc(mm, pgdp, addr);
+ if (WARN_ON(!p4dp))
+ goto fail_no_p4d;
+
+ pudp = pud_alloc(mm, p4dp, addr);
+ if (WARN_ON(!pudp))
+ goto fail_no_pud;
+
+ pmdp = pmd_alloc(mm, pudp, addr);
+ if (WARN_ON(!pmdp))
+ goto fail_no_pmd;
+
+ ptep = pte_alloc_map(mm, pmdp, addr);
+ if (WARN_ON(!ptep))
+ goto fail_no_pte;
+
+ this_cpu_write(cpu_patching_mm, mm);
+ this_cpu_write(cpu_patching_addr, addr);
+ this_cpu_write(cpu_patching_pte, ptep);
+
+ return 0;
+
+fail_no_pte:
+ pmd_free(mm, pmdp);
+ mm_dec_nr_pmds(mm);
+fail_no_pmd:
+ pud_free(mm, pudp);
+ mm_dec_nr_puds(mm);
+fail_no_pud:
+ p4d_free(patching_mm, p4dp);
+fail_no_p4d:
+ mmput(mm);
+fail_no_mm:
+ return -ENOMEM;
+}
+
+static int text_area_cpu_down_mm(unsigned int cpu)
+{
+ struct mm_struct *mm;
+ unsigned long addr;
+ pte_t *ptep;
+ pmd_t *pmdp;
+ pud_t *pudp;
+ p4d_t *p4dp;
+ pgd_t *pgdp;
+
+ mm = this_cpu_read(cpu_patching_mm);
+ addr = this_cpu_read(cpu_patching_addr);
+
+ pgdp = pgd_offset(mm, addr);
+ p4dp = p4d_offset(pgdp, addr);
+ pudp = pud_offset(p4dp, addr);
+ pmdp = pmd_offset(pudp, addr);
+ ptep = pte_offset_map(pmdp, addr);
+
+ pte_free(mm, ptep);
+ pmd_free(mm, pmdp);
+ pud_free(mm, pudp);
+ p4d_free(mm, p4dp);
+ /* pgd is dropped in mmput */
+
+ mm_dec_nr_ptes(mm);
+ mm_dec_nr_pmds(mm);
+ mm_dec_nr_puds(mm);
+
+ mmput(mm);
+
+ this_cpu_write(cpu_patching_mm, NULL);
+ this_cpu_write(cpu_patching_addr, 0);
+ this_cpu_write(cpu_patching_pte, NULL);
+
+ return 0;
+}
+
static __ro_after_init DEFINE_STATIC_KEY_FALSE(poking_init_done);
void __init poking_init(void)
{
- WARN_ON(cpuhp_setup_state(CPUHP_AP_ONLINE_DYN,
- "powerpc/text_poke:online",
- text_area_cpu_up,
- text_area_cpu_down) < 0);
+ int ret;
+
+ if (mm_patch_enabled())
+ ret = cpuhp_setup_state(CPUHP_AP_ONLINE_DYN,
+ "powerpc/text_poke_mm:online"
,
+ text_area_cpu_up_mm,
+ text_area_cpu_down_mm);
+ else
+ ret = cpuhp_setup_state(CPUHP_AP_ONLINE_DYN,
+ "powerpc/text_poke:online",
+ text_area_cpu_up,
+ text_area_cpu_down);
+
+ /* cpuhp_setup_state returns >= 0 on success */
+ WARN_ON(ret < 0);
static_branch_enable(&poking_init_done);
}
@@ -144,6 +310,53 @@ static void unmap_patch_area(unsigned long addr)
flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
}
+static int __do_patch_instruction_mm(u32 *addr, ppc_inst_t instr)
+{
+ int err;
+ u32 *patch_addr;
+ unsigned long text_poke_addr;
+ pte_t *pte;
+ unsigned long pfn = get_patch_pfn(addr);
+ struct mm_struct *patching_mm;
+ struct temp_mm_state prev;
Reverse christmas tree? If we care
Rest looks good to me.
quoted hunk
+
+ patching_mm = __this_cpu_read(cpu_patching_mm);
+ pte = __this_cpu_read(cpu_patching_pte);
+ text_poke_addr = __this_cpu_read(cpu_patching_addr);
+ patch_addr = (u32 *)(text_poke_addr + offset_in_page(addr));
+
+ if (unlikely(!patching_mm))
+ return -ENOMEM;
+
+ set_pte_at(patching_mm, text_poke_addr, pte, pfn_pte(pfn,
PAGE_KERNEL));
+
+ /* order PTE update before use, also serves as the hwsync */
+ asm volatile("ptesync": : :"memory");
+
+ /* order context switch after arbitrary prior code */
+ isync();
+
+ prev = start_using_temp_mm(patching_mm);
+
+ err = __patch_instruction(addr, instr, patch_addr);
+
+ /* hwsync performed by __patch_instruction (sync) if
successful */
+ if (err)
+ mb(); /* sync */
+
+ /* context synchronisation performed by __patch_instruction
(isync or exception) */
+ stop_using_temp_mm(patching_mm, prev);
+
+ pte_clear(patching_mm, text_poke_addr, pte);
+ /*
+ * ptesync to order PTE update before TLB invalidation done
+ * by radix__local_flush_tlb_page_psize (in _tlbiel_va)
+ */
+ local_flush_tlb_page_psize(patching_mm, text_poke_addr,
mmu_virtual_psize);
+
+ return err;
+}
+
static int __do_patch_instruction(u32 *addr, ppc_inst_t instr)
{
int err;
@@ -183,7 +396,10 @@ static int do_patch_instruction(u32 *addr,
From: Russell Currey <hidden> Date: 2022-10-24 04:23:07
On Fri, 2022-10-21 at 16:22 +1100, Benjamin Gray wrote:
quoted hunk
Adds a local TLB flush operation that works given an mm_struct, VA to
flush, and page size representation.
This removes the need to create a vm_area_struct, which the temporary
patching mm work does not need.
Signed-off-by: Benjamin Gray <redacted>
---
arch/powerpc/include/asm/book3s/32/tlbflush.h | 9 +++++++++
arch/powerpc/include/asm/book3s/64/tlbflush-hash.h | 5 +++++
arch/powerpc/include/asm/book3s/64/tlbflush.h | 8 ++++++++
arch/powerpc/include/asm/nohash/tlbflush.h | 1 +
4 files changed, 23 insertions(+)
long start, unsigned long end
extern void flush_tlb_kernel_range(unsigned long start, unsigned
long end);
extern void local_flush_tlb_mm(struct mm_struct *mm);
extern void local_flush_tlb_page(struct vm_area_struct *vma,
unsigned long vmaddr);
+extern void local_flush_tlb_page_psize(struct mm_struct *mm,
unsigned long vmaddr, int psize);
From: Benjamin Gray <hidden> Date: 2022-10-24 05:18:32
On Mon, 2022-10-24 at 14:45 +1100, Russell Currey wrote:
On Fri, 2022-10-21 at 16:22 +1100, Benjamin Gray wrote:
quoted
From: "Christopher M. Riedl" <redacted>
x86 supports the notion of a temporary mm which restricts access to
temporary PTEs to a single CPU. A temporary mm is useful for
situations
where a CPU needs to perform sensitive operations (such as patching
a
STRICT_KERNEL_RWX kernel) requiring temporary mappings without
exposing
said mappings to other CPUs. Another benefit is that other CPU TLBs
do
not need to be flushed when the temporary mm is torn down.
Mappings in the temporary mm can be set in the userspace portion of
the
address-space.
Interrupts must be disabled while the temporary mm is in use. HW
breakpoints, which may have been set by userspace as watchpoints on
addresses now within the temporary mm, are saved and disabled when
loading the temporary mm. The HW breakpoints are restored when
unloading
the temporary mm. All HW breakpoints are indiscriminately disabled
while
the temporary mm is in use - this may include breakpoints set by
perf.
Use the `poking_init` init hook to prepare a temporary mm and
patching
address. Initialize the temporary mm by copying the init mm. Choose
a
randomized patching address inside the temporary mm userspace
address
space. The patching address is randomized between PAGE_SIZE and
DEFAULT_MAP_WINDOW-PAGE_SIZE.
Bits of entropy with 64K page size on BOOK3S_64:
bits of entropy = log2(DEFAULT_MAP_WINDOW_USER64 /
PAGE_SIZE)
PAGE_SIZE=64K, DEFAULT_MAP_WINDOW_USER64=128TB
bits of entropy = log2(128TB / 64K)
bits of entropy = 31
The upper limit is DEFAULT_MAP_WINDOW due to how the Book3s64 Hash
MMU
operates - by default the space above DEFAULT_MAP_WINDOW is not
available. Currently the Hash MMU does not use a temporary mm so
technically this upper limit isn't necessary; however, a larger
randomization range does not further "harden" this overall approach
and
future work may introduce patching with a temporary mm on Hash as
well.
Randomization occurs only once during initialization for each CPU
as
it
comes online.
The patching page is mapped with PAGE_KERNEL to set EAA[0] for the
PTE
which ignores the AMR (so no need to unlock/lock KUAP) according to
PowerISA v3.0b Figure 35 on Radix.
Based on x86 implementation:
commit 4fc19708b165
("x86/alternatives: Initialize temporary mm for patching")
and:
commit b3fd8e83ada0
("x86/alternatives: Use temporary mm for text poking")
---
Is the section following the --- your addendum to Chris' patch? That
cuts it off from git, including your signoff. It'd be better to have
it together as one commit message and note the bits you contributed
below the --- after your signoff.
Commits where you're modifying someone else's previous work should
include their signoff above yours, as well.
Addendum to his wording, to break it off from the "From..." section
(which is me splicing together his comments from previous patches with
some minor changes to account for the patch changes). I found out
earlier today that Git will treat it as a comment :(
I'll add the signed off by back, I wasn't sure whether to leave it
there after making changes (same in patch 2).
quoted
+static int __do_patch_instruction_mm(u32 *addr, ppc_inst_t instr)
+{
+ int err;
+ u32 *patch_addr;
+ unsigned long text_poke_addr;
+ pte_t *pte;
+ unsigned long pfn = get_patch_pfn(addr);
+ struct mm_struct *patching_mm;
+ struct temp_mm_state prev;
Reverse christmas tree? If we care
Currently it's mirroring the __do_patch_instruction declarations, with
extra ones at the bottom.
From: Benjamin Gray <hidden> Date: 2022-10-24 05:23:17
On Mon, 2022-10-24 at 14:30 +1100, Russell Currey wrote:
On Fri, 2022-10-21 at 16:22 +1100, Benjamin Gray wrote:
quoted
Adds a local TLB flush operation that works given an mm_struct, VA
to
flush, and page size representation.
This removes the need to create a vm_area_struct, which the
temporary
patching mm work does not need.
Signed-off-by: Benjamin Gray <redacted>
---
arch/powerpc/include/asm/book3s/32/tlbflush.h | 9 +++++++++
arch/powerpc/include/asm/book3s/64/tlbflush-hash.h | 5 +++++
arch/powerpc/include/asm/book3s/64/tlbflush.h | 8 ++++++++
arch/powerpc/include/asm/nohash/tlbflush.h | 1 +
4 files changed, 23 insertions(+)
Is there any utility in adding this for 32bit if the following
patches
are only for Radix?
It needs some kind of definition to avoid #ifdef's. I figured I may as
well provide a correct implementation, given the functions around it
are implemented. The BUILD_BUG_ON specifically is just defensive in
case my assumptions are wrong. I don't know anything about these
machines, just what the kernel defines. I can remove the check, or
replace the whole implementation with a BUILD_BUG?
arch/powerpc/lib/code-patching.c:355:9: error: implicit declaration of function 'local_flush_tlb_page_psize'; did you mean 'local_flush_tlb_page'? [-Werror=implicit-function-declaration]
355 | local_flush_tlb_page_psize(patching_mm, text_poke_addr, mmu_virtual_psize);
| ^~~~~~~~~~~~~~~~~~~~~~~~~~
| local_flush_tlb_page
cc1: all warnings being treated as errors
vim +355 arch/powerpc/lib/code-patching.c
312
313 static int __do_patch_instruction_mm(u32 *addr, ppc_inst_t instr)
314 {
315 int err;
316 u32 *patch_addr;
317 unsigned long text_poke_addr;
318 pte_t *pte;
319 unsigned long pfn = get_patch_pfn(addr);
320 struct mm_struct *patching_mm;
321 struct temp_mm_state prev;
322
323 patching_mm = __this_cpu_read(cpu_patching_mm);
324 pte = __this_cpu_read(cpu_patching_pte);
325 text_poke_addr = __this_cpu_read(cpu_patching_addr);
326 patch_addr = (u32 *)(text_poke_addr + offset_in_page(addr));
327
328 if (unlikely(!patching_mm))
329 return -ENOMEM;
330
331 set_pte_at(patching_mm, text_poke_addr, pte, pfn_pte(pfn, PAGE_KERNEL));
332
333 /* order PTE update before use, also serves as the hwsync */
334 asm volatile("ptesync": : :"memory");
335
336 /* order context switch after arbitrary prior code */
337 isync();
338
339 prev = start_using_temp_mm(patching_mm);
340
341 err = __patch_instruction(addr, instr, patch_addr);
342
343 /* hwsync performed by __patch_instruction (sync) if successful */
344 if (err)
345 mb(); /* sync */
346
347 /* context synchronisation performed by __patch_instruction (isync or exception) */
348 stop_using_temp_mm(patching_mm, prev);
349
350 pte_clear(patching_mm, text_poke_addr, pte);
351 /*
352 * ptesync to order PTE update before TLB invalidation done
353 * by radix__local_flush_tlb_page_psize (in _tlbiel_va)
354 */
> 355 local_flush_tlb_page_psize(patching_mm, text_poke_addr, mmu_virtual_psize);
356
357 return err;
358 }
359
--
0-DAY CI Kernel Test Service
https://01.org/lkp
From: Christopher M. Riedl <hidden> Date: 2022-10-24 20:02:17
On Mon Oct 24, 2022 at 12:17 AM CDT, Benjamin Gray wrote:
On Mon, 2022-10-24 at 14:45 +1100, Russell Currey wrote:
quoted
On Fri, 2022-10-21 at 16:22 +1100, Benjamin Gray wrote:
quoted
From: "Christopher M. Riedl" <redacted>
-----%<------
quoted
quoted
---
Is the section following the --- your addendum to Chris' patch? That
cuts it off from git, including your signoff. It'd be better to have
it together as one commit message and note the bits you contributed
below the --- after your signoff.
Commits where you're modifying someone else's previous work should
include their signoff above yours, as well.
Addendum to his wording, to break it off from the "From..." section
(which is me splicing together his comments from previous patches with
some minor changes to account for the patch changes). I found out
earlier today that Git will treat it as a comment :(
I'll add the signed off by back, I wasn't sure whether to leave it
there after making changes (same in patch 2).
This commit has lots of my words so should probably keep the sign-off - if only
to guarantee that blame is properly directed at me for any nonsense therein ^^.
Patch 2 probably doesn't need my sign-off any more - iirc, I actually defended
the BUG_ON()s (which are WARN_ON()s now) at some point.
As a side note, I had a look at test-code-patching.c and it doesn't
look like we don't have a test for ppc_inst_equal() with prefixed
instructions. We should fix that.
Yeah, for a different series though I assume. And I think it would be
better suited in a suite dedicated to testing asm/inst.h functions.