From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:54:15
Changelog:
v16:
- Merged with 5.14-rc4
v15:
- Changed trans_pgd_copy_el2_vectors() to use vector table that
only shared by kexec and hibernate. This way sync does not have
dangling branch that was recently introduced. (Reported by Marc
Zyngier)
- Renamed is_hyp_callable() to is_hyp_nvhe() as requested by Marc
Zyngier
- Clean-ups, comment fixes.
- Sync with upstream 368094df48e680fa51cedb68537408cfa64b788e
v14:
- Fixed a bug in "arm64: hyp-stub: Move elx_sync into the vectors"
that was noticed by Marc Zyngier
- Merged with upstream
v13:
- Fixed a hang on ThunderX2, thank you Pingfan Liu for reporting
the problem. In relocation function we need civac not ivac, we
need to clean data in addition to invalidating it.
Since I was using ThunderX2 machine I also measured the new
performance data on this large ARM64 server. The MMU improves
kexec relocation 190 times on this machine! (see below for
raw data). Saves 7.5s during CentOS kexec reboot.
v12:
- A major change compared to previous version. Instead of using
contiguous VA range a copy of linear map is now used to perform
copying of segments during relocation as it was agreed in the
discussion of version 11 of this project.
- In addition to using linear map, I also took several ideas from
James Morse to better organize the kexec relocation:
1. skip relocation function entirely if that is not needed
2. remove the PoC flushing function since it is not needed
anymore with MMU enabled.
v11:
- Fixed missing KEXEC_CORE dependency for trans_pgd.c
- Removed useless "if(rc) return rc" statement (thank you Tyler Hicks)
- Another 12 patches were accepted into maintainer's get.
Re-based patches against:
https://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux.git
Branch: for-next/kexec
v10:
- Addressed a lot of comments form James Morse and from Marc Zyngier
- Added review-by's
- Synchronized with mainline
v9: - 9 patches from previous series landed in upstream, so now series
is smaller
- Added two patches from James Morse to address idmap issues for machines
with high physical addresses.
- Addressed comments from Selin Dag about compiling issues. He also tested
my series and got similar performance results: ~60 ms instead of ~580 ms
with an initramfs size of ~120MB.
v8:
- Synced with mainline to keep series up-to-date
v7:
-- Addressed comments from James Morse
- arm64: hibernate: pass the allocated pgdp to ttbr0
Removed "Fixes" tag, and added Added Reviewed-by: James Morse
- arm64: hibernate: check pgd table allocation
Sent out as a standalone patch so it can be sent to stable
Series applies on mainline + this patch
- arm64: hibernate: add trans_pgd public functions
Remove second allocation of tmp_pg_dir in swsusp_arch_resume
Added Reviewed-by: James Morse [off-list ref]
- arm64: kexec: move relocation function setup and clean up
Fixed typo in commit log
Changed kern_reloc to phys_addr_t types.
Added explanation why kern_reloc is needed.
Split into four patches:
arm64: kexec: make dtb_mem always enabled
arm64: kexec: remove unnecessary debug prints
arm64: kexec: call kexec_image_info only once
arm64: kexec: move relocation function setup
- arm64: kexec: add expandable argument to relocation function
Changed types of new arguments from unsigned long to phys_addr_t.
Changed offset prefix to KEXEC_*
Split into four patches:
arm64: kexec: cpu_soft_restart change argument types
arm64: kexec: arm64_relocate_new_kernel clean-ups
arm64: kexec: arm64_relocate_new_kernel don't use x0 as temp
arm64: kexec: add expandable argument to relocation function
- arm64: kexec: configure trans_pgd page table for kexec
Added invalid entries into EL2 vector table
Removed KEXEC_EL2_VECTOR_TABLE_SIZE and KEXEC_EL2_VECTOR_TABLE_OFFSET
Copy relocation functions and table into separate pages
Changed types in kern_reloc_arg.
Split into three patches:
arm64: kexec: offset for relocation function
arm64: kexec: kexec EL2 vectors
arm64: kexec: configure trans_pgd page table for kexec
- arm64: kexec: enable MMU during kexec relocation
Split into two patches:
arm64: kexec: enable MMU during kexec relocation
arm64: kexec: remove head from relocation argument
v6:
- Sync with mainline tip
- Added Acked's from Dave Young
v5:
- Addressed comments from Matthias Brugger: added review-by's, improved
comments, and made cleanups to swsusp_arch_resume() in addition to
create_safe_exec_page().
- Synced with mainline tip.
v4:
- Addressed comments from James Morse.
- Split "check pgd table allocation" into two patches, and moved to
the beginning of series for simpler backport of the fixes.
Added "Fixes:" tags to commit logs.
- Changed "arm64, hibernate:" to "arm64: hibernate:"
- Added Reviewed-by's
- Moved "add PUD_SECT_RDONLY" earlier in series to be with other
clean-ups
- Added "Derived from:" to arch/arm64/mm/trans_pgd.c
- Removed "flags" from trans_info
- Changed .trans_alloc_page assumption to return zeroed page.
- Simplify changes to trans_pgd_map_page(), by keeping the old
code.
- Simplify changes to trans_pgd_create_copy, by keeping the old
code.
- Removed: "add trans_pgd_create_empty"
- replace init_mm with NULL, and keep using non "__" version of
populate functions.
v3:
- Split changes to create_safe_exec_page() into several patches for
easier review as request by Mark Rutland. This is why this series
has 3 more patches.
- Renamed trans_table to tans_pgd as agreed with Mark. The header
comment in trans_pgd.c explains that trans stands for
transitional page tables. Meaning they are used in transition
between two kernels.
v2:
- Fixed hibernate bug reported by James Morse
- Addressed comments from James Morse:
* More incremental changes to trans_table
* Removed TRANS_FORCEMAP
* Added kexec reboot data for image with 380M in size.
Enable MMU during kexec relocation in order to improve reboot performance.
If kexec functionality is used for a fast system update, with a minimal
downtime, the relocation of kernel + initramfs takes a significant portion
of reboot.
The reason for slow relocation is because it is done without MMU, and thus
not benefiting from D-Cache.
Performance data
----------------
Cavium ThunderX2:
Kernel Image size: 38M Iniramfs size: 46M Total relocation size: 84M
MMU-disabled:
relocation 7.489539915s
MMU-enabled:
relocation 0.03946095s
Relocation performance is improved 190 times.
Broadcom Stingray:
For this experiment, the size of kernel plus initramfs is small, only 25M.
If initramfs was larger, than the improvements would be greater, as time
spent in relocation is proportional to the size of relocation.
MMU-disabled::
kernel shutdown 0.022131328s
relocation 0.440510736s
kernel startup 0.294706768s
Relocation was taking: 58.2% of reboot time
MMU-enabled:
kernel shutdown 0.032066576s
relocation 0.022158152s
kernel startup 0.296055880s
Now: Relocation takes 6.3% of reboot time
Total reboot is x2.16 times faster.
With bigger userland (fitImage 380M), the reboot time is improved by 3.57s,
and is reduced from 3.9s down to 0.33s
Previous approaches and discussions
-----------------------------------
v15: https://lore.kernel.org/lkml/20210609004419.936873-1-pasha.tatashin@soleen.com
v14: https://lore.kernel.org/lkml/20210527150526.271941-1-pasha.tatashin@soleen.com
v13: https://lore.kernel.org/lkml/20210408040537.2703241-1-pasha.tatashin@soleen.com
v12: https://lore.kernel.org/lkml/20210303002230.1083176-1-pasha.tatashin@soleen.com
v11: https://lore.kernel.org/lkml/20210127172706.617195-1-pasha.tatashin@soleen.com
v10: https://lore.kernel.org/linux-arm-kernel/20210125191923.1060122-1-pasha.tatashin@soleen.com
v9: https://lore.kernel.org/lkml/20200326032420.27220-1-pasha.tatashin@soleen.com
v8: https://lore.kernel.org/lkml/20191204155938.2279686-1-pasha.tatashin@soleen.com
v7: https://lore.kernel.org/lkml/20191016200034.1342308-1-pasha.tatashin@soleen.com
v6: https://lore.kernel.org/lkml/20191004185234.31471-1-pasha.tatashin@soleen.com
v5: https://lore.kernel.org/lkml/20190923203427.294286-1-pasha.tatashin@soleen.com
v4: https://lore.kernel.org/lkml/20190909181221.309510-1-pasha.tatashin@soleen.com
v3: https://lore.kernel.org/lkml/20190821183204.23576-1-pasha.tatashin@soleen.com
v2: https://lore.kernel.org/lkml/20190817024629.26611-1-pasha.tatashin@soleen.com
v1: https://lore.kernel.org/lkml/20190801152439.11363-1-pasha.tatashin@soleen.com
Pavel Tatashin (15):
arm64: kernel: add helper for booted at EL2 and not VHE
arm64: trans_pgd: hibernate: Add trans_pgd_copy_el2_vectors
arm64: hibernate: abstract ttrb0 setup function
arm64: kexec: flush image and lists during kexec load time
arm64: kexec: skip relocation code for inplace kexec
arm64: kexec: Use dcache ops macros instead of open-coding
arm64: kexec: pass kimage as the only argument to relocation function
arm64: kexec: configure EL2 vectors for kexec
arm64: kexec: relocate in EL1 mode
arm64: kexec: use ld script for relocation function
arm64: kexec: install a copy of the linear-map
arm64: kexec: keep MMU enabled during kexec relocation
arm64: kexec: remove the pre-kexec PoC maintenance
arm64: kexec: remove cpu-reset.h
arm64: trans_pgd: remove trans_pgd_map_page()
arch/arm64/Kconfig | 2 +-
arch/arm64/include/asm/assembler.h | 49 ++++++--
arch/arm64/include/asm/kexec.h | 12 ++
arch/arm64/include/asm/mmu_context.h | 24 ++++
arch/arm64/include/asm/sections.h | 1 +
arch/arm64/include/asm/trans_pgd.h | 12 +-
arch/arm64/include/asm/virt.h | 7 ++
arch/arm64/kernel/asm-offsets.c | 11 ++
arch/arm64/kernel/cpu-reset.S | 7 +-
arch/arm64/kernel/cpu-reset.h | 32 -----
arch/arm64/kernel/hibernate-asm.S | 72 -----------
arch/arm64/kernel/hibernate.c | 49 ++------
arch/arm64/kernel/machine_kexec.c | 177 ++++++++++++++-------------
arch/arm64/kernel/relocate_kernel.S | 70 +++++------
arch/arm64/kernel/sdei.c | 2 +-
arch/arm64/kernel/vmlinux.lds.S | 19 +++
arch/arm64/mm/Makefile | 1 +
arch/arm64/mm/trans_pgd-asm.S | 65 ++++++++++
arch/arm64/mm/trans_pgd.c | 82 ++++---------
19 files changed, 356 insertions(+), 338 deletions(-)
delete mode 100644 arch/arm64/kernel/cpu-reset.h
create mode 100644 arch/arm64/mm/trans_pgd-asm.S
base-commit: c500bee1c5b2f1d59b1081ac879d73268ab0ff17
--
2.25.1
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:54:17
Replace places that contain logic like this:
is_hyp_mode_available() && !is_kernel_in_hyp_mode()
With a dedicated boolean function is_hyp_nvhe(). This will be needed
later in kexec in order to sooner switch back to EL2.
Suggested-by: James Morse <james.morse@arm.com>
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/include/asm/virt.h | 5 +++++
arch/arm64/kernel/cpu-reset.h | 3 +--
arch/arm64/kernel/hibernate.c | 2 +-
arch/arm64/kernel/sdei.c | 2 +-
4 files changed, 8 insertions(+), 4 deletions(-)
@@ -49,7 +49,7 @@externintin_suspend;/* Do we need to reset el2? */-#define el2_reset_needed() (is_hyp_mode_available() && !is_kernel_in_hyp_mode())+#define el2_reset_needed() (is_hyp_nvhe())/* temporary el2 vectors in the __hibernate_exit_text section. */externcharhibernate_el2_vectors[];
@@ -202,7 +202,7 @@ unsigned long sdei_arch_get_entry_point(int conduit)*droppedtoEL1becausewedon'tsupportVHE,thenwecan'tsupport*SDEI.*/-if(is_hyp_mode_available()&&!is_kernel_in_hyp_mode()){+if(is_hyp_nvhe()){pr_err("Not supported on this hardware/boot configuration\n");gotoout_err;}
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:54:24
Users of trans_pgd may also need a copy of vector table because it is
also may be overwritten if a linear map can be overwritten.
Move setup of EL2 vectors from hibernate to trans_pgd, so it can be
later shared with kexec as well.
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/include/asm/trans_pgd.h | 7 +++-
arch/arm64/include/asm/virt.h | 2 ++
arch/arm64/kernel/hibernate-asm.S | 52 ---------------------------
arch/arm64/kernel/hibernate.c | 26 ++++++--------
arch/arm64/mm/Makefile | 1 +
arch/arm64/mm/trans_pgd-asm.S | 58 ++++++++++++++++++++++++++++++
arch/arm64/mm/trans_pgd.c | 25 ++++++++++++-
7 files changed, 101 insertions(+), 70 deletions(-)
create mode 100644 arch/arm64/mm/trans_pgd-asm.S
@@ -51,9 +51,6 @@ extern int in_suspend;/* Do we need to reset el2? */#define el2_reset_needed() (is_hyp_nvhe())-/* temporary el2 vectors in the __hibernate_exit_text section. */-externcharhibernate_el2_vectors[];-/* hyp-stub vectors, used to restore el2 during resume from hibernate. */externchar__hyp_stub_vectors[];
@@ -434,6 +431,7 @@ int swsusp_arch_resume(void)void*zero_page;size_texit_size;pgd_t*tmp_pg_dir;+phys_addr_tel2_vectors;void__noreturn(*hibernate_exit)(phys_addr_t,phys_addr_t,void*,void*,phys_addr_t,phys_addr_t);structtrans_pgd_infotrans_info={
@@ -461,6 +459,14 @@ int swsusp_arch_resume(void)return-ENOMEM;}+if(el2_reset_needed()){+rc=trans_pgd_copy_el2_vectors(&trans_info,&el2_vectors);+if(rc){+pr_err("Failed to setup el2 vectors\n");+returnrc;+}+}+exit_size=__hibernate_exit_text_end-__hibernate_exit_text_start;/**Copyswsusp_arch_suspend_exit()toasafepage.Thiswillgenerate
@@ -473,26 +479,14 @@ int swsusp_arch_resume(void)returnrc;}-/*-*Thehibernateexittextcontainsasetofel2vectors,thatwill-*beexecutedatel2withthemmuoffinordertoreloadhyp-stub.-*/-dcache_clean_inval_poc((unsignedlong)hibernate_exit,-(unsignedlong)hibernate_exit+exit_size);-/**KASLRwillcausetheel2vectorstobeinadifferentlocationin*theresumedkernel.Loadhibernate'stemporarycopyintoel2.**WecanskipthisstepifwebootedatEL1,orarerunningwithVHE.*/-if(el2_reset_needed()){-phys_addr_tel2_vectors=(phys_addr_t)hibernate_exit;-el2_vectors+=hibernate_el2_vectors--__hibernate_exit_text_start;/* offset */-+if(el2_reset_needed())__hyp_set_vectors(el2_vectors);-}hibernate_exit(virt_to_phys(tmp_pg_dir),resume_hdr.ttbr1_el1,resume_hdr.reenter_kernel,restore_pblist,
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:54:27
Currently, only hibernate sets custom ttbr0 with safe idmaped function.
Kexec, is also going to be using this functionality when relocation code
is going to be idmapped.
Move the setup sequence to a dedicated cpu_install_ttbr0() for custom
ttbr0.
Suggested-by: James Morse <james.morse@arm.com>
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/include/asm/mmu_context.h | 24 ++++++++++++++++++++++++
arch/arm64/kernel/hibernate.c | 21 +--------------------
2 files changed, 25 insertions(+), 20 deletions(-)
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:54:28
In case of kdump or when segments are already in place the relocation
is not needed, therefore the setup of relocation function and call to
it can be skipped.
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
Suggested-by: James Morse <james.morse@arm.com>
---
arch/arm64/kernel/machine_kexec.c | 34 ++++++++++++++++++-----------
arch/arm64/kernel/relocate_kernel.S | 3 ---
2 files changed, 21 insertions(+), 16 deletions(-)
@@ -144,16 +144,16 @@ int machine_kexec_post_load(struct kimage *kimage){void*reloc_code=page_to_virt(kimage->control_code_page);-/* If in place flush new kernel image, else flush lists and buffers */-if(kimage->head&IND_DONE)+/* If in place, relocation is not used, only flush next kernel */+if(kimage->head&IND_DONE){kexec_segment_flush(kimage);-else-kexec_list_flush(kimage);+kexec_image_info(kimage);+return0;+}memcpy(reloc_code,arm64_relocate_new_kernel,arm64_relocate_new_kernel_size);kimage->arch.kern_reloc=__pa(reloc_code);-kexec_image_info(kimage);/* Flush the reloc_code in preparation for its execution. */dcache_clean_inval_poc((unsignedlong)reloc_code,
@@ -162,6 +162,8 @@ int machine_kexec_post_load(struct kimage *kimage)icache_inval_pou((uintptr_t)reloc_code,(uintptr_t)reloc_code+arm64_relocate_new_kernel_size);+kexec_list_flush(kimage);+kexec_image_info(kimage);return0;}
@@ -188,19 +190,25 @@ void machine_kexec(struct kimage *kimage)local_daif_mask();/*-*cpu_soft_restartwillshutdowntheMMU,disabledatacaches,then-*transfercontroltothekern_relocwhichcontainsacopyof-*thearm64_relocate_new_kernelroutine.arm64_relocate_new_kernel-*usesphysicaladdressingtorelocatethenewimagetoitsfinal-*positionandtransferscontroltotheimageentrypointwhenthe-*relocationiscomplete.+*Bothrestartandcpu_soft_restartwillshutdowntheMMU,disabledata+*caches.However,restartwillstartnewkernelorpurgatorydirectly,+*cpu_soft_restartwilltransfercontroltoarm64_relocate_new_kernel*Inkexeccase,kimage->startpointstopurgatoryassumingthat*kernelentryanddtbaddressareembeddedinpurgatoryby*userspace(kexec-tools).*Inkexec_filecase,thekernelstartsdirectlywithoutpurgatory.*/-cpu_soft_restart(kimage->arch.kern_reloc,kimage->head,kimage->start,-kimage->arch.dtb_mem);+if(kimage->head&IND_DONE){+typeof(__cpu_soft_restart)*restart;++cpu_install_idmap();+restart=(void*)__pa_symbol(function_nocfi(__cpu_soft_restart));+restart(is_hyp_nvhe(),kimage->start,kimage->arch.dtb_mem,+0,0);+}else{+cpu_soft_restart(kimage->arch.kern_reloc,kimage->head,+kimage->start,kimage->arch.dtb_mem);+}BUG();/* Should never get here. */}
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:54:35
Currently, during kexec load we are copying relocation function and
flushing it. However, we can also flush kexec relocation buffers and
if new kernel image is already in place (i.e. crash kernel), we can
also flush the new kernel image itself.
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/kernel/machine_kexec.c | 58 ++++++++++++++-----------------
1 file changed, 26 insertions(+), 32 deletions(-)
@@ -163,6 +140,32 @@ static void kexec_segment_flush(const struct kimage *kimage)}}+intmachine_kexec_post_load(structkimage*kimage)+{+void*reloc_code=page_to_virt(kimage->control_code_page);++/* If in place flush new kernel image, else flush lists and buffers */+if(kimage->head&IND_DONE)+kexec_segment_flush(kimage);+else+kexec_list_flush(kimage);++memcpy(reloc_code,arm64_relocate_new_kernel,+arm64_relocate_new_kernel_size);+kimage->arch.kern_reloc=__pa(reloc_code);+kexec_image_info(kimage);++/* Flush the reloc_code in preparation for its execution. */+dcache_clean_inval_poc((unsignedlong)reloc_code,+(unsignedlong)reloc_code++arm64_relocate_new_kernel_size);+icache_inval_pou((uintptr_t)reloc_code,+(uintptr_t)reloc_code++arm64_relocate_new_kernel_size);++return0;+}+/***machine_kexec-Dothekexecreboot.*
@@ -180,13 +183,6 @@ void machine_kexec(struct kimage *kimage)WARN(in_kexec_crash&&(stuck_cpus||smp_crash_stop_failed()),"Some CPUs may be stale, kdump will be unreliable.\n");-/* Flush the kimage list and its buffers. */-kexec_list_flush(kimage);--/* Flush the new image if already in place. */-if((kimage!=kexec_crash_image)&&(kimage->head&IND_DONE))-kexec_segment_flush(kimage);-pr_info("Bye!\n");local_daif_mask();
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:54:40
Currently, kexec relocation function (arm64_relocate_new_kernel) accepts
the following arguments:
head: start of array that contains relocation information.
entry: entry point for new kernel or purgatory.
dtb_mem: first and only argument to entry.
The number of arguments cannot be easily expended, because this
function is also called from HVC_SOFT_RESTART, which preserves only
three arguments. And, also arm64_relocate_new_kernel is written in
assembly but called without stack, thus no place to move extra arguments
to free registers.
Soon, we will need to pass more arguments: once we enable MMU we
will need to pass information about page tables.
Pass kimage to arm64_relocate_new_kernel, and teach it to get the
required fields from kimage.
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/kernel/asm-offsets.c | 7 +++++++
arch/arm64/kernel/machine_kexec.c | 7 +++++--
arch/arm64/kernel/relocate_kernel.S | 10 ++++------
3 files changed, 16 insertions(+), 8 deletions(-)
@@ -206,8 +209,8 @@ void machine_kexec(struct kimage *kimage)restart(is_hyp_nvhe(),kimage->start,kimage->arch.dtb_mem,0,0);}else{-cpu_soft_restart(kimage->arch.kern_reloc,kimage->head,-kimage->start,kimage->arch.dtb_mem);+cpu_soft_restart(kimage->arch.kern_reloc,virt_to_phys(kimage),+0,0);}BUG();/* Should never get here. */
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:54:45
kexec does dcache maintenance when it re-writes all memory. Our
dcache_by_line_op macro depends on reading the sanitized DminLine
from memory. Kexec may have overwritten this, so open-codes the
sequence.
dcache_by_line_op is a whole set of macros, it uses dcache_line_size
which uses read_ctr for the sanitsed DminLine. Reading the DminLine
is the first thing the dcache_by_line_op does.
Rename dcache_by_line_op dcache_by_myline_op and take DminLine as
an argument. Kexec can now use the slightly smaller macro.
This makes up-coming changes to the dcache maintenance easier on
the eye.
Code generated by the existing callers is unchanged.
Suggested-by: James Morse <james.morse@arm.com>
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/include/asm/assembler.h | 30 ++++++++++++++++++++++-------
arch/arm64/kernel/relocate_kernel.S | 13 +++----------
2 files changed, 26 insertions(+), 17 deletions(-)
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:54:50
If we have a EL2 mode without VHE, the EL2 vectors are needed in order
to switch to EL2 and jump to new world with hypervisor privileges.
In preparation to MMU enabled relocation, configure our EL2 table now.
Kexec uses #HVC_SOFT_RESTART to branch to the new world, so extend
el1_sync vector that is provided by trans_pgd_copy_el2_vectors() to
support this case.
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/Kconfig | 2 +-
arch/arm64/include/asm/kexec.h | 1 +
arch/arm64/kernel/asm-offsets.c | 1 +
arch/arm64/kernel/machine_kexec.c | 31 +++++++++++++++++++++++++++++++
arch/arm64/mm/trans_pgd-asm.S | 9 ++++++++-
5 files changed, 42 insertions(+), 2 deletions(-)
@@ -143,9 +146,27 @@ static void kexec_segment_flush(const struct kimage *kimage)}}+/* Allocates pages for kexec page table */+staticvoid*kexec_page_alloc(void*arg)+{+structkimage*kimage=(structkimage*)arg;+structpage*page=kimage_alloc_control_pages(kimage,0);++if(!page)+returnNULL;++memset(page_address(page),0,PAGE_SIZE);++returnpage_address(page);+}+intmachine_kexec_post_load(structkimage*kimage){void*reloc_code=page_to_virt(kimage->control_code_page);+structtrans_pgd_infoinfo={+.trans_alloc_page=kexec_page_alloc,+.trans_alloc_arg=kimage,+};/* If in place, relocation is not used, only flush next kernel */if(kimage->head&IND_DONE){
@@ -154,6 +175,14 @@ int machine_kexec_post_load(struct kimage *kimage)return0;}+kimage->arch.el2_vectors=0;+if(is_hyp_nvhe()){+intrc=trans_pgd_copy_el2_vectors(&info,+&kimage->arch.el2_vectors);+if(rc)+returnrc;+}+memcpy(reloc_code,arm64_relocate_new_kernel,arm64_relocate_new_kernel_size);kimage->arch.kern_reloc=__pa(reloc_code);
@@ -24,7 +24,14 @@ SYM_CODE_START_LOCAL(el1_sync)msrvbar_el2,x1movx0,xzreret-1:/*Unexpectedargument,setanerror*/+1:cmpx0,#HVC_SOFT_RESTART /* Called from kexec */+b.ne2f+movx0,x2+movx2,x4+movx4,x1+movx1,x3+brx4+2:/*Unexpectedargument,setanerror*/mov_qx0,HVC_STUB_ERReretSYM_CODE_END(el1_sync)
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:54:53
Since we are going to keep MMU enabled during relocation, we need to
keep EL1 mode throughout the relocation.
Keep EL1 enabled, and switch EL2 only before entering the new world.
Suggested-by: James Morse <james.morse@arm.com>
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/kernel/cpu-reset.h | 3 +--
arch/arm64/kernel/machine_kexec.c | 4 ++--
arch/arm64/kernel/relocate_kernel.S | 13 +++++++++++--
3 files changed, 14 insertions(+), 6 deletions(-)
@@ -240,8 +240,8 @@ void machine_kexec(struct kimage *kimage)}else{if(is_hyp_nvhe())__hyp_set_vectors(kimage->arch.el2_vectors);-cpu_soft_restart(kimage->arch.kern_reloc,virt_to_phys(kimage),-0,0);+cpu_soft_restart(kimage->arch.kern_reloc,+virt_to_phys(kimage),0,0);}BUG();/* Should never get here. */
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:54:54
Currently, relocation code declares start and end variables
which are used to compute its size.
The better way to do this is to use ld script incited, and put relocation
function in its own section.
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/include/asm/sections.h | 1 +
arch/arm64/kernel/machine_kexec.c | 16 ++++++----------
arch/arm64/kernel/relocate_kernel.S | 15 ++-------------
arch/arm64/kernel/vmlinux.lds.S | 19 +++++++++++++++++++
4 files changed, 28 insertions(+), 23 deletions(-)
@@ -21,14 +21,11 @@#include<asm/mmu.h>#include<asm/mmu_context.h>#include<asm/page.h>+#include<asm/sections.h>#include<asm/trans_pgd.h>#include"cpu-reset.h"-/* Global variables for the arm64_relocate_new_kernel routine. */-externconstunsignedchararm64_relocate_new_kernel[];-externconstunsignedlongarm64_relocate_new_kernel_size;-/***kexec_image_info-Fordebuggingoutput.*/
@@ -183,17 +181,15 @@ int machine_kexec_post_load(struct kimage *kimage)returnrc;}-memcpy(reloc_code,arm64_relocate_new_kernel,-arm64_relocate_new_kernel_size);+reloc_size=__relocate_new_kernel_end-__relocate_new_kernel_start;+memcpy(reloc_code,__relocate_new_kernel_start,reloc_size);kimage->arch.kern_reloc=__pa(reloc_code);/* Flush the reloc_code in preparation for its execution. */dcache_clean_inval_poc((unsignedlong)reloc_code,-(unsignedlong)reloc_code+-arm64_relocate_new_kernel_size);+(unsignedlong)reloc_code+reloc_size);icache_inval_pou((uintptr_t)reloc_code,-(uintptr_t)reloc_code+-arm64_relocate_new_kernel_size);+(uintptr_t)reloc_code+reloc_size);kexec_list_flush(kimage);kexec_image_info(kimage);
@@ -348,3 +360,10 @@ ASSERT(swapper_pg_dir - reserved_pg_dir == RESERVED_SWAPPER_OFFSET,ASSERT(swapper_pg_dir-tramp_pg_dir==TRAMP_SWAPPER_OFFSET,"TRAMP_SWAPPER_OFFSET is wrong!")#endif++#ifdef CONFIG_KEXEC_CORE+/*kexecrelocationcodeshouldfitintooneKEXEC_CONTROL_PAGE_SIZE*/+ASSERT(__relocate_new_kernel_end-(__relocate_new_kernel_start&~(SZ_4K-1))+<=SZ_4K,"kexec relocation code is too big or misaligned")+ASSERT(KEXEC_CONTROL_PAGE_SIZE>=SZ_4K,"KEXEC_CONTROL_PAGE_SIZE is brokern")+#endif
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:54:59
To perform the kexec relocation with the MMU enabled, we need a copy
of the linear map.
Create one, and install it from the relocation code. This has to be done
from the assembly code as it will be idmapped with TTBR0. The kernel
runs in TTRB1, so can't use the break-before-make sequence on the mapping
it is executing from.
The makes no difference yet as the relocation code runs with the MMU
disabled.
Suggested-by: James Morse <james.morse@arm.com>
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/include/asm/assembler.h | 19 +++++++++++++++++++
arch/arm64/include/asm/kexec.h | 2 ++
arch/arm64/kernel/asm-offsets.c | 2 ++
arch/arm64/kernel/hibernate-asm.S | 20 --------------------
arch/arm64/kernel/machine_kexec.c | 16 ++++++++++++++--
arch/arm64/kernel/relocate_kernel.S | 3 +++
6 files changed, 40 insertions(+), 22 deletions(-)
@@ -175,12 +177,22 @@ int machine_kexec_post_load(struct kimage *kimage)kimage->arch.el2_vectors=0;if(is_hyp_nvhe()){-intrc=trans_pgd_copy_el2_vectors(&info,-&kimage->arch.el2_vectors);+rc=trans_pgd_copy_el2_vectors(&info,+&kimage->arch.el2_vectors);if(rc)returnrc;}+/* Create a copy of the linear map */+trans_pgd=kexec_page_alloc(kimage);+if(!trans_pgd)+return-ENOMEM;+rc=trans_pgd_create_copy(&info,&trans_pgd,PAGE_OFFSET,PAGE_END);+if(rc)+returnrc;+kimage->arch.ttbr1=__pa(trans_pgd);+kimage->arch.zero_page=__pa(empty_zero_page);+reloc_size=__relocate_new_kernel_end-__relocate_new_kernel_start;memcpy(reloc_code,__relocate_new_kernel_start,reloc_size);kimage->arch.kern_reloc=__pa(reloc_code);
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:55:06
Now, that we have linear map page tables configured, keep MMU enabled
to allow faster relocation of segments to final destination.
Cavium ThunderX2:
Kernel Image size: 38M Iniramfs size: 46M Total relocation size: 84M
MMU-disabled:
relocation 7.489539915s
MMU-enabled:
relocation 0.03946095s
Broadcom Stingray:
The performance data: for a moderate size kernel + initramfs: 25M the
relocation was taking 0.382s, with enabled MMU it now takes
0.019s only or x20 improvement.
The time is proportional to the size of relocation, therefore if initramfs
is larger, 100M it could take over a second.
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/include/asm/kexec.h | 3 +++
arch/arm64/kernel/asm-offsets.c | 1 +
arch/arm64/kernel/machine_kexec.c | 16 +++++++++++----
arch/arm64/kernel/relocate_kernel.S | 31 +++++++++++++++++++----------
4 files changed, 36 insertions(+), 15 deletions(-)
@@ -196,6 +196,11 @@ int machine_kexec_post_load(struct kimage *kimage)reloc_size=__relocate_new_kernel_end-__relocate_new_kernel_start;memcpy(reloc_code,__relocate_new_kernel_start,reloc_size);kimage->arch.kern_reloc=__pa(reloc_code);+rc=trans_pgd_idmap_page(&info,&kimage->arch.ttbr0,+&kimage->arch.t0sz,reloc_code);+if(rc)+returnrc;+kimage->arch.phys_offset=virt_to_phys(kimage)-(long)kimage;/* Flush the reloc_code in preparation for its execution. */dcache_clean_inval_poc((unsignedlong)reloc_code,
@@ -246,10 +251,13 @@ void machine_kexec(struct kimage *kimage)restart(is_hyp_nvhe(),kimage->start,kimage->arch.dtb_mem,0,0);}else{+void(*kernel_reloc)(structkimage*kimage);+if(is_hyp_nvhe())__hyp_set_vectors(kimage->arch.el2_vectors);-cpu_soft_restart(kimage->arch.kern_reloc,-virt_to_phys(kimage),0,0);+cpu_install_ttbr0(kimage->arch.ttbr0,kimage->arch.t0sz);+kernel_reloc=(void*)kimage->arch.kern_reloc;+kernel_reloc(kimage);}BUG();/* Should never get here. */
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:55:09
Now that kexec does its relocations with the MMU enabled, we no longer
need to clean the relocation data to the PoC.
Suggested-by: James Morse <james.morse@arm.com>
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/kernel/machine_kexec.c | 43 -------------------------------
1 file changed, 43 deletions(-)
@@ -77,48 +77,6 @@ int machine_kexec_prepare(struct kimage *kimage)return0;}-/**-*kexec_list_flush-HelpertoflushthekimagelistandsourcepagestoPoC.-*/-staticvoidkexec_list_flush(structkimage*kimage)-{-kimage_entry_t*entry;--dcache_clean_inval_poc((unsignedlong)kimage,-(unsignedlong)kimage+sizeof(*kimage));--for(entry=&kimage->head;;entry++){-unsignedintflag;-unsignedlongaddr;--/* flush the list entries. */-dcache_clean_inval_poc((unsignedlong)entry,-(unsignedlong)entry+-sizeof(kimage_entry_t));--flag=*entry&IND_FLAGS;-if(flag==IND_DONE)-break;--addr=(unsignedlong)phys_to_virt(*entry&PAGE_MASK);--switch(flag){-caseIND_INDIRECTION:-/* Set entry point just before the new list page. */-entry=(kimage_entry_t*)addr-1;-break;-caseIND_SOURCE:-/* flush the source pages. */-dcache_clean_inval_poc(addr,addr+PAGE_SIZE);-break;-caseIND_DESTINATION:-break;-default:-BUG();-}-}-}-/***kexec_segment_flush-HelpertoflushthekimagesegmentstoPoC.*/
@@ -207,7 +165,6 @@ int machine_kexec_post_load(struct kimage *kimage)(unsignedlong)reloc_code+reloc_size);icache_inval_pou((uintptr_t)reloc_code,(uintptr_t)reloc_code+reloc_size);-kexec_list_flush(kimage);kexec_image_info(kimage);return0;
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:55:13
This header contains only cpu_soft_restart() which is never used directly
anymore. So, remove this header, and rename the helper to be
cpu_soft_restart().
Suggested-by: James Morse <james.morse@arm.com>
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/include/asm/kexec.h | 6 ++++++
arch/arm64/kernel/cpu-reset.S | 7 +++----
arch/arm64/kernel/cpu-reset.h | 30 ------------------------------
arch/arm64/kernel/machine_kexec.c | 6 ++----
4 files changed, 11 insertions(+), 38 deletions(-)
delete mode 100644 arch/arm64/kernel/cpu-reset.h
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-02 21:55:17
The intend of trans_pgd_map_page() was to map contiguous range of VA
memory to the memory that is getting relocated during kexec. However,
since we are now using linear map instead of contiguous range this
function is not needed
Suggested-by: Pingfan Liu <redacted>
Signed-off-by: Pavel Tatashin <pasha.tatashin@soleen.com>
---
arch/arm64/include/asm/trans_pgd.h | 5 +--
arch/arm64/mm/trans_pgd.c | 57 ------------------------------
2 files changed, 1 insertion(+), 61 deletions(-)
Hi Pavel,
This series is still missing reviews from those who understand kexec
better than me.
On Mon, Aug 02, 2021 at 05:53:53PM -0400, Pavel Tatashin wrote:
Enable MMU during kexec relocation in order to improve reboot performance.
If kexec functionality is used for a fast system update, with a minimal
downtime, the relocation of kernel + initramfs takes a significant portion
of reboot.
The reason for slow relocation is because it is done without MMU, and thus
not benefiting from D-Cache.
The performance improvements are indeed significant on some platforms
(going from 7s to ~40ms), so I think the merging the series is worth it.
Some general questions so I better understand the impact:
- Is the kdump path affected in any way? IIUC that doesn't need any
relocation but we should also make sure we don't create the additional
page table unnecessarily (should keep as much memory intact as
possible). Maybe that's already handled.
- What happens if trans_pgd_create_copy() fails to allocate memory. Does
it fall back to an MMU-off relocation?
And I presume this series does not introduce any changes to the kexec
tools ABI.
Thanks.
--
Catalin
From: Pavel Tatashin <pasha.tatashin@soleen.com> Date: 2021-08-26 15:04:03
On Tue, Aug 24, 2021 at 2:06 PM Catalin Marinas [off-list ref] wrote:
Hi Pavel,
This series is still missing reviews from those who understand kexec
better than me.
Hi Catalin,
Yes, I am looking for reviewers.
On Mon, Aug 02, 2021 at 05:53:53PM -0400, Pavel Tatashin wrote:
quoted
Enable MMU during kexec relocation in order to improve reboot performance.
If kexec functionality is used for a fast system update, with a minimal
downtime, the relocation of kernel + initramfs takes a significant portion
of reboot.
The reason for slow relocation is because it is done without MMU, and thus
not benefiting from D-Cache.
The performance improvements are indeed significant on some platforms
(going from 7s to ~40ms), so I think the merging the series is worth it.
Some general questions so I better understand the impact:
- Is the kdump path affected in any way? IIUC that doesn't need any
relocation but we should also make sure we don't create the additional
page table unnecessarily (should keep as much memory intact as
possible). Maybe that's already handled.
Because kdump does not need relocation, we do not reserve pages for
the page table in the kdump reboot case. In fact, with this series,
kdump reboot becomes more straightforward as we skip the relocation
function entirely, and jump directly into the crash kernel (or
purgatory if kexec tools loaded them).
- What happens if trans_pgd_create_copy() fails to allocate memory. Does
it fall back to an MMU-off relocation?
In case we are so low on memory that trans_pgd_create_copy() fails to
allocate the linear map that uses the large pages (the size of the
page table is tiny) the kexec fails during kexec load time (not during
reboot time), as out of memory. The MMU enabled kexec reboot is always
on, and we should not have several ways to do kexec reboot as it makes
the kexec reboot unpredictable in terms of performance, and also prone
to bugs by having a common MMU enabled path and less common path when
we are low on memory which is never tested.
And I presume this series does not introduce any changes to the kexec
tools ABI.
Correct.
Thanks for taking a look at this series.
Pasha
From: Pingfan Liu <hidden> Date: 2021-09-08 08:59:55
On Mon, Aug 02, 2021 at 05:53:53PM -0400, Pavel Tatashin wrote:
Changelog:
v16:
- Merged with 5.14-rc4
v15:
- Changed trans_pgd_copy_el2_vectors() to use vector table that
only shared by kexec and hibernate. This way sync does not have
dangling branch that was recently introduced. (Reported by Marc
Zyngier)
- Renamed is_hyp_callable() to is_hyp_nvhe() as requested by Marc
Zyngier
- Clean-ups, comment fixes.
- Sync with upstream 368094df48e680fa51cedb68537408cfa64b788e
v14:
- Fixed a bug in "arm64: hyp-stub: Move elx_sync into the vectors"
that was noticed by Marc Zyngier
- Merged with upstream
v13:
- Fixed a hang on ThunderX2, thank you Pingfan Liu for reporting
the problem. In relocation function we need civac not ivac, we
need to clean data in addition to invalidating it.
Since I was using ThunderX2 machine I also measured the new
performance data on this large ARM64 server. The MMU improves
kexec relocation 190 times on this machine! (see below for
raw data). Saves 7.5s during CentOS kexec reboot.
v12:
- A major change compared to previous version. Instead of using
contiguous VA range a copy of linear map is now used to perform
copying of segments during relocation as it was agreed in the
discussion of version 11 of this project.
- In addition to using linear map, I also took several ideas from
James Morse to better organize the kexec relocation:
1. skip relocation function entirely if that is not needed
2. remove the PoC flushing function since it is not needed
anymore with MMU enabled.
v11:
- Fixed missing KEXEC_CORE dependency for trans_pgd.c
- Removed useless "if(rc) return rc" statement (thank you Tyler Hicks)
- Another 12 patches were accepted into maintainer's get.
Re-based patches against:
https://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux.git
Branch: for-next/kexec
v10:
- Addressed a lot of comments form James Morse and from Marc Zyngier
- Added review-by's
- Synchronized with mainline
v9: - 9 patches from previous series landed in upstream, so now series
is smaller
- Added two patches from James Morse to address idmap issues for machines
with high physical addresses.
- Addressed comments from Selin Dag about compiling issues. He also tested
my series and got similar performance results: ~60 ms instead of ~580 ms
with an initramfs size of ~120MB.
v8:
- Synced with mainline to keep series up-to-date
v7:
-- Addressed comments from James Morse
- arm64: hibernate: pass the allocated pgdp to ttbr0
Removed "Fixes" tag, and added Added Reviewed-by: James Morse
- arm64: hibernate: check pgd table allocation
Sent out as a standalone patch so it can be sent to stable
Series applies on mainline + this patch
- arm64: hibernate: add trans_pgd public functions
Remove second allocation of tmp_pg_dir in swsusp_arch_resume
Added Reviewed-by: James Morse [off-list ref]
- arm64: kexec: move relocation function setup and clean up
Fixed typo in commit log
Changed kern_reloc to phys_addr_t types.
Added explanation why kern_reloc is needed.
Split into four patches:
arm64: kexec: make dtb_mem always enabled
arm64: kexec: remove unnecessary debug prints
arm64: kexec: call kexec_image_info only once
arm64: kexec: move relocation function setup
- arm64: kexec: add expandable argument to relocation function
Changed types of new arguments from unsigned long to phys_addr_t.
Changed offset prefix to KEXEC_*
Split into four patches:
arm64: kexec: cpu_soft_restart change argument types
arm64: kexec: arm64_relocate_new_kernel clean-ups
arm64: kexec: arm64_relocate_new_kernel don't use x0 as temp
arm64: kexec: add expandable argument to relocation function
- arm64: kexec: configure trans_pgd page table for kexec
Added invalid entries into EL2 vector table
Removed KEXEC_EL2_VECTOR_TABLE_SIZE and KEXEC_EL2_VECTOR_TABLE_OFFSET
Copy relocation functions and table into separate pages
Changed types in kern_reloc_arg.
Split into three patches:
arm64: kexec: offset for relocation function
arm64: kexec: kexec EL2 vectors
arm64: kexec: configure trans_pgd page table for kexec
- arm64: kexec: enable MMU during kexec relocation
Split into two patches:
arm64: kexec: enable MMU during kexec relocation
arm64: kexec: remove head from relocation argument
v6:
- Sync with mainline tip
- Added Acked's from Dave Young
v5:
- Addressed comments from Matthias Brugger: added review-by's, improved
comments, and made cleanups to swsusp_arch_resume() in addition to
create_safe_exec_page().
- Synced with mainline tip.
v4:
- Addressed comments from James Morse.
- Split "check pgd table allocation" into two patches, and moved to
the beginning of series for simpler backport of the fixes.
Added "Fixes:" tags to commit logs.
- Changed "arm64, hibernate:" to "arm64: hibernate:"
- Added Reviewed-by's
- Moved "add PUD_SECT_RDONLY" earlier in series to be with other
clean-ups
- Added "Derived from:" to arch/arm64/mm/trans_pgd.c
- Removed "flags" from trans_info
- Changed .trans_alloc_page assumption to return zeroed page.
- Simplify changes to trans_pgd_map_page(), by keeping the old
code.
- Simplify changes to trans_pgd_create_copy, by keeping the old
code.
- Removed: "add trans_pgd_create_empty"
- replace init_mm with NULL, and keep using non "__" version of
populate functions.
v3:
- Split changes to create_safe_exec_page() into several patches for
easier review as request by Mark Rutland. This is why this series
has 3 more patches.
- Renamed trans_table to tans_pgd as agreed with Mark. The header
comment in trans_pgd.c explains that trans stands for
transitional page tables. Meaning they are used in transition
between two kernels.
v2:
- Fixed hibernate bug reported by James Morse
- Addressed comments from James Morse:
* More incremental changes to trans_table
* Removed TRANS_FORCEMAP
* Added kexec reboot data for image with 380M in size.
Enable MMU during kexec relocation in order to improve reboot performance.
If kexec functionality is used for a fast system update, with a minimal
downtime, the relocation of kernel + initramfs takes a significant portion
of reboot.
The reason for slow relocation is because it is done without MMU, and thus
not benefiting from D-Cache.
Performance data
----------------
Cavium ThunderX2:
Kernel Image size: 38M Iniramfs size: 46M Total relocation size: 84M
MMU-disabled:
relocation 7.489539915s
MMU-enabled:
relocation 0.03946095s
Relocation performance is improved 190 times.
Broadcom Stingray:
For this experiment, the size of kernel plus initramfs is small, only 25M.
If initramfs was larger, than the improvements would be greater, as time
spent in relocation is proportional to the size of relocation.
MMU-disabled::
kernel shutdown 0.022131328s
relocation 0.440510736s
kernel startup 0.294706768s
Relocation was taking: 58.2% of reboot time
MMU-enabled:
kernel shutdown 0.032066576s
relocation 0.022158152s
kernel startup 0.296055880s
Now: Relocation takes 6.3% of reboot time
Total reboot is x2.16 times faster.
With bigger userland (fitImage 380M), the reboot time is improved by 3.57s,
and is reduced from 3.9s down to 0.33s
Previous approaches and discussions
-----------------------------------
v15: https://lore.kernel.org/lkml/20210609004419.936873-1-pasha.tatashin@soleen.com
v14: https://lore.kernel.org/lkml/20210527150526.271941-1-pasha.tatashin@soleen.com
v13: https://lore.kernel.org/lkml/20210408040537.2703241-1-pasha.tatashin@soleen.com
v12: https://lore.kernel.org/lkml/20210303002230.1083176-1-pasha.tatashin@soleen.com
v11: https://lore.kernel.org/lkml/20210127172706.617195-1-pasha.tatashin@soleen.com
v10: https://lore.kernel.org/linux-arm-kernel/20210125191923.1060122-1-pasha.tatashin@soleen.com
v9: https://lore.kernel.org/lkml/20200326032420.27220-1-pasha.tatashin@soleen.com
v8: https://lore.kernel.org/lkml/20191204155938.2279686-1-pasha.tatashin@soleen.com
v7: https://lore.kernel.org/lkml/20191016200034.1342308-1-pasha.tatashin@soleen.com
v6: https://lore.kernel.org/lkml/20191004185234.31471-1-pasha.tatashin@soleen.com
v5: https://lore.kernel.org/lkml/20190923203427.294286-1-pasha.tatashin@soleen.com
v4: https://lore.kernel.org/lkml/20190909181221.309510-1-pasha.tatashin@soleen.com
v3: https://lore.kernel.org/lkml/20190821183204.23576-1-pasha.tatashin@soleen.com
v2: https://lore.kernel.org/lkml/20190817024629.26611-1-pasha.tatashin@soleen.com
v1: https://lore.kernel.org/lkml/20190801152439.11363-1-pasha.tatashin@soleen.com
Pavel Tatashin (15):
arm64: kernel: add helper for booted at EL2 and not VHE
arm64: trans_pgd: hibernate: Add trans_pgd_copy_el2_vectors
arm64: hibernate: abstract ttrb0 setup function
arm64: kexec: flush image and lists during kexec load time
arm64: kexec: skip relocation code for inplace kexec
arm64: kexec: Use dcache ops macros instead of open-coding
arm64: kexec: pass kimage as the only argument to relocation function
arm64: kexec: configure EL2 vectors for kexec
arm64: kexec: relocate in EL1 mode
arm64: kexec: use ld script for relocation function
arm64: kexec: install a copy of the linear-map
arm64: kexec: keep MMU enabled during kexec relocation
arm64: kexec: remove the pre-kexec PoC maintenance
arm64: kexec: remove cpu-reset.h
arm64: trans_pgd: remove trans_pgd_map_page()
arch/arm64/Kconfig | 2 +-
arch/arm64/include/asm/assembler.h | 49 ++++++--
arch/arm64/include/asm/kexec.h | 12 ++
arch/arm64/include/asm/mmu_context.h | 24 ++++
arch/arm64/include/asm/sections.h | 1 +
arch/arm64/include/asm/trans_pgd.h | 12 +-
arch/arm64/include/asm/virt.h | 7 ++
arch/arm64/kernel/asm-offsets.c | 11 ++
arch/arm64/kernel/cpu-reset.S | 7 +-
arch/arm64/kernel/cpu-reset.h | 32 -----
arch/arm64/kernel/hibernate-asm.S | 72 -----------
arch/arm64/kernel/hibernate.c | 49 ++------
arch/arm64/kernel/machine_kexec.c | 177 ++++++++++++++-------------
arch/arm64/kernel/relocate_kernel.S | 70 +++++------
arch/arm64/kernel/sdei.c | 2 +-
arch/arm64/kernel/vmlinux.lds.S | 19 +++
arch/arm64/mm/Makefile | 1 +
arch/arm64/mm/trans_pgd-asm.S | 65 ++++++++++
arch/arm64/mm/trans_pgd.c | 82 ++++---------
19 files changed, 356 insertions(+), 338 deletions(-)
delete mode 100644 arch/arm64/kernel/cpu-reset.h
create mode 100644 arch/arm64/mm/trans_pgd-asm.S
base-commit: c500bee1c5b2f1d59b1081ac879d73268ab0ff17
--
On Thu, Aug 26, 2021 at 11:03:21AM -0400, Pavel Tatashin wrote:
On Tue, Aug 24, 2021 at 2:06 PM Catalin Marinas [off-list ref] wrote:
quoted
quoted
Enable MMU during kexec relocation in order to improve reboot performance.
If kexec functionality is used for a fast system update, with a minimal
downtime, the relocation of kernel + initramfs takes a significant portion
of reboot.
The reason for slow relocation is because it is done without MMU, and thus
not benefiting from D-Cache.
The performance improvements are indeed significant on some platforms
(going from 7s to ~40ms), so I think the merging the series is worth it.
Some general questions so I better understand the impact:
- Is the kdump path affected in any way? IIUC that doesn't need any
relocation but we should also make sure we don't create the additional
page table unnecessarily (should keep as much memory intact as
possible). Maybe that's already handled.
Because kdump does not need relocation, we do not reserve pages for
the page table in the kdump reboot case. In fact, with this series,
kdump reboot becomes more straightforward as we skip the relocation
function entirely, and jump directly into the crash kernel (or
purgatory if kexec tools loaded them).
quoted
- What happens if trans_pgd_create_copy() fails to allocate memory. Does
it fall back to an MMU-off relocation?
In case we are so low on memory that trans_pgd_create_copy() fails to
allocate the linear map that uses the large pages (the size of the
page table is tiny) the kexec fails during kexec load time (not during
reboot time), as out of memory. The MMU enabled kexec reboot is always
on, and we should not have several ways to do kexec reboot as it makes
the kexec reboot unpredictable in terms of performance, and also prone
to bugs by having a common MMU enabled path and less common path when
we are low on memory which is never tested.
I think this makes sense, especially since it will fail during the kexec
load time rather than reboot.
I'm ok in principle with this series but I'd need to convince James
Morse to have a another look since he followed it more closely than me.
Could you please rebase it against 5.15-rc1?
Thanks.
--
Catalin
In case we are so low on memory that trans_pgd_create_copy() fails to
allocate the linear map that uses the large pages (the size of the
page table is tiny) the kexec fails during kexec load time (not during
reboot time), as out of memory. The MMU enabled kexec reboot is always
on, and we should not have several ways to do kexec reboot as it makes
the kexec reboot unpredictable in terms of performance, and also prone
to bugs by having a common MMU enabled path and less common path when
we are low on memory which is never tested.
I think this makes sense, especially since it will fail during the kexec
load time rather than reboot.
I'm ok in principle with this series but I'd need to convince James
Morse to have a another look since he followed it more closely than me.
Could you please rebase it against 5.15-rc1?