From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 11:34:07
Firmware-Assisted Dump (FADump) is currently supported only on pSeries
platform. This patch series adds support for PowerNV platform too.
The first few patches refactor the FADump code to make use of common
code across multiple platforms. Then basic FADump support is added for
PowerNV platform. Followed by patches to honour reserved-ranges DT node
while reserving/releasing memory used by FADump. The subsequent patch
processes CPU state data provided by firmware to create and append core
notes to the ELF core file and the next patch adds support to preserve
crash data for subsequent boots (useful in cases like petitboot). The
subsequent patches add support to export opalcore. opalcore makes
debugging of failures in OPAL code easier. Firmware-Assisted Dump
documentation is also updated appropriately.
The patch series is tested with the latest firmware plus the below skiboot
changes for MPIPL support:
https://patchwork.ozlabs.org/project/skiboot/list/?series=119169
("MPIPL support")
Changes in v4:
* Split the patches.
* Rebased to latest upstream kernel version.
* Updated according to latest OPAL changes.
---
Hari Bathini (25):
powerpc/fadump: move internal macros/definitions to a new header
powerpc/fadump: move internal code to a new file
powerpc/fadump: Improve fadump documentation
pseries/fadump: move rtas specific definitions to platform code
pseries/fadump: introduce callbacks for platform specific operations
pseries/fadump: define register/un-register callback functions
pseries/fadump: move out platform specific support from generic code
powerpc/fadump: use FADump instead of fadump for how it is pronounced
opal: add MPIPL interface definitions
powernv/fadump: add fadump support on powernv
powernv/fadump: register kernel metadata address with opal
powernv/fadump: define register/un-register callback functions
powernv/fadump: support copying multiple kernel memory regions
powernv/fadump: process the crashdump by exporting it as /proc/vmcore
powerpc/fadump: Update documentation about OPAL platform support
powerpc/fadump: consider reserved ranges while reserving memory
powerpc/fadump: consider reserved ranges while releasing memory
powernv/fadump: process architected register state data provided by firmware
powernv/fadump: add support to preserve crash data on FADUMP disabled kernel
powerpc/fadump: update documentation about CONFIG_PRESERVE_FA_DUMP
powernv/opalcore: export /sys/firmware/opal/core for analysing opal crashes
powernv/fadump: Warn before processing partial crashdump
powernv/opalcore: provide an option to invalidate /sys/firmware/opal/core file
powernv/fadump: consider f/w load area
powernv/fadump: update documentation about option to release opalcore
Documentation/powerpc/firmware-assisted-dump.txt | 224 +++-
arch/powerpc/Kconfig | 23
arch/powerpc/include/asm/fadump.h | 190 ----
arch/powerpc/include/asm/opal-api.h | 50 +
arch/powerpc/include/asm/opal.h | 6
arch/powerpc/kernel/Makefile | 6
arch/powerpc/kernel/fadump-common.c | 153 +++
arch/powerpc/kernel/fadump-common.h | 203 ++++
arch/powerpc/kernel/fadump.c | 1181 ++++++++--------------
arch/powerpc/kernel/prom.c | 4
arch/powerpc/platforms/powernv/Makefile | 3
arch/powerpc/platforms/powernv/opal-call.c | 3
arch/powerpc/platforms/powernv/opal-core.c | 637 ++++++++++++
arch/powerpc/platforms/powernv/opal-fadump.c | 671 ++++++++++++
arch/powerpc/platforms/powernv/opal-fadump.h | 154 +++
arch/powerpc/platforms/pseries/Makefile | 1
arch/powerpc/platforms/pseries/rtas-fadump.c | 595 +++++++++++
arch/powerpc/platforms/pseries/rtas-fadump.h | 123 ++
18 files changed, 3231 insertions(+), 996 deletions(-)
create mode 100644 arch/powerpc/kernel/fadump-common.c
create mode 100644 arch/powerpc/kernel/fadump-common.h
create mode 100644 arch/powerpc/platforms/powernv/opal-core.c
create mode 100644 arch/powerpc/platforms/powernv/opal-fadump.c
create mode 100644 arch/powerpc/platforms/powernv/opal-fadump.h
create mode 100644 arch/powerpc/platforms/pseries/rtas-fadump.c
create mode 100644 arch/powerpc/platforms/pseries/rtas-fadump.h
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 11:36:21
Though asm/fadump.h is meant to be used by other components dealing
with FADump, it also has macros/definitions internal to FADump code.
Move them to a new header file used within FADump code. This also
makes way for refactoring platform specific FADump code.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/include/asm/fadump.h | 71 ---------------------------
arch/powerpc/kernel/fadump-common.h | 93 +++++++++++++++++++++++++++++++++++
arch/powerpc/kernel/fadump.c | 2 +
3 files changed, 95 insertions(+), 71 deletions(-)
create mode 100644 arch/powerpc/kernel/fadump-common.h
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 11:38:08
Make way for refactoring platform specific FADump code by moving code
that could be referenced from multiple places to fadump-common.c file.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/kernel/Makefile | 2
arch/powerpc/kernel/fadump-common.c | 144 +++++++++++++++++++++++++++++++++++
arch/powerpc/kernel/fadump-common.h | 8 ++
arch/powerpc/kernel/fadump.c | 146 ++---------------------------------
4 files changed, 162 insertions(+), 138 deletions(-)
create mode 100644 arch/powerpc/kernel/fadump-common.c
@@ -0,0 +1,144 @@+/*+*Firmware-AssistedDumpinternalcode.+*+*Copyright2011,IBMCorporation+*Author:MaheshSalgaonkar<mahesh@linux.ibm.com>+*+*Copyright2019,IBMCorp.+*Author:HariBathini<hbathini@linux.ibm.com>+*+*Thisprogramisfreesoftware;youcanredistributeitand/or+*modifyitunderthetermsoftheGNUGeneralPublicLicense+*aspublishedbytheFreeSoftwareFoundation;eitherversion+*2oftheLicense,or(atyouroption)anylaterversion.+*/++#undef DEBUG+#define pr_fmt(fmt) "fadump: " fmt++#include<linux/memblock.h>+#include<linux/elf.h>+#include<linux/mm.h>+#include<linux/crash_core.h>++#include"fadump-common.h"++void*fadump_cpu_notes_buf_alloc(unsignedlongsize)+{+void*vaddr;+structpage*page;+unsignedlongorder,count,i;++order=get_order(size);+vaddr=(void*)__get_free_pages(GFP_KERNEL|__GFP_ZERO,order);+if(!vaddr)+returnNULL;++count=1<<order;+page=virt_to_page(vaddr);+for(i=0;i<count;i++)+SetPageReserved(page+i);+returnvaddr;+}++voidfadump_cpu_notes_buf_free(unsignedlongvaddr,unsignedlongsize)+{+structpage*page;+unsignedlongorder,count,i;++order=get_order(size);+count=1<<order;+page=virt_to_page(vaddr);+for(i=0;i<count;i++)+ClearPageReserved(page+i);+__free_pages(page,order);+}++u32*fadump_regs_to_elf_notes(u32*buf,structpt_regs*regs)+{+structelf_prstatusprstatus;++memset(&prstatus,0,sizeof(prstatus));+/*+*FIXME:HowdoigetPID?DoIreallyneedit?+*prstatus.pr_pid=????+*/+elf_core_copy_kernel_regs(&prstatus.pr_reg,regs);+buf=append_elf_note(buf,CRASH_CORE_NOTE_NAME,NT_PRSTATUS,+&prstatus,sizeof(prstatus));+returnbuf;+}++voidfadump_update_elfcore_header(structfw_dump*fadump_conf,char*bufp)+{+structelfhdr*elf;+structelf_phdr*phdr;++elf=(structelfhdr*)bufp;+bufp+=sizeof(structelfhdr);++/* First note is a place holder for cpu notes info. */+phdr=(structelf_phdr*)bufp;++if(phdr->p_type==PT_NOTE){+phdr->p_paddr=fadump_conf->cpu_notes_buf;+phdr->p_offset=phdr->p_paddr;+phdr->p_memsz=fadump_conf->cpu_notes_buf_size;+phdr->p_filesz=phdr->p_memsz;+}+}++/*+*Returns1,iftherearenoholesinmemoryareabetweend_starttod_end,+*0otherwise.+*/+staticintis_fadump_memory_area_contiguous(unsignedlongd_start,+unsignedlongd_end)+{+structmemblock_region*reg;+unsignedlongstart,end;+intret=0;++for_each_memblock(memory,reg){+start=max_t(unsignedlong,d_start,reg->base);+end=min_t(unsignedlong,d_end,(reg->base+reg->size));+if(d_start<end){+/* Memory hole from d_start to start */+if(start>d_start)+break;++if(end==d_end){+ret=1;+break;+}++d_start=end+1;+}+}++returnret;+}++/*+*Returns1,iftherearenoholesinbootmemoryarea,+*0otherwise.+*/+intis_fadump_boot_mem_contiguous(structfw_dump*fadump_conf)+{+unsignedlongd_start=RMA_START;+unsignedlongd_end=RMA_START+fadump_conf->boot_memory_size;++returnis_fadump_memory_area_contiguous(d_start,d_end);+}++/*+*Returns1,iftherearenoholesinreservedmemoryarea,+*0otherwise.+*/+intis_fadump_reserved_mem_contiguous(structfw_dump*fadump_conf)+{+unsignedlongd_start=fadump_conf->reserve_dump_area_start;+unsignedlongd_end=d_start+fadump_conf->reserve_dump_area_size;++returnis_fadump_memory_area_contiguous(d_start,d_end);+}
@@ -201,67 +200,6 @@ int is_fadump_active(void)returnfw_dump.dump_active;}-/*-*Returns1,iftherearenoholesinbootmemoryarea,-*0otherwise.-*/-staticintis_boot_memory_area_contiguous(void)-{-structmemblock_region*reg;-unsignedlongtstart,tend;-unsignedlongstart_pfn=PHYS_PFN(RMA_START);-unsignedlongend_pfn=PHYS_PFN(RMA_START+fw_dump.boot_memory_size);-unsignedintret=0;--for_each_memblock(memory,reg){-tstart=max(start_pfn,memblock_region_memory_base_pfn(reg));-tend=min(end_pfn,memblock_region_memory_end_pfn(reg));-if(tstart<tend){-/* Memory hole from start_pfn to tstart */-if(tstart>start_pfn)-break;--if(tend==end_pfn){-ret=1;-break;-}--start_pfn=tend+1;-}-}--returnret;-}--/*-*Returnstrue,iftherearenoholesinreservedmemoryarea,-*falseotherwise.-*/-staticboolis_reserved_memory_area_contiguous(void)-{-structmemblock_region*reg;-unsignedlongstart,end;-unsignedlongd_start=fw_dump.reserve_dump_area_start;-unsignedlongd_end=d_start+fw_dump.reserve_dump_area_size;--for_each_memblock(memory,reg){-start=max(d_start,(unsignedlong)reg->base);-end=min(d_end,(unsignedlong)(reg->base+reg->size));-if(d_start<end){-/* Memory hole from d_start to start */-if(start>d_start)-break;--if(end==d_end)-returntrue;--d_start=end+1;-}-}--returnfalse;-}-/* Print firmware assisted dump configurations for debugging purpose. */staticvoidfadump_show_config(void){
@@ -611,9 +549,9 @@ static int register_fw_dump(struct fadump_mem_struct *fdm)" dump. Hardware Error(%d).\n",rc);break;case-3:-if(!is_boot_memory_area_contiguous())+if(!is_fadump_boot_mem_contiguous(&fw_dump))pr_err("Can't have holes in boot memory area while registering fadump\n");-elseif(!is_reserved_memory_area_contiguous())+elseif(!is_fadump_reserved_mem_contiguous(&fw_dump))pr_err("Can't have holes in reserved memory area while"" registering fadump\n");
@@ -743,72 +681,6 @@ fadump_read_registers(struct fadump_reg_entry *reg_entry, struct pt_regs *regs)returnreg_entry;}-staticu32*fadump_regs_to_elf_notes(u32*buf,structpt_regs*regs)-{-structelf_prstatusprstatus;--memset(&prstatus,0,sizeof(prstatus));-/*-*FIXME:HowdoigetPID?DoIreallyneedit?-*prstatus.pr_pid=????-*/-elf_core_copy_kernel_regs(&prstatus.pr_reg,regs);-buf=append_elf_note(buf,CRASH_CORE_NOTE_NAME,NT_PRSTATUS,-&prstatus,sizeof(prstatus));-returnbuf;-}--staticvoidfadump_update_elfcore_header(char*bufp)-{-structelfhdr*elf;-structelf_phdr*phdr;--elf=(structelfhdr*)bufp;-bufp+=sizeof(structelfhdr);--/* First note is a place holder for cpu notes info. */-phdr=(structelf_phdr*)bufp;--if(phdr->p_type==PT_NOTE){-phdr->p_paddr=fw_dump.cpu_notes_buf;-phdr->p_offset=phdr->p_paddr;-phdr->p_filesz=fw_dump.cpu_notes_buf_size;-phdr->p_memsz=fw_dump.cpu_notes_buf_size;-}-return;-}--staticvoid*fadump_cpu_notes_buf_alloc(unsignedlongsize)-{-void*vaddr;-structpage*page;-unsignedlongorder,count,i;--order=get_order(size);-vaddr=(void*)__get_free_pages(GFP_KERNEL|__GFP_ZERO,order);-if(!vaddr)-returnNULL;--count=1<<order;-page=virt_to_page(vaddr);-for(i=0;i<count;i++)-SetPageReserved(page+i);-returnvaddr;-}--staticvoidfadump_cpu_notes_buf_free(unsignedlongvaddr,unsignedlongsize)-{-structpage*page;-unsignedlongorder,count,i;--order=get_order(size);-count=1<<order;-page=virt_to_page(vaddr);-for(i=0;i<count;i++)-ClearPageReserved(page+i);-__free_pages(page,order);-}-/**ReadCPUstatedumpdataandconvertitintoELFnotes.*TheCPUdumpstartswithmagicnumber"REGSAVE".NumCpusOffsetshouldbe
@@ -898,9 +770,9 @@ static int __init fadump_build_cpu_notes(const struct fadump_mem_struct *fdm)final_note(note_buf);if(fdh){-pr_debug("Updating elfcore header (%llx) with cpu notes\n",-fdh->elfcorehdr_addr);-fadump_update_elfcore_header((char*)__va(fdh->elfcorehdr_addr));+addr=fdh->elfcorehdr_addr;+pr_debug("Updating elfcore header(%lx) with cpu notes\n",addr);+fadump_update_elfcore_header(&fw_dump,(char*)__va(addr));}return0;
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 11:39:52
The figures depicting FADump's (Firmware-Assisted Dump) memory layout
are missing some finer details like different memory regions and what
they represent. Improve the documentation by updating those details.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
Documentation/powerpc/firmware-assisted-dump.txt | 65 ++++++++++++----------
1 file changed, 35 insertions(+), 30 deletions(-)
@@ -74,8 +74,9 @@ as follows: there is crash data available from a previous boot. During the early boot OS will reserve rest of the memory above boot memory size effectively booting with restricted memory- size. This will make sure that the second kernel will not- touch any of the dump memory area.+ size. This will make sure that this kernel (also, referred+ to as second kernel or capture kernel) will not touch any+ of the dump memory area. -- User-space tools will read /proc/vmcore to obtain the contents of memory, which holds the previous crashed kernel dump in ELF
@@ -125,48 +126,52 @@ space memory except the user pages that were present in CMA region. o Memory Reservation during first kernel- Low memory Top of memory- 0 boot memory size |- | | |<--Reserved dump area -->| |- V V | Permanent Reservation | V- +-----------+----------/ /---+---+----+-----------+----+------+- | | |CPU|HPTE| DUMP |ELF | |- +-----------+----------/ /---+---+----+-----------+----+------+- | ^- | |- \ /- -------------------------------------------- Boot memory content gets transferred to- reserved area by firmware at the time of- crash+ Low memory Top of memory+ 0 boot memory size |<--Reserved dump area --->| |+ | | | Permanent Reservation | |+ V V | (Preserve area) | V+ +-----------+----------/ /---+---+----+--------+---+----+------++ | | |CPU|HPTE| DUMP |HDR|ELF | |+ +-----------+----------/ /---+---+----+--------+---+----+------++ | ^ ^+ | | |+ \ / |+ ----------------------------------- FADump Header+ Boot memory content gets transferred (meta area)+ to reserved area by firmware at the+ time of crash+ Fig. 1+ o Memory Reservation during second kernel after crash- Low memory Top of memory- 0 boot memory size |- | |<------------- Reserved dump area ----------- -->|- V V V- +-----------+----------/ /---+---+----+-----------+----+------+- | | |CPU|HPTE| DUMP |ELF | |- +-----------+----------/ /---+---+----+-----------+----+------++ Low memory Top of memory+ 0 boot memory size |+ | |<------------- Reserved dump area --------------->|+ V V |<---- Preserve area ----->| V+ +-----------+----------/ /---+---+----+--------+---+----+------++ | | |CPU|HPTE| DUMP |HDR|ELF | |+ +-----------+----------/ /---+---+----+--------+---+----+------+ | | V V Used by second /proc/vmcore kernel to boot Fig. 2-Currently the dump will be copied from /proc/vmcore to a-a new file upon user intervention. The dump data available through-/proc/vmcore will be in ELF format. Hence the existing kdump-infrastructure (kdump scripts) to save the dump works fine with-minor modifications.+Currently the dump will be copied from /proc/vmcore to a new file upon+user intervention. The dump data available through /proc/vmcore will be+in ELF format. Hence the existing kdump infrastructure (kdump scripts)+to save the dump works fine with minor modifications. KDump scripts on+major Distro releases have already been modified to work seemlessly (no+user intervention in saving the dump) when FADump is used, instead of+KDump, as dump mechanism. The tools to examine the dump will be same as the ones used for kdump. How to enable firmware-assisted dump (fadump):--------------------------------------+--------------------------------------------- 1. Set config option CONFIG_FA_DUMP=y and build kernel. 2. Boot into linux kernel with 'fadump=on' kernel cmdline option.
@@ -189,7 +194,7 @@ NOTE: 1. 'fadump_reserve_mem=' parameter has been deprecated. Instead old behaviour. Sysfs/debugfs files:-------------+------------------- Firmware-assisted dump feature uses sysfs file system to hold the control files and debugfs file to display memory reserved region.
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 11:41:58
Currently, FADump is only supported on pSeries but that is going to
change soon with FADump support being added on PowerNV platform. So,
move rtas specific definitions to platform code to allow FADump
to have multiple platforms support.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/include/asm/fadump.h | 112 --------------------------
arch/powerpc/kernel/fadump-common.h | 20 ++++-
arch/powerpc/kernel/fadump.c | 90 +++++++++++----------
arch/powerpc/platforms/pseries/rtas-fadump.h | 107 +++++++++++++++++++++++++
4 files changed, 174 insertions(+), 155 deletions(-)
create mode 100644 arch/powerpc/platforms/pseries/rtas-fadump.h
@@ -248,24 +249,24 @@ static unsigned long init_fadump_mem_struct(struct fadump_mem_struct *fdm,/* Kernel dump sections *//* cpu state data section. */-fdm->cpu_state_data.request_flag=cpu_to_be32(FADUMP_REQUEST_FLAG);-fdm->cpu_state_data.source_data_type=cpu_to_be16(FADUMP_CPU_STATE_DATA);+fdm->cpu_state_data.request_flag=cpu_to_be32(RTAS_FADUMP_REQUEST_FLAG);+fdm->cpu_state_data.source_data_type=cpu_to_be16(RTAS_FADUMP_CPU_STATE_DATA);fdm->cpu_state_data.source_address=0;fdm->cpu_state_data.source_len=cpu_to_be64(fw_dump.cpu_state_data_size);fdm->cpu_state_data.destination_address=cpu_to_be64(addr);addr+=fw_dump.cpu_state_data_size;/* hpte region section */-fdm->hpte_region.request_flag=cpu_to_be32(FADUMP_REQUEST_FLAG);-fdm->hpte_region.source_data_type=cpu_to_be16(FADUMP_HPTE_REGION);+fdm->hpte_region.request_flag=cpu_to_be32(RTAS_FADUMP_REQUEST_FLAG);+fdm->hpte_region.source_data_type=cpu_to_be16(RTAS_FADUMP_HPTE_REGION);fdm->hpte_region.source_address=0;fdm->hpte_region.source_len=cpu_to_be64(fw_dump.hpte_region_size);fdm->hpte_region.destination_address=cpu_to_be64(addr);addr+=fw_dump.hpte_region_size;/* RMA region section */-fdm->rmr_region.request_flag=cpu_to_be32(FADUMP_REQUEST_FLAG);-fdm->rmr_region.source_data_type=cpu_to_be16(FADUMP_REAL_MODE_REGION);+fdm->rmr_region.request_flag=cpu_to_be32(RTAS_FADUMP_REQUEST_FLAG);+fdm->rmr_region.source_data_type=cpu_to_be16(RTAS_FADUMP_REAL_MODE_REGION);fdm->rmr_region.source_address=cpu_to_be64(RMA_START);fdm->rmr_region.source_len=cpu_to_be64(fw_dump.boot_memory_size);fdm->rmr_region.destination_address=cpu_to_be64(addr);
@@ -520,7 +521,7 @@ static int __init early_fadump_reserve_mem(char *p)}early_param("fadump_reserve_mem",early_fadump_reserve_mem);-staticintregister_fw_dump(structfadump_mem_struct*fdm)+staticintregister_fw_dump(structrtas_fadump_mem_struct*fdm){intrc,err;unsignedintwait_time;
@@ -531,7 +532,7 @@ static int register_fw_dump(struct fadump_mem_struct *fdm)do{rc=rtas_call(fw_dump.ibm_configure_kernel_dump,3,1,NULL,FADUMP_REGISTER,fdm,-sizeof(structfadump_mem_struct));+sizeof(structrtas_fadump_mem_struct));wait_time=rtas_busy_delay_time(rc);if(wait_time)
@@ -627,7 +628,7 @@ static inline int fadump_gpr_index(u64 id)inti=-1;charstr[3];-if((id&GPR_MASK)==REG_ID("GPR")){+if((id&GPR_MASK)==fadump_str_to_u64("GPR")){/* get the digits at the end */id&=~GPR_MASK;id>>=24;
@@ -713,7 +714,8 @@ static int __init fadump_build_cpu_notes(const struct fadump_mem_struct *fdm)vaddr=__va(addr);reg_header=vaddr;-if(be64_to_cpu(reg_header->magic_number)!=REGSAVE_AREA_MAGIC){+if(be64_to_cpu(reg_header->magic_number)!=+fadump_str_to_u64("REGSAVE")){printk(KERN_ERR"Unable to read register save area.\n");return-ENOENT;}
@@ -725,7 +727,7 @@ static int __init fadump_build_cpu_notes(const struct fadump_mem_struct *fdm)num_cpus=be32_to_cpu(*((__be32*)(vaddr)));pr_debug("NumCpus : %u\n",num_cpus);vaddr+=sizeof(u32);-reg_entry=(structfadump_reg_entry*)vaddr;+reg_entry=(structrtas_fadump_reg_entry*)vaddr;/* Allocate buffer to hold cpu crash notes. */fw_dump.cpu_notes_buf_size=num_cpus*sizeof(note_buf_t);
@@ -745,22 +747,22 @@ static int __init fadump_build_cpu_notes(const struct fadump_mem_struct *fdm)fdh=__va(fw_dump.fadumphdr_addr);for(i=0;i<num_cpus;i++){-if(be64_to_cpu(reg_entry->reg_id)!=REG_ID("CPUSTRT")){+if(be64_to_cpu(reg_entry->reg_id)!=fadump_str_to_u64("CPUSTRT")){printk(KERN_ERR"Unable to read CPU state data\n");rc=-ENOENT;gotoerror_out;}/* Lower 4 bytes of reg_value contains logical cpu id */-cpu=be64_to_cpu(reg_entry->reg_value)&FADUMP_CPU_ID_MASK;+cpu=be64_to_cpu(reg_entry->reg_value)&RTAS_FADUMP_CPU_ID_MASK;if(fdh&&!cpumask_test_cpu(cpu,&fdh->online_mask)){-SKIP_TO_NEXT_CPU(reg_entry);+RTAS_FADUMP_SKIP_TO_NEXT_CPU(reg_entry);continue;}pr_debug("Reading register data for cpu %d...\n",cpu);if(fdh&&fdh->crashing_cpu==cpu){regs=fdh->regs;note_buf=fadump_regs_to_elf_notes(note_buf,®s);-SKIP_TO_NEXT_CPU(reg_entry);+RTAS_FADUMP_SKIP_TO_NEXT_CPU(reg_entry);}else{reg_entry++;reg_entry=fadump_read_registers(reg_entry,®s);
@@ -798,7 +800,7 @@ static int __init process_fadump(const struct fadump_mem_struct *fdm_active)return-EINVAL;/* Check if the dump data is valid. */-if((be16_to_cpu(fdm_active->header.dump_status_flag)==FADUMP_ERROR_FLAG)||+if((be16_to_cpu(fdm_active->header.dump_status_flag)==RTAS_FADUMP_ERROR_FLAG)||(fdm_active->cpu_state_data.error_flags!=0)||(fdm_active->rmr_region.error_flags!=0)){printk(KERN_ERR"Dump taken by platform is not valid\n");
@@ -1129,7 +1131,7 @@ static unsigned long init_fadump_header(unsigned long addr)fdh->magic_number=FADUMP_CRASH_INFO_MAGIC;fdh->elfcorehdr_addr=addr;/* We will set the crashing cpu id in crash_fadump() during crash. */-fdh->crashing_cpu=CPU_UNKNOWN;+fdh->crashing_cpu=FADUMP_CPU_UNKNOWN;returnaddr;}
@@ -1163,7 +1165,7 @@ static int register_fadump(void)returnregister_fw_dump(&fdm);}-staticintfadump_unregister_dump(structfadump_mem_struct*fdm)+staticintfadump_unregister_dump(structrtas_fadump_mem_struct*fdm){intrc=0;unsignedintwait_time;
@@ -1174,7 +1176,7 @@ static int fadump_unregister_dump(struct fadump_mem_struct *fdm)do{rc=rtas_call(fw_dump.ibm_configure_kernel_dump,3,1,NULL,FADUMP_UNREGISTER,fdm,-sizeof(structfadump_mem_struct));+sizeof(structrtas_fadump_mem_struct));wait_time=rtas_busy_delay_time(rc);if(wait_time)
@@ -1190,7 +1192,7 @@ static int fadump_unregister_dump(struct fadump_mem_struct *fdm)return0;}-staticintfadump_invalidate_dump(conststructfadump_mem_struct*fdm)+staticintfadump_invalidate_dump(conststructrtas_fadump_mem_struct*fdm){intrc=0;unsignedintwait_time;
@@ -1201,7 +1203,7 @@ static int fadump_invalidate_dump(const struct fadump_mem_struct *fdm)do{rc=rtas_call(fw_dump.ibm_configure_kernel_dump,3,1,NULL,FADUMP_INVALIDATE,fdm,-sizeof(structfadump_mem_struct));+sizeof(structrtas_fadump_mem_struct));wait_time=rtas_busy_delay_time(rc);if(wait_time)
@@ -139,36 +127,7 @@ int __init early_init_dt_scan_fw_dump(unsigned long node, const char *uname,if(fdm_active)fw_dump.dump_active=1;-/* Get the sizes required to store dump data for the firmware provided-*dumpsections.-*Foreachdumpsectiontypesupported,a32bitcellwhichdefines-*theIDofasupportedsectionfollowedbytwo32bitcellswhich-*givestehsizeofthesectioninbytes.-*/-sections=of_get_flat_dt_prop(node,"ibm,configure-kernel-dump-sizes",-&size);--if(!sections)-return1;--num_sections=size/(3*sizeof(u32));--for(i=0;i<num_sections;i++,sections+=3){-u32type=(u32)of_read_number(sections,1);--switch(type){-caseRTAS_FADUMP_CPU_STATE_DATA:-fw_dump.cpu_state_data_size=-of_read_ulong(§ions[1],2);-break;-caseRTAS_FADUMP_HPTE_REGION:-fw_dump.hpte_region_size=-of_read_ulong(§ions[1],2);-break;-}-}--return1;+returnret;}/*
@@ -0,0 +1,134 @@+/*+*Firmware-AssistedDumpsupportonPOWERVMplatform.+*+*Copyright2011,IBMCorporation+*Author:MaheshSalgaonkar<mahesh@linux.ibm.com>+*+*Copyright2019,IBMCorp.+*Author:HariBathini<hbathini@linux.ibm.com>+*+*Thisprogramisfreesoftware;youcanredistributeitand/or+*modifyitunderthetermsoftheGNUGeneralPublicLicense+*aspublishedbytheFreeSoftwareFoundation;eitherversion+*2oftheLicense,or(atyouroption)anylaterversion.+*/++#undef DEBUG+#define pr_fmt(fmt) "rtas fadump: " fmt++#include<linux/string.h>+#include<linux/memblock.h>+#include<linux/delay.h>+#include<linux/seq_file.h>+#include<linux/crash_dump.h>++#include<asm/page.h>+#include<asm/prom.h>+#include<asm/rtas.h>+#include<asm/fadump.h>++#include"../../kernel/fadump-common.h"+#include"rtas-fadump.h"++staticulongrtas_fadump_init_mem_struct(structfw_dump*fadump_conf)+{+returnfadump_conf->reserve_dump_area_start;+}++staticintrtas_fadump_register_fadump(structfw_dump*fadump_conf)+{+return-EIO;+}++staticintrtas_fadump_unregister_fadump(structfw_dump*fadump_conf)+{+return-EIO;+}++staticintrtas_fadump_invalidate_fadump(structfw_dump*fadump_conf)+{+return-EIO;+}++/*+*Validateandprocessthedumpdatastoredbyfirmwarebeforeexporting+*itthrough'/proc/vmcore'.+*/+staticint__initrtas_fadump_process_fadump(structfw_dump*fadump_conf)+{+return-EINVAL;+}++staticvoidrtas_fadump_region_show(structfw_dump*fadump_conf,+structseq_file*m)+{+}++staticvoidrtas_fadump_trigger(structfadump_crash_info_header*fdh,+constchar*msg)+{+/* Call ibm,os-term rtas call to trigger firmware assisted dump */+rtas_os_term((char*)msg);+}++staticstructfadump_opsrtas_fadump_ops={+.init_fadump_mem_struct=rtas_fadump_init_mem_struct,+.register_fadump=rtas_fadump_register_fadump,+.unregister_fadump=rtas_fadump_unregister_fadump,+.invalidate_fadump=rtas_fadump_invalidate_fadump,+.process_fadump=rtas_fadump_process_fadump,+.fadump_region_show=rtas_fadump_region_show,+.fadump_trigger=rtas_fadump_trigger,+};++int__initrtas_fadump_dt_scan(structfw_dump*fadump_conf,ulongnode)+{+const__be32*sections;+inti,num_sections;+intsize;+const__be32*token;++/*+*CheckifFirmwareAssisteddumpissupported.ifyes,check+*ifdumphasbeeninitiatedonlastreboot.+*/+token=of_get_flat_dt_prop(node,"ibm,configure-kernel-dump",NULL);+if(!token)+return1;++fadump_conf->ibm_configure_kernel_dump=be32_to_cpu(*token);+fadump_conf->ops=&rtas_fadump_ops;+fadump_conf->fadump_platform=FADUMP_PLATFORM_PSERIES;+fadump_conf->fadump_supported=1;++/* Get the sizes required to store dump data for the firmware provided+*dumpsections.+*Foreachdumpsectiontypesupported,a32bitcellwhichdefines+*theIDofasupportedsectionfollowedbytwo32bitcellswhich+*givesthesizeofthesectioninbytes.+*/+sections=of_get_flat_dt_prop(node,"ibm,configure-kernel-dump-sizes",+&size);++if(!sections)+return1;++num_sections=size/(3*sizeof(u32));++for(i=0;i<num_sections;i++,sections+=3){+u32type=(u32)of_read_number(sections,1);++switch(type){+caseRTAS_FADUMP_CPU_STATE_DATA:+fadump_conf->cpu_state_data_size=+of_read_ulong(§ions[1],2);+break;+caseRTAS_FADUMP_HPTE_REGION:+fadump_conf->hpte_region_size=+of_read_ulong(§ions[1],2);+break;+}+}++return1;+}
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 11:46:04
Make RTAS calls to register and un-register for FADump. Also, update
how fadump_region contents are diplayed to provide more information.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/kernel/fadump-common.h | 2
arch/powerpc/kernel/fadump.c | 164 ++------------------------
arch/powerpc/platforms/pseries/rtas-fadump.c | 163 +++++++++++++++++++++++++-
3 files changed, 176 insertions(+), 153 deletions(-)
@@ -179,61 +178,6 @@ static void fadump_show_config(void)pr_debug("Boot memory size : %lx\n",fw_dump.boot_memory_size);}-staticunsignedlonginit_fadump_mem_struct(structrtas_fadump_mem_struct*fdm,-unsignedlongaddr)-{-if(!fdm)-return0;--memset(fdm,0,sizeof(structrtas_fadump_mem_struct));-addr=addr&PAGE_MASK;--fdm->header.dump_format_version=cpu_to_be32(0x00000001);-fdm->header.dump_num_sections=cpu_to_be16(3);-fdm->header.dump_status_flag=0;-fdm->header.offset_first_dump_section=-cpu_to_be32((u32)offsetof(structrtas_fadump_mem_struct,cpu_state_data));--/*-*Fieldsfordiskdumpoption.-*Wearenotusingdiskdumpoption,hencesetthesefieldsto0.-*/-fdm->header.dd_block_size=0;-fdm->header.dd_block_offset=0;-fdm->header.dd_num_blocks=0;-fdm->header.dd_offset_disk_path=0;--/* set 0 to disable an automatic dump-reboot. */-fdm->header.max_time_auto=0;--/* Kernel dump sections */-/* cpu state data section. */-fdm->cpu_state_data.request_flag=cpu_to_be32(RTAS_FADUMP_REQUEST_FLAG);-fdm->cpu_state_data.source_data_type=cpu_to_be16(RTAS_FADUMP_CPU_STATE_DATA);-fdm->cpu_state_data.source_address=0;-fdm->cpu_state_data.source_len=cpu_to_be64(fw_dump.cpu_state_data_size);-fdm->cpu_state_data.destination_address=cpu_to_be64(addr);-addr+=fw_dump.cpu_state_data_size;--/* hpte region section */-fdm->hpte_region.request_flag=cpu_to_be32(RTAS_FADUMP_REQUEST_FLAG);-fdm->hpte_region.source_data_type=cpu_to_be16(RTAS_FADUMP_HPTE_REGION);-fdm->hpte_region.source_address=0;-fdm->hpte_region.source_len=cpu_to_be64(fw_dump.hpte_region_size);-fdm->hpte_region.destination_address=cpu_to_be64(addr);-addr+=fw_dump.hpte_region_size;--/* RMA region section */-fdm->rmr_region.request_flag=cpu_to_be32(RTAS_FADUMP_REQUEST_FLAG);-fdm->rmr_region.source_data_type=cpu_to_be16(RTAS_FADUMP_REAL_MODE_REGION);-fdm->rmr_region.source_address=cpu_to_be64(RMA_START);-fdm->rmr_region.source_len=cpu_to_be64(fw_dump.boot_memory_size);-fdm->rmr_region.destination_address=cpu_to_be64(addr);-addr+=fw_dump.boot_memory_size;--returnaddr;-}-/***fadump_calculate_reserve_size():reservevariablebootarea5%ofSystemRAM*
@@ -480,61 +424,6 @@ static int __init early_fadump_reserve_mem(char *p)}early_param("fadump_reserve_mem",early_fadump_reserve_mem);-staticintregister_fw_dump(structrtas_fadump_mem_struct*fdm)-{-intrc,err;-unsignedintwait_time;--pr_debug("Registering for firmware-assisted kernel dump...\n");--/* TODO: Add upper time limit for the delay */-do{-rc=rtas_call(fw_dump.ibm_configure_kernel_dump,3,1,NULL,-FADUMP_REGISTER,fdm,-sizeof(structrtas_fadump_mem_struct));--wait_time=rtas_busy_delay_time(rc);-if(wait_time)-mdelay(wait_time);--}while(wait_time);--err=-EIO;-switch(rc){-default:-pr_err("Failed to register. Unknown Error(%d).\n",rc);-break;-case-1:-printk(KERN_ERR"Failed to register firmware-assisted kernel"-" dump. Hardware Error(%d).\n",rc);-break;-case-3:-if(!is_fadump_boot_mem_contiguous(&fw_dump))-pr_err("Can't have holes in boot memory area while registering fadump\n");-elseif(!is_fadump_reserved_mem_contiguous(&fw_dump))-pr_err("Can't have holes in reserved memory area while"-" registering fadump\n");--printk(KERN_ERR"Failed to register firmware-assisted kernel"-" dump. Parameter Error(%d).\n",rc);-err=-EINVAL;-break;-case-9:-printk(KERN_ERR"firmware-assisted kernel dump is already "-" registered.");-fw_dump.dump_registered=1;-err=-EEXIST;-break;-case0:-printk(KERN_INFO"firmware-assisted kernel dump registration"-" is successful\n");-fw_dump.dump_registered=1;-err=0;-break;-}-returnerr;-}-voidcrash_fadump(structpt_regs*regs,constchar*str){structfadump_crash_info_header*fdh=NULL;
@@ -987,7 +875,7 @@ static int fadump_setup_crash_memory_ranges(void)staticinlineunsignedlongfadump_relocate(unsignedlongpaddr){if(paddr>RMA_START&&paddr<fw_dump.boot_memory_size)-returnbe64_to_cpu(fdm.rmr_region.destination_address)+paddr;+returnfw_dump.boot_mem_dest_addr+paddr;elsereturnpaddr;}
@@ -1060,7 +948,7 @@ static int fadump_create_elfcore_headers(char *bufp)*tothespecifieddestination_address.Henceset*thecorrectoffset.*/-phdr->p_offset=be64_to_cpu(fdm.rmr_region.destination_address);+phdr->p_offset=fw_dump.boot_mem_dest_addr;}phdr->p_paddr=mbase;
@@ -1112,7 +1000,8 @@ static int register_fadump(void)if(ret)returnret;-addr=be64_to_cpu(fdm.rmr_region.destination_address)+be64_to_cpu(fdm.rmr_region.source_len);+addr=fw_dump.fadumphdr_addr;+/* Initialize fadump crash info header. */addr=init_fadump_header(addr);vaddr=__va(addr);
@@ -1121,34 +1010,8 @@ static int register_fadump(void)fadump_create_elfcore_headers(vaddr);/* register the future kernel dump with firmware. */-returnregister_fw_dump(&fdm);-}--staticintfadump_unregister_dump(structrtas_fadump_mem_struct*fdm)-{-intrc=0;-unsignedintwait_time;--pr_debug("Un-register firmware-assisted dump\n");--/* TODO: Add upper time limit for the delay */-do{-rc=rtas_call(fw_dump.ibm_configure_kernel_dump,3,1,NULL,-FADUMP_UNREGISTER,fdm,-sizeof(structrtas_fadump_mem_struct));--wait_time=rtas_busy_delay_time(rc);-if(wait_time)-mdelay(wait_time);-}while(wait_time);--if(rc){-printk(KERN_ERR"Failed to un-register firmware-assisted dump."-" unexpected error(%d).\n",rc);-returnrc;-}-fw_dump.dump_registered=0;-return0;+pr_debug("Registering for firmware-assisted kernel dump...\n");+returnfw_dump.ops->register_fadump(&fw_dump);}staticintfadump_invalidate_dump(conststructrtas_fadump_mem_struct*fdm)
@@ -1186,7 +1049,7 @@ void fadump_cleanup(void)fadump_invalidate_dump(fdm_active);}elseif(fw_dump.dump_registered){/* Un-register Firmware-assisted dump if it was registered. */-fadump_unregister_dump(&fdm);+fw_dump.ops->unregister_fadump(&fw_dump);free_crash_memory_ranges();}}
@@ -1296,7 +1159,7 @@ static void fadump_invalidate_release_mem(void)fw_dump.cpu_notes_buf_size=0;}/* Initialize the kernel dump memory structure for FAD registration. */-init_fadump_mem_struct(&fdm,fw_dump.reserve_dump_area_start);+fw_dump.ops->init_fadump_mem_struct(&fw_dump);}staticssize_tfadump_release_memory_store(structkobject*kobj,
@@ -30,19 +30,152 @@#include"../../kernel/fadump-common.h"#include"rtas-fadump.h"+staticstructrtas_fadump_mem_structfdm;++staticvoidrtas_fadump_update_config(structfw_dump*fadump_conf,+conststructrtas_fadump_mem_struct*fdm)+{+fadump_conf->boot_mem_dest_addr=+be64_to_cpu(fdm->rmr_region.destination_address);++fadump_conf->fadumphdr_addr=(fadump_conf->boot_mem_dest_addr++fadump_conf->boot_memory_size);+}+staticulongrtas_fadump_init_mem_struct(structfw_dump*fadump_conf){-returnfadump_conf->reserve_dump_area_start;+ulongaddr=fadump_conf->reserve_dump_area_start;++memset(&fdm,0,sizeof(structrtas_fadump_mem_struct));+addr=addr&PAGE_MASK;++fdm.header.dump_format_version=cpu_to_be32(0x00000001);+fdm.header.dump_num_sections=cpu_to_be16(3);+fdm.header.dump_status_flag=0;+fdm.header.offset_first_dump_section=+cpu_to_be32((u32)offsetof(structrtas_fadump_mem_struct,+cpu_state_data));++/*+*Fieldsfordiskdumpoption.+*Wearenotusingdiskdumpoption,hencesetthesefieldsto0.+*/+fdm.header.dd_block_size=0;+fdm.header.dd_block_offset=0;+fdm.header.dd_num_blocks=0;+fdm.header.dd_offset_disk_path=0;++/* set 0 to disable an automatic dump-reboot. */+fdm.header.max_time_auto=0;++/* Kernel dump sections */+/* cpu state data section. */+fdm.cpu_state_data.request_flag=+cpu_to_be32(RTAS_FADUMP_REQUEST_FLAG);+fdm.cpu_state_data.source_data_type=+cpu_to_be16(RTAS_FADUMP_CPU_STATE_DATA);+fdm.cpu_state_data.source_address=0;+fdm.cpu_state_data.source_len=+cpu_to_be64(fadump_conf->cpu_state_data_size);+fdm.cpu_state_data.destination_address=cpu_to_be64(addr);+addr+=fadump_conf->cpu_state_data_size;++/* hpte region section */+fdm.hpte_region.request_flag=cpu_to_be32(RTAS_FADUMP_REQUEST_FLAG);+fdm.hpte_region.source_data_type=+cpu_to_be16(RTAS_FADUMP_HPTE_REGION);+fdm.hpte_region.source_address=0;+fdm.hpte_region.source_len=+cpu_to_be64(fadump_conf->hpte_region_size);+fdm.hpte_region.destination_address=cpu_to_be64(addr);+addr+=fadump_conf->hpte_region_size;++/* RMA region section */+fdm.rmr_region.request_flag=cpu_to_be32(RTAS_FADUMP_REQUEST_FLAG);+fdm.rmr_region.source_data_type=+cpu_to_be16(RTAS_FADUMP_REAL_MODE_REGION);+fdm.rmr_region.source_address=cpu_to_be64(RMA_START);+fdm.rmr_region.source_len=+cpu_to_be64(fadump_conf->boot_memory_size);+fdm.rmr_region.destination_address=cpu_to_be64(addr);+addr+=fadump_conf->boot_memory_size;++rtas_fadump_update_config(fadump_conf,&fdm);++returnaddr;}staticintrtas_fadump_register_fadump(structfw_dump*fadump_conf){-return-EIO;+intrc,err=-EIO;+unsignedintwait_time;++/* TODO: Add upper time limit for the delay */+do{+rc=rtas_call(fadump_conf->ibm_configure_kernel_dump,3,1,+NULL,FADUMP_REGISTER,&fdm,+sizeof(structrtas_fadump_mem_struct));++wait_time=rtas_busy_delay_time(rc);+if(wait_time)+mdelay(wait_time);++}while(wait_time);++switch(rc){+case0:+pr_info("Registration is successful!\n");+fadump_conf->dump_registered=1;+err=0;+break;+case-1:+pr_err("Failed to register. Hardware Error(%d).\n",rc);+break;+case-3:+if(!is_fadump_boot_mem_contiguous(fadump_conf))+pr_err("Can't hot-remove boot memory area.\n");+elseif(!is_fadump_reserved_mem_contiguous(fadump_conf))+pr_err("Can't hot-remove reserved memory area.\n");++pr_err("Failed to register. Parameter Error(%d).\n",rc);+err=-EINVAL;+break;+case-9:+pr_err("Already registered!\n");+fadump_conf->dump_registered=1;+err=-EEXIST;+break;+default:+pr_err("Failed to register. Unknown Error(%d).\n",rc);+break;+}++returnerr;}staticintrtas_fadump_unregister_fadump(structfw_dump*fadump_conf){-return-EIO;+intrc;+unsignedintwait_time;++/* TODO: Add upper time limit for the delay */+do{+rc=rtas_call(fadump_conf->ibm_configure_kernel_dump,3,1,+NULL,FADUMP_UNREGISTER,&fdm,+sizeof(structrtas_fadump_mem_struct));++wait_time=rtas_busy_delay_time(rc);+if(wait_time)+mdelay(wait_time);+}while(wait_time);++if(rc){+pr_err("Failed to un-register - unexpected error(%d).\n",rc);+return-EIO;+}++fadump_conf->dump_registered=0;+return0;}staticintrtas_fadump_invalidate_fadump(structfw_dump*fadump_conf)
@@ -8,18 +8,18 @@ a crashed system, and to do so from a fully-reset system, and to minimize the total elapsed time until the system is back in production use.-- Firmware assisted dump (fadump) infrastructure is intended to replace+- Firmware-Assisted Dump (FADump) infrastructure is intended to replace the existing phyp assisted dump. - Fadump uses the same firmware interfaces and memory reservation model as phyp assisted dump.-- Unlike phyp dump, fadump exports the memory dump through /proc/vmcore+- Unlike phyp dump, FADump exports the memory dump through /proc/vmcore in the ELF format in the same way as kdump. This helps us reuse the kdump infrastructure for dump capture and filtering. - Unlike phyp dump, userspace tool does not need to refer any sysfs interface while reading /proc/vmcore.-- Unlike phyp dump, fadump allows user to release all the memory reserved+- Unlike phyp dump, FADump allows user to release all the memory reserved for dump, with a single operation of echo 1 > /sys/kernel/fadump_release_mem.-- Once enabled through kernel boot parameter, fadump can be+- Once enabled through kernel boot parameter, FADump can be started/stopped through /sys/kernel/fadump_registered interface (see sysfs files section below) and can be easily integrated with kdump service start/stop init scripts.
@@ -33,7 +33,7 @@ dump offers several strong, practical advantages: in a clean, consistent state. -- Once the dump is copied out, the memory that held the dump is immediately available to the running kernel. And therefore,- unlike kdump, fadump doesn't need a 2nd reboot to get back+ unlike kdump, FADump doesn't need a 2nd reboot to get back the system to the production configuration. The above can only be accomplished by coordination with,
@@ -61,7 +61,7 @@ as follows: boot successfully. For syntax of crashkernel= parameter, refer to Documentation/kdump/kdump.rst. If any offset is provided in crashkernel= parameter, it will be ignored- as fadump uses a predefined offset to reserve memory+ as FADump uses a predefined offset to reserve memory for boot memory dump preservation in case of a crash. -- After the low memory (boot memory) area has been saved, the
@@ -120,7 +120,7 @@ blocking this significant chunk of memory from production kernel. Hence, the implementation uses the Linux kernel's Contiguous Memory Allocator (CMA) for memory reservation if CMA is configured for kernel. With CMA reservation this memory will be available for applications to-use it, while kernel is prevented from using it. With this fadump will+use it, while kernel is prevented from using it. With this FADump will still be able to capture all of the kernel memory and most of the user space memory except the user pages that were present in CMA region.
@@ -170,14 +170,14 @@ KDump, as dump mechanism. The tools to examine the dump will be same as the ones used for kdump.-How to enable firmware-assisted dump (fadump):+How to enable firmware-assisted dump (FADump): --------------------------------------------- 1. Set config option CONFIG_FA_DUMP=y and build kernel.-2. Boot into linux kernel with 'fadump=on' kernel cmdline option.- By default, fadump reserved memory will be initialized as CMA area.- Alternatively, user can boot linux kernel with 'fadump=nocma' to- prevent fadump to use CMA.+2. Boot into linux kernel with 'FADump=on' kernel cmdline option.+ By default, FADump reserved memory will be initialized as CMA area.+ Alternatively, user can boot linux kernel with 'FADump=nocma' to+ prevent FADump to use CMA. 3. Optionally, user can also set 'crashkernel=' kernel cmdline to specify size of the memory to reserve for boot memory dump preservation.
@@ -190,7 +190,7 @@ NOTE: 1. 'fadump_reserve_mem=' parameter has been deprecated. Instead option is set at kernel cmdline. 3. if user wants to capture all of user space memory and ok with reserved memory not available to production system, then- 'fadump=nocma' kernel parameter can be used to fallback to+ 'FADump=nocma' kernel parameter can be used to fallback to old behaviour. Sysfs/debugfs files:
@@ -203,29 +203,29 @@ Here is the list of files under kernel sysfs: /sys/kernel/fadump_enabled- This is used to display the fadump status.- 0 = fadump is disabled- 1 = fadump is enabled+ This is used to display the FADump status.+ 0 = FADump is disabled+ 1 = FADump is enabled This interface can be used by kdump init scripts to identify if- fadump is enabled in the kernel and act accordingly.+ FADump is enabled in the kernel and act accordingly. /sys/kernel/fadump_registered- This is used to display the fadump registration status as well- as to control (start/stop) the fadump registration.- 0 = fadump is not registered.- 1 = fadump is registered and ready to handle system crash.+ This is used to display the FADump registration status as well+ as to control (start/stop) the FADump registration.+ 0 = FADump is not registered.+ 1 = FADump is registered and ready to handle system crash.- To register fadump echo 1 > /sys/kernel/fadump_registered and+ To register FADump echo 1 > /sys/kernel/fadump_registered and echo 0 > /sys/kernel/fadump_registered for un-register and stop the- fadump. Once the fadump is un-registered, the system crash will not+ FADump. Once the FADump is un-registered, the system crash will not be handled and vmcore will not be captured. This interface can be easily integrated with kdump service start/stop. /sys/kernel/fadump_release_mem- This file is available only when fadump is active during+ This file is available only when FADump is active during second kernel. This is used to release the reserved memory region that are held for saving crash dump. To release the reserved memory echo 1 to it:
@@ -244,26 +244,33 @@ Here is the list of files under powerpc debugfs: /sys/kernel/debug/powerpc/fadump_region- This file shows the reserved memory regions if fadump is+ This file shows the reserved memory regions if FADump is enabled otherwise this file is empty. The output format- is:+ for regions provided by f/w is: <region>: [<start>-<end>] <reserved-size> bytes, Dumped: <dump-size>+ and for kernel DUMP region is:++ DUMP: Src: <src-addr>, Dest: <dest-addr>, Size: <size>, Dumped: # bytes+ e.g.- Contents when fadump is registered during first kernel+ Contents when FADump is registered during first kernel # cat /sys/kernel/debug/powerpc/fadump_region CPU : [0x0000006ffb0000-0x0000006fff001f] 0x40020 bytes, Dumped: 0x0 HPTE: [0x0000006fff0020-0x0000006fff101f] 0x1000 bytes, Dumped: 0x0- DUMP: [0x0000006fff1020-0x0000007fff101f] 0x10000000 bytes, Dumped: 0x0+ DUMP: Src: 0x00000000000000, Dest: 0x0000006fff1020, Size: 0x10000000, Dumped: 0x0 bytes+ #- Contents when fadump is active during second kernel+ Contents when FADump is active during second kernel # cat /sys/kernel/debug/powerpc/fadump_region CPU : [0x0000006ffb0000-0x0000006fff001f] 0x40020 bytes, Dumped: 0x40020 HPTE: [0x0000006fff0020-0x0000006fff101f] 0x1000 bytes, Dumped: 0x1000- DUMP: [0x0000006fff1020-0x0000007fff101f] 0x10000000 bytes, Dumped: 0x10000000- : [0x00000010000000-0x0000006ffaffff] 0x5ffb0000 bytes, Dumped: 0x5ffb0000+ DUMP: Src: 0x00000000000000, Dest: 0x0000006fff1020, Size: 0x10000000, Dumped: 0x10000000 bytes++ Memory above 0x0000000010000000 is reserved for saving crash dump+ # NOTE: Please refer to Documentation/filesystems/debugfs.txt on how to mount the debugfs filesystem.
@@ -274,7 +281,7 @@ TODO: o Need to come up with the better approach to find out more accurate boot memory size that is required for a kernel to boot successfully when booted with restricted memory.- o The fadump implementation introduces a fadump crash info structure+ o The FADump implementation introduces a FADump crash info structure in the scratch area before the ELF core header. The idea of introducing this structure is to pass some important crash info data to the second kernel which will help second kernel to populate ELF core header with
@@ -577,7 +577,8 @@ config FA_DUMPismeanttobeakdumpreplacementofferingrobustnessandspeednotpossiblewithoutsystemfirmwareassistance.-Ifunsure,say"N"+Ifunsure,say"y".Onlyspecialkernelslikepetitbootmay+needtosay"N"here.configIRQ_ALL_CPUSbool"Distribute interrupts on all CPUs by default"
@@ -0,0 +1,102 @@+/*+*Firmware-AssistedDumpsupportonPOWERplatform(OPAL).+*+*Copyright2019,IBMCorp.+*Author:HariBathini<hbathini@linux.ibm.com>+*+*Thisprogramisfreesoftware;youcanredistributeitand/or+*modifyitunderthetermsoftheGNUGeneralPublicLicense+*aspublishedbytheFreeSoftwareFoundation;eitherversion+*2oftheLicense,or(atyouroption)anylaterversion.+*/++#undef DEBUG+#define pr_fmt(fmt) "opal fadump: " fmt++#include<linux/string.h>+#include<linux/seq_file.h>+#include<linux/of_fdt.h>+#include<linux/libfdt.h>++#include<asm/opal.h>++#include"../../kernel/fadump-common.h"++staticulongopal_fadump_init_mem_struct(structfw_dump*fadump_conf)+{+returnfadump_conf->reserve_dump_area_start;+}++staticintopal_fadump_register_fadump(structfw_dump*fadump_conf)+{+return-EIO;+}++staticintopal_fadump_unregister_fadump(structfw_dump*fadump_conf)+{+return-EIO;+}++staticintopal_fadump_invalidate_fadump(structfw_dump*fadump_conf)+{+return-EIO;+}++staticint__initopal_fadump_process_fadump(structfw_dump*fadump_conf)+{+return-EINVAL;+}++staticvoidopal_fadump_region_show(structfw_dump*fadump_conf,+structseq_file*m)+{+}++staticvoidopal_fadump_trigger(structfadump_crash_info_header*fdh,+constchar*msg)+{+intrc;++rc=opal_cec_reboot2(OPAL_REBOOT_MPIPL,msg);+if(rc==OPAL_UNSUPPORTED){+pr_emerg("Reboot type %d not supported.\n",+OPAL_REBOOT_MPIPL);+}elseif(rc==OPAL_HARDWARE)+pr_emerg("No backend support for MPIPL!\n");+}++staticstructfadump_opsopal_fadump_ops={+.init_fadump_mem_struct=opal_fadump_init_mem_struct,+.register_fadump=opal_fadump_register_fadump,+.unregister_fadump=opal_fadump_unregister_fadump,+.invalidate_fadump=opal_fadump_invalidate_fadump,+.process_fadump=opal_fadump_process_fadump,+.fadump_region_show=opal_fadump_region_show,+.fadump_trigger=opal_fadump_trigger,+};++int__initopal_fadump_dt_scan(structfw_dump*fadump_conf,ulongnode)+{+unsignedlongdn;++/*+*CheckifFirmware-AssistedDumpissupported.ifyes,check+*ifdumphasbeeninitiatedonlastreboot.+*/+dn=of_get_flat_dt_subnode_by_name(node,"dump");+if(dn==-FDT_ERR_NOTFOUND){+pr_debug("FADump support is missing!\n");+return1;+}++if(!of_flat_dt_is_compatible(dn,"ibm,opal-dump")){+pr_err("Support missing for this f/w version!\n");+return1;+}++fadump_conf->ops=&opal_fadump_ops;+fadump_conf->fadump_platform=FADUMP_PLATFORM_POWERNV;+fadump_conf->fadump_supported=1;++return1;+}
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 11:53:56
OPAL allows registering address with it in the first kernel and
retrieving it after MPIPL. Setup kernel metadata and register its
address with OPAL to use it for processing the crash dump.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/kernel/fadump-common.h | 4 +
arch/powerpc/kernel/fadump.c | 65 ++++++++++++++---------
arch/powerpc/platforms/powernv/opal-fadump.c | 73 ++++++++++++++++++++++++++
arch/powerpc/platforms/powernv/opal-fadump.h | 37 +++++++++++++
arch/powerpc/platforms/pseries/rtas-fadump.c | 32 +++++++++--
5 files changed, 177 insertions(+), 34 deletions(-)
create mode 100644 arch/powerpc/platforms/powernv/opal-fadump.h
@@ -258,6 +258,9 @@ static unsigned long get_fadump_area_size(void)size+=sizeof(structelf_phdr)*(memblock_num_regions(memory)+2);size=PAGE_ALIGN(size);++/* This is to hold kernel metadata on platforms that support it */+size+=fw_dump.ops->get_kernel_metadata_size();returnsize;}
@@ -283,17 +286,17 @@ static void __init fadump_reserve_crash_area(unsigned long base,int__initfadump_reserve_mem(void){+intret=1;unsignedlongbase,size,memory_boundary;if(!fw_dump.fadump_enabled)return0;if(!fw_dump.fadump_supported){-printk(KERN_INFO"Firmware-assisted dump is not supported on"-" this hardware\n");-fw_dump.fadump_enabled=0;-return0;+pr_info("Firmware-Assisted Dump is not supported on this hardware\n");+gotoerror_out;}+/**Initializebootmemorysize*Ifdumpisactivethenwehavealreadycalculatedthesizeduring
@@ -310,11 +313,13 @@ int __init fadump_reserve_mem(void)}size=get_fadump_area_size();+fw_dump.reserve_dump_area_size=size;if(memory_limit)memory_boundary=memory_limit;elsememory_boundary=memblock_end_of_DRAM();+base=fw_dump.boot_memory_size;if(fw_dump.dump_active){pr_info("Firmware-assisted dump is active.\n");
@@ -332,13 +337,11 @@ int __init fadump_reserve_mem(void)*dumpiswrittentodiskbyuserspacetool.Thismemory*willbereleasedforgeneraluseoncethedumpissaved.*/-base=fw_dump.boot_memory_size;size=memory_boundary-base;fadump_reserve_crash_area(base,size);pr_debug("fadumphdr_addr = %#016lx\n",fw_dump.fadumphdr_addr);fw_dump.reserve_dump_area_start=base;-fw_dump.reserve_dump_area_size=size;}else{/**ReservememoryatanoffsetclosertobottomoftheRAMto
@@ -346,30 +349,42 @@ int __init fadump_reserve_mem(void)*usememblock_find_in_range()heresinceitdoesn'tallocate*frombottomtotop.*/-for(base=fw_dump.boot_memory_size;-base<=(memory_boundary-size);-base+=size){+while(base<=(memory_boundary-size)){if(memblock_is_region_memory(base,size)&&!memblock_is_region_reserved(base,size))break;++base+=size;}-if((base>(memory_boundary-size))||-memblock_reserve(base,size)){++if(base>(memory_boundary-size)){+pr_err("Failed to find memory chunk for reservation\n");+gotoerror_out;+}+fw_dump.reserve_dump_area_start=base;++/*+*Calculatethekernelmetadataaddressandregisteritwith+*f/wiftheplatformsupports.+*/+if(fw_dump.ops->setup_kernel_metadata(&fw_dump)<0)+gotoerror_out;++if(memblock_reserve(base,size)){pr_err("Failed to reserve memory\n");-return0;+gotoerror_out;}-pr_info("Reserved %ldMB of memory at %ldMB for firmware-"-"assisted dump (System RAM: %ldMB)\n",-(unsignedlong)(size>>20),-(unsignedlong)(base>>20),+pr_info("Reserved %ldMB of memory at %#016lx (System RAM: %ldMB)\n",+(unsignedlong)(size>>20),base,(unsignedlong)(memblock_phys_mem_size()>>20));-fw_dump.reserve_dump_area_start=base;-fw_dump.reserve_dump_area_size=size;-returnfadump_cma_init();+ret=fadump_cma_init();}-return1;+returnret;+error_out:+fw_dump.fadump_enabled=0;+return0;}unsignedlong__initarch_reserved_kernel_pages(void)
@@ -0,0 +1,37 @@+/*+*Firmware-AssistedDumpsupportonPOWERplatform(OPAL).+*+*Copyright2019,IBMCorp.+*Author:HariBathini<hbathini@linux.ibm.com>+*+*Thisprogramisfreesoftware;youcanredistributeitand/or+*modifyitunderthetermsoftheGNUGeneralPublicLicense+*aspublishedbytheFreeSoftwareFoundation;eitherversion+*2oftheLicense,or(atyouroption)anylaterversion.+*/++#ifndef __PPC64_OPAL_FA_DUMP_H__+#define __PPC64_OPAL_FA_DUMP_H__++/* OPAL FADump structure format version */+#define OPAL_FADUMP_VERSION 0x1++/* Maximum number of memory regions kernel supports */+#define OPAL_FADUMP_MAX_MEM_REGS 128++/*+*FADumpmemorystructureforstoringkernelmetadataneededto+*register-for/processcrashdump.Theaddressofthisstructurewill+*beregisteredwithf/wforretrievingduringcrashdump.+*/+structopal_fadump_mem_struct{++u8version;+u8reserved[3];+u16region_cnt;/* number of regions */+u16registered_regions;/* Regions registered for MPIPL */+u64fadumphdr_addr;+structopal_mpipl_regionrgn[OPAL_FADUMP_MAX_MEM_REGS];+}__attribute__((packed));++#endif /* __PPC64_OPAL_FA_DUMP_H__ */
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 11:56:22
Make OPAL calls to register and un-register with firmware for MPIPL.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/platforms/powernv/opal-fadump.c | 71 +++++++++++++++++++++++++-
1 file changed, 69 insertions(+), 2 deletions(-)
@@ -88,12 +104,63 @@ static int opal_fadump_setup_kernel_metadata(struct fw_dump *fadump_conf)staticintopal_fadump_register_fadump(structfw_dump*fadump_conf){-return-EIO;+inti,err=-EIO;+s64rc;++for(i=0;i<opal_fdm->region_cnt;i++){+rc=opal_mpipl_update(OPAL_MPIPL_ADD_RANGE,+opal_fdm->rgn[i].src,+opal_fdm->rgn[i].dest,+opal_fdm->rgn[i].size);+if(rc!=OPAL_SUCCESS)+break;++opal_fdm->registered_regions++;+}++switch(rc){+caseOPAL_SUCCESS:+pr_info("Registration is successful!\n");+fadump_conf->dump_registered=1;+err=0;+break;+caseOPAL_UNSUPPORTED:+pr_err("Support not available.\n");+fadump_conf->fadump_supported=0;+fadump_conf->fadump_enabled=0;+break;+caseOPAL_INTERNAL_ERROR:+pr_err("Failed to register. Hardware Error(%lld).\n",rc);+break;+caseOPAL_PARAMETER:+pr_err("Failed to register. Parameter Error(%lld).\n",rc);+break;+caseOPAL_PERMISSION:+pr_err("Already registered!\n");+fadump_conf->dump_registered=1;+err=-EEXIST;+break;+default:+pr_err("Failed to register. Unknown Error(%lld).\n",rc);+break;+}++returnerr;}staticintopal_fadump_unregister_fadump(structfw_dump*fadump_conf){-return-EIO;+s64rc;++rc=opal_mpipl_update(OPAL_MPIPL_REMOVE_ALL,0,0,0);+if(rc){+pr_err("Failed to un-register - unexpected Error(%lld).\n",rc);+return-EIO;+}++opal_fdm->registered_regions=0;+fadump_conf->dump_registered=0;+return0;}staticintopal_fadump_invalidate_fadump(structfw_dump*fadump_conf)
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 11:58:19
Firmware uses 32-bit field for region size while copying/backing-up
memory during MPIPL. So, the maximum copy size for a region would
be a page less than 4GB (aligned to pagesize) but FADump capture
kernel usually needs more memory than that to be preserved to avoid
running into out of memory errors.
So, request firmware to copy multiple kernel memory regions instead
of just one (which worked fine for pseries as 64-bit field was used
for size there). With support to copy multiple kernel memory regions,
also handle holes in the memory area to be preserved. Support as many
as 128 kernel memory regions. This allows having an adequate FADump
capture kernel size for different scenarios.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/kernel/fadump-common.c | 15 ++
arch/powerpc/kernel/fadump-common.h | 16 ++
arch/powerpc/kernel/fadump.c | 173 ++++++++++++++++++++++----
arch/powerpc/platforms/powernv/opal-fadump.c | 25 +++-
arch/powerpc/platforms/powernv/opal-fadump.h | 5 -
arch/powerpc/platforms/pseries/rtas-fadump.c | 12 ++
arch/powerpc/platforms/pseries/rtas-fadump.h | 5 +
7 files changed, 211 insertions(+), 40 deletions(-)
@@ -125,10 +125,19 @@ static int is_fadump_memory_area_contiguous(unsigned long d_start,*/intis_fadump_boot_mem_contiguous(structfw_dump*fadump_conf){-unsignedlongd_start=RMA_START;-unsignedlongd_end=RMA_START+fadump_conf->boot_memory_size;+inti,ret=0;+unsignedlongd_start,d_end;-returnis_fadump_memory_area_contiguous(d_start,d_end);+for(i=0;i<fadump_conf->boot_mem_regs_cnt;i++){+d_start=fadump_conf->boot_mem_addr[i];+d_end=d_start+fadump_conf->boot_mem_size[i];++ret=is_fadump_memory_area_contiguous(d_start,d_end);+if(!ret)+break;+}++returnret;}/*
@@ -310,6 +401,10 @@ int __init fadump_reserve_mem(void)ALIGN(fw_dump.boot_memory_size,FADUMP_CMA_ALIGNMENT);#endif+if(!fadump_get_boot_mem_regions()){+pr_err("Too many holes in boot memory area to enable fadump\n");+gotoerror_out;+}}size=get_fadump_area_size();
@@ -319,7 +414,8 @@ int __init fadump_reserve_mem(void)elsememory_boundary=memblock_end_of_DRAM();-base=fw_dump.boot_memory_size;+base=fw_dump.boot_mem_top;+base=PAGE_ALIGN(base);if(fw_dump.dump_active){pr_info("Firmware-assisted dump is active.\n");
@@ -16,9 +16,6 @@/* OPAL FADump structure format version */#define OPAL_FADUMP_VERSION 0x1-/* Maximum number of memory regions kernel supports */-#define OPAL_FADUMP_MAX_MEM_REGS 128-/**FADumpmemorystructureforstoringkernelmetadataneededto*register-for/processcrashdump.Theaddressofthisstructurewill
@@ -31,7 +28,7 @@ struct opal_fadump_mem_struct {u16region_cnt;/* number of regions */u16registered_regions;/* Regions registered for MPIPL */u64fadumphdr_addr;-structopal_mpipl_regionrgn[OPAL_FADUMP_MAX_MEM_REGS];+structopal_mpipl_regionrgn[FADUMP_MAX_MEM_REGS];}__attribute__((packed));#endif /* __PPC64_OPAL_FA_DUMP_H__ */
@@ -535,6 +542,9 @@ int __init rtas_fadump_dt_scan(struct fw_dump *fadump_conf, ulong node)fadump_conf->fadump_platform=FADUMP_PLATFORM_PSERIES;fadump_conf->fadump_supported=1;+/* Firmware supports 64-bit value for size, align it to pagesize. */+fadump_conf->max_copy_size=_ALIGN_DOWN(U64_MAX,PAGE_SIZE);+/**The'ibm,kernel-dump'rtasnodeispresentonlyifthereis*dumpdatawaitingforus.
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 12:00:09
Add support in the kernel to process the crash'ed kernel's memory
preserved during MPIPL and export it as /proc/vmcore file for the
userland scripts to filter and analyze it later.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/platforms/powernv/opal-fadump.c | 190 ++++++++++++++++++++++++++
1 file changed, 187 insertions(+), 3 deletions(-)
@@ -41,6 +43,50 @@ static void opal_fadump_update_config(struct fw_dump *fadump_conf,fadump_conf->boot_mem_dest_addr);fadump_conf->fadumphdr_addr=fdm->fadumphdr_addr;++/* Start address of preserve area (permanent reservation) */+fadump_conf->preserv_area_start=fadump_conf->boot_mem_dest_addr;+pr_debug("Preserve area start address: 0x%lx\n",+fadump_conf->preserv_area_start);+}++/*+*Thisfunctioniscalledinthecapturekerneltogetconfigurationdetails+*frommetadatasetupbythefirstkernel.+*/+staticvoidopal_fadump_get_config(structfw_dump*fadump_conf,+conststructopal_fadump_mem_struct*fdm)+{+unsignedlongbase,size,last_end,hole_size;+inti;++if(!fadump_conf->dump_active)+return;++last_end=0;+hole_size=0;+fadump_conf->boot_memory_size=0;++if(fdm->region_cnt)+pr_debug("Boot memory regions:\n");++for(i=0;i<fdm->region_cnt;i++){+base=fdm->rgn[i].src;+size=fdm->rgn[i].size;+pr_debug("\t%d. base: 0x%lx, size: 0x%lx\n",+(i+1),base,size);++fadump_conf->boot_mem_addr[i]=base;+fadump_conf->boot_mem_size[i]=size;+fadump_conf->boot_memory_size+=size;+hole_size+=(base-last_end);++last_end=base+size;+}++fadump_conf->boot_mem_top=(fadump_conf->boot_memory_size+hole_size);+fadump_conf->boot_mem_regs_cnt=fdm->region_cnt;+opal_fadump_update_config(fadump_conf,fdm);}staticulongopal_fadump_init_mem_struct(structfw_dump*fadump_conf)
@@ -174,27 +220,127 @@ static int opal_fadump_unregister_fadump(struct fw_dump *fadump_conf)staticintopal_fadump_invalidate_fadump(structfw_dump*fadump_conf){-return-EIO;+s64rc;++rc=opal_mpipl_update(OPAL_MPIPL_FREE_PRESERVED_MEMORY,0,0,0);+if(rc){+pr_err("Failed to invalidate - unexpected Error(%lld).\n",rc);+return-EIO;+}++fadump_conf->dump_active=0;+opal_fdm_active=NULL;+return0;+}++/*+*ConvertCPUstatedatasavedatthetimeofcrashintoELFnotes.+*/+staticint__initopal_fadump_build_cpu_notes(structfw_dump*fadump_conf)+{+u32num_cpus,*note_buf;+structfadump_crash_info_header*fdh=NULL;++num_cpus=1;+/* Allocate buffer to hold cpu crash notes. */+fadump_conf->cpu_notes_buf_size=num_cpus*sizeof(note_buf_t);+fadump_conf->cpu_notes_buf_size=+PAGE_ALIGN(fadump_conf->cpu_notes_buf_size);+note_buf=fadump_cpu_notes_buf_alloc(fadump_conf->cpu_notes_buf_size);+if(!note_buf){+pr_err("Failed to allocate 0x%lx bytes for cpu notes buffer\n",+fadump_conf->cpu_notes_buf_size);+return-ENOMEM;+}+fadump_conf->cpu_notes_buf=__pa(note_buf);++pr_debug("Allocated buffer for cpu notes of size %ld at %p\n",+(num_cpus*sizeof(note_buf_t)),note_buf);++if(fadump_conf->fadumphdr_addr)+fdh=__va(fadump_conf->fadumphdr_addr);++if(fdh&&(fdh->crashing_cpu!=FADUMP_CPU_UNKNOWN)){+note_buf=fadump_regs_to_elf_notes(note_buf,&(fdh->regs));+final_note(note_buf);++pr_debug("Updating elfcore header (%llx) with cpu notes\n",+fdh->elfcorehdr_addr);+fadump_update_elfcore_header(fadump_conf,+__va(fdh->elfcorehdr_addr));+}++return0;}staticint__initopal_fadump_process_fadump(structfw_dump*fadump_conf){-return-EINVAL;+structfadump_crash_info_header*fdh;+intrc=0;++if(!opal_fdm_active||!fadump_conf->fadumphdr_addr)+return-EINVAL;++/* Validate the fadump crash info header */+fdh=__va(fadump_conf->fadumphdr_addr);+if(fdh->magic_number!=FADUMP_CRASH_INFO_MAGIC){+pr_err("Crash info header is not valid.\n");+return-EINVAL;+}++/*+*TODO:Tobuildcpunotes,findawaytomapPIRtologicalid.+*Also,wemayneeddifferentmethodforpseriesandpowernv.+*ThecurrentlybootedkernelcouldhaveadifferentPIRto+*logicalidmapping.So,trysavinginfoofpreviouskernel's+*pacatogettherightPIRtologicalidmapping.+*/+rc=opal_fadump_build_cpu_notes(fadump_conf);+if(rc)+returnrc;++/*+*Wearedonevalidatingdumpinfoandelfcoreheaderisnowready+*tobeexported.setelfcorehdr_addrsothatvmcoremodulewill+*exporttheelfcoreheaderthrough'/proc/vmcore'.+*/+elfcorehdr_addr=fdh->elfcorehdr_addr;++returnrc;}staticvoidopal_fadump_region_show(structfw_dump*fadump_conf,structseq_file*m){inti;-conststructopal_fadump_mem_struct*fdm_ptr=opal_fdm;+conststructopal_fadump_mem_struct*fdm_ptr;u64dumped_bytes=0;+if(fadump_conf->dump_active)+fdm_ptr=opal_fdm_active;+else+fdm_ptr=opal_fdm;+for(i=0;i<fdm_ptr->region_cnt;i++){+/*+*OnlyregionsthatareregisteredforMPIPL+*wouldhavedumpdata.+*/+if((fadump_conf->dump_active)&&+(i<fdm_ptr->registered_regions))+dumped_bytes=fdm_ptr->rgn[i].size;+seq_printf(m,"DUMP: Src: %#016llx, Dest: %#016llx, ",fdm_ptr->rgn[i].src,fdm_ptr->rgn[i].dest);seq_printf(m,"Size: %#llx, Dumped: %#llx bytes\n",fdm_ptr->rgn[i].size,dumped_bytes);}++/* Dump is active. Show reserved area start address. */+if(fadump_conf->dump_active){+seq_printf(m,"\nMemory above %#016lx is reserved for saving crash dump\n",+fadump_conf->reserve_dump_area_start);+}}staticvoidopal_fadump_trigger(structfadump_crash_info_header*fdh,
@@ -251,5 +398,42 @@ int __init opal_fadump_dt_scan(struct fw_dump *fadump_conf, ulong node)*/fadump_conf->max_copy_size=_ALIGN_DOWN(U32_MAX,PAGE_SIZE);+/*+*Checkifdumphasbeeninitiatedonlastreboot.+*/+prop=of_get_flat_dt_prop(dn,"mpipl-boot",NULL);+if(prop){+u64addr=0;+s64ret;+conststructopal_fadump_mem_struct*r_opal_fdm_active;++ret=opal_mpipl_query_tag(OPAL_MPIPL_TAG_KERNEL,&addr);+if((ret!=OPAL_SUCCESS)||!addr){+pr_err("Failed to get Kernel metadata (%lld)\n",ret);+return1;+}++addr=be64_to_cpu(addr);+pr_debug("Kernel metadata addr: %llx\n",addr);++opal_fdm_active=__va(addr);+r_opal_fdm_active=(void*)addr;+if(r_opal_fdm_active->version!=OPAL_FADUMP_VERSION){+pr_err("FADump active but version (%u) unsupported!\n",+r_opal_fdm_active->version);+return1;+}++/* Kernel regions not registered with f/w for MPIPL */+if(r_opal_fdm_active->registered_regions==0){+opal_fdm_active=NULL;+return1;+}++pr_info("Firmware-assisted dump is active.\n");+fadump_conf->dump_active=1;+opal_fadump_get_config(fadump_conf,r_opal_fdm_active);+}+return1;}
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 12:02:36
With FADump support now available on both pseries and OPAL platforms,
update FADump documentation with these details.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
Documentation/powerpc/firmware-assisted-dump.txt | 104 +++++++++++++---------
1 file changed, 63 insertions(+), 41 deletions(-)
@@ -70,7 +70,8 @@ as follows: normal. -- The freshly booted kernel will notice that there is a new- node (ibm,dump-kernel) in the device tree, indicating that+ node (ibm,dump-kernel on PSeries or ibm,opal/dump/result-table+ on OPAL platform) in the device tree, indicating that there is crash data available from a previous boot. During the early boot OS will reserve rest of the memory above boot memory size effectively booting with restricted memory
@@ -93,7 +94,9 @@ as follows: Please note that the firmware-assisted dump feature is only available on Power6 and above systems with recent-firmware versions.+firmware versions on PSeries (PowerVM) platform and Power9+and above systems with recent firmware versions on PowerNV+(OPAL) platform. Implementation details: ----------------------
@@ -108,57 +111,76 @@ that are run. If there is dump data, then the /sys/kernel/fadump_release_mem file is created, and the reserved memory is held.-If there is no waiting dump data, then only the memory required-to hold CPU state, HPTE region, boot memory dump and elfcore-header, is usually reserved at an offset greater than boot memory-size (see Fig. 1). This area is *not* released: this region will-be kept permanently reserved, so that it can act as a receptacle-for a copy of the boot memory content in addition to CPU state-and HPTE region, in the case a crash does occur. Since this reserved-memory area is used only after the system crash, there is no point in-blocking this significant chunk of memory from production kernel.-Hence, the implementation uses the Linux kernel's Contiguous Memory-Allocator (CMA) for memory reservation if CMA is configured for kernel.-With CMA reservation this memory will be available for applications to-use it, while kernel is prevented from using it. With this FADump will-still be able to capture all of the kernel memory and most of the user-space memory except the user pages that were present in CMA region.+If there is no waiting dump data, then only the memory required to+hold CPU state, HPTE region, boot memory dump, FADump header and+elfcore header, is usually reserved at an offset greater than boot+memory size (see Fig. 1). This area is *not* released: this region+will be kept permanently reserved, so that it can act as a receptacle+for a copy of the boot memory content in addition to CPU state and+HPTE region, in the case a crash does occur.++Since this reserved memory area is used only after the system crash,+there is no point in blocking this significant chunk of memory from+production kernel. Hence, the implementation uses the Linux kernel's+Contiguous Memory Allocator (CMA) for memory reservation if CMA is+configured for kernel. With CMA reservation this memory will be+available for applications to use it, while kernel is prevented from+using it. With this FADump will still be able to capture all of the+kernel memory and most of the user space memory except the user pages+that were present in CMA region. o Memory Reservation during first kernel- Low memory Top of memory- 0 boot memory size |<--Reserved dump area --->| |- | | | Permanent Reservation | |- V V | (Preserve area) | V- +-----------+----------/ /---+---+----+--------+---+----+------+- | | |CPU|HPTE| DUMP |HDR|ELF | |- +-----------+----------/ /---+---+----+--------+---+----+------+- | ^ ^- | | |- \ / |- ----------------------------------- FADump Header- Boot memory content gets transferred (meta area)- to reserved area by firmware at the- time of crash+ Low memory Top of memory+ 0 boot memory size |<--- Reserved dump area --->| |+ | | | Permanent Reservation | |+ V V | (Preserve area) | V+ +-----------+-----/ /---+---+----+-------+-----+-----+----+--++ | | |///|////| DUMP | HDR | ELF |////| |+ +-----------+-----/ /---+---+----+-------+-----+-----+----+--++ | ^ ^ ^ ^ ^+ | | | | | |+ \ CPU HPTE / | |+ ------------------------------ | |+ Boot memory content gets transferred | |+ to reserved area by firmware at the | |+ time of crash. | |+ FADump Header |+ (meta area) |+ |+ |+ Metadata: This area holds a metadata struture whose+ address is registered with f/w and retrieved in the+ second kernel after crash, on platforms that support+ tags (OPAL). Having such structure with info needed+ to process the crashdump eases dump capture process. Fig. 1 o Memory Reservation during second kernel after crash- Low memory Top of memory- 0 boot memory size |- | |<------------- Reserved dump area --------------->|- V V |<---- Preserve area ----->| V- +-----------+----------/ /---+---+----+--------+---+----+------+- | | |CPU|HPTE| DUMP |HDR|ELF | |- +-----------+----------/ /---+---+----+--------+---+----+------+- | |- V V- Used by second /proc/vmcore+ Low memory Top of memory+ 0 boot memory size |+ | |<------------ Reserved dump area -------------->|+ V V |<---- Preserve area ------->| |+ +-----------+-----/ /---+---+----+-------+-----+-----+----+--++ | | |///|////| DUMP | HDR | ELF |////| |+ +-----------+-----/ /---+---+----+-------+-----+-----+----+--++ | |+ V V+ Used by second /proc/vmcore kernel to boot++ +---++ |///| -> Regions (CPU, HPTE & Metadata) marked like this in the above+ +---+ figures are not always present. For example, OPAL platform+ does not have CPU & HPTE regions while Metadata region is+ not supported on pSeries currently.+ Fig. 2+ Currently the dump will be copied from /proc/vmcore to a new file upon user intervention. The dump data available through /proc/vmcore will be in ELF format. Hence the existing kdump infrastructure (kdump scripts)
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 12:04:22
Commit 0962e8004e97 ("powerpc/prom: Scan reserved-ranges node for
memory reservations") enabled support to parse reserved-ranges DT
node and reserve kernel memory falling in these ranges for F/W
purposes. Ensure memory in these ranges is not overlapped with
memory reserved for FADump.
Also, use a smaller offset, instead of the size of the memory to
be reserved, by which to skip memory before making another attempt
at reserving memory, after the previous attempt to reserve memory
for FADump failed due to memory holes and/or reserved ranges, to
reduce the likelihood of memory reservation failure.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/kernel/fadump-common.h | 13 +++
arch/powerpc/kernel/fadump.c | 143 ++++++++++++++++++++++++++++++++++-
2 files changed, 149 insertions(+), 7 deletions(-)
@@ -94,6 +94,17 @@ struct fad_crash_memory_ranges {/* Platform specific callback functions */structfadump_ops;+/*+*Amountofmemory(1024MB)toskipbeforemakinganotherattemptat+*reservingmemory(afterthepreviousattempttoreservememoryfor+*FADumpfailedduetomemoryholesand/orreservedranges)toreduce+*thelikelihoodofmemoryreservationfailure.+*/+#define FADUMP_OFFSET_SIZE 0x40000000U++/* Maximum no. of reserved ranges supported for processing. */+#define FADUMP_MAX_RESERVED_RANGES 128+/* Maximum number of memory regions kernel supports */#define FADUMP_MAX_MEM_REGS 128
@@ -469,218 +451,6 @@ void crash_fadump(struct pt_regs *regs, const char *str)fw_dump.ops->fadump_trigger(fdh,str);}-#define GPR_MASK 0xffffff0000000000-staticinlineintfadump_gpr_index(u64id)-{-inti=-1;-charstr[3];--if((id&GPR_MASK)==fadump_str_to_u64("GPR")){-/* get the digits at the end */-id&=~GPR_MASK;-id>>=24;-str[2]='\0';-str[1]=id&0xff;-str[0]=(id>>8)&0xff;-sscanf(str,"%d",&i);-if(i>31)-i=-1;-}-returni;-}--staticinlinevoidfadump_set_regval(structpt_regs*regs,u64reg_id,-u64reg_val)-{-inti;--i=fadump_gpr_index(reg_id);-if(i>=0)-regs->gpr[i]=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("NIA"))-regs->nip=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("MSR"))-regs->msr=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("CTR"))-regs->ctr=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("LR"))-regs->link=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("XER"))-regs->xer=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("CR"))-regs->ccr=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("DAR"))-regs->dar=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("DSISR"))-regs->dsisr=(unsignedlong)reg_val;-}--staticstructrtas_fadump_reg_entry*-fadump_read_registers(structrtas_fadump_reg_entry*reg_entry,structpt_regs*regs)-{-memset(regs,0,sizeof(structpt_regs));--while(be64_to_cpu(reg_entry->reg_id)!=fadump_str_to_u64("CPUEND")){-fadump_set_regval(regs,be64_to_cpu(reg_entry->reg_id),-be64_to_cpu(reg_entry->reg_value));-reg_entry++;-}-reg_entry++;-returnreg_entry;-}--/*-*ReadCPUstatedumpdataandconvertitintoELFnotes.-*TheCPUdumpstartswithmagicnumber"REGSAVE".NumCpusOffsetshouldbe-*usedtoaccessthedatatoallowforadditionalfieldstobeaddedwithout-*affectingcompatibility.EachlistofregistersforaCPUstartswith-*"CPUSTRT"andendswith"CPUEND".Eachregisterentryisof16bytes,-*8ByteASCIIidentifierand8Byteregistervalue.Theregisterentry-*withidentifier"CPUSTRT"and"CPUEND"contains4bytecpuidaspart-*ofregistervalue.FormoredetailsrefertoPAPRdocument.-*-*OnlyforthecrashingcpuweignoretheCPUdumpdataandgetexact-*statefromfadumpcrashinfostructurepopulatedbyfirstkernelatthe-*timeofcrash.-*/-staticint__initfadump_build_cpu_notes(conststructrtas_fadump_mem_struct*fdm)-{-structrtas_fadump_reg_save_area_header*reg_header;-structrtas_fadump_reg_entry*reg_entry;-structfadump_crash_info_header*fdh=NULL;-void*vaddr;-unsignedlongaddr;-u32num_cpus,*note_buf;-structpt_regsregs;-inti,rc=0,cpu=0;--if(!fdm->cpu_state_data.bytes_dumped)-return-EINVAL;--addr=be64_to_cpu(fdm->cpu_state_data.destination_address);-vaddr=__va(addr);--reg_header=vaddr;-if(be64_to_cpu(reg_header->magic_number)!=-fadump_str_to_u64("REGSAVE")){-printk(KERN_ERR"Unable to read register save area.\n");-return-ENOENT;-}-pr_debug("--------CPU State Data------------\n");-pr_debug("Magic Number: %llx\n",be64_to_cpu(reg_header->magic_number));-pr_debug("NumCpuOffset: %x\n",be32_to_cpu(reg_header->num_cpu_offset));--vaddr+=be32_to_cpu(reg_header->num_cpu_offset);-num_cpus=be32_to_cpu(*((__be32*)(vaddr)));-pr_debug("NumCpus : %u\n",num_cpus);-vaddr+=sizeof(u32);-reg_entry=(structrtas_fadump_reg_entry*)vaddr;--/* Allocate buffer to hold cpu crash notes. */-fw_dump.cpu_notes_buf_size=num_cpus*sizeof(note_buf_t);-fw_dump.cpu_notes_buf_size=PAGE_ALIGN(fw_dump.cpu_notes_buf_size);-note_buf=fadump_cpu_notes_buf_alloc(fw_dump.cpu_notes_buf_size);-if(!note_buf){-printk(KERN_ERR"Failed to allocate 0x%lx bytes for "-"cpu notes buffer\n",fw_dump.cpu_notes_buf_size);-return-ENOMEM;-}-fw_dump.cpu_notes_buf=__pa(note_buf);--pr_debug("Allocated buffer for cpu notes of size %ld at %p\n",-(num_cpus*sizeof(note_buf_t)),note_buf);--if(fw_dump.fadumphdr_addr)-fdh=__va(fw_dump.fadumphdr_addr);--for(i=0;i<num_cpus;i++){-if(be64_to_cpu(reg_entry->reg_id)!=fadump_str_to_u64("CPUSTRT")){-printk(KERN_ERR"Unable to read CPU state data\n");-rc=-ENOENT;-gotoerror_out;-}-/* Lower 4 bytes of reg_value contains logical cpu id */-cpu=be64_to_cpu(reg_entry->reg_value)&RTAS_FADUMP_CPU_ID_MASK;-if(fdh&&!cpumask_test_cpu(cpu,&fdh->online_mask)){-RTAS_FADUMP_SKIP_TO_NEXT_CPU(reg_entry);-continue;-}-pr_debug("Reading register data for cpu %d...\n",cpu);-if(fdh&&fdh->crashing_cpu==cpu){-regs=fdh->regs;-note_buf=fadump_regs_to_elf_notes(note_buf,®s);-RTAS_FADUMP_SKIP_TO_NEXT_CPU(reg_entry);-}else{-reg_entry++;-reg_entry=fadump_read_registers(reg_entry,®s);-note_buf=fadump_regs_to_elf_notes(note_buf,®s);-}-}-final_note(note_buf);--if(fdh){-addr=fdh->elfcorehdr_addr;-pr_debug("Updating elfcore header(%lx) with cpu notes\n",addr);-fadump_update_elfcore_header(&fw_dump,(char*)__va(addr));-}-return0;--error_out:-fadump_cpu_notes_buf_free((unsignedlong)__va(fw_dump.cpu_notes_buf),-fw_dump.cpu_notes_buf_size);-fw_dump.cpu_notes_buf=0;-fw_dump.cpu_notes_buf_size=0;-returnrc;--}--/*-*Validateandprocessthedumpdatastoredbyfirmwarebeforeexporting-*itthrough'/proc/vmcore'.-*/-staticint__initprocess_fadump(conststructrtas_fadump_mem_struct*fdm_active)-{-structfadump_crash_info_header*fdh;-intrc=0;--if(!fdm_active||!fw_dump.fadumphdr_addr)-return-EINVAL;--/* Check if the dump data is valid. */-if((be16_to_cpu(fdm_active->header.dump_status_flag)==RTAS_FADUMP_ERROR_FLAG)||-(fdm_active->cpu_state_data.error_flags!=0)||-(fdm_active->rmr_region.error_flags!=0)){-printk(KERN_ERR"Dump taken by platform is not valid\n");-return-EINVAL;-}-if((fdm_active->rmr_region.bytes_dumped!=-fdm_active->rmr_region.source_len)||-!fdm_active->cpu_state_data.bytes_dumped){-printk(KERN_ERR"Dump taken by platform is incomplete\n");-return-EINVAL;-}--/* Validate the fadump crash info header */-fdh=__va(fw_dump.fadumphdr_addr);-if(fdh->magic_number!=FADUMP_CRASH_INFO_MAGIC){-printk(KERN_ERR"Crash info header is not valid.\n");-return-EINVAL;-}--rc=fadump_build_cpu_notes(fdm_active);-if(rc)-returnrc;--/*-*Wearedonevalidatingdumpinfoandelfcoreheaderisnowready-*tobeexported.setelfcorehdr_addrsothatvmcoremodulewill-*exporttheelfcoreheaderthrough'/proc/vmcore'.-*/-elfcorehdr_addr=fdh->elfcorehdr_addr;--return0;-}-staticvoidfree_crash_memory_ranges(void){kfree(crash_memory_ranges);
@@ -970,7 +740,6 @@ static unsigned long init_fadump_header(unsigned long addr)if(!addr)return0;-fw_dump.fadumphdr_addr=addr;fdh=__va(addr);addr+=sizeof(structfadump_crash_info_header);
@@ -1014,39 +783,12 @@ static int register_fadump(void)returnfw_dump.ops->register_fadump(&fw_dump);}-staticintfadump_invalidate_dump(conststructrtas_fadump_mem_struct*fdm)-{-intrc=0;-unsignedintwait_time;--pr_debug("Invalidating firmware-assisted dump registration\n");--/* TODO: Add upper time limit for the delay */-do{-rc=rtas_call(fw_dump.ibm_configure_kernel_dump,3,1,NULL,-FADUMP_INVALIDATE,fdm,-sizeof(structrtas_fadump_mem_struct));--wait_time=rtas_busy_delay_time(rc);-if(wait_time)-mdelay(wait_time);-}while(wait_time);--if(rc){-pr_err("Failed to invalidate firmware-assisted dump registration. Unexpected error (%d).\n",rc);-returnrc;-}-fw_dump.dump_active=0;-fdm_active=NULL;-return0;-}-voidfadump_cleanup(void){/* Invalidate the registration only if dump is active. */if(fw_dump.dump_active){-/* pass the same memory dump structure provided by platform */-fadump_invalidate_dump(fdm_active);+pr_debug("Invalidating firmware-assisted dump registration\n");+fw_dump.ops->invalidate_fadump(&fw_dump);}elseif(fw_dump.dump_registered){/* Un-register Firmware-assisted dump if it was registered. */fw_dump.ops->unregister_fadump(&fw_dump);
@@ -40,6 +41,23 @@ static void rtas_fadump_update_config(struct fw_dump *fadump_conf,fadump_conf->fadumphdr_addr=(fadump_conf->boot_mem_dest_addr+fadump_conf->boot_memory_size);++/* Start address of preserve area (permanent reservation) */+fadump_conf->preserv_area_start=+be64_to_cpu(fdm->cpu_state_data.destination_address);+pr_debug("Preserve area start address: 0x%lx\n",+fadump_conf->preserv_area_start);+}++/*+*Thisfunctioniscalledinthecapturekerneltogetconfigurationdetails+*setupinthefirstkernelandpassedtothef/w.+*/+staticvoidrtas_fadump_get_config(structfw_dump*fadump_conf,+conststructrtas_fadump_mem_struct*fdm)+{+fadump_conf->boot_memory_size=be64_to_cpu(fdm->rmr_region.source_len);+rtas_fadump_update_config(fadump_conf,fdm);}staticulongrtas_fadump_init_mem_struct(structfw_dump*fadump_conf)
@@ -180,7 +198,196 @@ static int rtas_fadump_unregister_fadump(struct fw_dump *fadump_conf)staticintrtas_fadump_invalidate_fadump(structfw_dump*fadump_conf){-return-EIO;+intrc;+unsignedintwait_time;++/* TODO: Add upper time limit for the delay */+do{+rc=rtas_call(fadump_conf->ibm_configure_kernel_dump,3,1,+NULL,FADUMP_INVALIDATE,fdm_active,+sizeof(structrtas_fadump_mem_struct));++wait_time=rtas_busy_delay_time(rc);+if(wait_time)+mdelay(wait_time);+}while(wait_time);++if(rc){+pr_err("Failed to invalidate - unexpected error (%d).\n",rc);+return-EIO;+}++fadump_conf->dump_active=0;+fdm_active=NULL;+return0;+}++#define RTAS_FADUMP_GPR_MASK 0xffffff0000000000+staticinlineintrtas_fadump_gpr_index(u64id)+{+inti=-1;+charstr[3];++if((id&RTAS_FADUMP_GPR_MASK)==fadump_str_to_u64("GPR")){+/* get the digits at the end */+id&=~RTAS_FADUMP_GPR_MASK;+id>>=24;+str[2]='\0';+str[1]=id&0xff;+str[0]=(id>>8)&0xff;+if(kstrtoint(str,10,&i))+i=-EINVAL;+if(i>31)+i=-1;+}+returni;+}++voidrtas_fadump_set_regval(structpt_regs*regs,u64reg_id,u64reg_val)+{+inti;++i=rtas_fadump_gpr_index(reg_id);+if(i>=0)+regs->gpr[i]=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("NIA"))+regs->nip=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("MSR"))+regs->msr=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("CTR"))+regs->ctr=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("LR"))+regs->link=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("XER"))+regs->xer=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("CR"))+regs->ccr=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("DAR"))+regs->dar=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("DSISR"))+regs->dsisr=(unsignedlong)reg_val;+}++staticstructrtas_fadump_reg_entry*+rtas_fadump_read_regs(structrtas_fadump_reg_entry*reg_entry,+structpt_regs*regs)+{+memset(regs,0,sizeof(structpt_regs));++while(be64_to_cpu(reg_entry->reg_id)!=fadump_str_to_u64("CPUEND")){+rtas_fadump_set_regval(regs,be64_to_cpu(reg_entry->reg_id),+be64_to_cpu(reg_entry->reg_value));+reg_entry++;+}+reg_entry++;+returnreg_entry;+}++/*+*ReadCPUstatedumpdataandconvertitintoELFnotes.+*TheCPUdumpstartswithmagicnumber"REGSAVE".NumCpusOffsetshouldbe+*usedtoaccessthedatatoallowforadditionalfieldstobeaddedwithout+*affectingcompatibility.EachlistofregistersforaCPUstartswith+*"CPUSTRT"andendswith"CPUEND".Eachregisterentryisof16bytes,+*8ByteASCIIidentifierand8Byteregistervalue.Theregisterentry+*withidentifier"CPUSTRT"and"CPUEND"contains4bytecpuidaspart+*ofregistervalue.FormoredetailsrefertoPAPRdocument.+*+*OnlyforthecrashingcpuweignoretheCPUdumpdataandgetexact+*statefromfadumpcrashinfostructurepopulatedbyfirstkernelatthe+*timeofcrash.+*/+staticint__initrtas_fadump_build_cpu_notes(structfw_dump*fadump_conf)+{+structrtas_fadump_reg_save_area_header*reg_header;+structrtas_fadump_reg_entry*reg_entry;+structfadump_crash_info_header*fdh=NULL;+void*vaddr;+unsignedlongaddr;+u32num_cpus,*note_buf;+structpt_regsregs;+inti,rc=0,cpu=0;++addr=be64_to_cpu(fdm_active->cpu_state_data.destination_address);+vaddr=__va(addr);++reg_header=vaddr;+if(be64_to_cpu(reg_header->magic_number)!=+fadump_str_to_u64("REGSAVE")){+pr_err("Unable to read register save area.\n");+return-ENOENT;+}++pr_debug("--------CPU State Data------------\n");+pr_debug("Magic Number: %llx\n",be64_to_cpu(reg_header->magic_number));+pr_debug("NumCpuOffset: %x\n",be32_to_cpu(reg_header->num_cpu_offset));++vaddr+=be32_to_cpu(reg_header->num_cpu_offset);+num_cpus=be32_to_cpu(*((__be32*)(vaddr)));+pr_debug("NumCpus : %u\n",num_cpus);+vaddr+=sizeof(u32);+reg_entry=(structrtas_fadump_reg_entry*)vaddr;++/* Allocate buffer to hold cpu crash notes. */+fadump_conf->cpu_notes_buf_size=num_cpus*sizeof(note_buf_t);+fadump_conf->cpu_notes_buf_size=+PAGE_ALIGN(fadump_conf->cpu_notes_buf_size);+note_buf=fadump_cpu_notes_buf_alloc(fadump_conf->cpu_notes_buf_size);+if(!note_buf){+pr_err("Failed to allocate 0x%lx bytes for cpu notes buffer\n",+fadump_conf->cpu_notes_buf_size);+return-ENOMEM;+}+fadump_conf->cpu_notes_buf=__pa(note_buf);++pr_debug("Allocated buffer for cpu notes of size %ld at %p\n",+(num_cpus*sizeof(note_buf_t)),note_buf);++if(fadump_conf->fadumphdr_addr)+fdh=__va(fadump_conf->fadumphdr_addr);++for(i=0;i<num_cpus;i++){+if(be64_to_cpu(reg_entry->reg_id)!=+fadump_str_to_u64("CPUSTRT")){+pr_err("Unable to read CPU state data\n");+rc=-ENOENT;+gotoerror_out;+}+/* Lower 4 bytes of reg_value contains logical cpu id */+cpu=(be64_to_cpu(reg_entry->reg_value)&+RTAS_FADUMP_CPU_ID_MASK);+if(fdh&&!cpumask_test_cpu(cpu,&fdh->online_mask)){+RTAS_FADUMP_SKIP_TO_NEXT_CPU(reg_entry);+continue;+}+pr_debug("Reading register data for cpu %d...\n",cpu);+if(fdh&&fdh->crashing_cpu==cpu){+regs=fdh->regs;+note_buf=fadump_regs_to_elf_notes(note_buf,®s);+RTAS_FADUMP_SKIP_TO_NEXT_CPU(reg_entry);+}else{+reg_entry++;+reg_entry=rtas_fadump_read_regs(reg_entry,®s);+note_buf=fadump_regs_to_elf_notes(note_buf,®s);+}+}+final_note(note_buf);++if(fdh){+pr_debug("Updating elfcore header (%llx) with cpu notes\n",+fdh->elfcorehdr_addr);+fadump_update_elfcore_header(fadump_conf,+__va(fdh->elfcorehdr_addr));+}+return0;++error_out:+fadump_cpu_notes_buf_free((ulong)__va(fadump_conf->cpu_notes_buf),+fadump_conf->cpu_notes_buf_size);+fadump_conf->cpu_notes_buf=0;+fadump_conf->cpu_notes_buf_size=0;+returnrc;+}/*
@@ -189,15 +396,62 @@ static int rtas_fadump_invalidate_fadump(struct fw_dump *fadump_conf)*/staticint__initrtas_fadump_process_fadump(structfw_dump*fadump_conf){-return-EINVAL;+structfadump_crash_info_header*fdh;+intrc=0;++if(!fdm_active||!fadump_conf->fadumphdr_addr)+return-EINVAL;++/* Check if the dump data is valid. */+if((be16_to_cpu(fdm_active->header.dump_status_flag)==+RTAS_FADUMP_ERROR_FLAG)||+(fdm_active->cpu_state_data.error_flags!=0)||+(fdm_active->rmr_region.error_flags!=0)){+pr_err("Dump taken by platform is not valid\n");+return-EINVAL;+}+if((fdm_active->rmr_region.bytes_dumped!=+fdm_active->rmr_region.source_len)||+!fdm_active->cpu_state_data.bytes_dumped){+pr_err("Dump taken by platform is incomplete\n");+return-EINVAL;+}++/* Validate the fadump crash info header */+fdh=__va(fadump_conf->fadumphdr_addr);+if(fdh->magic_number!=FADUMP_CRASH_INFO_MAGIC){+pr_err("Crash info header is not valid.\n");+return-EINVAL;+}++if(!fdm_active->cpu_state_data.bytes_dumped)+return-EINVAL;++rc=rtas_fadump_build_cpu_notes(fadump_conf);+if(rc)+returnrc;++/*+*Wearedonevalidatingdumpinfoandelfcoreheaderisnowready+*tobeexported.setelfcorehdr_addrsothatvmcoremodulewill+*exporttheelfcoreheaderthrough'/proc/vmcore'.+*/+elfcorehdr_addr=fdh->elfcorehdr_addr;++return0;}staticvoidrtas_fadump_region_show(structfw_dump*fadump_conf,structseq_file*m){-conststructrtas_fadump_mem_struct*fdm_ptr=&fdm;+conststructrtas_fadump_mem_struct*fdm_ptr;conststructrtas_fadump_section*cpu_data_section;+if(fdm_active)+fdm_ptr=fdm_active;+else+fdm_ptr=&fdm;+cpu_data_section=&(fdm_ptr->cpu_state_data);seq_printf(m,"CPU :[%#016llx-%#016llx] %#llx bytes, Dumped: %#llx\n",be64_to_cpu(cpu_data_section->destination_address),
@@ -219,6 +473,12 @@ static void rtas_fadump_region_show(struct fw_dump *fadump_conf,seq_printf(m,"Size: %#llx, Dumped: %#llx bytes\n",be64_to_cpu(fdm_ptr->rmr_region.source_len),be64_to_cpu(fdm_ptr->rmr_region.bytes_dumped));++/* Dump is active. Show reserved area start address. */+if(fdm_active){+seq_printf(m,"\nMemory above %#016lx is reserved for saving crash dump\n",+fadump_conf->reserve_dump_area_start);+}}staticvoidrtas_fadump_trigger(structfadump_crash_info_header*fdh,
@@ -258,6 +519,17 @@ int __init rtas_fadump_dt_scan(struct fw_dump *fadump_conf, ulong node)fadump_conf->fadump_platform=FADUMP_PLATFORM_PSERIES;fadump_conf->fadump_supported=1;+/*+*The'ibm,kernel-dump'rtasnodeispresentonlyifthereis+*dumpdatawaitingforus.+*/+fdm_active=of_get_flat_dt_prop(node,"ibm,kernel-dump",NULL);+if(fdm_active){+pr_info("Firmware-assisted dump is active.\n");+fadump_conf->dump_active=1;+rtas_fadump_get_config(fadump_conf,(void*)__pa(fdm_active));+}+/* Get the sizes required to store dump data for the firmware provided*dumpsections.*Foreachdumpsectiontypesupported,a32bitcellwhichdefines
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 12:10:01
Commit 0962e8004e97 ("powerpc/prom: Scan reserved-ranges node for
memory reservations") enabled support to parse 'reserved-ranges' DT
node to reserve kernel memory falling in these ranges for firmware
purposes. Along with the preserved area memory, also ensure memory
in reserved ranges is not overlapped with memory released by capture
kernel aftering saving vmcore. Also, fix the off-by-one error in
fadump_release_reserved_area function while releasing memory.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/kernel/fadump.c | 61 +++++++++++++++++++++++++++++-------------
1 file changed, 42 insertions(+), 19 deletions(-)
@@ -876,7 +875,7 @@ static int fadump_setup_crash_memory_ranges(void)continue;}-/* add this range excluding the reserved dump area. */+/* add this range excluding the preserve area. */ret=fadump_exclude_reserved_area(start,end);if(ret)returnret;
@@ -1106,33 +1105,57 @@ static void fadump_release_reserved_area(unsigned long start, unsigned long end)if(tend==end_pfn)break;-start_pfn=tend+1;+start_pfn=tend;}}}/*-*Releasethememorythatwasreservedinearlyboottopreservethememory-*contents.Thereleasedmemorywillbeavailableforgeneraluse.+*Releasethememorythatwasreservedduringearlyboottopreservethe+*crash'edkernel'smemorycontentsexceptpreservearea(permanent+*reservation)andreservedrangesusedbyF/W.Thereleasedmemorywill+*beavailableforgeneraluse.*/staticvoidfadump_release_memory(unsignedlongbegin,unsignedlongend){+inti;unsignedlongra_start,ra_end;--ra_start=fw_dump.reserve_dump_area_start;-ra_end=ra_start+fw_dump.reserve_dump_area_size;+unsignedlongtstart;/*-*excludethedumpreservearea.Willreuseitfornext-*fadumpregistration.+*Addmemorytopermanentlypreservetoreservedrangeslist+*andexcludealltheserangeswhilereleasingmemory.*/-if(begin<ra_end&&end>ra_start){-if(begin<ra_start)-fadump_release_reserved_area(begin,ra_start);-if(end>ra_end)-fadump_release_reserved_area(ra_end,end);-}else-fadump_release_reserved_area(begin,end);+i=add_reserved_range(fw_dump.reserve_dump_area_start,+fw_dump.reserve_dump_area_size);+if(i==0){+/*+*ReachedtheMAXreservedrangescount.Toensurereserved+*dumpareaisexcluded(asitwillbereusedfornext+*FADumpregistration),ignorethelastreservedrangeand+*addreserveddumpareainstead.+*/+reserved_ranges_cnt--;+add_reserved_range(fw_dump.reserve_dump_area_start,+fw_dump.reserve_dump_area_size);+}+sort_and_merge_reserved_ranges();++tstart=begin;+for(i=0;i<reserved_ranges_cnt;i++){+ra_start=reserved_ranges[i].base;+ra_end=ra_start+reserved_ranges[i].size;++if(tstart>=ra_end)+continue;++if(tstart<ra_start)+fadump_release_reserved_area(tstart,ra_start);+tstart=ra_end;+}++if(tstart<end)+fadump_release_reserved_area(tstart,end);}staticvoidfadump_invalidate_release_mem(void)
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 12:13:04
From: Hari Bathini <redacted>
Firmware provides architected register state data at the time of crash.
Process this data and build CPU notes to append to ELF core.
Signed-off-by: Hari Bathini <redacted>
Signed-off-by: Vasant Hegde <redacted>
---
arch/powerpc/kernel/fadump-common.h | 4 +
arch/powerpc/platforms/powernv/opal-fadump.c | 197 ++++++++++++++++++++++++--
arch/powerpc/platforms/powernv/opal-fadump.h | 39 +++++
3 files changed, 228 insertions(+), 12 deletions(-)
@@ -233,15 +234,115 @@ static int opal_fadump_invalidate_fadump(struct fw_dump *fadump_conf)return0;}+staticinlinevoidopal_fadump_set_regval_regnum(structpt_regs*regs,+u32reg_type,u32reg_num,+u64reg_val)+{+if(reg_type==HDAT_FADUMP_REG_TYPE_GPR){+if(reg_num<32)+regs->gpr[reg_num]=reg_val;+return;+}++switch(reg_num){+caseSPRN_CTR:+regs->ctr=reg_val;+break;+caseSPRN_LR:+regs->link=reg_val;+break;+caseSPRN_XER:+regs->xer=reg_val;+break;+caseSPRN_DAR:+regs->dar=reg_val;+break;+caseSPRN_DSISR:+regs->dsisr=reg_val;+break;+caseHDAT_FADUMP_REG_ID_NIP:+regs->nip=reg_val;+break;+caseHDAT_FADUMP_REG_ID_MSR:+regs->msr=reg_val;+break;+caseHDAT_FADUMP_REG_ID_CCR:+regs->ccr=reg_val;+break;+}+}++staticinlinevoidopal_fadump_read_regs(char*bufp,unsignedintregs_cnt,+unsignedintreg_entry_size,+structpt_regs*regs)+{+inti;+structhdat_fadump_reg_entry*reg_entry;++memset(regs,0,sizeof(structpt_regs));++for(i=0;i<regs_cnt;i++,bufp+=reg_entry_size){+reg_entry=(structhdat_fadump_reg_entry*)bufp;+opal_fadump_set_regval_regnum(regs,+be32_to_cpu(reg_entry->reg_type),+be32_to_cpu(reg_entry->reg_num),+be64_to_cpu(reg_entry->reg_val));+}+}++staticinlinebool__initis_thread_core_inactive(u8core_state)+{+boolis_inactive=false;++if(core_state==HDAT_FADUMP_CORE_INACTIVE)+is_inactive=true;++returnis_inactive;+}+/**ConvertCPUstatedatasavedatthetimeofcrashintoELFnotes.+*+*Eachregisterentryisof16bytes,Anumericalidentifieralongwith+*aGPR/SPRflaginthefirst8bytesandtheregistervalueinthenext+*8bytes.FormoredetailsrefertoF/Wdocumentation.*/staticint__initopal_fadump_build_cpu_notes(structfw_dump*fadump_conf){u32num_cpus,*note_buf;structfadump_crash_info_header*fdh=NULL;+structhdat_fadump_thread_hdr*thdr;+unsignedlongaddr;+u32thread_pir;+char*bufp;+structpt_regsregs;+unsignedintsize_of_each_thread;+unsignedintregs_offset,regs_cnt,reg_esize;+inti;++if((fadump_conf->cpu_state_destination_addr==0)||+(fadump_conf->cpu_state_entry_size==0)){+pr_err("CPU state data not available for processing!\n");+return-ENODEV;+}++size_of_each_thread=fadump_conf->cpu_state_entry_size;+num_cpus=(fadump_conf->cpu_state_data_size/size_of_each_thread);++addr=fadump_conf->cpu_state_destination_addr;+bufp=__va(addr);++/*+*Offsetforregisterentries,entrysizeandregisterscountis+*duplicatedineverythreadheaderinkeepingwithHDATformat.+*Usethesevaluesfromthefirstthreadheader.+*/+thdr=(structhdat_fadump_thread_hdr*)bufp;+regs_offset=(offsetof(structhdat_fadump_thread_hdr,offset)++be32_to_cpu(thdr->offset));+reg_esize=be32_to_cpu(thdr->esize);+regs_cnt=be32_to_cpu(thdr->ecnt);-num_cpus=1;/* Allocate buffer to hold cpu crash notes. */fadump_conf->cpu_notes_buf_size=num_cpus*sizeof(note_buf_t);fadump_conf->cpu_notes_buf_size=
@@ -260,10 +361,53 @@ static int __init opal_fadump_build_cpu_notes(struct fw_dump *fadump_conf)if(fadump_conf->fadumphdr_addr)fdh=__va(fadump_conf->fadumphdr_addr);-if(fdh&&(fdh->crashing_cpu!=FADUMP_CPU_UNKNOWN)){-note_buf=fadump_regs_to_elf_notes(note_buf,&(fdh->regs));-final_note(note_buf);+pr_debug("--------CPU State Data------------\n");+pr_debug("NumCpus : %u\n",num_cpus);+pr_debug("\tOffset: %u, Entry size: %u, Cnt: %u\n",+regs_offset,reg_esize,regs_cnt);++for(i=0;i<num_cpus;i++,bufp+=size_of_each_thread){+thdr=(structhdat_fadump_thread_hdr*)bufp;++thread_pir=be32_to_cpu(thdr->pir);+pr_debug("%04d) PIR: 0x%x, core state: 0x%02x\n",+(i+1),thread_pir,thdr->core_state);++/*+*RegisterstatedataofMAXcoresisprovidedbyfirmware,+*butsomeofthiscoresmaynotbeactive.So,while+*processingregisterstatedata,checkcorestateand+*skipthreadsthatbelongtoinactivecores.+*/+if(is_thread_core_inactive(thdr->core_state))+continue;++/*+*Ifthisiskernelinitiatedcrash,crashing_cpuwouldbeset+*appropriatelyandregisterdataofthecrashingCPUsavedby+*crashingkernel.AddthissavedregisterdataofcrashingCPU+*toelfnotesandpopulatethept_regsfortheremainingCPUs+*fromregisterstatedataprovidedbyfirmware.+*/+if(fdh&&(fdh->crashing_cpu==thread_pir)){+note_buf=fadump_regs_to_elf_notes(note_buf,+&fdh->regs);+pr_debug("Crashing CPU PIR: 0x%x - R1 : 0x%lx, NIP : 0x%lx\n",+fdh->crashing_cpu,fdh->regs.gpr[1],+fdh->regs.nip);+continue;+}++opal_fadump_read_regs((bufp+regs_offset),regs_cnt,+reg_esize,®s);+note_buf=fadump_regs_to_elf_notes(note_buf,®s);+pr_debug("CPU PIR: 0x%x - R1 : 0x%lx, NIP : 0x%lx\n",+thread_pir,regs.gpr[1],regs.nip);+}+final_note(note_buf);++if(fdh){pr_debug("Updating elfcore header (%llx) with cpu notes\n",fdh->elfcorehdr_addr);fadump_update_elfcore_header(fadump_conf,
@@ -278,7 +422,8 @@ static int __init opal_fadump_process_fadump(struct fw_dump *fadump_conf)structfadump_crash_info_header*fdh;intrc=0;-if(!opal_fdm_active||!fadump_conf->fadumphdr_addr)+if(!opal_fdm_active||!opal_cpu_metadata||+!fadump_conf->fadumphdr_addr)return-EINVAL;/* Validate the fadump crash info header */
@@ -288,13 +433,6 @@ static int __init opal_fadump_process_fadump(struct fw_dump *fadump_conf)return-EINVAL;}-/*-*TODO:Tobuildcpunotes,findawaytomapPIRtologicalid.-*Also,wemayneeddifferentmethodforpseriesandpowernv.-*ThecurrentlybootedkernelcouldhaveadifferentPIRto-*logicalidmapping.So,trysavinginfoofpreviouskernel's-*pacatogettherightPIRtologicalidmapping.-*/rc=opal_fadump_build_cpu_notes(fadump_conf);if(rc)returnrc;
@@ -348,6 +486,14 @@ static void opal_fadump_trigger(struct fadump_crash_info_header *fdh,{intrc;+/*+*UnlikeonpSeriesplatform,logicalCPUnumberisnotprovided+*witharchitectedregisterstatedata.So,storethecrashing+*CPU'sPIRinsteadtoplugtheappropriateregisterdatafor+*crashingCPUinthevmcorefile.+*/+fdh->crashing_cpu=(u32)mfspr(SPRN_PIR);+rc=opal_cec_reboot2(OPAL_REBOOT_MPIPL,msg);if(rc==OPAL_UNSUPPORTED){pr_emerg("Reboot type %d not supported.\n",
@@ -430,6 +577,32 @@ int __init opal_fadump_dt_scan(struct fw_dump *fadump_conf, ulong node)return1;}+ret=opal_mpipl_query_tag(OPAL_MPIPL_TAG_CPU,&addr);+if((ret!=OPAL_SUCCESS)||!addr){+pr_err("Failed to get CPU metadata (%lld)\n",ret);+return1;+}++addr=be64_to_cpu(addr);+pr_debug("CPU metadata addr: %llx\n",addr);++opal_cpu_metadata=__va(addr);+r_opal_cpu_metadata=(void*)addr;+fadump_conf->cpu_state_data_version=+be32_to_cpu(r_opal_cpu_metadata->cpu_data_version);+if(fadump_conf->cpu_state_data_version!=+HDAT_FADUMP_CPU_DATA_VERSION){+pr_err("CPU data format version (%lu) mismatch!\n",+fadump_conf->cpu_state_data_version);+return1;+}+fadump_conf->cpu_state_entry_size=+be32_to_cpu(r_opal_cpu_metadata->cpu_data_size);+fadump_conf->cpu_state_destination_addr=+be64_to_cpu(r_opal_cpu_metadata->region[0].dest);+fadump_conf->cpu_state_data_size=+be64_to_cpu(r_opal_cpu_metadata->region[0].size);+pr_info("Firmware-assisted dump is active.\n");fadump_conf->dump_active=1;opal_fadump_get_config(fadump_conf,r_opal_fdm_active);
@@ -31,4 +31,43 @@ struct opal_fadump_mem_struct {structopal_mpipl_regionrgn[FADUMP_MAX_MEM_REGS];}__attribute__((packed));+/*+*CPUstatedataisprovidedbyf/w.Belowarethedefinitions+*providedinHDATspec.RefertolatestHDATspecificationfor+*anyupdatetothisformat.+*/++#define HDAT_FADUMP_CPU_DATA_VERSION 1++#define HDAT_FADUMP_CORE_INACTIVE (0x0F)++/* HDAT thread header for register entries */+structhdat_fadump_thread_hdr{+__be32pir;+/* 0x00 - 0x0F - The corresponding stop state of the core */+u8core_state;+u8reserved[3];++__be32offset;/* Offset to Register Entries array */+__be32ecnt;/* Number of entries */+__be32esize;/* Alloc size of each array entry in bytes */+__be32eactsz;/* Actual size of each array entry in bytes */+}__attribute__((packed));++/* Register types populated by f/w */+#define HDAT_FADUMP_REG_TYPE_GPR 0x01+#define HDAT_FADUMP_REG_TYPE_SPR 0x02++/* ID numbers used by f/w while populating certain registers */+#define HDAT_FADUMP_REG_ID_NIP 0x7D0+#define HDAT_FADUMP_REG_ID_MSR 0x7D1+#define HDAT_FADUMP_REG_ID_CCR 0x7D2++/* HDAT register entry. */+structhdat_fadump_reg_entry{+__be32reg_type;+__be32reg_num;+__be64reg_val;+}__attribute__((packed));+#endif /* __PPC64_OPAL_FA_DUMP_H__ */
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 12:15:17
Add a new kernel config option, CONFIG_PRESERVE_FA_DUMP that ensures
that crash data, from previously crash'ed kernel, is preserved. This
helps in cases where FADump is not enabled but the subsequent memory
preserving kernel boot is likely to process this crash data. One
typical usecase for this config option is petitboot kernel.
As OPAL allows registering address with it in the first kernel and
retrieving it after MPIPL, use it to store the top of boot memory.
A kernel that intends to preserve crash data retrieves it and avoids
using memory beyond this address.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/Kconfig | 9 ++
arch/powerpc/include/asm/fadump.h | 9 +-
arch/powerpc/kernel/Makefile | 6 +
arch/powerpc/kernel/fadump-common.h | 13 ++-
arch/powerpc/kernel/fadump.c | 128 ++++++++++++++++----------
arch/powerpc/kernel/prom.c | 4 -
arch/powerpc/platforms/powernv/Makefile | 1
arch/powerpc/platforms/powernv/opal-fadump.c | 59 ++++++++++++
arch/powerpc/platforms/powernv/opal-fadump.h | 3 +
9 files changed, 176 insertions(+), 56 deletions(-)
@@ -206,26 +209,6 @@ static void __init early_init_dt_scan_reserved_ranges(unsigned long node)}}-/* Scan the Firmware Assisted dump configuration details. */-int__initearly_init_dt_scan_fw_dump(unsignedlongnode,constchar*uname,-intdepth,void*data)-{-if(depth!=1){-if(depth==0)-early_init_dt_scan_reserved_ranges(node);--return0;-}--if(strcmp(uname,"rtas")==0)-returnrtas_fadump_dt_scan(&fw_dump,node);--if(strcmp(uname,"ibm,opal")==0)-returnopal_fadump_dt_scan(&fw_dump,node);--return0;-}-/**Iffadumpisregistered,checkifthememoryprovided*fallswithinbootmemoryareaandreservedmemoryarea.
@@ -481,26 +464,6 @@ static bool overlaps_with_reserved_ranges(ulong base, ulong end)returnret;}-staticvoid__initfadump_reserve_crash_area(unsignedlongbase,-unsignedlongsize)-{-structmemblock_region*reg;-unsignedlongmstart,mend,msize;--for_each_memblock(memory,reg){-mstart=max_t(unsignedlong,base,reg->base);-mend=reg->base+reg->size;-mend=min(base+size,mend);--if(mstart<mend){-msize=mend-mstart;-memblock_reserve(mstart,msize);-pr_info("Reserved %ldMB of memory at %#016lx for saving crash dump\n",-(msize>>20),mstart);-}-}-}-int__initfadump_reserve_mem(void){intret=1;
@@ -558,12 +521,11 @@ int __init fadump_reserve_mem(void)#endif/**Iflastboothascrashedthenreserveallthememory-*aboveboot_memory_sizesothatwedon'ttouchituntil+*abovebootmemorysizesothatwedon'ttouchituntil*dumpiswrittentodiskbyuserspacetool.Thismemory-*willbereleasedforgeneraluseoncethedumpissaved.+*canbereleasedforgeneralusebyinvalidatingfadump.*/-size=memory_boundary-base;-fadump_reserve_crash_area(base,size);+fadump_reserve_crash_area(base);pr_debug("fadumphdr_addr = %#016lx\n",fw_dump.fadumphdr_addr);fw_dump.reserve_dump_area_start=base;
@@ -613,11 +575,6 @@ int __init fadump_reserve_mem(void)return0;}-unsignedlong__initarch_reserved_kernel_pages(void)-{-returnmemblock_reserved_size()/PAGE_SIZE;-}-/* Look for fadump= cmdline option. */staticint__initearly_fadump_param(char*p){
@@ -1375,3 +1332,76 @@ int __init setup_fadump(void)return1;}subsys_initcall(setup_fadump);+#else /* !CONFIG_PRESERVE_FA_DUMP */++staticinlinevoidearly_init_dt_scan_reserved_ranges(unsignedlongnode){}++/*+*WhendumpisactivebutPRESERVE_FA_DUMPisenabledonthekernel,+*preservecrashdata.Thesubsequentmemorypreservingkernelboot+*islikelytoprocessthiscrashdata.+*/+int__initfadump_reserve_mem(void)+{+if(fw_dump.dump_active){+/*+*Iflastboothascrashedthenreserveallthememory+*abovebootmemorytopreservecrashdata.+*/+pr_info("Preserving crash data for processing in next boot.\n");+fadump_reserve_crash_area(PAGE_ALIGN(fw_dump.boot_mem_top));+}else+pr_debug("FADump-aware kernel..\n");++return1;+}+#endif /* CONFIG_PRESERVE_FA_DUMP */++/* Preserve everything above the base address */+staticvoid__initfadump_reserve_crash_area(unsignedlongbase)+{+structmemblock_region*reg;+unsignedlongmstart,msize;++for_each_memblock(memory,reg){+mstart=reg->base;+msize=reg->size;++if((mstart+msize)<base)+continue;++if(mstart<base){+msize-=(base-mstart);+mstart=base;+}++pr_info("Reserving %luMB of memory at %#016lx for preserving crash data",+(msize>>20),mstart);+memblock_reserve(mstart,msize);+}+}++unsignedlong__initarch_reserved_kernel_pages(void)+{+returnmemblock_reserved_size()/PAGE_SIZE;+}++/* Scan the Firmware Assisted dump configuration details. */+int__initearly_init_dt_scan_fw_dump(unsignedlongnode,constchar*uname,+intdepth,void*data)+{+if(depth!=1){+if(depth==0)+early_init_dt_scan_reserved_ranges(node);++return0;+}++if(strcmp(uname,"rtas")==0)+returnrtas_fadump_dt_scan(&fw_dump,node);++if(strcmp(uname,"ibm,opal")==0)+returnopal_fadump_dt_scan(&fw_dump,node);++return0;+}
@@ -704,7 +704,7 @@ void __init early_init_devtree(void *params)of_scan_flat_dt(early_init_dt_scan_opal,NULL);#endif-#ifdef CONFIG_FA_DUMP+#if defined(CONFIG_FA_DUMP) || defined(CONFIG_PRESERVE_FA_DUMP)/* scan tree to see if dump is active during last boot */of_scan_flat_dt(early_init_dt_scan_fw_dump,NULL);#endif
@@ -26,6 +26,53 @@#include"../../kernel/fadump-common.h"#include"opal-fadump.h"++#ifdef CONFIG_PRESERVE_FA_DUMP+/*+*WhendumpisactivebutPRESERVE_FA_DUMPisenabledonthekernel,+*ensurecrashdataispreservedinhopethatthesubsequentmemory+*preservingkernelbootisgoingtoprocessthiscrashdata.+*/+int__initopal_fadump_dt_scan(structfw_dump*fadump_conf,ulongnode)+{+unsignedlongdn;+const__be32*prop;++dn=of_get_flat_dt_subnode_by_name(node,"dump");+if(dn==-FDT_ERR_NOTFOUND)+return1;++/*+*Checkifdumphasbeeninitiatedonlastreboot.+*/+prop=of_get_flat_dt_prop(dn,"mpipl-boot",NULL);+if(prop){+u64addr=0;+s64ret;++ret=opal_mpipl_query_tag(OPAL_MPIPL_TAG_BOOT_MEM,&addr);+if((ret!=OPAL_SUCCESS)||!addr){+pr_err("Failed to get boot memory tag (%lld)\n",ret);+return1;+}++/*+*Anythingbelowthisaddresscanbeusedforbootinga+*capturekernelorpetitbootkernel.Preserveeverything+*abovethisaddressforprocessingcrashdump.+*/+fadump_conf->boot_mem_top=be64_to_cpu(addr);+pr_debug("Preserve everything above %lx\n",+fadump_conf->boot_mem_top);++pr_info("Firmware-assisted dump is active.\n");+fadump_conf->dump_active=1;+}++return1;+}++#else /* CONFIG_PRESERVE_FA_DUMP */staticconststructopal_fadump_mem_struct*opal_fdm_active;staticconststructopal_mpipl_fadump*opal_cpu_metadata;staticstructopal_fadump_mem_struct*opal_fdm;
@@ -155,6 +202,17 @@ static int opal_fadump_setup_kernel_metadata(struct fw_dump *fadump_conf)err=-EPERM;}+/*+*Registerbootmemorytopaddresswithf/w.Shouldberetrieved+*byakernelthatintendstopreservecrash'edkernel'smemory.+*/+ret=opal_mpipl_register_tag(OPAL_MPIPL_TAG_BOOT_MEM,+fadump_conf->boot_mem_top);+if(ret!=OPAL_SUCCESS){+pr_err("Failed to set boot memory tag!\n");+err=-EPERM;+}+returnerr;}
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 12:18:24
Kernel config option CONFIG_PRESERVE_FA_DUMP is introduced to ensure
crash data, from previously crash'ed kernel, is preserved. Update
documentation with this details.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
Documentation/powerpc/firmware-assisted-dump.txt | 9 +++++++++
1 file changed, 9 insertions(+)
@@ -98,6 +98,15 @@ firmware versions on PSeries (PowerVM) platform and Power9 and above systems with recent firmware versions on PowerNV (OPAL) platform.+On OPAL based machines, system first boots into an intermittent+kernel (referred to as petitboot kernel) before booting into the+capture kernel. This kernel would have minimal kernel and/or+userspace support to process crash data. Such kernel needs to+preserve previously crash'ed kernel's memory for the subsequent+capture kernel boot to process this crash data. Kernel config+option CONFIG_PRESERVE_FA_DUMP has to be enabled on such kernel+to ensure that crash data is preserved to process later.+ Implementation details: ----------------------
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 12:20:28
From: Hari Bathini <redacted>
Export /sys/firmware/opal/core file to analyze opal crashes. Since OPAL
core can be generated independent of CONFIG_FA_DUMP support in kernel,
add this support under a new kernel config option CONFIG_OPAL_CORE.
Also, avoid code duplication by moving common code used while exporting
/proc/vmcore and/or /sys/firmware/opal/core file(s).
Signed-off-by: Hari Bathini <redacted>
---
arch/powerpc/Kconfig | 9
arch/powerpc/platforms/powernv/Makefile | 1
arch/powerpc/platforms/powernv/opal-core.c | 599 ++++++++++++++++++++++++++
arch/powerpc/platforms/powernv/opal-fadump.c | 84 +---
arch/powerpc/platforms/powernv/opal-fadump.h | 71 +++
5 files changed, 697 insertions(+), 67 deletions(-)
create mode 100644 arch/powerpc/platforms/powernv/opal-core.c
@@ -589,6 +589,15 @@ config PRESERVE_FA_DUMPmemorypreservingkernelbootwouldprocessthiscrashdata.Petitbootkernelisthetypicalusecaseforthisoption.+configOPAL_CORE+bool"Export OPAL memory as /sys/firmware/opal/core"+depends onPPC64&&PPC_POWERNV+help+ThisoptionusestheMPIPLsupportinfirmwaretoprovidean+ELFcoreofOPALmemoryafteracrash.TheELFcoreisexported+as/sys/firmware/opal/corefilewhichishelpfulindebugging+OPALcrashesusingGDB.+configIRQ_ALL_CPUSbool"Distribute interrupts on all CPUs by default"depends onSMP
@@ -0,0 +1,599 @@+/*+*InterfaceforexportingtheOPALELFcore.+*Heavilyinspiredfromfs/proc/vmcore.c+*+*Copyright2019,IBMCorp.+*Author:HariBathini<hbathini@linux.ibm.com>+*+*Thisprogramisfreesoftware;youcanredistributeitand/or+*modifyitunderthetermsoftheGNUGeneralPublicLicense+*aspublishedbytheFreeSoftwareFoundation;eitherversion+*2oftheLicense,or(atyouroption)anylaterversion.+*/++#undef DEBUG+#define pr_fmt(fmt) "opalcore: " fmt++#include<linux/memblock.h>+#include<linux/uaccess.h>+#include<linux/proc_fs.h>+#include<linux/elf.h>+#include<linux/elfcore.h>+#include<linux/slab.h>+#include<linux/crash_core.h>+#include<linux/of.h>++#include<asm/page.h>+#include<asm/opal.h>++#include"../../kernel/fadump-common.h"+#include"opal-fadump.h"++#define MAX_PT_LOAD_CNT 8++/* NT_AUXV note related info */+#define AUXV_CNT 1+#define AUXV_DESC_SZ (((2 * AUXV_CNT) + 1) * sizeof(Elf64_Off))++structopalcore_config{+unsignedintnum_cpus;+/* PIR value of crashing CPU */+unsignedintcrashing_cpu;++/* CPU state data info from F/W */+unsignedlongcpu_state_destination_addr;+unsignedlongcpu_state_data_size;+unsignedlongcpu_state_entry_size;++/* OPAL memory to be exported as PT_LOAD segments */+unsignedlongptload_addr[MAX_PT_LOAD_CNT];+unsignedlongptload_size[MAX_PT_LOAD_CNT];+unsignedlongptload_cnt;++/* Pointer to the first PT_LOAD in the ELF core file */+Elf64_Phdr*ptload_phdr;++/* Total size of opalcore file. */+size_topalcore_size;++/* Buffer for all the ELF core headers and the PT_NOTE */+size_topalcorebuf_sz;+char*opalcorebuf;++/* NT_AUXV buffer */+charauxv_buf[AUXV_DESC_SZ];+};++structopalcore{+structlist_headlist;+unsignedlonglongpaddr;+unsignedlonglongsize;+loff_toffset;+};++staticLIST_HEAD(opalcore_list);+staticstructopalcore_config*oc_conf;+staticconststructopal_mpipl_fadump*opalc_metadata;+staticconststructopal_mpipl_fadump*opalc_cpu_metadata;++/*+*SetcrashingCPU'ssignaltoSIGUSR1.ifthekernelistriggered+*bykernel,SIGTERMotherwise.+*/+boolkernel_initiated;++staticstructopalcore*__initget_new_element(void)+{+returnkzalloc(sizeof(structopalcore),GFP_KERNEL);+}++staticinlineintis_opalcore_usable(void)+{+return(oc_conf&&oc_conf->opalcorebuf!=NULL)?1:0;+}++staticElf64_Word*append_elf64_note(Elf64_Word*buf,char*name,+unsignedinttype,void*data,+size_tdata_len)+{+Elf64_Nhdr*note=(Elf64_Nhdr*)buf;+Elf64_Wordnamesz=strlen(name)+1;++note->n_namesz=cpu_to_be32(namesz);+note->n_descsz=cpu_to_be32(data_len);+note->n_type=cpu_to_be32(type);+buf+=DIV_ROUND_UP(sizeof(*note),sizeof(Elf64_Word));+memcpy(buf,name,namesz);+buf+=DIV_ROUND_UP(namesz,sizeof(Elf64_Word));+memcpy(buf,data,data_len);+buf+=DIV_ROUND_UP(data_len,sizeof(Elf64_Word));++returnbuf;+}++staticvoidfill_prstatus(structelf_prstatus*prstatus,intpir,+structpt_regs*regs)+{+memset(prstatus,0,sizeof(structelf_prstatus));+elf_core_copy_kernel_regs(&(prstatus->pr_reg),regs);++/*+*OverloadPIDwithPIRvalue.+*AsaPIRvaluecouldalsobe'0',addanoffsetof'100'+*toeveryPIRtoavoidmisinterpretationsinGDB.+*/+prstatus->pr_pid=cpu_to_be32(100+pir);+prstatus->pr_ppid=cpu_to_be32(1);++/*+*IndicateSIGUSR1forcrashinitiatedfromkernel.+*SIGTERMotherwise.+*/+if(pir==oc_conf->crashing_cpu){+shortsig;++sig=kernel_initiated?SIGUSR1:SIGTERM;+prstatus->pr_cursig=cpu_to_be16(sig);+}+}++staticElf64_Word*auxv_to_elf64_notes(Elf64_Word*buf,+uint64_topal_boot_entry)+{+intidx=0;+Elf64_Off*bufp=(Elf64_Off*)oc_conf->auxv_buf;++memset(bufp,0,AUXV_DESC_SZ);++/* Entry point of OPAL */+bufp[idx++]=cpu_to_be64(AT_ENTRY);+bufp[idx++]=cpu_to_be64(opal_boot_entry);++/* end of vector */+bufp[idx++]=cpu_to_be64(AT_NULL);++buf=append_elf64_note(buf,CRASH_CORE_NOTE_NAME,NT_AUXV,+oc_conf->auxv_buf,AUXV_DESC_SZ);+returnbuf;+}++/*+*ReadfromtheELFheaderandthenthecrashdump.+*Returnsnumberofbytesreadonsuccess,-errnoonfailure.+*/+staticssize_tread_opalcore(structfile*file,structkobject*kobj,+structbin_attribute*bin_attr,char*to,+loff_tpos,size_tcount)+{+structopalcore*m;+ssize_ttsz,avail;+loff_ttpos=pos;++if(pos>=oc_conf->opalcore_size)+return0;++/* Adjust count if it goes beyond opalcore size */+avail=oc_conf->opalcore_size-pos;+if(count>avail)+count=avail;++if(count==0)+return0;++/* Read ELF core header and/or PT_NOTE segment */+if(tpos<oc_conf->opalcorebuf_sz){+tsz=min_t(size_t,oc_conf->opalcorebuf_sz-tpos,count);+memcpy(to,oc_conf->opalcorebuf+tpos,tsz);+to+=tsz;+tpos+=tsz;+count-=tsz;+}++list_for_each_entry(m,&opalcore_list,list){+/* nothing more to read here */+if(count==0)+break;++if(tpos<m->offset+m->size){+void*addr;++tsz=min_t(size_t,m->offset+m->size-tpos,count);+addr=(void*)(m->paddr+tpos-m->offset);+memcpy(to,__va(addr),tsz);+to+=tsz;+tpos+=tsz;+count-=tsz;+}+}++return(tpos-pos);+}++staticstructbin_attributeopal_core_attr={+.attr={.name="core",.mode=0400},+.read=read_opalcore+};++/*+*ReadCPUstatedumpdataandconvertitintoELFnotes.+*+*Eachregisterentryisof16bytes,Anumericalidentifieralongwith+*aGPR/SPRflaginthefirst8bytesandtheregistervalueinthenext+*8bytes.FormoredetailsrefertoF/Wdocumentation.+*/+staticElf64_Word*__initopalcore_append_cpu_notes(Elf64_Word*buf)+{+structhdat_fadump_thread_hdr*thdr;+unsignedlongaddr;+u32thread_pir;+char*bufp;+Elf64_Word*first_cpu_note;+structpt_regsregs;+structelf_prstatusprstatus;+unsignedintsize_of_each_thread;+unsignedintregs_offset,regs_cnt,reg_esize;+inti;++size_of_each_thread=oc_conf->cpu_state_entry_size;++addr=oc_conf->cpu_state_destination_addr;+bufp=__va(addr);++/*+*Offsetforregisterentries,entrysizeandregisterscountis+*duplicatedineverythreadheaderinkeepingwithHDATformat.+*Usethesevaluesfromthefirstthreadheader.+*/+thdr=(structhdat_fadump_thread_hdr*)bufp;+regs_offset=(offsetof(structhdat_fadump_thread_hdr,offset)++be32_to_cpu(thdr->offset));+reg_esize=be32_to_cpu(thdr->esize);+regs_cnt=be32_to_cpu(thdr->ecnt);++pr_debug("--------CPU State Data------------\n");+pr_debug("NumCpus : %u\n",oc_conf->num_cpus);+pr_debug("\tOffset: %u, Entry size: %u, Cnt: %u\n",+regs_offset,reg_esize,regs_cnt);++/*+*SkippastthefirstCPUnote.Fillthisnotewiththe+*crashingCPU'sprstatus.+*/+first_cpu_note=buf;+buf=append_elf64_note(buf,CRASH_CORE_NOTE_NAME,NT_PRSTATUS,+&prstatus,sizeof(prstatus));++for(i=0;i<oc_conf->num_cpus;i++,bufp+=size_of_each_thread){+thdr=(structhdat_fadump_thread_hdr*)bufp;+thread_pir=be32_to_cpu(thdr->pir);++pr_debug("%04d) PIR: 0x%x, core state: 0x%02x\n",+(i+1),thread_pir,thdr->core_state);++/*+*RegisterstatedataofMAXcoresisprovidedbyfirmware,+*butsomeofthiscoresmaynotbeactive.So,while+*processingregisterstatedata,checkcorestateand+*skipthreadsthatbelongtoinactivecores.+*/+if(is_thread_core_inactive(thdr->core_state))+continue;++opal_fadump_read_regs((bufp+regs_offset),regs_cnt,+reg_esize,false,®s);++pr_debug("PIR 0x%x - R1 : 0x%llx, NIP : 0x%llx\n",thread_pir,+be64_to_cpu(regs.gpr[1]),be64_to_cpu(regs.nip));+fill_prstatus(&prstatus,thread_pir,®s);++if(thread_pir!=oc_conf->crashing_cpu){+buf=append_elf64_note(buf,CRASH_CORE_NOTE_NAME,+NT_PRSTATUS,&prstatus,+sizeof(prstatus));+}else{+/*+*AddcrashingCPUasthefirstNT_PRSTATUSnotefor+*GDBtoprocessthecorefileappropriately.+*/+append_elf64_note(first_cpu_note,CRASH_CORE_NOTE_NAME,+NT_PRSTATUS,&prstatus,+sizeof(prstatus));+}+}++returnbuf;+}++staticint__initcreate_opalcore(void)+{+inthdr_size,cpu_notes_size,order,count;+inti,ret;+unsignedintnumcpus;+unsignedlongpaddr;+Elf64_Ehdr*elf;+Elf64_Phdr*phdr;+loff_topalcore_off;+structopalcore*new;+structpage*page;+char*bufp;+structdevice_node*dn;+uint64_topal_base_addr;+uint64_topal_boot_entry;+++if((oc_conf->ptload_cnt==0)||+(oc_conf->ptload_cnt>MAX_PT_LOAD_CNT)){+pr_err("Invalid PT_LOAD count: %lu\n",oc_conf->ptload_cnt);+return-EINVAL;+}++numcpus=oc_conf->num_cpus;+hdr_size=(sizeof(Elf64_Ehdr)++((oc_conf->ptload_cnt+1)*sizeof(Elf64_Phdr)));+cpu_notes_size=((numcpus*(CRASH_CORE_NOTE_HEAD_BYTES++CRASH_CORE_NOTE_NAME_BYTES++CRASH_CORE_NOTE_DESC_BYTES))++(CRASH_CORE_NOTE_HEAD_BYTES++CRASH_CORE_NOTE_NAME_BYTES+AUXV_DESC_SZ));+oc_conf->opalcorebuf_sz=(hdr_size+cpu_notes_size);+order=get_order(oc_conf->opalcorebuf_sz);+oc_conf->opalcorebuf=+(char*)__get_free_pages(GFP_KERNEL|__GFP_ZERO,order);+if(!oc_conf->opalcorebuf){+pr_err("Not enough memory to setup opalcore (size: %lu)\n",+oc_conf->opalcorebuf_sz);+oc_conf->opalcorebuf_sz=0;+return-ENOMEM;+}++pr_debug("opalcorebuf = 0x%lx\n",(unsignedlong)oc_conf->opalcorebuf);++count=1<<order;+page=virt_to_page(oc_conf->opalcorebuf);+for(i=0;i<count;i++)+SetPageReserved(page+i);++/* Read OPAL related device-tree entries */+dn=of_find_node_by_name(NULL,"ibm,opal");+if(dn){+ret=of_property_read_u64(dn,"opal-base-address",+&opal_base_addr);+pr_debug("opal-base-address: %llx\n",opal_base_addr);+ret|=of_property_read_u64(dn,"opal-boot-address",+&opal_boot_entry);+pr_debug("opal-boot-address: %llx\n",opal_boot_entry);+}+if(!dn||ret)+pr_warn("WARNING: Failed to read OPAL base & entry values\n");++/* Use count to keep track of the program headers */+count=0;++bufp=oc_conf->opalcorebuf;+elf=(Elf64_Ehdr*)bufp;+bufp+=sizeof(Elf64_Ehdr);+memcpy(elf->e_ident,ELFMAG,SELFMAG);+elf->e_ident[EI_CLASS]=ELF_CLASS;+elf->e_ident[EI_DATA]=ELFDATA2MSB;+elf->e_ident[EI_VERSION]=EV_CURRENT;+elf->e_ident[EI_OSABI]=ELF_OSABI;+memset(elf->e_ident+EI_PAD,0,EI_NIDENT-EI_PAD);+elf->e_type=cpu_to_be16(ET_CORE);+elf->e_machine=cpu_to_be16(ELF_ARCH);+elf->e_version=cpu_to_be32(EV_CURRENT);+elf->e_entry=0;+elf->e_phoff=cpu_to_be64(sizeof(Elf64_Ehdr));+elf->e_shoff=0;+elf->e_flags=0;++elf->e_ehsize=cpu_to_be16(sizeof(Elf64_Ehdr));+elf->e_phentsize=cpu_to_be16(sizeof(Elf64_Phdr));+elf->e_phnum=0;+elf->e_shentsize=0;+elf->e_shnum=0;+elf->e_shstrndx=0;++phdr=(Elf64_Phdr*)bufp;+bufp+=sizeof(Elf64_Phdr);+phdr->p_type=cpu_to_be32(PT_NOTE);+phdr->p_flags=0;+phdr->p_align=0;+phdr->p_paddr=phdr->p_vaddr=0;+phdr->p_offset=cpu_to_be64(hdr_size);+phdr->p_filesz=phdr->p_memsz=cpu_to_be64(cpu_notes_size);+count++;++opalcore_off=oc_conf->opalcorebuf_sz;+oc_conf->ptload_phdr=(Elf64_Phdr*)bufp;+paddr=0;+for(i=0;i<oc_conf->ptload_cnt;i++){+phdr=(Elf64_Phdr*)bufp;+bufp+=sizeof(Elf64_Phdr);+phdr->p_type=cpu_to_be32(PT_LOAD);+phdr->p_flags=cpu_to_be32(PF_R|PF_W|PF_X);+phdr->p_align=0;++new=get_new_element();+if(!new)+return-ENOMEM;+new->paddr=oc_conf->ptload_addr[i];+new->size=oc_conf->ptload_size[i];+new->offset=opalcore_off;+list_add_tail(&new->list,&opalcore_list);++phdr->p_paddr=cpu_to_be64(paddr);+phdr->p_vaddr=cpu_to_be64(opal_base_addr+paddr);+phdr->p_filesz=phdr->p_memsz=+cpu_to_be64(oc_conf->ptload_size[i]);+phdr->p_offset=cpu_to_be64(opalcore_off);++count++;+opalcore_off+=oc_conf->ptload_size[i];+paddr+=oc_conf->ptload_size[i];+}++elf->e_phnum=cpu_to_be16(count);++bufp=(char*)opalcore_append_cpu_notes((Elf64_Word*)bufp);+bufp=(char*)auxv_to_elf64_notes((Elf64_Word*)bufp,opal_boot_entry);++oc_conf->opalcore_size=opalcore_off;+return0;+}++staticvoid__initopalcore_config_init(void)+{+structdevice_node*np;+const__be32*prop;+uint64_taddr=0;+uint32_tidx,cpu_data_version;+inti,ret;+++np=of_find_node_by_path("/ibm,opal/dump");+if(np==NULL)+return;++if(!of_device_is_compatible(np,"ibm,opal-dump")){+pr_err("Support missing for this f/w version!\n");+return;+}++/*+*Checkifdumphasbeeninitiatedonlastreboot.+*/+prop=of_get_property(np,"mpipl-boot",NULL);+if(!prop)+gotoout;++ret=opal_mpipl_query_tag(OPAL_MPIPL_TAG_OPAL,&addr);+if((ret!=OPAL_SUCCESS)||!addr){+pr_err("Failed to get OPAL metadata (%d)\n",ret);+gotoout;+}++addr=be64_to_cpu(addr);+pr_debug("OPAL metadata addr: %llx\n",addr);+opalc_metadata=__va(addr);+if(opalc_metadata->version!=MPIPL_FADUMP_VERSION){+pr_err("OPAL metadata version (%u) not supported by kernel!\n",+opalc_metadata->version);+gotoout;+}++ret=opal_mpipl_query_tag(OPAL_MPIPL_TAG_CPU,&addr);+if((ret!=OPAL_SUCCESS)||!addr){+pr_err("Failed to get OPAL CPU metadata (%d)\n",ret);+gotoout;+}++addr=be64_to_cpu(addr);+pr_debug("CPU metadata addr: %llx\n",addr);+opalc_cpu_metadata=__va(addr);+cpu_data_version=be32_to_cpu(opalc_cpu_metadata->cpu_data_version);+if(cpu_data_version!=HDAT_FADUMP_CPU_DATA_VERSION){+pr_err("CPU data version (%u) not supported by kernel!\n",+cpu_data_version);+gotoout;+}++oc_conf=kzalloc(sizeof(structopalcore_config),GFP_KERNEL);+if(oc_conf==NULL)+gotoout;++oc_conf->ptload_cnt=0;+idx=be32_to_cpu(opalc_metadata->region_cnt);+if(idx>MAX_PT_LOAD_CNT){+pr_warn("OPAL regions count (%d) adjusted to limit (%d)",+MAX_PT_LOAD_CNT,idx);+idx=MAX_PT_LOAD_CNT;+}+for(i=0;i<idx;i++){+oc_conf->ptload_addr[oc_conf->ptload_cnt]=+be64_to_cpu(opalc_metadata->region[i].dest);+oc_conf->ptload_size[oc_conf->ptload_cnt++]=+be64_to_cpu(opalc_metadata->region[i].size);+}+oc_conf->ptload_cnt=i;+oc_conf->crashing_cpu=be32_to_cpu(opalc_metadata->crashing_pir);++oc_conf->cpu_state_destination_addr=+be64_to_cpu(opalc_cpu_metadata->region[0].dest);+oc_conf->cpu_state_data_size=+be64_to_cpu(opalc_cpu_metadata->region[0].size);+oc_conf->cpu_state_entry_size=+be32_to_cpu(opalc_cpu_metadata->cpu_data_size);++oc_conf->num_cpus=(oc_conf->cpu_state_data_size/+oc_conf->cpu_state_entry_size);++out:+of_node_put(np);+}++/* Cleanup function for opalcore module. */+staticvoidopalcore_cleanup(void)+{+unsignedlongorder,count,i;+structpage*page;++if(oc_conf==NULL)+return;++sysfs_remove_bin_file(opal_kobj,&opal_core_attr);+oc_conf->ptload_phdr=NULL;+oc_conf->ptload_cnt=0;++/* free core buffer */+if((oc_conf->opalcorebuf!=NULL)&&(oc_conf->opalcorebuf_sz!=0)){+order=get_order(oc_conf->opalcorebuf_sz);+count=1<<order;+page=virt_to_page(oc_conf->opalcorebuf);+for(i=0;i<count;i++)+ClearPageReserved(page+i);+__free_pages(page,order);++oc_conf->opalcorebuf=NULL;+oc_conf->opalcorebuf_sz=0;+}++kfree(oc_conf);+oc_conf=NULL;+}+__exitcall(opalcore_cleanup);++/* Init function for opalcore module. */+staticint__initopalcore_init(void)+{+intrc=-1;++opalcore_config_init();++if(oc_conf==NULL)+returnrc;++create_opalcore();++/*+*Ifoc_conf->opalcorebuf=issetinthe2ndkernel,+*thencapturethedump.+*/+if(!(is_opalcore_usable())){+pr_err("Failed to export /sys/firmware/opal/core\n");+opalcore_cleanup();+returnrc;+}++/* Set opal core size */+opal_core_attr.size=oc_conf->opalcore_size;++rc=sysfs_create_bin_file(opal_kobj,&opal_core_attr);+if(rc!=0){+pr_err("Failed to export /sys/firmware/opal/core\n");+opalcore_cleanup();+returnrc;+}++return0;+}+fs_initcall(opalcore_init);
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 12:23:13
If not all kernel boot memory regions are registered for MPIPL before
system crashes, try processing the partial crashdump but warn the user
before proceeding.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/platforms/powernv/opal-fadump.c | 21 +++++++++++++++++++++
1 file changed, 21 insertions(+)
@@ -136,6 +136,27 @@ static void opal_fadump_get_config(struct fw_dump *fadump_conf,last_end=base+size;}+/*+*Rarely,butitcansohappenthatsystemcrashesbeforeall+*bootmemoryregionsareregisteredforMPIPL.Insuch+*cases,warnthatthevmcoremaynotbeaccurateandproceed+*anywayasthatisthebestbetconsideringfreepages,cache+*pages,userpages,etcareusuallyfilteredout.+*+*Hopethememorythatcouldnotbepreservedonlyhaspages+*thatareusuallyfilteredoutwhilesavingthevmcore.+*/+if(fdm->region_cnt<fdm->registered_regions){+pr_warn("The crashdump may not be accurate as the below boot memory regions could not be preserved:\n");+i=fdm->registered_regions;+while(i<fdm->region_cnt){+pr_warn("\t%d. base: 0x%llx, size: 0x%llx\n",+(i+1),fdm->rgn[i].src,+fdm->rgn[i].size);+i++;+}+}+fadump_conf->boot_mem_top=(fadump_conf->boot_memory_size+hole_size);fadump_conf->boot_mem_regs_cnt=fdm->region_cnt;opal_fadump_update_config(fadump_conf,fdm);
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 12:25:52
Writing '1' to /sys/kernel/fadump_release_opalcore would release the
memory held by kernel in exporting /sys/firmware/opal/core file.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/platforms/powernv/opal-core.c | 38 ++++++++++++++++++++++++++++
1 file changed, 38 insertions(+)
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 12:28:38
OPAL loads kernel & initrd at 512MB offset (256MB size), also exported
as ibm,opal/dump/fw-load-area. So, if boot memory size of FADump is
less than 768MB, kernel memory to be exported as '/proc/vmcore' would
be overwritten by f/w while loading kernel & initrd. To avoid such a
scenario, enforce a minimum boot memory size of 768MB on OPAL platform.
Also, skip using FADump if a newer F/W version loads kernel & initrd
above 768MB.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/kernel/fadump-common.h | 11 +---------
arch/powerpc/kernel/fadump.c | 11 +++++++++-
arch/powerpc/platforms/powernv/opal-fadump.c | 29 ++++++++++++++++++++++++++
arch/powerpc/platforms/powernv/opal-fadump.h | 7 ++++++
arch/powerpc/platforms/pseries/rtas-fadump.c | 6 +++++
arch/powerpc/platforms/pseries/rtas-fadump.h | 11 ++++++++++
6 files changed, 64 insertions(+), 11 deletions(-)
@@ -335,7 +335,8 @@ static inline unsigned long fadump_calculate_reserve_size(void)if(memory_limit&&size>memory_limit)size=memory_limit;-return(size>MIN_BOOT_MEM?size:MIN_BOOT_MEM);+return(size>fw_dump.ops->get_bootmem_min()?size:+fw_dump.ops->get_bootmem_min());}/*
@@ -493,6 +494,14 @@ int __init fadump_reserve_mem(void)ALIGN(fw_dump.boot_memory_size,FADUMP_CMA_ALIGNMENT);#endif++if(fw_dump.boot_memory_size<fw_dump.ops->get_bootmem_min()){+pr_err("Can't enable fadump with boot memory size (0x%lx) less than 0x%lx\n",+fw_dump.boot_memory_size,+fw_dump.ops->get_bootmem_min());+gotoerror_out;+}+if(!fadump_get_boot_mem_regions()){pr_err("Too many holes in boot memory area to enable fadump\n");gotoerror_out;
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-07-16 12:31:31
With /sys/firmware/opal/core support available on OPAL based machines
and an option to the release memory used by kernel in exporting this
core file, update FADump documentation with these details.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
Documentation/powerpc/firmware-assisted-dump.txt | 19 +++++++++++++++++++
1 file changed, 19 insertions(+)
@@ -107,6 +107,16 @@ capture kernel boot to process this crash data. Kernel config option CONFIG_PRESERVE_FA_DUMP has to be enabled on such kernel to ensure that crash data is preserved to process later.+-- On OPAL based machines (PowerNV), if the kernel is build with+ CONFIG_OPAL_CORE=y, OPAL memory at the time of crash is also+ exported as /sys/firmware/opal/core file. This procfs file is+ helpful in debugging OPAL crashes with GDB. The kernel memory+ used for exporting this procfs file can be released by echo'ing+ '1' to /sys/kernel/fadump_release_opalcore node.++ e.g.+ # echo 1 > /sys/kernel/fadump_release_opalcore+ Implementation details: ----------------------
@@ -270,6 +280,15 @@ Here is the list of files under kernel sysfs: enhanced to use this interface to release the memory reserved for dump and continue without 2nd reboot.+ /sys/kernel/fadump_release_opalcore++ This file is available only on OPAL based machines when FADump is+ active during capture kernel. This is used to release the memory+ used by the kernel to export /sys/firmware/opal/core file. To+ release this memory, echo '1' to it:++ echo 1 > /sys/kernel/fadump_release_opalcore+ Here is the list of files under powerpc debugfs: (Assuming debugfs is mounted on /sys/kernel/debug directory.)
Firmware-Assisted Dump (FADump) is currently supported only on pSeries
platform. This patch series adds support for PowerNV platform too.
The first few patches refactor the FADump code to make use of common
code across multiple platforms. Then basic FADump support is added for
PowerNV platform. Followed by patches to honour reserved-ranges DT node
while reserving/releasing memory used by FADump. The subsequent patch
processes CPU state data provided by firmware to create and append core
notes to the ELF core file and the next patch adds support to preserve
crash data for subsequent boots (useful in cases like petitboot). The
subsequent patches add support to export opalcore. opalcore makes
debugging of failures in OPAL code easier. Firmware-Assisted Dump
documentation is also updated appropriately.
The patch series is tested with the latest firmware plus the below skiboot
changes for MPIPL support:
https://patchwork.ozlabs.org/project/skiboot/list/?series=119169
("MPIPL support")
Changes in v4:
* Split the patches.
* Rebased to latest upstream kernel version.
* Updated according to latest OPAL changes.
---
Hari Bathini (25):
powerpc/fadump: move internal macros/definitions to a new header
powerpc/fadump: move internal code to a new file
powerpc/fadump: Improve fadump documentation
pseries/fadump: move rtas specific definitions to platform code
pseries/fadump: introduce callbacks for platform specific operations
pseries/fadump: define register/un-register callback functions
pseries/fadump: move out platform specific support from generic code
powerpc/fadump: use FADump instead of fadump for how it is pronounced
opal: add MPIPL interface definitions
powernv/fadump: add fadump support on powernv
powernv/fadump: register kernel metadata address with opal
powernv/fadump: define register/un-register callback functions
powernv/fadump: support copying multiple kernel memory regions
powernv/fadump: process the crashdump by exporting it as /proc/vmcore
powerpc/fadump: Update documentation about OPAL platform support
powerpc/fadump: consider reserved ranges while reserving memory
powerpc/fadump: consider reserved ranges while releasing memory
powernv/fadump: process architected register state data provided by firmware
powernv/fadump: add support to preserve crash data on FADUMP disabled kernel
powerpc/fadump: update documentation about CONFIG_PRESERVE_FA_DUMP
powernv/opalcore: export /sys/firmware/opal/core for analysing opal crashes
powernv/fadump: Warn before processing partial crashdump
powernv/opalcore: provide an option to invalidate /sys/firmware/opal/core file
powernv/fadump: consider f/w load area
powernv/fadump: update documentation about option to release opalcore
Documentation/powerpc/firmware-assisted-dump.txt | 224 +++-
arch/powerpc/Kconfig | 23
arch/powerpc/include/asm/fadump.h | 190 ----
arch/powerpc/include/asm/opal-api.h | 50 +
arch/powerpc/include/asm/opal.h | 6
arch/powerpc/kernel/Makefile | 6
arch/powerpc/kernel/fadump-common.c | 153 +++
arch/powerpc/kernel/fadump-common.h | 203 ++++
arch/powerpc/kernel/fadump.c | 1181 ++++++++--------------
arch/powerpc/kernel/prom.c | 4
arch/powerpc/platforms/powernv/Makefile | 3
arch/powerpc/platforms/powernv/opal-call.c | 3
arch/powerpc/platforms/powernv/opal-core.c | 637 ++++++++++++
arch/powerpc/platforms/powernv/opal-fadump.c | 671 ++++++++++++
arch/powerpc/platforms/powernv/opal-fadump.h | 154 +++
arch/powerpc/platforms/pseries/Makefile | 1
arch/powerpc/platforms/pseries/rtas-fadump.c | 595 +++++++++++
arch/powerpc/platforms/pseries/rtas-fadump.h | 123 ++
18 files changed, 3231 insertions(+), 996 deletions(-)
create mode 100644 arch/powerpc/kernel/fadump-common.c
create mode 100644 arch/powerpc/kernel/fadump-common.h
create mode 100644 arch/powerpc/platforms/powernv/opal-core.c
create mode 100644 arch/powerpc/platforms/powernv/opal-fadump.c
create mode 100644 arch/powerpc/platforms/powernv/opal-fadump.h
create mode 100644 arch/powerpc/platforms/pseries/rtas-fadump.c
create mode 100644 arch/powerpc/platforms/pseries/rtas-fadump.h
The figures depicting FADump's (Firmware-Assisted Dump) memory layout
are missing some finer details like different memory regions and what
they represent. Improve the documentation by updating those details.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
Documentation/powerpc/firmware-assisted-dump.txt | 65 ++++++++++++----------
1 file changed, 35 insertions(+), 30 deletions(-)
This will have to be rebased now on firmware-assisted-dump.rst. However
Changes looks good to me.
Thanks,
-Mahesh.
quoted hunk
@@ -74,8 +74,9 @@ as follows: there is crash data available from a previous boot. During the early boot OS will reserve rest of the memory above boot memory size effectively booting with restricted memory- size. This will make sure that the second kernel will not- touch any of the dump memory area.+ size. This will make sure that this kernel (also, referred+ to as second kernel or capture kernel) will not touch any+ of the dump memory area. -- User-space tools will read /proc/vmcore to obtain the contents of memory, which holds the previous crashed kernel dump in ELF
@@ -125,48 +126,52 @@ space memory except the user pages that were present in CMA region. o Memory Reservation during first kernel- Low memory Top of memory- 0 boot memory size |- | | |<--Reserved dump area -->| |- V V | Permanent Reservation | V- +-----------+----------/ /---+---+----+-----------+----+------+- | | |CPU|HPTE| DUMP |ELF | |- +-----------+----------/ /---+---+----+-----------+----+------+- | ^- | |- \ /- -------------------------------------------- Boot memory content gets transferred to- reserved area by firmware at the time of- crash+ Low memory Top of memory+ 0 boot memory size |<--Reserved dump area --->| |+ | | | Permanent Reservation | |+ V V | (Preserve area) | V+ +-----------+----------/ /---+---+----+--------+---+----+------++ | | |CPU|HPTE| DUMP |HDR|ELF | |+ +-----------+----------/ /---+---+----+--------+---+----+------++ | ^ ^+ | | |+ \ / |+ ----------------------------------- FADump Header+ Boot memory content gets transferred (meta area)+ to reserved area by firmware at the+ time of crash+ Fig. 1+ o Memory Reservation during second kernel after crash- Low memory Top of memory- 0 boot memory size |- | |<------------- Reserved dump area ----------- -->|- V V V- +-----------+----------/ /---+---+----+-----------+----+------+- | | |CPU|HPTE| DUMP |ELF | |- +-----------+----------/ /---+---+----+-----------+----+------++ Low memory Top of memory+ 0 boot memory size |+ | |<------------- Reserved dump area --------------->|+ V V |<---- Preserve area ----->| V+ +-----------+----------/ /---+---+----+--------+---+----+------++ | | |CPU|HPTE| DUMP |HDR|ELF | |+ +-----------+----------/ /---+---+----+--------+---+----+------+ | | V V Used by second /proc/vmcore kernel to boot Fig. 2-Currently the dump will be copied from /proc/vmcore to a-a new file upon user intervention. The dump data available through-/proc/vmcore will be in ELF format. Hence the existing kdump-infrastructure (kdump scripts) to save the dump works fine with-minor modifications.+Currently the dump will be copied from /proc/vmcore to a new file upon+user intervention. The dump data available through /proc/vmcore will be+in ELF format. Hence the existing kdump infrastructure (kdump scripts)+to save the dump works fine with minor modifications. KDump scripts on+major Distro releases have already been modified to work seemlessly (no+user intervention in saving the dump) when FADump is used, instead of+KDump, as dump mechanism. The tools to examine the dump will be same as the ones used for kdump. How to enable firmware-assisted dump (fadump):--------------------------------------+--------------------------------------------- 1. Set config option CONFIG_FA_DUMP=y and build kernel. 2. Boot into linux kernel with 'fadump=on' kernel cmdline option.
@@ -189,7 +194,7 @@ NOTE: 1. 'fadump_reserve_mem=' parameter has been deprecated. Instead old behaviour. Sysfs/debugfs files:-------------+------------------- Firmware-assisted dump feature uses sysfs file system to hold the control files and debugfs file to display memory reserved region.
Do we really need these ? Aren't we hiding all platform specific things
under fadump_ops functions ? I see that these values are used only for
assignements and not making any decision in code flow. Am I missing
anything here ?
Thanks,
-Mahesh.
quoted hunk
+
/*
* Copy the ascii values for first 8 characters from a string into u64
* variable at their respective indexes.
@@ -84,6 +90,9 @@ struct fad_crash_memory_ranges { unsigned long long size; };+/* Platform specific callback functions */+struct fadump_ops;+ /* Firmware-assisted dump configuration details. */ struct fw_dump { unsigned long reserve_dump_area_start;
@@ -106,6 +115,21 @@ struct fw_dump { unsigned long dump_active:1; unsigned long dump_registered:1; unsigned long nocma:1;++ enum fadump_platform_type fadump_platform;+ struct fadump_ops *ops;+};++struct fadump_ops {+ ulong (*init_fadump_mem_struct)(struct fw_dump *fadump_config);+ int (*register_fadump)(struct fw_dump *fadump_config);+ int (*unregister_fadump)(struct fw_dump *fadump_config);+ int (*invalidate_fadump)(struct fw_dump *fadump_config);+ int (*process_fadump)(struct fw_dump *fadump_config);+ void (*fadump_region_show)(struct fw_dump *fadump_config,+ struct seq_file *m);+ void (*fadump_trigger)(struct fadump_crash_info_header *fdh,+ const char *msg); }; /* Helper functions */
@@ -116,4 +140,13 @@ void fadump_update_elfcore_header(struct fw_dump *fadump_config, char *bufp); int is_fadump_boot_mem_contiguous(struct fw_dump *fadump_conf); int is_fadump_reserved_mem_contiguous(struct fw_dump *fadump_conf);+#ifdef CONFIG_PPC_PSERIES+extern int rtas_fadump_dt_scan(struct fw_dump *fadump_config, ulong node);+#else+static inline int rtas_fadump_dt_scan(struct fw_dump *fadump_config, ulong node)+{+ return 1;+}+#endif+ #endif /* __PPC64_FA_DUMP_INTERNAL_H__ */
@@ -139,36 +127,7 @@ int __init early_init_dt_scan_fw_dump(unsigned long node, const char *uname,if(fdm_active)fw_dump.dump_active=1;-/* Get the sizes required to store dump data for the firmware provided-*dumpsections.-*Foreachdumpsectiontypesupported,a32bitcellwhichdefines-*theIDofasupportedsectionfollowedbytwo32bitcellswhich-*givestehsizeofthesectioninbytes.-*/-sections=of_get_flat_dt_prop(node,"ibm,configure-kernel-dump-sizes",-&size);--if(!sections)-return1;--num_sections=size/(3*sizeof(u32));--for(i=0;i<num_sections;i++,sections+=3){-u32type=(u32)of_read_number(sections,1);--switch(type){-caseRTAS_FADUMP_CPU_STATE_DATA:-fw_dump.cpu_state_data_size=-of_read_ulong(§ions[1],2);-break;-caseRTAS_FADUMP_HPTE_REGION:-fw_dump.hpte_region_size=-of_read_ulong(§ions[1],2);-break;-}-}--return1;+returnret;}/*
@@ -0,0 +1,134 @@+/*+*Firmware-AssistedDumpsupportonPOWERVMplatform.+*+*Copyright2011,IBMCorporation+*Author:MaheshSalgaonkar<mahesh@linux.ibm.com>+*+*Copyright2019,IBMCorp.+*Author:HariBathini<hbathini@linux.ibm.com>+*+*Thisprogramisfreesoftware;youcanredistributeitand/or+*modifyitunderthetermsoftheGNUGeneralPublicLicense+*aspublishedbytheFreeSoftwareFoundation;eitherversion+*2oftheLicense,or(atyouroption)anylaterversion.+*/++#undef DEBUG+#define pr_fmt(fmt) "rtas fadump: " fmt++#include<linux/string.h>+#include<linux/memblock.h>+#include<linux/delay.h>+#include<linux/seq_file.h>+#include<linux/crash_dump.h>++#include<asm/page.h>+#include<asm/prom.h>+#include<asm/rtas.h>+#include<asm/fadump.h>++#include"../../kernel/fadump-common.h"+#include"rtas-fadump.h"++staticulongrtas_fadump_init_mem_struct(structfw_dump*fadump_conf)+{+returnfadump_conf->reserve_dump_area_start;+}++staticintrtas_fadump_register_fadump(structfw_dump*fadump_conf)+{+return-EIO;+}++staticintrtas_fadump_unregister_fadump(structfw_dump*fadump_conf)+{+return-EIO;+}++staticintrtas_fadump_invalidate_fadump(structfw_dump*fadump_conf)+{+return-EIO;+}++/*+*Validateandprocessthedumpdatastoredbyfirmwarebeforeexporting+*itthrough'/proc/vmcore'.+*/+staticint__initrtas_fadump_process_fadump(structfw_dump*fadump_conf)+{+return-EINVAL;+}++staticvoidrtas_fadump_region_show(structfw_dump*fadump_conf,+structseq_file*m)+{+}++staticvoidrtas_fadump_trigger(structfadump_crash_info_header*fdh,+constchar*msg)+{+/* Call ibm,os-term rtas call to trigger firmware assisted dump */+rtas_os_term((char*)msg);+}++staticstructfadump_opsrtas_fadump_ops={+.init_fadump_mem_struct=rtas_fadump_init_mem_struct,+.register_fadump=rtas_fadump_register_fadump,+.unregister_fadump=rtas_fadump_unregister_fadump,+.invalidate_fadump=rtas_fadump_invalidate_fadump,+.process_fadump=rtas_fadump_process_fadump,+.fadump_region_show=rtas_fadump_region_show,+.fadump_trigger=rtas_fadump_trigger,+};++int__initrtas_fadump_dt_scan(structfw_dump*fadump_conf,ulongnode)+{+const__be32*sections;+inti,num_sections;+intsize;+const__be32*token;++/*+*CheckifFirmwareAssisteddumpissupported.ifyes,check+*ifdumphasbeeninitiatedonlastreboot.+*/+token=of_get_flat_dt_prop(node,"ibm,configure-kernel-dump",NULL);+if(!token)+return1;++fadump_conf->ibm_configure_kernel_dump=be32_to_cpu(*token);+fadump_conf->ops=&rtas_fadump_ops;+fadump_conf->fadump_platform=FADUMP_PLATFORM_PSERIES;+fadump_conf->fadump_supported=1;++/* Get the sizes required to store dump data for the firmware provided+*dumpsections.+*Foreachdumpsectiontypesupported,a32bitcellwhichdefines+*theIDofasupportedsectionfollowedbytwo32bitcellswhich+*givesthesizeofthesectioninbytes.+*/+sections=of_get_flat_dt_prop(node,"ibm,configure-kernel-dump-sizes",+&size);++if(!sections)+return1;++num_sections=size/(3*sizeof(u32));++for(i=0;i<num_sections;i++,sections+=3){+u32type=(u32)of_read_number(sections,1);++switch(type){+caseRTAS_FADUMP_CPU_STATE_DATA:+fadump_conf->cpu_state_data_size=+of_read_ulong(§ions[1],2);+break;+caseRTAS_FADUMP_HPTE_REGION:+fadump_conf->hpte_region_size=+of_read_ulong(§ions[1],2);+break;+}+}++return1;+}
Make RTAS calls to register and un-register for FADump. Also, update
how fadump_region contents are diplayed to provide more information.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/kernel/fadump-common.h | 2
arch/powerpc/kernel/fadump.c | 164 ++------------------------
arch/powerpc/platforms/pseries/rtas-fadump.c | 163 +++++++++++++++++++++++++-
3 files changed, 176 insertions(+), 153 deletions(-)
[...]
static int rtas_fadump_register_fadump(struct fw_dump *fadump_conf)
{
- return -EIO;
+ int rc, err = -EIO;
+ unsigned int wait_time;
+
+ /* TODO: Add upper time limit for the delay */
+ do {
+ rc = rtas_call(fadump_conf->ibm_configure_kernel_dump, 3, 1,
+ NULL, FADUMP_REGISTER, &fdm,
+ sizeof(struct rtas_fadump_mem_struct));
+
+ wait_time = rtas_busy_delay_time(rc);
+ if (wait_time)
+ mdelay(wait_time);
+
+ } while (wait_time);
+
+ switch (rc) {
+ case 0:
+ pr_info("Registration is successful!\n");
+ fadump_conf->dump_registered = 1;
+ err = 0;
+ break;
+ case -1:
+ pr_err("Failed to register. Hardware Error(%d).\n", rc);
+ break;
+ case -3:
+ if (!is_fadump_boot_mem_contiguous(fadump_conf))
+ pr_err("Can't hot-remove boot memory area.\n");
+ else if (!is_fadump_reserved_mem_contiguous(fadump_conf))
+ pr_err("Can't hot-remove reserved memory area.\n");
Any reason why we changed the error messages here ? it gives an impression as
if fadump reservation tried to hot remove memory and failed.
Rest looks fine to me..
Reviewed-by: Mahesh Salgaonkar <redacted>
Thanks,
-Mahesh.
@@ -469,218 +451,6 @@ void crash_fadump(struct pt_regs *regs, const char *str)fw_dump.ops->fadump_trigger(fdh,str);}-#define GPR_MASK 0xffffff0000000000-staticinlineintfadump_gpr_index(u64id)-{-inti=-1;-charstr[3];--if((id&GPR_MASK)==fadump_str_to_u64("GPR")){-/* get the digits at the end */-id&=~GPR_MASK;-id>>=24;-str[2]='\0';-str[1]=id&0xff;-str[0]=(id>>8)&0xff;-sscanf(str,"%d",&i);-if(i>31)-i=-1;-}-returni;-}--staticinlinevoidfadump_set_regval(structpt_regs*regs,u64reg_id,-u64reg_val)-{-inti;--i=fadump_gpr_index(reg_id);-if(i>=0)-regs->gpr[i]=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("NIA"))-regs->nip=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("MSR"))-regs->msr=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("CTR"))-regs->ctr=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("LR"))-regs->link=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("XER"))-regs->xer=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("CR"))-regs->ccr=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("DAR"))-regs->dar=(unsignedlong)reg_val;-elseif(reg_id==fadump_str_to_u64("DSISR"))-regs->dsisr=(unsignedlong)reg_val;-}--staticstructrtas_fadump_reg_entry*-fadump_read_registers(structrtas_fadump_reg_entry*reg_entry,structpt_regs*regs)-{-memset(regs,0,sizeof(structpt_regs));--while(be64_to_cpu(reg_entry->reg_id)!=fadump_str_to_u64("CPUEND")){-fadump_set_regval(regs,be64_to_cpu(reg_entry->reg_id),-be64_to_cpu(reg_entry->reg_value));-reg_entry++;-}-reg_entry++;-returnreg_entry;-}--/*-*ReadCPUstatedumpdataandconvertitintoELFnotes.-*TheCPUdumpstartswithmagicnumber"REGSAVE".NumCpusOffsetshouldbe-*usedtoaccessthedatatoallowforadditionalfieldstobeaddedwithout-*affectingcompatibility.EachlistofregistersforaCPUstartswith-*"CPUSTRT"andendswith"CPUEND".Eachregisterentryisof16bytes,-*8ByteASCIIidentifierand8Byteregistervalue.Theregisterentry-*withidentifier"CPUSTRT"and"CPUEND"contains4bytecpuidaspart-*ofregistervalue.FormoredetailsrefertoPAPRdocument.-*-*OnlyforthecrashingcpuweignoretheCPUdumpdataandgetexact-*statefromfadumpcrashinfostructurepopulatedbyfirstkernelatthe-*timeofcrash.-*/-staticint__initfadump_build_cpu_notes(conststructrtas_fadump_mem_struct*fdm)-{-structrtas_fadump_reg_save_area_header*reg_header;-structrtas_fadump_reg_entry*reg_entry;-structfadump_crash_info_header*fdh=NULL;-void*vaddr;-unsignedlongaddr;-u32num_cpus,*note_buf;-structpt_regsregs;-inti,rc=0,cpu=0;--if(!fdm->cpu_state_data.bytes_dumped)-return-EINVAL;--addr=be64_to_cpu(fdm->cpu_state_data.destination_address);-vaddr=__va(addr);--reg_header=vaddr;-if(be64_to_cpu(reg_header->magic_number)!=-fadump_str_to_u64("REGSAVE")){-printk(KERN_ERR"Unable to read register save area.\n");-return-ENOENT;-}-pr_debug("--------CPU State Data------------\n");-pr_debug("Magic Number: %llx\n",be64_to_cpu(reg_header->magic_number));-pr_debug("NumCpuOffset: %x\n",be32_to_cpu(reg_header->num_cpu_offset));--vaddr+=be32_to_cpu(reg_header->num_cpu_offset);-num_cpus=be32_to_cpu(*((__be32*)(vaddr)));-pr_debug("NumCpus : %u\n",num_cpus);-vaddr+=sizeof(u32);-reg_entry=(structrtas_fadump_reg_entry*)vaddr;--/* Allocate buffer to hold cpu crash notes. */-fw_dump.cpu_notes_buf_size=num_cpus*sizeof(note_buf_t);-fw_dump.cpu_notes_buf_size=PAGE_ALIGN(fw_dump.cpu_notes_buf_size);-note_buf=fadump_cpu_notes_buf_alloc(fw_dump.cpu_notes_buf_size);-if(!note_buf){-printk(KERN_ERR"Failed to allocate 0x%lx bytes for "-"cpu notes buffer\n",fw_dump.cpu_notes_buf_size);-return-ENOMEM;-}-fw_dump.cpu_notes_buf=__pa(note_buf);--pr_debug("Allocated buffer for cpu notes of size %ld at %p\n",-(num_cpus*sizeof(note_buf_t)),note_buf);--if(fw_dump.fadumphdr_addr)-fdh=__va(fw_dump.fadumphdr_addr);--for(i=0;i<num_cpus;i++){-if(be64_to_cpu(reg_entry->reg_id)!=fadump_str_to_u64("CPUSTRT")){-printk(KERN_ERR"Unable to read CPU state data\n");-rc=-ENOENT;-gotoerror_out;-}-/* Lower 4 bytes of reg_value contains logical cpu id */-cpu=be64_to_cpu(reg_entry->reg_value)&RTAS_FADUMP_CPU_ID_MASK;-if(fdh&&!cpumask_test_cpu(cpu,&fdh->online_mask)){-RTAS_FADUMP_SKIP_TO_NEXT_CPU(reg_entry);-continue;-}-pr_debug("Reading register data for cpu %d...\n",cpu);-if(fdh&&fdh->crashing_cpu==cpu){-regs=fdh->regs;-note_buf=fadump_regs_to_elf_notes(note_buf,®s);-RTAS_FADUMP_SKIP_TO_NEXT_CPU(reg_entry);-}else{-reg_entry++;-reg_entry=fadump_read_registers(reg_entry,®s);-note_buf=fadump_regs_to_elf_notes(note_buf,®s);-}-}-final_note(note_buf);--if(fdh){-addr=fdh->elfcorehdr_addr;-pr_debug("Updating elfcore header(%lx) with cpu notes\n",addr);-fadump_update_elfcore_header(&fw_dump,(char*)__va(addr));-}-return0;--error_out:-fadump_cpu_notes_buf_free((unsignedlong)__va(fw_dump.cpu_notes_buf),-fw_dump.cpu_notes_buf_size);-fw_dump.cpu_notes_buf=0;-fw_dump.cpu_notes_buf_size=0;-returnrc;--}--/*-*Validateandprocessthedumpdatastoredbyfirmwarebeforeexporting-*itthrough'/proc/vmcore'.-*/-staticint__initprocess_fadump(conststructrtas_fadump_mem_struct*fdm_active)-{-structfadump_crash_info_header*fdh;-intrc=0;--if(!fdm_active||!fw_dump.fadumphdr_addr)-return-EINVAL;--/* Check if the dump data is valid. */-if((be16_to_cpu(fdm_active->header.dump_status_flag)==RTAS_FADUMP_ERROR_FLAG)||-(fdm_active->cpu_state_data.error_flags!=0)||-(fdm_active->rmr_region.error_flags!=0)){-printk(KERN_ERR"Dump taken by platform is not valid\n");-return-EINVAL;-}-if((fdm_active->rmr_region.bytes_dumped!=-fdm_active->rmr_region.source_len)||-!fdm_active->cpu_state_data.bytes_dumped){-printk(KERN_ERR"Dump taken by platform is incomplete\n");-return-EINVAL;-}--/* Validate the fadump crash info header */-fdh=__va(fw_dump.fadumphdr_addr);-if(fdh->magic_number!=FADUMP_CRASH_INFO_MAGIC){-printk(KERN_ERR"Crash info header is not valid.\n");-return-EINVAL;-}--rc=fadump_build_cpu_notes(fdm_active);-if(rc)-returnrc;--/*-*Wearedonevalidatingdumpinfoandelfcoreheaderisnowready-*tobeexported.setelfcorehdr_addrsothatvmcoremodulewill-*exporttheelfcoreheaderthrough'/proc/vmcore'.-*/-elfcorehdr_addr=fdh->elfcorehdr_addr;--return0;-}-staticvoidfree_crash_memory_ranges(void){kfree(crash_memory_ranges);
@@ -970,7 +740,6 @@ static unsigned long init_fadump_header(unsigned long addr)if(!addr)return0;-fw_dump.fadumphdr_addr=addr;fdh=__va(addr);addr+=sizeof(structfadump_crash_info_header);
@@ -1014,39 +783,12 @@ static int register_fadump(void)returnfw_dump.ops->register_fadump(&fw_dump);}-staticintfadump_invalidate_dump(conststructrtas_fadump_mem_struct*fdm)-{-intrc=0;-unsignedintwait_time;--pr_debug("Invalidating firmware-assisted dump registration\n");--/* TODO: Add upper time limit for the delay */-do{-rc=rtas_call(fw_dump.ibm_configure_kernel_dump,3,1,NULL,-FADUMP_INVALIDATE,fdm,-sizeof(structrtas_fadump_mem_struct));--wait_time=rtas_busy_delay_time(rc);-if(wait_time)-mdelay(wait_time);-}while(wait_time);--if(rc){-pr_err("Failed to invalidate firmware-assisted dump registration. Unexpected error (%d).\n",rc);-returnrc;-}-fw_dump.dump_active=0;-fdm_active=NULL;-return0;-}-voidfadump_cleanup(void){/* Invalidate the registration only if dump is active. */if(fw_dump.dump_active){-/* pass the same memory dump structure provided by platform */-fadump_invalidate_dump(fdm_active);+pr_debug("Invalidating firmware-assisted dump registration\n");+fw_dump.ops->invalidate_fadump(&fw_dump);}elseif(fw_dump.dump_registered){/* Un-register Firmware-assisted dump if it was registered. */fw_dump.ops->unregister_fadump(&fw_dump);
@@ -40,6 +41,23 @@ static void rtas_fadump_update_config(struct fw_dump *fadump_conf,fadump_conf->fadumphdr_addr=(fadump_conf->boot_mem_dest_addr+fadump_conf->boot_memory_size);++/* Start address of preserve area (permanent reservation) */+fadump_conf->preserv_area_start=+be64_to_cpu(fdm->cpu_state_data.destination_address);+pr_debug("Preserve area start address: 0x%lx\n",+fadump_conf->preserv_area_start);+}++/*+*Thisfunctioniscalledinthecapturekerneltogetconfigurationdetails+*setupinthefirstkernelandpassedtothef/w.+*/+staticvoidrtas_fadump_get_config(structfw_dump*fadump_conf,+conststructrtas_fadump_mem_struct*fdm)+{+fadump_conf->boot_memory_size=be64_to_cpu(fdm->rmr_region.source_len);+rtas_fadump_update_config(fadump_conf,fdm);}staticulongrtas_fadump_init_mem_struct(structfw_dump*fadump_conf)
@@ -180,7 +198,196 @@ static int rtas_fadump_unregister_fadump(struct fw_dump *fadump_conf)staticintrtas_fadump_invalidate_fadump(structfw_dump*fadump_conf){-return-EIO;+intrc;+unsignedintwait_time;++/* TODO: Add upper time limit for the delay */+do{+rc=rtas_call(fadump_conf->ibm_configure_kernel_dump,3,1,+NULL,FADUMP_INVALIDATE,fdm_active,+sizeof(structrtas_fadump_mem_struct));++wait_time=rtas_busy_delay_time(rc);+if(wait_time)+mdelay(wait_time);+}while(wait_time);++if(rc){+pr_err("Failed to invalidate - unexpected error (%d).\n",rc);+return-EIO;+}++fadump_conf->dump_active=0;+fdm_active=NULL;+return0;+}++#define RTAS_FADUMP_GPR_MASK 0xffffff0000000000+staticinlineintrtas_fadump_gpr_index(u64id)+{+inti=-1;+charstr[3];++if((id&RTAS_FADUMP_GPR_MASK)==fadump_str_to_u64("GPR")){+/* get the digits at the end */+id&=~RTAS_FADUMP_GPR_MASK;+id>>=24;+str[2]='\0';+str[1]=id&0xff;+str[0]=(id>>8)&0xff;+if(kstrtoint(str,10,&i))+i=-EINVAL;+if(i>31)+i=-1;+}+returni;+}++voidrtas_fadump_set_regval(structpt_regs*regs,u64reg_id,u64reg_val)+{+inti;++i=rtas_fadump_gpr_index(reg_id);+if(i>=0)+regs->gpr[i]=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("NIA"))+regs->nip=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("MSR"))+regs->msr=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("CTR"))+regs->ctr=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("LR"))+regs->link=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("XER"))+regs->xer=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("CR"))+regs->ccr=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("DAR"))+regs->dar=(unsignedlong)reg_val;+elseif(reg_id==fadump_str_to_u64("DSISR"))+regs->dsisr=(unsignedlong)reg_val;+}++staticstructrtas_fadump_reg_entry*+rtas_fadump_read_regs(structrtas_fadump_reg_entry*reg_entry,+structpt_regs*regs)+{+memset(regs,0,sizeof(structpt_regs));++while(be64_to_cpu(reg_entry->reg_id)!=fadump_str_to_u64("CPUEND")){+rtas_fadump_set_regval(regs,be64_to_cpu(reg_entry->reg_id),+be64_to_cpu(reg_entry->reg_value));+reg_entry++;+}+reg_entry++;+returnreg_entry;+}++/*+*ReadCPUstatedumpdataandconvertitintoELFnotes.+*TheCPUdumpstartswithmagicnumber"REGSAVE".NumCpusOffsetshouldbe+*usedtoaccessthedatatoallowforadditionalfieldstobeaddedwithout+*affectingcompatibility.EachlistofregistersforaCPUstartswith+*"CPUSTRT"andendswith"CPUEND".Eachregisterentryisof16bytes,+*8ByteASCIIidentifierand8Byteregistervalue.Theregisterentry+*withidentifier"CPUSTRT"and"CPUEND"contains4bytecpuidaspart+*ofregistervalue.FormoredetailsrefertoPAPRdocument.+*+*OnlyforthecrashingcpuweignoretheCPUdumpdataandgetexact+*statefromfadumpcrashinfostructurepopulatedbyfirstkernelatthe+*timeofcrash.+*/+staticint__initrtas_fadump_build_cpu_notes(structfw_dump*fadump_conf)+{+structrtas_fadump_reg_save_area_header*reg_header;+structrtas_fadump_reg_entry*reg_entry;+structfadump_crash_info_header*fdh=NULL;+void*vaddr;+unsignedlongaddr;+u32num_cpus,*note_buf;+structpt_regsregs;+inti,rc=0,cpu=0;++addr=be64_to_cpu(fdm_active->cpu_state_data.destination_address);+vaddr=__va(addr);++reg_header=vaddr;+if(be64_to_cpu(reg_header->magic_number)!=+fadump_str_to_u64("REGSAVE")){+pr_err("Unable to read register save area.\n");+return-ENOENT;+}++pr_debug("--------CPU State Data------------\n");+pr_debug("Magic Number: %llx\n",be64_to_cpu(reg_header->magic_number));+pr_debug("NumCpuOffset: %x\n",be32_to_cpu(reg_header->num_cpu_offset));++vaddr+=be32_to_cpu(reg_header->num_cpu_offset);+num_cpus=be32_to_cpu(*((__be32*)(vaddr)));+pr_debug("NumCpus : %u\n",num_cpus);+vaddr+=sizeof(u32);+reg_entry=(structrtas_fadump_reg_entry*)vaddr;++/* Allocate buffer to hold cpu crash notes. */+fadump_conf->cpu_notes_buf_size=num_cpus*sizeof(note_buf_t);+fadump_conf->cpu_notes_buf_size=+PAGE_ALIGN(fadump_conf->cpu_notes_buf_size);+note_buf=fadump_cpu_notes_buf_alloc(fadump_conf->cpu_notes_buf_size);+if(!note_buf){+pr_err("Failed to allocate 0x%lx bytes for cpu notes buffer\n",+fadump_conf->cpu_notes_buf_size);+return-ENOMEM;+}+fadump_conf->cpu_notes_buf=__pa(note_buf);++pr_debug("Allocated buffer for cpu notes of size %ld at %p\n",+(num_cpus*sizeof(note_buf_t)),note_buf);++if(fadump_conf->fadumphdr_addr)+fdh=__va(fadump_conf->fadumphdr_addr);++for(i=0;i<num_cpus;i++){+if(be64_to_cpu(reg_entry->reg_id)!=+fadump_str_to_u64("CPUSTRT")){+pr_err("Unable to read CPU state data\n");+rc=-ENOENT;+gotoerror_out;+}+/* Lower 4 bytes of reg_value contains logical cpu id */+cpu=(be64_to_cpu(reg_entry->reg_value)&+RTAS_FADUMP_CPU_ID_MASK);+if(fdh&&!cpumask_test_cpu(cpu,&fdh->online_mask)){+RTAS_FADUMP_SKIP_TO_NEXT_CPU(reg_entry);+continue;+}+pr_debug("Reading register data for cpu %d...\n",cpu);+if(fdh&&fdh->crashing_cpu==cpu){+regs=fdh->regs;+note_buf=fadump_regs_to_elf_notes(note_buf,®s);+RTAS_FADUMP_SKIP_TO_NEXT_CPU(reg_entry);+}else{+reg_entry++;+reg_entry=rtas_fadump_read_regs(reg_entry,®s);+note_buf=fadump_regs_to_elf_notes(note_buf,®s);+}+}+final_note(note_buf);++if(fdh){+pr_debug("Updating elfcore header (%llx) with cpu notes\n",+fdh->elfcorehdr_addr);+fadump_update_elfcore_header(fadump_conf,+__va(fdh->elfcorehdr_addr));+}+return0;++error_out:+fadump_cpu_notes_buf_free((ulong)__va(fadump_conf->cpu_notes_buf),+fadump_conf->cpu_notes_buf_size);+fadump_conf->cpu_notes_buf=0;+fadump_conf->cpu_notes_buf_size=0;+returnrc;+}/*
@@ -189,15 +396,62 @@ static int rtas_fadump_invalidate_fadump(struct fw_dump *fadump_conf)*/staticint__initrtas_fadump_process_fadump(structfw_dump*fadump_conf){-return-EINVAL;+structfadump_crash_info_header*fdh;+intrc=0;++if(!fdm_active||!fadump_conf->fadumphdr_addr)+return-EINVAL;++/* Check if the dump data is valid. */+if((be16_to_cpu(fdm_active->header.dump_status_flag)==+RTAS_FADUMP_ERROR_FLAG)||+(fdm_active->cpu_state_data.error_flags!=0)||+(fdm_active->rmr_region.error_flags!=0)){+pr_err("Dump taken by platform is not valid\n");+return-EINVAL;+}+if((fdm_active->rmr_region.bytes_dumped!=+fdm_active->rmr_region.source_len)||+!fdm_active->cpu_state_data.bytes_dumped){+pr_err("Dump taken by platform is incomplete\n");+return-EINVAL;+}++/* Validate the fadump crash info header */+fdh=__va(fadump_conf->fadumphdr_addr);+if(fdh->magic_number!=FADUMP_CRASH_INFO_MAGIC){+pr_err("Crash info header is not valid.\n");+return-EINVAL;+}++if(!fdm_active->cpu_state_data.bytes_dumped)+return-EINVAL;++rc=rtas_fadump_build_cpu_notes(fadump_conf);+if(rc)+returnrc;++/*+*Wearedonevalidatingdumpinfoandelfcoreheaderisnowready+*tobeexported.setelfcorehdr_addrsothatvmcoremodulewill+*exporttheelfcoreheaderthrough'/proc/vmcore'.+*/+elfcorehdr_addr=fdh->elfcorehdr_addr;++return0;}staticvoidrtas_fadump_region_show(structfw_dump*fadump_conf,structseq_file*m){-conststructrtas_fadump_mem_struct*fdm_ptr=&fdm;+conststructrtas_fadump_mem_struct*fdm_ptr;conststructrtas_fadump_section*cpu_data_section;+if(fdm_active)+fdm_ptr=fdm_active;+else+fdm_ptr=&fdm;+cpu_data_section=&(fdm_ptr->cpu_state_data);seq_printf(m,"CPU :[%#016llx-%#016llx] %#llx bytes, Dumped: %#llx\n",be64_to_cpu(cpu_data_section->destination_address),
@@ -219,6 +473,12 @@ static void rtas_fadump_region_show(struct fw_dump *fadump_conf,seq_printf(m,"Size: %#llx, Dumped: %#llx bytes\n",be64_to_cpu(fdm_ptr->rmr_region.source_len),be64_to_cpu(fdm_ptr->rmr_region.bytes_dumped));++/* Dump is active. Show reserved area start address. */+if(fdm_active){+seq_printf(m,"\nMemory above %#016lx is reserved for saving crash dump\n",+fadump_conf->reserve_dump_area_start);+}}staticvoidrtas_fadump_trigger(structfadump_crash_info_header*fdh,
@@ -258,6 +519,17 @@ int __init rtas_fadump_dt_scan(struct fw_dump *fadump_conf, ulong node)fadump_conf->fadump_platform=FADUMP_PLATFORM_PSERIES;fadump_conf->fadump_supported=1;+/*+*The'ibm,kernel-dump'rtasnodeispresentonlyifthereis+*dumpdatawaitingforus.+*/+fdm_active=of_get_flat_dt_prop(node,"ibm,kernel-dump",NULL);+if(fdm_active){+pr_info("Firmware-assisted dump is active.\n");+fadump_conf->dump_active=1;+rtas_fadump_get_config(fadump_conf,(void*)__pa(fdm_active));+}+/* Get the sizes required to store dump data for the firmware provided*dumpsections.*Foreachdumpsectiontypesupported,a32bitcellwhichdefines
OPAL allows registering address with it in the first kernel and
retrieving it after MPIPL. Setup kernel metadata and register its
address with OPAL to use it for processing the crash dump.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/kernel/fadump-common.h | 4 +
arch/powerpc/kernel/fadump.c | 65 ++++++++++++++---------
arch/powerpc/platforms/powernv/opal-fadump.c | 73 ++++++++++++++++++++++++++
arch/powerpc/platforms/powernv/opal-fadump.h | 37 +++++++++++++
arch/powerpc/platforms/pseries/rtas-fadump.c | 32 +++++++++--
5 files changed, 177 insertions(+), 34 deletions(-)
create mode 100644 arch/powerpc/platforms/powernv/opal-fadump.h
[...]
quoted hunk
@@ -346,30 +349,42 @@ int __init fadump_reserve_mem(void) * use memblock_find_in_range() here since it doesn't allocate * from bottom to top. */- for (base = fw_dump.boot_memory_size;- base <= (memory_boundary - size);- base += size) {+ while (base <= (memory_boundary - size)) { if (memblock_is_region_memory(base, size) && !memblock_is_region_reserved(base, size)) break;++ base += size; }- if ((base > (memory_boundary - size)) ||- memblock_reserve(base, size)) {++ if (base > (memory_boundary - size)) {+ pr_err("Failed to find memory chunk for reservation\n");+ goto error_out;+ }+ fw_dump.reserve_dump_area_start = base;++ /*+ * Calculate the kernel metadata address and register it with+ * f/w if the platform supports.+ */+ if (fw_dump.ops->setup_kernel_metadata(&fw_dump) < 0)+ goto error_out;
I see setup_kernel_metadata() registers the metadata address with opal without
having any minimum data initialized in it. Secondaly, why can't this wait until
registration ? I think we should defer this until fadump registration.
What if kernel crashes before metadata area is initialized ?
+
+ if (memblock_reserve(base, size)) {
pr_err("Failed to reserve memory\n");
- return 0;
+ goto error_out;
}
Can you make the tab space changes in your previous patch where these
were initially introduced ? So that this patch can only show new members
that are added.
Thanks,
-Mahesh.
Make OPAL calls to register and un-register with firmware for MPIPL.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/platforms/powernv/opal-fadump.c | 71 +++++++++++++++++++++++++-
1 file changed, 69 insertions(+), 2 deletions(-)
[...]
quoted hunk
@@ -88,12 +104,63 @@ static int opal_fadump_setup_kernel_metadata(struct fw_dump *fadump_conf) static int opal_fadump_register_fadump(struct fw_dump *fadump_conf) {- return -EIO;+ int i, err = -EIO;+ s64 rc;++ for (i = 0; i < opal_fdm->region_cnt; i++) {+ rc = opal_mpipl_update(OPAL_MPIPL_ADD_RANGE,+ opal_fdm->rgn[i].src,+ opal_fdm->rgn[i].dest,+ opal_fdm->rgn[i].size);+ if (rc != OPAL_SUCCESS)
You may want to remove ranges which has been added so far on error and reset
opal_fdm->registered_regions.
+ break;
+
+ opal_fdm->registered_regions++;
+ }
+
+ switch (rc) {
+ case OPAL_SUCCESS:
+ pr_info("Registration is successful!\n");
+ fadump_conf->dump_registered = 1;
+ err = 0;
+ break;
+ case OPAL_UNSUPPORTED:
+ pr_err("Support not available.\n");
+ fadump_conf->fadump_supported = 0;
+ fadump_conf->fadump_enabled = 0;
+ break;
+ case OPAL_INTERNAL_ERROR:
+ pr_err("Failed to register. Hardware Error(%lld).\n", rc);
+ break;
+ case OPAL_PARAMETER:
+ pr_err("Failed to register. Parameter Error(%lld).\n", rc);
+ break;
+ case OPAL_PERMISSION:
You may want to remove this check. With latest opal mpipl patches
opal_mpipl_update() no more returns OPAL_PERMISSION.
Even if opal does, we can not say fadump already registered just by
looking at return status of single entry addition.
Thanks,
-Mahesh.
Firmware uses 32-bit field for region size while copying/backing-up
memory during MPIPL. So, the maximum copy size for a region would
be a page less than 4GB (aligned to pagesize) but FADump capture
kernel usually needs more memory than that to be preserved to avoid
running into out of memory errors.
So, request firmware to copy multiple kernel memory regions instead
of just one (which worked fine for pseries as 64-bit field was used
for size there). With support to copy multiple kernel memory regions,
also handle holes in the memory area to be preserved. Support as many
as 128 kernel memory regions. This allows having an adequate FADump
capture kernel size for different scenarios.
Can you split this patch into 2 ? One for handling holes in boot memory
and other for handling 4Gb region size ? So that it will be easy to
review changes.
Thanks,
-Mahesh.
Do we really need these ? Aren't we hiding all platform specific things
under fadump_ops functions ? I see that these values are used only for
assignements and not making any decision in code flow. Am I missing
anything here ?
True. This isn't really useful. will drop it..
Thanks
Hari
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-08-14 06:43:35
On 12/08/19 9:31 PM, Mahesh J Salgaonkar wrote:
On 2019-07-16 17:02:38 Tue, Hari Bathini wrote:
quoted
Make RTAS calls to register and un-register for FADump. Also, update
how fadump_region contents are diplayed to provide more information.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/kernel/fadump-common.h | 2
arch/powerpc/kernel/fadump.c | 164 ++------------------------
arch/powerpc/platforms/pseries/rtas-fadump.c | 163 +++++++++++++++++++++++++-
3 files changed, 176 insertions(+), 153 deletions(-)
[...]
quoted
static int rtas_fadump_register_fadump(struct fw_dump *fadump_conf)
{
- return -EIO;
+ int rc, err = -EIO;
+ unsigned int wait_time;
+
+ /* TODO: Add upper time limit for the delay */
+ do {
+ rc = rtas_call(fadump_conf->ibm_configure_kernel_dump, 3, 1,
+ NULL, FADUMP_REGISTER, &fdm,
+ sizeof(struct rtas_fadump_mem_struct));
+
+ wait_time = rtas_busy_delay_time(rc);
+ if (wait_time)
+ mdelay(wait_time);
+
+ } while (wait_time);
+
+ switch (rc) {
+ case 0:
+ pr_info("Registration is successful!\n");
+ fadump_conf->dump_registered = 1;
+ err = 0;
+ break;
+ case -1:
+ pr_err("Failed to register. Hardware Error(%d).\n", rc);
+ break;
+ case -3:
+ if (!is_fadump_boot_mem_contiguous(fadump_conf))
+ pr_err("Can't hot-remove boot memory area.\n");
+ else if (!is_fadump_reserved_mem_contiguous(fadump_conf))
+ pr_err("Can't hot-remove reserved memory area.\n");
Any reason why we changed the error messages here ? it gives an impression as
if fadump reservation tried to hot remove memory and failed.
Yeah, the message is indeed a bit confusing. Will stick with old message..
Thanks
Hari
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-08-14 07:08:28
On 13/08/19 4:11 PM, Mahesh J Salgaonkar wrote:
On 2019-07-16 17:03:15 Tue, Hari Bathini wrote:
quoted
OPAL allows registering address with it in the first kernel and
retrieving it after MPIPL. Setup kernel metadata and register its
address with OPAL to use it for processing the crash dump.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/kernel/fadump-common.h | 4 +
arch/powerpc/kernel/fadump.c | 65 ++++++++++++++---------
arch/powerpc/platforms/powernv/opal-fadump.c | 73 ++++++++++++++++++++++++++
arch/powerpc/platforms/powernv/opal-fadump.h | 37 +++++++++++++
arch/powerpc/platforms/pseries/rtas-fadump.c | 32 +++++++++--
5 files changed, 177 insertions(+), 34 deletions(-)
create mode 100644 arch/powerpc/platforms/powernv/opal-fadump.h
[...]
quoted
@@ -346,30 +349,42 @@ int __init fadump_reserve_mem(void) * use memblock_find_in_range() here since it doesn't allocate * from bottom to top. */- for (base = fw_dump.boot_memory_size;- base <= (memory_boundary - size);- base += size) {+ while (base <= (memory_boundary - size)) { if (memblock_is_region_memory(base, size) && !memblock_is_region_reserved(base, size)) break;++ base += size; }- if ((base > (memory_boundary - size)) ||- memblock_reserve(base, size)) {++ if (base > (memory_boundary - size)) {+ pr_err("Failed to find memory chunk for reservation\n");+ goto error_out;+ }+ fw_dump.reserve_dump_area_start = base;++ /*+ * Calculate the kernel metadata address and register it with+ * f/w if the platform supports.+ */+ if (fw_dump.ops->setup_kernel_metadata(&fw_dump) < 0)+ goto error_out;
I see setup_kernel_metadata() registers the metadata address with opal without
having any minimum data initialized in it. Secondaly, why can't this wait until> registration ? I think we should defer this until fadump registration.
If setting up metadata address fails (it should ideally not fail, but..), everything else
is useless. So, we might as well try that early and fall back to KDump in case of an error..
What if kernel crashes before metadata area is initialized ?
registered_regions would be '0'. So, it is treated as fadump is not registered case. Let me
initialize metadata explicitly before registering the address with f/w to avoid any assumption...
quoted
+
+ if (memblock_reserve(base, size)) {
pr_err("Failed to reserve memory\n");
- return 0;
+ goto error_out;
}
Can you make the tab space changes in your previous patch where these
were initially introduced ? So that this patch can only show new members
that are added.
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-08-14 07:13:48
On 13/08/19 8:04 PM, Mahesh J Salgaonkar wrote:
On 2019-07-16 17:03:23 Tue, Hari Bathini wrote:
quoted
Make OPAL calls to register and un-register with firmware for MPIPL.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/platforms/powernv/opal-fadump.c | 71 +++++++++++++++++++++++++-
1 file changed, 69 insertions(+), 2 deletions(-)
[...]
quoted
@@ -88,12 +104,63 @@ static int opal_fadump_setup_kernel_metadata(struct fw_dump *fadump_conf) static int opal_fadump_register_fadump(struct fw_dump *fadump_conf) {- return -EIO;+ int i, err = -EIO;+ s64 rc;++ for (i = 0; i < opal_fdm->region_cnt; i++) {+ rc = opal_mpipl_update(OPAL_MPIPL_ADD_RANGE,+ opal_fdm->rgn[i].src,+ opal_fdm->rgn[i].dest,+ opal_fdm->rgn[i].size);+ if (rc != OPAL_SUCCESS)
You may want to remove ranges which has been added so far on error and reset
opal_fdm->registered_regions.
Thanks for catching this, Mahesh.
Will update..
quoted
+ break;
+
+ opal_fdm->registered_regions++;
+ }
+
+ switch (rc) {
+ case OPAL_SUCCESS:
+ pr_info("Registration is successful!\n");
+ fadump_conf->dump_registered = 1;
+ err = 0;
+ break;
+ case OPAL_UNSUPPORTED:
+ pr_err("Support not available.\n");
+ fadump_conf->fadump_supported = 0;
+ fadump_conf->fadump_enabled = 0;
+ break;
+ case OPAL_INTERNAL_ERROR:
+ pr_err("Failed to register. Hardware Error(%lld).\n", rc);
+ break;
+ case OPAL_PARAMETER:
+ pr_err("Failed to register. Parameter Error(%lld).\n", rc);
+ break;
+ case OPAL_PERMISSION:
You may want to remove this check. With latest opal mpipl patches
opal_mpipl_update() no more returns OPAL_PERMISSION.
Even if opal does, we can not say fadump already registered just by
looking at return status of single entry addition.
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-08-14 07:17:49
On 13/08/19 8:33 PM, Mahesh J Salgaonkar wrote:
On 2019-07-16 17:03:30 Tue, Hari Bathini wrote:
quoted
Firmware uses 32-bit field for region size while copying/backing-up
memory during MPIPL. So, the maximum copy size for a region would
be a page less than 4GB (aligned to pagesize) but FADump capture
kernel usually needs more memory than that to be preserved to avoid
running into out of memory errors.
So, request firmware to copy multiple kernel memory regions instead
of just one (which worked fine for pseries as 64-bit field was used
for size there). With support to copy multiple kernel memory regions,
also handle holes in the memory area to be preserved. Support as many
as 128 kernel memory regions. This allows having an adequate FADump
capture kernel size for different scenarios.
Can you split this patch into 2 ? One for handling holes in boot memory
and other for handling 4Gb region size ? So that it will be easy to
review changes.
Sure. Let me split and have the patch that handles holes in boot memory
as the last patch in the series.
Add support in the kernel to process the crash'ed kernel's memory
preserved during MPIPL and export it as /proc/vmcore file for the
userland scripts to filter and analyze it later.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/platforms/powernv/opal-fadump.c | 190 ++++++++++++++++++++++++++
1 file changed, 187 insertions(+), 3 deletions(-)
[...]
+ ret = opal_mpipl_query_tag(OPAL_MPIPL_TAG_KERNEL, &addr);
+ if ((ret != OPAL_SUCCESS) || !addr) {
+ pr_err("Failed to get Kernel metadata (%lld)\n", ret);
+ return 1;
+ }
+
+ addr = be64_to_cpu(addr);
+ pr_debug("Kernel metadata addr: %llx\n", addr);
+
+ opal_fdm_active = __va(addr);
+ r_opal_fdm_active = (void *)addr;
+ if (r_opal_fdm_active->version != OPAL_FADUMP_VERSION) {
+ pr_err("FADump active but version (%u) unsupported!\n",
+ r_opal_fdm_active->version);
+ return 1;
+ }
+
+ /* Kernel regions not registered with f/w for MPIPL */
+ if (r_opal_fdm_active->registered_regions == 0) {
+ opal_fdm_active = NULL;
What about partial dump capture scenario ? What if opal crashes while
kernel was in middle of registering ranges ? We may have partial dump
captured which won't be useful.
e,g. If we have total of 4 ranges to be registered and opal crashes
after successful registration of only 2 ranges with 2 pending, we will get a
partial dump which needs to be ignored.
I think check shuold be comparing registered_regions against total number of
regions. What do you think ?
Thanks,
-Mahesh.
OPAL allows registering address with it in the first kernel and
retrieving it after MPIPL. Setup kernel metadata and register its
address with OPAL to use it for processing the crash dump.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/kernel/fadump-common.h | 4 +
arch/powerpc/kernel/fadump.c | 65 ++++++++++++++---------
arch/powerpc/platforms/powernv/opal-fadump.c | 73 ++++++++++++++++++++++++++
arch/powerpc/platforms/powernv/opal-fadump.h | 37 +++++++++++++
arch/powerpc/platforms/pseries/rtas-fadump.c | 32 +++++++++--
5 files changed, 177 insertions(+), 34 deletions(-)
create mode 100644 arch/powerpc/platforms/powernv/opal-fadump.h
[...]
quoted
@@ -346,30 +349,42 @@ int __init fadump_reserve_mem(void) * use memblock_find_in_range() here since it doesn't allocate * from bottom to top. */- for (base = fw_dump.boot_memory_size;- base <= (memory_boundary - size);- base += size) {+ while (base <= (memory_boundary - size)) { if (memblock_is_region_memory(base, size) && !memblock_is_region_reserved(base, size)) break;++ base += size; }- if ((base > (memory_boundary - size)) ||- memblock_reserve(base, size)) {++ if (base > (memory_boundary - size)) {+ pr_err("Failed to find memory chunk for reservation\n");+ goto error_out;+ }+ fw_dump.reserve_dump_area_start = base;++ /*+ * Calculate the kernel metadata address and register it with+ * f/w if the platform supports.+ */+ if (fw_dump.ops->setup_kernel_metadata(&fw_dump) < 0)+ goto error_out;
I see setup_kernel_metadata() registers the metadata address with opal without
having any minimum data initialized in it. Secondaly, why can't this wait until> registration ? I think we should defer this until fadump registration.
If setting up metadata address fails (it should ideally not fail, but..), everything else
is useless.
That's less likely.. so is true with opal_mpipl_update() as well.
So, we might as well try that early and fall back to KDump in case of an error..
ok. Yeah but not uninitialized metadata.
quoted
What if kernel crashes before metadata area is initialized ?
registered_regions would be '0'. So, it is treated as fadump is not registered case.
Let me
initialize metadata explicitly before registering the address with f/w to avoid any assumption...
Do you want to do that before memblock reservation ? Should we move this
to setup_fadump() ?
Thanks,
-Mahesh.
quoted
quoted
+
+ if (memblock_reserve(base, size)) {
pr_err("Failed to reserve memory\n");
- return 0;
+ goto error_out;
}
Can you make the tab space changes in your previous patch where these
were initially introduced ? So that this patch can only show new members
that are added.
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-08-14 11:13:47
On 14/08/19 3:48 PM, Mahesh J Salgaonkar wrote:
On 2019-07-16 17:03:38 Tue, Hari Bathini wrote:
quoted
Add support in the kernel to process the crash'ed kernel's memory
preserved during MPIPL and export it as /proc/vmcore file for the
userland scripts to filter and analyze it later.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/platforms/powernv/opal-fadump.c | 190 ++++++++++++++++++++++++++
1 file changed, 187 insertions(+), 3 deletions(-)
[...]
quoted
+ ret = opal_mpipl_query_tag(OPAL_MPIPL_TAG_KERNEL, &addr);
+ if ((ret != OPAL_SUCCESS) || !addr) {
+ pr_err("Failed to get Kernel metadata (%lld)\n", ret);
+ return 1;
+ }
+
+ addr = be64_to_cpu(addr);
+ pr_debug("Kernel metadata addr: %llx\n", addr);
+
+ opal_fdm_active = __va(addr);
+ r_opal_fdm_active = (void *)addr;
+ if (r_opal_fdm_active->version != OPAL_FADUMP_VERSION) {
+ pr_err("FADump active but version (%u) unsupported!\n",
+ r_opal_fdm_active->version);
+ return 1;
+ }
+
+ /* Kernel regions not registered with f/w for MPIPL */
+ if (r_opal_fdm_active->registered_regions == 0) {
+ opal_fdm_active = NULL;
What about partial dump capture scenario ? What if opal crashes while
kernel was in middle of registering ranges ? We may have partial dump
captured which won't be useful.
e,g. If we have total of 4 ranges to be registered and opal crashes
after successful registration of only 2 ranges with 2 pending, we will get a
partial dump which needs to be ignored.
I think check shuold be comparing registered_regions against total number of
regions. What do you think ?
Yes, Mahesh.
Taking care of that in 22/25
Thanks
Hari
From: Hari Bathini <redacted>
Firmware provides architected register state data at the time of crash.
Process this data and build CPU notes to append to ELF core.
Signed-off-by: Hari Bathini <redacted>
Signed-off-by: Vasant Hegde <redacted>
---
arch/powerpc/kernel/fadump-common.h | 4 +
arch/powerpc/platforms/powernv/opal-fadump.c | 197 ++++++++++++++++++++++++--
arch/powerpc/platforms/powernv/opal-fadump.h | 39 +++++
3 files changed, 228 insertions(+), 12 deletions(-)
[...]
quoted hunk
@@ -430,6 +577,32 @@ int __init opal_fadump_dt_scan(struct fw_dump *fadump_conf, ulong node) return 1; }+ ret = opal_mpipl_query_tag(OPAL_MPIPL_TAG_CPU, &addr);+ if ((ret != OPAL_SUCCESS) || !addr) {+ pr_err("Failed to get CPU metadata (%lld)\n", ret);+ return 1;+ }++ addr = be64_to_cpu(addr);+ pr_debug("CPU metadata addr: %llx\n", addr);++ opal_cpu_metadata = __va(addr);+ r_opal_cpu_metadata = (void *)addr;+ fadump_conf->cpu_state_data_version =+ be32_to_cpu(r_opal_cpu_metadata->cpu_data_version);+ if (fadump_conf->cpu_state_data_version !=+ HDAT_FADUMP_CPU_DATA_VERSION) {+ pr_err("CPU data format version (%lu) mismatch!\n",+ fadump_conf->cpu_state_data_version);+ return 1;+ }+ fadump_conf->cpu_state_entry_size =+ be32_to_cpu(r_opal_cpu_metadata->cpu_data_size);+ fadump_conf->cpu_state_destination_addr =+ be64_to_cpu(r_opal_cpu_metadata->region[0].dest);+ fadump_conf->cpu_state_data_size =+ be64_to_cpu(r_opal_cpu_metadata->region[0].size);+
opal_fadump_dt_scan isn't the right place to do this. Can you please move above
cpu related data processing to opal_fadump_build_cpu_notes() ?
Thanks,
-Mahesh.
pr_info("Firmware-assisted dump is active.\n");
fadump_conf->dump_active = 1;
opal_fadump_get_config(fadump_conf, r_opal_fdm_active);
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-08-16 02:40:32
On 14/08/19 10:45 PM, Mahesh J Salgaonkar wrote:
On 2019-07-16 17:04:08 Tue, Hari Bathini wrote:
quoted
From: Hari Bathini <redacted>
Firmware provides architected register state data at the time of crash.
Process this data and build CPU notes to append to ELF core.
Signed-off-by: Hari Bathini <redacted>
Signed-off-by: Vasant Hegde <redacted>
---
arch/powerpc/kernel/fadump-common.h | 4 +
arch/powerpc/platforms/powernv/opal-fadump.c | 197 ++++++++++++++++++++++++--
arch/powerpc/platforms/powernv/opal-fadump.h | 39 +++++
3 files changed, 228 insertions(+), 12 deletions(-)
[...]
quoted
@@ -430,6 +577,32 @@ int __init opal_fadump_dt_scan(struct fw_dump *fadump_conf, ulong node) return 1; }+ ret = opal_mpipl_query_tag(OPAL_MPIPL_TAG_CPU, &addr);+ if ((ret != OPAL_SUCCESS) || !addr) {+ pr_err("Failed to get CPU metadata (%lld)\n", ret);+ return 1;+ }++ addr = be64_to_cpu(addr);+ pr_debug("CPU metadata addr: %llx\n", addr);++ opal_cpu_metadata = __va(addr);+ r_opal_cpu_metadata = (void *)addr;+ fadump_conf->cpu_state_data_version =+ be32_to_cpu(r_opal_cpu_metadata->cpu_data_version);+ if (fadump_conf->cpu_state_data_version !=+ HDAT_FADUMP_CPU_DATA_VERSION) {+ pr_err("CPU data format version (%lu) mismatch!\n",+ fadump_conf->cpu_state_data_version);+ return 1;+ }
I think cpu data version mismatch check should still be done early on?
Add a new kernel config option, CONFIG_PRESERVE_FA_DUMP that ensures
that crash data, from previously crash'ed kernel, is preserved. This
helps in cases where FADump is not enabled but the subsequent memory
preserving kernel boot is likely to process this crash data. One
typical usecase for this config option is petitboot kernel.
As OPAL allows registering address with it in the first kernel and
retrieving it after MPIPL, use it to store the top of boot memory.
A kernel that intends to preserve crash data retrieves it and avoids
using memory beyond this address.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/Kconfig | 9 ++
arch/powerpc/include/asm/fadump.h | 9 +-
arch/powerpc/kernel/Makefile | 6 +
arch/powerpc/kernel/fadump-common.h | 13 ++-
arch/powerpc/kernel/fadump.c | 128 ++++++++++++++++----------
arch/powerpc/kernel/prom.c | 4 -
arch/powerpc/platforms/powernv/Makefile | 1
arch/powerpc/platforms/powernv/opal-fadump.c | 59 ++++++++++++
arch/powerpc/platforms/powernv/opal-fadump.h | 3 +
9 files changed, 176 insertions(+), 56 deletions(-)
[...]
quoted hunk
#include "../../kernel/fadump-common.h"
#include "opal-fadump.h"
+
+#ifdef CONFIG_PRESERVE_FA_DUMP
+/*
+ * When dump is active but PRESERVE_FA_DUMP is enabled on the kernel,
+ * ensure crash data is preserved in hope that the subsequent memory
+ * preserving kernel boot is going to process this crash data.
+ */
+int __init opal_fadump_dt_scan(struct fw_dump *fadump_conf, ulong node)
+{
+ unsigned long dn;
+ const __be32 *prop;
+
+ dn = of_get_flat_dt_subnode_by_name(node, "dump");
+ if (dn == -FDT_ERR_NOTFOUND)
+ return 1;
+
+ /*
+ * Check if dump has been initiated on last reboot.
+ */
+ prop = of_get_flat_dt_prop(dn, "mpipl-boot", NULL);
+ if (prop) {
+ u64 addr = 0;
+ s64 ret;
+
+ ret = opal_mpipl_query_tag(OPAL_MPIPL_TAG_BOOT_MEM, &addr);
+ if ((ret != OPAL_SUCCESS) || !addr) {
+ pr_err("Failed to get boot memory tag (%lld)\n", ret);
+ return 1;
+ }
+
+ /*
+ * Anything below this address can be used for booting a
+ * capture kernel or petitboot kernel. Preserve everything
+ * above this address for processing crashdump.
+ */
+ fadump_conf->boot_mem_top = be64_to_cpu(addr);
+ pr_debug("Preserve everything above %lx\n",
+ fadump_conf->boot_mem_top);
+
+ pr_info("Firmware-assisted dump is active.\n");
+ fadump_conf->dump_active = 1;
+ }
+
+ return 1;
+}
+
+#else /* CONFIG_PRESERVE_FA_DUMP */
static const struct opal_fadump_mem_struct *opal_fdm_active;
static const struct opal_mpipl_fadump *opal_cpu_metadata;
static struct opal_fadump_mem_struct *opal_fdm;
@@ -155,6 +202,17 @@ static int opal_fadump_setup_kernel_metadata(struct fw_dump *fadump_conf) err = -EPERM; }+ /*+ * Register boot memory top address with f/w. Should be retrieved+ * by a kernel that intends to preserve crash'ed kernel's memory.+ */+ ret = opal_mpipl_register_tag(OPAL_MPIPL_TAG_BOOT_MEM,+ fadump_conf->boot_mem_top);
Looks like we only register tag but never de-register ot set them to
NULL when we don't need it. Same for kernel TAG. i.e if we kexec into
new kernel which may not do fadump and if opal crashes it will present
stale tags to next kernel. I think we should set bootmem/kernel tag to
NULL in fadump_cleanup() path so that kexec path can be taken care of.
Thanks,
-Mahesh.
If not all kernel boot memory regions are registered for MPIPL before
system crashes, try processing the partial crashdump but warn the user
before proceeding.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
arch/powerpc/platforms/powernv/opal-fadump.c | 21 +++++++++++++++++++++
1 file changed, 21 insertions(+)
@@ -136,6 +136,27 @@ static void opal_fadump_get_config(struct fw_dump *fadump_conf,last_end=base+size;}+/*+*Rarely,butitcansohappenthatsystemcrashesbeforeall+*bootmemoryregionsareregisteredforMPIPL.Insuch+*cases,warnthatthevmcoremaynotbeaccurateandproceed+*anywayasthatisthebestbetconsideringfreepages,cache+*pages,userpages,etcareusuallyfilteredout.+*+*Hopethememorythatcouldnotbepreservedonlyhaspages+*thatareusuallyfilteredoutwhilesavingthevmcore.+*/+if(fdm->region_cnt<fdm->registered_regions){+pr_warn("The crashdump may not be accurate as the below boot memory regions could not be preserved:\n");
This would be opal crashing while kernel is middle of gearing itself for
fadump. If you decide to still go ahead with partial dump then you will need to
have nice warning message about dump capture (makedmpfile capture) may
fail, but we will still have full opal core that can help in analysis.
Thanks,
-Mahesh.
From: Hari Bathini <hbathini@linux.ibm.com> Date: 2019-08-19 15:54:01
On 14/08/19 3:51 PM, Mahesh Jagannath Salgaonkar wrote:
On 8/14/19 12:36 PM, Hari Bathini wrote:
quoted
On 13/08/19 4:11 PM, Mahesh J Salgaonkar wrote:
quoted
On 2019-07-16 17:03:15 Tue, Hari Bathini wrote:
quoted
OPAL allows registering address with it in the first kernel and
retrieving it after MPIPL. Setup kernel metadata and register its
address with OPAL to use it for processing the crash dump.
Signed-off-by: Hari Bathini <hbathini@linux.ibm.com>
---
[...]
quoted
quoted
What if kernel crashes before metadata area is initialized ?
registered_regions would be '0'. So, it is treated as fadump is not registered case.
Let me
initialize metadata explicitly before registering the address with f/w to avoid any assumption...
Do you want to do that before memblock reservation ? Should we move this
to setup_fadump() ?
Better here as failing early would mean we could fall back to KDump..