The following series of patches implement a basic framework
for hypervisor-assisted dump. The very first patch provides
documentation explaining what this is :-) . Yes, its supposed
to be an improvement over kdump.
A list of open issues / todo list is included in the documentation.
It also appears that the not-yet-released firmware versions this was tested
on are still, ahem, incomplete; this work is also pending.
I have included most of the changes requested. Although, I did find
one or two, fixed in a later patch file rather than the first location
they appeared at.
-- Manish & Linas.
@@ -0,0 +1,129 @@++ Hypervisor-Assisted Dump+ ------------------------+ November 2007++The goal of hypervisor-assisted dump is to enable the dump of+a crashed system, and to do so from a fully-reset system, and+to minimize the total elapsed time until the system is back+in production use.++As compared to kdump or other strategies, hypervisor-assisted+dump offers several strong, practical advantages:++-- Unlike kdump, the system has been reset, and loaded+ with a fresh copy of the kernel. In particular,+ PCI and I/O devices have been reinitialized and are+ in a clean, consistent state.+-- As the dump is performed, the dumped memory becomes+ immediately available to the system for normal use.+-- After the dump is completed, no further reboots are+ required; the system will be fully usable, and running+ in it's normal, production mode on it normal kernel.++The above can only be accomplished by coordination with,+and assistance from the hypervisor. The procedure is+as follows:++-- When a system crashes, the hypervisor will save+ the low 256MB of RAM to a previously registered+ save region. It will also save system state, system+ registers, and hardware PTE's.++-- After the low 256MB area has been saved, the+ hypervisor will reset PCI and other hardware state.+ It will *not* clear RAM. It will then launch the+ bootloader, as normal.++-- The freshly booted kernel will notice that there+ is a new node (ibm,dump-kernel) in the device tree,+ indicating that there is crash data available from+ a previous boot. It will boot into only 256MB of RAM,+ reserving the rest of system memory.++-- Userspace tools will parse /sys/kernel/release_region+ and read /proc/vmcore to obtain the contents of memory,+ which holds the previous crashed kernel. The userspace+ tools may copy this info to disk, or network, nas, san,+ iscsi, etc. as desired.++ For Example: the values in /sys/kernel/release-region+ would look something like this (address-range pairs).+ CPU:0x177fee000-0x10000: HPTE:0x177ffe020-0x1000: /+ DUMP:0x177fff020-0x10000000, 0x10000000-0x16F1D370A++-- As the userspace tools complete saving a portion of+ dump, they echo an offset and size to+ /sys/kernel/release_region to release the reserved+ memory back to general use.++ An example of this is:+ "echo 0x40000000 0x10000000 > /sys/kernel/release_region"+ which will release 256MB at the 1GB boundary.++Please note that the hypervisor-assisted dump feature+is only available on Power6-based systems with recent+firmware versions.++Implementation details:+----------------------+In order for this scheme to work, memory needs to be reserved+quite early in the boot cycle. However, access to the device+tree this early in the boot cycle is difficult, and device-tree+access is needed to determine if there is a crash data waiting.+To work around this problem, all but 256MB of RAM is reserved+during early boot. A short while later in boot, a check is made+to determine if there is dump data waiting. If there isn't,+then the reserved memory is released to general kernel use.+If there is dump data, then the /sys/kernel/release_region+file is created, and the reserved memory is held.++If there is no waiting dump data, then all but 256MB of the+reserved ram will be released for general kernel use. The+highest 256 MB of RAM will *not* be released: this region+will be kept permanently reserved, so that it can act as+a receptacle for a copy of the low 256MB in the case a crash+does occur. See, however, "open issues" below, as to whether+such a reserved region is really needed.++Currently the dump will be copied from /proc/vmcore to a+a new file upon user intervention. The starting address+to be read and the range for each data point in provided+in /sys/kernel/release_region.++The tools to examine the dump will be same as the ones+used for kdump.+++General notes:+--------------+Security: please note that there are potential security issues+with any sort of dump mechanism. In particular, plaintext+(unencrypted) data, and possibly passwords, may be present in+the dump data. Userspace tools must take adequate precautions to+preserve security.++Open issues/ToDo:+------------+ o The various code paths that tell the hypervisor that a crash+ occurred, vs. it simply being a normal reboot, should be+ reviewed, and possibly clarified/fixed.++ o Instead of using /sys/kernel, should there be a /sys/dump+ instead? There is a dump_subsys being created by the s390 code,+ perhaps the pseries code should use a similar layout as well.++ o Is reserving a 256MB region really required? The goal of+ reserving a 256MB scratch area is to make sure that no+ important crash data is clobbered when the hypervisor+ save low mem to the scratch area. But, if one could assure+ that nothing important is located in some 256MB area, then+ it would not need to be reserved. Something that can be+ improved in subsequent versions.++ o Still working the kdump team to integrate this with kdump,+ some work remains but this would not affect the current+ patches.++ o Still need to write a shell script, to copy the dump away.+ Currently I am parsing it manually.
Check to see if there actually is data from a previously
crashed kernel waiting. If so, Allow user-sapce tools to
grab the data (by reading /proc/kcore). When user-space
finishes dumping a section, it must release that memory
by writing to sysfs. For example,
echo "0x40000000 0x10000000" > /sys/kernel/release_region
will release 256MB starting at the 1GB. The released memory
becomes free for general use.
Signed-off-by: Linas Vepstas <linasvepstas@gmail.com>
------
arch/powerpc/platforms/pseries/phyp_dump.c | 102 +++++++++++++++++++++++++++--
1 file changed, 96 insertions(+), 6 deletions(-)
Index: 2.6.24-rc5/arch/powerpc/platforms/pseries/phyp_dump.c
===================================================================
@@ -12,17 +12,24 @@*/#include<linux/init.h>+#include<linux/kobject.h>#include<linux/mm.h>+#include<linux/of.h>#include<linux/pfn.h>#include<linux/swap.h>+#include<linux/sysfs.h>#include<asm/page.h>#include<asm/phyp_dump.h>+#include<asm/rtas.h>/* Global, used to communicate data between early boot and late boot */staticstructphyp_dumpphyp_dump_global;structphyp_dump*phyp_dump_info=&phyp_dump_global;+staticintibm_configure_kernel_dump;++/* ------------------------------------------------- *//***release_memory_range--releasememorypreviouslylmb_reserved*@start_pfn:startingphysicalframenumber
@@ -52,20 +59,103 @@ release_memory_range(unsigned long start}}-staticint__initphyp_dump_setup(void)+/* ------------------------------------------------- */+/**+*sysfs_release_region--sysfsinterfacetoreleasememoryrange.+*+*Usage:+*"echo <start addr> <length> > /sys/kernel/release_region"+*+*Example:+*"echo 0x40000000 0x10000000 > /sys/kernel/release_region"+*+*willrelease256MBstartingat1GB.+*/+staticssize_t+store_release_region(structkset*kset,constchar*buf,size_tcount){+unsignedlongstart_addr,length,end_addr;unsignedlongstart_pfn,nr_pages;+ssize_tret;-/* If no memory was reserved in early boot, there is nothing to do */-if(phyp_dump_info->init_reserve_size==0)-return0;+ret=sscanf(buf,"%lx %lx",&start_addr,&length);+if(ret!=2)+return-EINVAL;++/* Range-check - don't free any reserved memory that+*wasn'treservedforphyp-dump*/+if(start_addr<phyp_dump_info->init_reserve_start)+start_addr=phyp_dump_info->init_reserve_start;++end_addr=phyp_dump_info->init_reserve_start++phyp_dump_info->init_reserve_size;+if(start_addr+length>end_addr)+length=end_addr-start_addr;++/* Release the region of memory assed in by user */+start_pfn=PFN_DOWN(start_addr);+nr_pages=PFN_DOWN(length);+release_memory_range(start_pfn,nr_pages);-/* Release memory that was reserved in early boot */+returncount;+}++staticssize_t+show_release_region(structkset*kset,char*buf)+{+returnsprintf(buf,"ola\n");+}++staticstructsubsys_attributerr=__ATTR(release_region,0600,+show_release_region,+store_release_region);++/* ------------------------------------------------- */++staticvoidrelease_all(void)+{+unsignedlongstart_pfn,nr_pages;++/* Release all memory that was reserved in early boot */start_pfn=PFN_DOWN(phyp_dump_info->init_reserve_start);nr_pages=PFN_DOWN(phyp_dump_info->init_reserve_size);release_memory_range(start_pfn,nr_pages);+}++staticint__initphyp_dump_setup(void)+{+structdevice_node*rtas;+constint*dump_header;+intheader_len=0;+intrc;++/* If no memory was reserved in early boot, there is nothing to do */+if(phyp_dump_info->init_reserve_size==0)+return0;++/* Return if phyp dump not supported */+ibm_configure_kernel_dump=rtas_token("ibm,configure-kernel-dump");+if(ibm_configure_kernel_dump==RTAS_UNKNOWN_SERVICE){+release_all();+return-ENOSYS;+}++/* Is there dump data waiting for us? */+rtas=of_find_node_by_path("/rtas");+dump_header=of_get_property(rtas,"ibm,kernel-dump",&header_len);+if(dump_header==NULL){+release_all();+return0;+}++/* Should we create a dump_subsys, analogous to s390/ipl.c ? */+rc=subsys_create_file(&kernel_subsys,&rr);+if(rc){+printk(KERN_ERR"phyp-dump: unable to create sysfs file (%d)\n",rc);+release_all();+return0;+}return0;}-subsys_initcall(phyp_dump_setup);
Initial patch for reserving memory in early boot, and freeing it later.
If the previous boot had ended with a crash, the reserved memory would contain
a copy of the crashed kernel data.
Signed-off-by: Manish Ahuja <redacted>
Signed-off-by: Linas Vepstas <redacted>
----
arch/powerpc/kernel/prom.c | 46 ++++++++++++++++++
arch/powerpc/kernel/rtas.c | 27 +++++++++++
arch/powerpc/platforms/pseries/Makefile | 1
arch/powerpc/platforms/pseries/phyp_dump.c | 71 +++++++++++++++++++++++++++++
include/asm-powerpc/phyp_dump.h | 37 +++++++++++++++
include/asm/rtas.h | 3 +
6 files changed, 185 insertions(+)
Index: 2.6.24-rc5/include/asm-powerpc/phyp_dump.h
===================================================================
@@ -0,0 +1,37 @@+/*+*Hypervisor-assisteddump+*+*LinasVepstas,ManishAhuja2007+*Copyright(c)2007IBMCorp.+*+*Thisprogramisfreesoftware;youcanredistributeitand/or+*modifyitunderthetermsoftheGNUGeneralPublicLicense+*aspublishedbytheFreeSoftwareFoundation;eitherversion+*2oftheLicense,or(atyouroption)anylaterversion.+*/++#ifndef _PPC64_PHYP_DUMP_H+#define _PPC64_PHYP_DUMP_H++#ifdef CONFIG_PHYP_DUMP++/* The RMR region will be saved for later dumping+*wheneverthekernelcrashes.Setthisto256MB.*/+#define PHYP_DUMP_RMR_START 0x0+#define PHYP_DUMP_RMR_END (1UL<<28)++structphyp_dump{+/* Memory that is reserved during very early boot. */+unsignedlonginit_reserve_start;+unsignedlonginit_reserve_size;+/* Check status during boot if dump active & present*/+unsignedlongphyp_dump_is_active;+/* store cpu & hpte size */+unsignedlongcpu_state_size;+unsignedlonghpte_region_size;+};++externstructphyp_dump*phyp_dump_info;++#endif /* CONFIG_PHYP_DUMP */+#endif /* _PPC64_PHYP_DUMP_H */
@@ -0,0 +1,71 @@+/*+*Hypervisor-assisteddump+*+*LinasVepstas,ManishAhuja2007+*Copyrhgit(c)2007IBMCorp.+*+*Thisprogramisfreesoftware;youcanredistributeitand/or+*modifyitunderthetermsoftheGNUGeneralPublicLicense+*aspublishedbytheFreeSoftwareFoundation;eitherversion+*2oftheLicense,or(atyouroption)anylaterversion.+*+*/++#include<linux/init.h>+#include<linux/mm.h>+#include<linux/pfn.h>+#include<linux/swap.h>++#include<asm/page.h>+#include<asm/phyp_dump.h>++/* Global, used to communicate data between early boot and late boot */+staticstructphyp_dumpphyp_dump_global;+structphyp_dump*phyp_dump_info=&phyp_dump_global;++/**+*release_memory_range--releasememorypreviouslylmb_reserved+*@start_pfn:startingphysicalframenumber+*@nr_pages:numberofpagestofree.+*+*Thisroutinewillreleasememorythathadbeenpreviously+*lmb_reservedinearlyboot.Thereleasedmemorybecomes+*availableforgenrealuse.+*/+staticvoid+release_memory_range(unsignedlongstart_pfn,unsignedlongnr_pages)+{+structpage*rpage;+unsignedlongend_pfn;+longi;++end_pfn=start_pfn+nr_pages;++for(i=start_pfn;i<=end_pfn;i++){+rpage=pfn_to_page(i);+if(PageReserved(rpage)){+ClearPageReserved(rpage);+init_page_count(rpage);+__free_page(rpage);+totalram_pages++;+}+}+}++staticint__initphyp_dump_setup(void)+{+unsignedlongstart_pfn,nr_pages;++/* If no memory was reserved in early boot, there is nothing to do */+if(phyp_dump_info->init_reserve_size==0)+return0;++/* Release memory that was reserved in early boot */+start_pfn=PFN_DOWN(phyp_dump_info->init_reserve_start);+nr_pages=PFN_DOWN(phyp_dump_info->init_reserve_size);+release_memory_range(start_pfn,nr_pages);++return0;+}++subsys_initcall(phyp_dump_setup);
@@ -1011,6 +1012,48 @@ static void __init early_reserve_mem(voi#endif}+#ifdef CONFIG_PHYP_DUMP++/**+*reserve_crashed_mem()-reserveallnot-yet-dumpedmmemory+*+*Thisroutinewillreservealmostallofthememoryinthe+*system,exceptforafewhundredmegabytesusedtobootthe+*newkernel.Asthereservedmemoryisdumpedtothedump+*device(byuserlandtools),itwillbefreedandmadeavailable.+*/+staticvoid__initreserve_crashed_mem(void)+{+unsignedlongbase,size;++if(phyp_dump_info->phyp_dump_is_active){+/* Reserve *everything* above RMR. We'll free this real soon.*/+base=PHYP_DUMP_RMR_END;+size=lmb_end_of_DRAM()-base;++/* XXX crashed_ram_end is wrong, since it may be beyond+*thememory_limit,itwillneedtobeadjusted.*/+lmb_reserve(base,size);++phyp_dump_info->init_reserve_start=base;+phyp_dump_info->init_reserve_size=size;+}+else{+size=phyp_dump_info->cpu_state_size++phyp_dump_info->hpte_region_size++PHYP_DUMP_RMR_END;+base=lmb_end_of_DRAM()-size;+printk(KERN_ERR"Manish reserve regular kernel space is %ld %ld\n",base,size);+lmb_reserve(base,size);+phyp_dump_info->init_reserve_start=base;+phyp_dump_info->init_reserve_size=size;+}+}+#else+staticinlinevoid__initreserve_crashed_mem(void){}+#endif /* CONFIG_PHYP_DUMP */++void__initearly_init_devtree(void*params){DBG(" -> early_init_devtree(%p)\n",params);
@@ -1022,6 +1065,8 @@ void __init early_init_devtree(void *par/* Some machines might need RTAS info for debugging, grab it now. */of_scan_flat_dt(early_init_dt_scan_rtas,NULL);#endif+/* scan tree to see if dump occured during last boot */+of_scan_flat_dt(early_init_dt_scan_phyp_dump,NULL);/* Retrieve various informations from the /chosen node of the*device-tree,includingtheplatformtype,initrdlocationand
Initial patch for reserving memory in early boot, and freeing it later.
If the previous boot had ended with a crash, the reserved memory would contain
a copy of the crashed kernel data.
Signed-off-by: Manish Ahuja <redacted>
Signed-off-by: Linas Vepstas <linasvepstas@gmail.com>
----
arch/powerpc/kernel/prom.c | 46 ++++++++++++++++++
arch/powerpc/kernel/rtas.c | 27 +++++++++++
arch/powerpc/platforms/pseries/Makefile | 1
arch/powerpc/platforms/pseries/phyp_dump.c | 71 +++++++++++++++++++++++++++++
include/asm-powerpc/phyp_dump.h | 37 +++++++++++++++
include/asm/rtas.h | 3 +
6 files changed, 185 insertions(+)
Index: 2.6.24-rc5/include/asm-powerpc/phyp_dump.h
===================================================================
@@ -0,0 +1,37 @@+/*+*Hypervisor-assisteddump+*+*LinasVepstas,ManishAhuja2007+*Copyright(c)2007IBMCorp.+*+*Thisprogramisfreesoftware;youcanredistributeitand/or+*modifyitunderthetermsoftheGNUGeneralPublicLicense+*aspublishedbytheFreeSoftwareFoundation;eitherversion+*2oftheLicense,or(atyouroption)anylaterversion.+*/++#ifndef _PPC64_PHYP_DUMP_H+#define _PPC64_PHYP_DUMP_H++#ifdef CONFIG_PHYP_DUMP++/* The RMR region will be saved for later dumping+*wheneverthekernelcrashes.Setthisto256MB.*/+#define PHYP_DUMP_RMR_START 0x0+#define PHYP_DUMP_RMR_END (1UL<<28)++structphyp_dump{+/* Memory that is reserved during very early boot. */+unsignedlonginit_reserve_start;+unsignedlonginit_reserve_size;+/* Check status during boot if dump active & present*/+unsignedlongphyp_dump_is_active;+/* store cpu & hpte size */+unsignedlongcpu_state_size;+unsignedlonghpte_region_size;+};++externstructphyp_dump*phyp_dump_info;++#endif /* CONFIG_PHYP_DUMP */+#endif /* _PPC64_PHYP_DUMP_H */
@@ -0,0 +1,71 @@+/*+*Hypervisor-assisteddump+*+*LinasVepstas,ManishAhuja2007+*Copyrhgit(c)2007IBMCorp.+*+*Thisprogramisfreesoftware;youcanredistributeitand/or+*modifyitunderthetermsoftheGNUGeneralPublicLicense+*aspublishedbytheFreeSoftwareFoundation;eitherversion+*2oftheLicense,or(atyouroption)anylaterversion.+*+*/++#include<linux/init.h>+#include<linux/mm.h>+#include<linux/pfn.h>+#include<linux/swap.h>++#include<asm/page.h>+#include<asm/phyp_dump.h>++/* Global, used to communicate data between early boot and late boot */+staticstructphyp_dumpphyp_dump_global;+structphyp_dump*phyp_dump_info=&phyp_dump_global;++/**+*release_memory_range--releasememorypreviouslylmb_reserved+*@start_pfn:startingphysicalframenumber+*@nr_pages:numberofpagestofree.+*+*Thisroutinewillreleasememorythathadbeenpreviously+*lmb_reservedinearlyboot.Thereleasedmemorybecomes+*availableforgenrealuse.+*/+staticvoid+release_memory_range(unsignedlongstart_pfn,unsignedlongnr_pages)+{+structpage*rpage;+unsignedlongend_pfn;+longi;++end_pfn=start_pfn+nr_pages;++for(i=start_pfn;i<=end_pfn;i++){+rpage=pfn_to_page(i);+if(PageReserved(rpage)){+ClearPageReserved(rpage);+init_page_count(rpage);+__free_page(rpage);+totalram_pages++;+}+}+}++staticint__initphyp_dump_setup(void)+{+unsignedlongstart_pfn,nr_pages;++/* If no memory was reserved in early boot, there is nothing to do */+if(phyp_dump_info->init_reserve_size==0)+return0;++/* Release memory that was reserved in early boot */+start_pfn=PFN_DOWN(phyp_dump_info->init_reserve_start);+nr_pages=PFN_DOWN(phyp_dump_info->init_reserve_size);+release_memory_range(start_pfn,nr_pages);++return0;+}++subsys_initcall(phyp_dump_setup);
@@ -1011,6 +1012,48 @@ static void __init early_reserve_mem(voi#endif}+#ifdef CONFIG_PHYP_DUMP++/**+*reserve_crashed_mem()-reserveallnot-yet-dumpedmmemory+*+*Thisroutinewillreservealmostallofthememoryinthe+*system,exceptforafewhundredmegabytesusedtobootthe+*newkernel.Asthereservedmemoryisdumpedtothedump+*device(byuserlandtools),itwillbefreedandmadeavailable.+*/+staticvoid__initreserve_crashed_mem(void)+{+unsignedlongbase,size;++if(phyp_dump_info->phyp_dump_is_active){+/* Reserve *everything* above RMR. We'll free this real soon.*/+base=PHYP_DUMP_RMR_END;+size=lmb_end_of_DRAM()-base;++/* XXX crashed_ram_end is wrong, since it may be beyond+*thememory_limit,itwillneedtobeadjusted.*/+lmb_reserve(base,size);++phyp_dump_info->init_reserve_start=base;+phyp_dump_info->init_reserve_size=size;+}+else{+size=phyp_dump_info->cpu_state_size++phyp_dump_info->hpte_region_size++PHYP_DUMP_RMR_END;+base=lmb_end_of_DRAM()-size;+printk(KERN_ERR"Manish reserve regular kernel space is %ld %ld\n",base,size);+lmb_reserve(base,size);+phyp_dump_info->init_reserve_start=base;+phyp_dump_info->init_reserve_size=size;+}+}+#else+staticinlinevoid__initreserve_crashed_mem(void){}+#endif /* CONFIG_PHYP_DUMP */++void__initearly_init_devtree(void*params){DBG(" -> early_init_devtree(%p)\n",params);
@@ -1022,6 +1065,8 @@ void __init early_init_devtree(void *par/* Some machines might need RTAS info for debugging, grab it now. */of_scan_flat_dt(early_init_dt_scan_rtas,NULL);#endif+/* scan tree to see if dump occured during last boot */+of_scan_flat_dt(early_init_dt_scan_phyp_dump,NULL);/* Retrieve various informations from the /chosen node of the*device-tree,includingtheplatformtype,initrdlocationand
Set up the actual dump header, register it with the hypervisor.
Signed-off-by: Manish Ahuja <redacted>
Signed-off-by: Linas Vepstas <linasvepstas@gmail.com>
------
arch/powerpc/platforms/pseries/phyp_dump.c | 136 +++++++++++++++++++++++++++--
1 file changed, 129 insertions(+), 7 deletions(-)
Index: 2.6.24-rc5/arch/powerpc/platforms/pseries/phyp_dump.c
===================================================================
@@ -30,6 +30,117 @@ struct phyp_dump *phyp_dump_info = &phypstaticintibm_configure_kernel_dump;/* ------------------------------------------------- */+/* RTAS interfaces to declare the dump regions */++structdump_section{+u32dump_flags;+u16source_type;+u16error_flags;+u64source_address;+u64source_length;+u64length_copied;+u64destination_address;+};++structphyp_dump_header{+u32version;+u16num_of_sections;+u16status;++u32first_offset_section;+u32dump_disk_section;+u64block_num_dd;+u64num_of_blocks_dd;+u32offset_dd;+u32maxtime_to_auto;+/* No dump disk path string used */++structdump_sectioncpu_data;+structdump_sectionhpte_data;+structdump_sectionkernel_data;+};++/* The dump header *must be* in low memory, so .bss it */+staticstructphyp_dump_headerphdr;++#define NUM_DUMP_SECTIONS 3+#define DUMP_HEADER_VERSION 0x1+#define DUMP_REQUEST_FLAG 0x1+#define DUMP_SOURCE_CPU 0x0001+#define DUMP_SOURCE_HPTE 0x0002+#define DUMP_SOURCE_RMO 0x0011++/**+*init_dump_header()-initializetheheaderdeclaringadump+*Returns:lengthofdumpsavearea.+*+*Whenthehypervisorsavescrashedstate,itneedstoput+*itsomewhere.Thedumpheadertellsthehypervisorwhere+*thedatacanbesaved.+*/+staticunsignedlonginit_dump_header(structphyp_dump_header*ph)+{+unsignedlongaddr_offset=0;++/* Set up the dump header */+ph->version=DUMP_HEADER_VERSION;+ph->num_of_sections=NUM_DUMP_SECTIONS;+ph->status=0;++ph->first_offset_section=+(u32)offsetof(structphyp_dump_header,cpu_data);+ph->dump_disk_section=0;+ph->block_num_dd=0;+ph->num_of_blocks_dd=0;+ph->offset_dd=0;++ph->maxtime_to_auto=0;/* disabled */++/* The first two sections are mandatory */+ph->cpu_data.dump_flags=DUMP_REQUEST_FLAG;+ph->cpu_data.source_type=DUMP_SOURCE_CPU;+ph->cpu_data.source_address=0;+ph->cpu_data.source_length=phyp_dump_info->cpu_state_size;+ph->cpu_data.destination_address=addr_offset;+addr_offset+=phyp_dump_info->cpu_state_size;++ph->hpte_data.dump_flags=DUMP_REQUEST_FLAG;+ph->hpte_data.source_type=DUMP_SOURCE_HPTE;+ph->hpte_data.source_address=0;+ph->hpte_data.source_length=phyp_dump_info->hpte_region_size;+ph->hpte_data.destination_address=addr_offset;+addr_offset+=phyp_dump_info->hpte_region_size;++/* This section describes the low kernel region */+ph->kernel_data.dump_flags=DUMP_REQUEST_FLAG;+ph->kernel_data.source_type=DUMP_SOURCE_RMO;+ph->kernel_data.source_address=PHYP_DUMP_RMR_START;+ph->kernel_data.source_length=PHYP_DUMP_RMR_END;+ph->kernel_data.destination_address=addr_offset;+addr_offset+=ph->kernel_data.source_length;++returnaddr_offset;+}++staticvoidregister_dump_area(structphyp_dump_header*ph,unsignedlongaddr)+{+intrc;+ph->cpu_data.destination_address+=addr;+ph->hpte_data.destination_address+=addr;+ph->kernel_data.destination_address+=addr;++do{+rc=rtas_call(ibm_configure_kernel_dump,3,1,NULL,+1,ph,sizeof(structphyp_dump_header));+}while(rtas_busy_delay(rc));++if(rc)+{+printk(KERN_ERR"phyp-dump: unexpected error (%d) on register\n",rc);+}+}++/* ------------------------------------------------- *//***release_memory_range--releasememorypreviouslylmb_reserved*@start_pfn:startingphysicalframenumber
@@ -140,22 +253,31 @@ static int __init phyp_dump_setup(void)return-ENOSYS;}-/* Is there dump data waiting for us? */+/* Is there dump data waiting for us? If there isn't,+*thenregisteranewdumparea,andreleaseallof+*therestofthereservedram.+*+*The/rtas/ibm,kernel-dumprtasnodeispresentonly+*ifthereisdumpdatawaitingforus.+*/rtas=of_find_node_by_path("/rtas");dump_header=of_get_property(rtas,"ibm,kernel-dump",&header_len);+of_node_put(rtas);++dump_area_length=init_dump_header(&phdr);+dump_area_start=phyp_dump_info->init_reserve_start&PAGE_MASK;/* align down */+if(dump_header==NULL){-release_all();+register_dump_area(&phdr,dump_area_start);return0;}/* Should we create a dump_subsys, analogous to s390/ipl.c ? */rc=subsys_create_file(&kernel_subsys,&rr);-if(rc){+if(rc)printk(KERN_ERR"phyp-dump: unable to create sysfs file (%d)\n",rc);-release_all();-return0;-}+/* ToDo: re-register the dump area, for next time. */return0;}subsys_initcall(phyp_dump_setup);
@@ -122,6 +122,61 @@ static unsigned long init_dump_header(streturnaddr_offset;}+staticvoidprint_dump_header(conststructphyp_dump_header*ph)+{+#ifdef DEBUG+printk(KERN_INFO"dump header:\n");+/* setup some ph->sections required */+printk(KERN_INFO"version = %d\n",ph->version);+printk(KERN_INFO"Sections = %d\n",ph->num_of_sections);+printk(KERN_INFO"Status = 0x%x\n",ph->status);++/* No ph->disk, so all should be set to 0 */+printk(KERN_INFO"Offset to first section 0x%x\n",+ph->first_offset_section);+printk(KERN_INFO"dump disk sections should be zero\n");+printk(KERN_INFO"dump disk section = %d\n",ph->dump_disk_section);+printk(KERN_INFO"block num = %ld\n",ph->block_num_dd);+printk(KERN_INFO"number of blocks = %ld\n",ph->num_of_blocks_dd);+printk(KERN_INFO"dump disk offset = %d\n",ph->offset_dd);+printk(KERN_INFO"Max auto time= %d\n",ph->maxtime_to_auto);++/*set cpu state and hpte states as well scratch pad area */+printk(KERN_INFO" CPU AREA \n");+printk(KERN_INFO"cpu dump_flags =%d\n",ph->cpu_data.dump_flags);+printk(KERN_INFO"cpu source_type =%d\n",ph->cpu_data.source_type);+printk(KERN_INFO"cpu error_flags =%d\n",ph->cpu_data.error_flags);+printk(KERN_INFO"cpu source_address =%lx\n",+ph->cpu_data.source_address);+printk(KERN_INFO"cpu source_length =%lx\n",+ph->cpu_data.source_length);+printk(KERN_INFO"cpu length_copied =%lx\n",+ph->cpu_data.length_copied);++printk(KERN_INFO" HPTE AREA \n");+printk(KERN_INFO"HPTE dump_flags =%d\n",ph->hpte_data.dump_flags);+printk(KERN_INFO"HPTE source_type =%d\n",ph->hpte_data.source_type);+printk(KERN_INFO"HPTE error_flags =%d\n",ph->hpte_data.error_flags);+printk(KERN_INFO"HPTE source_address =%lx\n",+ph->hpte_data.source_address);+printk(KERN_INFO"HPTE source_length =%lx\n",+ph->hpte_data.source_length);+printk(KERN_INFO"HPTE length_copied =%lx\n",+ph->hpte_data.length_copied);++printk(KERN_INFO" SRSD AREA \n");+printk(KERN_INFO"SRSD dump_flags =%d\n",ph->kernel_data.dump_flags);+printk(KERN_INFO"SRSD source_type =%d\n",ph->kernel_data.source_type);+printk(KERN_INFO"SRSD error_flags =%d\n",ph->kernel_data.error_flags);+printk(KERN_INFO"SRSD source_address =%lx\n",+ph->kernel_data.source_address);+printk(KERN_INFO"SRSD source_length =%lx\n",+ph->kernel_data.source_length);+printk(KERN_INFO"SRSD length_copied =%lx\n",+ph->kernel_data.length_copied);+#endif+}+staticvoidregister_dump_area(structphyp_dump_header*ph,unsignedlongaddr){intrc;
@@ -134,9 +189,9 @@ static void register_dump_area(struct ph1,ph,sizeof(structphyp_dump_header));}while(rtas_busy_delay(rc));-if(rc)-{-printk(KERN_ERR"phyp-dump: unexpected error (%d) on register\n",rc);+if(rc){+printk(KERN_ERR"phyp-dump: unexpected error (%d) on register\n",rc);+print_dump_header(ph);}}
@@ -249,6 +304,7 @@ static int __init phyp_dump_setup(void)release_all();return-ENOSYS;}+print_dump_header(dump_header);/* Is there dump data waiting for us? If there isn't,*thenregisteranewdumparea,andreleaseallof
Routines to invalidate and unregister dump routines.
Unregister has not been used yet, I will release another
patch for that at a later stage with the kdump integration patches.
There is also a routine which calculates the regions to be
freed and exports that through sysfs.
Signed-off-by: Manish Ahuja <redacted>
-----
---
arch/powerpc/platforms/pseries/phyp_dump.c | 101 +++++++++++++++++++++++++----
include/asm/phyp_dump.h | 3
2 files changed, 93 insertions(+), 11 deletions(-)
Index: 2.6.24-rc5/arch/powerpc/platforms/pseries/phyp_dump.c
===================================================================
@@ -180,9 +184,15 @@ static void print_dump_header(const strustaticvoidregister_dump_area(structphyp_dump_header*ph,unsignedlongaddr){intrc;-ph->cpu_data.destination_address+=addr;-ph->hpte_data.destination_address+=addr;-ph->kernel_data.destination_address+=addr;++/* Add addr value if not initialized before */+if(ph->cpu_data.destination_address==0){+ph->cpu_data.destination_address+=addr;+ph->hpte_data.destination_address+=addr;+ph->kernel_data.destination_address+=addr;+}++/* ToDo Invalidate kdump and free memory range. */do{rc=rtas_call(ibm_configure_kernel_dump,3,1,NULL,
@@ -195,6 +205,46 @@ static void register_dump_area(struct ph}}+static+voidinvalidate_last_dump(structphyp_dump_header*ph,unsignedlongaddr)+{+intrc;++/* Add addr value if not initialized before */+if(ph->cpu_data.destination_address==0){+ph->cpu_data.destination_address+=addr;+ph->hpte_data.destination_address+=addr;+ph->kernel_data.destination_address+=addr;+}++do{+rc=rtas_call(ibm_configure_kernel_dump,3,1,NULL,+2,ph,sizeof(structphyp_dump_header));+}while(rtas_busy_delay(rc));++if(rc){+printk(KERN_ERR"phyp-dump: unexpected error (%d) "+"on invalidate\n",rc);+print_dump_header(ph);+}+}++staticvoidunregister_dump_area(structphyp_dump_header*ph)+{+intrc;++do{+rc=rtas_call(ibm_configure_kernel_dump,3,1,NULL,+3,ph,sizeof(structphyp_dump_header));+}while(rtas_busy_delay(rc));++if(rc){+printk(KERN_ERR"phyp-dump: unexpected error (%d) "+"on unregister\n",rc);+print_dump_header(ph);+}+}+/* ------------------------------------------------- *//***release_memory_range--releasememorypreviouslylmb_reserved
@@ -237,8 +287,8 @@ release_memory_range(unsigned long start**willrelease256MBstartingat1GB.*/-staticssize_t-store_release_region(structkset*kset,constchar*buf,size_tcount)+static+ssize_tstore_release_region(structkset*kset,constchar*buf,size_tcount){unsignedlongstart_addr,length,end_addr;unsignedlongstart_pfn,nr_pages;
@@ -266,10 +316,23 @@ store_release_region(struct kset *kset, returncount;}-staticssize_t-show_release_region(structkset*kset,char*buf)+staticssize_tshow_release_region(structkset*kset,char*buf){-returnsprintf(buf,"ola\n");+u64second_addr_range;++/* total reserved size - start of scratch area */+second_addr_range=phyp_dump_info->init_reserve_size-+phyp_dump_info->reserved_scratch_size;+returnsprintf(buf,"CPU:0x%lx-0x%lx: HPTE:0x%lx-0x%lx:"+" DUMP:0x%lx-0x%lx, 0x%lx-0x%lx:\n",+phdr.cpu_data.destination_address,+phdr.cpu_data.length_copied,+phdr.hpte_data.destination_address,+phdr.hpte_data.length_copied,+phdr.kernel_data.destination_address,+phdr.kernel_data.length_copied,+phyp_dump_info->init_reserve_start,+second_addr_range);}staticstructsubsys_attributerr=__ATTR(release_region,0600,
@@ -307,7 +370,6 @@ static int __init phyp_dump_setup(void)release_all();return-ENOSYS;}-print_dump_header(dump_header);/* Is there dump data waiting for us? If there isn't,*thenregisteranewdumparea,andreleaseallof
@@ -319,6 +381,7 @@ static int __init phyp_dump_setup(void)rtas=of_find_node_by_path("/rtas");dump_header=of_get_property(rtas,"ibm,kernel-dump",&header_len);of_node_put(rtas);+print_dump_header(dump_header);dump_area_length=init_dump_header(&phdr);dump_area_start=phyp_dump_info->init_reserve_start&PAGE_MASK;/* align down */
@@ -328,6 +391,22 @@ static int __init phyp_dump_setup(void)return0;}+/* re-register the dump area, if old dump was invalid */+if((dump_header)&&(dump_header->status&DUMP_ERROR_FLAG)){+invalidate_last_dump(&phdr,dump_area_start);+register_dump_area(&phdr,dump_area_start);+return0;+}++if(dump_header){+phyp_dump_info->reserved_scratch_addr=+dump_header->cpu_data.destination_address;+phyp_dump_info->reserved_scratch_size=+dump_header->cpu_data.source_length++dump_header->hpte_data.source_length++dump_header->kernel_data.source_length;+}+/* Should we create a dump_subsys, analogous to s390/ipl.c ? */rc=subsys_create_file(&kernel_subsys,&rr);if(rc)
This patch tracks the size freed. For now it does a simple
rudimentary calculation of the ranges freed. The idea is
to keep it simple at the external shell script level and
send in large chunks for now.
Signed-off-by: Manish Ahuja <redacted>
-----
---
arch/powerpc/platforms/pseries/phyp_dump.c | 35 +++++++++++++++++++++++++++++
1 file changed, 35 insertions(+)
Index: 2.6.24-rc5/arch/powerpc/platforms/pseries/phyp_dump.c
===================================================================
From: Paul Mackerras <hidden> Date: 2008-02-07 00:42:05
Manish Ahuja writes:
Initial patch for reserving memory in early boot, and freeing it later.
If the previous boot had ended with a crash, the reserved memory would contain
a copy of the crashed kernel data.
[snip]
+static void __init reserve_crashed_mem(void)
+{
+ unsigned long base, size;
+
+ if (phyp_dump_info->phyp_dump_is_active) {
+ /* Reserve *everything* above RMR. We'll free this real soon.*/
+ base = PHYP_DUMP_RMR_END;
+ size = lmb_end_of_DRAM() - base;
+
+ /* XXX crashed_ram_end is wrong, since it may be beyond
+ * the memory_limit, it will need to be adjusted. */
+ lmb_reserve(base, size);
+
+ phyp_dump_info->init_reserve_start = base;
+ phyp_dump_info->init_reserve_size = size;
+ }
+ else {
+ size = phyp_dump_info->cpu_state_size +
+ phyp_dump_info->hpte_region_size +
+ PHYP_DUMP_RMR_END;
+ base = lmb_end_of_DRAM() - size;
+ printk(KERN_ERR "Manish reserve regular kernel space is %ld %ld\n", base, size);
+ lmb_reserve(base, size);
This is still reserving memory even on systems that aren't running on
pHyp at all. Please rework this so that no memory is reserved if the
system doesn't support phyp-assisted dump.
Paul.
Sorry,
I think i sent the wrong patch file, it shouldn't have my printk statement in there. Let me re-send
the correct file and let me test it once more to make sure it does the right thing.
-Manish
Paul Mackerras wrote:
Manish Ahuja writes:
quoted
Initial patch for reserving memory in early boot, and freeing it later.
If the previous boot had ended with a crash, the reserved memory would contain
a copy of the crashed kernel data.
[snip]
quoted
+static void __init reserve_crashed_mem(void)
+{
+ unsigned long base, size;
+
+ if (phyp_dump_info->phyp_dump_is_active) {
+ /* Reserve *everything* above RMR. We'll free this real soon.*/
+ base = PHYP_DUMP_RMR_END;
+ size = lmb_end_of_DRAM() - base;
+
+ /* XXX crashed_ram_end is wrong, since it may be beyond
+ * the memory_limit, it will need to be adjusted. */
+ lmb_reserve(base, size);
+
+ phyp_dump_info->init_reserve_start = base;
+ phyp_dump_info->init_reserve_size = size;
+ }
+ else {
+ size = phyp_dump_info->cpu_state_size +
+ phyp_dump_info->hpte_region_size +
+ PHYP_DUMP_RMR_END;
+ base = lmb_end_of_DRAM() - size;
+ printk(KERN_ERR "Manish reserve regular kernel space is %ld %ld\n", base, size);
+ lmb_reserve(base, size);
This is still reserving memory even on systems that aren't running on
pHyp at all. Please rework this so that no memory is reserved if the
system doesn't support phyp-assisted dump.
Paul.