Hi,
Please find the patch set that performs the machine check handling inside linux
host. The design is to be able to handle re-entrancy so that we do not clobber
the machine check information during nested machine check interrupt.
The patch 2 introduces separate emergency stack in paca structure exclusively
for machine check exception handling. Patch 3 implements the logic to save the
raw MCE info onto the emergency stack and prepares to take another exception.
Patch 4 implements the additional checks that helps to decide whether to
deliver machine check event to host kernel right away or queue it up and
return. Patch 5 and 6 adds CPU-side hooks for early machine check handler
and TLB flush. The patch 7 and 8 is responsible to detect SLB/TLB errors and
flush them off in the real mode. The patch 9 implements the logic to decode
and save high level MCE information to per cpu buffer without clobbering.
Patch 10 implements mechanism to queue up MCE event in cases where early
handler can not deliver the event to host kernel right away. The patch 12
adds the basic error handling to the high level C code with MMU on.
I have tested SLB multihit, MC coming from opal context on powernv.
Please review and let me know your comments.
Changes in v3:
- Rebased to v3.11-rc7
- Handle MCE coming from opal context, secondary thread nap and return
from interrupt. Queue up the MCE event in this scenario and log it
later during syscall exit path. (patch 4 and 10).
Changes in v2:
- Moved early machine check handling code under CPU_FTR_HVMODE section.
This makes sure that the early machine check handler will get executed
only in hypervisor kernel.
- Add dedicated emergency stack for machine check so that we don't end up
disturbing others who use same emergency stack.
- Fixed the machine check early handle where it used to assume that r1 always
contains the valid stack pointer.
- Fixed an issue where per-cpu mce_nest_count variable underflows when kvm
fails to handle MC error and exit the guest.
- Fixed the code to restore r13 while before exiting early handler.
Thanks,
-Mahesh.
---
Mahesh Salgaonkar (12):
powerpc/book3s: Split the common exception prolog logic into two section.
powerpc/book3s: Introduce exclusive emergency stack for machine check exception.
powerpc/book3s: handle machine check in Linux host.
Validate r1 value before going to host kernel in virtual mode.
powerpc/book3s: Introduce a early machine check hook in cpu_spec.
powerpc/book3s: Add flush_tlb operation in cpu_spec.
powerpc/book3s: Flush SLB/TLBs if we get SLB/TLB machine check errors on power7.
powerpc/book3s: Flush SLB/TLBs if we get SLB/TLB machine check errors on power8.
powerpc/book3s: Decode and save machine check event.
Queue up and process delayed MCE events.
powerpc/powernv: Remove machine check handling in OPAL.
powerpc/powernv: Machine check exception handling.
arch/powerpc/include/asm/bitops.h | 5
arch/powerpc/include/asm/cputable.h | 12 +
arch/powerpc/include/asm/exception-64s.h | 67 +++---
arch/powerpc/include/asm/mce.h | 198 +++++++++++++++++
arch/powerpc/include/asm/paca.h | 9 +
arch/powerpc/kernel/Makefile | 1
arch/powerpc/kernel/asm-offsets.c | 4
arch/powerpc/kernel/cpu_setup_power.S | 38 ++-
arch/powerpc/kernel/cputable.c | 16 +
arch/powerpc/kernel/entry_64.S | 5
arch/powerpc/kernel/exceptions-64s.S | 181 ++++++++++++++++
arch/powerpc/kernel/mce.c | 345 ++++++++++++++++++++++++++++++
arch/powerpc/kernel/mce_power.c | 287 +++++++++++++++++++++++++
arch/powerpc/kernel/setup_64.c | 10 +
arch/powerpc/kernel/traps.c | 15 +
arch/powerpc/kvm/book3s_hv_ras.c | 50 ++--
arch/powerpc/platforms/powernv/opal.c | 161 ++++----------
arch/powerpc/xmon/xmon.c | 4
18 files changed, 1228 insertions(+), 180 deletions(-)
create mode 100644 arch/powerpc/include/asm/mce.h
create mode 100644 arch/powerpc/kernel/mce.c
create mode 100644 arch/powerpc/kernel/mce_power.c
--
-Mahesh
From: Mahesh Salgaonkar <redacted>
This patch splits the common exception prolog logic into two parts to
facilitate reuse of existing code in the next patch. The second part will
be reused in the machine check exception routine in the next patch.
Please note that this patch does not introduce or change existing code
logic. Instead it is just a code movement.
Signed-off-by: Mahesh Salgaonkar <redacted>
---
arch/powerpc/include/asm/exception-64s.h | 67 ++++++++++++++++--------------
1 file changed, 35 insertions(+), 32 deletions(-)
@@ -248,6 +248,40 @@ do_kvm_##n: \#define NOTEST(n)+#define EXCEPTION_PROLOG_COMMON_2(n, area) \+stdr2,GPR2(r1);/* save r2 in stackframe */\+SAVE_4GPRS(3,r1);/* save r3 - r6 in stackframe */\+SAVE_2GPRS(7,r1);/* save r7, r8 in stackframe */\+ldr9,area+EX_R9(r13);/* move r9, r10 to stackframe */\+ldr10,area+EX_R10(r13);\+stdr9,GPR9(r1);\+stdr10,GPR10(r1);\+ldr9,area+EX_R11(r13);/* move r11 - r13 to stackframe */\+ldr10,area+EX_R12(r13);\+ldr11,area+EX_R13(r13);\+stdr9,GPR11(r1);\+stdr10,GPR12(r1);\+stdr11,GPR13(r1);\+BEGIN_FTR_SECTION_NESTED(66);\+ldr10,area+EX_CFAR(r13);\+stdr10,ORIG_GPR3(r1);\+END_FTR_SECTION_NESTED(CPU_FTR_CFAR,CPU_FTR_CFAR,66);\+GET_LR(r9,area);/* Get LR, later save to stack */\+ldr2,PACATOC(r13);/* get kernel TOC into r2 */\+stdr9,_LINK(r1);\+mfctrr10;/* save CTR in stackframe */\+stdr10,_CTR(r1);\+lbzr10,PACASOFTIRQEN(r13);\+mfsprr11,SPRN_XER;/* save XER in stackframe */\+stdr10,SOFTE(r1);\+stdr11,_XER(r1);\+lir9,(n)+1;\+stdr9,_TRAP(r1);/* set trap number */\+lir10,0;\+ldr11,exception_marker@toc(r2);\+stdr10,RESULT(r1);/* clear regs->result */\+stdr11,STACK_FRAME_OVERHEAD-16(r1);/* mark the frame */+/**Thecommonexceptionprologisusedforallexceptafewexceptions*suchasasegmentmissonakerneladdress.Wehavetobeprepared
@@ -281,38 +315,7 @@ do_kvm_##n: \beq4f;/* if from kernel mode */\ACCOUNT_CPU_USER_ENTRY(r9,r10);\SAVE_PPR(area,r9,r10);\-4:stdr2,GPR2(r1);/* save r2 in stackframe */\-SAVE_4GPRS(3,r1);/* save r3 - r6 in stackframe */\-SAVE_2GPRS(7,r1);/* save r7, r8 in stackframe */\-ldr9,area+EX_R9(r13);/* move r9, r10 to stackframe */\-ldr10,area+EX_R10(r13);\-stdr9,GPR9(r1);\-stdr10,GPR10(r1);\-ldr9,area+EX_R11(r13);/* move r11 - r13 to stackframe */\-ldr10,area+EX_R12(r13);\-ldr11,area+EX_R13(r13);\-stdr9,GPR11(r1);\-stdr10,GPR12(r1);\-stdr11,GPR13(r1);\-BEGIN_FTR_SECTION_NESTED(66);\-ldr10,area+EX_CFAR(r13);\-stdr10,ORIG_GPR3(r1);\-END_FTR_SECTION_NESTED(CPU_FTR_CFAR,CPU_FTR_CFAR,66);\-GET_LR(r9,area);/* Get LR, later save to stack */\-ldr2,PACATOC(r13);/* get kernel TOC into r2 */\-stdr9,_LINK(r1);\-mfctrr10;/* save CTR in stackframe */\-stdr10,_CTR(r1);\-lbzr10,PACASOFTIRQEN(r13);\-mfsprr11,SPRN_XER;/* save XER in stackframe */\-stdr10,SOFTE(r1);\-stdr11,_XER(r1);\-lir9,(n)+1;\-stdr9,_TRAP(r1);/* set trap number */\-lir10,0;\-ldr11,exception_marker@toc(r2);\-stdr10,RESULT(r1);/* clear regs->result */\-stdr11,STACK_FRAME_OVERHEAD-16(r1);/* mark the frame */\+4:EXCEPTION_PROLOG_COMMON_2(n,area)\ACCOUNT_STOLEN_TIME/*
From: Mahesh Salgaonkar <redacted>
This patch introduces exclusive emergency stack for machine check exception.
We use emergency stack to handle machine check exception so that we can save
MCE information (srr1, srr0, dar and dsisr) before turning on ME bit and be
ready for re-entrancy. This helps us to prevent clobbering of MCE information
in case of nested machine checks.
The reason for using emergency stack over normal kernel stack is that the
machine check might occur in the middle of setting up a stack frame which may
result into improper use of kernel stack.
Signed-off-by: Mahesh Salgaonkar <redacted>
---
arch/powerpc/include/asm/paca.h | 9 +++++++++
arch/powerpc/kernel/setup_64.c | 10 +++++++++-
arch/powerpc/xmon/xmon.c | 4 ++++
3 files changed, 22 insertions(+), 1 deletion(-)
From: Mahesh Salgaonkar <redacted>
Move machine check entry point into Linux. So far we were dependent on
firmware to decode MCE error details and handover the high level info to OS.
This patch introduces early machine check routine that saves the MCE
information (srr1, srr0, dar and dsisr) to the emergency stack. We allocate
stack frame on emergency stack and set the r1 accordingly. This allows us to be
prepared to take another exception without loosing context. One thing to note
here that, if we get another machine check while ME bit is off then we risk a
checkstop. Hence we restrict ourselves to save only MCE information and turn
the ME bit on. We use paca->in_mce flag to differentiate between first entry
and nested machine check entry which helps proper use of emergency stack. We
increment paca->in_mce every time we enter in early machine check handler and
decrement it while leaving. When we enter machine check early handler first
time (paca->in_mce == 0), we are sure nobody is using MC emergency stack and
allocate a stack frame at the start of the emergency stack. During subsequent
entry (paca->in_mce > 0), we know that r1 points inside emergency stack and we
allocate separate stack frame accordingly. This prevents us from clobbering MCE
information during nested machine checks.
The early machine check handler changes are placed under CPU_FTR_HVMODE
section. This makes sure that the early machine check handler will get executed
only in hypervisor kernel.
This is the code flow:
Machine Check Interrupt
|
V
0x200 vector ME=0, IR=0, DR=0
|
V
+-----------------------------------------------+
|machine_check_pSeries_early: | ME=0, IR=0, DR=0
| Alloc frame on emergency stack |
| Save srr1, srr0, dar and dsisr on stack |
+-----------------------------------------------+
|
(ME=1, IR=0, DR=0, RFID)
|
V
machine_check_handle_early ME=1, IR=0, DR=0
|
V
+-----------------------------------------------+
| machine_check_early (r3=pt_regs) | ME=1, IR=0, DR=0
| Things to do: (in next patches) |
| Flush SLB for SLB errors |
| Flush TLB for TLB errors |
| Decode and save MCE info |
+-----------------------------------------------+
|
(Fall through existing exception handler routine.)
|
V
machine_check_pSerie ME=1, IR=0, DR=0
|
(ME=1, IR=1, DR=1, RFID)
|
V
machine_check_common ME=1, IR=1, DR=1
.
.
.
Signed-off-by: Mahesh Salgaonkar <redacted>
---
arch/powerpc/kernel/asm-offsets.c | 4 +
arch/powerpc/kernel/exceptions-64s.S | 109 ++++++++++++++++++++++++++++++++++
arch/powerpc/kernel/traps.c | 12 ++++
3 files changed, 125 insertions(+)
@@ -404,6 +408,61 @@ denorm_exception_hv:.align7/*movedfrom0x200*/+machine_check_pSeries_early:+BEGIN_FTR_SECTION+EXCEPTION_PROLOG_1(PACA_EXMC,NOTEST,0x200)+/*+*Registercontents:+*R12=interruptvector+*R13=PACA+*R9=CR+*R11&R12issavedonPACA_EXMC+*+*Switchtomc_emergencystackandhandlere-entrancy (thoughwe+*currentlydon't test for overflow). Save MCE registers srr1,+*srr0,daranddsisrandthensetME=1+*+*Weusepaca->in_mcetocheckwhetherthisisthefirstentryor+*nestedmachinecheck.Weincrementpaca->in_mcetotracknested+*machinechecks.+*+*Ifthisisthefirstentrythensetstackpointerto+*paca->mc_emergency_sp,otherwiser1isalreadypointingto+*stackframeonmc_emergencystack.+*+*NOTE:WeareherewithMSR_ME=0(off),whichmeansweriska+*checkstopifwegetanothermachinecheckexceptionbeforewedo+*rfidwithMSR_ME=1.+*/+mrr11,r1/*Saver1*/+lhzr10,PACA_IN_MCE(r13)+cmpwir10,0/*Areweinnestedmachinecheck*/+bne0f/*Yes,weare.*/+/*Firstmachinecheckentry*/+ldr1,PACAMCEMERGSP(r13)/*UseMCemergencystack*/+0:subir1,r1,INT_FRAME_SIZE/*allocstackframe*/+addir10,r10,1/*incrementpaca->in_mce*/+sthr10,PACA_IN_MCE(r13)+stdr11,GPR1(r1)/*Saver1onthestack.*/+stdr11,0(r1)/*makestackchainpointer*/+mfsprr11,SPRN_SRR0/*SaveSRR0*/+stdr11,_NIP(r1)+mfsprr11,SPRN_SRR1/*SaveSRR1*/+stdr11,_MSR(r1)+mfsprr11,SPRN_DAR/*SaveDAR*/+stdr11,_DAR(r1)+mfsprr11,SPRN_DSISR/*SaveDSISR*/+stdr11,_DSISR(r1)+mfmsrr11/*getMSRvalue*/+orir11,r11,MSR_ME/*turnonMEbit*/+ldr12,PACAKBASE(r13)/*gethighpartof&label*/+LOAD_HANDLER(r12,machine_check_handle_early)+mtsprSPRN_SRR0,r12+mtsprSPRN_SRR1,r11+rfid+b./*preventspeculativeexecution*/+END_FTR_SECTION_IFSET(CPU_FTR_HVMODE)+machine_check_pSeries:.globlmachine_check_fwnmimachine_check_fwnmi:
@@ -284,6 +284,18 @@ void system_reset_exception(struct pt_regs *regs)/* What should we do here? We could issue a shutdown or hard reset. */}++/*+*Thisfunctioniscalledinrealmode.Strictlynoprintk'splease.+*+*regs->nipandregs->msrcontainssrr0andssr1.+*/+longmachine_check_early(structpt_regs*regs)+{+/* TODO: handle/decode machine check reason */+return0;+}+#endif/*
From: Mahesh Salgaonkar <redacted>
We can get machine checks from any context. We need to make sure that
we handle all of them correctly. Once we decode MCE reason and generate
MCE event, we continue in host kernel in virtual mode so that we can
log/display it later. But before going to virtual mode we need to make
sure that r1 points to host kernel stack. But machine check can occur
in any context and r1 may not always point to host kernel stack. In cases
where we can not trust r1 value, we should queue up the MCE event and return
from interrupt. This patch implements the additional checks that helps to
decide whether to deleiver machine check event to host kernel right away
or queue it up and return.
Signed-off-by: Mahesh Salgaonkar <redacted>
---
arch/powerpc/kernel/exceptions-64s.S | 72 ++++++++++++++++++++++++++++++++++
1 file changed, 71 insertions(+), 1 deletion(-)
From: Mahesh Salgaonkar <redacted>
This patch adds the early machine check function pointer in cputable for
CPU specific early machine check handling. The early machine handle routine
will be called in real mode to handle SLB and TLB errors. This patch just
sets up a mechanism invoke CPU specific handler. The subsequent patches
will populate the function pointer.
Signed-off-by: Mahesh Salgaonkar <redacted>
---
arch/powerpc/include/asm/cputable.h | 7 +++++++
arch/powerpc/kernel/traps.c | 7 +++++--
2 files changed, 12 insertions(+), 2 deletions(-)
From: Mahesh Salgaonkar <redacted>
This patch introduces flush_tlb operation in cpu_spec structure. This will
help us to invoke appropriate CPU-side flush tlb routine. This patch
adds the foundation to invoke CPU specific flush routine for respective
architectures. Currently this patch introduce flush_tlb for p7 and p8.
Signed-off-by: Mahesh Salgaonkar <redacted>
---
arch/powerpc/include/asm/cputable.h | 5 ++++
arch/powerpc/kernel/cpu_setup_power.S | 38 +++++++++++++++++++++++----------
arch/powerpc/kernel/cputable.c | 8 +++++++
arch/powerpc/kvm/book3s_hv_ras.c | 18 +++-------------
4 files changed, 44 insertions(+), 25 deletions(-)
@@ -96,7 +84,8 @@ static long kvmppc_realmode_mc_power7(struct kvm_vcpu *vcpu)DSISR_MC_SLB_PARITY|DSISR_MC_DERAT_MULTI);}if(dsisr&DSISR_MC_TLB_MULTI){-flush_tlb_power7(vcpu);+if(cur_cpu_spec&&cur_cpu_spec->flush_tlb)+cur_cpu_spec->flush_tlb(TLBIEL_INVAL_SET_LPID);dsisr&=~DSISR_MC_TLB_MULTI;}/* Any other errors we don't understand? */
@@ -113,7 +102,8 @@ static long kvmppc_realmode_mc_power7(struct kvm_vcpu *vcpu)reload_slb(vcpu);break;caseSRR1_MC_IFETCH_TLBMULTI:-flush_tlb_power7(vcpu);+if(cur_cpu_spec&&cur_cpu_spec->flush_tlb)+cur_cpu_spec->flush_tlb(TLBIEL_INVAL_SET_LPID);break;default:handled=0;
From: Mahesh Salgaonkar <redacted>
If we get a machine check exception due to SLB or TLB errors, then flush
SLBs/TLBs and reload SLBs to recover. We do this in real mode before turning
on MMU. Otherwise we would run into nested machine checks.
If we get a machine check when we are in guest, then just flush the
SLBs and continue. This patch handles errors for power7. The next
patch will handle errors for power8
Signed-off-by: Mahesh Salgaonkar <redacted>
---
arch/powerpc/include/asm/bitops.h | 5 +
arch/powerpc/include/asm/mce.h | 67 ++++++++++++++++
arch/powerpc/kernel/Makefile | 1
arch/powerpc/kernel/cputable.c | 4 +
arch/powerpc/kernel/mce_power.c | 153 +++++++++++++++++++++++++++++++++++++
5 files changed, 230 insertions(+)
create mode 100644 arch/powerpc/include/asm/mce.h
create mode 100644 arch/powerpc/kernel/mce_power.c
From: Mahesh Salgaonkar <redacted>
This patch handles the memory errors on power8. If we get a machine check
exception due to SLB or TLB errors, then flush SLBs/TLBs and reload SLBs to
recover.
Signed-off-by: Mahesh Salgaonkar <redacted>
---
arch/powerpc/include/asm/mce.h | 3 +++
arch/powerpc/kernel/cputable.c | 4 ++++
arch/powerpc/kernel/mce_power.c | 34 ++++++++++++++++++++++++++++++++++
3 files changed, 41 insertions(+)
From: Mahesh Salgaonkar <redacted>
Now that we handle machine check in linux, the MCE decoding should also
take place in linux host. This info is crucial to log before we go down
in case we can not handle the machine check errors. This patch decodes
and populates a machine check event which contain high level meaning full
MCE information.
We do this in real mode C code with ME bit on. The MCE information is still
available on emergency stack (in pt_regs structure format). Even if we take
another exception at this point the MCE early handler will allocate a new
stack frame on top of current one. So when we return back here we still have
our MCE information safe on current stack.
We use per cpu buffer to save high level MCE information. Each per cpu buffer
is an array of machine check event structure indexed by per cpu counter
mce_nest_count. The mce_nest_count is incremented every time we enter
machine check early handler in real mode to get the current free slot
(index = mce_nest_count - 1). The mce_nest_count is decremented once the
MCE info is consumed by virtual mode machine exception handler.
This patch provides save_mce_event(), get_mce_event() and release_mce_event()
generic routines that can be used by machine check handlers to populate and
retrieve the event. The routine release_mce_event() will free the event slot so
that it can be reused. Caller can invoke get_mce_event() with a release flag
either to release the event slot immediately OR keep it so that it can be
fetched again. The event slot can be also released anytime by invoking
release_mce_event().
This patch also updates kvm code to invoke get_mce_event to retrieve generic
mce event rather than paca->opal_mce_evt.
The KVM code always calls get_mce_event() with release flags set to false so
that event is available for linus host machine
If machine check occurs while we are in guest, KVM tries to handle the error.
If KVM is able to handle MC error successfully, it enters the guest and
delivers the machine check to guest. If KVM is not able to handle MC error, it
exists the guest and passes the control to linux host machine check handler
which then logs MC event and decides how to handle it in linux host. In failure
case, KVM needs to make sure that the MC event is available for linux host to
consume. Hence KVM always calls get_mce_event() with release flags set to false
and later it invokes release_mce_event() only if it succeeds to handle error.
Signed-off-by: Mahesh Salgaonkar <redacted>
---
arch/powerpc/include/asm/mce.h | 124 +++++++++++++++++++++++++
arch/powerpc/kernel/Makefile | 2
arch/powerpc/kernel/mce.c | 164 +++++++++++++++++++++++++++++++++
arch/powerpc/kernel/mce_power.c | 116 ++++++++++++++++++++++-
arch/powerpc/kvm/book3s_hv_ras.c | 32 ++++--
arch/powerpc/platforms/powernv/opal.c | 35 +++----
6 files changed, 434 insertions(+), 39 deletions(-)
create mode 100644 arch/powerpc/kernel/mce.c
@@ -0,0 +1,164 @@+/*+*Machinecheckexceptionhandling.+*+*Thisprogramisfreesoftware;youcanredistributeitand/ormodify+*itunderthetermsoftheGNUGeneralPublicLicenseaspublishedby+*theFreeSoftwareFoundation;eitherversion2oftheLicense,or+*(atyouroption)anylaterversion.+*+*Thisprogramisdistributedinthehopethatitwillbeuseful,+*butWITHOUTANYWARRANTY;withouteventheimpliedwarrantyof+*MERCHANTABILITYorFITNESSFORAPARTICULARPURPOSE.Seethe+*GNUGeneralPublicLicenseformoredetails.+*+*YoushouldhavereceivedacopyoftheGNUGeneralPublicLicense+*alongwiththisprogram;ifnot,writetotheFreeSoftware+*Foundation,Inc.,59TemplePlace-Suite330,Boston,MA02111-1307,USA.+*+*Copyright2013IBMCorporation+*Author:MaheshSalgaonkar<mahesh@linux.vnet.ibm.com>+*/++#undef DEBUG+#define pr_fmt(fmt) "mce: " fmt++#include<linux/types.h>+#include<linux/ptrace.h>+#include<linux/percpu.h>+#include<linux/export.h>+#include<asm/mce.h>++staticDEFINE_PER_CPU(int,mce_nest_count);+staticDEFINE_PER_CPU(structmachine_check_event[MAX_MC_EVT],mce_event);++staticvoidmce_set_error_info(structmachine_check_event*mce,+structmce_error_info*mce_err)+{+mce->error_type=mce_err->error_type;+switch(mce_err->error_type){+caseMCE_ERROR_TYPE_UE:+mce->u.ue_error.ue_error_type=mce_err->u.ue_error_type;+break;+caseMCE_ERROR_TYPE_SLB:+mce->u.slb_error.slb_error_type=mce_err->u.slb_error_type;+break;+caseMCE_ERROR_TYPE_ERAT:+mce->u.erat_error.erat_error_type=mce_err->u.erat_error_type;+break;+caseMCE_ERROR_TYPE_TLB:+mce->u.tlb_error.tlb_error_type=mce_err->u.tlb_error_type;+break;+caseMCE_ERROR_TYPE_UNKNOWN:+default:+break;+}+}++/*+*DecodeandsavehighlevelMCEinformationintopercpubufferwhich+*isanarrayofmachine_check_eventstructure.+*/+voidsave_mce_event(structpt_regs*regs,longhandled,+structmce_error_info*mce_err,+uint64_taddr)+{+uint64_tsrr1;+intindex=__get_cpu_var(mce_nest_count)++;+structmachine_check_event*mce=&__get_cpu_var(mce_event[index]);++/*+*Returnifwedon'thaveenoughspacetologmceevent.+*mce_nest_countmaygobeyondMAX_MC_EVTbutthat'sok,+*thecheckbelowwillstopbufferoverrun.+*/+if(index>=MAX_MC_EVT)+return;++/* Populate generic machine check info */+mce->version=MCE_V1;+mce->srr0=regs->nip;+mce->srr1=regs->msr;+mce->gpr3=regs->gpr[3];+mce->in_use=1;++mce->initiator=MCE_INITIATOR_CPU;+if(handled)+mce->disposition=MCE_DISPOSITION_RECOVERED;+else+mce->disposition=MCE_DISPOSITION_NOT_RECOVERED;+mce->severity=MCE_SEV_ERROR_SYNC;++srr1=regs->msr;++/*+*Populatethemceerror_typeandtype-specificerror_type.+*/+mce_set_error_info(mce,mce_err);++if(!addr)+return;++if(mce->error_type==MCE_ERROR_TYPE_TLB){+mce->u.tlb_error.effective_address_provided=true;+mce->u.tlb_error.effective_address=addr;+}elseif(mce->error_type==MCE_ERROR_TYPE_SLB){+mce->u.slb_error.effective_address_provided=true;+mce->u.slb_error.effective_address=addr;+}elseif(mce->error_type==MCE_ERROR_TYPE_ERAT){+mce->u.erat_error.effective_address_provided=true;+mce->u.erat_error.effective_address=addr;+}elseif(mce->error_type==MCE_ERROR_TYPE_UE){+mce->u.ue_error.effective_address_provided=true;+mce->u.ue_error.effective_address=addr;+}+return;+}++/*+*get_mce_event:+*mcePointertomachine_check_eventstructuretobefilled.+*releaseFlagtoindicatewhethertofreetheeventslotornot.+*0<=donotreleasethemceevent.Callerwillinvoke+*release_mce_event()onceeventhasbeenconsumed.+*1<=releasetheslot.+*+*return1=success+*0=failure+*+*get_mce_event()willbecalledbyplatformspecificmachinecheck+*handleroutineandinKVM.+*Whenwecallget_mce_event(),wearestillininterruptcontextand+*preemptionwillnotbescheduleduntilret_from_expect()routine+*iscalled.+*/+intget_mce_event(structmachine_check_event*mce,boolrelease)+{+intindex=__get_cpu_var(mce_nest_count)-1;+structmachine_check_event*mc_evt;+intret=0;++/* Sanity check */+if(index<0)+returnret;++/* Check if we have MCE info to process. */+if(index<MAX_MC_EVT){+mc_evt=&__get_cpu_var(mce_event[index]);+/* Copy the event structure and release the original */+if(mce)+*mce=*mc_evt;+if(release)+mc_evt->in_use=0;+ret=1;+}+/* Decrement the count to free the slot. */+if(release)+__get_cpu_var(mce_nest_count)--;++returnret;+}++voidrelease_mce_event(void)+{+get_mce_event(NULL,true);+}
@@ -245,8 +246,7 @@ int opal_put_chars(uint32_t vtermno, const char *data, int total_len)intopal_machine_check(structpt_regs*regs){-structopal_machine_check_event*opal_evt=get_paca()->opal_mc_evt;-structopal_machine_check_eventevt;+structmachine_check_eventevt;constchar*level,*sevstr,*subtype;staticconstchar*opal_mc_ue_types[]={"Indeterminate",
@@ -271,30 +271,29 @@ int opal_machine_check(struct pt_regs *regs)"Multihit",};-/* Copy the event structure and release the original */-evt=*opal_evt;-opal_evt->in_use=0;+if(!get_mce_event(&evt,MCE_EVENT_RELEASE))+return0;/* Print things out */-if(evt.version!=OpalMCE_V1){+if(evt.version!=MCE_V1){pr_err("Machine Check Exception, Unknown event version %d !\n",evt.version);return0;}switch(evt.severity){-caseOpalMCE_SEV_NO_ERROR:+caseMCE_SEV_NO_ERROR:level=KERN_INFO;sevstr="Harmless";break;-caseOpalMCE_SEV_WARNING:+caseMCE_SEV_WARNING:level=KERN_WARNING;sevstr="";break;-caseOpalMCE_SEV_ERROR_SYNC:+caseMCE_SEV_ERROR_SYNC:level=KERN_ERR;sevstr="Severe";break;-caseOpalMCE_SEV_FATAL:+caseMCE_SEV_FATAL:default:level=KERN_ERR;sevstr="Fatal";
From: Mahesh Salgaonkar <redacted>
When machine check real mode handler can not continue into host kernel
in V mode, it returns from the interrupt and we loose MCE event which
never gets logged. In such a situation queue up the MCE event so that
we can log it later when we get back into host kernel with r1 pointing to
kernel stack e.g. during syscall exit.
Signed-off-by: Mahesh Salgaonkar <redacted>
---
arch/powerpc/include/asm/mce.h | 3 +
arch/powerpc/kernel/entry_64.S | 5 +
arch/powerpc/kernel/exceptions-64s.S | 6 +
arch/powerpc/kernel/mce.c | 154 +++++++++++++++++++++++++++++++++
arch/powerpc/platforms/powernv/opal.c | 97 ---------------------
5 files changed, 167 insertions(+), 98 deletions(-)
@@ -247,29 +247,6 @@ int opal_put_chars(uint32_t vtermno, const char *data, int total_len)intopal_machine_check(structpt_regs*regs){structmachine_check_eventevt;-constchar*level,*sevstr,*subtype;-staticconstchar*opal_mc_ue_types[]={-"Indeterminate",-"Instruction fetch",-"Page table walk ifetch",-"Load/Store",-"Page table walk Load/Store",-};-staticconstchar*opal_mc_slb_types[]={-"Indeterminate",-"Parity",-"Multihit",-};-staticconstchar*opal_mc_erat_types[]={-"Indeterminate",-"Parity",-"Multihit",-};-staticconstchar*opal_mc_tlb_types[]={-"Indeterminate",-"Parity",-"Multihit",-};if(!get_mce_event(&evt,MCE_EVENT_RELEASE))return0;
From: Mahesh Salgaonkar <redacted>
Now that we are ready to handle machine check directly in linux, do not
register with firmware to handle machine check exception.
Signed-off-by: Mahesh Salgaonkar <redacted>
---
arch/powerpc/platforms/powernv/opal.c | 8 ++------
1 file changed, 2 insertions(+), 6 deletions(-)
@@ -83,14 +83,10 @@ static int __init opal_register_exception_handlers(void)if(!(powerpc_firmware_features&FW_FEATURE_OPAL))return-ENODEV;-/* Hookup some exception handlers. We use the fwnmi area at 0x7000-*toprovidethegluespacetoOPAL+/* Hookup some exception handlers except machine check. We use the+*fwnmiareaat0x7000toprovidethegluespacetoOPAL*/glue=0x7000;-opal_register_exception_handler(OPAL_MACHINE_CHECK_HANDLER,-__pa(opal_mc_secondary_handler[0]),-glue);-glue+=128;opal_register_exception_handler(OPAL_HYPERVISOR_MAINTENANCE_HANDLER,0,glue);glue+=128;
From: Mahesh Salgaonkar <redacted>
Add basic error handling in machine check exception handler.
- If MSR_RI isn't set, we can not recover.
- Check if disposition set to OpalMCE_DISPOSITION_RECOVERED.
- Check if address at fault is inside kernel address space, if not then send
SIGBUS to process if we hit exception when in userspace.
- If address at fault is not provided then and if we get a synchronous machine
check while in userspace then kill the task.
Signed-off-by: Mahesh Salgaonkar <redacted>
---
arch/powerpc/include/asm/mce.h | 1 +
arch/powerpc/kernel/mce.c | 27 +++++++++++++++++++++
arch/powerpc/platforms/powernv/opal.c | 43 ++++++++++++++++++++++++++++++++-
3 files changed, 70 insertions(+), 1 deletion(-)
From: Paul Mackerras <hidden> Date: 2013-09-09 04:29:02
On Tue, Aug 27, 2013 at 01:01:24AM +0530, Mahesh J Salgaonkar wrote:
From: Mahesh Salgaonkar <redacted>
This patch splits the common exception prolog logic into two parts to
facilitate reuse of existing code in the next patch. The second part will
be reused in the machine check exception routine in the next patch.
Please note that this patch does not introduce or change existing code
logic. Instead it is just a code movement.
Looks OK. Note however that the diff would be a lot smaller if you
had put EXCEPTION_PROLOG_COMMON_2 after EXCEPTION_PROLOG_COMMON instead
of before it.
Acked-by: Paul Mackerras <redacted>
From: Paul Mackerras <hidden> Date: 2013-09-09 04:30:16
On Tue, Aug 27, 2013 at 01:01:32AM +0530, Mahesh J Salgaonkar wrote:
From: Mahesh Salgaonkar <redacted>
This patch introduces exclusive emergency stack for machine check exception.
We use emergency stack to handle machine check exception so that we can save
MCE information (srr1, srr0, dar and dsisr) before turning on ME bit and be
ready for re-entrancy. This helps us to prevent clobbering of MCE information
in case of nested machine checks.
The reason for using emergency stack over normal kernel stack is that the
machine check might occur in the middle of setting up a stack frame which may
result into improper use of kernel stack.
Signed-off-by: Mahesh Salgaonkar <redacted>
From: Paul Mackerras <hidden> Date: 2013-09-09 04:52:37
On Tue, Aug 27, 2013 at 01:01:40AM +0530, Mahesh J Salgaonkar wrote:
From: Mahesh Salgaonkar <redacted>
Move machine check entry point into Linux. So far we were dependent on
firmware to decode MCE error details and handover the high level info to OS.
This patch introduces early machine check routine that saves the MCE
information (srr1, srr0, dar and dsisr) to the emergency stack. We allocate
stack frame on emergency stack and set the r1 accordingly. This allows us to be
prepared to take another exception without loosing context. One thing to note
here that, if we get another machine check while ME bit is off then we risk a
checkstop. Hence we restrict ourselves to save only MCE information and turn
the ME bit on. We use paca->in_mce flag to differentiate between first entry
and nested machine check entry which helps proper use of emergency stack. We
increment paca->in_mce every time we enter in early machine check handler and
decrement it while leaving. When we enter machine check early handler first
time (paca->in_mce == 0), we are sure nobody is using MC emergency stack and
allocate a stack frame at the start of the emergency stack. During subsequent
entry (paca->in_mce > 0), we know that r1 points inside emergency stack and we
allocate separate stack frame accordingly. This prevents us from clobbering MCE
information during nested machine checks.
The early machine check handler changes are placed under CPU_FTR_HVMODE
section. This makes sure that the early machine check handler will get executed
only in hypervisor kernel.
Actually r12 doesn't contain anything except the value it had when the
machine check occurred.
+ * R13 = PACA
+ * R9 = CR
+ * R11 & R12 is saved on PACA_EXMC
All of r9 to r12 are saved in the PACA_EXMC save area.
+ *
+ * Switch to mc_emergency stack and handle re-entrancy (though we
+ * currently don't test for overflow). Save MCE registers srr1,
+ * srr0, dar and dsisr and then set ME=1
+ *
+ * We use paca->in_mce to check whether this is the first entry or
+ * nested machine check. We increment paca->in_mce to track nested
+ * machine checks.
+ *
+ * If this is the first entry then set stack pointer to
+ * paca->mc_emergency_sp, otherwise r1 is already pointing to
+ * stack frame on mc_emergency stack.
+ *
+ * NOTE: We are here with MSR_ME=0 (off), which means we risk a
+ * checkstop if we get another machine check exception before we do
+ * rfid with MSR_ME=1.
+ */
+ mr r11,r1 /* Save r1 */
+ lhz r10,PACA_IN_MCE(r13)
+ cmpwi r10,0 /* Are we in nested machine check */
+ bne 0f /* Yes, we are. */
+ /* First machine check entry */
+ ld r1,PACAMCEMERGSP(r13) /* Use MC emergency stack */
+0: subi r1,r1,INT_FRAME_SIZE /* alloc stack frame */
+ addi r10,r10,1 /* increment paca->in_mce */
+ sth r10,PACA_IN_MCE(r13)
+ std r11,GPR1(r1) /* Save r1 on the stack. */
+ std r11,0(r1) /* make stack chain pointer */
+ mfspr r11,SPRN_SRR0 /* Save SRR0 */
+ std r11,_NIP(r1)
+ mfspr r11,SPRN_SRR1 /* Save SRR1 */
+ std r11,_MSR(r1)
+ mfspr r11,SPRN_DAR /* Save DAR */
+ std r11,_DAR(r1)
+ mfspr r11,SPRN_DSISR /* Save DSISR */
+ std r11,_DSISR(r1)
+ mfmsr r11 /* get MSR value */
+ ori r11,r11,MSR_ME /* turn on ME bit */
At this point you should turn on MSR_RI as well, so that if we get
another machine check after the rfid we don't consider it
unrecoverable. Also, since we could in principle get another machine
check soon after the rfid, and that would use the EX_MC save area,
we need to copy everything out of the EX_MC save area onto the
stack before turning on ME.
quoted hunk
+ ld r12,PACAKBASE(r13) /* get high part of &label */
+ LOAD_HANDLER(r12, machine_check_handle_early)
+ mtspr SPRN_SRR0,r12
+ mtspr SPRN_SRR1,r11
+ rfid
+ b . /* prevent speculative execution */
+END_FTR_SECTION_IFSET(CPU_FTR_HVMODE)
+
machine_check_pSeries:
.globl machine_check_fwnmi
machine_check_fwnmi:
This is basically fast_exception_return without the rfid, and with the
decrementing of paca->in_mce. You forgot to clear MSR_RI before
setting SRR0 and SRR1. Also, there's no point restoring CFAR if the
next thing you are going to do is either an rfid or a branch, both of
which will set CFAR. Further, you don't need the REST_NVGPRS unless
you are expecting that your machine check handler will want to modify
some of the GPR values in regs->gpr[14 ... 31], which I don't believe
is the case.
Paul.
From: Paul Mackerras <hidden> Date: 2013-09-09 05:29:30
On Tue, Aug 27, 2013 at 01:01:48AM +0530, Mahesh J Salgaonkar wrote:
From: Mahesh Salgaonkar <redacted>
We can get machine checks from any context. We need to make sure that
we handle all of them correctly. Once we decode MCE reason and generate
MCE event, we continue in host kernel in virtual mode so that we can
log/display it later. But before going to virtual mode we need to make
sure that r1 points to host kernel stack. But machine check can occur
in any context and r1 may not always point to host kernel stack. In cases
where we can not trust r1 value, we should queue up the MCE event and return
from interrupt. This patch implements the additional checks that helps to
decide whether to deleiver machine check event to host kernel right away
or queue it up and return.
Some comments below...
+ /*
+ * We are now going to host kernel in V mode. We need to make sure
+ * that r1 points to host kernel stack.
+ *
+ * If we are coming from userspace then we can continue in host kernel
+ * in V mode.
+ * But if we are coming from kernel and r1 does not point to kernel
+ * stack then we can not continue, instead we return from here.
+ */
+
+ ld r12,_MSR(r1)
+ andi. r11,r12,MSR_PR /* See if coming from user. */
+ bne 3f /* continue if we are. */
+
+#ifdef CONFIG_KVM_BOOK3S_64_HV
+ /*
+ * We are coming from kernel context. Check if we are coming from
+ * guest. if yes, then we can continue. We will fall through
+ * do_kvm_200->kvmppc_interrupt which will setup r1 correctly.
+ */
It seems fragile to have to check various conditions to know whether
r1 is actually a kernel stack pointer, but I guess it's the best we
can do at present.
+ lbz r11,HSTATE_IN_GUEST(r13)
+ cmpwi r11,0 /* Check if coming from guest */
+ bne 3f /* continue if we are. */
+
+ /*
+ * So, we did not come from guest. That leaves three possibilities:
+ * a. We come from secondary thread which just came out of nap and
+ * about to call kvm_start_guest.
+ * b. We come from secondary thread which is about to go to nap
+ * state (see kvm_no_guest()).
+ * c. We come from opal context and r1 may be pointing to opal
+ * kernel stack.
+ */
+
+ lbz r11,HSTATE_HWTHREAD_STATE(r13)
+ cmpwi r11,KVM_HWTHREAD_IN_NAP /* Was it nap-ing? or about to */
+ beq 0f /* Queue up event and return from interrupt */
Two comments here: first, we change the hwthread_state to
KVM_HWTHREAD_IN_KERNEL before loading up r1 -- this is in
system_reset_pSeries in exceptions-64s.S. So this test isn't really
safe. It would be possible to add ld r1, PACAR1(r13) before setting
the hwthread_state, and I think that would fix it.
Secondly, if the CPU is napping when the machine check comes along,
it doesn't jump to the machine check vector. It restarts the CPU at
the system reset vector, with a particular wakeup code in SRR1, which
we currently don't handle. So you need to add code to do that.
+ * So far we checked all possible situations where we can not
+ * trust r1. Now we can trust r1.
+ * r1 < 0 r1 points to host kernel stack
+ * r1 > 0 r1 points to opal stack
Are we guaranteed that Sapphire will keep the stack pointer positive
at all times? (More a question for Ben H than you.)
Paul.
From: Paul Mackerras <hidden> Date: 2013-09-09 05:33:50
On Tue, Aug 27, 2013 at 01:01:56AM +0530, Mahesh J Salgaonkar wrote:
From: Mahesh Salgaonkar <redacted>
This patch adds the early machine check function pointer in cputable for
CPU specific early machine check handling. The early machine handle routine
will be called in real mode to handle SLB and TLB errors. This patch just
sets up a mechanism invoke CPU specific handler. The subsequent patches
will populate the function pointer.
Signed-off-by: Mahesh Salgaonkar <redacted>
Your patch description should talk about how this new hook is
different from the existing machine_check hook and why you need a new
one.
Apart from that:
Acked-by: Paul Mackerras <redacted>
From: Paul Mackerras <hidden> Date: 2013-09-09 05:36:15
On Tue, Aug 27, 2013 at 01:02:04AM +0530, Mahesh J Salgaonkar wrote:
From: Mahesh Salgaonkar <redacted>
This patch introduces flush_tlb operation in cpu_spec structure. This will
help us to invoke appropriate CPU-side flush tlb routine. This patch
adds the foundation to invoke CPU specific flush routine for respective
architectures. Currently this patch introduce flush_tlb for p7 and p8.
Signed-off-by: Mahesh Salgaonkar <redacted>
From: Paul Mackerras <hidden> Date: 2013-09-09 06:00:29
On Tue, Aug 27, 2013 at 01:02:12AM +0530, Mahesh J Salgaonkar wrote:
From: Mahesh Salgaonkar <redacted>
If we get a machine check exception due to SLB or TLB errors, then flush
SLBs/TLBs and reload SLBs to recover. We do this in real mode before turning
on MMU. Otherwise we would run into nested machine checks.
If we get a machine check when we are in guest, then just flush the
SLBs and continue. This patch handles errors for power7. The next
patch will handle errors for power8
This will break if anyone is so incautious as to use this in a 32-bit
kernel. It would be kinder to do:
#define PPC_BITLSHIFT(be) (BITS_PER_LONG - 1 - (be))
#define PPC_BIT(bit) (1ul << PPC_BITLSHIFT(bit))
+/* SRR1 bits for machine check (On Power7 and Power8) */
+#define P7_SRR1_MC_IFETCH(srr1) ((srr1) & PPC_BITMASK(43, 45)) /* P8 too */
+
+#define P7_SRR1_MC_IFETCH_UE (0x1 << PPC_BITLSHIFT(45)) /* P8 too */
+#define P7_SRR1_MC_IFETCH_SLB_PARITY (0x2 << PPC_BITLSHIFT(45)) /* P8 too */
+#define P7_SRR1_MC_IFETCH_SLB_MULTIHIT (0x3 << PPC_BITLSHIFT(45)) /* P8 too */
+#define P7_SRR1_MC_IFETCH_SLB_BOTH (0x4 << PPC_BITLSHIFT(45)) /* P8 too */
+#define P7_SRR1_MC_IFETCH_TLB_MULTIHIT (0x5 << PPC_BITLSHIFT(45)) /* P8 too */
+#define P7_SRR1_MC_IFETCH_UE_TLB_RELOAD (0x6 << PPC_BITLSHIFT(45)) /* P8 too */
+#define P7_SRR1_MC_IFETCH_UE_IFU_INTERNAL (0x7 << PPC_BITLSHIFT(45))
+
+/* SRR1 bits for machine check (On Power8) */
+#define P8_SRR1_MC_IFETCH_ERAT_MULTIHIT (0x4 << PPC_BITLSHIFT(45))
How do we tell the difference between that and P7_SRR1_MC_IFETCH_SLB_BOTH?
+/* flush SLBs and reload */
+static void flush_and_reload_slb(void)
+{
+ struct slb_shadow *slb;
+ unsigned long i, n;
+
+ if (!mmu_has_feature(MMU_FTR_SLB))
+ return;
This seems unnecessary when we can only call this on POWER7.
Paul.
From: Paul Mackerras <hidden> Date: 2013-09-09 06:01:25
On Tue, Aug 27, 2013 at 01:02:20AM +0530, Mahesh J Salgaonkar wrote:
From: Mahesh Salgaonkar <redacted>
This patch handles the memory errors on power8. If we get a machine check
exception due to SLB or TLB errors, then flush SLBs/TLBs and reload SLBs to
recover.
Signed-off-by: Mahesh Salgaonkar <redacted>