I'm trying to flush my patch queue again for 2.6.20, this is
most of what has been coming in since the last submission
to powerpc.git. Please review for inclusion in powerpc.git.
As a quick overview, the patches contain:
* new spufs file interfaces to support more gdb functionality,
to look into the state of various HW registers without changing
data.
* core dump support for SPUs. When core dumping a task that has
created SPU contexts, information about them will be included
in the dump.
* A rework of the SPU HW isolation support. Reusing an isolated
SPU for another program now doesn't require a new spufs interface
any more, which simplifies the implementation significantly.
* A new oprofile model for using the profile event mechanism on
the Cell Broadband Engine. This is only for PPE profiling so far,
profiling SPU code is still work-in-progress.
* Lots of bug fixes.
I have some other patches in my cell kernel tree currently, which
I'm not submitting myself, because I assume they will find their own
way into the kernel:
* Patches to the spidernet driver, these go through netdev
* Support for the Playstation 3, the patches are currently under
discussion, and I expect that Paul will take them directly once
they are done.
* xmon add-ons for debugging with SPUs. Waiting for a new version
to be submitted.
Arnd <><
From: Dwayne Grant McConnell <redacted>
We need to check the channel count of the signal notification registers
before reading them, because it can be undefined when the count is
zero. In order to read count and data atomically, we read from the
saved context.
This patch uses spu_acquire_saved() to force a context save before a
/signal1 or /signal2 read. Because of this it is no longer necessary to
have backing_ops and hw_ops versions of this function so they have been
removed.
Regular applications should not rely on reading this register
to be fast, as it's conceptually a write-only file from the PPE
perspective.
Signed-off-by: Dwayne Grant McConnell <redacted>
Signed-off-by: Arnd Bergmann <redacted>
---
Dwayne Grant McConnell [off-list ref]
Lotus Notes Mail: Dwayne McConnell [Mail]/Austin/IBM@IBMUS
Lotus Notes Calendar: Dwayne McConnell [Calendar]/Austin/IBM@IBMUS
Index: linux-2.6/arch/powerpc/platforms/cell/spufs/file.c
===================================================================
From: Kevin Corry <redacted>
More macros for manipulating bits in the Cell PMU control registers.
Signed-off-by: Kevin Corry <redacted>
Signed-off-by: Carl Love <redacted>
Signed-off-by: Arnd Bergmann <redacted>
Index: linux-2.6/arch/powerpc/platforms/cell/cbe_regs.h
===================================================================
Add symbol-exports for the new routines in arch/powerpc/platforms/cell/pmu.c.
They are needed for Oprofile, which can be built as a module.
Patch is against 2.6.18-arnd5.
Signed-Off-By: Kevin Corry <redacted>
Signed-off-by: Arnd Bergmann <redacted>
Index: linux-2.6/arch/powerpc/platforms/cell/pmu.c
===================================================================
From: Christoph Hellwig <hch@infradead.org>
When one of the spufs files is mapped into a process address
space, regular users can use ptrace to attempt accessing
them with access_process_vm(). With the way that the
mappings currently work, this likely causes an oops.
Setting the vm_flags to VM_IO makes sure that ptrace can
not access them but returns an error code. This is not
the perfect solution in case of the local store mapping,
but it fixes the oops in a well-defined way.
Also remove leftover VM_RESERVED flags in spufs. The
VM_RESERVED flag is on it's way out and not checked by
the memory managment code anymore.
Signed-off-by: Arnd Bergmann <redacted>
Signed-off-by: Christoph Hellwig <redacted>
---
Index: linux-2.6/arch/powerpc/platforms/cell/spufs/file.c
===================================================================
From: Maynard Johnson <redacted>
Add PPU event-based and cycle-based profiling support to Oprofile for Cell.
Oprofile is expected to collect data on all CPUs simultaneously.
However, there is one set of performance counters per node. There are
two hardware threads or virtual CPUs on each node. Hence, OProfile must
multiplex in time the performance counter collection on the two virtual
CPUs.
The multiplexing of the performance counters is done by a virtual
counter routine. Initially, the counters are configured to collect data
on the even CPUs in the system, one CPU per node. In order to capture
the PC for the virtual CPU when the performance counter interrupt occurs
(the specified number of events between samples has occurred), the even
processors are configured to handle the performance counter interrupts
for their node. The virtual counter routine is called via a kernel
timer after the virtual sample time. The routine stops the counters,
saves the current counts, loads the last counts for the other virtual
CPU on the node, sets interrupts to be handled by the other virtual CPU
and restarts the counters, the virtual timer routine is scheduled to run
again. The virtual sample time is kept relatively small to make sure
sampling occurs on both CPUs on the node with a relatively small
granularity. Whenever the counters overflow, the performance counter
interrupt is called to collect the PC for the CPU where data is being
collected.
The oprofile driver relies on a firmware RTAS call to setup the debug bus
to route the desired signals to the performance counter hardware to be
counted. The RTAS call must set the routing registers appropriately in
each of the islands to pass the signals down the debug bus as well as
routing the signals from a particular island onto the bus. There is a
second firmware RTAS call to reset the debug bus to the non pass thru
state when the counters are not in use.
Signed-off-by: Carl Love <redacted>
Signed-off-by: Maynard Johnson <redacted>
Signed-off-by: Arnd Bergmann <redacted>
Index: linux-2.6/arch/powerpc/configs/cell_defconfig
===================================================================
@@ -1121,7 +1121,8 @@ CONFIG_PLIST=y # # Instrumentation Support #-# CONFIG_PROFILING is not set+CONFIG_PROFILING=y+CONFIG_OPROFILE=y # CONFIG_KPROBES is not set #
@@ -0,0 +1,724 @@+/*+*CellBroadbandEngineOProfileSupport+*+*(C)CopyrightIBMCorporation2006+*+*Author:DavidErb(djerb@us.ibm.com)+*Modifications:+*CarlLove<carll@us.ibm.com>+*MaynardJohnson<maynardj@us.ibm.com>+*+*Thisprogramisfreesoftware;youcanredistributeitand/or+*modifyitunderthetermsoftheGNUGeneralPublicLicense+*aspublishedbytheFreeSoftwareFoundation;eitherversion+*2oftheLicense,or(atyouroption)anylaterversion.+*/++#include<linux/cpufreq.h>+#include<linux/delay.h>+#include<linux/init.h>+#include<linux/jiffies.h>+#include<linux/kthread.h>+#include<linux/oprofile.h>+#include<linux/percpu.h>+#include<linux/smp.h>+#include<linux/spinlock.h>+#include<linux/timer.h>+#include<asm/cell-pmu.h>+#include<asm/cputable.h>+#include<asm/firmware.h>+#include<asm/io.h>+#include<asm/oprofile_impl.h>+#include<asm/processor.h>+#include<asm/prom.h>+#include<asm/ptrace.h>+#include<asm/reg.h>+#include<asm/rtas.h>+#include<asm/system.h>++#include"../platforms/cell/interrupt.h"++#define PPU_CYCLES_EVENT_NUM 1 /* event number for CYCLES */+#define CBE_COUNT_ALL_CYCLES 0x42800000 /* PPU cycle event specifier */++#define NUM_THREADS 2+#define VIRT_CNTR_SW_TIME_NS 100000000 // 0.5 seconds++structpmc_cntrl_data{+unsignedlongvcntr;+unsignedlongevnts;+unsignedlongmasks;+unsignedlongenabled;+};++/*+*ibm,cbe-perftoolsrtasparameters+*/++structpm_signal{+u16cpu;/* Processor to modify */+u16sub_unit;/* hw subunit this applies to (if applicable) */+u16signal_group;/* Signal Group to Enable/Disable */+u8bus_word;/* Enable/Disable on this Trace/Trigger/Event+*BusWord(s)(bitmask)+*/+u8bit;/* Trigger/Event bit (if applicable) */+};++/*+*rtascallarguments+*/+enum{+SUBFUNC_RESET=1,+SUBFUNC_ACTIVATE=2,+SUBFUNC_DEACTIVATE=3,++PASSTHRU_IGNORE=0,+PASSTHRU_ENABLE=1,+PASSTHRU_DISABLE=2,+};++structpm_cntrl{+u16enable;+u16stop_at_max;+u16trace_mode;+u16freeze;+u16count_mode;+};++staticstruct{+u32group_control;+u32debug_bus_control;+structpm_cntrlpm_cntrl;+u32pm07_cntrl[NR_PHYS_CTRS];+}pm_regs;+++#define GET_SUB_UNIT(x) ((x & 0x0000f000) >> 12)+#define GET_BUS_WORD(x) ((x & 0x000000f0) >> 4)+#define GET_BUS_TYPE(x) ((x & 0x00000300) >> 8)+#define GET_POLARITY(x) ((x & 0x00000002) >> 1)+#define GET_COUNT_CYCLES(x) (x & 0x00000001)+#define GET_INPUT_CONTROL(x) ((x & 0x00000004) >> 2)+++staticDEFINE_PER_CPU(unsignedlong[NR_PHYS_CTRS],pmc_values);++staticstructpmc_cntrl_datapmc_cntrl[NUM_THREADS][NR_PHYS_CTRS];++/* Interpetation of hdw_thread:+*0-evenvirtualcpus0,2,4,...+*1-oddvirtualcpus1,3,5,...+*/+staticu32hdw_thread;++staticu32virt_cntr_inter_mask;+staticstructtimer_listtimer_virt_cntr;++/* pm_signal needs to be global since it is initialized in+*cell_reg_setupatthetimewhenthenecessaryinformation+*isavailable.+*/+staticstructpm_signalpm_signal[NR_PHYS_CTRS];+staticintpm_rtas_token;++staticu32reset_value[NR_PHYS_CTRS];+staticintnum_counters;+staticintoprofile_running;+staticspinlock_tvirt_cntr_lock=SPIN_LOCK_UNLOCKED;++staticu32ctr_enabled;++staticunsignedchartrace_bus[4];+staticunsignedcharinput_bus[2];++/*+*Firmwareinterfacefunctions+*/+staticint+rtas_ibm_cbe_perftools(intsubfunc,intpassthru,+void*address,unsignedlonglength)+{+u64paddr=__pa(address);++returnrtas_call(pm_rtas_token,5,1,NULL,subfunc,passthru,+paddr>>32,paddr&0xffffffff,length);+}++staticvoidpm_rtas_reset_signals(u32node)+{+intret;+structpm_signalpm_signal_local;++/* The debug bus is being set to the passthru disable state.+*However,theFWstillexpectsatleastonelegalsignalrouting+*entryoritwillreturnanerroronthearguments.Ifwedon't+*supplyavalidentry,wemustignoreallreturnvalues.Ignoring+*allreturnvaluesmeanswemightmissanerrorweshouldbe+*concernedabout.+*/++/* fw expects physical cpu #. */+pm_signal_local.cpu=node;+pm_signal_local.signal_group=21;+pm_signal_local.bus_word=1;+pm_signal_local.sub_unit=0;+pm_signal_local.bit=0;++ret=rtas_ibm_cbe_perftools(SUBFUNC_RESET,PASSTHRU_DISABLE,+&pm_signal_local,+sizeof(structpm_signal));++if(ret)+printk(KERN_WARNING"%s: rtas returned: %d\n",+__FUNCTION__,ret);+}++staticvoidpm_rtas_activate_signals(u32node,u32count)+{+intret;+intj;+structpm_signalpm_signal_local[NR_PHYS_CTRS];++for(j=0;j<count;j++){+/* fw expects physical cpu # */+pm_signal_local[j].cpu=node;+pm_signal_local[j].signal_group=pm_signal[j].signal_group;+pm_signal_local[j].bus_word=pm_signal[j].bus_word;+pm_signal_local[j].sub_unit=pm_signal[j].sub_unit;+pm_signal_local[j].bit=pm_signal[j].bit;+}++ret=rtas_ibm_cbe_perftools(SUBFUNC_ACTIVATE,PASSTHRU_ENABLE,+pm_signal_local,+count*sizeof(structpm_signal));++if(ret)+printk(KERN_WARNING"%s: rtas returned: %d\n",+__FUNCTION__,ret);+}++/*+*PMSignalfunctions+*/+staticvoidset_pm_event(u32ctr,intevent,u32unit_mask)+{+structpm_signal*p;+u32signal_bit;+u32bus_word,bus_type,count_cycles,polarity,input_control;+intj,i;++if(event==PPU_CYCLES_EVENT_NUM){+/* Special Event: Count all cpu cycles */+pm_regs.pm07_cntrl[ctr]=CBE_COUNT_ALL_CYCLES;+p=&(pm_signal[ctr]);+p->signal_group=21;+p->bus_word=1;+p->sub_unit=0;+p->bit=0;+gotoout;+}else{+pm_regs.pm07_cntrl[ctr]=0;+}++bus_word=GET_BUS_WORD(unit_mask);+bus_type=GET_BUS_TYPE(unit_mask);+count_cycles=GET_COUNT_CYCLES(unit_mask);+polarity=GET_POLARITY(unit_mask);+input_control=GET_INPUT_CONTROL(unit_mask);+signal_bit=(event%100);++p=&(pm_signal[ctr]);++p->signal_group=event/100;+p->bus_word=bus_word;+p->sub_unit=unit_mask&0x0000f000;++pm_regs.pm07_cntrl[ctr]=0;+pm_regs.pm07_cntrl[ctr]|=PM07_CTR_COUNT_CYCLES(count_cycles);+pm_regs.pm07_cntrl[ctr]|=PM07_CTR_POLARITY(polarity);+pm_regs.pm07_cntrl[ctr]|=PM07_CTR_INPUT_CONTROL(input_control);++if(input_control==0){+if(signal_bit>31){+signal_bit-=32;+if(bus_word==0x3)+bus_word=0x2;+elseif(bus_word==0xc)+bus_word=0x8;+}++if((bus_type==0)&&p->signal_group>=60)+bus_type=2;+if((bus_type==1)&&p->signal_group>=50)+bus_type=0;++pm_regs.pm07_cntrl[ctr]|=PM07_CTR_INPUT_MUX(signal_bit);+}else{+pm_regs.pm07_cntrl[ctr]=0;+p->bit=signal_bit;+}++for(i=0;i<4;i++){+if(bus_word&(1<<i)){+pm_regs.debug_bus_control|=+(bus_type<<(31-(2*i)+1));++for(j=0;j<2;j++){+if(input_bus[j]==0xff){+input_bus[j]=i;+pm_regs.group_control|=+(i<<(31-i));+break;+}+}+}+}+out:+;+}++staticvoidwrite_pm_cntrl(intcpu,structpm_cntrl*pm_cntrl)+{+/* Oprofile will use 32 bit counters, set bits 7:10 to 0 */+u32val=0;+if(pm_cntrl->enable==1)+val|=CBE_PM_ENABLE_PERF_MON;++if(pm_cntrl->stop_at_max==1)+val|=CBE_PM_STOP_AT_MAX;++if(pm_cntrl->trace_mode==1)+val|=CBE_PM_TRACE_MODE_SET(pm_cntrl->trace_mode);++if(pm_cntrl->freeze==1)+val|=CBE_PM_FREEZE_ALL_CTRS;++/* Routine set_count_mode must be called previously to set+*thecountmodebasedontheuserselectionofuserandkernel.+*/+val|=CBE_PM_COUNT_MODE_SET(pm_cntrl->count_mode);+cbe_write_pm(cpu,pm_control,val);+}++staticinlinevoid+set_count_mode(u32kernel,u32user,structpm_cntrl*pm_cntrl)+{+/* The user must specify user and kernel if they want them. If+*neitherisspecified,OProfilewillcountinhypervisormode+*/+if(kernel){+if(user)+pm_cntrl->count_mode=CBE_COUNT_ALL_MODES;+else+pm_cntrl->count_mode=CBE_COUNT_SUPERVISOR_MODE;+}else{+if(user)+pm_cntrl->count_mode=CBE_COUNT_PROBLEM_MODE;+else+pm_cntrl->count_mode=CBE_COUNT_HYPERVISOR_MODE;+}+}++staticinlinevoidenable_ctr(u32cpu,u32ctr,u32*pm07_cntrl)+{++pm07_cntrl[ctr]|=PM07_CTR_ENABLE(1);+cbe_write_pm07_control(cpu,ctr,pm07_cntrl[ctr]);+}++/*+*OprofileisexpectedtocollectdataonallCPUssimultaneously.+*However,thereisonesetofperformancecounterspernode.Thereare+*twohardwarethreadsorvirtualCPUsoneachnode.Hence,OProfilemust+*multiplexintimetheperformancecountercollectiononthetwovirtual+*CPUs.Themultiplexingoftheperformancecountersisdonebythis+*virtualcounterroutine.+*+*Thepmc_valuesusedbelowisdefinedas'per-cpu'butitsuseis+*moreakinto'per-node'.Weneedtostoretwosetsofcounter+*valuespernode--oneforthepreviousrunandoneforthenext.+*Theper-cpu[NR_PHYS_CTRS]givesusthestorageweneed.Eachodd/even+*pairofper-cpuarraysisusedforstoringthepreviousandnext+*pmcvaluesforagivennode.+*NOTE:Weusetheper-cpuvariabletoimprovecacheperformance.+*/+staticvoidcell_virtual_cntr(unsignedlongdata)+{+/* This routine will alternate loading the virtual counters for+*virtualCPUs+*/+inti,prev_hdw_thread,next_hdw_thread;+u32cpu;+unsignedlongflags;++/* Make sure that the interrupt_hander and+*thevirtcounterarenotbothplayingwith+*thecountersonthesamenode.+*/++spin_lock_irqsave(&virt_cntr_lock,flags);++prev_hdw_thread=hdw_thread;++/* switch the cpu handling the interrupts */+hdw_thread=1^hdw_thread;+next_hdw_thread=hdw_thread;++/* The following is done only once per each node, but+*weneedcpu#,notnode#,topasstothecbe_xxxfunctions.+*/+for_each_online_cpu(cpu){+if(cbe_get_hw_thread_id(cpu))+continue;++/* stop counters, save counter values, restore counts+*forpreviousthread+*/+cbe_disable_pm(cpu);+cbe_disable_pm_interrupts(cpu);+for(i=0;i<num_counters;i++){+per_cpu(pmc_values,cpu+prev_hdw_thread)[i]+=cbe_read_ctr(cpu,i);++if(per_cpu(pmc_values,cpu+next_hdw_thread)[i]+==0xFFFFFFFF)+/* If the cntr value is 0xffffffff, we must+*resetthatto0xfffffff0whenthecurrent+*threadisrestarted.Thiswillgenerateanew+*interruptandmakesurethatweneverrestore+*thecounterstothemaxvalue.Ifthecounters+*wererestoredtothemaxvalue,theydonot+*incrementandnointerruptsaregenerated.Hence+*nomoresampleswillbecollectedonthatcpu.+*/+cbe_write_ctr(cpu,i,0xFFFFFFF0);+else+cbe_write_ctr(cpu,i,+per_cpu(pmc_values,+cpu++next_hdw_thread)[i]);+}++/* Switch to the other thread. Change the interrupt+*andcontrolregstobescheduledontheCPU+*correspondingtothethreadtoexecute.+*/+for(i=0;i<num_counters;i++){+if(pmc_cntrl[next_hdw_thread][i].enabled){+/* There are some per thread events.+*Mustdothesetevent,enable_cntr+*foreachcpu.+*/+set_pm_event(i,+pmc_cntrl[next_hdw_thread][i].evnts,+pmc_cntrl[next_hdw_thread][i].masks);+enable_ctr(cpu,i,+pm_regs.pm07_cntrl);+}else{+cbe_write_pm07_control(cpu,i,0);+}+}++/* Enable interrupts on the CPU thread that is starting */+cbe_enable_pm_interrupts(cpu,next_hdw_thread,+virt_cntr_inter_mask);+cbe_enable_pm(cpu);+}++spin_unlock_irqrestore(&virt_cntr_lock,flags);++mod_timer(&timer_virt_cntr,jiffies+HZ/10);+}++staticvoidstart_virt_cntrs(void)+{+init_timer(&timer_virt_cntr);+timer_virt_cntr.function=cell_virtual_cntr;+timer_virt_cntr.data=0UL;+timer_virt_cntr.expires=jiffies+HZ/10;+add_timer(&timer_virt_cntr);+}++/* This function is called once for all cpus combined */+staticvoid+cell_reg_setup(structop_counter_config*ctr,+structop_system_config*sys,intnum_ctrs)+{+inti,j,cpu;++pm_rtas_token=rtas_token("ibm,cbe-perftools");+if(pm_rtas_token==RTAS_UNKNOWN_SERVICE){+printk(KERN_WARNING"%s: RTAS_UNKNOWN_SERVICE\n",+__FUNCTION__);+gotoout;+}++num_counters=num_ctrs;++pm_regs.group_control=0;+pm_regs.debug_bus_control=0;++/* setup the pm_control register */+memset(&pm_regs.pm_cntrl,0,sizeof(structpm_cntrl));+pm_regs.pm_cntrl.stop_at_max=1;+pm_regs.pm_cntrl.trace_mode=0;+pm_regs.pm_cntrl.freeze=1;++set_count_mode(sys->enable_kernel,sys->enable_user,+&pm_regs.pm_cntrl);++/* Setup the thread 0 events */+for(i=0;i<num_ctrs;++i){++pmc_cntrl[0][i].evnts=ctr[i].event;+pmc_cntrl[0][i].masks=ctr[i].unit_mask;+pmc_cntrl[0][i].enabled=ctr[i].enabled;+pmc_cntrl[0][i].vcntr=i;++for_each_possible_cpu(j)+per_cpu(pmc_values,j)[i]=0;+}++/* Setup the thread 1 events, map the thread 0 event to the+*equivalentthread1event.+*/+for(i=0;i<num_ctrs;++i){+if((ctr[i].event>=2100)&&(ctr[i].event<=2111))+pmc_cntrl[1][i].evnts=ctr[i].event+19;+elseif(ctr[i].event==2203)+pmc_cntrl[1][i].evnts=ctr[i].event;+elseif((ctr[i].event>=2200)&&(ctr[i].event<=2215))+pmc_cntrl[1][i].evnts=ctr[i].event+16;+else+pmc_cntrl[1][i].evnts=ctr[i].event;++pmc_cntrl[1][i].masks=ctr[i].unit_mask;+pmc_cntrl[1][i].enabled=ctr[i].enabled;+pmc_cntrl[1][i].vcntr=i;+}++for(i=0;i<4;i++)+trace_bus[i]=0xff;++for(i=0;i<2;i++)+input_bus[i]=0xff;++/* Our counters count up, and "count" refers to+*howmuchbeforethenextinterrupt,andweinterrupt+*onoverflow.Sowecalculatethestartingvalue+*whichwillgiveus"count"untiloverflow.+*Thenwesettheeventsontheenabledcounters.+*/+for(i=0;i<num_counters;++i){+/* start with virtual counter set 0 */+if(pmc_cntrl[0][i].enabled){+/* Using 32bit counters, reset max - count */+reset_value[i]=0xFFFFFFFF-ctr[i].count;+set_pm_event(i,+pmc_cntrl[0][i].evnts,+pmc_cntrl[0][i].masks);++/* global, used by cell_cpu_setup */+ctr_enabled|=(1<<i);+}+}++/* initialize the previous counts for the virtual cntrs */+for_each_online_cpu(cpu)+for(i=0;i<num_counters;++i){+per_cpu(pmc_values,cpu)[i]=reset_value[i];+}+out:+;+}++/* This function is called once for each cpu */+staticvoidcell_cpu_setup(structop_counter_config*cntr)+{+u32cpu=smp_processor_id();+u32num_enabled=0;+inti;++/* There is one performance monitor per processor chip (i.e. node),+*soweonlyneedtoperformthisfunctiononcepernode.+*/+if(cbe_get_hw_thread_id(cpu))+gotoout;++if(pm_rtas_token==RTAS_UNKNOWN_SERVICE){+printk(KERN_WARNING"%s: RTAS_UNKNOWN_SERVICE\n",+__FUNCTION__);+gotoout;+}++/* Stop all counters */+cbe_disable_pm(cpu);+cbe_disable_pm_interrupts(cpu);++cbe_write_pm(cpu,pm_interval,0);+cbe_write_pm(cpu,pm_start_stop,0);+cbe_write_pm(cpu,group_control,pm_regs.group_control);+cbe_write_pm(cpu,debug_bus_control,pm_regs.debug_bus_control);+write_pm_cntrl(cpu,&pm_regs.pm_cntrl);++for(i=0;i<num_counters;++i){+if(ctr_enabled&(1<<i)){+pm_signal[num_enabled].cpu=cbe_cpu_to_node(cpu);+num_enabled++;+}+}++pm_rtas_activate_signals(cbe_cpu_to_node(cpu),num_enabled);+out:+;+}++staticvoidcell_global_start(structop_counter_config*ctr)+{+u32cpu;+u32interrupt_mask=0;+u32i;++/* This routine gets called once for the system.+*Thereisoneperformancemonitorpernode,sowe+*onlyneedtoperformthisfunctiononcepernode.+*/+for_each_online_cpu(cpu){+if(cbe_get_hw_thread_id(cpu))+continue;++interrupt_mask=0;++for(i=0;i<num_counters;++i){+if(ctr_enabled&(1<<i)){+cbe_write_ctr(cpu,i,reset_value[i]);+enable_ctr(cpu,i,pm_regs.pm07_cntrl);+interrupt_mask|=+CBE_PM_CTR_OVERFLOW_INTR(i);+}else{+/* Disable counter */+cbe_write_pm07_control(cpu,i,0);+}+}++cbe_clear_pm_interrupts(cpu);+cbe_enable_pm_interrupts(cpu,hdw_thread,interrupt_mask);+cbe_enable_pm(cpu);+}++virt_cntr_inter_mask=interrupt_mask;+oprofile_running=1;+smp_wmb();++/* NOTE: start_virt_cntrs will result in cell_virtual_cntr() being+*executedwhichmanipulatesthePMU.Westartthe"virtual counter"+*heresothatwedonotneedtosynchronizeaccesstothePMUin+*theabovefor-loop.+*/+start_virt_cntrs();+}++staticvoidcell_global_stop(void)+{+intcpu;++/* This routine will be called once for the system.+*Thereisoneperformancemonitorpernode,sowe+*onlyneedtoperformthisfunctiononcepernode.+*/+del_timer_sync(&timer_virt_cntr);+oprofile_running=0;+smp_wmb();++for_each_online_cpu(cpu){+if(cbe_get_hw_thread_id(cpu))+continue;++cbe_sync_irq(cbe_cpu_to_node(cpu));+/* Stop the counters */+cbe_disable_pm(cpu);++/* Deactivate the signals */+pm_rtas_reset_signals(cbe_cpu_to_node(cpu));++/* Deactivate interrupts */+cbe_disable_pm_interrupts(cpu);+}+}++staticvoid+cell_handle_interrupt(structpt_regs*regs,structop_counter_config*ctr)+{+u32cpu;+u64pc;+intis_kernel;+unsignedlongflags=0;+u32interrupt_mask;+inti;++cpu=smp_processor_id();++/* Need to make sure the interrupt handler and the virt counter+*routinearenotrunningatthesametime.Seethe+*cell_virtual_cntr()routineforadditionalcomments.+*/+spin_lock_irqsave(&virt_cntr_lock,flags);++/* Need to disable and reenable the performance counters+*togetthedesiredbehaviorfromthehardware.This+*ishardwarespecific.+*/++cbe_disable_pm(cpu);++interrupt_mask=cbe_clear_pm_interrupts(cpu);++/* If the interrupt mask has been cleared, then the virt cntr+*hasclearedtheinterrupt.Whenthethreadthatgenerated+*theinterruptisrestored,thedatacountwillberestoredto+*0xffffff0tocausetheinterrupttoberegenerated.+*/++if((oprofile_running==1)&&(interrupt_mask!=0)){+pc=regs->nip;+is_kernel=is_kernel_addr(pc);++for(i=0;i<num_counters;++i){+if((interrupt_mask&CBE_PM_CTR_OVERFLOW_INTR(i))+&&ctr[i].enabled){+oprofile_add_pc(pc,is_kernel,i);+cbe_write_ctr(cpu,i,reset_value[i]);+}+}++/* The counters were frozen by the interrupt.+*Reenabletheinterruptandrestartthecounters.+*Iftherewasaracebetweentheinterrupthandlerand+*thevirtualcounterroutine.Thevirutalcounter+*routinemayhaveclearedtheinterrupts.Hencemust+*usethevirt_cntr_inter_masktore-enabletheinterrupts.+*/+cbe_enable_pm_interrupts(cpu,hdw_thread,+virt_cntr_inter_mask);++/* The writes to the various performance counters only writes+*toalatch.Thenewvalues(interruptsettingbits,reset+*countervalueetc.)arenotcopiedtotheactualregisters+*untiltheperformancemonitorisenabled.Inordertoget+*thistoworkasdesired,thepermormancemonitorneedsto+*bedisabledwhilewrittingtothelatches.Thisisa+*HWdesignissue.+*/+cbe_enable_pm(cpu);+}+spin_unlock_irqrestore(&virt_cntr_lock,flags);+}++structop_powerpc_modelop_model_cell={+.reg_setup=cell_reg_setup,+.cpu_setup=cell_cpu_setup,+.global_start=cell_global_start,+.global_stop=cell_global_stop,+.handle_interrupt=cell_handle_interrupt,+};
From: Kevin Corry <redacted>
Move some PMU-related macros and function prototypes from cbe_regs.h
and pmu.h in arch/powerpc/platforms/cell/ to a new header at
include/asm-powerpc/cell-pmu.h
This is cleaner to use from the oprofile code, since that sits in
arch/powerpc/oprofile, not in the cell platform directory.
Signed-off-by: Kevin Corry <redacted>
Signed-off-by: Arnd Bergmann <redacted>
Index: linux-2.6/arch/powerpc/platforms/cell/cbe_regs.h
===================================================================
@@ -1,57 +0,0 @@-/*- * Cell Broadband Engine Performance Monitor- *- * (C) Copyright IBM Corporation 2001,2006- *- * Author:- * David Erb (djerb@us.ibm.com)- * Kevin Corry (kevcorry@us.ibm.com)- *- * This program is free software; you can redistribute it and/or modify- * it under the terms of the GNU General Public License as published by- * the Free Software Foundation; either version 2, or (at your option)- * any later version.- *- * This program is distributed in the hope that it will be useful,- * but WITHOUT ANY WARRANTY; without even the implied warranty of- * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the- * GNU General Public License for more details.- *- * You should have received a copy of the GNU General Public License- * along with this program; if not, write to the Free Software- * Foundation, Inc., 675 Mass Ave, Cambridge, MA 02139, USA.- */--#ifndef __PERFMON_H__-#define __PERFMON_H__--enum pm_reg_name {- group_control,- debug_bus_control,- trace_address,- ext_tr_timer,- pm_status,- pm_control,- pm_interval,- pm_start_stop,-};--extern u32 cbe_read_phys_ctr(u32 cpu, u32 phys_ctr);-extern void cbe_write_phys_ctr(u32 cpu, u32 phys_ctr, u32 val);-extern u32 cbe_read_ctr(u32 cpu, u32 ctr);-extern void cbe_write_ctr(u32 cpu, u32 ctr, u32 val);--extern u32 cbe_read_pm07_control(u32 cpu, u32 ctr);-extern void cbe_write_pm07_control(u32 cpu, u32 ctr, u32 val);-extern u32 cbe_read_pm (u32 cpu, enum pm_reg_name reg);-extern void cbe_write_pm (u32 cpu, enum pm_reg_name reg, u32 val);--extern u32 cbe_get_ctr_size(u32 cpu, u32 phys_ctr);-extern void cbe_set_ctr_size(u32 cpu, u32 phys_ctr, u32 ctr_size);--extern void cbe_enable_pm(u32 cpu);-extern void cbe_disable_pm(u32 cpu);--extern void cbe_read_trace_buffer(u32 cpu, u64 *buf);--#endif
From: Dwayne Grant McConnell <redacted>
This patch implements read only access to
/mbox_info - SPU Write Outbound Mailbox
/ibox_info - SPU Write Outbound Interrupt Mailbox
/wbox_info - SPU Read Inbound Mailbox
These files are used by gdb in order to look into the current mailbox
queues without changing the contents at the same time. They are
not meant for general programming use, since the access requires
a context save and is therefore rather slow.
It would be good to complement this patch with one that adds
write support as well.
Signed-off-by: Dwayne Grant McConnell <redacted>
Signed-off-by: Arnd Bergmann <redacted>
---
Dwayne Grant McConnell [off-list ref]
Lotus Notes Mail: Dwayne McConnell [Mail]/Austin/IBM@IBMUS
Lotus Notes Calendar: Dwayne McConnell [Calendar]/Austin/IBM@IBMUS
Index: linux-2.6/arch/powerpc/platforms/cell/spufs/file.c
===================================================================
From: Jeremy Kerr <redacted>
In order to fit with the "don't-run-spus-outside-of-spu_run" model, this
patch starts the isolated-mode loader in spu_run, rather than
spu_create. If spu_run is passed an isolated-mode context that isn't in
isolated mode state, it will run the loader.
This fixes potential races with the isolated SPE app doing a
stop-and-signal before the PPE has called spu_run: bugzilla #29111.
Also (in conjunction with a mambo patch), this addresses #28565, as we
always set the runcntrl register when entering spu_run.
It is up to libspe to ensure that isolated-mode apps are cleaned up
after running to completion - ie, put the app through the "ISOLATE EXIT"
state (see Ch11 of the CBEA).
Signed-off-by: Jeremy Kerr <redacted>
Signed-off-by: Arnd Bergmann <redacted>
---
Index: linux-2.6/arch/powerpc/platforms/cell/spufs/file.c
===================================================================
@@ -235,102 +233,6 @@ struct file_operations spufs_context_fop.fsync=simple_sync_file,};-staticintspu_setup_isolated(structspu_context*ctx)-{-intret;-u64__iomem*mfc_cntl;-u64sr1;-u32status;-unsignedlongtimeout;-constu32status_loading=SPU_STATUS_RUNNING-|SPU_STATUS_ISOLATED_STATE|SPU_STATUS_ISOLATED_LOAD_STATUS;--if(!isolated_loader)-return-ENODEV;--/* prevent concurrent operation with spu_run */-down(&ctx->run_sema);-ctx->ops->master_start(ctx);--ret=spu_acquire_exclusive(ctx);-if(ret)-gotoout;--mfc_cntl=&ctx->spu->priv2->mfc_control_RW;--/* purge the MFC DMA queue to ensure no spurious accesses before we-*enterkernelmode*/-timeout=jiffies+HZ;-out_be64(mfc_cntl,MFC_CNTL_PURGE_DMA_REQUEST);-while((in_be64(mfc_cntl)&MFC_CNTL_PURGE_DMA_STATUS_MASK)-!=MFC_CNTL_PURGE_DMA_COMPLETE){-if(time_after(jiffies,timeout)){-printk(KERN_ERR"%s: timeout flushing MFC DMA queue\n",-__FUNCTION__);-ret=-EIO;-gotoout_unlock;-}-cond_resched();-}--/* put the SPE in kernel mode to allow access to the loader */-sr1=spu_mfc_sr1_get(ctx->spu);-sr1&=~MFC_STATE1_PROBLEM_STATE_MASK;-spu_mfc_sr1_set(ctx->spu,sr1);--/* start the loader */-ctx->ops->signal1_write(ctx,(unsignedlong)isolated_loader>>32);-ctx->ops->signal2_write(ctx,-(unsignedlong)isolated_loader&0xffffffff);--ctx->ops->runcntl_write(ctx,-SPU_RUNCNTL_RUNNABLE|SPU_RUNCNTL_ISOLATE);--ret=0;-timeout=jiffies+HZ;-while(((status=ctx->ops->status_read(ctx))&status_loading)==-status_loading){-if(time_after(jiffies,timeout)){-printk(KERN_ERR"%s: timeout waiting for loader\n",-__FUNCTION__);-ret=-EIO;-gotoout_drop_priv;-}-cond_resched();-}--if(!(status&SPU_STATUS_RUNNING)){-/* If isolated LOAD has failed: run SPU, we will get a stop-and-*signallater.*/-pr_debug("%s: isolated LOAD failed\n",__FUNCTION__);-ctx->ops->runcntl_write(ctx,SPU_RUNCNTL_RUNNABLE);-ret=-EACCES;--}elseif(!(status&SPU_STATUS_ISOLATED_STATE)){-/* This isn't allowed by the CBEA, but check anyway */-pr_debug("%s: SPU fell out of isolated mode?\n",__FUNCTION__);-ctx->ops->runcntl_write(ctx,SPU_RUNCNTL_STOP);-ret=-EINVAL;-}--out_drop_priv:-/* Finished accessing the loader. Drop kernel mode */-sr1|=MFC_STATE1_PROBLEM_STATE_MASK;-spu_mfc_sr1_set(ctx->spu,sr1);--out_unlock:-spu_release_exclusive(ctx);-out:-ctx->ops->master_stop(ctx);-up(&ctx->run_sema);-returnret;-}--intspu_recycle_isolated(structspu_context*ctx)-{-returnspu_setup_isolated(ctx);-}-staticintspufs_mkdir(structinode*dir,structdentry*dentry,unsignedintflags,intmode)
@@ -439,15 +341,6 @@ static int spufs_create_context(struct iout_unlock:mutex_unlock(&inode->i_mutex);out:-if(ret>=0&&(flags&SPU_CREATE_ISOLATE)){-intsetup_err=spu_setup_isolated(-SPUFS_I(dentry->d_inode)->i_ctx);-/* FIXME: clean up context again on failure to avoid-leak.*/-if(setup_err)-ret=setup_err;-}-dput(dentry);returnret;}
@@ -51,21 +53,122 @@ static inline int spu_stopped(struct spureturn(!(*stat&0x1)||pte_fault||spu->class_0_pending)?1:0;}+staticintspu_setup_isolated(structspu_context*ctx)+{+intret;+u64__iomem*mfc_cntl;+u64sr1;+u32status;+unsignedlongtimeout;+constu32status_loading=SPU_STATUS_RUNNING+|SPU_STATUS_ISOLATED_STATE|SPU_STATUS_ISOLATED_LOAD_STATUS;++if(!isolated_loader)+return-ENODEV;++ret=spu_acquire_exclusive(ctx);+if(ret)+gotoout;++mfc_cntl=&ctx->spu->priv2->mfc_control_RW;++/* purge the MFC DMA queue to ensure no spurious accesses before we+*enterkernelmode*/+timeout=jiffies+HZ;+out_be64(mfc_cntl,MFC_CNTL_PURGE_DMA_REQUEST);+while((in_be64(mfc_cntl)&MFC_CNTL_PURGE_DMA_STATUS_MASK)+!=MFC_CNTL_PURGE_DMA_COMPLETE){+if(time_after(jiffies,timeout)){+printk(KERN_ERR"%s: timeout flushing MFC DMA queue\n",+__FUNCTION__);+ret=-EIO;+gotoout_unlock;+}+cond_resched();+}++/* put the SPE in kernel mode to allow access to the loader */+sr1=spu_mfc_sr1_get(ctx->spu);+sr1&=~MFC_STATE1_PROBLEM_STATE_MASK;+spu_mfc_sr1_set(ctx->spu,sr1);++/* start the loader */+ctx->ops->signal1_write(ctx,(unsignedlong)isolated_loader>>32);+ctx->ops->signal2_write(ctx,+(unsignedlong)isolated_loader&0xffffffff);++ctx->ops->runcntl_write(ctx,+SPU_RUNCNTL_RUNNABLE|SPU_RUNCNTL_ISOLATE);++ret=0;+timeout=jiffies+HZ;+while(((status=ctx->ops->status_read(ctx))&status_loading)==+status_loading){+if(time_after(jiffies,timeout)){+printk(KERN_ERR"%s: timeout waiting for loader\n",+__FUNCTION__);+ret=-EIO;+gotoout_drop_priv;+}+cond_resched();+}++if(!(status&SPU_STATUS_RUNNING)){+/* If isolated LOAD has failed: run SPU, we will get a stop-and+*signallater.*/+pr_debug("%s: isolated LOAD failed\n",__FUNCTION__);+ctx->ops->runcntl_write(ctx,SPU_RUNCNTL_RUNNABLE);+ret=-EACCES;++}elseif(!(status&SPU_STATUS_ISOLATED_STATE)){+/* This isn't allowed by the CBEA, but check anyway */+pr_debug("%s: SPU fell out of isolated mode?\n",__FUNCTION__);+ctx->ops->runcntl_write(ctx,SPU_RUNCNTL_STOP);+ret=-EINVAL;+}++out_drop_priv:+/* Finished accessing the loader. Drop kernel mode */+sr1|=MFC_STATE1_PROBLEM_STATE_MASK;+spu_mfc_sr1_set(ctx->spu,sr1);++out_unlock:+spu_release_exclusive(ctx);+out:+returnret;+}+staticinlineintspu_run_init(structspu_context*ctx,u32*npc){intret;unsignedlongruncntl=SPU_RUNCNTL_RUNNABLE;-if((ret=spu_acquire_runnable(ctx))!=0)+ret=spu_acquire_runnable(ctx);+if(ret)returnret;-/* if we're in isolated mode, we would have started the SPU-*earlier,sodon'tdoitagainnow.*/-if(!(ctx->flags&SPU_CREATE_ISOLATE)){+if(ctx->flags&SPU_CREATE_ISOLATE){+if(!(ctx->ops->status_read(ctx)&SPU_STATUS_ISOLATED_STATE)){+/* Need to release ctx, because spu_setup_isolated will+*acquireitexclusively.+*/+spu_release(ctx);+ret=spu_setup_isolated(ctx);+if(!ret)+ret=spu_acquire_runnable(ctx);+}++/* if userspace has set the runcntrl register (eg, to issue an+*isolatedexit),weneedtore-setithere*/+runcntl=ctx->ops->runcntl_read(ctx)&+(SPU_RUNCNTL_RUNNABLE|SPU_RUNCNTL_ISOLATE);+if(runcntl==0)+runcntl=SPU_RUNCNTL_RUNNABLE;+}elsectx->ops->npc_write(ctx,*npc);-ctx->ops->runcntl_write(ctx,runcntl);-}-return0;++ctx->ops->runcntl_write(ctx,runcntl);+returnret;}staticinlineintspu_run_fini(structspu_context*ctx,u32*npc,
From: Dwayne Grant McConnell <redacted>
The /lslr file gives read access to the SPU_LSLR register in hex; 0x3fff
for example The /dma_info file provides read access to the SPU Command
Queue in a binary format. The /proxydma_info files provides read access
access to the Proxy Command Queue in a binary format. The spu_info.h
file provides data structures for interpreting the binary format of
/dma_info and /proxydma_info.
Signed-off-by: Dwayne Grant McConnell <redacted>
Signed-off-by: Arnd Bergmann <redacted>
---
Index: linux-2.6/arch/powerpc/platforms/cell/spufs/backing_ops.c
===================================================================
From: Geoff Levand <redacted>
Change the definition of powerpc's cond_syscall() to use the standard gcc
weak attribute specifier which provides proper support for C linkage as
needed by spu_syscall_table[].
Fixes this powerpc build error with CONFIG_SPU_FS=y, CONFIG_PPC_RTAS=n:
arch/powerpc/platforms/built-in.o: undefined reference to `ppc_rtas'
Signed-off-by: Geoff Levand <redacted>
Signed-off-by: Arnd Bergmann <redacted>
---
include/asm-powerpc/unistd.h | 12 ++----------
1 file changed, 2 insertions(+), 10 deletions(-)
Index: linux-2.6/include/asm-powerpc/unistd.h
===================================================================
When we attempt an MFC DMA to an unmapped address, the event
returned from spu_run should be SPE_EVENT_SPE_DATA_STORAGE,
not SPE_EVENT_INVALID_DMA.
Signed-off-by: Arnd Bergmann <redacted>
---
Index: linux-2.6/arch/powerpc/platforms/cell/spufs/run.c
===================================================================
From: Dwayne Grant McConnell <redacted>
This patch removes the /spu_tag_mask file from spufs. The data provided by
this file is also available from the /dma_info file in the dma_info_mask
of the spu_dma_info struct.
The file was intended to be used by gdb, but that never used it, and
now it has been replaced with the more verbose dma_info file.
Signed-off-by: Dwayne Grant McConnell <redacted>
Signed-off-by: Arnd Bergmann <redacted>
---
Index: linux-2.6/arch/powerpc/platforms/cell/spufs/file.c
===================================================================
From: Masato Noguchi <redacted>
When there is pending signals, current spufs_run_spu() always returns
-ERESTARTSYS and it is called again automatically.
But, if spe already stopped by stop-and-signal or halt instruction,
returning -ERESTARTSYS makes stop-and-signal/halt lost and
spu run over the end-point.
For your convenience, I attached a sample code to restage this bug.
If there is no bug, printed NPC will be 0x4000.
Signed-off-by: Masato Noguchi <redacted>
Signed-off-by: Arnd Bergmann <redacted>
---
run.c | 28 ++++++++++++++++++----------
1 files changed, 18 insertions(+), 10 deletions(-)
Index: linux-2.6/arch/powerpc/platforms/cell/spufs/run.c
===================================================================
I got a bug report that I believe might be fixed by this
patch. The problem seems to be that with soft-disabled
interrupts in power_save, we can still get external exceptions
on Cell, even if we are in pause(0) a.k.a. sleep state.
When the CPU really wakes up through the 0x100 (system reset)
vector, while we have already started processing the 0x500
(external) exception, we get a panic in unrecoverable_exception()
because of the lost state.
This occurred in Systemsim for Cell, but as far as I can see,
it can theoretically occur on any machine that uses the
system reset exception to get out of sleep state.
Signed-off-by: Arnd Bergmann <redacted>
Index: linux-2.6/arch/powerpc/kernel/idle.c
===================================================================
From: Kevin Corry <redacted>
The following routines are added to arch/powerpc/platforms/cell/pmu.c:
cbe_clear_pm_interrupts()
cbe_enable_pm_interrupts()
cbe_disable_pm_interrupts()
cbe_query_pm_interrupts()
cbe_pm_irq()
cbe_init_pm_irq()
This also adds a routine in arch/powerpc/platforms/cell/interrupt.c and
some macros in cbe_regs.h to manipulate the IIC_IR register:
iic_set_interrupt_routing()
Signed-off-by: Kevin Corry <redacted>
Signed-off-by: Carl Love <redacted>
Signed-off-by: Arnd Bergmann <redacted>
Index: linux-2.6/arch/powerpc/platforms/cell/pmu.c
===================================================================
@@ -338,3 +340,71 @@ void cbe_read_trace_buffer(u32 cpu, u64 }EXPORT_SYMBOL_GPL(cbe_read_trace_buffer);+/*+*Enabling/disablinginterruptsfortheentireperformancemonitoringunit.+*/++u32cbe_query_pm_interrupts(u32cpu)+{+returncbe_read_pm(cpu,pm_status);+}+EXPORT_SYMBOL_GPL(cbe_query_pm_interrupts);++u32cbe_clear_pm_interrupts(u32cpu)+{+/* Reading pm_status clears the interrupt bits. */+returncbe_query_pm_interrupts(cpu);+}+EXPORT_SYMBOL_GPL(cbe_clear_pm_interrupts);++voidcbe_enable_pm_interrupts(u32cpu,u32thread,u32mask)+{+/* Set which node and thread will handle the next interrupt. */+iic_set_interrupt_routing(cpu,thread,0);++/* Enable the interrupt bits in the pm_status register. */+if(mask)+cbe_write_pm(cpu,pm_status,mask);+}+EXPORT_SYMBOL_GPL(cbe_enable_pm_interrupts);++voidcbe_disable_pm_interrupts(u32cpu)+{+cbe_clear_pm_interrupts(cpu);+cbe_write_pm(cpu,pm_status,0);+}+EXPORT_SYMBOL_GPL(cbe_disable_pm_interrupts);++staticirqreturn_tcbe_pm_irq(intirq,void*dev_id,structpt_regs*regs)+{+perf_irq(regs);+returnIRQ_HANDLED;+}++int__initcbe_init_pm_irq(void)+{+unsignedintirq;+intrc,node;++for_each_node(node){+irq=irq_create_mapping(NULL,IIC_IRQ_IOEX_PMI|+(node<<IIC_IRQ_NODE_SHIFT));+if(irq==NO_IRQ){+printk("ERROR: Unable to allocate irq for node %d\n",+node);+return-EINVAL;+}++rc=request_irq(irq,cbe_pm_irq,+IRQF_DISABLED,"cbe-pmu-0",NULL);+if(rc){+printk("ERROR: Request for irq on node %d failed\n",+node);+returnrc;+}+}++return0;+}+arch_initcall(cbe_init_pm_irq);+
@@ -396,3 +396,19 @@ void __init iic_init_IRQ(void)/* Enable on current CPU */iic_setup_cpu();}++voidiic_set_interrupt_routing(intcpu,intthread,intpriority)+{+structcbe_iic_regs__iomem*iic_regs=cbe_get_cpu_iic_regs(cpu);+u64iic_ir=0;+intnode=cpu>>1;++/* Set which node and thread will handle the next interrupt */+iic_ir|=CBE_IIC_IR_PRIO(priority)|+CBE_IIC_IR_DEST_NODE(node);+if(thread==0)+iic_ir|=CBE_IIC_IR_DEST_UNIT(CBE_IIC_IR_PT_0);+else+iic_ir|=CBE_IIC_IR_DEST_UNIT(CBE_IIC_IR_PT_1);+out_be64(&iic_regs->iic_ir,iic_ir);+}
From: Dwayne Grant McConnell <redacted>
This patch adds SPU elf notes to the coredump. It creates a separate note
for each of /regs, /fpcr, /lslr, /decr, /decr_status, /mem, /signal1,
/signal1_type, /signal2, /signal2_type, /event_mask, /event_status,
/mbox_info, /ibox_info, /wbox_info, /dma_info, /proxydma_info, /object-id.
A new macro, ARCH_HAVE_EXTRA_NOTES, was created for architectures to
specify they have extra elf core notes.
A new macro, ELF_CORE_EXTRA_NOTES_SIZE, was created so the size of the
additional notes could be calculated and added to the notes phdr entry.
A new macro, ELF_CORE_WRITE_EXTRA_NOTES, was created so the new notes
would be written after the existing notes.
The SPU coredump code resides in spufs. Stub functions are provided in the
kernel which are hooked into the spufs code which does the actual work via
register_arch_coredump_calls().
A new set of __spufs_<file>_read/get() functions was provided to allow the
coredump code to read from the spufs files without having to lock the
SPU context for each file read from.
Cc: <redacted>
Signed-off-by: Dwayne Grant McConnell <redacted>
Signed-off-by: Arnd Bergmann <redacted>
---
Cc'ing linux-arch because I couldn't find the right maintainer
for fs/binfmt_elf.c, and the patch adds architecture specific hooks.
Index: linux-2.6/arch/powerpc/platforms/cell/Makefile
===================================================================
@@ -1499,12 +1570,18 @@ static u64 spufs_id_get(void *data)}DEFINE_SIMPLE_ATTRIBUTE(spufs_id_ops,spufs_id_get,NULL,"0x%llx\n")-staticu64spufs_object_id_get(void*data)+staticu64__spufs_object_id_get(void*data){structspu_context*ctx=data;returnctx->object_id;}+staticu64spufs_object_id_get(void*data)+{+/* FIXME: Should there really be no locking here? */+return__spufs_object_id_get(data);+}+staticvoidspufs_object_id_set(void*data,u64id){structspu_context*ctx=data;
@@ -1,7 +1,7 @@obj-y+=switch.oobj-$(CONFIG_SPU_FS)+=spufs.o-spufs-y+=inode.ofile.ocontext.osyscalls.o+spufs-y+=inode.ofile.ocontext.osyscalls.ocoredump.ospufs-y+=sched.obacking_ops.ohw_ops.orun.ogang.o# Rules to build switch.o with the help of SPU tool chain
@@ -1582,6 +1582,10 @@ static int elf_core_dump(long signr, strsz+=thread_status_size;+#ifdef ELF_CORE_WRITE_EXTRA_NOTES+sz+=ELF_CORE_EXTRA_NOTES_SIZE;+#endif+fill_elf_note_phdr(&phdr,sz,offset);offset+=sz;DUMP_WRITE(&phdr,sizeof(phdr));
@@ -1622,6 +1626,10 @@ static int elf_core_dump(long signr, strif(!writenote(notes+i,file,&foffset))gotoend_coredump;+#ifdef ELF_CORE_WRITE_EXTRA_NOTES+ELF_CORE_WRITE_EXTRA_NOTES;+#endif+/* write out the thread status notes section */list_for_each(t,&thread_list){structelf_thread_status*tmp=
@@ -411,4 +411,17 @@ do { \/* Keep this the last entry. */#define R_PPC64_NUM 107+#ifdef CONFIG_PPC_CELL+/* Notes used in ET_CORE. Note name is "SPU/<fd>/<filename>". */+#define NT_SPU 1++externintarch_notes_size(void);+externvoidarch_write_notes(structfile*file);++#define ELF_CORE_EXTRA_NOTES_SIZE arch_notes_size()+#define ELF_CORE_WRITE_EXTRA_NOTES arch_write_notes(file)++#define ARCH_HAVE_EXTRA_ELF_NOTES+#endif /* CONFIG_PPC_CELL */+#endif /* _ASM_POWERPC_ELF_H */
When the user changes the runcontrol register, an SPU might be
running without a process being attached to it and waiting for
events. In order to prevent this, make sure we always disable
the priv1 master control when we're not inside of spu_run.
Signed-off-by: Arnd Bergmann <redacted>
---
Index: linux-2.6/arch/powerpc/platforms/cell/spufs/hw_ops.c
===================================================================
@@ -122,29 +122,29 @@ void spu_unmap_mappings(struct spu_conteintspu_acquire_exclusive(structspu_context*ctx){-intret=0;+intret=0;-down_write(&ctx->state_sema);-/* ctx is about to be freed, can't acquire any more */-if(!ctx->owner){-ret=-EINVAL;-gotoout;-}--if(ctx->state==SPU_STATE_SAVED){-ret=spu_activate(ctx,0);-if(ret)-gotoout;-ctx->state=SPU_STATE_RUNNABLE;-}else{-/* We need to exclude userspace access to the context. */-spu_unmap_mappings(ctx);-}+down_write(&ctx->state_sema);+/* ctx is about to be freed, can't acquire any more */+if(!ctx->owner){+ret=-EINVAL;+gotoout;+}++if(ctx->state==SPU_STATE_SAVED){+ret=spu_activate(ctx,0);+if(ret)+gotoout;+ctx->state=SPU_STATE_RUNNABLE;+}else{+/* We need to exclude userspace access to the context. */+spu_unmap_mappings(ctx);+}out:-if(ret)-up_write(&ctx->state_sema);-returnret;+if(ret)+up_write(&ctx->state_sema);+returnret;}intspu_acquire_runnable(structspu_context*ctx)
@@ -435,6 +442,8 @@ out:if(ret>=0&&(flags&SPU_CREATE_ISOLATE)){intsetup_err=spu_setup_isolated(SPUFS_I(dentry->d_inode)->i_ctx);+/* FIXME: clean up context again on failure to avoid+leak.*/if(setup_err)ret=setup_err;}
From: Geoff Levand <redacted>
Replace the use of the platform specific variable spu.nid with the
platform independednt variable spu.node.
Signed-off-by: Geoff Levand <redacted>
Signed-off-by: Arnd Bergmann <redacted>
---
Index: linux-2.6/arch/powerpc/platforms/cell/spu_base.c
===================================================================
From: Jeremy Kerr <redacted>
This change adds a read accessor for the SPE problem-state run control
register.
This is required for for applying (userspace) changes made to the run
control register while the SPE is stopped - simply asserting the master
run control bit is not sufficient. My next patch for isolated-mode
setup requires this.
Signed-off-by: Jeremy Kerr <jk@ozlabs.org>
Signed-off-by: Arnd Bergmann <redacted>
---
arch/powerpc/platforms/cell/spufs/backing_ops.c | 6 ++++++
arch/powerpc/platforms/cell/spufs/hw_ops.c | 6 ++++++
arch/powerpc/platforms/cell/spufs/spufs.h | 1 +
3 files changed, 13 insertions(+)
Index: linux-2.6/arch/powerpc/platforms/cell/spufs/backing_ops.c
===================================================================
From: Dwayne Grant McConnell <redacted>
This patches changes /npc, /decr, /decr_status, /spu_tag_mask,
/event_mask, /event_status, and /srr0 files to provide output according to
the format string "0x%llx" instead of "%llx".
Before this patch some files used "0x%llx" and other used "%llx" which is
inconsistent and potentially confusing. A user might assume "%llx" numbers
were decimal if they happened to not contain any a-f digits. This change
will break any code cannot tolerate a leading 0x in the file contents. The
only known users of these files are the libspe but there might also be
some scripts which access these files. This risk is deemed acceptable for
future consistency.
Signed-off-by: Dwayne Grant McConnell <redacted>
Signed-off-by: Arnd Bergmann <redacted>
---
Dwayne Grant McConnell [off-list ref]
Lotus Notes Mail: Dwayne McConnell [Mail]/Austin/IBM@IBMUS
Lotus Notes Calendar: Dwayne McConnell [Calendar]/Austin/IBM@IBMUS
Index: linux-2.6/arch/powerpc/platforms/cell/spufs/file.c
===================================================================
When fixing spufs to map the 'mem' file backing store cacheable,
I incorrectly set the physical mapping to use both cache-inhibited
and guarded mapping, which resulted in a serious performance
degradation.
Debugged-by: Michael Ellerman [off-list ref]
Signed-off-by: Arnd Bergmann <redacted>
---
Index: linux-2.6/arch/powerpc/platforms/cell/spufs/file.c
===================================================================
From: Paul Mackerras <hidden> Date: 2006-11-20 22:19:34
Arnd Bergmann writes:
I got a bug report that I believe might be fixed by this
patch. The problem seems to be that with soft-disabled
interrupts in power_save, we can still get external exceptions
on Cell, even if we are in pause(0) a.k.a. sleep state.
[snip]
- local_irq_disable();
+ hard_irq_disable();
This would mean that any platform-specific power_save function that
wants to re-enable interrupts (as the pseries ones do) would have to
do hard_irq_enable instead of local_irq_enable. Also, I don't think
this change will be good on iSeries.
What we want is an irq-disable function that is like local_irq_disable
but also clears MSR_EE and the hard irq enabled flag (provided we
aren't running on iSeries).
Paul.
From: Benjamin Herrenschmidt <benh@kernel.crashing.org> Date: 2006-11-21 00:54:05
I got a bug report that I believe might be fixed by this
patch. The problem seems to be that with soft-disabled
interrupts in power_save, we can still get external exceptions
on Cell, even if we are in pause(0) a.k.a. sleep state.
When the CPU really wakes up through the 0x100 (system reset)
vector, while we have already started processing the 0x500
(external) exception, we get a panic in unrecoverable_exception()
because of the lost state.
This occurred in Systemsim for Cell, but as far as I can see,
it can theoretically occur on any machine that uses the
system reset exception to get out of sleep state.
Signed-off-by: Benjamin Herrenschmidt <benh@kernel.crashing.org>
---
What about that patch instead ?
Index: linux-cell/arch/powerpc/platforms/cell/pervasive.c
===================================================================
@@ -41,6 +41,15 @@staticvoidcbe_power_save(void){unsignedlongctrl,thread_switch_control;++/*+*Weneedtoharddisableinterrupts,butwealsoneedtomarkthem+*harddisabledinthePACAsothatthelocal_irq_enable()doneby+*ourcalleruponreturnpropertlyhardenables.+*/+hard_irq_disable();+get_paca()->hard_enabled=0;+ctrl=mfspr(SPRN_CTRLF);/* Enable DEC and EE interrupt request */
This is broken for !CELL. arch_write_notes(void) can't be called as
arch_write_notes(file).
cheers
--
Michael Ellerman
OzLabs, IBM Australia Development Lab
wwweb: http://michael.ellerman.id.au
phone: +61 2 6212 1183 (tie line 70 21183)
We do not inherit the earth from our ancestors,
we borrow it from our children. - S.M.A.R.T Person
On Tuesday 21 November 2006 01:53, Benjamin Herrenschmidt wrote:
+
+=A0=A0=A0=A0=A0=A0=A0/*
+=A0=A0=A0=A0=A0=A0=A0 * We need to hard disable interrupts, but we also =
need to mark them
+=A0=A0=A0=A0=A0=A0=A0 * hard disabled in the PACA so that the local_irq_=
enable() done by
+=A0=A0=A0=A0=A0=A0=A0 * our caller upon return propertly hard enables.
+=A0=A0=A0=A0=A0=A0=A0 */
+=A0=A0=A0=A0=A0=A0=A0hard_irq_disable();
+=A0=A0=A0=A0=A0=A0=A0get_paca()->hard_enabled =3D 0;
+
Yes, this looks good. Paul, please use this patch instead of mine.
Do we need to do the same thing for any of the other power_save functions?
IIRC, all new CPUs are supposed to use the same mechanism based on the
0x100 vector.
Arnd <><
From: Olof Johansson <hidden> Date: 2006-11-21 16:59:26
On Tue, 21 Nov 2006 11:14:45 +0100 Arnd Bergmann [off-list ref] wrote:
IIRC, all new CPUs are supposed to use the same mechanism based on the
0x100 vector.
It's only really affecting platforms without hypervisors though, which
aren't all that many (yet).
I don't have a problem dealing with it locally in my platform.
-Olof
From: Michael Ellerman <hidden> Date: 2007-03-01 06:18:14
On Mon, 2006-11-20 at 18:45 +0100, Arnd Bergmann wrote:
plain text document attachment (spufs-master-control.diff)
When the user changes the runcontrol register, an SPU might be
running without a process being attached to it and waiting for
events. In order to prevent this, make sure we always disable
the priv1 master control when we're not inside of spu_run.
Hi Arnd,
Sorry I didn't comment on this when you sent it, I wasn't paying enough
attention. This patch confuses me, you say we should make sure we always
disable the master control when we're not inside spu_run, but I see
several exit paths where we leave the master run bit enabled - or maybe
I'm reading it wrong.
I think I've also seen it happen:
[root@localhost dma5]# ./put-test
10963.13
[root@localhost dma5]# find /spu
/spu
[root@localhost dma5]# echo x > /proc/sysrq-trigger
SysRq : Entering xmon
0:mon> ss
..
Stopped spu 06, was running (mfc_sr1: 0x32 runcntl: 0x1)
Stopped spu 07, was running (mfc_sr1: 0x32 runcntl: 0x1)
..
cheers
--
Michael Ellerman
OzLabs, IBM Australia Development Lab
wwweb: http://michael.ellerman.id.au
phone: +61 2 6212 1183 (tie line 70 21183)
We do not inherit the earth from our ancestors,
we borrow it from our children. - S.M.A.R.T Person
On Thursday 01 March 2007, Michael Ellerman wrote:
On Mon, 2006-11-20 at 18:45 +0100, Arnd Bergmann wrote:
quoted
plain text document attachment (spufs-master-control.diff)
When the user changes the runcontrol register, an SPU might be
running without a process being attached to it and waiting for
events. In order to prevent this, make sure we always disable
the priv1 master control when we're not inside of spu_run.
Hi Arnd,
Sorry I didn't comment on this when you sent it, I wasn't paying enough
attention. This patch confuses me, you say we should make sure we always
disable the master control when we're not inside spu_run, but I see
several exit paths where we leave the master run bit enabled - or maybe
I'm reading it wrong.
I think you're right, there is at least one path that I now saw
getting out of spufs_run_spu incorrectly. In particular, when
spu_reacquire_runnable() fails, we never call the master stop,
which is a bug, but should happen very infrequently in practice.
Do you see another case where we end up with the same problem?
If not, I'll prepare a patch to fix this one case.
Arnd <><
From: Michael Ellerman <hidden> Date: 2007-03-02 10:14:02
On Thu, 2007-03-01 at 14:50 +0100, Arnd Bergmann wrote:
On Thursday 01 March 2007, Michael Ellerman wrote:
quoted
On Mon, 2006-11-20 at 18:45 +0100, Arnd Bergmann wrote:
quoted
plain text document attachment (spufs-master-control.diff)
When the user changes the runcontrol register, an SPU might be
running without a process being attached to it and waiting for
events. In order to prevent this, make sure we always disable
the priv1 master control when we're not inside of spu_run.
Hi Arnd,
Sorry I didn't comment on this when you sent it, I wasn't paying enough
attention. This patch confuses me, you say we should make sure we always
disable the master control when we're not inside spu_run, but I see
several exit paths where we leave the master run bit enabled - or maybe
I'm reading it wrong.
I think you're right, there is at least one path that I now saw
getting out of spufs_run_spu incorrectly. In particular, when
spu_reacquire_runnable() fails, we never call the master stop,
which is a bug, but should happen very infrequently in practice.
Do you see another case where we end up with the same problem?
That was the first one that caught my eye, but I wasn't sure about the
semantics of the error case.
There's also the error case for spu_run_init() which skips the master
stop. I guess that's ok because we've only set the master control in the
backing store, and the only way that will ever get propagated to an
actual spu is by coming back thorough spufs_run_spu().
What originally caught my eye on this was the output from xmon. When we
drop into xmon with no spu programs running and stop the spus, it
reports that they _all_ have the master run enabled, and some of them
have the runcntl enabled (those that have had spu programs run on them
since boot it seems).
It looks like the save/restore code sets the master bit in several
places, but never sets/clears the runcntl, which seems bogus to me.
So when we leave spufs_spu_run we do the master stop call:
spu_mfc_sr1_set: spu: c00000007ffdfc80 (15) sr1: 0x1b runcntl: 0x1
Call Trace:
[C00000000196BAA0] [C00000000000F920] .show_stack+0x68/0x1b0 (unreliable)
[C00000000196BB40] [D0000000001475C0] .spu_hw_master_stop+0xa8/0x170 [spufs]
[C00000000196BBE0] [D000000000148598] .spufs_run_spu+0x5ec/0x770 [spufs]
[C00000000196BCC0] [D000000000144BA0] .do_spu_run+0xb4/0x180 [spufs]
[C00000000196BD80] [C00000000003905C] .sys_spu_run+0xb0/0x108
[C00000000196BE30] [C000000000008634] syscall_exit+0x0/0x40
But then the save/restore code sets it back on?
spu_mfc_sr1_set: spu: c00000007ffdfc80 (15) sr1: 0x32 runcntl: 0x1
Call Trace:
[C00000000FF9B790] [C00000000000F920] .show_stack+0x68/0x1b0 (unreliable)
[C00000000FF9B830] [C00000000003B310] .spu_save+0x9b0/0x156c
[C00000000FF9B950] [D0000000001457FC] .spu_unbind_context+0x124/0x1a4 [spufs]
[C00000000FF9B9F0] [D0000000001458B0] .spu_deactivate+0x34/0x138 [spufs]
[C00000000FF9BA80] [D00000000014448C] .spu_acquire_saved+0x34/0x4c [spufs]
[C00000000FF9BB10] [D00000000014493C] .spu_forget+0x18/0x4c [spufs]
[C00000000FF9BBA0] [D00000000014030C] .spufs_dir_close+0x78/0xb4 [spufs]
[C00000000FF9BC50] [C0000000000B8594] .__fput+0x110/0x200
[C00000000FF9BD00] [C0000000000B4F38] .filp_close+0xac/0xd4
[C00000000FF9BD90] [C0000000000B68BC] .sys_close+0xc4/0x130
[C00000000FF9BE30] [C000000000008634] syscall_exit+0x0/0x40
cheers
--
Michael Ellerman
OzLabs, IBM Australia Development Lab
wwweb: http://michael.ellerman.id.au
phone: +61 2 6212 1183 (tie line 70 21183)
We do not inherit the earth from our ancestors,
we borrow it from our children. - S.M.A.R.T Person
There's also the error case for spu_run_init() which skips the master
stop. I guess that's ok because we've only set the master control in the
backing store, and the only way that will ever get propagated to an
actual spu is by coming back thorough spufs_run_spu().
Hmm, the correct way would be to switch off the master control in there,
afaics. Fixing it only in spu_run_init would mean that we also handle
the case of spu_reacquire_runnable along with it.
What originally caught my eye on this was the output from xmon. When we
drop into xmon with no spu programs running and stop the spus, it
reports that they _all_ have the master run enabled,
That looks right, there is no problem to have master control enabled,
as long as user space can't access the spu through a context that is
bound to it.
and some of them
have the runcntl enabled (those that have had spu programs run on them
since boot it seems).
While this sounds wrong. Maybe the runcntl is active on those that have
_not_ run since boot, which would make more sense. We should investigate
this.
It looks like the save/restore code sets the master bit in several
places, but never sets/clears the runcntl, which seems bogus to me.
So when we leave spufs_spu_run we do the master stop call:
spu_mfc_sr1_set: spu: c00000007ffdfc80 (15) sr1: 0x1b runcntl: 0x1
Call Trace:
[C00000000196BAA0] [C00000000000F920] .show_stack+0x68/0x1b0 (unreliable)
[C00000000196BB40] [D0000000001475C0] .spu_hw_master_stop+0xa8/0x170 [spufs]
[C00000000196BBE0] [D000000000148598] .spufs_run_spu+0x5ec/0x770 [spufs]
[C00000000196BCC0] [D000000000144BA0] .do_spu_run+0xb4/0x180 [spufs]
[C00000000196BD80] [C00000000003905C] .sys_spu_run+0xb0/0x108
[C00000000196BE30] [C000000000008634] syscall_exit+0x0/0x40
But then the save/restore code sets it back on?
Right, the context save code needs to enable master control in order to
run on the spu. However, that should be after all mappings to user space
have been discarded.
Arnd <><
From: Michael Ellerman <hidden> Date: 2007-03-07 08:58:33
On Mon, 2007-03-05 at 02:02 +0100, Arnd Bergmann wrote:
On Friday 02 March 2007, Michael Ellerman wrote:
quoted
There's also the error case for spu_run_init() which skips the master
stop. I guess that's ok because we've only set the master control in the
backing store, and the only way that will ever get propagated to an
actual spu is by coming back thorough spufs_run_spu().
Hmm, the correct way would be to switch off the master control in there,
afaics. Fixing it only in spu_run_init would mean that we also handle
the case of spu_reacquire_runnable along with it.
quoted
What originally caught my eye on this was the output from xmon. When we
drop into xmon with no spu programs running and stop the spus, it
reports that they _all_ have the master run enabled,
That looks right, there is no problem to have master control enabled,
as long as user space can't access the spu through a context that is
bound to it.
quoted
and some of them
have the runcntl enabled (those that have had spu programs run on them
since boot it seems).
While this sounds wrong. Maybe the runcntl is active on those that have
_not_ run since boot, which would make more sense. We should investigate
this.
No I'm pretty sure it's enabled on the ones that _have_ run since boot.
I'm booting up fresh, running two spu programs, and then I see two spus
with master and runcntl set.
cheers
--
Michael Ellerman
OzLabs, IBM Australia Development Lab
wwweb: http://michael.ellerman.id.au
phone: +61 2 6212 1183 (tie line 70 21183)
We do not inherit the earth from our ancestors,
we borrow it from our children. - S.M.A.R.T Person