From: Bryant G. Ly <hidden> Date: 2017-12-13 15:33:25
This patch series will enable SR-IOV on PowerVM. A specific set of
lids for PFW/PHYP is required. They are planned to release with
920 at the moment.
For IBM internal testers let me know of a system you want to test on
and we can put on the lids required or we can provide a system to run
the tests.
This patch depends on the three patches:
988fc3ba5653278a8c14d6ccf687371775930d2b
dae7253f9f78a731755ca20c66b2d2c40b86baea
608c0d8804ef3ca4cda8ec6ad914e47deb283d7b
Bryant G. Ly (7):
platform/pseries: Update VF config space after EEH
powerpc/kernel: Add uevents in EEH error/resume
platforms/pseries: Set eeh_pe of EEH_PE_VF type
powerpc/kernel Add EEH operations to notify resume
powerpc/kernel: Add EEH notify resume sysfs
pseries/pci: Associate PEs to VFs in configure SR-IOV
pseries/setup: Add Initialization of VF Bars
arch/powerpc/include/asm/eeh.h | 1 +
arch/powerpc/include/asm/pci-bridge.h | 5 +-
arch/powerpc/include/asm/pci.h | 2 +
arch/powerpc/kernel/eeh_driver.c | 9 +-
arch/powerpc/kernel/eeh_sysfs.c | 46 ++++++-
arch/powerpc/kernel/pci_of_scan.c | 2 +-
arch/powerpc/platforms/powernv/eeh-powernv.c | 3 +-
arch/powerpc/platforms/pseries/eeh_pseries.c | 192 ++++++++++++++++++++++++++-
arch/powerpc/platforms/pseries/pci.c | 156 +++++++++++++++++++++-
arch/powerpc/platforms/pseries/setup.c | 183 +++++++++++++++++++++++++
10 files changed, 589 insertions(+), 10 deletions(-)
--
2.14.3 (Apple Git-98)
From: Bryant G. Ly <hidden> Date: 2017-12-13 15:33:25
Add EEH platform operations for pseries to update VF
config space. With this change after EEH, the VF
will have updated config space for pseries platform.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/platforms/pseries/eeh_pseries.c | 85 +++++++++++++++++++++++++++-
1 file changed, 84 insertions(+), 1 deletion(-)
From: Bryant G. Ly <hidden> Date: 2017-12-13 15:33:30
To correctly use EEH code one has to make
sure that the EEH_PE_VF is set for dynamic created
VFs. Therefore this patch allocates an eeh_pe of
eeh type EEH_PE_VF and associates PE with parent.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/include/asm/pci-bridge.h | 5 ++++-
arch/powerpc/platforms/pseries/eeh_pseries.c | 9 ++++++++-
2 files changed, 12 insertions(+), 2 deletions(-)
@@ -211,7 +211,10 @@ struct pci_dn {unsignedint*pe_num_map;/* PE# for the first VF PE or array */boolm64_single_mode;/* Use M64 BAR in Single Mode */#define IODA_INVALID_M64 (-1)-int(*m64_map)[PCI_SRIOV_NUM_BARS];+union{+int(*m64_map)[PCI_SRIOV_NUM_BARS];+intlast_allow_rc;+};#endif /* CONFIG_PCI_IOV */intmps;/* Maximum Payload Size */structlist_headchild_list;
From: Bryant G. Ly <hidden> Date: 2017-12-13 15:33:34
Introduce a method for notify resume to be
called from sysfs. In this patch one can
now call notify resume from sysfs when
is supported by platform.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/kernel/eeh_sysfs.c | 46 ++++++++++++++++++++++++++++++++++++++++-
1 file changed, 45 insertions(+), 1 deletion(-)
From: Bryant G. Ly <hidden> Date: 2017-12-13 15:33:35
When pseries SR-IOV is enabled and after a PF driver
has resumed from EEH, platform has to be notified
of the event so the child VFs can be allowed to
resume their normal recovery path.
This patch makes the EEH operation allow unfreeze
platform dependent code and adds the call to
pseries EEH code.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/include/asm/eeh.h | 1 +
arch/powerpc/kernel/eeh_driver.c | 4 ++
arch/powerpc/platforms/powernv/eeh-powernv.c | 3 +-
arch/powerpc/platforms/pseries/eeh_pseries.c | 100 ++++++++++++++++++++++++++-
4 files changed, 106 insertions(+), 2 deletions(-)
@@ -798,6 +798,103 @@ static int pseries_eeh_restore_config(struct pci_dn *pdn)return0;}+#ifdef CONFIG_PCI_IOV+intpseries_send_allow_unfreeze(structeeh_pe*pe,+u16*vf_pe_array,intcur_vfs)+{+intrc,config_addr;+intibm_allow_unfreeze=rtas_token("ibm,open-sriov-allow-unfreeze");++config_addr=pe->config_addr;+spin_lock(&rtas_data_buf_lock);+memcpy(rtas_data_buf,vf_pe_array,RTAS_DATA_BUF_SIZE);+rc=rtas_call(ibm_allow_unfreeze,5,1,NULL,+config_addr,+BUID_HI(pe->phb->buid),+BUID_LO(pe->phb->buid),+rtas_data_buf,cur_vfs*sizeof(u16));+spin_unlock(&rtas_data_buf_lock);+if(rc)+pr_warn("%s: Failed to allow unfreeze for PHB#%x-PE#%x, rc=%x\n",+__func__,+pe->phb->global_number,+pe->config_addr,rc);+returnrc;+}++staticintpseries_call_allow_unfreeze(structeeh_dev*edev)+{+structeeh_pe*pe;+structpci_dn*pdn,*tmp,*parent,*physfn_pdn;+intcur_vfs,rc,vf_index;+u16*vf_pe_array;++vf_pe_array=kzalloc(RTAS_DATA_BUF_SIZE,GFP_KERNEL);+if(!vf_pe_array)+return-ENOMEM;++memset(vf_pe_array,0,RTAS_DATA_BUF_SIZE);+cur_vfs=0;+rc=0;+if(edev->pdev->is_physfn){+pe=eeh_dev_to_pe(edev);+cur_vfs=pci_num_vf(edev->pdev);+pdn=eeh_dev_to_pdn(edev);+parent=pdn->parent;+/* For each of its VF+*callallowunfreeze+*/+for(vf_index=0;vf_index<cur_vfs;vf_index++)+vf_pe_array[vf_index]=+be16_to_cpu(pdn->pe_num_map[vf_index]);++rc=pseries_send_allow_unfreeze(pe,vf_pe_array,cur_vfs);+pdn->last_allow_rc=rc;+for(vf_index=0;vf_index<cur_vfs;vf_index++){+list_for_each_entry_safe(pdn,tmp,&parent->child_list,+list){+if(pdn->busno+!=pci_iov_virtfn_bus(edev->pdev,+vf_index)||+pdn->devfn+!=pci_iov_virtfn_devfn(edev->pdev,+vf_index))+continue;+pdn->last_allow_rc=rc;+}+}+}else{+pdn=pci_get_pdn(edev->pdev);+vf_pe_array[0]=be16_to_cpu(pdn->pe_number);+physfn_pdn=pci_get_pdn(edev->physfn);+edev=pdn_to_eeh_dev(physfn_pdn);+pe=eeh_dev_to_pe(edev);+rc=pseries_send_allow_unfreeze(pe,vf_pe_array,1);+pdn->last_allow_rc=rc;+}++kfree(vf_pe_array);+returnrc;+}++staticintpseries_notify_resume(structpci_dn*pdn)+{+structeeh_dev*edev=pdn_to_eeh_dev(pdn);++if(!edev)+return-EEXIST;++if(rtas_token("ibm,open-sriov-allow-unfreeze")+==RTAS_UNKNOWN_SERVICE)+return-EINVAL;++if(edev->pdev->is_physfn||edev->pdev->is_virtfn)+returnpseries_call_allow_unfreeze(edev);++return0;+}+#endif+staticstructeeh_opspseries_eeh_ops={.name="pseries",.init=pseries_eeh_init,
From: Bryant G. Ly <hidden> Date: 2017-12-13 15:33:38
When enabling SR-IOV in pseries platform,
the VF bar properties for a PF are reported on
the device node in the device tree.
This patch adds the IOV Bar resources to Linux
structures from the device tree for later use
when configuring SR-IOV by PF driver.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/include/asm/pci.h | 2 +
arch/powerpc/kernel/pci_of_scan.c | 2 +-
arch/powerpc/platforms/pseries/setup.c | 183 +++++++++++++++++++++++++++++++++
3 files changed, 186 insertions(+), 1 deletion(-)
@@ -459,6 +459,181 @@ static void __init find_and_init_phbs(void)of_pci_check_probe_only();}+#ifdef CONFIG_PCI_IOV+enumrtas_iov_fw_value_map{+NUM_RES_PROPERTY=0,///< Number of Resources+LOW_INT=1,///< Lowest 32 bits of Address+START_OF_ENTRIES=2,///< Always start of entry+APERTURE_PROPERTY=2,///< Start of entry+ to Aperture Size+WDW_SIZE_PROPERTY=4,///< Start of entry+ to Window Size+NEXT_ENTRY=7///< Go to next entry on array+};++enumget_iov_fw_value_index{+BAR_ADDRS=1,///< Get Bar Address+APERTURE_SIZE=2,///< Get Aperture Size+WDW_SIZE=3///< Get Window Size+};++resource_size_tpseries_get_iov_fw_values(structpci_dev*dev,intresno,+enumget_iov_fw_value_indexvalue)+{+structvf_bar_wdw{+__be64addr;+__be64aperture_size;+__be64wdw_size;+};++structvf_bar_wdwwindow_avail[PCI_SRIOV_NUM_BARS];+constint*indexes;+structdevice_node*dn=pci_device_to_OF_node(dev);+inti,r,num_res;+resource_size_treturn_value;++indexes=of_get_property(dn,"ibm,open-sriov-vf-bar-info",NULL);+if(!indexes)+return0;++memset(window_avail,+0,sizeof(structvf_bar_wdw)*PCI_SRIOV_NUM_BARS);+return_value=0;+/*+*FirstelementinthearrayisthenumberofBars+*returned.Searchthroughthelisttofindthematching+*bar+*/+num_res=of_read_number(&indexes[NUM_RES_PROPERTY],1);+for(i=START_OF_ENTRIES,r=0;r<num_res&&r<PCI_SRIOV_NUM_BARS;+i+=NEXT_ENTRY,r++){+window_avail[r].addr=of_read_number(&indexes[i],2);+window_avail[r].aperture_size=+of_read_number(&indexes[i+APERTURE_PROPERTY],2);+window_avail[r].wdw_size=+of_read_number(&indexes[i+WDW_SIZE_PROPERTY],2);+}++switch(value){+caseBAR_ADDRS:+return_value=window_avail[resno].addr;+break;+caseAPERTURE_SIZE:+return_value=window_avail[resno].aperture_size;+break;+caseWDW_SIZE:+return_value=window_avail[resno].wdw_size;+break;+default:+break;+}+returnreturn_value;+}++voidof_pci_parse_vf_bar_size(structpci_dev*dev,constint*indexes)+{+structresource*res;+resource_size_tbase,size;+inti,r,num_res;++num_res=of_read_number(&indexes[NUM_RES_PROPERTY],1);+for(i=START_OF_ENTRIES,r=0;r<num_res&&r<PCI_SRIOV_NUM_BARS;+i+=NEXT_ENTRY,r++){+res=&dev->resource[r+PCI_IOV_RESOURCES];+base=of_read_number(&indexes[i],2);+size=of_read_number(&indexes[i+APERTURE_PROPERTY],2);+res->flags=pci_parse_of_flags(of_read_number+(&indexes[i+LOW_INT],1),0);+res->flags|=(IORESOURCE_MEM_64|IORESOURCE_PCI_FIXED);+res->name=pci_name(dev);+res->start=base;+res->end=base+size-1;+}+}++voidof_pci_parse_iov_addrs(structpci_dev*dev,constint*indexes)+{+structresource*res,*root,*conflict;+resource_size_tbase,size;+inti,r,num_res;++/*+*FirstelementinthearrayisthenumberofBars+*returned.Searchthroughthelisttofindthematching+*barsassignthemfromfirmwareintoresourcesstructure.+*/+num_res=of_read_number(&indexes[NUM_RES_PROPERTY],1);+for(i=START_OF_ENTRIES,r=0;r<num_res&&r<PCI_SRIOV_NUM_BARS;+i+=NEXT_ENTRY,r++){+res=&dev->resource[r+PCI_IOV_RESOURCES];+base=of_read_number(&indexes[i],2);+size=of_read_number(&indexes[i+WDW_SIZE_PROPERTY],2);+res->name=pci_name(dev);+res->start=base;+res->end=base+size-1;+root=pci_find_parent_resource(dev,res);++if(!root)+root=&iomem_resource;+dev_dbg(&dev->dev,+"Pseries IOV BAR %d: trying firmware assignment %pR\n",+r+PCI_IOV_RESOURCES,res);+conflict=request_resource_conflict(root,res);+if(conflict){+dev_info(&dev->dev,+"BAR %d: %pR conflicts with %s %pR\n",+r+PCI_IOV_RESOURCES,res,+conflict->name,conflict);+res->flags|=IORESOURCE_UNSET;+}+}+}++staticvoidpseries_pci_fixup_resources(structpci_dev*pdev)+{+constint*indexes;+structdevice_node*dn=pci_device_to_OF_node(pdev);++/*Firmware must support open sriov otherwise dont configure*/+indexes=of_get_property(dn,"ibm,open-sriov-vf-bar-info",NULL);+if(!indexes)+return;+/* Assign the addresses from device tree*/+of_pci_parse_vf_bar_size(pdev,indexes);+}++staticvoidpseries_pci_fixup_iov_resources(structpci_dev*pdev)+{+constint*indexes;+structdevice_node*dn=pci_device_to_OF_node(pdev);++if(!pdev->is_physfn||pdev->is_added)+return;+/*Firmware must support open sriov otherwise dont configure*/+indexes=of_get_property(dn,"ibm,open-sriov-vf-bar-info",NULL);+if(!indexes)+return;+/* Assign the addresses from device tree*/+of_pci_parse_iov_addrs(pdev,indexes);+}++staticresource_size_tpseries_pci_iov_resource_alignment(structpci_dev*pdev,+intresno)+{+const__be32*reg;+structdevice_node*dn=pci_device_to_OF_node(pdev);++/*Firmware must support open sriov otherwise report regular alignment*/+reg=of_get_property(dn,"ibm,is-open-sriov-pf",NULL);+if(!reg)+returnpci_iov_resource_size(pdev,resno);++if(!pdev->is_physfn)+return0;+returnpseries_get_iov_fw_values(pdev,+resno-PCI_IOV_RESOURCES,+APERTURE_SIZE);+}+#endif+staticvoid__initpSeries_setup_arch(void){set_arch_panic_timeout(10,ARCH_PANIC_TIMEOUT);
@@ -490,6 +665,14 @@ static void __init pSeries_setup_arch(void)vpa_init(boot_cpuid);ppc_md.power_save=pseries_lpar_idle;ppc_md.enable_pmcs=pseries_lpar_enable_pmcs;+#ifdef CONFIG_PCI_IOV+ppc_md.pcibios_fixup_resources=+pseries_pci_fixup_resources;+ppc_md.pcibios_fixup_sriov=+pseries_pci_fixup_iov_resources;+ppc_md.pcibios_iov_resource_alignment=+pseries_pci_iov_resource_alignment;+#endif}else{/* No special idle routine */ppc_md.enable_pmcs=power4_enable_pmcs;
From: Bryant G. Ly <hidden> Date: 2017-12-13 15:33:38
After initial validation of SR-IOV resources, firmware will
associate PEs to the dynamic VFs created within this call. This
patch adds the association of PEs to the PF array of PE numbers
indexed by VF.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/platforms/pseries/pci.c | 156 ++++++++++++++++++++++++++++++++++-
1 file changed, 153 insertions(+), 3 deletions(-)
@@ -57,18 +57,168 @@ void pcibios_name_device(struct pci_dev *dev)}DECLARE_PCI_FIXUP_HEADER(PCI_ANY_ID,PCI_ANY_ID,pcibios_name_device);#endif-#ifdef CONFIG_PCI_IOV+#define MAX_VFS_FOR_MAP_PE 256+structpe_map_bar_entry{+__be64bar;///< Input: Virtual Function BAR+__be16rid;///< Input: Virtual Function Router ID+__be16pe_num;///< Output: Virtual Function PE Number+__be32reserved;///< Reserved Space+};++intpseries_send_map_pe(structpci_dev*pdev,+u16num_vfs,+structpe_map_bar_entry*vf_pe_array)+{+structpci_dn*pdn;+intrc;+unsignedlongbuid,addr;+intibm_map_pes=rtas_token("ibm,open-sriov-map-pe-number");++if(ibm_map_pes==RTAS_UNKNOWN_SERVICE)+return-EINVAL;++pdn=pci_get_pdn(pdev);+addr=rtas_config_addr(pdn->busno,pdn->devfn,0);+buid=pdn->phb->buid;+spin_lock(&rtas_data_buf_lock);+memcpy(rtas_data_buf,vf_pe_array,+RTAS_DATA_BUF_SIZE);+rc=rtas_call(ibm_map_pes,5,1,NULL,addr,+BUID_HI(buid),BUID_LO(buid),+rtas_data_buf,+num_vfs*sizeof(structpe_map_bar_entry));+memcpy(vf_pe_array,rtas_data_buf,+RTAS_DATA_BUF_SIZE);+spin_unlock(&rtas_data_buf_lock);++if(rc)+dev_err(&pdev->dev,+"%s: Failed to associate pes PE#%lx, rc=%x\n",+__func__,addr,rc);++returnrc;+}++voidpseries_set_pe_num(structpci_dev*pdev,+u16vf_index,__be16pe_num)+{+structpci_dn*pdn;++pdn=pci_get_pdn(pdev);+pdn->pe_num_map[vf_index]=be16_to_cpu(pe_num);+dev_dbg(&pdev->dev,"VF %04x:%02x:%02x.%x associated with PE#%x\n",+pci_domain_nr(pdev->bus),+pdev->bus->number,+PCI_SLOT(pci_iov_virtfn_devfn(pdev,vf_index)),+PCI_FUNC(pci_iov_virtfn_devfn(pdev,vf_index)),+pdn->pe_num_map[vf_index]);+}++intpseries_associate_pes(structpci_dev*pdev,u16num_vfs)+{+structpci_dn*pdn;+inti,rc,vf_index;+structpe_map_bar_entry*vf_pe_array;+structresource*res;+u64size;++vf_pe_array=kzalloc(RTAS_DATA_BUF_SIZE,GFP_KERNEL);+if(!vf_pe_array)+return-ENOMEM;++memset(vf_pe_array,0,RTAS_DATA_BUF_SIZE);+pdn=pci_get_pdn(pdev);+/* create firmware structure to associate pes */+for(vf_index=0;vf_index<num_vfs&&vf_index<MAX_VFS_FOR_MAP_PE;+vf_index++){+pdn->pe_num_map[vf_index]=IODA_INVALID_PE;+for(i=0;i<PCI_SRIOV_NUM_BARS;i++){+res=&pdev->resource[i+PCI_IOV_RESOURCES];+if(!res->parent)+continue;+size=pcibios_iov_resource_alignment(pdev,i++PCI_IOV_RESOURCES+);+vf_pe_array[vf_index].bar=+be64_to_cpu(res->start+size*vf_index);+vf_pe_array[vf_index].rid=+be16_to_cpu((pci_iov_virtfn_bus(pdev,vf_index)+<<8)|pci_iov_virtfn_devfn(pdev,+vf_index));+vf_pe_array[vf_index].pe_num=+be16_to_cpu(IODA_INVALID_PE);+}+}++rc=pseries_send_map_pe(pdev,num_vfs,vf_pe_array);+/* Only zero is success */+if(!rc)+for(vf_index=0;vf_index<num_vfs&&vf_index<+MAX_VFS_FOR_MAP_PE;vf_index++)+pseries_set_pe_num(pdev,vf_index,+vf_pe_array[vf_index].pe_num);++kfree(vf_pe_array);+returnrc;+}++intpseries_pci_sriov_enable(structpci_dev*pdev,u16num_vfs)+{+structpci_dn*pdn;+intrc;+constint*max_vfs;+intmax_config_vfs;+structdevice_node*dn=pci_device_to_OF_node(pdev);++max_vfs=of_get_property(dn,"ibm,number-of-configurable-vfs",NULL);++if(!max_vfs)+return-EINVAL;++/* First integer stores max config */+max_config_vfs=of_read_number(&max_vfs[0],1);+if(max_config_vfs<num_vfs){+dev_err(&pdev->dev,+"Num VFs %x > %x Configurable VFs\n",+num_vfs,max_config_vfs);+return-EINVAL;+}++pdn=pci_get_pdn(pdev);+pdn->pe_num_map=kmalloc_array(num_vfs,+sizeof(*pdn->pe_num_map),+GFP_KERNEL);+if(!pdn->pe_num_map)+return-ENOMEM;++rc=pseries_associate_pes(pdev,num_vfs);++/* Anything other than zero is failure */+if(rc){+dev_err(&pdev->dev,"Failure to enable sriov: %x\n",rc);+kfree(pdn->pe_num_map);+}else{+pci_vf_drivers_autoprobe(pdev,false);+}++returnrc;+}+intpseries_pcibios_sriov_enable(structpci_dev*pdev,u16num_vfs){/* Allocate PCI data */add_dev_pci_data(pdev);-pci_vf_drivers_autoprobe(pdev,false);-return0;+returnpseries_pci_sriov_enable(pdev,num_vfs);}intpseries_pcibios_sriov_disable(structpci_dev*pdev){+structpci_dn*pdn;++pdn=pci_get_pdn(pdev);+/* Releasing pe_num_map */+kfree(pdn->pe_num_map);/* Release PCI data */remove_dev_pci_data(pdev);pci_vf_drivers_autoprobe(pdev,true);
From: Bryant G. Ly <hidden> Date: 2017-12-13 15:33:42
Devices can go offline when EEH is reported. This patch adds
a change to the kernel object and lets udev know of error.
When device resumes a change is also set reporting device as
online. Therefore, EEH events are better propagated to user
space for devices in powerpc arch.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/kernel/eeh_driver.c | 5 ++++-
1 file changed, 4 insertions(+), 1 deletion(-)
Add EEH platform operations for pseries to update VF
config space. With this change after EEH, the VF
will have updated config space for pseries platform.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/platforms/pseries/eeh_pseries.c | 85 +++++++++++++++++++++++++++-
1 file changed, 84 insertions(+), 1 deletion(-)
@@ -708,6 +708,89 @@ static int pseries_eeh_write_config(struct pci_dn *pdn, int where, int size, u32returnrtas_write_config(pdn,where,size,val);}+staticintpseries_eeh_restore_vf_config(structpci_dn*pdn)
This particular function is just a copy of its powernv counterpart -
pnv_eeh_restore_vf_config(), it could go to arch/powerpc/kernel/eeh.c, for
example. Or I am missing something here?
Devices can go offline when EEH is reported. This patch adds
a change to the kernel object and lets udev know of error.
When device resumes a change is also set reporting device as
online. Therefore, EEH events are better propagated to user
space for devices in powerpc arch.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/kernel/eeh_driver.c | 5 ++++-
1 file changed, 4 insertions(+), 1 deletion(-)
From: Russell Currey <hidden> Date: 2017-12-18 04:15:12
On Wed, 2017-12-13 at 09:32 -0600, Bryant G. Ly wrote:
Devices can go offline when EEH is reported. This patch adds
a change to the kernel object and lets udev know of error.
When device resumes a change is also set reporting device as
online. Therefore, EEH events are better propagated to user
space for devices in powerpc arch.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
It would probably also be useful to communicate when recovery fails and
a device is no longer usable, so userspace knows not to keep waiting
for recovery to complete.
From: Russell Currey <hidden> Date: 2017-12-18 04:29:13
On Wed, 2017-12-13 at 09:32 -0600, Bryant G. Ly wrote:
When pseries SR-IOV is enabled and after a PF driver
has resumed from EEH, platform has to be notified
of the event so the child VFs can be allowed to
resume their normal recovery path.
This patch makes the EEH operation allow unfreeze
platform dependent code and adds the call to
pseries EEH code.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
Just some nitpicks, there's a lot of weird whitespace in this patch
u32 val);
int (*next_error)(struct eeh_pe **pe);
int (*restore_config)(struct pci_dn *pdn);
+ int (*notify_resume)(struct pci_dn *pdn);
};
extern int eeh_subsystem_flags;
diff --git a/arch/powerpc/kernel/eeh_driver.c
b/arch/powerpc/kernel/eeh_driver.c
index c61bf770282b..dbda0cda559b 100644
From: Russell Currey <hidden> Date: 2017-12-18 04:31:42
On Wed, 2017-12-13 at 09:32 -0600, Bryant G. Ly wrote:
quoted hunk
To correctly use EEH code one has to make
sure that the EEH_PE_VF is set for dynamic created
VFs. Therefore this patch allocates an eeh_pe of
eeh type EEH_PE_VF and associates PE with parent.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/include/asm/pci-bridge.h | 5 ++++-
arch/powerpc/platforms/pseries/eeh_pseries.c | 9 ++++++++-
2 files changed, 12 insertions(+), 2 deletions(-)
@@ -211,7 +211,10 @@ struct pci_dn {unsignedint*pe_num_map;/* PE# for the first VF PE
or array */
bool m64_single_mode; /* Use M64 BAR in Single
Mode */
#define IODA_INVALID_M64 (-1)
- int (*m64_map)[PCI_SRIOV_NUM_BARS];
+ union {
+ int (*m64_map)[PCI_SRIOV_NUM_BARS];
+ int last_allow_rc;
+ };
A comment would be useful here. Why are these mutually exclusive,
last_allow_rc isn't amazingly self-documenting.
quoted hunk
#endif /* CONFIG_PCI_IOV */
int mps; /* Maximum Payload
Size */
struct list_head child_list;
To correctly use EEH code one has to make
sure that the EEH_PE_VF is set for dynamic created
VFs. Therefore this patch allocates an eeh_pe of
eeh type EEH_PE_VF and associates PE with parent.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/include/asm/pci-bridge.h | 5 ++++-
arch/powerpc/platforms/pseries/eeh_pseries.c | 9 ++++++++-
2 files changed, 12 insertions(+), 2 deletions(-)
@@ -211,7 +211,10 @@ struct pci_dn {unsignedint*pe_num_map;/* PE# for the first VF PE or array */boolm64_single_mode;/* Use M64 BAR in Single Mode */#define IODA_INVALID_M64 (-1)-int(*m64_map)[PCI_SRIOV_NUM_BARS];+union{+int(*m64_map)[PCI_SRIOV_NUM_BARS];+intlast_allow_rc;
I'd suggest defining it where is used, easier to follow what this actually
does.
quoted hunk
+ };
#endif /* CONFIG_PCI_IOV */
int mps; /* Maximum Payload Size */
struct list_head child_list;
When pseries SR-IOV is enabled and after a PF driver
has resumed from EEH, platform has to be notified
of the event so the child VFs can be allowed to
resume their normal recovery path.
This patch makes the EEH operation allow unfreeze
platform dependent code and adds the call to
pseries EEH code.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/include/asm/eeh.h | 1 +
arch/powerpc/kernel/eeh_driver.c | 4 ++
arch/powerpc/platforms/powernv/eeh-powernv.c | 3 +-
arch/powerpc/platforms/pseries/eeh_pseries.c | 100 ++++++++++++++++++++++++++-
4 files changed, 106 insertions(+), 2 deletions(-)
eeh_ops->notify_resume(eeh_dev_to_pdn(edev));
otherwise the compiler will complain at @pdn declaration if
!defined(CONFIG_PCI_IOV). Just try compiling without CONFIG_PCI_IOV.
+ cur_vfs = 0;
+ rc = 0;
+ if (edev->pdev->is_physfn) {
+ pe = eeh_dev_to_pe(edev);
+ cur_vfs = pci_num_vf(edev->pdev);
+ pdn = eeh_dev_to_pdn(edev);
+ parent = pdn->parent;
+ /* For each of its VF
+ * call allow unfreeze
+ */
+ for (vf_index = 0; vf_index < cur_vfs; vf_index++)
+ vf_pe_array[vf_index] =
+ be16_to_cpu(pdn->pe_num_map[vf_index]);
It is kind of assumed that the number of VFs is always less than 2048 minus
rtas call parameters (RTAS_DATA_BUF_SIZE==4096 now)?
Is the upper limit of cur_vfs checked anywhere?
We can afford 4K on stack if we know for sure this is all we n
I am missing the point of copying last_allow_rc - cannot
eeh_notify_resume_show() just return it from the PF?
May be just add another flag to eeh_pe::state for last_allow_rc? How many
different @rc do we expect here?
After initial validation of SR-IOV resources, firmware will
associate PEs to the dynamic VFs created within this call. This
patch adds the association of PEs to the PF array of PE numbers
indexed by VF.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/platforms/pseries/pci.c | 156 ++++++++++++++++++++++++++++++++++-
1 file changed, 153 insertions(+), 3 deletions(-)
As mentioned elsewhere, kzalloc() above resets memory.
+ pdn = pci_get_pdn(pdev);
+ /* create firmware structure to associate pes */
+ for (vf_index = 0; vf_index < num_vfs && vf_index < MAX_VFS_FOR_MAP_PE;
It would make the code and your life easier if you check for
num_vfs<=MAX_VFS_FOR_MAP_PE in pseries_pci_sriov_enable() and then you
won't have to check for MAX_VFS_FOR_MAP_PE. As for now, if
num_vfs>=MAX_VFS_FOR_MAP_PE, all VFs above MAX_VFS_FOR_MAP_PE will be ignored.
+ vf_index++) {
+ pdn->pe_num_map[vf_index] = IODA_INVALID_PE;
+ for (i = 0; i < PCI_SRIOV_NUM_BARS; i++) {
+ res = &pdev->resource[i + PCI_IOV_RESOURCES];
+ if (!res->parent)
+ continue;
+ size = pcibios_iov_resource_alignment(pdev, i +
+ PCI_IOV_RESOURCES
+ );
afaik the kernel coding style is tolerant to 2 tabs indents so it is not
necessary to align under the opening bracket and you can do:
When enabling SR-IOV in pseries platform,
the VF bar properties for a PF are reported on
the device node in the device tree.
This patch adds the IOV Bar resources to Linux
structures from the device tree for later use
when configuring SR-IOV by PF driver.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/include/asm/pci.h | 2 +
arch/powerpc/kernel/pci_of_scan.c | 2 +-
arch/powerpc/platforms/pseries/setup.c | 183 +++++++++++++++++++++++++++++++++
3 files changed, 186 insertions(+), 1 deletion(-)
@@ -459,6 +459,181 @@ static void __init find_and_init_phbs(void)of_pci_check_probe_only();}+#ifdef CONFIG_PCI_IOV+enumrtas_iov_fw_value_map{+NUM_RES_PROPERTY=0,///< Number of Resources+LOW_INT=1,///< Lowest 32 bits of Address+START_OF_ENTRIES=2,///< Always start of entry+APERTURE_PROPERTY=2,///< Start of entry+ to Aperture Size+WDW_SIZE_PROPERTY=4,///< Start of entry+ to Window Size+NEXT_ENTRY=7///< Go to next entry on array+};++enumget_iov_fw_value_index{+BAR_ADDRS=1,///< Get Bar Address+APERTURE_SIZE=2,///< Get Aperture Size+WDW_SIZE=3///< Get Window Size+};++resource_size_tpseries_get_iov_fw_values(structpci_dev*dev,intresno,
s/pseries_get_iov_fw_values/pseries_get_iov_fw_value/ as it returns a
single value.
This is more common way of doing the same initialization:
struct vf_bar_wdw {
__be64 addr;
__be64 aperture_size;
__be64 wdw_size;
} window_avail[PCI_SRIOV_NUM_BARS] = { 0 };
+ return_value = 0;
+ /*
+ * First element in the array is the number of Bars
+ * returned. Search through the list to find the matching
+ * bar
+ */
+ num_res = of_read_number(&indexes[NUM_RES_PROPERTY], 1);
if (resno >= num_res)
return 0; /* or an errror */
i = START_OF_ENTRIES + NEXT_ENTRY * resno;
switch (value) {
case BAR_ADDRS:
ret = f_read_number(&indexes[i], 2);
break;
case APERTURE_SIZE:
ret = of_read_number(&indexes[i + APERTURE_PROPERTY], 2);
break;
case WDW_SIZE:
ret = of_read_number(&indexes[i + WDW_SIZE_PROPERTY], 2);
break;
}
return ret;
}
and remove the reminder of the function, and window_avail, and vf_bar_wdw?
+ for (i = START_OF_ENTRIES, r = 0; r < num_res && r < PCI_SRIOV_NUM_BARS;
+ i += NEXT_ENTRY, r++) {
+ window_avail[r].addr = of_read_number(&indexes[i], 2);
+ window_avail[r].aperture_size =
+ of_read_number(&indexes[i + APERTURE_PROPERTY], 2);
+ window_avail[r].wdw_size =
+ of_read_number(&indexes[i + WDW_SIZE_PROPERTY], 2);
+ }
+
+ switch (value) {
+ case BAR_ADDRS:
+ return_value = window_avail[resno].addr;
+ break;
+ case APERTURE_SIZE:
+ return_value = window_avail[resno].aperture_size;
+ break;
+ case WDW_SIZE:
+ return_value = window_avail[resno].wdw_size;
+ break;
+ default:
+ break;
+ }
+ return return_value;
+}
+
+void of_pci_parse_vf_bar_size(struct pci_dev *dev, const int *indexes)
It does not seem to make a lot of sense as a separate function imho...
+{
+ struct resource *res;
+ resource_size_t base, size;
+ int i, r, num_res;
+
+ num_res = of_read_number(&indexes[NUM_RES_PROPERTY], 1);
num_res = min_t(int, num_res, PCI_SRIOV_NUM_BARS) ? Matter of personal
taste though.
+ for (i = START_OF_ENTRIES, r = 0; r < num_res && r < PCI_SRIOV_NUM_BARS;
+ i += NEXT_ENTRY, r++) {
+ res = &dev->resource[r + PCI_IOV_RESOURCES];
+ base = of_read_number(&indexes[i], 2);
+ size = of_read_number(&indexes[i + APERTURE_PROPERTY], 2);
+ res->flags = pci_parse_of_flags(of_read_number
+ (&indexes[i + LOW_INT], 1), 0);
+ res->flags |= (IORESOURCE_MEM_64 | IORESOURCE_PCI_FIXED);
+ res->name = pci_name(dev);
+ res->start = base;
+ res->end = base + size - 1;
The function name suggests it only parses sizes but just above it assigns
all resource parameters - size, address, flags.
+ }
+}
+
+void of_pci_parse_iov_addrs(struct pci_dev *dev, const int *indexes)
+{
+ struct resource *res, *root, *conflict;
+ resource_size_t base, size;
+ int i, r, num_res;
+
+ /*
+ * First element in the array is the number of Bars
+ * returned. Search through the list to find the matching
+ * bars assign them from firmware into resources structure.
+ */
+ num_res = of_read_number(&indexes[NUM_RES_PROPERTY], 1);
+ for (i = START_OF_ENTRIES, r = 0; r < num_res && r < PCI_SRIOV_NUM_BARS;
+ i += NEXT_ENTRY, r++) {
+ res = &dev->resource[r + PCI_IOV_RESOURCES];
+ base = of_read_number(&indexes[i], 2);
+ size = of_read_number(&indexes[i + WDW_SIZE_PROPERTY], 2);
+ res->name = pci_name(dev);
+ res->start = base;
+ res->end = base + size - 1;
+ root = pci_find_parent_resource(dev, res);
+
+ if (!root)
+ root = &iomem_resource;
@dev here is a VF, right? I am not familiar with powervn much but from what
I see - the devices are sitting on a root bus of their own PHB and they all
either have a root returned from pci_find_parent_resource() or none of them
has a root and will fall back to &iomem_resource, or both cases are possible?
From: Bryant G. Ly <hidden> Date: 2017-12-18 18:45:26
On 12/17/17 9:54 PM, Alexey Kardashevskiy wrote:
On 14/12/17 02:32, Bryant G. Ly wrote:
quoted
Devices can go offline when EEH is reported. This patch adds
a change to the kernel object and lets udev know of error.
When device resumes a change is also set reporting device as
online. Therefore, EEH events are better propagated to user
space for devices in powerpc arch.
Signed-off-by: Bryant G. Ly <redacted>
Signed-off-by: Juan J. Alvarez <redacted>
---
arch/powerpc/kernel/eeh_driver.c | 5 ++++-
1 file changed, 4 insertions(+), 1 deletion(-)
From: Juan Alvarez <hidden> Date: 2017-12-18 19:29:26
Here we need to set the config_addr as PHYP (platform)
does not enable the PE until the PE is bound to a VM,
reason why we disable VF autoprobe.
On 12/17/17 10:34 PM, Alexey Kardashevskiy wrote:
powernv does this from eeh_ops::probe, and so does pseries_eeh_probe(), do
you still need this here?
From: Juan Alvarez <hidden> Date: 2017-12-18 19:29:49
Yes, way less. So our current design only supports less than or equal to
256 VFs per PF. That is 256*2 bytes.
On 12/17/17 11:02 PM, Alexey Kardashevskiy wrote:
It is kind of assumed that the number of VFs is always less than 2048 minus
rtas call parameters (RTAS_DATA_BUF_SIZE==4096 now)?
From: Juan Alvarez <hidden> Date: 2017-12-18 19:30:07
This is PF only path. Yes either we have a root returned otherwise
will fall back to iomem_resource.
On 12/18/17 1:21 AM, Alexey Kardashevskiy wrote:
@dev here is a VF, right? I am not familiar with powervn much but from what
I see - the devices are sitting on a root bus of their own PHB and they all
either have a root returned from pci_find_parent_resource() or none of them
has a root and will fall back to &iomem_resource, or both cases are possible?
This is PF only path. Yes either we have a root returned otherwise
will fall back to iomem_resource.
You have removed context from my response, do not do that please.
When will you have root and when you won't? imho it should always be either
one or another.
On 12/18/17 1:21 AM, Alexey Kardashevskiy wrote:
quoted
@dev here is a VF, right? I am not familiar with powervn much but from what
I see - the devices are sitting on a root bus of their own PHB and they all
either have a root returned from pci_find_parent_resource() or none of them
has a root and will fall back to &iomem_resource, or both cases are possible?
From: Juan Alvarez <hidden> Date: 2017-12-21 03:04:50
On 12/19/17 12:38 AM, Alexey Kardashevskiy wrote:
On 19/12/17 06:29, Juan Alvarez wrote:
quoted
This is PF only path. Yes either we have a root returned otherwise
will fall back to iomem_resource.
You have removed context from my response, do not do that please.
My apologies. I will not do that.
When will you have root and when you won't? imho it should always be either
one or another.
Yes you are correct. The resource is carved out of a different mmio
space and will never be passed in the assigned-addresses property in
the device node of PF.
We will remove that function call, conditional check and set root accordingly.
quoted
On 12/18/17 1:21 AM, Alexey Kardashevskiy wrote:
quoted
@dev here is a VF, right? I am not familiar with powervn much but from what
I see - the devices are sitting on a root bus of their own PHB and they all
either have a root returned from pci_find_parent_resource() or none of them
has a root and will fall back to &iomem_resource, or both cases are possible?