This patchset enables EEH on SRIOV VFs. The general idea is to create proper
VF edev and VF PE and handle them properly.
Different from the Bus PE, VF PE just contain one VF. This introduces the
difference of EEH error handling on a VF PE. Generally, it has several
differences.
First, the VF's removal and re-enumerate rely on its PF. VF has a tight
relationship between its PF. This is not proper to enumerate a VF by usual
scan procedure. That's why virtfn_add/virtfn_remove are exported in this patch
set.
Second, the reset/restore of a VF is done in kernel space. FW is not aware of
the VF, this means the usual reset function done in FW will not work. One of
the patch will imitate the reset/restore function in kernel space.
Third, the VF may be removed during the PF's error_detected function. In this
case, the original error_detected->slot_reset->resume sequence is not proper
to those removed VFs, since they are re-created by PF in a fresh state. A flag
in eeh_dev is introduce to mark the eeh_dev is in error state. By doing so, we
track whether this device needs to be reset or not.
This has been tested both on host and in guest on Power8 with latest kernel
version.
v4:
* refine the change logs, comment and code style
* change pnv_pci_fixup_vf_eeh() to pnv_eeh_vf_final_fixup() and remove the
CONFIG_PCI_IOV macro
* reorder patch 5/6 to make the logic more reasonable
* remove remove_dev_pci_data()
* remove the EEH_DEV_VF flag, use edev->physfn to identify a VF EEH DEV and
remove related CONFIG_PCI_IOV macro
* add the option for VF reset
* fix the pnv_eeh_cfg_blocked() logic
* replace pnv_pci_cfg_{read,write} with eeh_ops->{read,write}_config in
pnv_eeh_vf_restore_config()
* rename pnv_eeh_vf_restore_config() to pnv_eeh_restore_vf_config()
* rename pnv_pci_fixup_vf_caps() to pnv_pci_vf_header_fixup() and move it
to arch/powerpc/platforms/powernv/pci.c
* add a field compound in pnv_ioda_pe to link compound PEs
* handle compound PE for VF PEs
v3:
* add back vf_index in pci_dn to track the VF's index
* rename ppdev in eeh_dev to physfn for consistency
* move edev->physfn assignment before dev->dev.archdata.edev is set
* move pnv_pci_fixup_vf_eeh() and pnv_pci_fixup_vf_caps() to eeh-powernv.c
* more clear and detail in commit log and comment in code
* merge eeh_rmv_virt_device() with eeh_rmv_device()
* move the cfg_blocked check logic from pnv_eeh_read/write_config() to
pnv_eeh_cfg_blocked()
* move the vf reset/restore logic into its own patch, two patches are
created.
powerpc/powernv: Support PCI config restore for VFs
powerpc/powernv: Support EEH reset for VFs
* simplify the vf reset logic
v2:
* add prefix pci_iov_ to virtfn_add/virtfn_remove
* use EEH_DEV_VF as a flag for a VF's eeh_dev
* use eeh_dev instead of edev in change log
* remove vf_index in eeh_dev, calculate it from pdn->busno and devfn
* do eeh_add_device_late() and eeh_sysfs_add_device() both after pci_dev is
well initialized
* do FLR to reset a VF PE
* imitate the restore function in FW for VF
* remove the reverse order patch, since it is still under discussion
Wei Yang (11):
pci/iov: rename and export virtfn_add/virtfn_remove
powerpc/pci_dn: cache vf_index in pci_dn
powerpc/pci: remove PCI devices in reverse order
powerpc/eeh: cache address range just for normal device
powerpc/powernv: create/release eeh_dev for VF
powerpc/eeh: create EEH_PE_VF for VF PE
powerpc/powernv: Support EEH reset for VFs
powerpc/powernv: Support PCI config restore for VFs
powerpc/eeh: handle VF PE properly
powerpc/powernv: use "compound" as the child's list_head for compound
PE
powerpc/powernv: compound PE for VFs
arch/powerpc/include/asm/eeh.h | 4 +
arch/powerpc/include/asm/pci-bridge.h | 2 +
arch/powerpc/kernel/eeh.c | 5 +
arch/powerpc/kernel/eeh_cache.c | 2 +-
arch/powerpc/kernel/eeh_driver.c | 103 +++++++++++---
arch/powerpc/kernel/eeh_pe.c | 13 +-
arch/powerpc/kernel/pci-hotplug.c | 2 +-
arch/powerpc/kernel/pci_dn.c | 15 +-
arch/powerpc/platforms/powernv/eeh-powernv.c | 196 +++++++++++++++++++++++++-
arch/powerpc/platforms/powernv/pci-ioda.c | 35 ++++-
arch/powerpc/platforms/powernv/pci.c | 16 +++
arch/powerpc/platforms/powernv/pci.h | 1 +
drivers/pci/iov.c | 10 +-
include/linux/pci.h | 2 +
14 files changed, 366 insertions(+), 40 deletions(-)
--
1.7.9.5
During the EEH recovery, when a device's driver is not EEH aware or no
driver is bound with a device, EEH core would do hotplug on this device.
While it isn't feasible for a VF with usual hotplug procedure. During
removal of a VF, virtual bus should be removed if necessary. During the
re-creation, the pci_scan_slot() doesn't work on a VF.
This patch exports two functions to handle the hotplug case for VF
properly. They will be invoked when the EEH core does the hotplug case for
VFs.
Signed-off-by: Wei Yang <redacted>
---
drivers/pci/iov.c | 10 +++++-----
include/linux/pci.h | 2 ++
2 files changed, 7 insertions(+), 5 deletions(-)
@@ -1679,6 +1679,8 @@ int pci_iov_virtfn_devfn(struct pci_dev *dev, int id);intpci_enable_sriov(structpci_dev*dev,intnr_virtfn);voidpci_disable_sriov(structpci_dev*dev);+intpci_iov_virtfn_add(structpci_dev*dev,intid,intreset);+voidpci_iov_virtfn_remove(structpci_dev*dev,intid,intreset);intpci_num_vf(structpci_dev*dev);intpci_vfs_assigned(structpci_dev*dev);intpci_sriov_set_totalvfs(structpci_dev*dev,u16numvfs);
The patch caches the VF index in pci_dn, which can be used to calculate
VF's bus, device and function number. Those information helps to locate
the VF's PCI device instance when doing hotplug during EEH recovery if
necessary.
Signed-off-by: Wei Yang <redacted>
---
arch/powerpc/include/asm/pci-bridge.h | 1 +
arch/powerpc/kernel/pci_dn.c | 4 +++-
2 files changed, 4 insertions(+), 1 deletion(-)
@@ -199,6 +199,7 @@ struct pci_dn {#ifdef CONFIG_PCI_IOVu16vfs_expanded;/* number of VFs IOV BAR expanded */u16num_vfs;/* number of VFs enabled*/+intvf_index;/* VF index in the PF */intoffset;/* PE# for the first VF PE */#define M64_PER_IOV 4intm64_per_iov;
As commit ac205b7b ("PCI: make sriov work with hotplug remove") indicates,
when removing PCI devices on a bus which has VFs, we need to remove them
in the reverse order.
This patch applies this pattern to the hotplug removal code for the powerpc
arch.
Signed-off-by: Wei Yang <redacted>
Acked-by: Gavin Shan <redacted>
---
arch/powerpc/kernel/pci-hotplug.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
The address cache is used to find the related eeh_dev for a given MMIO
address. From the definition of pci_dev.resource[], it keeps MMIO address
in following order: 6 normal BAR, ROM BAR, 6 IOV BAR, 4 Bridge window.
In the address cache, first it doesn't cache bridge device, second the IOV
BAR range should map to their own VFs separately. This means it just need
to cache the first 7 BARs for a normal device.
This patch restricts the address cache to save the first 7 BARs for a pci
device.
Signed-off-by: Wei Yang <redacted>
Acked-by: Gavin Shan <redacted>
---
arch/powerpc/kernel/eeh_cache.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
@@ -196,7 +196,7 @@ static void __eeh_addr_cache_insert_dev(struct pci_dev *dev)}/* Walk resources on this device, poke them into the tree */-for(i=0;i<DEVICE_COUNT_RESOURCE;i++){+for(i=0;i<=PCI_ROM_RESOURCE;i++){resource_size_tstart=pci_resource_start(dev,i);resource_size_tend=pci_resource_end(dev,i);unsignedlongflags=pci_resource_flags(dev,i);
EEH on powerpc platform needs eeh_dev structure to track the PCI device
status. Since VFs are created/released dynamically, VF's eeh_dev is also
dynamically created/released in system.
This patch creates/removes eeh_dev when pci_dn is created/removed for VFs,
and marks it with EEH_DEV_VF type.
Signed-off-by: Wei Yang <redacted>
---
arch/powerpc/include/asm/eeh.h | 1 +
arch/powerpc/kernel/eeh.c | 4 ++++
arch/powerpc/kernel/pci_dn.c | 11 +++++++++++
3 files changed, 16 insertions(+)
@@ -180,7 +180,9 @@ static struct pci_dn *add_one_dev_pci_data(struct pci_dn *parent,structpci_dn*add_dev_pci_data(structpci_dev*pdev){#ifdef CONFIG_PCI_IOV+structpci_controller*hose=pci_bus_to_host(pdev->bus);structpci_dn*parent,*pdn;+structeeh_dev*edev;inti;/* Only support IOV for now */
On powernv platform, VF PE is a special PE which is different from the Bus
PE. On the EEH side, it needs a corresponding concept to handle the VF PE
properly. For example, we need to create VF PE when VF's pci_dev is
initialized in kernel. And add a flag to mark it is a VF PE.
This patch introduces the EEH_PE_VF type for VF PE and creates it for a VF.
At the mean time, it creates the sysfs and address cache for VF PE at PCI
device final fixup time.
Signed-off-by: Wei Yang <redacted>
---
arch/powerpc/include/asm/eeh.h | 1 +
arch/powerpc/kernel/eeh_pe.c | 10 ++++++++--
arch/powerpc/platforms/powernv/eeh-powernv.c | 14 ++++++++++++++
3 files changed, 23 insertions(+), 2 deletions(-)
@@ -299,7 +299,10 @@ static struct eeh_pe *eeh_pe_get_parent(struct eeh_dev *edev)*EEHdevicealreadyhavingassociatedPE,but*thedirectparentEEHdevicedoesn'thaveyet.*/-pdn=pdn?pdn->parent:NULL;+if(edev->physfn)+pdn=pci_get_pdn(edev->physfn);+else+pdn=pdn?pdn->parent:NULL;while(pdn){/* We're poking out of PCI territory */parent=pdn_to_eeh_dev(pdn);
@@ -382,7 +385,10 @@ int eeh_add_to_parent_pe(struct eeh_dev *edev)}/* Create a new EEH PE */-pe=eeh_pe_alloc(edev->phb,EEH_PE_DEVICE);+if(edev->physfn)+pe=eeh_pe_alloc(edev->phb,EEH_PE_VF);+else+pe=eeh_pe_alloc(edev->phb,EEH_PE_DEVICE);if(!pe){pr_err("%s: out of memory!\n",__func__);return-ENOMEM;
Before VF PE is introduced, there isn't a method to reset an individual PCI
function. And since skiboot firmware is not aware of the VF, the VF's reset
should be done in kernel.
This patch introduces a function pnv_eeh_vf_pe_reset() to do the FLR or AF
FLR to a VF.
Signed-off-by: Wei Yang <redacted>
---
arch/powerpc/include/asm/eeh.h | 1 +
arch/powerpc/platforms/powernv/eeh-powernv.c | 123 +++++++++++++++++++++++++-
2 files changed, 123 insertions(+), 1 deletion(-)
@@ -134,6 +134,7 @@ struct eeh_dev {intpcix_cap;/* Saved PCIx capability */intpcie_cap;/* Saved PCIe capability */intaer_cap;/* Saved AER capability */+intaf_cap;/* Saved AF capability */structeeh_pe*pe;/* Associated PE */structlist_headlist;/* Form link list in the PE */structpci_controller*phb;/* Associated PHB */
@@ -891,6 +892,117 @@ static int pnv_eeh_bridge_reset(struct pci_dev *dev, int option)return0;}+staticintpnv_pci_wait_for_pending(structpci_dn*pdn,intpos,u16mask)+{+inti;++/* Wait for Transaction Pending bit clean */+for(i=0;i<4;i++){+u32status;+if(i)+msleep((1<<(i-1))*100);++eeh_ops->read_config(pdn,pos,2,&status);+if(!(status&mask))+return1;+}++return0;+}++staticintpnv_eeh_do_flr(structpci_dn*pdn,intoption)+{+u32cap;+structeeh_dev*edev=pdn_to_eeh_dev(pdn);++eeh_ops->read_config(pdn,edev->pcie_cap+PCI_EXP_DEVCAP,4,&cap);+if(!(cap&PCI_EXP_DEVCAP_FLR))+return-ENOTTY;++if(!pnv_pci_wait_for_pending(pdn,edev->pcie_cap+PCI_EXP_DEVSTA,+PCI_EXP_DEVSTA_TRPND))+pr_err("%04x:%02x:%02x:%01x timed out waiting for pending "+"transaction; performing function level reset anyway\n",+edev->phb->global_number,pdn->busno,+PCI_SLOT(pdn->devfn),PCI_FUNC(pdn->devfn));++eeh_ops->read_config(pdn,edev->pcie_cap+PCI_EXP_DEVCTL,4,&cap);+if(option==EEH_RESET_DEACTIVATE)+cap&=~PCI_EXP_DEVCTL_BCR_FLR;+else+cap|=PCI_EXP_DEVCTL_BCR_FLR;+eeh_ops->write_config(pdn,edev->pcie_cap+PCI_EXP_DEVCTL,4,cap);+msleep(100);+return0;+}++staticintpnv_eeh_do_af_flr(structpci_dn*pdn,intoption)+{+u32cap;+structeeh_dev*edev=pdn_to_eeh_dev(pdn);++if(!edev->af_cap)+return-ENOTTY;++eeh_ops->read_config(pdn,edev->af_cap+PCI_AF_CAP,1,&cap);+if(!(cap&PCI_AF_CAP_TP)||!(cap&PCI_AF_CAP_FLR))+return-ENOTTY;++/*+*WaitforTransactionPendingbittoclear.Aword-alignedtest+*isused,soweusetheconroloffsetratherthanstatusandshift+*thetestbittomatch.+*/+if(!pnv_pci_wait_for_pending(pdn,edev->af_cap+PCI_AF_CTRL,+PCI_AF_STATUS_TP<<8))+pr_err("%04x:%02x:%02x:%01x timed out waiting for pending "+"transaction; performing AF function level reset anyway\n",+edev->phb->global_number,pdn->busno,+PCI_SLOT(pdn->devfn),PCI_FUNC(pdn->devfn));++if(option==EEH_RESET_DEACTIVATE)+eeh_ops->write_config(pdn,edev->af_cap+PCI_AF_CTRL,1,0);+else+eeh_ops->write_config(pdn,edev->af_cap+PCI_AF_CTRL,1,+PCI_AF_CTRL_FLR);+msleep(100);+return0;+}++staticintpnv_eeh_reset_vf(structpci_dn*pdn,intoption)+{+intrc;++might_sleep();++rc=pnv_eeh_do_flr(pdn,option);+if(rc!=-ENOTTY)+gotodone;++rc=pnv_eeh_do_af_flr(pdn,option);+if(rc!=-ENOTTY)+gotodone;++done:+returnrc;+}++staticintpnv_eeh_vf_pe_reset(structeeh_pe*pe,intoption)+{+structeeh_dev*edev,*tmp;+structpci_dn*pdn;+intret=0;++eeh_pe_for_each_dev(pe,edev,tmp){+pdn=eeh_dev_to_pdn(edev);+ret|=pnv_eeh_reset_vf(pdn,option);+if(ret)+returnret;+}++returnret;+}+voidpnv_pci_reset_secondary_bus(structpci_dev*dev){structpci_controller*hose;
@@ -966,7 +1078,9 @@ static int pnv_eeh_reset(struct eeh_pe *pe, int option)}bus=eeh_pe_bus_get(pe);-if(pci_is_root_bus(bus)||+if(pe->type&EEH_PE_VF)+ret=pnv_eeh_vf_pe_reset(pe,option);+elseif(pci_is_root_bus(bus)||pci_is_root_bus(bus->parent))ret=pnv_eeh_root_reset(hose,option);else
Since skiboot firmware is not aware of VFs, the restore action for VF
should be done in kernel.
The patch introduces function pnv_eeh_restore_vf_config() to restore PCI
config space for VFs after reset.
Signed-off-by: Wei Yang <redacted>
---
arch/powerpc/include/asm/pci-bridge.h | 1 +
arch/powerpc/platforms/powernv/eeh-powernv.c | 59 +++++++++++++++++++++++++-
arch/powerpc/platforms/powernv/pci.c | 16 +++++++
3 files changed, 75 insertions(+), 1 deletion(-)
@@ -1601,6 +1601,59 @@ static int pnv_eeh_next_error(struct eeh_pe **pe)returnret;}+#ifdef CONFIG_PCI_IOV+staticintpnv_eeh_restore_vf_config(structpci_dn*pdn)+{+intpcie_cap,aer_cap,old_mps;+u32devctl,cmd,cap2,aer_capctl;++/* Restore MPS */+pcie_cap=pnv_eeh_find_cap(pdn,PCI_CAP_ID_EXP);+if(pcie_cap){+old_mps=(ffs(pdn->mps)-8)<<5;+eeh_ops->read_config(pdn,pcie_cap+PCI_EXP_DEVCTL,2,&devctl);+devctl&=~PCI_EXP_DEVCTL_PAYLOAD;+devctl|=old_mps;+eeh_ops->write_config(pdn,pcie_cap+PCI_EXP_DEVCTL,2,devctl);+}++/* Disable Completion Timeout */+if(pcie_cap){+eeh_ops->read_config(pdn,pcie_cap+PCI_EXP_DEVCAP2,4,&cap2);+if(cap2&0x10){+eeh_ops->read_config(pdn,pcie_cap+PCI_EXP_DEVCTL2,4,&cap2);+cap2|=0x10;+eeh_ops->write_config(pdn,pcie_cap+PCI_EXP_DEVCTL2,4,cap2);+}+}++/* Enable SERR and parity checking */+eeh_ops->read_config(pdn,PCI_COMMAND,2,&cmd);+cmd|=(PCI_COMMAND_PARITY|PCI_COMMAND_SERR);+eeh_ops->write_config(pdn,PCI_COMMAND,2,cmd);++/* Enable report various errors */+if(pcie_cap){+eeh_ops->read_config(pdn,pcie_cap+PCI_EXP_DEVCTL,2,&devctl);+devctl&=~PCI_EXP_DEVCTL_CERE;+devctl|=(PCI_EXP_DEVCTL_NFERE|+PCI_EXP_DEVCTL_FERE|+PCI_EXP_DEVCTL_URRE);+eeh_ops->write_config(pdn,pcie_cap+PCI_EXP_DEVCTL,2,devctl);+}++/* Enable ECRC generation and check */+if(pcie_cap){+aer_cap=pnv_eeh_find_ecap(pdn,PCI_EXT_CAP_ID_ERR);+eeh_ops->read_config(pdn,aer_cap+PCI_ERR_CAP,4,&aer_capctl);+aer_capctl|=(PCI_ERR_CAP_ECRC_GENE|PCI_ERR_CAP_ECRC_CHKE);+eeh_ops->write_config(pdn,aer_cap+PCI_ERR_CAP,4,aer_capctl);+}++return0;+}+#endif /* CONFIG_PCI_IOV */+staticintpnv_eeh_restore_config(structpci_dn*pdn){structeeh_dev*edev=pdn_to_eeh_dev(pdn);
@@ -1611,7 +1664,11 @@ static int pnv_eeh_restore_config(struct pci_dn *pdn)return-EEXIST;phb=edev->phb->private_data;-ret=opal_pci_reinit(phb->opal_id,+/* FW is not VF aware, we rely on OS to restore it */+if(edev->physfn)+ret=pnv_eeh_restore_vf_config(pdn);+else+ret=opal_pci_reinit(phb->opal_id,OPAL_REINIT_PCI_DEV,edev->config_addr);if(ret){pr_warn("%s: Can't reinit PCI dev 0x%x (%lld)\n",
Compared with Bus PE, VF PE just has one single pci function. This
introduces the difference of error handling on a VF PE.
For example in the hotplug case, EEH needs to remove and re-create the VF
properly. In the case when PF's error_detected() disable SRIOV, this patch
introduces a flag to mark the eeh_dev of a VF to avoid the slot_reset() and
resume(). Since the FW is not ware of the VF, this patch handles the VF
restore/reset in kernel directly.
This patch is to handle the VF PE properly in these cases.
Signed-off-by: Wei Yang <redacted>
---
arch/powerpc/include/asm/eeh.h | 1 +
arch/powerpc/kernel/eeh.c | 1 +
arch/powerpc/kernel/eeh_driver.c | 103 ++++++++++++++++++++++++++++++--------
arch/powerpc/kernel/eeh_pe.c | 3 +-
4 files changed, 85 insertions(+), 23 deletions(-)
@@ -548,6 +587,7 @@ static int eeh_reset_device(struct eeh_pe *pe, struct pci_bus *bus)structpci_bus*frozen_bus=eeh_pe_bus_get(pe);structtimevaltstamp;intcnt,rc,removed=0;+structeeh_dev*edev;/* pcibios will clear the counter; save the value */cnt=pe->freeze_count;
Commit 262af557dd75(powerpc/powernv: Enable M64 aperatus for PHB3)
introduces the concept of compound PE, and they are linked together to
master PE's slaves lish_head with the list field. While this field is
usually used to linked to the phb->ioda.pe_list to represents the PE is
used.
This patch introduces a field "compound" to link those compound PEs.
Signed-off-by: Wei Yang <redacted>
---
arch/powerpc/platforms/powernv/pci-ioda.c | 8 ++++----
arch/powerpc/platforms/powernv/pci.h | 1 +
2 files changed, 5 insertions(+), 4 deletions(-)
@@ -464,7 +464,7 @@ static int pnv_ioda_unfreeze_pe(struct pnv_phb *phb, int pe_no, int opt)return0;/* Clear frozen state for slave PEs */-list_for_each_entry(slave,&pe->slaves,list){+list_for_each_entry(slave,&pe->slaves,compound){rc=opal_pci_eeh_freeze_clear(phb->opal_id,slave->pe_number,opt);
@@ -516,7 +516,7 @@ static int pnv_ioda_get_pe_state(struct pnv_phb *phb, int pe_no)if(!(pe->flags&PNV_IODA_PE_MASTER))returnstate;-list_for_each_entry(slave,&pe->slaves,list){+list_for_each_entry(slave,&pe->slaves,compound){rc=opal_pci_eeh_freeze_status(phb->opal_id,slave->pe_number,&fstate,
@@ -73,6 +73,7 @@ struct pnv_ioda_pe {/* PEs in compound case */structpnv_ioda_pe*master;structlist_headslaves;+structlist_headcompound;/* Link in list of PE#s */structlist_headdma_link;
When VF BAR size is larger than 64MB, we group VFs in terms of M64 BAR,
which means those VFs in a group should form a compound PE.
This patch links those VF PEs into compound PE in this case.
Signed-off-by: Wei Yang <redacted>
---
arch/powerpc/platforms/powernv/pci-ioda.c | 27 ++++++++++++++++++++++++++-
1 file changed, 26 insertions(+), 1 deletion(-)
@@ -1360,6 +1361,11 @@ static void pnv_ioda_release_vf_PE(struct pci_dev *pdev, u16 num_vfs)__func__,pdn->offset+vf_index1,rc);}++/* Remove a Slave PE from Master PE */+pe=&phb->ioda.pe_array[pdn->offset+vf_index];+if(pe->flags&PNV_IODA_PE_SLAVE)+list_del(&pe->compound);}list_for_each_entry_safe(pe,pe_n,&phb->ioda.pe_list,list){
On Fri, May 15, 2015 at 01:46:16PM +0800, Wei Yang wrote:
During the EEH recovery, when a device's driver is not EEH aware or no
driver is bound with a device, EEH core would do hotplug on this device.
While it isn't feasible for a VF with usual hotplug procedure. During
removal of a VF, virtual bus should be removed if necessary. During the
re-creation, the pci_scan_slot() doesn't work on a VF.
This patch exports two functions to handle the hotplug case for VF
properly. They will be invoked when the EEH core does the hotplug case for
VFs.
Signed-off-by: Wei Yang <redacted>
@@ -106,7 +106,7 @@ resource_size_t pci_iov_resource_size(struct pci_dev *dev, int resno)
return dev->sriov->barsz[resno - PCI_IOV_RESOURCES];
}
-static int virtfn_add(struct pci_dev *dev, int id, int reset)
+int pci_iov_virtfn_add(struct pci_dev *dev, int id, int reset)
{
int i;
int rc = -ENOMEM;
@@ -181,7 +181,7 @@ failed:
return rc;
}
-static void virtfn_remove(struct pci_dev *dev, int id, int reset)
+void pci_iov_virtfn_remove(struct pci_dev *dev, int id, int reset)
{
struct pci_dev *virtfn;
@@ -302,7 +302,7 @@ static int sriov_enable(struct pci_dev *dev, int nr_virtfn)
}
for (i = 0; i < initial; i++) {
- rc = virtfn_add(dev, i, 0);
+ rc = pci_iov_virtfn_add(dev, i, 0);
if (rc)
goto failed;
}
@@ -314,7 +314,7 @@ static int sriov_enable(struct pci_dev *dev, int nr_virtfn)
@@ -1679,6 +1679,8 @@ int pci_iov_virtfn_devfn(struct pci_dev *dev, int id);
int pci_enable_sriov(struct pci_dev *dev, int nr_virtfn);
void pci_disable_sriov(struct pci_dev *dev);
+int pci_iov_virtfn_add(struct pci_dev *dev, int id, int reset);
+void pci_iov_virtfn_remove(struct pci_dev *dev, int id, int reset);
int pci_num_vf(struct pci_dev *dev);
int pci_vfs_assigned(struct pci_dev *dev);
int pci_sriov_set_totalvfs(struct pci_dev *dev, u16 numvfs);
--
1.7.9.5
On Fri, May 15, 2015 at 01:46:17PM +0800, Wei Yang wrote:
The patch caches the VF index in pci_dn, which can be used to calculate
VF's bus, device and function number. Those information helps to locate
the VF's PCI device instance when doing hotplug during EEH recovery if
necessary.
Signed-off-by: Wei Yang <redacted>
#ifdef CONFIG_PCI_IOV
u16 vfs_expanded; /* number of VFs IOV BAR expanded */
u16 num_vfs; /* number of VFs enabled*/
+ int vf_index; /* VF index in the PF */
int offset; /* PE# for the first VF PE */
#define M64_PER_IOV 4
int m64_per_iov;
On Fri, May 15, 2015 at 01:46:20PM +0800, Wei Yang wrote:
EEH on powerpc platform needs eeh_dev structure to track the PCI device
status. Since VFs are created/released dynamically, VF's eeh_dev is also
dynamically created/released in system.
This patch creates/removes eeh_dev when pci_dn is created/removed for VFs,
and marks it with EEH_DEV_VF type.
Signed-off-by: Wei Yang <redacted>
Acked-by: Gavin Shan <redacted>
After removing the unnecessary line of code as below.
Nothing is done to edev after getting it. So I think the last line of changes
here isn't needed. Could you check and remove it if I'm correct?
Thanks,
Gavin
On Fri, May 15, 2015 at 01:46:21PM +0800, Wei Yang wrote:
On powernv platform, VF PE is a special PE which is different from the Bus
PE. On the EEH side, it needs a corresponding concept to handle the VF PE
properly. For example, we need to create VF PE when VF's pci_dev is
initialized in kernel. And add a flag to mark it is a VF PE.
This patch introduces the EEH_PE_VF type for VF PE and creates it for a VF.
At the mean time, it creates the sysfs and address cache for VF PE at PCI
device final fixup time.
Signed-off-by: Wei Yang <redacted>
Acked-by: Gavin Shan <redacted>
With one thing fixed as below.
* EEH device already having associated PE, but
* the direct parent EEH device doesn't have yet.
*/
- pdn = pdn ? pdn->parent : NULL;
+ if (edev->physfn)
+ pdn = pci_get_pdn(edev->physfn);
+ else
+ pdn = pdn ? pdn->parent : NULL;
while (pdn) {
/* We're poking out of PCI territory */
parent = pdn_to_eeh_dev(pdn);
@@ -382,7 +385,10 @@ int eeh_add_to_parent_pe(struct eeh_dev *edev)
}
/* Create a new EEH PE */
- pe = eeh_pe_alloc(edev->phb, EEH_PE_DEVICE);
+ if (edev->physfn)
+ pe = eeh_pe_alloc(edev->phb, EEH_PE_VF);
+ else
+ pe = eeh_pe_alloc(edev->phb, EEH_PE_DEVICE);
if (!pe) {
pr_err("%s: out of memory!\n", __func__);
return -ENOMEM;
@@ -1540,3 +1540,17 @@ static int __init eeh_powernv_init(void)
return ret;
}
machine_early_initcall(powernv, eeh_powernv_init);
+
+static void pnv_eeh_vf_final_fixup(struct pci_dev *pdev)
+{
+ /*
+ * The following operations will fail if VF's sysfs files aren't
+ * created or its resources aren't finalized.
+ */
+ if (!pdev->is_virtfn)
+ return;
+
+ eeh_add_device_late(pdev);
+ eeh_sysfs_add_device(pdev);
It's worthy to have following code to make the logic here complete. Otherwise,
we will run into problem quickly once the eeh_add_device_{early,late}() get
changed in eeh.c:
eeh_add_device_early(pdn);
eeh_add_device_late(pdev);
eeh_sysfs_add_device(pdev);
Thanks,
Gavin
On Fri, May 15, 2015 at 01:46:22PM +0800, Wei Yang wrote:
quoted hunk
Before VF PE is introduced, there isn't a method to reset an individual PCI
function. And since skiboot firmware is not aware of the VF, the VF's reset
should be done in kernel.
This patch introduces a function pnv_eeh_vf_pe_reset() to do the FLR or AF
FLR to a VF.
Signed-off-by: Wei Yang <redacted>
---
arch/powerpc/include/asm/eeh.h | 1 +
arch/powerpc/platforms/powernv/eeh-powernv.c | 123 +++++++++++++++++++++++++-
2 files changed, 123 insertions(+), 1 deletion(-)
int pcix_cap; /* Saved PCIx capability */
int pcie_cap; /* Saved PCIe capability */
int aer_cap; /* Saved AER capability */
+ int af_cap; /* Saved AF capability */
struct eeh_pe *pe; /* Associated PE */
struct list_head list; /* Form link list in the PE */
struct pci_controller *phb; /* Associated PHB */
@@ -891,6 +892,117 @@ static int pnv_eeh_bridge_reset(struct pci_dev *dev, int option)
return 0;
}
+static int pnv_pci_wait_for_pending(struct pci_dn *pdn, int pos, u16 mask)
Could you change this function to something as below?
static bool pnv_eeh_wait_for_pending(struct pci_dn *pdn, int pos, u16 mask)
+{
+ int i;
u32 status;
int i;
You don't need the following "u32 status".
+
+ /* Wait for Transaction Pending bit clean */
+ for (i = 0; i < 4; i++) {
+ u32 status;
+ if (i)
+ msleep((1 << (i - 1)) * 100);
+
+ eeh_ops->read_config(pdn, pos, 2, &status);
+ if (!(status & mask))
+ return 1;
+ }
+
+ return 0;
+}
+
+static int pnv_eeh_do_flr(struct pci_dn *pdn, int option)
+{
+ u32 cap;
+ struct eeh_dev *edev = pdn_to_eeh_dev(pdn);
+
It's worthy to check if the device has PCIE cap though this function is
used by VFs who always have PCIE cap. However, it's still used for one
without PCIE cap.
+ eeh_ops->read_config(pdn, edev->pcie_cap + PCI_EXP_DEVCAP, 4, &cap);
+ if (!(cap & PCI_EXP_DEVCAP_FLR))
+ return -ENOTTY;
+
+ if (!pnv_pci_wait_for_pending(pdn, edev->pcie_cap + PCI_EXP_DEVSTA,
+ PCI_EXP_DEVSTA_TRPND))
+ pr_err("%04x:%02x:%02x:%01x timed out waiting for pending "
+ "transaction; performing function level reset anyway\n",
+ edev->phb->global_number, pdn->busno,
+ PCI_SLOT(pdn->devfn), PCI_FUNC(pdn->devfn));
Please print the function name and simplify the log into following one.
Also, the connection symbol between device and function number would
be ".", not ":".
pr_warn("%s: Pending transaction while issuing FLR to "
"%04x:%02x:%02x.%01x",
__func__, .....);
The hold and stablization delay has been standarized in EEH as below:
EEH_PE_RST_HOLD_TIME - After asserting reset
EEH_PE_RST_SETTLE_TIME - After deasserting reset
+ return 0;
+}
+
+static int pnv_eeh_do_af_flr(struct pci_dn *pdn, int option)
+{
+ u32 cap;
+ struct eeh_dev *edev = pdn_to_eeh_dev(pdn);
+
+ if (!edev->af_cap)
+ return -ENOTTY;
+
+ eeh_ops->read_config(pdn, edev->af_cap + PCI_AF_CAP, 1, &cap);
+ if (!(cap & PCI_AF_CAP_TP) || !(cap & PCI_AF_CAP_FLR))
+ return -ENOTTY;
+
+ /*
+ * Wait for Transaction Pending bit to clear. A word-aligned test
+ * is used, so we use the conrol offset rather than status and shift
+ * the test bit to match.
+ */
+ if (!pnv_pci_wait_for_pending(pdn, edev->af_cap + PCI_AF_CTRL,
+ PCI_AF_STATUS_TP << 8))
+ pr_err("%04x:%02x:%02x:%01x timed out waiting for pending "
+ "transaction; performing AF function level reset anyway\n",
+ edev->phb->global_number, pdn->busno,
+ PCI_SLOT(pdn->devfn), PCI_FUNC(pdn->devfn));
You can avoid using unnecessary tag:
rc = pnv_eeh_do_flr();
if (!rc)
return rc;
rc = pnv_eeh_do_af_flr();
return rc;
quoted hunk
+
+static int pnv_eeh_vf_pe_reset(struct eeh_pe *pe, int option)
+{
+ struct eeh_dev *edev, *tmp;
+ struct pci_dn *pdn;
+ int ret = 0;
+
+ eeh_pe_for_each_dev(pe, edev, tmp) {
+ pdn = eeh_dev_to_pdn(edev);
+ ret |= pnv_eeh_reset_vf(pdn, option);
+ if (ret)
+ return ret;
+ }
+
+ return ret;
+}
+
void pnv_pci_reset_secondary_bus(struct pci_dev *dev)
{
struct pci_controller *hose;
@@ -966,7 +1078,9 @@ static int pnv_eeh_reset(struct eeh_pe *pe, int option)
}
bus = eeh_pe_bus_get(pe);
- if (pci_is_root_bus(bus) ||
+ if (pe->type & EEH_PE_VF)
+ ret = pnv_eeh_vf_pe_reset(pe, option);
+ else if (pci_is_root_bus(bus) ||
pci_is_root_bus(bus->parent))
ret = pnv_eeh_root_reset(hose, option);
else
if (!edev || !edev->pe)
return false;
+ /*
+ * For VF's reset operation, we need to rely on the kernel to
+ * do those PCI config operations since firmware isn't aware of VFs.
+ */
+ if ((edev->physfn) && (edev->pe->state & EEH_PE_RESET))
+ return false;
+
if (edev->pe->state & EEH_PE_CFG_BLOCKED)
return true;
On Fri, May 15, 2015 at 01:46:23PM +0800, Wei Yang wrote:
quoted hunk
Since skiboot firmware is not aware of VFs, the restore action for VF
should be done in kernel.
The patch introduces function pnv_eeh_restore_vf_config() to restore PCI
config space for VFs after reset.
Signed-off-by: Wei Yang <redacted>
---
arch/powerpc/include/asm/pci-bridge.h | 1 +
arch/powerpc/platforms/powernv/eeh-powernv.c | 59 +++++++++++++++++++++++++-
arch/powerpc/platforms/powernv/pci.c | 16 +++++++
3 files changed, 75 insertions(+), 1 deletion(-)
@@ -1611,7 +1664,11 @@ static int pnv_eeh_restore_config(struct pci_dn *pdn)
return -EEXIST;
phb = edev->phb->private_data;
- ret = opal_pci_reinit(phb->opal_id,
+ /* FW is not VF aware, we rely on OS to restore it */
Please change the comment to:
/*
* We have to restore the PCI config space after reset since
* the firmware can't see SRIOV VFs.
*/
quoted hunk
+ if (edev->physfn)
+ ret = pnv_eeh_restore_vf_config(pdn);
+ else
+ ret = opal_pci_reinit(phb->opal_id,
OPAL_REINIT_PCI_DEV, edev->config_addr);
if (ret) {
pr_warn("%s: Can't reinit PCI dev 0x%x (%lld)\n",
On Fri, May 15, 2015 at 01:46:24PM +0800, Wei Yang wrote:
Compared with Bus PE, VF PE just has one single pci function. This
introduces the difference of error handling on a VF PE.
For example in the hotplug case, EEH needs to remove and re-create the VF
properly. In the case when PF's error_detected() disable SRIOV, this patch
introduces a flag to mark the eeh_dev of a VF to avoid the slot_reset() and
resume(). Since the FW is not ware of the VF, this patch handles the VF
restore/reset in kernel directly.
This patch is to handle the VF PE properly in these cases.
Signed-off-by: Wei Yang <redacted>
Acked-by: Gavin Shan <redacted>
With following things fixed:
* from the parent PE during the BAR resotre.
*/
edev->pdev = NULL;
+ edev->in_error = 0;
Could you please put detailed comments aboug the the usage of "in_error" here?
I may look into it later to remove it. For now, you don't need to do that since
we're almost run out of time.
quoted hunk
dev->dev.archdata.edev = NULL;
if (!(edev->pe->state & EEH_PE_KEEP))
eeh_rmv_from_parent_pe(edev);
On Fri, May 15, 2015 at 01:46:25PM +0800, Wei Yang wrote:
Commit 262af557dd75(powerpc/powernv: Enable M64 aperatus for PHB3)
introduces the concept of compound PE, and they are linked together to
master PE's slaves lish_head with the list field. While this field is
usually used to linked to the phb->ioda.pe_list to represents the PE is
used.
This patch introduces a field "compound" to link those compound PEs.
I don't think we needn't it with:
- VF PEs are classified to master and slave PEs.
- Master PEs is linked to phb->list;
- Slave PEs is linked to master->slaves;
- When iterating all PEs, you have to check PE's flag to cover all
(master & slave) VF PEs:
for_each_pe_in_phb_list {
/* Things we're doing */
if (pe_is_not_vf_pe ||
pe_is_not_master_pe)
continue;
for_each_pe_in_master_vf_pe_list {
/* slave VF PEs */
}
}
Thanks,
Gavin
/* PEs in compound case */
struct pnv_ioda_pe *master;
struct list_head slaves;
+ struct list_head compound;
/* Link in list of PE#s */
struct list_head dma_link;
--
1.7.9.5
On Fri, May 15, 2015 at 04:19:16PM +1000, Gavin Shan wrote:
On Fri, May 15, 2015 at 01:46:20PM +0800, Wei Yang wrote:
quoted
EEH on powerpc platform needs eeh_dev structure to track the PCI device
status. Since VFs are created/released dynamically, VF's eeh_dev is also
dynamically created/released in system.
This patch creates/removes eeh_dev when pci_dn is created/removed for VFs,
and marks it with EEH_DEV_VF type.
Signed-off-by: Wei Yang <redacted>
Acked-by: Gavin Shan <redacted>
After removing the unnecessary line of code as below.
On Fri, May 15, 2015 at 05:27:52PM +1000, Gavin Shan wrote:
On Fri, May 15, 2015 at 01:46:23PM +0800, Wei Yang wrote:
quoted
Since skiboot firmware is not aware of VFs, the restore action for VF
should be done in kernel.
The patch introduces function pnv_eeh_restore_vf_config() to restore PCI
config space for VFs after reset.
Signed-off-by: Wei Yang <redacted>
---
arch/powerpc/include/asm/pci-bridge.h | 1 +
arch/powerpc/platforms/powernv/eeh-powernv.c | 59 +++++++++++++++++++++++++-
arch/powerpc/platforms/powernv/pci.c | 16 +++++++
3 files changed, 75 insertions(+), 1 deletion(-)
@@ -1611,7 +1664,11 @@ static int pnv_eeh_restore_config(struct pci_dn *pdn)
return -EEXIST;
phb = edev->phb->private_data;
- ret = opal_pci_reinit(phb->opal_id,
+ /* FW is not VF aware, we rely on OS to restore it */
Please change the comment to:
/*
* We have to restore the PCI config space after reset since
* the firmware can't see SRIOV VFs.
*/
quoted
+ if (edev->physfn)
+ ret = pnv_eeh_restore_vf_config(pdn);
+ else
+ ret = opal_pci_reinit(phb->opal_id,
OPAL_REINIT_PCI_DEV, edev->config_addr);
if (ret) {
pr_warn("%s: Can't reinit PCI dev 0x%x (%lld)\n",