Hi Everyone,
Now that the patchset which creates a command line option to disable
ACS redirection has landed it's time to revisit the P2P patchset for
copy offoad in NVMe fabrics.
I present version 5 wihch no longer does any magic with the ACS bits and
instead will reject P2P transactions between devices that would be affected
by them. A few other cleanups were done which are described in the
changelog below.
This version is based on v4.19-rc1 and a git repo is here:
https://github.com/sbates130272/linux-p2pmem pci-p2p-v5
Thanks,
Logan
--
Changes in v5:
* Rebased on v4.19-rc1
* Drop changing ACS settings in this patchset. Now, the code
will only allow P2P transactions between devices whos
downstream ports do not restrict P2P TLPs.
* Drop the REQ_PCI_P2PDMA block flag and instead use
is_pci_p2pdma_page() to tell if a request is P2P or not. In that
case we check for queue support and enforce using REQ_NOMERGE.
Per feedback from Christoph.
* Drop the pci_p2pdma_unmap_sg() function as it was empty and only
there for symmetry and compatibility with dma_unmap_sg. Per feedback
from Christoph.
* Split off the logic to handle enabling P2P in NVMe fabrics' configfs
into specific helpers in the p2pdma code. Per feedback from Christoph.
* A number of other minor cleanups and fixes as pointed out by
Christoph and others.
Changes in v4:
* Change the original upstream_bridges_match() function to
upstream_bridge_distance() which calculates the distance between two
devices as long as they are behind the same root port. This should
address Bjorn's concerns that the code was to focused on
being behind a single switch.
* The disable ACS function now disables ACS for all bridge ports instead
of switch ports (ie. those that had two upstream_bridge ports).
* Change the pci_p2pmem_alloc_sgl() and pci_p2pmem_free_sgl()
API to be more like sgl_alloc() in that the alloc function returns
the allocated scatterlist and nents is not required bythe free
function.
* Moved the new documentation into the driver-api tree as requested
by Jonathan
* Add SGL alloc and free helpers in the nvmet code so that the
individual drivers can share the code that allocates P2P memory.
As requested by Christoph.
* Cleanup the nvmet_p2pmem_store() function as Christoph
thought my first attempt was ugly.
* Numerous commit message and comment fix-ups
Changes in v3:
* Many more fixes and minor cleanups that were spotted by Bjorn
* Additional explanation of the ACS change in both the commit message
and Kconfig doc. Also, the code that disables the ACS bits is surrounded
explicitly by an #ifdef
* Removed the flag we added to rdma_rw_ctx() in favour of using
is_pci_p2pdma_page(), as suggested by Sagi.
* Adjust pci_p2pmem_find() so that it prefers P2P providers that
are closest to (or the same as) the clients using them. In cases
of ties, the provider is randomly chosen.
* Modify the NVMe Target code so that the PCI device name of the provider
may be explicitly specified, bypassing the logic in pci_p2pmem_find().
(Note: it's still enforced that the provider must be behind the
same switch as the clients).
* As requested by Bjorn, added documentation for driver writers.
Changes in v2:
* Renamed everything to 'p2pdma' per the suggestion from Bjorn as well
as a bunch of cleanup and spelling fixes he pointed out in the last
series.
* To address Alex's ACS concerns, we change to a simpler method of
just disabling ACS behind switches for any kernel that has
CONFIG_PCI_P2PDMA.
* We also reject using devices that employ 'dma_virt_ops' which should
fairly simply handle Jason's concerns that this work might break with
the HFI, QIB and rxe drivers that use the virtual ops to implement
their own special DMA operations.
--
This is a continuation of our work to enable using Peer-to-Peer PCI
memory in the kernel with initial support for the NVMe fabrics target
subsystem. Many thanks go to Christoph Hellwig who provided valuable
feedback to get these patches to where they are today.
The concept here is to use memory that's exposed on a PCI BAR as
data buffers in the NVMe target code such that data can be transferred
from an RDMA NIC to the special memory and then directly to an NVMe
device avoiding system memory entirely. The upside of this is better
QoS for applications running on the CPU utilizing memory and lower
PCI bandwidth required to the CPU (such that systems could be designed
with fewer lanes connected to the CPU).
Due to these trade-offs we've designed the system to only enable using
the PCI memory in cases where the NIC, NVMe devices and memory are all
behind the same PCI switch hierarchy. This will mean many setups that
could likely work well will not be supported so that we can be more
confident it will work and not place any responsibility on the user to
understand their topology. (We chose to go this route based on feedback
we received at the last LSF). Future work may enable these transfers
using a white list of known good root complexes. However, at this time,
there is no reliable way to ensure that Peer-to-Peer transactions are
permitted between PCI Root Ports.
In order to enable this functionality, we introduce a few new PCI
functions such that a driver can register P2P memory with the system.
Struct pages are created for this memory using devm_memremap_pages()
and the PCI bus offset is stored in the corresponding pagemap structure.
When the PCI P2PDMA config option is selected the ACS bits in every
bridge port in the system are turned off to allow traffic to
pass freely behind the root port. At this time, the bit must be disabled
at boot so the IOMMU subsystem can correctly create the groups, though
this could be addressed in the future. There is no way to dynamically
disable the bit and alter the groups.
Another set of functions allow a client driver to create a list of
client devices that will be used in a given P2P transactions and then
use that list to find any P2P memory that is supported by all the
client devices.
In the block layer, we also introduce a P2P request flag to indicate a
given request targets P2P memory as well as a flag for a request queue
to indicate a given queue supports targeting P2P memory. P2P requests
will only be accepted by queues that support it. Also, P2P requests
are marked to not be merged seeing a non-homogenous request would
complicate the DMA mapping requirements.
In the PCI NVMe driver, we modify the existing CMB support to utilize
the new PCI P2P memory infrastructure and also add support for P2P
memory in its request queue. When a P2P request is received it uses the
pci_p2pmem_map_sg() function which applies the necessary transformation
to get the corrent pci_bus_addr_t for the DMA transactions.
In the RDMA core, we also adjust rdma_rw_ctx_init() and
rdma_rw_ctx_destroy() to take a flags argument which indicates whether
to use the PCI P2P mapping functions or not. To avoid odd RDMA devices
that don't use the proper DMA infrastructure this code rejects using
any device that employs the virt_dma_ops implementation.
Finally, in the NVMe fabrics target port we introduce a new
configuration boolean: 'allow_p2pmem'. When set, the port will attempt
to find P2P memory supported by the RDMA NIC and all namespaces. If
supported memory is found, it will be used in all IO transfers. And if
a port is using P2P memory, adding new namespaces that are not supported
by that memory will fail.
These patches have been tested on a number of Intel based systems and
for a variety of RDMA NICs (Mellanox, Broadcomm, Chelsio) and NVMe
SSDs (Intel, Seagate, Samsung) and p2pdma devices (Eideticom,
Microsemi, Chelsio and Everspin) using switches from both Microsemi
and Broadcomm.
Logan Gunthorpe (13):
PCI/P2PDMA: Support peer-to-peer memory
PCI/P2PDMA: Add sysfs group to display p2pmem stats
PCI/P2PDMA: Add PCI p2pmem DMA mappings to adjust the bus offset
PCI/P2PDMA: Introduce configfs/sysfs enable attribute helpers
docs-rst: Add a new directory for PCI documentation
PCI/P2PDMA: Add P2P DMA driver writer's documentation
block: Add PCI P2P flag for request queue and check support for
requests
IB/core: Ensure we map P2P memory correctly in
rdma_rw_ctx_[init|destroy]()
nvme-pci: Use PCI p2pmem subsystem to manage the CMB
nvme-pci: Add support for P2P memory in requests
nvme-pci: Add a quirk for a pseudo CMB
nvmet: Introduce helper functions to allocate and free request SGLs
nvmet: Optionally use PCI P2P memory
Documentation/ABI/testing/sysfs-bus-pci | 25 +
Documentation/driver-api/index.rst | 2 +-
Documentation/driver-api/pci/index.rst | 21 +
Documentation/driver-api/pci/p2pdma.rst | 170 ++++++
Documentation/driver-api/{ => pci}/pci.rst | 0
block/blk-core.c | 14 +
drivers/infiniband/core/rw.c | 11 +-
drivers/nvme/host/core.c | 4 +
drivers/nvme/host/nvme.h | 8 +
drivers/nvme/host/pci.c | 121 ++--
drivers/nvme/target/configfs.c | 36 ++
drivers/nvme/target/core.c | 149 +++++
drivers/nvme/target/nvmet.h | 15 +
drivers/nvme/target/rdma.c | 22 +-
drivers/pci/Kconfig | 17 +
drivers/pci/Makefile | 1 +
drivers/pci/p2pdma.c | 941 +++++++++++++++++++++++++++++
include/linux/blkdev.h | 3 +
include/linux/memremap.h | 6 +
include/linux/mm.h | 18 +
include/linux/pci-p2pdma.h | 124 ++++
include/linux/pci.h | 4 +
22 files changed, 1658 insertions(+), 54 deletions(-)
create mode 100644 Documentation/driver-api/pci/index.rst
create mode 100644 Documentation/driver-api/pci/p2pdma.rst
rename Documentation/driver-api/{ => pci}/pci.rst (100%)
create mode 100644 drivers/pci/p2pdma.c
create mode 100644 include/linux/pci-p2pdma.h
--
2.11.0
Add helpers to allocate and free the SGL in a struct nvmet_req:
int nvmet_req_alloc_sgl(struct nvmet_req *req, struct nvmet_sq *sq)
void nvmet_req_free_sgl(struct nvmet_req *req)
This will be expanded in a future patch to implement peer-to-peer
memory DMAs and should be common with all target drivers. The presently
unused 'sq' argument in the alloc function will be necessary to
decide whether to use peer-to-peer memory and obtain the correct
provider to allocate the memory.
The new helpers are used in nvmet-rdma. Seeing we use req.transfer_len
as the length of the SGL it is set earlier and cleared on any error.
It also seems to be unnecessary to accumulate the length as the map_sgl
functions should only ever be called once per request.
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
Cc: Christoph Hellwig <hch@lst.de>
Cc: Sagi Grimberg <sagi@grimberg.me>
---
drivers/nvme/target/core.c | 18 ++++++++++++++++++
drivers/nvme/target/nvmet.h | 2 ++
drivers/nvme/target/rdma.c | 20 ++++++++++++--------
3 files changed, 32 insertions(+), 8 deletions(-)
The DMA address used when mapping PCI P2P memory must be the PCI bus
address. Thus, introduce pci_p2pmem_map_sg() to map the correct
addresses when using P2P memory. Memory mapped in this way does not
need to be unmapped.
For this, we assume that an SGL passed to these functions contain all
P2P memory or no P2P memory.
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
---
drivers/pci/p2pdma.c | 43 +++++++++++++++++++++++++++++++++++++++++++
include/linux/memremap.h | 1 +
include/linux/pci-p2pdma.h | 7 +++++++
3 files changed, 51 insertions(+)
@@ -191,6 +191,8 @@ int pci_p2pdma_add_resource(struct pci_dev *pdev, int bar, size_t size,pgmap->res.flags=pci_resource_flags(pdev,bar);pgmap->ref=&pdev->p2pdma->devmap_ref;pgmap->type=MEMORY_DEVICE_PCI_P2PDMA;+pgmap->pci_p2pdma_bus_offset=pci_bus_address(pdev,bar)-+pci_resource_start(pdev,bar);addr=devm_memremap_pages(&pdev->dev,pgmap);if(IS_ERR(addr)){
Introduce a quirk to use CMB-like memory on older devices that have
an exposed BAR but do not advertise support for using CMBLOC and
CMBSIZE.
We'd like to use some of these older cards to test P2P memory.
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
---
drivers/nvme/host/nvme.h | 7 +++++++
drivers/nvme/host/pci.c | 24 ++++++++++++++++++++----
2 files changed, 27 insertions(+), 4 deletions(-)
Users of the P2PDMA infrastructure will typically need a way for
the user to tell the kernel to use P2P resources. Typically
this will be a simple on/off boolean operation but sometimes
it may be desirable for the user to specify the exact device to
use for the P2P operation.
Add new helpers for attributes which take a boolean or a PCI device.
Any boolean, or the word 'auto' turn P2P on or off. Specifying a full
PCI device name/BDF will select the specific device.
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
---
drivers/pci/p2pdma.c | 83 ++++++++++++++++++++++++++++++++++++++++++++++
include/linux/pci-p2pdma.h | 15 +++++++++
2 files changed, 98 insertions(+)
@@ -856,3 +857,85 @@ int pci_p2pdma_map_sg(struct device *dev, struct scatterlist *sg, int nents,returnnents;}EXPORT_SYMBOL_GPL(pci_p2pdma_map_sg);++/**+*pci_p2pdma_enable_store-parseaconfigfs/sysfsattributestore+*toenablep2pdma+*@page:contentsofthevaluetobestored+*@p2p_dev:returnsthePCIdevicethatwasselectedtobeused+*(if'auto','noneorabooleanisn'tthestorevalue)+*@use_p2pdma:returnswhethertoenablep2pdmaornot+*+*Parsesanattributevaluetodecidewhethertoenablep2pdma.+*ThevaluecanselectaPCIdevice(usingit'sfullBDFdevice+*name),aboolean,or'auto'.'auto'andatruebooleanvalue+*havethesamemeaning.Afalsevaluedisablesp2pdmaand+*aPCIdeviceenablesittouseaspecificdeviceasthe+*backingprovider.+*+*pci_p2pdma_enable_show()shouldbeusedastheshowoperationfor+*theattribute.+*+*Returns0onsuccess+*/+intpci_p2pdma_enable_store(constchar*page,structpci_dev**p2p_dev,+bool*use_p2pdma)+{+structdevice*dev;++dev=bus_find_device_by_name(&pci_bus_type,NULL,page);+if(dev){+*use_p2pdma=true;+*p2p_dev=to_pci_dev(dev);++if(!pci_has_p2pmem(*p2p_dev)){+pr_err("PCI device has no peer-to-peer memory: %s\n",+page);+pci_dev_put(*p2p_dev);+return-ENODEV;+}++return0;+}elseif(sysfs_streq(page,"auto")){+*use_p2pdma=true;+return0;+}elseif((page[0]=='0'||page[0]=='1')&&!iscntrl(page[1])){+/*+*IftheuserentersaPCIdevicethatdoesn'texist+*like"0000:01:00.1",wedon'twantstrtobooltothink+*it'sa'0'whenit'sclearlynotwhattheuserwanted.+*Sowerequire0'sand1'stobeexactlyonecharacter.+*/+}elseif(!strtobool(page,use_p2pdma)){+return0;+}++pr_err("No such PCI device: %.*s\n",(int)strcspn(page,"\n"),page);+return-ENODEV;+}+EXPORT_SYMBOL_GPL(pci_p2pdma_enable_store);++/**+*pci_p2pdma_enable_show-showaconfigfs/sysfsattributeindicating+*whetherp2pdmaisenabled+*@page:contentsofthestoredvalue+*@p2p_dev:theselectedp2pdevice(NULLifnodeviceisselected)+*@use_p2pdma:whetherp2pdmehasbeenenabled+*+*Attributesthatusepci_p2pdma_enable_store()shouldusethisfunction+*toshowthevalueoftheattribute.+*+*Returns0onsuccess+*/+ssize_tpci_p2pdma_enable_show(char*page,structpci_dev*p2p_dev,+booluse_p2pdma)+{+if(!use_p2pdma)+returnsprintf(page,"none\n");++if(!p2p_dev)+returnsprintf(page,"auto\n");++returnsprintf(page,"%s\n",pci_name(p2p_dev));+}+EXPORT_SYMBOL_GPL(pci_p2pdma_enable_show);
@@ -0,0 +1,20 @@+.. SPDX-License-Identifier: GPL-2.0+============================================+The Linux PCI driver implementer's API guide+============================================++..class:: toc-title++ Table of contents++..toctree::+:maxdepth: 2++ pci++..only:: subproject and html++ Indices+ =======++*:ref:`genindex`
diff --git a/Documentation/driver-api/pci.rst b/Documentation/driver-api/pci/pci.rstsimilarity index 100%rename from Documentation/driver-api/pci.rstrename to Documentation/driver-api/pci/pci.rst
--
2.11.0
We create a configfs attribute in each nvme-fabrics target port to
enable p2p memory use. When enabled, the port will only then use the
p2p memory if a p2p memory device can be found which is behind the
same switch hierarchy as the RDMA port and all the block devices in
use. If the user enabled it and no devices are found, then the system
will silently fall back on using regular memory.
If appropriate, that port will allocate memory for the RDMA buffers
for queues from the p2pmem device falling back to system memory should
anything fail.
Ideally, we'd want to use an NVME CMB buffer as p2p memory. This would
save an extra PCI transfer as the NVME card could just take the data
out of it's own memory. However, at this time, only a limited number
of cards with CMB buffers seem to be available.
Signed-off-by: Stephen Bates <redacted>
Signed-off-by: Steve Wise <redacted>
[hch: partial rewrite of the initial code]
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
---
drivers/nvme/target/configfs.c | 36 +++++++++++
drivers/nvme/target/core.c | 133 ++++++++++++++++++++++++++++++++++++++++-
drivers/nvme/target/nvmet.h | 13 ++++
drivers/nvme/target/rdma.c | 2 +
4 files changed, 183 insertions(+), 1 deletion(-)
Some PCI devices may have memory mapped in a BAR space that's
intended for use in peer-to-peer transactions. In order to enable
such transactions the memory must be registered with ZONE_DEVICE pages
so it can be used by DMA interfaces in existing drivers.
Add an interface for other subsystems to find and allocate chunks of P2P
memory as necessary to facilitate transfers between two PCI peers:
int pci_p2pdma_add_client();
struct pci_dev *pci_p2pmem_find();
void *pci_alloc_p2pmem();
The new interface requires a driver to collect a list of client devices
involved in the transaction with the pci_p2pmem_add_client*() functions
then call pci_p2pmem_find() to obtain any suitable P2P memory. Once
this is done the list is bound to the memory and the calling driver is
free to add and remove clients as necessary (adding incompatible clients
will fail). With a suitable p2pmem device, memory can then be
allocated with pci_alloc_p2pmem() for use in DMA transactions.
Depending on hardware, using peer-to-peer memory may reduce the bandwidth
of the transfer but can significantly reduce pressure on system memory.
This may be desirable in many cases: for example a system could be designed
with a small CPU connected to a PCIe switch by a small number of lanes
which would maximize the number of lanes available to connect to NVMe
devices.
The code is designed to only utilize the p2pmem device if all the devices
involved in a transfer are behind the same PCI bridge. This is because we
have no way of knowing whether peer-to-peer routing between PCIe Root Ports
is supported (PCIe r4.0, sec 1.3.1). Additionally, the benefits of P2P
transfers that go through the RC is limited to only reducing DRAM usage
and, in some cases, coding convenience. The PCI-SIG may be exploring
adding a new capability bit to advertise whether this is possible for
future hardware.
This commit includes significant rework and feedback from Christoph
Hellwig.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
---
drivers/pci/Kconfig | 17 +
drivers/pci/Makefile | 1 +
drivers/pci/p2pdma.c | 761 +++++++++++++++++++++++++++++++++++++++++++++
include/linux/memremap.h | 5 +
include/linux/mm.h | 18 ++
include/linux/pci-p2pdma.h | 102 ++++++
include/linux/pci.h | 4 +
7 files changed, 908 insertions(+)
create mode 100644 drivers/pci/p2pdma.c
create mode 100644 include/linux/pci-p2pdma.h
@@ -26,6 +26,7 @@ obj-$(CONFIG_PCI_SYSCALL) += syscall.oobj-$(CONFIG_PCI_STUB)+=pci-stub.oobj-$(CONFIG_PCI_PF_STUB)+=pci-pf-stub.oobj-$(CONFIG_PCI_ECAM)+=ecam.o+obj-$(CONFIG_PCI_P2PDMA)+=p2pdma.oobj-$(CONFIG_XEN_PCIDEV_FRONTEND)+=xen-pcifront.o# Endpoint library must be initialized before its users
@@ -0,0 +1,761 @@+// SPDX-License-Identifier: GPL-2.0+/*+*PCIPeer2PeerDMAsupport.+*+*Copyright(c)2016-2018,LoganGunthorpe+*Copyright(c)2016-2017,MicrosemiCorporation+*Copyright(c)2017,ChristophHellwig+*Copyright(c)2018,EideticomInc.+*/++#define pr_fmt(fmt) "pci-p2pdma: " fmt+#include<linux/pci-p2pdma.h>+#include<linux/module.h>+#include<linux/slab.h>+#include<linux/genalloc.h>+#include<linux/memremap.h>+#include<linux/percpu-refcount.h>+#include<linux/random.h>+#include<linux/seq_buf.h>++structpci_p2pdma{+structpercpu_refdevmap_ref;+structcompletiondevmap_ref_done;+structgen_pool*pool;+boolp2pmem_published;+};++staticvoidpci_p2pdma_percpu_release(structpercpu_ref*ref)+{+structpci_p2pdma*p2p=+container_of(ref,structpci_p2pdma,devmap_ref);++complete_all(&p2p->devmap_ref_done);+}++staticvoidpci_p2pdma_percpu_kill(void*data)+{+structpercpu_ref*ref=data;++if(percpu_ref_is_dying(ref))+return;++percpu_ref_kill(ref);+}++staticvoidpci_p2pdma_release(void*data)+{+structpci_dev*pdev=data;++if(!pdev->p2pdma)+return;++wait_for_completion(&pdev->p2pdma->devmap_ref_done);+percpu_ref_exit(&pdev->p2pdma->devmap_ref);++gen_pool_destroy(pdev->p2pdma->pool);+pdev->p2pdma=NULL;+}++staticintpci_p2pdma_setup(structpci_dev*pdev)+{+interror=-ENOMEM;+structpci_p2pdma*p2p;++p2p=devm_kzalloc(&pdev->dev,sizeof(*p2p),GFP_KERNEL);+if(!p2p)+return-ENOMEM;++p2p->pool=gen_pool_create(PAGE_SHIFT,dev_to_node(&pdev->dev));+if(!p2p->pool)+gotoout;++init_completion(&p2p->devmap_ref_done);+error=percpu_ref_init(&p2p->devmap_ref,+pci_p2pdma_percpu_release,0,GFP_KERNEL);+if(error)+gotoout_pool_destroy;++percpu_ref_switch_to_atomic_sync(&p2p->devmap_ref);++error=devm_add_action_or_reset(&pdev->dev,pci_p2pdma_release,pdev);+if(error)+gotoout_pool_destroy;++pdev->p2pdma=p2p;++return0;++out_pool_destroy:+gen_pool_destroy(p2p->pool);+out:+devm_kfree(&pdev->dev,p2p);+returnerror;+}++/**+*pci_p2pdma_add_resource-addmemoryforuseasp2pmemory+*@pdev:thedevicetoaddthememoryto+*@bar:PCIBARtoadd+*@size:sizeofthememorytoadd,maybezerotousethewholeBAR+*@offset:offsetintothePCIBAR+*+*ThememorywillbegivenZONE_DEVICEstructpagessothatitmay+*beusedwithanyDMArequest.+*/+intpci_p2pdma_add_resource(structpci_dev*pdev,intbar,size_tsize,+u64offset)+{+structdev_pagemap*pgmap;+void*addr;+interror;++if(!(pci_resource_flags(pdev,bar)&IORESOURCE_MEM))+return-EINVAL;++if(offset>=pci_resource_len(pdev,bar))+return-EINVAL;++if(!size)+size=pci_resource_len(pdev,bar)-offset;++if(size+offset>pci_resource_len(pdev,bar))+return-EINVAL;++if(!pdev->p2pdma){+error=pci_p2pdma_setup(pdev);+if(error)+returnerror;+}++pgmap=devm_kzalloc(&pdev->dev,sizeof(*pgmap),GFP_KERNEL);+if(!pgmap)+return-ENOMEM;++pgmap->res.start=pci_resource_start(pdev,bar)+offset;+pgmap->res.end=pgmap->res.start+size-1;+pgmap->res.flags=pci_resource_flags(pdev,bar);+pgmap->ref=&pdev->p2pdma->devmap_ref;+pgmap->type=MEMORY_DEVICE_PCI_P2PDMA;++addr=devm_memremap_pages(&pdev->dev,pgmap);+if(IS_ERR(addr)){+error=PTR_ERR(addr);+gotopgmap_free;+}++error=gen_pool_add_virt(pdev->p2pdma->pool,(unsignedlong)addr,+pci_bus_address(pdev,bar)+offset,+resource_size(&pgmap->res),dev_to_node(&pdev->dev));+if(error)+gotopgmap_free;++error=devm_add_action_or_reset(&pdev->dev,pci_p2pdma_percpu_kill,+&pdev->p2pdma->devmap_ref);+if(error)+gotopgmap_free;++pci_info(pdev,"added peer-to-peer DMA memory %pR\n",+&pgmap->res);++return0;++pgmap_free:+devres_free(pgmap);+returnerror;+}+EXPORT_SYMBOL_GPL(pci_p2pdma_add_resource);++staticstructpci_dev*find_parent_pci_dev(structdevice*dev)+{+structdevice*parent;++dev=get_device(dev);++while(dev){+if(dev_is_pci(dev))+returnto_pci_dev(dev);++parent=get_device(dev->parent);+put_device(dev);+dev=parent;+}++returnNULL;+}++/*+*CheckifaPCIbridgehasit'sACSredirectionbitssettoredirectP2P+*TLPsupstreamviaACS.Returns1ifthepacketswillberedirected+*upstream,0otherwise.+*/+staticintpci_bridge_has_acs_redir(structpci_dev*dev)+{+intpos;+u16ctrl;++pos=pci_find_ext_capability(dev,PCI_EXT_CAP_ID_ACS);+if(!pos)+return0;++pci_read_config_word(dev,pos+PCI_ACS_CTRL,&ctrl);++if(ctrl&(PCI_ACS_RR|PCI_ACS_CR|PCI_ACS_EC))+return1;++return0;+}++staticvoidseq_buf_print_bus_devfn(structseq_buf*buf,structpci_dev*dev)+{+if(!buf)+return;++seq_buf_printf(buf,"%04x:%02x:%02x.%x;",pci_domain_nr(dev->bus),+dev->bus->number,PCI_SLOT(dev->devfn),+PCI_FUNC(dev->devfn));+}++/*+*Findthedistancethroughthenearestcommonupstreambridgebetween+*twoPCIdevices.+*+*Ifthetwodevicesarethesamedevicethen0willbereturned.+*+*Iftherearetwovirtualfunctionsofthesamedevicebehindthesame+*bridgeportthen2willbereturned(onestepdowntothePCIeswitch,+*thenonestepbacktothesamedevice).+*+*InthecasewheretwodevicesareconnectedtothesamePCIeswitch,the+*value4willbereturned.ThiscorrespondstothefollowingPCItree:+*+*-+RootPort+*\+SwitchUpstreamPort+*+-+SwitchDownstreamPort+*+\-DeviceA+*\-+SwitchDownstreamPort+*\-DeviceB+*+*Thedistanceis4becausewetraversefromDeviceAthroughthedownstream+*portoftheswitch,tothecommonupstreamport,backuptothesecond+*downstreamportandthentoDeviceB.+*+*Anytwodevicesthatdon'thaveacommonupstreambridgewillreturn-1.+*InthiswaydevicesonseparatePCIerootportswillberejected,which+*iswhatwewantforpeer-to-peerseeingeachPCIerootportdefinesa+*separatehierarchydomainandthere'snowaytodeterminewhethertheroot+*complexsupportsforwardingbetweenthem.+*+*InthecasewheretwodevicesareconnectedtodifferentPCIeswitches,+*thisfunctionwillstillreturnapositivedistanceaslongasboth+*switchesevenutallyhaveacommonupstreambridge.Notethiscovers+*thecaseofusingmultiplePCIeswitchestoachieveadesiredlevelof+*fan-outfromarootport.Theexactdistancewillbeafunctionofthe+*numberofswitchesbetweenDeviceAandDeviceB.+*+*IfabridgewhichhasanyACSredirectionbitssetisinthepath+*thenthisfunctionswillreturn-2.Thisissowerejectany+*caseswheretheTLPsareforwardedupintotherootcomplex.+*Inthiscase,alistofallinfringingbridgeaddresseswillbe+*populatedinacs_list(assumingit'snon-null)forprintkpurposes.+*/+staticintupstream_bridge_distance(structpci_dev*a,+structpci_dev*b,+structseq_buf*acs_list)+{+intdist_a=0;+intdist_b=0;+structpci_dev*bb=NULL;+intacs_cnt=0;++/*+*Note,wedon'tneedtotakereferencestodevicesreturnedby+*pci_upstream_bridge()seeingweholdareferencetoachild+*devicewhichwillalreadyholdareferencetotheupstreambridge.+*/++while(a){+dist_b=0;++if(pci_bridge_has_acs_redir(a)){+seq_buf_print_bus_devfn(acs_list,a);+acs_cnt++;+}++bb=b;++while(bb){+if(a==bb)+gotocheck_b_path_acs;++bb=pci_upstream_bridge(bb);+dist_b++;+}++a=pci_upstream_bridge(a);+dist_a++;+}++return-1;++check_b_path_acs:+bb=b;++while(bb){+if(a==bb)+break;++if(pci_bridge_has_acs_redir(bb)){+seq_buf_print_bus_devfn(acs_list,bb);+acs_cnt++;+}++bb=pci_upstream_bridge(bb);+}++if(acs_cnt)+return-2;++returndist_a+dist_b;+}++staticintupstream_bridge_distance_warn(structpci_dev*provider,+structpci_dev*client)+{+structseq_bufacs_list;+intret;++seq_buf_init(&acs_list,kmalloc(PAGE_SIZE,GFP_KERNEL),PAGE_SIZE);++ret=upstream_bridge_distance(provider,client,&acs_list);+if(ret==-2){+pci_warn(client,"cannot be used for peer-to-peer DMA as ACS redirect is set between the client and provider\n");+/* Drop final semicolon */+acs_list.buffer[acs_list.len-1]=0;+pci_warn(client,"to disable ACS redirect for this path, add the kernel parameter: pci=disable_acs_redir=%s\n",+acs_list.buffer);++}elseif(ret<0){+pci_warn(client,"cannot be used for peer-to-peer DMA as the client and provider do not share an upstream bridge\n");+}++kfree(acs_list.buffer);++returnret;+}++structpci_p2pdma_client{+structlist_headlist;+structpci_dev*client;+structpci_dev*provider;+};++/**+*pci_p2pdma_add_client-allocateanewelementinaclientdevicelist+*@head:listheadofp2pdmaclients+*@dev:devicetoaddtothelist+*+*Thisadds@devtoalistofclientsusedbyap2pdmadevice.+*Thislistshouldbepassedtopci_p2pmem_find().Oncepci_p2pmem_find()has+*beencalledsuccessfully,thelistwillbeboundtoaspecificp2pdma+*deviceandnewclientscanonlybeaddedtothelistiftheyare+*supportedbythatp2pdmadevice.+*+*Thecallerisexpectedtohavealockwhichprotects@headasnecessary+*sothatnoneofthepci_p2pfunctionscanbecalledconcurrently+*onthatlist.+*+*Returns0iftheclientwassuccessfullyadded.+*/+intpci_p2pdma_add_client(structlist_head*head,structdevice*dev)+{+structpci_p2pdma_client*item,*new_item;+structpci_dev*provider=NULL;+structpci_dev*client;+intret;++if(IS_ENABLED(CONFIG_DMA_VIRT_OPS)&&dev->dma_ops==&dma_virt_ops){+dev_warn(dev,"cannot be used for peer-to-peer DMA because the driver makes use of dma_virt_ops\n");+return-ENODEV;+}++client=find_parent_pci_dev(dev);+if(!client){+dev_warn(dev,"cannot be used for peer-to-peer DMA as it is not a PCI device\n");+return-ENODEV;+}++item=list_first_entry_or_null(head,structpci_p2pdma_client,list);+if(item&&item->provider){+provider=item->provider;++ret=upstream_bridge_distance_warn(provider,client);+if(ret<0){+ret=-EXDEV;+gotoput_client;+}+}++new_item=kzalloc(sizeof(*new_item),GFP_KERNEL);+if(!new_item){+ret=-ENOMEM;+gotoput_client;+}++new_item->client=client;+new_item->provider=pci_dev_get(provider);++list_add_tail(&new_item->list,head);++return0;++put_client:+pci_dev_put(client);+returnret;+}+EXPORT_SYMBOL_GPL(pci_p2pdma_add_client);++staticvoidpci_p2pdma_client_free(structpci_p2pdma_client*item)+{+list_del(&item->list);+pci_dev_put(item->client);+pci_dev_put(item->provider);+kfree(item);+}++/**+*pci_p2pdma_remove_client-removeandfreeap2pdmaclient+*@head:listheadofp2pdmaclients+*@dev:devicetoremovefromthelist+*+*Thisremoves@devfromalistofclientsusedbyap2pdmadevice.+*Thecallerisexpectedtohavealockwhichprotects@headasnecessary+*sothatnoneofthepci_p2pfunctionscanbecalledconcurrently+*onthatlist.+*/+voidpci_p2pdma_remove_client(structlist_head*head,structdevice*dev)+{+structpci_p2pdma_client*pos,*tmp;+structpci_dev*pdev;++pdev=find_parent_pci_dev(dev);+if(!pdev)+return;++list_for_each_entry_safe(pos,tmp,head,list){+if(pos->client!=pdev)+continue;++pci_p2pdma_client_free(pos);+}++pci_dev_put(pdev);+}+EXPORT_SYMBOL_GPL(pci_p2pdma_remove_client);++/**+*pci_p2pdma_client_list_free-freeanentirelistofp2pdmaclients+*@head:listheadofp2pdmaclients+*+*Thisremovesalldevicesinalistofclientsusedbyap2pdmadevice.+*Thecallerisexpectedtohavealockwhichprotects@headasnecessary+*sothatnoneofthepci_p2pdmafunctionscanbecalledconcurrently+*onthatlist.+*/+voidpci_p2pdma_client_list_free(structlist_head*head)+{+structpci_p2pdma_client*pos,*tmp;++list_for_each_entry_safe(pos,tmp,head,list)+pci_p2pdma_client_free(pos);+}+EXPORT_SYMBOL_GPL(pci_p2pdma_client_list_free);++/**+*pci_p2pdma_distance-Determivethecumulativedistancebetween+*ap2pdmaproviderandtheclientsinuse.+*@provider:p2pdmaprovidertocheckagainsttheclientlist+*@clients:listofdevicestocheck(NULL-terminated)+*@verbose:iftrue,printwarningsfordeviceswhenwereturn-1+*+*Returns-1ifanyoftheclientsarenotcompatible(behindthesame+*rootportastheprovider),otherwisereturnsapositivenumberwhere+*thelowernumberisthepreferrablechoice.(Ifthere'soneclient+*that'sthesameastheprovideritwillreturn0,whichisbestchoice).+*+*Fornow,"compatible"meanstheproviderandtheclientsareallbehind+*thesamePCIrootport.Thiscutsoutcasesthatmayworkbutissafest+*fortheuser.Futureworkcanexpandthistowhite-listrootcomplexesthat+*cansafelyforwardbetweeneachports.+*/+intpci_p2pdma_distance(structpci_dev*provider,structlist_head*clients,+boolverbose)+{+structpci_p2pdma_client*pos;+intret;+intdistance=0;+boolnot_supported=false;++if(list_empty(clients))+return-1;++list_for_each_entry(pos,clients,list){+if(verbose)+ret=upstream_bridge_distance_warn(provider,+pos->client);+else+ret=upstream_bridge_distance(provider,pos->client,+NULL);++if(ret<0)+not_supported=true;++if(not_supported&&!verbose)+break;++distance+=ret;+}++if(not_supported)+return-1;++returndistance;+}+EXPORT_SYMBOL_GPL(pci_p2pdma_distance);++/**+*pci_p2pdma_assign_provider-Checkcompatibily(asperpci_p2pdma_distance)+*andassignaprovidertoalistofclients+*@provider:p2pdmaprovidertoassigntotheclientlist+*@clients:listofdevicestocheck(NULL-terminated)+*+*Returnsfalseifanyoftheclientsarenotcompatible,trueifthe+*providerwassuccessfullyassignedtotheclients.+*/+boolpci_p2pdma_assign_provider(structpci_dev*provider,+structlist_head*clients)+{+structpci_p2pdma_client*pos;++if(pci_p2pdma_distance(provider,clients,true)<0)+returnfalse;++list_for_each_entry(pos,clients,list)+pos->provider=provider;++returntrue;+}+EXPORT_SYMBOL_GPL(pci_p2pdma_assign_provider);++/**+*pci_has_p2pmem-checkifagivenPCIdevicehaspublishedanyp2pmem+*@pdev:PCIdevicetocheck+*/+boolpci_has_p2pmem(structpci_dev*pdev)+{+returnpdev->p2pdma&&pdev->p2pdma->p2pmem_published;+}+EXPORT_SYMBOL_GPL(pci_has_p2pmem);++/**+*pci_p2pmem_find-findapeer-to-peerDMAmemorydevicecompatiblewith+*thespecifiedlistofclientsandshortestdistance(asdetermined+*bypci_p2pmem_dma())+*@clients:listofdevicestocheck(NULL-terminated)+*+*Ifmultipledevicesarebehindthesameswitch,theone"closest"tothe+*clientdevicesinusewillbechosenfirst.(Soifoneoftheprovidersare+*thesameasoneoftheclients,thatproviderwillbeusedaheadofany+*otherprovidersthatareunrelated).Ifmultipleprovidersareanequal+*distanceaway,onewillbechosenatrandom.+*+*ReturnsapointertothePCIdevicewithareferencetaken(usepci_dev_put+*toreturnthereference)orNULLifnocompatibledeviceisfound.The+*foundproviderwillalsobeassignedtotheclientlist.+*/+structpci_dev*pci_p2pmem_find(structlist_head*clients)+{+structpci_dev*pdev=NULL;+structpci_p2pdma_client*pos;+intdistance;+intclosest_distance=INT_MAX;+structpci_dev**closest_pdevs;+intdev_cnt=0;+constintmax_devs=PAGE_SIZE/sizeof(*closest_pdevs);+inti;++closest_pdevs=kmalloc(PAGE_SIZE,GFP_KERNEL);++while((pdev=pci_get_device(PCI_ANY_ID,PCI_ANY_ID,pdev))){+if(!pci_has_p2pmem(pdev))+continue;++distance=pci_p2pdma_distance(pdev,clients,false);+if(distance<0||distance>closest_distance)+continue;++if(distance==closest_distance&&dev_cnt>=max_devs)+continue;++if(distance<closest_distance){+for(i=0;i<dev_cnt;i++)+pci_dev_put(closest_pdevs[i]);++dev_cnt=0;+closest_distance=distance;+}++closest_pdevs[dev_cnt++]=pci_dev_get(pdev);+}++if(dev_cnt)+pdev=pci_dev_get(closest_pdevs[prandom_u32_max(dev_cnt)]);++for(i=0;i<dev_cnt;i++)+pci_dev_put(closest_pdevs[i]);++if(pdev)+list_for_each_entry(pos,clients,list)+pos->provider=pdev;++kfree(closest_pdevs);+returnpdev;+}+EXPORT_SYMBOL_GPL(pci_p2pmem_find);++/**+*pci_alloc_p2p_mem-allocatepeer-to-peerDMAmemory+*@pdev:thedevicetoallocatememoryfrom+*@size:numberofbytestoallocate+*+*ReturnstheallocatedmemoryorNULLonerror.+*/+void*pci_alloc_p2pmem(structpci_dev*pdev,size_tsize)+{+void*ret;++if(unlikely(!pdev->p2pdma))+returnNULL;++if(unlikely(!percpu_ref_tryget_live(&pdev->p2pdma->devmap_ref)))+returnNULL;++ret=(void*)gen_pool_alloc(pdev->p2pdma->pool,size);++if(unlikely(!ret))+percpu_ref_put(&pdev->p2pdma->devmap_ref);++returnret;+}+EXPORT_SYMBOL_GPL(pci_alloc_p2pmem);++/**+*pci_free_p2pmem-allocatepeer-to-peerDMAmemory+*@pdev:thedevicethememorywasallocatedfrom+*@addr:addressofthememorythatwasallocated+*@size:numberofbytesthatwasallocated+*/+voidpci_free_p2pmem(structpci_dev*pdev,void*addr,size_tsize)+{+gen_pool_free(pdev->p2pdma->pool,(uintptr_t)addr,size);+percpu_ref_put(&pdev->p2pdma->devmap_ref);+}+EXPORT_SYMBOL_GPL(pci_free_p2pmem);++/**+*pci_virt_to_bus-returnthePCIbusaddressforagivenvirtual+*addressobtainedwithpci_alloc_p2pmem()+*@pdev:thedevicethememorywasallocatedfrom+*@addr:addressofthememorythatwasallocated+*/+pci_bus_addr_tpci_p2pmem_virt_to_bus(structpci_dev*pdev,void*addr)+{+if(!addr)+return0;+if(!pdev->p2pdma)+return0;++/*+*Note:whenweaddedthememorytothepoolweusedthePCI+*busaddressasthephysicaladdress.Sogen_pool_virt_to_phys()+*actuallyreturnsthebusaddressdespitethemisleadingname.+*/+returngen_pool_virt_to_phys(pdev->p2pdma->pool,(unsignedlong)addr);+}+EXPORT_SYMBOL_GPL(pci_p2pmem_virt_to_bus);++/**+*pci_p2pmem_alloc_sgl-allocatepeer-to-peerDMAmemoryinascatterlist+*@pdev:thedevicetoallocatememoryfrom+*@sgl:theallocatedscatterlist+*@nents:thenumberofSGentriesinthelist+*@length:numberofbytestoallocate+*+*Returns0onsuccess+*/+structscatterlist*pci_p2pmem_alloc_sgl(structpci_dev*pdev,+unsignedint*nents,u32length)+{+structscatterlist*sg;+void*addr;++sg=kzalloc(sizeof(*sg),GFP_KERNEL);+if(!sg)+returnNULL;++sg_init_table(sg,1);++addr=pci_alloc_p2pmem(pdev,length);+if(!addr)+gotoout_free_sg;++sg_set_buf(sg,addr,length);+*nents=1;+returnsg;++out_free_sg:+kfree(sg);+returnNULL;+}+EXPORT_SYMBOL_GPL(pci_p2pmem_alloc_sgl);++/**+*pci_p2pmem_free_sgl-freeascatterlistallocatedbypci_p2pmem_alloc_sgl()+*@pdev:thedevicetoallocatememoryfrom+*@sgl:theallocatedscatterlist+*@nents:thenumberofSGentriesinthelist+*/+voidpci_p2pmem_free_sgl(structpci_dev*pdev,structscatterlist*sgl)+{+structscatterlist*sg;+intcount;++for_each_sg(sgl,sg,INT_MAX,count){+if(!sg)+break;++pci_free_p2pmem(pdev,sg_virt(sg),sg->length);+}+kfree(sgl);+}+EXPORT_SYMBOL_GPL(pci_p2pmem_free_sgl);++/**+*pci_p2pmem_publish-publishthepeer-to-peerDMAmemoryforuseby+*otherdeviceswithpci_p2pmem_find()+*@pdev:thedevicewithpeer-to-peerDMAmemorytopublish+*@publish:settotruetopublishthememory,falsetounpublishit+*+*PublishedmemorycanbeusedbyotherPCIdevicedriversfor+*peer-2-peerDMAoperations.Non-publishedmemoryisreservedfor+*exlusiveuseofthedevicedriverthatregistersthepeer-to-peer+*memory.+*/+voidpci_p2pmem_publish(structpci_dev*pdev,boolpublish)+{+if(publish&&!pdev->p2pdma)+return;++pdev->p2pdma->p2pmem_published=publish;+}+EXPORT_SYMBOL_GPL(pci_p2pmem_publish);
@@ -439,6 +440,9 @@ struct pci_dev {#ifdef CONFIG_PCI_PASIDu16pasid_features;#endif+#ifdef CONFIG_PCI_P2PDMA+structpci_p2pdma*p2pdma;+#endifphys_addr_trom;/* Physical address if not from BAR */size_tromlen;/* Length if not from BAR */char*driver_override;/* Driver name to force a match */
Add a restructured text file describing how to write drivers
with support for P2P DMA transactions. The document describes
how to use the APIs that were added in the previous few
commits.
Also adds an index for the PCI documentation tree even though this
is the only PCI document that has been converted to restructured text
at this time.
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
Cc: Jonathan Corbet <corbet@lwn.net>
---
Documentation/driver-api/pci/index.rst | 1 +
Documentation/driver-api/pci/p2pdma.rst | 170 ++++++++++++++++++++++++++++++++
2 files changed, 171 insertions(+)
create mode 100644 Documentation/driver-api/pci/p2pdma.rst
@@ -0,0 +1,170 @@+.. SPDX-License-Identifier: GPL-2.0++============================+PCI Peer-to-Peer DMA Support+============================++The PCI bus has pretty decent support for performing DMA transfers+between two devices on the bus. This type of transaction is henceforth+called Peer-to-Peer (or P2P). However, there are a number of issues that+make P2P transactions tricky to do in a perfectly safe way.++One of the biggest issues is that PCI doesn't require forwarding+transactions between hierarchy domains, and in PCIe, each Root Port+defines a separate hierarchy domain. To make things worse, there is no+simple way to determine if a given Root Complex supports this or not.+(See PCIe r4.0, sec 1.3.1). Therefore, as of this writing, the kernel+only supports doing P2P when the endpoints involved are all behind the+same PCI bridge, as such devices are all in the same PCI hierarchy+domain, and the spec guarantees that all transacations within the+hierarchy will be routable, but it does not require routing+between hierarchies.++The second issue is that to make use of existing interfaces in Linux,+memory that is used for P2P transactions needs to be backed by struct+pages. However, PCI BARs are not typically cache coherent so there are+a few corner case gotchas with these pages so developers need to+be careful about what they do with them.+++Driver Writer's Guide+=====================++In a given P2P implementation there may be three or more different+types of kernel drivers in play:++* Provider - A driver which provides or publishes P2P resources like+ memory or doorbell registers to other drivers.+* Client - A driver which makes use of a resource by setting up a+ DMA transaction to or from it.+* Orchestrator - A driver which orchestrates the flow of data between+ clients and providers++In many cases there could be overlap between these three types (i.e.,+it may be typical for a driver to be both a provider and a client).++For example, in the NVMe Target Copy Offload implementation:++* The NVMe PCI driver is both a client, provider and orchestrator+ in that it exposes any CMB (Controller Memory Buffer) as a P2P memory+ resource (provider), it accepts P2P memory pages as buffers in requests+ to be used directly (client) and it can also make use the CMB as+ submission queue entries.+* The RDMA driver is a client in this arrangement so that an RNIC+ can DMA directly to the memory exposed by the NVMe device.+* The NVMe Target driver (nvmet) can orchestrate the data from the RNIC+ to the P2P memory (CMB) and then to the NVMe device (and vice versa).++This is currently the only arrangement supported by the kernel but+one could imagine slight tweaks to this that would allow for the same+functionality. For example, if a specific RNIC added a BAR with some+memory behind it, its driver could add support as a P2P provider and+then the NVMe Target could use the RNIC's memory instead of the CMB+in cases where the NVMe cards in use do not have CMB support.+++Provider Drivers+----------------++A provider simply needs to register a BAR (or a portion of a BAR)+as a P2P DMA resource using :c:func:`pci_p2pdma_add_resource()`.+This will register struct pages for all the specified memory.++After that it may optionally publish all of its resources as+P2P memory using :c:func:`pci_p2pmem_publish()`. This will allow+any orchestrator drivers to find and use the memory. When marked in+this way, the resource must be regular memory with no side effects.++For the time being this is fairly rudimentary in that all resources+are typically going to be P2P memory. Future work will likely expand+this to include other types of resources like doorbells.+++Client Drivers+--------------++A client driver typically only has to conditionally change its DMA map+routine to use the mapping function :c:func:`pci_p2pdma_map_sg()` instead+of the usual :c:func:`dma_map_sg()` function. Memory mapped in this+way does not need to be unmapped.++The client may also, optionally, make use of+:c:func:`is_pci_p2pdma_page()` to determine when to use the P2P mapping+functions and when to use the regular mapping functions. In some+situations, it may be more appropriate to use a flag to indicate a+given request is P2P memory and map appropriately (for example the+block layer uses a flag to keep P2P memory out of queues that do not+have P2P client support). It is important to ensure that struct pages that+back P2P memory stay out of code that does not have support for them.+++Orchestrator Drivers+--------------------++The first task an orchestrator driver must do is compile a list of+all client devices that will be involved in a given transaction. For+example, the NVMe Target driver creates a list including all NVMe+devices and the RNIC in use. The list is stored as an anonymous struct+list_head which must be initialized with the usual INIT_LIST_HEAD.+The following functions may then be used to add to, remove from and free+the list of clients with the functions :c:func:`pci_p2pdma_add_client()`,+:c:func:`pci_p2pdma_remove_client()` and+:c:func:`pci_p2pdma_client_list_free()`.++With the client list in hand, the orchestrator may then call+:c:func:`pci_p2pmem_find()` to obtain a published P2P memory provider+that is supported (behind the same root port) as all the clients. If more+than one provider is supported, the one nearest to all the clients will+be chosen first. If there are more than one provider is an equal distance+away, the one returned will be chosen at random. This function returns the PCI+device to use for the provider with a reference taken and therefore+when it's no longer needed it should be returned with pci_dev_put().++Alternatively, if the orchestrator knows (via some other means)+which provider it wants to use it may use :c:func:`pci_has_p2pmem()`+to determine if it has P2P memory and :c:func:`pci_p2pdma_distance()`+to determine the cumulative distance between it and a potential+list of clients.++With a supported provider in hand, the driver can then call+:c:func:`pci_p2pdma_assign_provider()` to assign the provider+to the client list. This function returns false if any of the+clients are unsupported by the provider.++Once a provider is assigned to a client list via either+:c:func:`pci_p2pmem_find()` or :c:func:`pci_p2pdma_assign_provider()`,+the list is permanently bound to the provider such that any new clients+added to the list must be supported by the already selected provider.+If they are not supported, :c:func:`pci_p2pdma_add_client()` will return+an error. In this way, orchestrators are free to add and remove devices+without having to recheck support or tear down existing transfers to+change P2P providers.++Once a provider is selected, the orchestrator can then use+:c:func:`pci_alloc_p2pmem()` and :c:func:`pci_free_p2pmem()` to+allocate P2P memory from the provider. :c:func:`pci_p2pmem_alloc_sgl()`+and :c:func:`pci_p2pmem_free_sgl()` are convenience functions for+allocating scatter-gather lists with P2P memory.++Struct Page Caveats+-------------------++Driver writers should be very careful about not passing these special+struct pages to code that isn't prepared for it. At this time, the kernel+interfaces do not have any checks for ensuring this. This obviously+precludes passing these pages to userspace.++P2P memory is also technically IO memory but should never have any side+effects behind it. Thus, the order of loads and stores should not be important+and ioreadX(), iowriteX() and friends should not be necessary.+However, as the memory is not cache coherent, if access ever needs to+be protected by a spinlock then :c:func:`mmiowb()` must be used before+unlocking the lock. (See ACQUIRES VS I/O ACCESSES in+Documentation/memory-barriers.txt)+++P2P DMA Support Library+=====================++..kernel-doc:: drivers/pci/p2pdma.c+:export:
In order to use PCI P2P memory the pci_p2pmem_map_sg() function must be
called to map the correct PCI bus address.
To do this, check the first page in the scatter list to see if it is P2P
memory or not. At the moment, scatter lists that contain P2P memory must
be homogeneous so if the first page is P2P the entire SGL should be P2P.
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
Reviewed-by: Christoph Hellwig <hch@lst.de>
---
drivers/infiniband/core/rw.c | 11 +++++++++--
1 file changed, 9 insertions(+), 2 deletions(-)
@@ -602,7 +607,9 @@ void rdma_rw_ctx_destroy(struct rdma_rw_ctx *ctx, struct ib_qp *qp, u8 port_num,break;}-ib_dma_unmap_sg(qp->pd->device,sg,sg_cnt,dir);+/* P2PDMA contexts do not need to be unmapped */+if(!is_pci_p2pdma_page(sg_page(sg)))+ib_dma_unmap_sg(qp->pd->device,sg,sg_cnt,dir);}EXPORT_SYMBOL(rdma_rw_ctx_destroy);
Add a sysfs group to display statistics about P2P memory that is
registered in each PCI device.
Attributes in the group display the total amount of P2P memory, the
amount available and whether it is published or not.
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
---
Documentation/ABI/testing/sysfs-bus-pci | 25 +++++++++++++++
drivers/pci/p2pdma.c | 54 +++++++++++++++++++++++++++++++++
2 files changed, 79 insertions(+)
@@ -323,3 +323,28 @@ Description: This is similar to /sys/bus/pci/drivers_autoprobe, but affects only the VFs associated with a specific PF.++What: /sys/bus/pci/devices/.../p2pmem/available+Date: November 2017+Contact: Logan Gunthorpe <logang@deltatee.com>+Description:+ If the device has any Peer-to-Peer memory registered, this+ file contains the amount of memory that has not been+ allocated (in decimal).++What: /sys/bus/pci/devices/.../p2pmem/size+Date: November 2017+Contact: Logan Gunthorpe <logang@deltatee.com>+Description:+ If the device has any Peer-to-Peer memory registered, this+ file contains the total amount of memory that the device+ provides (in decimal).++What: /sys/bus/pci/devices/.../p2pmem/published+Date: November 2017+Contact: Logan Gunthorpe <logang@deltatee.com>+Description:+ If the device has any Peer-to-Peer memory registered, this+ file contains a '1' if the memory has been published for+ use inside the kernel or a '0' if it is only intended+ for use within the driver that published it.
Register the CMB buffer as p2pmem and use the appropriate allocation
functions to create and destroy the IO submission queues.
If the CMB supports WDS and RDS, publish it for use as P2P memory
by other devices.
Kernels without CONFIG_PCI_P2PDMA will also no longer support NVMe CMB.
However, seeing the main use-cases for the CMB is P2P operations,
this seems like a reasonable dependency.
We drop the __iomem safety on the buffer seeing that, by convention, it's
safe to directly access memory mapped by memremap()/devm_memremap_pages().
Architectures where this is not safe will not be supported by memremap()
and therefore will not be support PCI P2P and have no support for CMB.
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
---
drivers/nvme/host/pci.c | 80 +++++++++++++++++++++++++++----------------------
1 file changed, 45 insertions(+), 35 deletions(-)
@@ -1315,12 +1321,21 @@ static int nvme_cmb_qdepth(struct nvme_dev *dev, int nr_io_queues,staticintnvme_alloc_sq_cmds(structnvme_dev*dev,structnvme_queue*nvmeq,intqid,intdepth){-/* CMB SQEs will be mapped before creation */-if(qid&&dev->cmb&&use_cmb_sqes&&(dev->cmbsz&NVME_CMBSZ_SQS))-return0;+structpci_dev*pdev=to_pci_dev(dev->dev);++if(qid&&dev->cmb_use_sqes&&(dev->cmbsz&NVME_CMBSZ_SQS)){+nvmeq->sq_cmds=pci_alloc_p2pmem(pdev,SQ_SIZE(depth));+nvmeq->sq_dma_addr=pci_p2pmem_virt_to_bus(pdev,+nvmeq->sq_cmds);+nvmeq->sq_cmds_is_io=true;+}++if(!nvmeq->sq_cmds){+nvmeq->sq_cmds=dma_alloc_coherent(dev->dev,SQ_SIZE(depth),+&nvmeq->sq_dma_addr,GFP_KERNEL);+nvmeq->sq_cmds_is_io=false;+}-nvmeq->sq_cmds=dma_alloc_coherent(dev->dev,SQ_SIZE(depth),-&nvmeq->sq_dma_addr,GFP_KERNEL);if(!nvmeq->sq_cmds)return-ENOMEM;return0;
@@ -1397,13 +1412,6 @@ static int nvme_create_queue(struct nvme_queue *nvmeq, int qid)intresult;s16vector;-if(dev->cmb&&use_cmb_sqes&&(dev->cmbsz&NVME_CMBSZ_SQS)){-unsignedoffset=(qid-1)*roundup(SQ_SIZE(nvmeq->q_depth),-dev->ctrl.page_size);-nvmeq->sq_dma_addr=dev->cmb_bus_addr+offset;-nvmeq->sq_cmds_io=dev->cmb+offset;-}-/**Aqueue'svectormatchesthequeueidentifierunlessthecontroller*hasonlyonevectoravailable.
QUEUE_FLAG_PCI_P2P is introduced meaning a driver's request queue
supports targeting P2P memory.
When a request is submitted we check if PCI P2PDMA memory is assigned
to the first page in the bio. If it is, we ensure the queue it's
submitted to supports it, and enforce REQ_NOMERGE.
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
---
block/blk-core.c | 14 ++++++++++++++
include/linux/blkdev.h | 3 +++
2 files changed, 17 insertions(+)
For P2P requests, we must use the pci_p2pmem_map_sg() function
instead of the dma_map_sg functions.
With that, we can then indicate PCI_P2P support in the request queue.
For this, we create an NVME_F_PCI_P2P flag which tells the core to
set QUEUE_FLAG_PCI_P2P in the request queue.
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Reviewed-by: Christoph Hellwig <hch@lst.de>
---
drivers/nvme/host/core.c | 4 ++++
drivers/nvme/host/nvme.h | 1 +
drivers/nvme/host/pci.c | 17 +++++++++++++----
3 files changed, 18 insertions(+), 4 deletions(-)
@@ -780,7 +785,10 @@ static void nvme_unmap_data(struct nvme_dev *dev, struct request *req)DMA_TO_DEVICE:DMA_FROM_DEVICE;if(iod->nents){-dma_unmap_sg(dev->dev,iod->sg,iod->nents,dma_dir);+/* P2PDMA requests do not need to be unmapped */+if(!is_pci_p2pdma_page(sg_page(iod->sg)))+dma_unmap_sg(dev->dev,iod->sg,iod->nents,dma_dir);+if(blk_integrity_rq(req))dma_unmap_sg(dev->dev,&iod->meta_sg,1,dma_dir);}
@@ -2392,7 +2400,8 @@ static int nvme_pci_get_address(struct nvme_ctrl *ctrl, char *buf, int size)staticconststructnvme_ctrl_opsnvme_pci_ctrl_ops={.name="pcie",.module=THIS_MODULE,-.flags=NVME_F_METADATA_SUPPORTED,+.flags=NVME_F_METADATA_SUPPORTED|+NVME_F_PCI_P2PDMA,.reg_read32=nvme_pci_reg_read32,.reg_write32=nvme_pci_reg_write32,.reg_read64=nvme_pci_reg_read64,
--
2.11.0
_______________________________________________
Linux-nvme mailing list
Linux-nvme@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-nvme
QUEUE_FLAG_PCI_P2P is introduced meaning a driver's request queue
supports targeting P2P memory.
When a request is submitted we check if PCI P2PDMA memory is assigned
to the first page in the bio. If it is, we ensure the queue it's
submitted to supports it, and enforce REQ_NOMERGE.
I think this belongs in the caller - both the validity check, and
passing in NOMERGE for this type of request. I don't want to impose
this overhead on everything, for a pretty niche case.
--
Jens Axboe
QUEUE_FLAG_PCI_P2P is introduced meaning a driver's request queue
supports targeting P2P memory.
When a request is submitted we check if PCI P2PDMA memory is assigned
to the first page in the bio. If it is, we ensure the queue it's
submitted to supports it, and enforce REQ_NOMERGE.
I think this belongs in the caller - both the validity check, and
passing in NOMERGE for this type of request. I don't want to impose
this overhead on everything, for a pretty niche case.
Well, the point was to prevent driver writers from doing the wrong
thing. The WARN_ON would be a bit pointless in the driver if we rely on
the driver to either do the right thing or add the WARN_ON themselves.
If I'm going to change anything I'd drop the warning entirely and move
the NO_MERGE back into the caller...
Note: the check will be compiled out if the kernel does not support PCI P2P.
Logan
QUEUE_FLAG_PCI_P2P is introduced meaning a driver's request queue
supports targeting P2P memory.
When a request is submitted we check if PCI P2PDMA memory is assigned
to the first page in the bio. If it is, we ensure the queue it's
submitted to supports it, and enforce REQ_NOMERGE.
I think this belongs in the caller - both the validity check, and
passing in NOMERGE for this type of request. I don't want to impose
this overhead on everything, for a pretty niche case.
Well, the point was to prevent driver writers from doing the wrong
thing. The WARN_ON would be a bit pointless in the driver if we rely on
the driver to either do the right thing or add the WARN_ON themselves.
If I'm going to change anything I'd drop the warning entirely and move
the NO_MERGE back into the caller...
Of course, if you move it into the caller, the warning makes no sense.
Note: the check will be compiled out if the kernel does not support PCI P2P.
Sure, but then distros tend to enable everything...
--
Jens Axboe
On Thu, Aug 30, 2018 at 12:53:39PM -0600, Logan Gunthorpe wrote:
[...]
When the PCI P2PDMA config option is selected the ACS bits in every
bridge port in the system are turned off to allow traffic to
pass freely behind the root port. At this time, the bit must be disabled
at boot so the IOMMU subsystem can correctly create the groups, though
this could be addressed in the future. There is no way to dynamically
disable the bit and alter the groups.
Can you provide an example on how to test this ? Like kernel command
line option, the doc patch does not have any such example. It would be
nice to add.
Maybe i have miss it in some of the patch. I just skimmed over for
now.
Cheers,
J�r�me
On Thu, Aug 30, 2018 at 12:53:39PM -0600, Logan Gunthorpe wrote:
[...]
quoted
When the PCI P2PDMA config option is selected the ACS bits in every
bridge port in the system are turned off to allow traffic to
pass freely behind the root port. At this time, the bit must be disabled
at boot so the IOMMU subsystem can correctly create the groups, though
this could be addressed in the future. There is no way to dynamically
disable the bit and alter the groups.
Oh, sorry this paragraph in the cover letter is wrong now. We now rely
on the disable_acs_redir command line option introduced in
aaca43fda742 ("PCI: Add "pci=disable_acs_redir=" parameter for
peer-to-peer support")
Can you provide an example on how to test this ? Like kernel command
line option, the doc patch does not have any such example. It would be
nice to add.
Do you mean to test the patchset or the ACS bits you quoted?
Testing the patchset is a matter of having the right hardware (ie an
RDMA NIC and CMB enabled NVMe behind a PCIe switch, with the ACS bits
set correctly by the above command line option) and setting the p2pmem
configfs attribute in an nvme-of port to 'yes'.
Logan
@@ -0,0 +1,170 @@+.. SPDX-License-Identifier: GPL-2.0++============================+PCI Peer-to-Peer DMA Support+============================++The PCI bus has pretty decent support for performing DMA transfers+between two devices on the bus. This type of transaction is henceforth+called Peer-to-Peer (or P2P). However, there are a number of issues that+make P2P transactions tricky to do in a perfectly safe way.++One of the biggest issues is that PCI doesn't require forwarding+transactions between hierarchy domains, and in PCIe, each Root Port+defines a separate hierarchy domain. To make things worse, there is no+simple way to determine if a given Root Complex supports this or not.+(See PCIe r4.0, sec 1.3.1). Therefore, as of this writing, the kernel+only supports doing P2P when the endpoints involved are all behind the+same PCI bridge, as such devices are all in the same PCI hierarchy+domain, and the spec guarantees that all transacations within the
transactions
+hierarchy will be routable, but it does not require routing
+between hierarchies.
+
+The second issue is that to make use of existing interfaces in Linux,
+memory that is used for P2P transactions needs to be backed by struct
+pages. However, PCI BARs are not typically cache coherent so there are
+a few corner case gotchas with these pages so developers need to
+be careful about what they do with them.
+
+
+Driver Writer's Guide
+=====================
+
+In a given P2P implementation there may be three or more different
+types of kernel drivers in play:
+
+* Provider - A driver which provides or publishes P2P resources like
+ memory or doorbell registers to other drivers.
+* Client - A driver which makes use of a resource by setting up a
+ DMA transaction to or from it.
+* Orchestrator - A driver which orchestrates the flow of data between
+ clients and providers
Might as well end that last one with a period since the other 2 are.
+
+In many cases there could be overlap between these three types (i.e.,
+it may be typical for a driver to be both a provider and a client).
+
[snip]
+
+Orchestrator Drivers
+--------------------
+
+The first task an orchestrator driver must do is compile a list of
+all client devices that will be involved in a given transaction. For
+example, the NVMe Target driver creates a list including all NVMe
+devices and the RNIC in use. The list is stored as an anonymous struct
+list_head which must be initialized with the usual INIT_LIST_HEAD.
+The following functions may then be used to add to, remove from and free
+the list of clients with the functions :c:func:`pci_p2pdma_add_client()`,
+:c:func:`pci_p2pdma_remove_client()` and
+:c:func:`pci_p2pdma_client_list_free()`.
+
+With the client list in hand, the orchestrator may then call> +:c:func:`pci_p2pmem_find()` to obtain a published P2P memory provider
+that is supported (behind the same root port) as all the clients. If more
+than one provider is supported, the one nearest to all the clients will
+be chosen first. If there are more than one provider is an equal distance
+away, the one returned will be chosen at random. This function returns the PCI
random or just arbitrarily?
+device to use for the provider with a reference taken and therefore
+when it's no longer needed it should be returned with pci_dev_put().
From: Christian König <christian.koenig@amd.com> Date: 2018-08-31 08:04:53
Am 30.08.2018 um 20:53 schrieb Logan Gunthorpe:
Some PCI devices may have memory mapped in a BAR space that's
intended for use in peer-to-peer transactions. In order to enable
such transactions the memory must be registered with ZONE_DEVICE pages
so it can be used by DMA interfaces in existing drivers.
We want to use that feature without ZONE_DEVICE pages for DMA-buf as well.
How hard would it be to separate enabling P2P detection (e.g. distance
between two devices) from this?
Regards,
Christian.
quoted hunk
Add an interface for other subsystems to find and allocate chunks of P2P
memory as necessary to facilitate transfers between two PCI peers:
int pci_p2pdma_add_client();
struct pci_dev *pci_p2pmem_find();
void *pci_alloc_p2pmem();
The new interface requires a driver to collect a list of client devices
involved in the transaction with the pci_p2pmem_add_client*() functions
then call pci_p2pmem_find() to obtain any suitable P2P memory. Once
this is done the list is bound to the memory and the calling driver is
free to add and remove clients as necessary (adding incompatible clients
will fail). With a suitable p2pmem device, memory can then be
allocated with pci_alloc_p2pmem() for use in DMA transactions.
Depending on hardware, using peer-to-peer memory may reduce the bandwidth
of the transfer but can significantly reduce pressure on system memory.
This may be desirable in many cases: for example a system could be designed
with a small CPU connected to a PCIe switch by a small number of lanes
which would maximize the number of lanes available to connect to NVMe
devices.
The code is designed to only utilize the p2pmem device if all the devices
involved in a transfer are behind the same PCI bridge. This is because we
have no way of knowing whether peer-to-peer routing between PCIe Root Ports
is supported (PCIe r4.0, sec 1.3.1). Additionally, the benefits of P2P
transfers that go through the RC is limited to only reducing DRAM usage
and, in some cases, coding convenience. The PCI-SIG may be exploring
adding a new capability bit to advertise whether this is possible for
future hardware.
This commit includes significant rework and feedback from Christoph
Hellwig.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
---
drivers/pci/Kconfig | 17 +
drivers/pci/Makefile | 1 +
drivers/pci/p2pdma.c | 761 +++++++++++++++++++++++++++++++++++++++++++++
include/linux/memremap.h | 5 +
include/linux/mm.h | 18 ++
include/linux/pci-p2pdma.h | 102 ++++++
include/linux/pci.h | 4 +
7 files changed, 908 insertions(+)
create mode 100644 drivers/pci/p2pdma.c
create mode 100644 include/linux/pci-p2pdma.h
@@ -26,6 +26,7 @@ obj-$(CONFIG_PCI_SYSCALL) += syscall.oobj-$(CONFIG_PCI_STUB)+=pci-stub.oobj-$(CONFIG_PCI_PF_STUB)+=pci-pf-stub.oobj-$(CONFIG_PCI_ECAM)+=ecam.o+obj-$(CONFIG_PCI_P2PDMA)+=p2pdma.oobj-$(CONFIG_XEN_PCIDEV_FRONTEND)+=xen-pcifront.o # Endpoint library must be initialized before its users
@@ -0,0 +1,761 @@+// SPDX-License-Identifier: GPL-2.0+/*+*PCIPeer2PeerDMAsupport.+*+*Copyright(c)2016-2018,LoganGunthorpe+*Copyright(c)2016-2017,MicrosemiCorporation+*Copyright(c)2017,ChristophHellwig+*Copyright(c)2018,EideticomInc.+*/++#define pr_fmt(fmt) "pci-p2pdma: " fmt+#include<linux/pci-p2pdma.h>+#include<linux/module.h>+#include<linux/slab.h>+#include<linux/genalloc.h>+#include<linux/memremap.h>+#include<linux/percpu-refcount.h>+#include<linux/random.h>+#include<linux/seq_buf.h>++structpci_p2pdma{+structpercpu_refdevmap_ref;+structcompletiondevmap_ref_done;+structgen_pool*pool;+boolp2pmem_published;+};++staticvoidpci_p2pdma_percpu_release(structpercpu_ref*ref)+{+structpci_p2pdma*p2p=+container_of(ref,structpci_p2pdma,devmap_ref);++complete_all(&p2p->devmap_ref_done);+}++staticvoidpci_p2pdma_percpu_kill(void*data)+{+structpercpu_ref*ref=data;++if(percpu_ref_is_dying(ref))+return;++percpu_ref_kill(ref);+}++staticvoidpci_p2pdma_release(void*data)+{+structpci_dev*pdev=data;++if(!pdev->p2pdma)+return;++wait_for_completion(&pdev->p2pdma->devmap_ref_done);+percpu_ref_exit(&pdev->p2pdma->devmap_ref);++gen_pool_destroy(pdev->p2pdma->pool);+pdev->p2pdma=NULL;+}++staticintpci_p2pdma_setup(structpci_dev*pdev)+{+interror=-ENOMEM;+structpci_p2pdma*p2p;++p2p=devm_kzalloc(&pdev->dev,sizeof(*p2p),GFP_KERNEL);+if(!p2p)+return-ENOMEM;++p2p->pool=gen_pool_create(PAGE_SHIFT,dev_to_node(&pdev->dev));+if(!p2p->pool)+gotoout;++init_completion(&p2p->devmap_ref_done);+error=percpu_ref_init(&p2p->devmap_ref,+pci_p2pdma_percpu_release,0,GFP_KERNEL);+if(error)+gotoout_pool_destroy;++percpu_ref_switch_to_atomic_sync(&p2p->devmap_ref);++error=devm_add_action_or_reset(&pdev->dev,pci_p2pdma_release,pdev);+if(error)+gotoout_pool_destroy;++pdev->p2pdma=p2p;++return0;++out_pool_destroy:+gen_pool_destroy(p2p->pool);+out:+devm_kfree(&pdev->dev,p2p);+returnerror;+}++/**+*pci_p2pdma_add_resource-addmemoryforuseasp2pmemory+*@pdev:thedevicetoaddthememoryto+*@bar:PCIBARtoadd+*@size:sizeofthememorytoadd,maybezerotousethewholeBAR+*@offset:offsetintothePCIBAR+*+*ThememorywillbegivenZONE_DEVICEstructpagessothatitmay+*beusedwithanyDMArequest.+*/+intpci_p2pdma_add_resource(structpci_dev*pdev,intbar,size_tsize,+u64offset)+{+structdev_pagemap*pgmap;+void*addr;+interror;++if(!(pci_resource_flags(pdev,bar)&IORESOURCE_MEM))+return-EINVAL;++if(offset>=pci_resource_len(pdev,bar))+return-EINVAL;++if(!size)+size=pci_resource_len(pdev,bar)-offset;++if(size+offset>pci_resource_len(pdev,bar))+return-EINVAL;++if(!pdev->p2pdma){+error=pci_p2pdma_setup(pdev);+if(error)+returnerror;+}++pgmap=devm_kzalloc(&pdev->dev,sizeof(*pgmap),GFP_KERNEL);+if(!pgmap)+return-ENOMEM;++pgmap->res.start=pci_resource_start(pdev,bar)+offset;+pgmap->res.end=pgmap->res.start+size-1;+pgmap->res.flags=pci_resource_flags(pdev,bar);+pgmap->ref=&pdev->p2pdma->devmap_ref;+pgmap->type=MEMORY_DEVICE_PCI_P2PDMA;++addr=devm_memremap_pages(&pdev->dev,pgmap);+if(IS_ERR(addr)){+error=PTR_ERR(addr);+gotopgmap_free;+}++error=gen_pool_add_virt(pdev->p2pdma->pool,(unsignedlong)addr,+pci_bus_address(pdev,bar)+offset,+resource_size(&pgmap->res),dev_to_node(&pdev->dev));+if(error)+gotopgmap_free;++error=devm_add_action_or_reset(&pdev->dev,pci_p2pdma_percpu_kill,+&pdev->p2pdma->devmap_ref);+if(error)+gotopgmap_free;++pci_info(pdev,"added peer-to-peer DMA memory %pR\n",+&pgmap->res);++return0;++pgmap_free:+devres_free(pgmap);+returnerror;+}+EXPORT_SYMBOL_GPL(pci_p2pdma_add_resource);++staticstructpci_dev*find_parent_pci_dev(structdevice*dev)+{+structdevice*parent;++dev=get_device(dev);++while(dev){+if(dev_is_pci(dev))+returnto_pci_dev(dev);++parent=get_device(dev->parent);+put_device(dev);+dev=parent;+}++returnNULL;+}++/*+*CheckifaPCIbridgehasit'sACSredirectionbitssettoredirectP2P+*TLPsupstreamviaACS.Returns1ifthepacketswillberedirected+*upstream,0otherwise.+*/+staticintpci_bridge_has_acs_redir(structpci_dev*dev)+{+intpos;+u16ctrl;++pos=pci_find_ext_capability(dev,PCI_EXT_CAP_ID_ACS);+if(!pos)+return0;++pci_read_config_word(dev,pos+PCI_ACS_CTRL,&ctrl);++if(ctrl&(PCI_ACS_RR|PCI_ACS_CR|PCI_ACS_EC))+return1;++return0;+}++staticvoidseq_buf_print_bus_devfn(structseq_buf*buf,structpci_dev*dev)+{+if(!buf)+return;++seq_buf_printf(buf,"%04x:%02x:%02x.%x;",pci_domain_nr(dev->bus),+dev->bus->number,PCI_SLOT(dev->devfn),+PCI_FUNC(dev->devfn));+}++/*+*Findthedistancethroughthenearestcommonupstreambridgebetween+*twoPCIdevices.+*+*Ifthetwodevicesarethesamedevicethen0willbereturned.+*+*Iftherearetwovirtualfunctionsofthesamedevicebehindthesame+*bridgeportthen2willbereturned(onestepdowntothePCIeswitch,+*thenonestepbacktothesamedevice).+*+*InthecasewheretwodevicesareconnectedtothesamePCIeswitch,the+*value4willbereturned.ThiscorrespondstothefollowingPCItree:+*+*-+RootPort+*\+SwitchUpstreamPort+*+-+SwitchDownstreamPort+*+\-DeviceA+*\-+SwitchDownstreamPort+*\-DeviceB+*+*Thedistanceis4becausewetraversefromDeviceAthroughthedownstream+*portoftheswitch,tothecommonupstreamport,backuptothesecond+*downstreamportandthentoDeviceB.+*+*Anytwodevicesthatdon'thaveacommonupstreambridgewillreturn-1.+*InthiswaydevicesonseparatePCIerootportswillberejected,which+*iswhatwewantforpeer-to-peerseeingeachPCIerootportdefinesa+*separatehierarchydomainandthere'snowaytodeterminewhethertheroot+*complexsupportsforwardingbetweenthem.+*+*InthecasewheretwodevicesareconnectedtodifferentPCIeswitches,+*thisfunctionwillstillreturnapositivedistanceaslongasboth+*switchesevenutallyhaveacommonupstreambridge.Notethiscovers+*thecaseofusingmultiplePCIeswitchestoachieveadesiredlevelof+*fan-outfromarootport.Theexactdistancewillbeafunctionofthe+*numberofswitchesbetweenDeviceAandDeviceB.+*+*IfabridgewhichhasanyACSredirectionbitssetisinthepath+*thenthisfunctionswillreturn-2.Thisissowerejectany+*caseswheretheTLPsareforwardedupintotherootcomplex.+*Inthiscase,alistofallinfringingbridgeaddresseswillbe+*populatedinacs_list(assumingit'snon-null)forprintkpurposes.+*/+staticintupstream_bridge_distance(structpci_dev*a,+structpci_dev*b,+structseq_buf*acs_list)+{+intdist_a=0;+intdist_b=0;+structpci_dev*bb=NULL;+intacs_cnt=0;++/*+*Note,wedon'tneedtotakereferencestodevicesreturnedby+*pci_upstream_bridge()seeingweholdareferencetoachild+*devicewhichwillalreadyholdareferencetotheupstreambridge.+*/++while(a){+dist_b=0;++if(pci_bridge_has_acs_redir(a)){+seq_buf_print_bus_devfn(acs_list,a);+acs_cnt++;+}++bb=b;++while(bb){+if(a==bb)+gotocheck_b_path_acs;++bb=pci_upstream_bridge(bb);+dist_b++;+}++a=pci_upstream_bridge(a);+dist_a++;+}++return-1;++check_b_path_acs:+bb=b;++while(bb){+if(a==bb)+break;++if(pci_bridge_has_acs_redir(bb)){+seq_buf_print_bus_devfn(acs_list,bb);+acs_cnt++;+}++bb=pci_upstream_bridge(bb);+}++if(acs_cnt)+return-2;++returndist_a+dist_b;+}++staticintupstream_bridge_distance_warn(structpci_dev*provider,+structpci_dev*client)+{+structseq_bufacs_list;+intret;++seq_buf_init(&acs_list,kmalloc(PAGE_SIZE,GFP_KERNEL),PAGE_SIZE);++ret=upstream_bridge_distance(provider,client,&acs_list);+if(ret==-2){+pci_warn(client,"cannot be used for peer-to-peer DMA as ACS redirect is set between the client and provider\n");+/* Drop final semicolon */+acs_list.buffer[acs_list.len-1]=0;+pci_warn(client,"to disable ACS redirect for this path, add the kernel parameter: pci=disable_acs_redir=%s\n",+acs_list.buffer);++}elseif(ret<0){+pci_warn(client,"cannot be used for peer-to-peer DMA as the client and provider do not share an upstream bridge\n");+}++kfree(acs_list.buffer);++returnret;+}++structpci_p2pdma_client{+structlist_headlist;+structpci_dev*client;+structpci_dev*provider;+};++/**+*pci_p2pdma_add_client-allocateanewelementinaclientdevicelist+*@head:listheadofp2pdmaclients+*@dev:devicetoaddtothelist+*+*Thisadds@devtoalistofclientsusedbyap2pdmadevice.+*Thislistshouldbepassedtopci_p2pmem_find().Oncepci_p2pmem_find()has+*beencalledsuccessfully,thelistwillbeboundtoaspecificp2pdma+*deviceandnewclientscanonlybeaddedtothelistiftheyare+*supportedbythatp2pdmadevice.+*+*Thecallerisexpectedtohavealockwhichprotects@headasnecessary+*sothatnoneofthepci_p2pfunctionscanbecalledconcurrently+*onthatlist.+*+*Returns0iftheclientwassuccessfullyadded.+*/+intpci_p2pdma_add_client(structlist_head*head,structdevice*dev)+{+structpci_p2pdma_client*item,*new_item;+structpci_dev*provider=NULL;+structpci_dev*client;+intret;++if(IS_ENABLED(CONFIG_DMA_VIRT_OPS)&&dev->dma_ops==&dma_virt_ops){+dev_warn(dev,"cannot be used for peer-to-peer DMA because the driver makes use of dma_virt_ops\n");+return-ENODEV;+}++client=find_parent_pci_dev(dev);+if(!client){+dev_warn(dev,"cannot be used for peer-to-peer DMA as it is not a PCI device\n");+return-ENODEV;+}++item=list_first_entry_or_null(head,structpci_p2pdma_client,list);+if(item&&item->provider){+provider=item->provider;++ret=upstream_bridge_distance_warn(provider,client);+if(ret<0){+ret=-EXDEV;+gotoput_client;+}+}++new_item=kzalloc(sizeof(*new_item),GFP_KERNEL);+if(!new_item){+ret=-ENOMEM;+gotoput_client;+}++new_item->client=client;+new_item->provider=pci_dev_get(provider);++list_add_tail(&new_item->list,head);++return0;++put_client:+pci_dev_put(client);+returnret;+}+EXPORT_SYMBOL_GPL(pci_p2pdma_add_client);++staticvoidpci_p2pdma_client_free(structpci_p2pdma_client*item)+{+list_del(&item->list);+pci_dev_put(item->client);+pci_dev_put(item->provider);+kfree(item);+}++/**+*pci_p2pdma_remove_client-removeandfreeap2pdmaclient+*@head:listheadofp2pdmaclients+*@dev:devicetoremovefromthelist+*+*Thisremoves@devfromalistofclientsusedbyap2pdmadevice.+*Thecallerisexpectedtohavealockwhichprotects@headasnecessary+*sothatnoneofthepci_p2pfunctionscanbecalledconcurrently+*onthatlist.+*/+voidpci_p2pdma_remove_client(structlist_head*head,structdevice*dev)+{+structpci_p2pdma_client*pos,*tmp;+structpci_dev*pdev;++pdev=find_parent_pci_dev(dev);+if(!pdev)+return;++list_for_each_entry_safe(pos,tmp,head,list){+if(pos->client!=pdev)+continue;++pci_p2pdma_client_free(pos);+}++pci_dev_put(pdev);+}+EXPORT_SYMBOL_GPL(pci_p2pdma_remove_client);++/**+*pci_p2pdma_client_list_free-freeanentirelistofp2pdmaclients+*@head:listheadofp2pdmaclients+*+*Thisremovesalldevicesinalistofclientsusedbyap2pdmadevice.+*Thecallerisexpectedtohavealockwhichprotects@headasnecessary+*sothatnoneofthepci_p2pdmafunctionscanbecalledconcurrently+*onthatlist.+*/+voidpci_p2pdma_client_list_free(structlist_head*head)+{+structpci_p2pdma_client*pos,*tmp;++list_for_each_entry_safe(pos,tmp,head,list)+pci_p2pdma_client_free(pos);+}+EXPORT_SYMBOL_GPL(pci_p2pdma_client_list_free);++/**+*pci_p2pdma_distance-Determivethecumulativedistancebetween+*ap2pdmaproviderandtheclientsinuse.+*@provider:p2pdmaprovidertocheckagainsttheclientlist+*@clients:listofdevicestocheck(NULL-terminated)+*@verbose:iftrue,printwarningsfordeviceswhenwereturn-1+*+*Returns-1ifanyoftheclientsarenotcompatible(behindthesame+*rootportastheprovider),otherwisereturnsapositivenumberwhere+*thelowernumberisthepreferrablechoice.(Ifthere'soneclient+*that'sthesameastheprovideritwillreturn0,whichisbestchoice).+*+*Fornow,"compatible"meanstheproviderandtheclientsareallbehind+*thesamePCIrootport.Thiscutsoutcasesthatmayworkbutissafest+*fortheuser.Futureworkcanexpandthistowhite-listrootcomplexesthat+*cansafelyforwardbetweeneachports.+*/+intpci_p2pdma_distance(structpci_dev*provider,structlist_head*clients,+boolverbose)+{+structpci_p2pdma_client*pos;+intret;+intdistance=0;+boolnot_supported=false;++if(list_empty(clients))+return-1;++list_for_each_entry(pos,clients,list){+if(verbose)+ret=upstream_bridge_distance_warn(provider,+pos->client);+else+ret=upstream_bridge_distance(provider,pos->client,+NULL);++if(ret<0)+not_supported=true;++if(not_supported&&!verbose)+break;++distance+=ret;+}++if(not_supported)+return-1;++returndistance;+}+EXPORT_SYMBOL_GPL(pci_p2pdma_distance);++/**+*pci_p2pdma_assign_provider-Checkcompatibily(asperpci_p2pdma_distance)+*andassignaprovidertoalistofclients+*@provider:p2pdmaprovidertoassigntotheclientlist+*@clients:listofdevicestocheck(NULL-terminated)+*+*Returnsfalseifanyoftheclientsarenotcompatible,trueifthe+*providerwassuccessfullyassignedtotheclients.+*/+boolpci_p2pdma_assign_provider(structpci_dev*provider,+structlist_head*clients)+{+structpci_p2pdma_client*pos;++if(pci_p2pdma_distance(provider,clients,true)<0)+returnfalse;++list_for_each_entry(pos,clients,list)+pos->provider=provider;++returntrue;+}+EXPORT_SYMBOL_GPL(pci_p2pdma_assign_provider);++/**+*pci_has_p2pmem-checkifagivenPCIdevicehaspublishedanyp2pmem+*@pdev:PCIdevicetocheck+*/+boolpci_has_p2pmem(structpci_dev*pdev)+{+returnpdev->p2pdma&&pdev->p2pdma->p2pmem_published;+}+EXPORT_SYMBOL_GPL(pci_has_p2pmem);++/**+*pci_p2pmem_find-findapeer-to-peerDMAmemorydevicecompatiblewith+*thespecifiedlistofclientsandshortestdistance(asdetermined+*bypci_p2pmem_dma())+*@clients:listofdevicestocheck(NULL-terminated)+*+*Ifmultipledevicesarebehindthesameswitch,theone"closest"tothe+*clientdevicesinusewillbechosenfirst.(Soifoneoftheprovidersare+*thesameasoneoftheclients,thatproviderwillbeusedaheadofany+*otherprovidersthatareunrelated).Ifmultipleprovidersareanequal+*distanceaway,onewillbechosenatrandom.+*+*ReturnsapointertothePCIdevicewithareferencetaken(usepci_dev_put+*toreturnthereference)orNULLifnocompatibledeviceisfound.The+*foundproviderwillalsobeassignedtotheclientlist.+*/+structpci_dev*pci_p2pmem_find(structlist_head*clients)+{+structpci_dev*pdev=NULL;+structpci_p2pdma_client*pos;+intdistance;+intclosest_distance=INT_MAX;+structpci_dev**closest_pdevs;+intdev_cnt=0;+constintmax_devs=PAGE_SIZE/sizeof(*closest_pdevs);+inti;++closest_pdevs=kmalloc(PAGE_SIZE,GFP_KERNEL);++while((pdev=pci_get_device(PCI_ANY_ID,PCI_ANY_ID,pdev))){+if(!pci_has_p2pmem(pdev))+continue;++distance=pci_p2pdma_distance(pdev,clients,false);+if(distance<0||distance>closest_distance)+continue;++if(distance==closest_distance&&dev_cnt>=max_devs)+continue;++if(distance<closest_distance){+for(i=0;i<dev_cnt;i++)+pci_dev_put(closest_pdevs[i]);++dev_cnt=0;+closest_distance=distance;+}++closest_pdevs[dev_cnt++]=pci_dev_get(pdev);+}++if(dev_cnt)+pdev=pci_dev_get(closest_pdevs[prandom_u32_max(dev_cnt)]);++for(i=0;i<dev_cnt;i++)+pci_dev_put(closest_pdevs[i]);++if(pdev)+list_for_each_entry(pos,clients,list)+pos->provider=pdev;++kfree(closest_pdevs);+returnpdev;+}+EXPORT_SYMBOL_GPL(pci_p2pmem_find);++/**+*pci_alloc_p2p_mem-allocatepeer-to-peerDMAmemory+*@pdev:thedevicetoallocatememoryfrom+*@size:numberofbytestoallocate+*+*ReturnstheallocatedmemoryorNULLonerror.+*/+void*pci_alloc_p2pmem(structpci_dev*pdev,size_tsize)+{+void*ret;++if(unlikely(!pdev->p2pdma))+returnNULL;++if(unlikely(!percpu_ref_tryget_live(&pdev->p2pdma->devmap_ref)))+returnNULL;++ret=(void*)gen_pool_alloc(pdev->p2pdma->pool,size);++if(unlikely(!ret))+percpu_ref_put(&pdev->p2pdma->devmap_ref);++returnret;+}+EXPORT_SYMBOL_GPL(pci_alloc_p2pmem);++/**+*pci_free_p2pmem-allocatepeer-to-peerDMAmemory+*@pdev:thedevicethememorywasallocatedfrom+*@addr:addressofthememorythatwasallocated+*@size:numberofbytesthatwasallocated+*/+voidpci_free_p2pmem(structpci_dev*pdev,void*addr,size_tsize)+{+gen_pool_free(pdev->p2pdma->pool,(uintptr_t)addr,size);+percpu_ref_put(&pdev->p2pdma->devmap_ref);+}+EXPORT_SYMBOL_GPL(pci_free_p2pmem);++/**+*pci_virt_to_bus-returnthePCIbusaddressforagivenvirtual+*addressobtainedwithpci_alloc_p2pmem()+*@pdev:thedevicethememorywasallocatedfrom+*@addr:addressofthememorythatwasallocated+*/+pci_bus_addr_tpci_p2pmem_virt_to_bus(structpci_dev*pdev,void*addr)+{+if(!addr)+return0;+if(!pdev->p2pdma)+return0;++/*+*Note:whenweaddedthememorytothepoolweusedthePCI+*busaddressasthephysicaladdress.Sogen_pool_virt_to_phys()+*actuallyreturnsthebusaddressdespitethemisleadingname.+*/+returngen_pool_virt_to_phys(pdev->p2pdma->pool,(unsignedlong)addr);+}+EXPORT_SYMBOL_GPL(pci_p2pmem_virt_to_bus);++/**+*pci_p2pmem_alloc_sgl-allocatepeer-to-peerDMAmemoryinascatterlist+*@pdev:thedevicetoallocatememoryfrom+*@sgl:theallocatedscatterlist+*@nents:thenumberofSGentriesinthelist+*@length:numberofbytestoallocate+*+*Returns0onsuccess+*/+structscatterlist*pci_p2pmem_alloc_sgl(structpci_dev*pdev,+unsignedint*nents,u32length)+{+structscatterlist*sg;+void*addr;++sg=kzalloc(sizeof(*sg),GFP_KERNEL);+if(!sg)+returnNULL;++sg_init_table(sg,1);++addr=pci_alloc_p2pmem(pdev,length);+if(!addr)+gotoout_free_sg;++sg_set_buf(sg,addr,length);+*nents=1;+returnsg;++out_free_sg:+kfree(sg);+returnNULL;+}+EXPORT_SYMBOL_GPL(pci_p2pmem_alloc_sgl);++/**+*pci_p2pmem_free_sgl-freeascatterlistallocatedbypci_p2pmem_alloc_sgl()+*@pdev:thedevicetoallocatememoryfrom+*@sgl:theallocatedscatterlist+*@nents:thenumberofSGentriesinthelist+*/+voidpci_p2pmem_free_sgl(structpci_dev*pdev,structscatterlist*sgl)+{+structscatterlist*sg;+intcount;++for_each_sg(sgl,sg,INT_MAX,count){+if(!sg)+break;++pci_free_p2pmem(pdev,sg_virt(sg),sg->length);+}+kfree(sgl);+}+EXPORT_SYMBOL_GPL(pci_p2pmem_free_sgl);++/**+*pci_p2pmem_publish-publishthepeer-to-peerDMAmemoryforuseby+*otherdeviceswithpci_p2pmem_find()+*@pdev:thedevicewithpeer-to-peerDMAmemorytopublish+*@publish:settotruetopublishthememory,falsetounpublishit+*+*PublishedmemorycanbeusedbyotherPCIdevicedriversfor+*peer-2-peerDMAoperations.Non-publishedmemoryisreservedfor+*exlusiveuseofthedevicedriverthatregistersthepeer-to-peer+*memory.+*/+voidpci_p2pmem_publish(structpci_dev*pdev,boolpublish)+{+if(publish&&!pdev->p2pdma)+return;++pdev->p2pdma->p2pmem_published=publish;+}+EXPORT_SYMBOL_GPL(pci_p2pmem_publish);
@@ -439,6 +440,9 @@ struct pci_dev {#ifdef CONFIG_PCI_PASIDu16pasid_features;#endif+#ifdef CONFIG_PCI_P2PDMA+structpci_p2pdma*p2pdma;+#endifphys_addr_trom;/* Physical address if not from BAR */size_tromlen;/* Length if not from BAR */char*driver_override;/* Driver name to force a match */
From: Christian König <christian.koenig@amd.com> Date: 2018-08-31 08:12:30
Am 30.08.2018 um 20:53 schrieb Logan Gunthorpe:
[SNIP]
+============================
+PCI Peer-to-Peer DMA Support
+============================
+
+The PCI bus has pretty decent support for performing DMA transfers
+between two devices on the bus. This type of transaction is henceforth
+called Peer-to-Peer (or P2P). However, there are a number of issues that
+make P2P transactions tricky to do in a perfectly safe way.
+
+One of the biggest issues is that PCI doesn't require forwarding
+transactions between hierarchy domains, and in PCIe, each Root Port
+defines a separate hierarchy domain. To make things worse, there is no
+simple way to determine if a given Root Complex supports this or not.
+(See PCIe r4.0, sec 1.3.1). Therefore, as of this writing, the kernel
+only supports doing P2P when the endpoints involved are all behind the
+same PCI bridge, as such devices are all in the same PCI hierarchy
+domain, and the spec guarantees that all transacations within the
+hierarchy will be routable, but it does not require routing
+between hierarchies.
Can we add a kernel command line switch and a whitelist to enable P2P
between separate hierarchies?
At least all newer AMD chipsets supports this and I'm pretty sure that
Intel has a list with PCI-IDs of the root hubs for this as well.
Regards,
Christian.
_______________________________________________
Linux-nvme mailing list
Linux-nvme@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-nvme
Thanks, for the review.
On 30/08/18 06:25 PM, Sagi Grimberg wrote:
quoted
+ if (req->port->p2p_dev) {
+ if (!pci_p2pdma_assign_provider(req->port->p2p_dev,
+ &ctrl->p2p_clients)) {
+ pr_info("peer-to-peer memory on %s is not supported\n",
+ pci_name(req->port->p2p_dev));
+ goto free_devices;
+ }
+ ctrl->p2p_dev = pci_dev_get(req->port->p2p_dev);
+ } else {
When is port->p2p_dev == NULL? a little more documentation would help
here...
In the configfs functions, if the user enables p2p (port->use_p2pmem)
using 'auto' or 'y' then port->p2p_dev will be NULL. If the user sets a
specific p2p_dev to use, port->p2p_dev will be set to that device. I can
add a couple comments in the next version.
Logan
Hey,
Thanks for the review. I'll make the fixes for the next version.
On 30/08/18 06:34 PM, Randy Dunlap wrote:
quoted
+With the client list in hand, the orchestrator may then call> +:c:func:`pci_p2pmem_find()` to obtain a published P2P memory provider
+that is supported (behind the same root port) as all the clients. If more
+than one provider is supported, the one nearest to all the clients will
+be chosen first. If there are more than one provider is an equal distance
+away, the one returned will be chosen at random. This function returns the PCI
random or just arbitrarily?
Randomly. See pci_p2pmem_find() in patch 1. We use prandom_u32_max() to
select any of the supported devices.
Logan
Some PCI devices may have memory mapped in a BAR space that's
intended for use in peer-to-peer transactions. In order to enable
such transactions the memory must be registered with ZONE_DEVICE pages
so it can be used by DMA interfaces in existing drivers.
We want to use that feature without ZONE_DEVICE pages for DMA-buf as well.
How hard would it be to separate enabling P2P detection (e.g. distance
between two devices) from this?
Pretty easy. P2P detection is pretty much just pci_p2pdma_distance() ,
which has nothing to do with the ZONE_DEVICE support.
(And the distance function makes use of a number of static functions
which could be combined into a simpler interface, should we need it.)
Logan
+One of the biggest issues is that PCI doesn't require forwarding
+transactions between hierarchy domains, and in PCIe, each Root Port
+defines a separate hierarchy domain. To make things worse, there is no
+simple way to determine if a given Root Complex supports this or not.
+(See PCIe r4.0, sec 1.3.1). Therefore, as of this writing, the kernel
+only supports doing P2P when the endpoints involved are all behind the
+same PCI bridge, as such devices are all in the same PCI hierarchy
+domain, and the spec guarantees that all transacations within the
+hierarchy will be routable, but it does not require routing
+between hierarchies.
Can we add a kernel command line switch and a whitelist to enable P2P
between separate hierarchies?
In future work, yes. But not for this patchset. This is definitely the
way I see things going, but we've chosen to start with what we've presented.
Logan
From: Jonathan Cameron <hidden> Date: 2018-08-31 16:19:30
On Thu, 30 Aug 2018 12:53:40 -0600
Logan Gunthorpe [off-list ref] wrote:
Some PCI devices may have memory mapped in a BAR space that's
intended for use in peer-to-peer transactions. In order to enable
such transactions the memory must be registered with ZONE_DEVICE pages
so it can be used by DMA interfaces in existing drivers.
Add an interface for other subsystems to find and allocate chunks of P2P
memory as necessary to facilitate transfers between two PCI peers:
int pci_p2pdma_add_client();
struct pci_dev *pci_p2pmem_find();
void *pci_alloc_p2pmem();
The new interface requires a driver to collect a list of client devices
involved in the transaction with the pci_p2pmem_add_client*() functions
then call pci_p2pmem_find() to obtain any suitable P2P memory. Once
this is done the list is bound to the memory and the calling driver is
free to add and remove clients as necessary (adding incompatible clients
will fail). With a suitable p2pmem device, memory can then be
allocated with pci_alloc_p2pmem() for use in DMA transactions.
Depending on hardware, using peer-to-peer memory may reduce the bandwidth
of the transfer but can significantly reduce pressure on system memory.
This may be desirable in many cases: for example a system could be designed
with a small CPU connected to a PCIe switch by a small number of lanes
which would maximize the number of lanes available to connect to NVMe
devices.
The code is designed to only utilize the p2pmem device if all the devices
involved in a transfer are behind the same PCI bridge. This is because we
have no way of knowing whether peer-to-peer routing between PCIe Root Ports
is supported (PCIe r4.0, sec 1.3.1). Additionally, the benefits of P2P
transfers that go through the RC is limited to only reducing DRAM usage
and, in some cases, coding convenience. The PCI-SIG may be exploring
adding a new capability bit to advertise whether this is possible for
future hardware.
This commit includes significant rework and feedback from Christoph
Hellwig.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
Apologies for being a late entrant to this conversation so I may be asking
about a topic that has been covered in detail in earlier patches!
---
...
+/*
+ * Find the distance through the nearest common upstream bridge between
+ * two PCI devices.
+ *
+ * If the two devices are the same device then 0 will be returned.
+ *
+ * If there are two virtual functions of the same device behind the same
+ * bridge port then 2 will be returned (one step down to the PCIe switch,
+ * then one step back to the same device).
+ *
+ * In the case where two devices are connected to the same PCIe switch, the
+ * value 4 will be returned. This corresponds to the following PCI tree:
+ *
+ * -+ Root Port
+ * \+ Switch Upstream Port
+ * +-+ Switch Downstream Port
+ * + \- Device A
+ * \-+ Switch Downstream Port
+ * \- Device B
+ *
+ * The distance is 4 because we traverse from Device A through the downstream
+ * port of the switch, to the common upstream port, back up to the second
+ * downstream port and then to Device B.
+ *
+ * Any two devices that don't have a common upstream bridge will return -1.
+ * In this way devices on separate PCIe root ports will be rejected, which
+ * is what we want for peer-to-peer seeing each PCIe root port defines a
+ * separate hierarchy domain and there's no way to determine whether the root
+ * complex supports forwarding between them.
+ *
+ * In the case where two devices are connected to different PCIe switches,
+ * this function will still return a positive distance as long as both
+ * switches evenutally have a common upstream bridge. Note this covers
+ * the case of using multiple PCIe switches to achieve a desired level of
+ * fan-out from a root port. The exact distance will be a function of the
+ * number of switches between Device A and Device B.
This feels like a somewhat simplistic starting point rather than a
generally correct estimate to use. Should we be taking the bandwidth of
those links into account for example, or any discoverable latencies?
Not all PCIe switches are alike - particularly when it comes to P2P.
I guess that can be a topic for future development if it turns out people
have horrible mixed systems.
+ *
+ * If a bridge which has any ACS redirection bits set is in the path
+ * then this functions will return -2. This is so we reject any
+ * cases where the TLPs are forwarded up into the root complex.
+ * In this case, a list of all infringing bridge addresses will be
+ * populated in acs_list (assuming it's non-null) for printk purposes.
+ */
This feels like a somewhat simplistic starting point rather than a
generally correct estimate to use. Should we be taking the bandwidth of
those links into account for example, or any discoverable latencies?
Not all PCIe switches are alike - particularly when it comes to P2P.
I don't think this is necessary. There won't typically be a ton of
choice in terms of devices to use and if there is, the hardware will
probably be fairly homogenous. For example, it would be unusual to have
an NVMe drive on a x4 and another one on an x8. Or mixing say Gen3
switches with Gen4 would also be very strange. In weird unusual cases
like this where the user specifically wants to use a faster device they
can specify the specific device in the configfs interface.
I think the latency would probably be proportional to the distance which
is what we are already using.
I guess that can be a topic for future development if it turns out people
have horrible mixed systems.
If you can separate out adding the detection I can take a look adding
this with my DMA-buf P2P efforts.
Oh, maybe my previous email wasn't clear, but I'd say that detection is
already separate from ZONE_DEVICE. Nothing really needs to be changed.
I just think you'll probably want to write you're own function similar
to pci_p2pdma_distance that perhaps just takes two pci_devs instead of
the list of clients as is needed by nvme-of-like users.
To enable a whitelist we just have to handle the case where
upstream_bridge_distance() returns -1 and check if the devices are in
the same root complex with supported root ports before deciding the
transaction is not supported.
Logan
From: Christoph Hellwig <hch@lst.de> Date: 2018-09-01 08:24:04
On Fri, Aug 31, 2018 at 09:48:40AM -0600, Logan Gunthorpe wrote:
Pretty easy. P2P detection is pretty much just pci_p2pdma_distance() ,
which has nothing to do with the ZONE_DEVICE support.
(And the distance function makes use of a number of static functions
which could be combined into a simpler interface, should we need it.)
I'd ѕay lets get things merged as-is, so that we can review the
non-ZONE_DEVICE users. I'm a little curious how that is going to work,
so having it as a full series would be useful.
From: Christoph Hellwig <hch@lst.de> Date: 2018-09-01 08:25:13
On Thu, Aug 30, 2018 at 01:11:18PM -0600, Jens Axboe wrote:
I think this belongs in the caller - both the validity check, and
passing in NOMERGE for this type of request. I don't want to impose
this overhead on everything, for a pretty niche case.
It is just a single branch, which will be predicted as not taken
for non-P2P users. The benefit is that we get proper error checking
by doing it in the block code.
On Thu, Aug 30, 2018 at 01:11:18PM -0600, Jens Axboe wrote:
quoted
I think this belongs in the caller - both the validity check, and
passing in NOMERGE for this type of request. I don't want to impose
this overhead on everything, for a pretty niche case.
It is just a single branch, which will be predicted as not taken
for non-P2P users. The benefit is that we get proper error checking
by doing it in the block code.
I personally agree with Christoph. But if there's consensus in the other
direction or this is a real blocker moving this forward, I can remove it
for the next version.
Logan
From: Jason Gunthorpe <hidden> Date: 2018-09-04 15:17:00
On Thu, Aug 30, 2018 at 12:53:49PM -0600, Logan Gunthorpe wrote:
quoted hunk
For P2P requests, we must use the pci_p2pmem_map_sg() function
instead of the dma_map_sg functions.
With that, we can then indicate PCI_P2P support in the request queue.
For this, we create an NVME_F_PCI_P2P flag which tells the core to
set QUEUE_FLAG_PCI_P2P in the request queue.
Signed-off-by: Logan Gunthorpe <logang@deltatee.com>
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Reviewed-by: Christoph Hellwig <hch@lst.de>
drivers/nvme/host/core.c | 4 ++++
drivers/nvme/host/nvme.h | 1 +
drivers/nvme/host/pci.c | 17 +++++++++++++----
3 files changed, 18 insertions(+), 4 deletions(-)
@@ -780,7 +785,10 @@ static void nvme_unmap_data(struct nvme_dev *dev, struct request *req)DMA_TO_DEVICE:DMA_FROM_DEVICE;if(iod->nents){-dma_unmap_sg(dev->dev,iod->sg,iod->nents,dma_dir);+/* P2PDMA requests do not need to be unmapped */+if(!is_pci_p2pdma_page(sg_page(iod->sg)))+dma_unmap_sg(dev->dev,iod->sg,iod->nents,dma_dir);
This seems like a poor direction, if we add IOMMU hairpin support we
will need unmapping.
Jason
if (iod->nents) {
- dma_unmap_sg(dev->dev, iod->sg, iod->nents, dma_dir);
+ /* P2PDMA requests do not need to be unmapped */
+ if (!is_pci_p2pdma_page(sg_page(iod->sg)))
+ dma_unmap_sg(dev->dev, iod->sg, iod->nents, dma_dir);
This seems like a poor direction, if we add IOMMU hairpin support we
will need unmapping.
It can always be added later. In any case, you'll have to convince
Christoph who requested the change; I'm not that invested in this decision.
Logan
From: Christoph Hellwig <hch@lst.de> Date: 2018-09-05 19:19:17
On Tue, Sep 04, 2018 at 09:47:07AM -0600, Logan Gunthorpe wrote:
On 04/09/18 09:16 AM, Jason Gunthorpe wrote:
quoted
quoted
if (iod->nents) {
- dma_unmap_sg(dev->dev, iod->sg, iod->nents, dma_dir);
+ /* P2PDMA requests do not need to be unmapped */
+ if (!is_pci_p2pdma_page(sg_page(iod->sg)))
+ dma_unmap_sg(dev->dev, iod->sg, iod->nents, dma_dir);
This seems like a poor direction, if we add IOMMU hairpin support we
will need unmapping.
It can always be added later. In any case, you'll have to convince
Christoph who requested the change; I'm not that invested in this decision.
Yes, no point to add dead code here. In the long run we should
aim for hiding the p2p address translation behind the normal DMA API
anyway, but we're not quite ready for it yet.
On Thu, Aug 30, 2018 at 01:11:18PM -0600, Jens Axboe wrote:
quoted
I think this belongs in the caller - both the validity check, and
passing in NOMERGE for this type of request. I don't want to impose
this overhead on everything, for a pretty niche case.
It is just a single branch, which will be predicted as not taken
for non-P2P users. The benefit is that we get proper error checking
by doing it in the block code.
I personally agree with Christoph. But if there's consensus in the other
direction or this is a real blocker moving this forward, I can remove it
for the next version.
It's a simple branch because the check isn't exhaustive. It just checks
the first page. At that point you may as well just require the caller to
flag the bio/rq as being P2P, and then do a check for P2P compatibility
with the queue.
--
Jens Axboe
I personally agree with Christoph. But if there's consensus in the other
direction or this is a real blocker moving this forward, I can remove it
for the next version.
It's a simple branch because the check isn't exhaustive. It just checks
the first page. At that point you may as well just require the caller to
flag the bio/rq as being P2P, and then do a check for P2P compatibility
with the queue.
Hmm, we had something like that in v4[1] but it just seemed redundant to
create a flag when the information was already in the bio and kind of
ugly for the caller to check for, then set, the flag. I'm not _that_
averse to going back to that though...
Logan
[1]
https://lore.kernel.org/lkml/20180423233046.21476-8-logang@deltatee.com/T/#u
I personally agree with Christoph. But if there's consensus in the other
direction or this is a real blocker moving this forward, I can remove it
for the next version.
It's a simple branch because the check isn't exhaustive. It just checks
the first page. At that point you may as well just require the caller to
flag the bio/rq as being P2P, and then do a check for P2P compatibility
with the queue.
Hmm, we had something like that in v4[1] but it just seemed redundant to
create a flag when the information was already in the bio and kind of
ugly for the caller to check for, then set, the flag. I'm not _that_
averse to going back to that though...
The point is that the caller doesn't necessarily know where the bio
will end up, hence the caller can't fully check if the whole stack
supports P2P.
What happens if a P2P request ends up with a driver that doesn't
support it?
--
Jens Axboe
From: Christoph Hellwig <hch@lst.de> Date: 2018-09-05 19:53:12
On Wed, Sep 05, 2018 at 01:45:04PM -0600, Jens Axboe wrote:
The point is that the caller doesn't necessarily know where the bio
will end up, hence the caller can't fully check if the whole stack
supports P2P.
The caller must necessarily know where the bio will end up, as for P2P
support we need to query if the bio target is P2P capable vs the
source of the P2P memory.
The point is that the caller doesn't necessarily know where the bio
will end up, hence the caller can't fully check if the whole stack
supports P2P.
What happens if a P2P request ends up with a driver that doesn't
support it?
Yes, that's the whole point this check. Although we expect the caller to
do other checks before submitting a P2P request to a queue, so if a
driver does submit a P2P request to an unsupported queue, it is
definitely a problem in the driver (which is why we want to WARN).
Queues that support P2P (only PCI NVMe at this time, see patch 10) must
set QUEUE_FLAG_PCI_P2PDMA to indicate it. The check we are adding in
blk-core is meant to ensure any broken drivers that submit requests with
P2P memory do not get sent to a queue that doesn't indicate support.
On top of that, the code in NVMe target ensures that all namespaces on a
port are backed by queues that support P2P and, if not, it never
allocates any P2P SGLs.
Logan
On Wed, Sep 05, 2018 at 01:45:04PM -0600, Jens Axboe wrote:
quoted
The point is that the caller doesn't necessarily know where the bio
will end up, hence the caller can't fully check if the whole stack
supports P2P.
The caller must necessarily know where the bio will end up, as for P2P
support we need to query if the bio target is P2P capable vs the
source of the P2P memory.
Then what's the point of having the check at all?
--
Jens Axboe
From: Christoph Hellwig <hch@lst.de> Date: 2018-09-05 20:08:16
On Wed, Sep 05, 2018 at 01:54:31PM -0600, Jens Axboe wrote:
On 9/5/18 1:56 PM, Christoph Hellwig wrote:
quoted
On Wed, Sep 05, 2018 at 01:45:04PM -0600, Jens Axboe wrote:
quoted
The point is that the caller doesn't necessarily know where the bio
will end up, hence the caller can't fully check if the whole stack
supports P2P.
The caller must necessarily know where the bio will end up, as for P2P
support we need to query if the bio target is P2P capable vs the
source of the P2P memory.
Then what's the point of having the check at all?
Just an additional little safe guard. If you think it isn't worth
it I guess we can just drop it for now.
On Wed, Sep 05, 2018 at 01:54:31PM -0600, Jens Axboe wrote:
quoted
On 9/5/18 1:56 PM, Christoph Hellwig wrote:
quoted
On Wed, Sep 05, 2018 at 01:45:04PM -0600, Jens Axboe wrote:
quoted
The point is that the caller doesn't necessarily know where the bio
will end up, hence the caller can't fully check if the whole stack
supports P2P.
The caller must necessarily know where the bio will end up, as for P2P
support we need to query if the bio target is P2P capable vs the
source of the P2P memory.
Then what's the point of having the check at all?
Just an additional little safe guard. If you think it isn't worth
it I guess we can just drop it for now.
Yes, the point is to prevent driver writers from doing the wrong thing
by not doing the necessary checks before submitting to the queue.
Logan
On Wed, Sep 05, 2018 at 01:54:31PM -0600, Jens Axboe wrote:
quoted
On 9/5/18 1:56 PM, Christoph Hellwig wrote:
quoted
On Wed, Sep 05, 2018 at 01:45:04PM -0600, Jens Axboe wrote:
quoted
The point is that the caller doesn't necessarily know where the bio
will end up, hence the caller can't fully check if the whole stack
supports P2P.
The caller must necessarily know where the bio will end up, as for P2P
support we need to query if the bio target is P2P capable vs the
source of the P2P memory.
Then what's the point of having the check at all?
Just an additional little safe guard. If you think it isn't worth
it I guess we can just drop it for now.
Yes, the point is to prevent driver writers from doing the wrong thing
by not doing the necessary checks before submitting to the queue.
But if the caller must absolutely know where the bio will end up, then
it seems super redundant. So I'd vote for killing this check, it buys
us absolutely nothing and isn't even exhaustive in its current form.
--
Jens Axboe
But if the caller must absolutely know where the bio will end up, then
it seems super redundant. So I'd vote for killing this check, it buys
us absolutely nothing and isn't even exhaustive in its current form.
But if the caller must absolutely know where the bio will end up, then
it seems super redundant. So I'd vote for killing this check, it buys
us absolutely nothing and isn't even exhaustive in its current form.
Ok, I'll remove it for v6.
Since the drivers needs to know it's doing it right, it might not
hurt to add a sanity check helper for that. Just have the driver
call it, and don't add it in the normal IO submission path.
--
Jens Axboe
But if the caller must absolutely know where the bio will end up, then
it seems super redundant. So I'd vote for killing this check, it buys
us absolutely nothing and isn't even exhaustive in its current form.
Ok, I'll remove it for v6.
Since the drivers needs to know it's doing it right, it might not
hurt to add a sanity check helper for that. Just have the driver
call it, and don't add it in the normal IO submission path.
I'm not sure I really see the value in that. It's the same principle in
asking the driver to do the WARN: if the developer knew enough to use
the special helper, they probably knew well enough to do the rest correctly.
I guess one other thing to point out is that, on x86, if a driver
submits P2P pages to a PCI device that doesn't have kernel support,
everything will likely just work. Even though the driver isn't doing any
of the work correctly and the requests are not being mapped with
pci_p2pdma_map() functions. Such code on other arches would likely
break. So developers may be lulled into thinking they're doing the
correct thing when in fact they are not and the WARN in the common code
would prevent that.
Logan
But if the caller must absolutely know where the bio will end up, then
it seems super redundant. So I'd vote for killing this check, it buys
us absolutely nothing and isn't even exhaustive in its current form.
Ok, I'll remove it for v6.
Since the drivers needs to know it's doing it right, it might not
hurt to add a sanity check helper for that. Just have the driver
call it, and don't add it in the normal IO submission path.
I'm not sure I really see the value in that. It's the same principle in
asking the driver to do the WARN: if the developer knew enough to use
the special helper, they probably knew well enough to do the rest correctly.
I don't agree with that at all. It's a "is my request valid" helper,
it's not some obscure and rarely used functionality. You're making up
this API right now, if you really want it done for every IO, make it
part of the p2p submission process. You could even hide it behind a
debug thing, if you like.
I guess one other thing to point out is that, on x86, if a driver
submits P2P pages to a PCI device that doesn't have kernel support,
everything will likely just work. Even though the driver isn't doing any
of the work correctly and the requests are not being mapped with
pci_p2pdma_map() functions. Such code on other arches would likely
break. So developers may be lulled into thinking they're doing the
correct thing when in fact they are not and the WARN in the common code
would prevent that.
If you're adamant about having it in common code, put it in your
common submission code. Most folks aren't going to care about P2P, let the
ones that do have the checks.
--
Jens Axboe
But if the caller must absolutely know where the bio will end up, then
it seems super redundant. So I'd vote for killing this check, it buys
us absolutely nothing and isn't even exhaustive in its current form.
Ok, I'll remove it for v6.
Since the drivers needs to know it's doing it right, it might not
hurt to add a sanity check helper for that. Just have the driver
call it, and don't add it in the normal IO submission path.
I'm not sure I really see the value in that. It's the same principle in
asking the driver to do the WARN: if the developer knew enough to use
the special helper, they probably knew well enough to do the rest correctly.
I don't agree with that at all. It's a "is my request valid" helper,
it's not some obscure and rarely used functionality. You're making up
this API right now, if you really want it done for every IO, make it
part of the p2p submission process. You could even hide it behind a
debug thing, if you like.
There is no special p2p submission process. In the nvme-of case we are
using the existing process and with the code in blk-core it didn't
change it's process at all. Creating a helper will create one and I can
look at making a pci_p2pdma_submit_bio() for v6; but if the developer
screws up and still calls the regular submit_bio() things will only be
very subtly broken and that won't be obvious.
Logan
From: Christoph Hellwig <hch@lst.de> Date: 2018-09-05 21:10:16
On Wed, Sep 05, 2018 at 03:03:18PM -0600, Logan Gunthorpe wrote:
There is no special p2p submission process. In the nvme-of case we are
using the existing process and with the code in blk-core it didn't
change it's process at all. Creating a helper will create one and I can
look at making a pci_p2pdma_submit_bio() for v6; but if the developer
screws up and still calls the regular submit_bio() things will only be
very subtly broken and that won't be obvious.
I thought about that when reviewing the previous series, and even
started hacking it up. In the end I decided against it for the above
reason - it just adds code, but doesn't actually help with anything
as it is trivial to forget, and not using it will in fact just work.
But if the caller must absolutely know where the bio will end up, then
it seems super redundant. So I'd vote for killing this check, it buys
us absolutely nothing and isn't even exhaustive in its current form.
Ok, I'll remove it for v6.
Since the drivers needs to know it's doing it right, it might not
hurt to add a sanity check helper for that. Just have the driver
call it, and don't add it in the normal IO submission path.
I'm not sure I really see the value in that. It's the same principle in
asking the driver to do the WARN: if the developer knew enough to use
the special helper, they probably knew well enough to do the rest correctly.
I don't agree with that at all. It's a "is my request valid" helper,
it's not some obscure and rarely used functionality. You're making up
this API right now, if you really want it done for every IO, make it
part of the p2p submission process. You could even hide it behind a
debug thing, if you like.
There is no special p2p submission process. In the nvme-of case we are
using the existing process and with the code in blk-core it didn't
change it's process at all. Creating a helper will create one and I can
look at making a pci_p2pdma_submit_bio() for v6; but if the developer
screws up and still calls the regular submit_bio() things will only be
very subtly broken and that won't be obvious.
I'm very sure that something that basic will be caught in review. I
don't care if you wrap the submission or just require the caller to
call some validity helper check first, fwiw.
And I think we're done beating the dead horse at this point.
--
Jens Axboe
From: Christoph Hellwig <hch@lst.de> Date: 2018-09-10 16:37:29
On Wed, Sep 05, 2018 at 03:03:18PM -0600, Logan Gunthorpe wrote:
There is no special p2p submission process. In the nvme-of case we are
using the existing process and with the code in blk-core it didn't
change it's process at all. Creating a helper will create one and I can
look at making a pci_p2pdma_submit_bio() for v6; but if the developer
screws up and still calls the regular submit_bio() things will only be
very subtly broken and that won't be obvious.
I just saw you added that "helper" in your tree. Please don't, it is
a negative value add as it doesn't help anything with the checking.
On Wed, Sep 05, 2018 at 03:03:18PM -0600, Logan Gunthorpe wrote:
quoted
There is no special p2p submission process. In the nvme-of case we are
using the existing process and with the code in blk-core it didn't
change it's process at all. Creating a helper will create one and I can
look at making a pci_p2pdma_submit_bio() for v6; but if the developer
screws up and still calls the regular submit_bio() things will only be
very subtly broken and that won't be obvious.
I just saw you added that "helper" in your tree. Please don't, it is
a negative value add as it doesn't help anything with the checking.