From: Kenneth Lee <hidden> Date: 2018-09-03 00:52:18
From: Kenneth Lee <redacted>
WarpDrive is an accelerator framework to expose the hardware capabilities
directly to the user space. It makes use of the exist vfio and vfio-mdev
facilities. So the user application can send request and DMA to the
hardware without interaction with the kernel. This removes the latency
of syscall.
WarpDrive is the name for the whole framework. The component in kernel
is called SDMDEV, Share Domain Mediated Device. Driver driver exposes its
hardware resource by registering to SDMDEV as a VFIO-Mdev. So the user
library of WarpDrive can access it via VFIO interface.
The patchset contains document for the detail. Please refer to it for more
information.
This patchset is intended to be used with Jean Philippe Brucker's SVA
patch [1], which enables not only IO side page fault, but also PASID
support to IOMMU and VFIO.
With these features, WarpDrive can support non-pinned memory and
multi-process in the same accelerator device. We tested it in our SoC
integrated Accelerator (board ID: D06, Chip ID: HIP08). A reference work
tree can be found here: [2].
But it is not mandatory. This patchset is tested in the latest mainline
kernel without the SVA patches. So it supports only one process for each
accelerator.
We have noticed the IOMMU aware mdev RFC announced recently [3].
The IOMMU aware mdev has similar idea but different intention comparing to
WarpDrive. It intends to dedicate part of the hardware resource to a VM.
And the design is supposed to be used with Scalable I/O Virtualization.
While sdmdev is intended to share the hardware resource with a big amount
of processes. It just requires the hardware supporting address
translation per process (PCIE's PASID or ARM SMMU's substream ID).
But we don't see serious confliction on both design. We believe they can be
normalized as one.
The patch 1 is document of the framework. The patch 2 and 3 add sdmdev
support. The patch 4, 5 and 6 is drivers for Hislicon's ZIP Accelerator
which is registered to both crypto and warpdrive(sdmdev) and can be
used from kernel or user space at the same time. The patch 7 is a user
space sample demonstrating how WarpDrive works.
Change History:
V2 changed from V1:
1. Change kernel framework name from SPIMDEV (Share Parent IOMMU
Mdev) to SDMDEV (Share Domain Mdev).
2. Allocate Hardware Resource when a new mdev is created (While
it is allocated when the mdev is openned)
3. Unmap pages from the shared domain when the sdmdev iommu group is
detached. (This procedure is necessary, but missed in V1)
4. Update document accordingly.
5. Rebase to the latest kernel (4.19.0-rc1)
According the review comment on RFCv1, We did try to use dma-buf
as back end of WarpDrive. It can work properly with the current
solution [4], but it cannot make use of process's
own memory address space directly. This is important to many
acceleration scenario. So dma-buf will be taken as a backup
alternative for noiommu scenario, it will be added in the future
version.
Refernces:
[1] https://www.spinics.net/lists/kernel/msg2651481.html
[2] https://github.com/Kenneth-Lee/linux-kernel-warpdrive/tree/warpdrive-sva-v0.5
[3] https://lkml.org/lkml/2018/7/22/34
[4] https://github.com/Kenneth-Lee/linux-kernel-warpdrive/tree/warpdrive-v0.7-dmabuf
Best Regards
Kenneth Lee
Kenneth Lee (7):
vfio/sdmdev: Add documents for WarpDrive framework
iommu: Add share domain interface in iommu for sdmdev
vfio: add sdmdev support
crypto: add hisilicon Queue Manager driver
crypto: Add Hisilicon Zip driver
crypto: add sdmdev support to Hisilicon QM
vfio/sdmdev: add user sample
Documentation/00-INDEX | 2 +
Documentation/warpdrive/warpdrive.rst | 100 +++
Documentation/warpdrive/wd-arch.svg | 728 ++++++++++++++++
drivers/crypto/Makefile | 2 +-
drivers/crypto/hisilicon/Kconfig | 25 +
drivers/crypto/hisilicon/Makefile | 2 +
drivers/crypto/hisilicon/qm.c | 979 ++++++++++++++++++++++
drivers/crypto/hisilicon/qm.h | 122 +++
drivers/crypto/hisilicon/zip/Makefile | 2 +
drivers/crypto/hisilicon/zip/zip.h | 57 ++
drivers/crypto/hisilicon/zip/zip_crypto.c | 353 ++++++++
drivers/crypto/hisilicon/zip/zip_crypto.h | 8 +
drivers/crypto/hisilicon/zip/zip_main.c | 195 +++++
drivers/iommu/iommu.c | 29 +-
drivers/vfio/Kconfig | 1 +
drivers/vfio/Makefile | 1 +
drivers/vfio/sdmdev/Kconfig | 10 +
drivers/vfio/sdmdev/Makefile | 3 +
drivers/vfio/sdmdev/vfio_sdmdev.c | 363 ++++++++
drivers/vfio/vfio_iommu_type1.c | 151 +++-
include/linux/iommu.h | 15 +
include/linux/vfio_sdmdev.h | 96 +++
include/uapi/linux/vfio_sdmdev.h | 29 +
samples/warpdrive/AUTHORS | 2 +
samples/warpdrive/ChangeLog | 1 +
samples/warpdrive/Makefile.am | 9 +
samples/warpdrive/NEWS | 1 +
samples/warpdrive/README | 32 +
samples/warpdrive/autogen.sh | 3 +
samples/warpdrive/cleanup.sh | 13 +
samples/warpdrive/configure.ac | 52 ++
samples/warpdrive/drv/hisi_qm_udrv.c | 223 +++++
samples/warpdrive/drv/hisi_qm_udrv.h | 53 ++
samples/warpdrive/test/Makefile.am | 7 +
samples/warpdrive/test/comp_hw.h | 23 +
samples/warpdrive/test/test_hisi_zip.c | 206 +++++
samples/warpdrive/wd.c | 309 +++++++
samples/warpdrive/wd.h | 154 ++++
samples/warpdrive/wd_adapter.c | 74 ++
samples/warpdrive/wd_adapter.h | 43 +
40 files changed, 4470 insertions(+), 8 deletions(-)
create mode 100644 Documentation/warpdrive/warpdrive.rst
create mode 100644 Documentation/warpdrive/wd-arch.svg
create mode 100644 drivers/crypto/hisilicon/qm.c
create mode 100644 drivers/crypto/hisilicon/qm.h
create mode 100644 drivers/crypto/hisilicon/zip/Makefile
create mode 100644 drivers/crypto/hisilicon/zip/zip.h
create mode 100644 drivers/crypto/hisilicon/zip/zip_crypto.c
create mode 100644 drivers/crypto/hisilicon/zip/zip_crypto.h
create mode 100644 drivers/crypto/hisilicon/zip/zip_main.c
create mode 100644 drivers/vfio/sdmdev/Kconfig
create mode 100644 drivers/vfio/sdmdev/Makefile
create mode 100644 drivers/vfio/sdmdev/vfio_sdmdev.c
create mode 100644 include/linux/vfio_sdmdev.h
create mode 100644 include/uapi/linux/vfio_sdmdev.h
create mode 100644 samples/warpdrive/AUTHORS
create mode 100644 samples/warpdrive/ChangeLog
create mode 100644 samples/warpdrive/Makefile.am
create mode 100644 samples/warpdrive/NEWS
create mode 100644 samples/warpdrive/README
create mode 100755 samples/warpdrive/autogen.sh
create mode 100755 samples/warpdrive/cleanup.sh
create mode 100644 samples/warpdrive/configure.ac
create mode 100644 samples/warpdrive/drv/hisi_qm_udrv.c
create mode 100644 samples/warpdrive/drv/hisi_qm_udrv.h
create mode 100644 samples/warpdrive/test/Makefile.am
create mode 100644 samples/warpdrive/test/comp_hw.h
create mode 100644 samples/warpdrive/test/test_hisi_zip.c
create mode 100644 samples/warpdrive/wd.c
create mode 100644 samples/warpdrive/wd.h
create mode 100644 samples/warpdrive/wd_adapter.c
create mode 100644 samples/warpdrive/wd_adapter.h
--
2.17.1
From: Kenneth Lee <hidden> Date: 2018-09-03 00:52:29
From: Kenneth Lee <redacted>
This patch add sharing interface for a iommu_group. The new interface:
iommu_group_share_domain()
iommu_group_unshare_domain()
can be used by some virtual iommu_group (such as iommu_group of sdmdev)
to share their parent's iommu_group.
When the domain of a group is shared, it cannot be changed before
being unshared. By this way, all domain users can assume the shared
IOMMU have the same configuration. In the future, notification can be
added if update is required.
Signed-off-by: Kenneth Lee <redacted>
---
drivers/iommu/iommu.c | 29 ++++++++++++++++++++++++++++-
include/linux/iommu.h | 15 +++++++++++++++
2 files changed, 43 insertions(+), 1 deletion(-)
@@ -58,6 +58,9 @@ struct iommu_group {intid;structiommu_domain*default_domain;structiommu_domain*domain;+atomic_tdomain_shared_ref;/* Number of user of current domain.+*Thedomaincannotbemodifiedifref>0+*/};structgroup_device{
@@ -518,6 +522,26 @@ int iommu_group_set_name(struct iommu_group *group, const char *name)}EXPORT_SYMBOL_GPL(iommu_group_set_name);+structiommu_domain*iommu_group_share_domain(structiommu_group*group)+{+/* the domain can be shared only when the default domain is used */+/* todo: more shareable check */+if(group->domain!=group->default_domain)+returnERR_PTR(-EINVAL);++atomic_inc(&group->domain_shared_ref);+returngroup->domain;+}+EXPORT_SYMBOL_GPL(iommu_group_share_domain);++structiommu_domain*iommu_group_unshare_domain(structiommu_group*group)+{+atomic_dec(&group->domain_shared_ref);+WARN_ON(atomic_read(&group->domain_shared_ref)<0);+returngroup->domain;+}+EXPORT_SYMBOL_GPL(iommu_group_unshare_domain);+staticintiommu_group_create_direct_mappings(structiommu_group*group,structdevice*dev){
@@ -1437,7 +1461,8 @@ static int __iommu_attach_group(struct iommu_domain *domain,{intret;-if(group->default_domain&&group->domain!=group->default_domain)+if((group->default_domain&&group->domain!=group->default_domain)||+atomic_read(&group->domain_shared_ref)>0)return-EBUSY;ret=__iommu_group_for_each_dev(group,domain,
From: Kenneth Lee <hidden> Date: 2018-09-03 00:52:29
From: Kenneth Lee <redacted>
WarpDrive is a common user space accelerator framework. Its main component
in Kernel is called sdmdev, Share Domain Mediated Device. It exposes
the hardware capabilities to the user space via vfio-mdev. So processes in
user land can obtain a "queue" by open the device and direct access the
hardware MMIO space or do DMA operation via VFIO interface.
WarpDrive is intended to be used with Jean Philippe Brucker's SVA
patchset to support multi-process. But This is not a must. Without the
SVA patches, WarpDrive can still work for one process for every hardware
device.
This patch add detail documents for the framework.
Signed-off-by: Kenneth Lee <redacted>
---
Documentation/00-INDEX | 2 +
Documentation/warpdrive/warpdrive.rst | 100 ++++
Documentation/warpdrive/wd-arch.svg | 728 ++++++++++++++++++++++++++
3 files changed, 830 insertions(+)
create mode 100644 Documentation/warpdrive/warpdrive.rst
create mode 100644 Documentation/warpdrive/wd-arch.svg
@@ -410,6 +410,8 @@ vm/ - directory with info on the Linux vm code. w1/ - directory with documents regarding the 1-wire (w1) subsystem.+warpdrive/+ - directory with documents about WarpDrive accelerator framework. watchdog/ - how to auto-reboot Linux if it has "fallen and can't get up". ;-) wimax/
@@ -0,0 +1,100 @@+Introduction of WarpDrive+=========================++*WarpDrive* is a general accelerator framework for user space. It intends to+provide interface for the user process to send request to hardware+accelerator without heavy user-kernel interaction cost.++The *WarpDrive* user library is supposed to provide a pipe-based API, such as:+ ::+ int wd_request_queue(struct wd_queue *q);+ void wd_release_queue(struct wd_queue *q);++ int wd_send(struct wd_queue *q, void *req);+ int wd_recv(struct wd_queue *q, void **req);+ int wd_recv_sync(struct wd_queue *q, void **req);+ int wd_flush(struct wd_queue *q);++*wd_request_queue* creates the pipe connection, *queue*, between the+application and the hardware. The application sends request and pulls the+answer back by asynchronized wd_send/wd_recv, which directly interact with the+hardware (by MMIO or share memory) without syscall.++*WarpDrive* maintains a unified application address space among all involved+accelerators. With the following APIs: ::++ int wd_mem_share(struct wd_queue *q, const void *addr,+ size_t size, int flags);+ void wd_mem_unshare(struct wd_queue *q, const void *addr, size_t size);++The referred process space shared by these APIs can be directly referred by the+hardware. The process can also dedicate its whole process space with flags,+*WD_SHARE_ALL* (not in this patch yet).++The name *WarpDrive* is simply a cool and general name meaning the framework+makes the application faster. As it will be explained in this text later, the+facility in kernel is called *SDMDEV*, namely "Share Domain Mediated Device".+++How does it work+================++*WarpDrive* is built upon *VFIO-MDEV*. The queue is wrapped as *mdev* in VFIO.+So memory sharing can be done via standard VFIO standard DMA interface.++The architecture is illustrated as follow figure:++..image:: wd-arch.svg+:alt: WarpDrive Architecture++Accelerator driver shares its capability via *SDMDEV* API: ::++ vfio_sdmdev_register(struct vfio_sdmdev *sdmdev);+ vfio_sdmdev_unregister(struct vfio_sdmdev *sdmdev);+ vfio_sdmdev_wake_up(struct spimdev_queue *q);++*vfio_sdmdev_register* is a helper function to register the hardware to the+*VFIO_MDEV* framework. The queue creation is done by *mdev* creation interface.++*WarpDrive* User library mmap the mdev to access its mmio space and shared+memory. Request can be sent to, or receive from, hardware in this mmap-ed+space until the queue is full or empty.++The user library can wait on the queue by ioctl(VFIO_SDMDEV_CMD_WAIT) the mdev+if the queue is full or empty. If the queue status is changed, the hardware+driver use *vfio_sdmdev_wake_up* to wake up the waiting process.+++Multiple processes support+==========================++In the latest mainline kernel (4.18) when this document is written,+multi-process is not supported in VFIO yet.++Jean Philippe Brucker has a patchset to enable it[1]_. We have tested it+with our hardware (which is known as *D06*). It works well. *WarpDrive* rely+on them to support multiple processes. If it is not enabled, *WarpDrive* can+still work, but it support only one mdev for a process, which will share the+same io map table with kernel. (But it is not going to be a security problem,+since the user application cannot access the kernel address space)++When multiprocess is support, mdev can be created based on how many+hardware resource (queue) is available. Because the VFIO framework accepts only+one open from one mdev iommu_group. Mdev become the smallest unit for process+to use queue. And the mdev will not be released if the user process exist. So+it will need a resource agent to manage the mdev allocation for the user+process. This is not in this document's range.+++Legacy Mode Support+===================+For the hardware on which IOMMU is not support, WarpDrive can run on *NOIOMMU*+mode. That require some update to the mdev driver, which is not included in+this version yet.+++References+==========+..[1] https://patchwork.kernel.org/patch/10394851/++.. vim: tw=78
From: Kenneth Lee <hidden> Date: 2018-09-03 00:52:35
From: Kenneth Lee <redacted>
SDMDEV is "Share Domain Mdev". It is a vfio-mdev. But differ from
the general vfio-mdev, it shares its parent's IOMMU. If Multi-PASID
support is enabled in the IOMMU (not yet in the current kernel HEAD),
multiple process can share the IOMMU by different PASID. If it is not
support, only one process can share the IOMMU with the kernel driver.
Currently only the vfio type-1 driver is updated to make it to be aware
of.
Signed-off-by: Kenneth Lee <redacted>
Signed-off-by: Zaibo Xu <redacted>
Signed-off-by: Zhou Wang <wangzhou1@hisilicon.com>
---
drivers/vfio/Kconfig | 1 +
drivers/vfio/Makefile | 1 +
drivers/vfio/sdmdev/Kconfig | 10 +
drivers/vfio/sdmdev/Makefile | 3 +
drivers/vfio/sdmdev/vfio_sdmdev.c | 363 ++++++++++++++++++++++++++++++
drivers/vfio/vfio_iommu_type1.c | 151 ++++++++++++-
include/linux/vfio_sdmdev.h | 96 ++++++++
include/uapi/linux/vfio_sdmdev.h | 29 +++
8 files changed, 648 insertions(+), 6 deletions(-)
create mode 100644 drivers/vfio/sdmdev/Kconfig
create mode 100644 drivers/vfio/sdmdev/Makefile
create mode 100644 drivers/vfio/sdmdev/vfio_sdmdev.c
create mode 100644 include/linux/vfio_sdmdev.h
create mode 100644 include/uapi/linux/vfio_sdmdev.h
@@ -0,0 +1,363 @@+// SPDX-License-Identifier: GPL-2.0++#include<linux/module.h>+#include<linux/vfio_sdmdev.h>++staticstructclass*sdmdev_class;++staticintvfio_sdmdev_dev_exist(structdevice*dev,void*data)+{+return!strcmp(dev_name(dev),dev_name((structdevice*)data));+}++#ifdef CONFIG_IOMMU_SVA+staticboolvfio_sdmdev_is_valid_pasid(intpasid)+{+structmm_struct*mm;++mm=iommu_sva_find(pasid);+if(mm){+mmput(mm);+returnmm==current->mm;+}++returnfalse;+}+#endif++/* Check if the device is a mediated device belongs to vfio_sdmdev */+intvfio_sdmdev_is_sdmdev(structdevice*dev)+{+structmdev_device*mdev;+structdevice*pdev;++mdev=mdev_from_dev(dev);+if(!mdev)+return0;++pdev=mdev_parent_dev(mdev);+if(!pdev)+return0;++returnclass_for_each_device(sdmdev_class,NULL,pdev,+vfio_sdmdev_dev_exist);+}+EXPORT_SYMBOL_GPL(vfio_sdmdev_is_sdmdev);++structvfio_sdmdev*vfio_sdmdev_pdev_sdmdev(structdevice*dev)+{+structdevice*class_dev;++if(!dev)+returnERR_PTR(-EINVAL);++class_dev=class_find_device(sdmdev_class,NULL,dev,+(int(*)(structdevice*,constvoid*))vfio_sdmdev_dev_exist);+if(!class_dev)+returnERR_PTR(-ENODEV);++returncontainer_of(class_dev,structvfio_sdmdev,cls_dev);+}+EXPORT_SYMBOL_GPL(vfio_sdmdev_pdev_sdmdev);++structvfio_sdmdev*mdev_sdmdev(structmdev_device*mdev)+{+structdevice*pdev=mdev_parent_dev(mdev);++returnvfio_sdmdev_pdev_sdmdev(pdev);+}+EXPORT_SYMBOL_GPL(mdev_sdmdev);++staticssize_tiommu_type_show(structdevice*dev,+structdevice_attribute*attr,char*buf)+{+structvfio_sdmdev*sdmdev=vfio_sdmdev_pdev_sdmdev(dev);++if(!sdmdev)+return-ENODEV;++returnsprintf(buf,"%d\n",sdmdev->iommu_type);+}++staticDEVICE_ATTR_RO(iommu_type);++staticssize_tdma_flag_show(structdevice*dev,+structdevice_attribute*attr,char*buf)+{+structvfio_sdmdev*sdmdev=vfio_sdmdev_pdev_sdmdev(dev);++if(!sdmdev)+return-ENODEV;++returnsprintf(buf,"%d\n",sdmdev->dma_flag);+}++staticDEVICE_ATTR_RO(dma_flag);++/* mdev->dev_attr_groups */+staticstructattribute*vfio_sdmdev_attrs[]={+&dev_attr_iommu_type.attr,+&dev_attr_dma_flag.attr,+NULL,+};+staticconststructattribute_groupvfio_sdmdev_group={+.name=VFIO_SDMDEV_PDEV_ATTRS_GRP_NAME,+.attrs=vfio_sdmdev_attrs,+};+conststructattribute_group*vfio_sdmdev_groups[]={+&vfio_sdmdev_group,+NULL,+};++/* default attributes for mdev->supported_type_groups, used by registerer*/+#define MDEV_TYPE_ATTR_RO_EXPORT(name) \+MDEV_TYPE_ATTR_RO(name);\+EXPORT_SYMBOL_GPL(mdev_type_attr_##name);++#define DEF_SIMPLE_SDMDEV_ATTR(_name, sdmdev_member, format) \+staticssize_t_name##_show(structkobject*kobj,structdevice*dev,\+char*buf)\+{\+structvfio_sdmdev*sdmdev=vfio_sdmdev_pdev_sdmdev(dev);\+if(!sdmdev)\+return-ENODEV;\+returnsprintf(buf,format,sdmdev->sdmdev_member);\+}\+MDEV_TYPE_ATTR_RO_EXPORT(_name)++DEF_SIMPLE_SDMDEV_ATTR(flags,flags,"%d");+DEF_SIMPLE_SDMDEV_ATTR(name,name,"%s");/* this should be algorithm name, */+/* but you would not care if you have only one algorithm */+DEF_SIMPLE_SDMDEV_ATTR(device_api,api_ver,"%s");++staticssize_t+available_instances_show(structkobject*kobj,structdevice*dev,char*buf)+{+structvfio_sdmdev*sdmdev=vfio_sdmdev_pdev_sdmdev(dev);+intnr_inst=0;++nr_inst=sdmdev->ops->get_available_instances?+sdmdev->ops->get_available_instances(sdmdev):0;+returnsprintf(buf,"%d",nr_inst);+}+MDEV_TYPE_ATTR_RO_EXPORT(available_instances);++staticintvfio_sdmdev_mdev_create(structkobject*kobj,+structmdev_device*mdev)+{+structdevice*pdev=mdev_parent_dev(mdev);+structvfio_sdmdev_queue*q;+structvfio_sdmdev*sdmdev=mdev_sdmdev(mdev);+intret;++if(!sdmdev->ops->get_queue)+return-ENODEV;++ret=sdmdev->ops->get_queue(sdmdev,&q);+if(ret)+returnret;++q->sdmdev=sdmdev;+q->mdev=mdev;+init_waitqueue_head(&q->wait);++mdev_set_drvdata(mdev,q);+get_device(pdev);++return0;+}++staticintvfio_sdmdev_mdev_remove(structmdev_device*mdev)+{+structvfio_sdmdev_queue*q=+(structvfio_sdmdev_queue*)mdev_get_drvdata(mdev);+structvfio_sdmdev*sdmdev=q->sdmdev;+structdevice*pdev=mdev_parent_dev(mdev);++put_device(pdev);++if(sdmdev->ops->put_queue);+sdmdev->ops->put_queue(q);++return0;+}++/* Wake up the process who is waiting this queue */+voidvfio_sdmdev_wake_up(structvfio_sdmdev_queue*q)+{+wake_up_all(&q->wait);+}+EXPORT_SYMBOL_GPL(vfio_sdmdev_wake_up);++staticintvfio_sdmdev_mdev_mmap(structmdev_device*mdev,+structvm_area_struct*vma)+{+structvfio_sdmdev_queue*q=+(structvfio_sdmdev_queue*)mdev_get_drvdata(mdev);+structvfio_sdmdev*sdmdev=q->sdmdev;++if(sdmdev->ops->mmap)+returnsdmdev->ops->mmap(q,vma);++dev_err(sdmdev->dev,"no driver mmap!\n");+return-EINVAL;+}++staticinlineintvfio_sdmdev_wait(structvfio_sdmdev_queue*q,+unsignedlongtimeout)+{+intret;+structvfio_sdmdev*sdmdev=q->sdmdev;++if(!sdmdev->ops->mask_notify)+return-ENODEV;++sdmdev->ops->mask_notify(q,VFIO_SDMDEV_EVENT_Q_UPDATE);++ret=timeout?wait_event_interruptible_timeout(q->wait,+sdmdev->ops->is_q_updated(q),timeout):+wait_event_interruptible(q->wait,+sdmdev->ops->is_q_updated(q));++sdmdev->ops->mask_notify(q,0);++returnret;+}++staticlongvfio_sdmdev_mdev_ioctl(structmdev_device*mdev,unsignedintcmd,+unsignedlongarg)+{+structvfio_sdmdev_queue*q=+(structvfio_sdmdev_queue*)mdev_get_drvdata(mdev);+structvfio_sdmdev*sdmdev=q->sdmdev;++switch(cmd){+caseVFIO_SDMDEV_CMD_WAIT:+returnvfio_sdmdev_wait(q,arg);++#ifdef CONFIG_IOMMU_SVA+caseVFIO_SDMDEV_CMD_BIND_PASID:+intret;++if(!vfio_sdmdev_is_valid_pasid(arg))+return-EINVAL;++mutex_lock(&q->mutex);+q->pasid=arg;++if(sdmdev->ops->start_queue)+ret=sdmdev->ops->start_queue(q);++mutex_unlock(&q->mutex);++returnret;+#endif++default:+if(sdmdev->ops->ioctl)+returnsdmdev->ops->ioctl(q,cmd,arg);++dev_err(sdmdev->dev,"ioctl cmd (%d) is not supported!\n",cmd);+return-EINVAL;+}+}++staticvoidvfio_sdmdev_release(structdevice*dev){}++staticvoidvfio_sdmdev_mdev_release(structmdev_device*mdev)+{+structvfio_sdmdev_queue*q=+(structvfio_sdmdev_queue*)mdev_get_drvdata(mdev);+structvfio_sdmdev*sdmdev=q->sdmdev;++if(sdmdev->ops->stop_queue)+sdmdev->ops->stop_queue(q);+}++staticintvfio_sdmdev_mdev_open(structmdev_device*mdev)+{+#ifndef CONFIG_IOMMU_SVA+structvfio_sdmdev_queue*q=+(structvfio_sdmdev_queue*)mdev_get_drvdata(mdev);+structvfio_sdmdev*sdmdev=q->sdmdev;++if(sdmdev->ops->start_queue)+sdmdev->ops->start_queue(q);+#endif++return0;+}++/**+*vfio_sdmdev_register-registerasdmdev+*@sdmdev:devicestructure+*/+intvfio_sdmdev_register(structvfio_sdmdev*sdmdev)+{+intret;++if(!sdmdev->dev)+return-ENODEV;++atomic_set(&sdmdev->ref,0);+sdmdev->cls_dev.parent=sdmdev->dev;+sdmdev->cls_dev.class=sdmdev_class;+sdmdev->cls_dev.release=vfio_sdmdev_release;+dev_set_name(&sdmdev->cls_dev,"%s",dev_name(sdmdev->dev));+ret=device_register(&sdmdev->cls_dev);+if(ret)+gotoerr;++sdmdev->mdev_fops.owner=THIS_MODULE;+sdmdev->mdev_fops.dev_attr_groups=vfio_sdmdev_groups;+WARN_ON(!sdmdev->mdev_fops.supported_type_groups);+sdmdev->mdev_fops.create=vfio_sdmdev_mdev_create;+sdmdev->mdev_fops.remove=vfio_sdmdev_mdev_remove;+sdmdev->mdev_fops.ioctl=vfio_sdmdev_mdev_ioctl;+sdmdev->mdev_fops.open=vfio_sdmdev_mdev_open;+sdmdev->mdev_fops.release=vfio_sdmdev_mdev_release;+sdmdev->mdev_fops.mmap=vfio_sdmdev_mdev_mmap,++ret=mdev_register_device(sdmdev->dev,&sdmdev->mdev_fops);+if(ret)+gotoerr_with_cls_dev;++return0;++err_with_cls_dev:+device_unregister(&sdmdev->cls_dev);+err:+returnret;+}+EXPORT_SYMBOL_GPL(vfio_sdmdev_register);++/**+*vfio_sdmdev_unregister-unregistersasdmdev+*@sdmdev:devicetounregister+*+*Unregisterasdmdevthatwatpreviouslysuccessullyregisteredwith+*vfio_sdmdev_register().+*/+voidvfio_sdmdev_unregister(structvfio_sdmdev*sdmdev)+{+mdev_unregister_device(sdmdev->dev);+device_unregister(&sdmdev->cls_dev);+}+EXPORT_SYMBOL_GPL(vfio_sdmdev_unregister);++staticint__initvfio_sdmdev_init(void)+{+sdmdev_class=class_create(THIS_MODULE,VFIO_SDMDEV_CLASS_NAME);+returnPTR_ERR_OR_ZERO(sdmdev_class);+}++static__exitvoidvfio_sdmdev_exit(void)+{+class_destroy(sdmdev_class);+}++module_init(vfio_sdmdev_init);+module_exit(vfio_sdmdev_exit);++MODULE_LICENSE("GPL");+MODULE_AUTHOR("Hisilicon Tech. Co., Ltd.");+MODULE_DESCRIPTION("VFIO Share Domain Mediated Device");
@@ -1327,6 +1330,109 @@ static bool vfio_iommu_has_sw_msi(struct iommu_group *group, phys_addr_t *base)returnret;}+/* return 0 if the device is not sdmdev.+*return1ifthedeviceissdmdev,thedatawillbeupdatedwithparent+*device'sgroup.+*return-errnoifothererror.+*/+staticintvfio_sdmdev_type(structdevice*dev,void*data)+{+structiommu_group**group=data;+structiommu_group*pgroup;+int(*_is_sdmdev)(structdevice*dev);+structdevice*pdev;+intret=1;++/* vfio_sdmdev module is not configurated */+_is_sdmdev=symbol_get(vfio_sdmdev_is_sdmdev);+if(!_is_sdmdev)+return0;++/* check if it belongs to vfio_sdmdev device */+if(!_is_sdmdev(dev)){+ret=0;+gotoout;+}++pdev=dev->parent;+pgroup=iommu_group_get(pdev);+if(!pgroup){+ret=-ENODEV;+gotoout;+}++if(group){+/* check if all parent devices is the same */+if(*group&&*group!=pgroup)+ret=-ENODEV;+else+*group=pgroup;+}++iommu_group_put(pgroup);++out:+symbol_put(vfio_sdmdev_is_sdmdev);++returnret;+}++/* return 0 or -errno */+staticintvfio_sdmdev_bus(structdevice*dev,void*data)+{+structbus_type**bus=data;++if(!dev->bus)+return-ENODEV;++/* ensure all devices has the same bus_type */+if(*bus&&*bus!=dev->bus)+return-EINVAL;++*bus=dev->bus;+return0;+}++/* return 0 means it is not sd group, 1 means it is, or -EXXX for error */+staticintvfio_iommu_type1_attach_sdgroup(structvfio_domain*domain,+structvfio_group*group,+structiommu_group*iommu_group)+{+intret;+structbus_type*pbus=NULL;+structiommu_group*pgroup=NULL;++ret=iommu_group_for_each_dev(iommu_group,&pgroup,+vfio_sdmdev_type);+if(ret<0)+gotoout;+elseif(ret>0){+domain->domain=iommu_group_share_domain(pgroup);+if(IS_ERR(domain->domain))+gotoout;+ret=iommu_group_for_each_dev(pgroup,&pbus,+vfio_sdmdev_bus);+if(ret<0)+gotoerr_with_share_domain;++if(pbus&&iommu_capable(pbus,IOMMU_CAP_CACHE_COHERENCY))+domain->prot|=IOMMU_CACHE;++group->parent_group=pgroup;+INIT_LIST_HEAD(&domain->group_list);+list_add(&group->next,&domain->group_list);++return1;+}++return0;++err_with_share_domain:+iommu_group_unshare_domain(pgroup);+out:+returnret;+}+staticintvfio_iommu_type1_attach_group(void*iommu_data,structiommu_group*iommu_group){
@@ -1335,8 +1441,8 @@ static int vfio_iommu_type1_attach_group(void *iommu_data,structvfio_domain*domain,*d;structbus_type*bus=NULL,*mdev_bus;intret;-boolresv_msi,msi_remap;-phys_addr_tresv_msi_base;+boolresv_msi=false,msi_remap;+phys_addr_tresv_msi_base=0;mutex_lock(&iommu->lock);
@@ -1373,6 +1479,14 @@ static int vfio_iommu_type1_attach_group(void *iommu_data,if(mdev_bus){if((bus==mdev_bus)&&!iommu_present(bus)){symbol_put(mdev_bus_type);++ret=vfio_iommu_type1_attach_sdgroup(domain,group,+iommu_group);+if(ret<0)+gotoout_free;+elseif(ret>0)+gotoreplay_check;+if(!iommu->external_domain){INIT_LIST_HEAD(&domain->group_list);iommu->external_domain=domain;
@@ -1451,12 +1565,13 @@ static int vfio_iommu_type1_attach_group(void *iommu_data,vfio_test_domain_fgsp(domain);+replay_check:/* replay mappings on new domains */ret=vfio_iommu_replay(iommu,domain);if(ret)gotoout_detach;-if(resv_msi){+if(!group->parent_group&&resv_msi){ret=iommu_get_msi_cookie(domain->domain,resv_msi_base);if(ret)gotoout_detach;
@@ -1471,7 +1586,10 @@ static int vfio_iommu_type1_attach_group(void *iommu_data,out_detach:iommu_detach_group(domain->domain,iommu_group);out_domain:-iommu_domain_free(domain->domain);+if(group->parent_group)+iommu_group_unshare_domain(group->parent_group);+else+iommu_domain_free(domain->domain);out_free:kfree(domain);kfree(group);
From: Kenneth Lee <hidden> Date: 2018-09-03 00:52:41
From: Kenneth Lee <redacted>
Hisilicon QM is a general IP used by some Hisilicon accelerators. It
provides a general PCIE interface for the CPU and the accelerator to share
a group of queues.
This commit includes a library used by the accelerator driver to access
the QM hardware.
Signed-off-by: Kenneth Lee <redacted>
Signed-off-by: Zhou Wang <wangzhou1@hisilicon.com>
Signed-off-by: Hao Fang <redacted>
---
drivers/crypto/Makefile | 2 +-
drivers/crypto/hisilicon/Kconfig | 8 +
drivers/crypto/hisilicon/Makefile | 1 +
drivers/crypto/hisilicon/qm.c | 820 ++++++++++++++++++++++++++++++
drivers/crypto/hisilicon/qm.h | 110 ++++
5 files changed, 940 insertions(+), 1 deletion(-)
create mode 100644 drivers/crypto/hisilicon/qm.c
create mode 100644 drivers/crypto/hisilicon/qm.h
From: Kenneth Lee <hidden> Date: 2018-09-03 00:52:47
From: Kenneth Lee <redacted>
The Hisilicon ZIP accelerator implements zlib and gzip algorithm support
for the software. It uses Hisilicon QM as the interface to the CPU, so it
is shown up as a PCIE device to the CPU with a group of queues.
This commit provides PCIE driver to the accelerator and register it to
the crypto subsystem.
Signed-off-by: Kenneth Lee <redacted>
Signed-off-by: Zhou Wang <wangzhou1@hisilicon.com>
Signed-off-by: Hao Fang <redacted>
---
drivers/crypto/hisilicon/Kconfig | 7 +
drivers/crypto/hisilicon/Makefile | 1 +
drivers/crypto/hisilicon/zip/Makefile | 2 +
drivers/crypto/hisilicon/zip/zip.h | 57 ++++
drivers/crypto/hisilicon/zip/zip_crypto.c | 353 ++++++++++++++++++++++
drivers/crypto/hisilicon/zip/zip_crypto.h | 8 +
drivers/crypto/hisilicon/zip/zip_main.c | 195 ++++++++++++
7 files changed, 623 insertions(+)
create mode 100644 drivers/crypto/hisilicon/zip/Makefile
create mode 100644 drivers/crypto/hisilicon/zip/zip.h
create mode 100644 drivers/crypto/hisilicon/zip/zip_crypto.c
create mode 100644 drivers/crypto/hisilicon/zip/zip_crypto.h
create mode 100644 drivers/crypto/hisilicon/zip/zip_main.c
@@ -0,0 +1,353 @@+/* SPDX-License-Identifier: GPL-2.0+ */+#include<linux/crypto.h>+#include<linux/dma-mapping.h>+#include<linux/pci.h>+#include<linux/topology.h>+#include"../qm.h"+#include"zip.h"++#define INPUT_BUFFER_SIZE (64 * 1024)+#define OUTPUT_BUFFER_SIZE (64 * 1024)++#define COMP_NAME_TO_TYPE(alg_name) \+(!strcmp((alg_name),"zlib-deflate")?0x02:\+!strcmp((alg_name),"gzip")?0x03:0)\++structhisi_zip_buffer{+u8*input;+dma_addr_tinput_dma;+u8*output;+dma_addr_toutput_dma;+};++structhisi_zip_qp_ctx{+structhisi_zip_bufferbuffer;+structhisi_qp*qp;+structhisi_zip_sqezip_sqe;+};++structhisi_zip_ctx{+#define QPC_COMP 0+#define QPC_DECOMP 1+structhisi_zip_qp_ctxqp_ctx[2];+};++staticstructhisi_zip*find_zip_device(intnode)+{+structhisi_zip*hisi_zip,*ret=NULL;+structdevice*dev;+intmin_distance=100;++list_for_each_entry(hisi_zip,&hisi_zip_list,list){+dev=&hisi_zip->qm.pdev->dev;+if(node_distance(dev->numa_node,node)<min_distance){+ret=hisi_zip;+min_distance=node_distance(dev->numa_node,node);+}+}++returnret;+}++staticvoidhisi_zip_qp_event_notifier(structhisi_qp*qp)+{+complete(&qp->completion);+}++staticinthisi_zip_fill_sqe_v1(void*sqe,void*q_parm,u32len)+{+structhisi_zip_sqe*zip_sqe=(structhisi_zip_sqe*)sqe;+structhisi_zip_qp_ctx*qp_ctx=(structhisi_zip_qp_ctx*)q_parm;+structhisi_zip_buffer*buffer=&qp_ctx->buffer;++memset(zip_sqe,0,sizeof(structhisi_zip_sqe));++zip_sqe->input_data_length=len;+zip_sqe->dw9=qp_ctx->qp->req_type;+zip_sqe->dest_avail_out=OUTPUT_BUFFER_SIZE;+zip_sqe->source_addr_l=lower_32_bits(buffer->input_dma);+zip_sqe->source_addr_h=upper_32_bits(buffer->input_dma);+zip_sqe->dest_addr_l=lower_32_bits(buffer->output_dma);+zip_sqe->dest_addr_h=upper_32_bits(buffer->output_dma);++return0;+}++/* let's allocate one buffer now, may have problem in async case */+staticinthisi_zip_alloc_qp_buffer(structhisi_zip_qp_ctx*hisi_zip_qp_ctx)+{+structhisi_zip_buffer*buffer=&hisi_zip_qp_ctx->buffer;+structhisi_qp*qp=hisi_zip_qp_ctx->qp;+structdevice*dev=&qp->qm->pdev->dev;+intret;++buffer->input=dma_alloc_coherent(dev,INPUT_BUFFER_SIZE,+&buffer->input_dma,GFP_KERNEL);+if(!buffer->input)+return-ENOMEM;++buffer->output=dma_alloc_coherent(dev,OUTPUT_BUFFER_SIZE,+&buffer->output_dma,GFP_KERNEL);+if(!buffer->output){+ret=-ENOMEM;+gotoerr_alloc_output_buffer;+}++return0;++err_alloc_output_buffer:+dma_free_coherent(dev,INPUT_BUFFER_SIZE,buffer->input,+buffer->input_dma);+returnret;+}++staticvoidhisi_zip_free_qp_buffer(structhisi_zip_qp_ctx*hisi_zip_qp_ctx)+{+structhisi_zip_buffer*buffer=&hisi_zip_qp_ctx->buffer;+structhisi_qp*qp=hisi_zip_qp_ctx->qp;+structdevice*dev=&qp->qm->pdev->dev;++dma_free_coherent(dev,INPUT_BUFFER_SIZE,buffer->input,+buffer->input_dma);+dma_free_coherent(dev,OUTPUT_BUFFER_SIZE,buffer->output,+buffer->output_dma);+}++staticinthisi_zip_create_qp(structqm_info*qm,structhisi_zip_qp_ctx*ctx,+intalg_type,intreq_type)+{+structhisi_qp*qp;+intret;++qp=hisi_qm_create_qp(qm,alg_type);++if(IS_ERR(qp))+returnPTR_ERR(qp);++qp->event_cb=hisi_zip_qp_event_notifier;+qp->req_type=req_type;++qp->qp_ctx=ctx;+ctx->qp=qp;++ret=hisi_zip_alloc_qp_buffer(ctx);+if(ret)+gotoerr_with_qp;++ret=hisi_qm_start_qp(qp,0);+if(ret<0)+gotoerr_with_qp_buffer;++return0;+err_with_qp_buffer:+hisi_zip_free_qp_buffer(ctx);+err_with_qp:+hisi_qm_release_qp(qp);+returnret;+}++staticvoidhisi_zip_release_qp(structhisi_zip_qp_ctx*ctx)+{+hisi_qm_release_qp(ctx->qp);+hisi_zip_free_qp_buffer(ctx);+}++staticinthisi_zip_alloc_comp_ctx(structcrypto_tfm*tfm)+{+structhisi_zip_ctx*hisi_zip_ctx=crypto_tfm_ctx(tfm);+constchar*alg_name=crypto_tfm_alg_name(tfm);+structhisi_zip*hisi_zip;+structqm_info*qm;+intret,i,j;++u8req_type=COMP_NAME_TO_TYPE(alg_name);++/* find the proper zip device */+hisi_zip=find_zip_device(cpu_to_node(smp_processor_id()));+if(!hisi_zip){+pr_err("Can not find proper ZIP device!\n");+return-ENODEV;+}+qm=&hisi_zip->qm;++for(i=0;i<2;i++){+/* it is just happen that 0 is compress, 1 is decompress on alg_type */+ret=hisi_zip_create_qp(qm,&hisi_zip_ctx->qp_ctx[i],i,+req_type);+if(ret)+gotoerr;+}++return0;+err:+for(j=i-1;j>=0;j--)+hisi_zip_release_qp(&hisi_zip_ctx->qp_ctx[j]);++returnret;+}++staticvoidhisi_zip_free_comp_ctx(structcrypto_tfm*tfm)+{+structhisi_zip_ctx*hisi_zip_ctx=crypto_tfm_ctx(tfm);+inti;++/* release the qp */+for(i=1;i>=0;i--)+hisi_zip_release_qp(&hisi_zip_ctx->qp_ctx[i]);+}++staticinthisi_zip_copy_data_to_buffer(structhisi_zip_qp_ctx*qp_ctx,+constu8*src,unsignedintslen)+{+structhisi_zip_buffer*buffer=&qp_ctx->buffer;++if(slen>INPUT_BUFFER_SIZE)+return-EINVAL;++memcpy(buffer->input,src,slen);++return0;+}++staticstructhisi_zip_sqe*hisi_zip_get_writeback_sqe(structhisi_qp*qp)+{+structhisi_acc_qp_status*qp_status=&qp->qp_status;+structhisi_zip_sqe*sq_base=QP_SQE_ADDR(qp);+u16sq_head=qp_status->sq_head;++returnsq_base+sq_head;+}++staticinthisi_zip_copy_data_from_buffer(structhisi_zip_qp_ctx*qp_ctx,+u8*dst,unsignedint*dlen)+{+structhisi_zip_buffer*buffer=&qp_ctx->buffer;+structhisi_qp*qp=qp_ctx->qp;+structhisi_zip_sqe*zip_sqe=hisi_zip_get_writeback_sqe(qp);+u32status=zip_sqe->dw3&0xff;+u16sq_head;++if(status!=0){+pr_err("hisi zip: %s fail!\n",(qp->alg_type==0)?+"compression":"decompression");+returnstatus;+}++if(zip_sqe->produced>OUTPUT_BUFFER_SIZE)+return-ENOMEM;++memcpy(dst,buffer->output,zip_sqe->produced);+*dlen=zip_sqe->produced;++sq_head=qp->qp_status.sq_head;+if(sq_head==QM_Q_DEPTH-1)+qp->qp_status.sq_head=0;+else+qp->qp_status.sq_head++;++return0;+}++staticinthisi_zip_compress(structcrypto_tfm*tfm,constu8*src,+unsignedintslen,u8*dst,unsignedint*dlen)+{+structhisi_zip_ctx*hisi_zip_ctx=crypto_tfm_ctx(tfm);+structhisi_zip_qp_ctx*qp_ctx=&hisi_zip_ctx->qp_ctx[QPC_COMP];+structhisi_qp*qp=qp_ctx->qp;+structhisi_zip_sqe*zip_sqe=&qp_ctx->zip_sqe;+intret;++ret=hisi_zip_copy_data_to_buffer(qp_ctx,src,slen);+if(ret<0)+returnret;++hisi_zip_fill_sqe_v1(zip_sqe,qp_ctx,slen);++/* send command to start the compress job */+hisi_qp_send(qp,zip_sqe);++returnhisi_zip_copy_data_from_buffer(qp_ctx,dst,dlen);+}++staticinthisi_zip_decompress(structcrypto_tfm*tfm,constu8*src,+unsignedintslen,u8*dst,unsignedint*dlen)+{+structhisi_zip_ctx*hisi_zip_ctx=crypto_tfm_ctx(tfm);+structhisi_zip_qp_ctx*qp_ctx=&hisi_zip_ctx->qp_ctx[QPC_DECOMP];+structhisi_qp*qp=qp_ctx->qp;+structhisi_zip_sqe*zip_sqe=&qp_ctx->zip_sqe;+intret;++ret=hisi_zip_copy_data_to_buffer(qp_ctx,src,slen);+if(ret<0)+returnret;++hisi_zip_fill_sqe_v1(zip_sqe,qp_ctx,slen);++/* send command to start the decompress job */+hisi_qp_send(qp,zip_sqe);++returnhisi_zip_copy_data_from_buffer(qp_ctx,dst,dlen);+}++staticstructcrypto_alghisi_zip_zlib={+.cra_name="zlib-deflate",+.cra_flags=CRYPTO_ALG_TYPE_COMPRESS,+.cra_ctxsize=sizeof(structhisi_zip_ctx),+.cra_priority=300,+.cra_module=THIS_MODULE,+.cra_init=hisi_zip_alloc_comp_ctx,+.cra_exit=hisi_zip_free_comp_ctx,+.cra_u={+.compress={+.coa_compress=hisi_zip_compress,+.coa_decompress=hisi_zip_decompress+}+}+};++staticstructcrypto_alghisi_zip_gzip={+.cra_name="gzip",+.cra_flags=CRYPTO_ALG_TYPE_COMPRESS,+.cra_ctxsize=sizeof(structhisi_zip_ctx),+.cra_priority=300,+.cra_module=THIS_MODULE,+.cra_init=hisi_zip_alloc_comp_ctx,+.cra_exit=hisi_zip_free_comp_ctx,+.cra_u={+.compress={+.coa_compress=hisi_zip_compress,+.coa_decompress=hisi_zip_decompress+}+}+};++inthisi_zip_register_to_crypto(void)+{+intret;++ret=crypto_register_alg(&hisi_zip_zlib);+if(ret<0){+pr_err("Zlib algorithm registration failed\n");+returnret;+}++ret=crypto_register_alg(&hisi_zip_gzip);+if(ret<0){+pr_err("Gzip algorithm registration failed\n");+gotoerr_unregister_zlib;+}++return0;++err_unregister_zlib:+crypto_unregister_alg(&hisi_zip_zlib);++returnret;+}++voidhisi_zip_unregister_from_crypto(void)+{+crypto_unregister_alg(&hisi_zip_zlib);+crypto_unregister_alg(&hisi_zip_gzip);+}
From: Kenneth Lee <hidden> Date: 2018-09-03 00:52:53
From: Kenneth Lee <redacted>
This commit add spimdev support to the Hislicon QM driver, any
accelerator that use QM can expose its queues to the user space.
Signed-off-by: Kenneth Lee <redacted>
Signed-off-by: Zhou Wang <wangzhou1@hisilicon.com>
Signed-off-by: Hao Fang <redacted>
Signed-off-by: Zaibo Xu <redacted>
---
drivers/crypto/hisilicon/Kconfig | 10 ++
drivers/crypto/hisilicon/qm.c | 159 +++++++++++++++++++++++++++++++
drivers/crypto/hisilicon/qm.h | 12 +++
3 files changed, 181 insertions(+)
@@ -639,6 +639,155 @@ int hisi_qp_send(struct hisi_qp *qp, void *msg)}EXPORT_SYMBOL_GPL(hisi_qp_send);+#ifdef CONFIG_CRYPTO_DEV_HISI_SDMDEV+/* mdev->supported_type_groups */+staticstructattribute*hisi_qm_type_attrs[]={+VFIO_SDMDEV_DEFAULT_MDEV_TYPE_ATTRS,+NULL,+};+staticstructattribute_grouphisi_qm_type_group={+.attrs=hisi_qm_type_attrs,+};+staticstructattribute_group*mdev_type_groups[]={+&hisi_qm_type_group,+NULL,+};++staticvoidqm_qp_event_notifier(structhisi_qp*qp)+{+vfio_sdmdev_wake_up(qp->sdmdev_q);+}++staticinthisi_qm_get_queue(structvfio_sdmdev*sdmdev,+structvfio_sdmdev_queue**q)+{+structqm_info*qm=sdmdev->priv;+structhisi_qp*qp=NULL;+structvfio_sdmdev_queue*wd_q;+u8alg_type=0;/* fix me here */+intret;++qp=hisi_qm_create_qp(qm,alg_type);+if(IS_ERR(qp))+returnPTR_ERR(qp);++wd_q=kzalloc(sizeof(structvfio_sdmdev_queue),GFP_KERNEL);+if(!wd_q){+ret=-ENOMEM;+gotoerr_with_qp;+}++wd_q->priv=qp;+wd_q->sdmdev=sdmdev;+*q=wd_q;+qp->sdmdev_q=wd_q;+qp->event_cb=qm_qp_event_notifier;++return0;++err_with_qp:+hisi_qm_release_qp(qp);+returnret;+}++voidhisi_qm_put_queue(structvfio_sdmdev_queue*q)+{+structhisi_qp*qp=q->priv;++hisi_qm_release_qp(qp);+kfree(q);+}++/* map sq/cq/doorbell to user space */+staticinthisi_qm_mmap(structvfio_sdmdev_queue*q,+structvm_area_struct*vma)+{+structhisi_qp*qp=(structhisi_qp*)q->priv;+structqm_info*qm=qp->qm;+structdevice*dev=&qm->pdev->dev;+size_tsz=vma->vm_end-vma->vm_start;+u8region;++vma->vm_flags|=(VM_IO|VM_LOCKED|VM_DONTEXPAND|VM_DONTDUMP);+region=_VFIO_SDMDEV_REGION(vma->vm_pgoff);++switch(region){+case0:+if(sz>PAGE_SIZE)+return-EINVAL;+/*+*Warning:Thisisnotsafeasmultiplequeuesusethesame+*doorbell,v1hardwareinterfaceproblem.v2willfixit+*/+returnremap_pfn_range(vma,vma->vm_start,+qm->phys_base>>PAGE_SHIFT,+sz,pgprot_noncached(vma->vm_page_prot));+case1:+vma->vm_pgoff=0;+if(sz>qp->scqe.size)+return-EINVAL;++returndma_mmap_coherent(dev,vma,qp->scqe.addr,qp->scqe.dma,+sz);++default:+return-EINVAL;+}+}++staticinthisi_qm_start_queue(structvfio_sdmdev_queue*q)+{+structhisi_qp*qp=q->priv;++#ifdef CONFIG_IOMMU_SVA+returnhisi_qm_start_qp(qp,q->pasid);+#else+returnhisi_qm_start_qp(qp,0);+#endif+}++staticvoidhisi_qm_stop_queue(structvfio_sdmdev_queue*q)+{+/* need to stop hardware, but can not support in v1 */+}++staticconststructvfio_sdmdev_opsqm_ops={+.get_queue=hisi_qm_get_queue,+.put_queue=hisi_qm_put_queue,+.start_queue=hisi_qm_start_queue,+.stop_queue=hisi_qm_stop_queue,+.mmap=hisi_qm_mmap,+};++staticintqm_register_sdmdev(structqm_info*qm)+{+structpci_dev*pdev=qm->pdev;+structvfio_sdmdev*sdmdev=&qm->sdmdev;++sdmdev->iommu_type=VFIO_TYPE1_IOMMU;++#ifdef CONFIG_IOMMU_SVA+sdmdev->dma_flag=VFIO_SDMDEV_DMA_MULTI_PROC_MAP;+#else+sdmdev->dma_flag=VFIO_SDMDEV_DMA_SINGLE_PROC_MAP;+#endif++sdmdev->name=qm->dev_name;+sdmdev->dev=&pdev->dev;+sdmdev->is_vf=pdev->is_virtfn;+sdmdev->priv=qm;+sdmdev->api_ver="hisi_qm_v1";+sdmdev->flags=0;++sdmdev->mdev_fops.mdev_attr_groups=qm->mdev_dev_groups;+hisi_qm_type_group.name=qm->dev_name;+sdmdev->mdev_fops.supported_type_groups=mdev_type_groups;+sdmdev->ops=&qm_ops;++returnvfio_sdmdev_register(sdmdev);+}+#endif+inthisi_qm_init(constchar*dev_name,structqm_info*qm){structpci_dev*pdev=qm->pdev;
@@ -769,6 +918,12 @@ int hisi_qm_start(struct qm_info *qm)if(ret)gotoerr_with_cqc;+#ifdef CONFIG_CRYPTO_DEV_HISI_SDMDEV+ret=qm_register_sdmdev(qm);+if(ret)+gotoerr_with_cqc;+#endif+writel(0x0,QM_ADDR(qm,QM_VF_EQ_INT_MASK));return0;
@@ -0,0 +1,32 @@+WD User Land Demonstration+==========================++This directory contains some applications and libraries to demonstrate how a++WrapDrive application can be constructed.+++As a demo, we try to make it simple and clear for understanding. It is not++supposed to be used in business scenario.+++The directory contains the following elements:++wd.[ch]+ A demonstration WrapDrive fundamental library which wraps the basic+ operations to the WrapDrive-ed device.++wd_adapter.[ch]+ User driver adaptor for wd.[ch]++wd_utils.[ch]+ Some utitlities function used by WD and its drivers++drv/*+ User drivers. It helps to fulfill the semantic of wd.[ch] for+ particular hardware++test/*+ Test applications to use the wrapdrive library+
@@ -0,0 +1,52 @@+AC_PREREQ([2.69])+AC_INIT([wrapdrive], [0.1], [liguozhu@hisilicon.com])+AC_CONFIG_SRCDIR([wd.c])+AM_INIT_AUTOMAKE([1.10 no-define])++AC_CONFIG_MACRO_DIR([m4])+AC_CONFIG_HEADERS([config.h])++# Checks for programs.+AC_PROG_CXX+AC_PROG_AWK+AC_PROG_CC+AC_PROG_CPP+AC_PROG_INSTALL+AC_PROG_LN_S+AC_PROG_MAKE_SET+AC_PROG_RANLIB++AM_PROG_AR+AC_PROG_LIBTOOL+AM_PROG_LIBTOOL+LT_INIT+AM_PROG_CC_C_O++AC_DEFINE([HAVE_SVA], [0], [enable SVA support])+AC_ARG_ENABLE([sva],+ [ --enable-sva enable to support sva feature],+ AC_DEFINE([HAVE_SVA], [1]))++# Checks for libraries.++# Checks for header files.+AC_CHECK_HEADERS([fcntl.h stdint.h stdlib.h string.h sys/ioctl.h sys/time.h unistd.h])++# Checks for typedefs, structures, and compiler characteristics.+AC_CHECK_HEADER_STDBOOL+AC_C_INLINE+AC_TYPE_OFF_T+AC_TYPE_SIZE_T+AC_TYPE_UINT16_T+AC_TYPE_UINT32_T+AC_TYPE_UINT64_T+AC_TYPE_UINT8_T++# Checks for library functions.+AC_FUNC_MALLOC+AC_FUNC_MMAP+AC_CHECK_FUNCS([memset munmap])++AC_CONFIG_FILES([Makefile+ test/Makefile])+AC_OUTPUT
@@ -0,0 +1,309 @@+// SPDX-License-Identifier: GPL-2.0+#include"config.h"+#include<stdlib.h>+#include<unistd.h>+#include<sys/types.h>+#include<sys/stat.h>+#include<sys/queue.h>+#include<fcntl.h>+#include<sys/ioctl.h>+#include<errno.h>+#include<sys/mman.h>+#include<string.h>+#include<assert.h>+#include<dirent.h>+#include<sys/poll.h>+#include"wd.h"+#include"wd_adapter.h"++#if (defined(HAVE_SVA) & HAVE_SVA)+staticint_wd_bind_process(structwd_queue*q)+{+structbind_data{+structvfio_iommu_type1_bindbind;+structvfio_iommu_type1_bind_processdata;+}wd_bind;+intret;+__u32flags=0;++if(q->dma_flag&VFIO_SDMDEV_DMA_MULTI_PROC_MAP)+flags=VFIO_IOMMU_BIND_PRIV;+elseif(q->dma_flag&VFIO_SDMDEV_DMA_SVM_NO_FAULT)+flags=VFIO_IOMMU_BIND_NOPF;++wd_bind.bind.flags=VFIO_IOMMU_BIND_PROCESS;+wd_bind.bind.argsz=sizeof(wd_bind);+wd_bind.data.flags=flags;+ret=ioctl(q->container,VFIO_IOMMU_BIND,&wd_bind);+if(ret)+returnret;+q->pasid=wd_bind.data.pasid;+returnret;+}++staticint_wd_unbind_process(structwd_queue*q)+{+structbind_data{+structvfio_iommu_type1_bindbind;+structvfio_iommu_type1_bind_processdata;+}wd_bind;+__u32flags=0;++if(q->dma_flag&VFIO_SDMDEV_DMA_MULTI_PROC_MAP)+flags=VFIO_IOMMU_BIND_PRIV;+elseif(q->dma_flag&VFIO_SDMDEV_DMA_SVM_NO_FAULT)+flags=VFIO_IOMMU_BIND_NOPF;++wd_bind.bind.flags=VFIO_IOMMU_BIND_PROCESS;+wd_bind.data.pasid=q->pasid;+wd_bind.data.flags=flags;+wd_bind.bind.argsz=sizeof(wd_bind);++returnioctl(q->container,VFIO_IOMMU_UNBIND,&wd_bind);+}+#endif++intwd_request_queue(structwd_queue*q)+{+structvfio_group_statusgroup_status={+.argsz=sizeof(group_status)};+intiommu_ext;+intret;++if(!q->vfio_group_path||+!q->device_api_path||+!q->dmaflag_ext_path||+!q->iommu_ext_path){+WD_ERR("please set vfio_group_path, dmaflag_ext_path, "+"device_api_path, and iommu_ext_path before call %s",__func__);+return-EINVAL;+}++q->hw_type_id=0;/* this can be set according to the device api_version in the future */++q->group=open(q->vfio_group_path,O_RDWR);+if(q->group<0){+WD_ERR("open vfio group(%s) fail, errno=%d\n",+q->vfio_group_path,errno);+return-errno;+}++if(q->container<=0){+q->container=open("/dev/vfio/vfio",O_RDWR);+if(q->container<0){+WD_ERR("Create VFIO container fail!\n");+ret=-ENODEV;+gotoerr_with_group;+}+}++if(ioctl(q->container,VFIO_GET_API_VERSION)!=VFIO_API_VERSION){+WD_ERR("VFIO version check fail!\n");+ret=-EINVAL;+gotoerr_with_container;+}++q->dma_flag=_get_attr_int(q->dmaflag_ext_path);+if(q->dma_flag==INT_MIN){+ret=-EINVAL;+gotoerr_with_container;+}++iommu_ext=_get_attr_int(q->iommu_ext_path);+if(iommu_ext==INT_MIN){+ret=-EINVAL;+gotoerr_with_container;+}++ret=ioctl(q->container,VFIO_CHECK_EXTENSION,iommu_ext);+if(!ret){+WD_ERR("VFIO iommu check (%d) fail (%d)!\n",iommu_ext,ret);+gotoerr_with_container;+}++ret=_get_attr_str(q->device_api_path,q->hw_type);+if(ret)+gotoerr_with_container;++ret=ioctl(q->group,VFIO_GROUP_GET_STATUS,&group_status);+if(!(group_status.flags&VFIO_GROUP_FLAGS_VIABLE)){+WD_ERR("VFIO group is not viable\n");+gotoerr_with_container;+}++ret=ioctl(q->group,VFIO_GROUP_SET_CONTAINER,&q->container);+if(ret){+WD_ERR("VFIO group fail on VFIO_GROUP_SET_CONTAINER\n");+gotoerr_with_container;+}++ret=ioctl(q->container,VFIO_SET_IOMMU,iommu_ext);+if(ret){+WD_ERR("VFIO fail on VFIO_SET_IOMMU(%d)\n",iommu_ext);+gotoerr_with_container;+}++q->mdev=ioctl(q->group,VFIO_GROUP_GET_DEVICE_FD,q->mdev_name);+if(q->mdev<0){+WD_ERR("VFIO fail on VFIO_GROUP_GET_DEVICE_FD (%d)\n",q->mdev);+ret=q->mdev;+gotoerr_with_container;+}++#if (defined(HAVE_SVA) & HAVE_SVA)+if(!(q->dma_flag&(VFIO_SDMDEV_DMA_PHY|VFIO_SDMDEV_DMA_SINGLE_PROC_MAP))){+ret=_wd_bind_process(q);+if(ret){+close(q->mdev);+WD_ERR("VFIO fails to bind process!\n");+gotoerr_with_mdev;++}+}++ret=ioctl(q->mdev,VFIO_SDMDEV_CMD_BIND_PASID,(unsignedlong)q->pasid);+if(ret<0){+WD_ERR("fail to bind paisd to device,ret=%d\n",errno);+gotoerr_with_mdev;+}+#endif++ret=drv_open(q);+if(ret)+gotoerr_with_mdev;++return0;++err_with_mdev:+close(q->mdev);+err_with_container:+close(q->container);+err_with_group:+close(q->group);+returnret;+}++voidwd_release_queue(structwd_queue*q)+{+drv_close(q);++#if (defined(HAVE_SVA) & HAVE_SVA)+if(!(q->dma_flag&(VFIO_SDMDEV_DMA_PHY|VFIO_SDMDEV_DMA_SINGLE_PROC_MAP))){+if(q->pasid<=0){+WD_ERR("Wd queue pasid ! pasid=%d\n",q->pasid);+return;+}+if(_wd_unbind_process(q)){+WD_ERR("VFIO fails to unbind process!\n");+return;+}+}+#endif++close(q->mdev);+close(q->container);+close(q->group);+}++intwd_send(structwd_queue*q,void*req)+{+returndrv_send(q,req);+}++intwd_recv(structwd_queue*q,void**resp)+{+returndrv_recv(q,resp);+}++staticintwd_flush_and_wait(structwd_queue*q,intms)+{+wd_flush(q);+returnioctl(q->mdev,VFIO_SDMDEV_CMD_WAIT,ms);+}++intwd_recv_sync(structwd_queue*q,void**resp,__u16ms)+{+intret;++while(1){+ret=wd_recv(q,resp);+if(ret==-EBUSY){+ret=wd_flush_and_wait(q,ms);+if(ret)+returnret;+}else+returnret;+}+}++voidwd_flush(structwd_queue*q)+{+drv_flush(q);+}++staticint_wd_mem_share_type1(structwd_queue*q,constvoid*addr,+size_tsize,intflags)+{+structvfio_iommu_type1_dma_mapdma_map;++if(q->dma_flag&VFIO_SDMDEV_DMA_SVM_NO_FAULT)+returnmlock(addr,size);++#if (defined(HAVE_SVA) & HAVE_SVA)+elseif((q->dma_flag&VFIO_SDMDEV_DMA_MULTI_PROC_MAP)&&+(q->pasid>0))+dma_map.pasid=q->pasid;+#endif+elseif((q->dma_flag&VFIO_SDMDEV_DMA_SINGLE_PROC_MAP))+;//todo+else+return-1;++dma_map.vaddr=(__u64)addr;+dma_map.size=size;+dma_map.iova=(__u64)addr;+dma_map.flags=+VFIO_DMA_MAP_FLAG_READ|VFIO_DMA_MAP_FLAG_WRITE|flags;+dma_map.argsz=sizeof(dma_map);++returnioctl(q->container,VFIO_IOMMU_MAP_DMA,&dma_map);+}++staticvoid_wd_mem_unshare_type1(structwd_queue*q,constvoid*addr,+size_tsize)+{+#if (defined(HAVE_SVA) & HAVE_SVA)+structvfio_iommu_type1_dma_unmapdma_unmap;+#endif++if(q->dma_flag&VFIO_SDMDEV_DMA_SVM_NO_FAULT){+(void)munlock(addr,size);+return;+}++#if (defined(HAVE_SVA) & HAVE_SVA)+dma_unmap.iova=(__u64)addr;+if((q->dma_flag&VFIO_SDMDEV_DMA_MULTI_PROC_MAP)&&(q->pasid>0))+dma_unmap.flags=0;+dma_unmap.size=size;+dma_unmap.argsz=sizeof(dma_unmap);+ioctl(q->container,VFIO_IOMMU_UNMAP_DMA,&dma_unmap);+#endif+}++intwd_mem_share(structwd_queue*q,constvoid*addr,size_tsize,intflags)+{+if(drv_can_do_mem_share(q))+returndrv_share(q,addr,size,flags);+else+return_wd_mem_share_type1(q,addr,size,flags);+}++voidwd_mem_unshare(structwd_queue*q,constvoid*addr,size_tsize)+{+if(drv_can_do_mem_share(q))+drv_unshare(q,addr,size);+else+_wd_mem_unshare_type1(q,addr,size);+}+
@@ -0,0 +1,154 @@+// SPDX-License-Identifier: GPL-2.0+#ifndef __WD_H+#define __WD_H+#include<stdlib.h>+#include<errno.h>+#include<stdio.h>+#include<string.h>+#include<sys/types.h>+#include<sys/stat.h>+#include<fcntl.h>+#include<stdint.h>+#include<unistd.h>+#include<limits.h>+#include"../../include/uapi/linux/vfio.h"+#include"../../include/uapi/linux/vfio_sdmdev.h"++#define SYS_VAL_SIZE 16+#define PATH_STR_SIZE 256+#define WD_NAME_SIZE 64+#define WD_MAX_MEMLIST_SZ 128+++#ifndef dma_addr_t+#define dma_addr_t __u64+#endif++typedefintbool;++#ifndef true+#define true 1+#endif++#ifndef false+#define false 0+#endif++/* the flags used by wd_capa->flags, the high 16bits are for algorithm+*andthelow16bitsareforFramework+*/+#define WD_FLAGS_FW_PREFER_LOCAL_ZONE 1++#define WD_FLAGS_FW_MASK 0x0000FFFF+#ifndef WD_ERR+#define WD_ERR(format, args...) fprintf(stderr, format, ##args)+#endif++/* Default page size should be 4k size */+#define WDQ_MAP_REGION(region_index) ((region_index << 12) & 0xf000)+#define WDQ_MAP_Q(q_index) ((q_index << 16) & 0xffff0000)++staticinlinevoidwd_reg_write(void*reg_addr,uint32_tvalue)+{+*((volatileuint32_t*)reg_addr)=value;+}++staticinlineuint32_twd_reg_read(void*reg_addr)+{+uint32_ttemp;++temp=*((volatileuint32_t*)reg_addr);++returntemp;+}++staticinlineint_get_attr_str(constchar*path,charvalue[PATH_STR_SIZE])+{+intfd,ret;++fd=open(path,O_RDONLY);+if(fd<0){+WD_ERR("get_attr_str: open %s fail\n",path);+returnfd;+}+memset(value,0,PATH_STR_SIZE);+ret=read(fd,value,PATH_STR_SIZE);+if(ret>0){+close(fd);+return0;+}+close(fd);++WD_ERR("read nothing from %s\n",path);+return-EINVAL;+}++staticinlineint_get_attr_int(constchar*path)+{+charvalue[PATH_STR_SIZE];+if(_get_attr_str(path,value))+returnINT_MIN;+else+returnatoi(value);+}++/* Memory in accelerating message can be different */+enumwd_addr_flags{+WD_AATTR_INVALID=0,++/* Common user virtual memory */+_WD_AATTR_COM_VIRT=1,++/* Physical address*/+_WD_AATTR_PHYS=2,++/* I/O virtual address*/+_WD_AATTR_IOVA=4,++/* SGL, user cares for */+WD_AATTR_SGL=8,+};++#define WD_CAPA_PRIV_DATA_SIZE 64++#define alloc_obj(objp) do { \+objp=malloc(sizeof(*objp));\+memset(objp,0,sizeof(*objp));\+}while(0)+#define free_obj(objp) if (objp)free(objp)++structwd_queue{+constchar*mdev_name;+charhw_type[PATH_STR_SIZE];+inthw_type_id;+intdma_flag;+void*priv;/* private data used by the drv layer */+intcontainer;+intgroup;+intmdev;+intpasid;+intiommu_type;+char*vfio_group_path;+char*iommu_ext_path;+char*dmaflag_ext_path;+char*device_api_path;+};++externintwd_request_queue(structwd_queue*q);+externvoidwd_release_queue(structwd_queue*q);+externintwd_send(structwd_queue*q,void*req);+externintwd_recv(structwd_queue*q,void**resp);+externvoidwd_flush(structwd_queue*q);+externintwd_recv_sync(structwd_queue*q,void**resp,__u16ms);+externintwd_mem_share(structwd_queue*q,constvoid*addr,+size_tsize,intflags);+externvoidwd_mem_unshare(structwd_queue*q,constvoid*addr,size_tsize);++/* for debug only */+externintwd_dump_all_algos(void);++/* this is only for drv used */+externintwd_set_queue_attr(structwd_queue*q,constchar*name,+char*value);+externint__iommu_type(structwd_queue*q);+#endif
@@ -0,0 +1,74 @@+// SPDX-License-Identifier: GPL-2.0+#include<stdio.h>+#include<string.h>+#include<dirent.h>+++#include"wd_adapter.h"+#include"./drv/hisi_qm_udrv.h"++staticstructwd_drv_dio_ifhw_dio_tbl[]={{+.hw_type="hisi_qm_v1",+.open=hisi_qm_set_queue_dio,+.close=hisi_qm_unset_queue_dio,+.send=hisi_qm_add_to_dio_q,+.recv=hisi_qm_get_from_dio_q,+},+/* Add other drivers direct IO operations here */+};++/* todo: there should be some stable way to match the device and the driver */+#define MAX_HW_TYPE (sizeof(hw_dio_tbl) / sizeof(hw_dio_tbl[0]))++intdrv_open(structwd_queue*q)+{+inti;++//todo: try to find another dev if the user driver is not avaliable+for(i=0;i<MAX_HW_TYPE;i++){+if(!strcmp(q->hw_type,+hw_dio_tbl[i].hw_type)){+q->hw_type_id=i;+returnhw_dio_tbl[q->hw_type_id].open(q);+}+}+WD_ERR("No matching driver to use!\n");+errno=ENODEV;+return-ENODEV;+}++voiddrv_close(structwd_queue*q)+{+hw_dio_tbl[q->hw_type_id].close(q);+}++intdrv_send(structwd_queue*q,void*req)+{+returnhw_dio_tbl[q->hw_type_id].send(q,req);+}++intdrv_recv(structwd_queue*q,void**req)+{+returnhw_dio_tbl[q->hw_type_id].recv(q,req);+}++intdrv_share(structwd_queue*q,constvoid*addr,size_tsize,intflags)+{+returnhw_dio_tbl[q->hw_type_id].share(q,addr,size,flags);+}++voiddrv_unshare(structwd_queue*q,constvoid*addr,size_tsize)+{+hw_dio_tbl[q->hw_type_id].unshare(q,addr,size);+}++booldrv_can_do_mem_share(structwd_queue*q)+{+returnhw_dio_tbl[q->hw_type_id].share!=NULL;+}++voiddrv_flush(structwd_queue*q)+{+if(hw_dio_tbl[q->hw_type_id].flush)+hw_dio_tbl[q->hw_type_id].flush(q);+}
@@ -0,0 +1,43 @@+// SPDX-License-Identifier: GPL-2.0+/* the common drv header define the unified interface for wd */+#ifndef __WD_ADAPTER_H__+#define __WD_ADAPTER_H__++#include<stdio.h>+#include<string.h>+#include<stdlib.h>+#include<stdint.h>+#include<unistd.h>+#include<fcntl.h>+#include<sys/stat.h>+#include<sys/ioctl.h>+#include<sys/mman.h>+++#include"wd.h"++structwd_drv_dio_if{+char*hw_type;+int(*open)(structwd_queue*q);+void(*close)(structwd_queue*q);+int(*set_pasid)(structwd_queue*q);+int(*unset_pasid)(structwd_queue*q);+int(*send)(structwd_queue*q,void*req);+int(*recv)(structwd_queue*q,void**req);+void(*flush)(structwd_queue*q);+int(*share)(structwd_queue*q,constvoid*addr,+size_tsize,intflags);+int(*unshare)(structwd_queue*q,constvoid*addr,size_tsize);+};++externintdrv_open(structwd_queue*q);+externvoiddrv_close(structwd_queue*q);+externintdrv_send(structwd_queue*q,void*req);+externintdrv_recv(structwd_queue*q,void**req);+externvoiddrv_flush(structwd_queue*q);+externintdrv_share(structwd_queue*q,constvoid*addr,+size_tsize,intflags);+externvoiddrv_unshare(structwd_queue*q,constvoid*addr,size_tsize);+externbooldrv_can_do_mem_share(structwd_queue*q);++#endif
At a minimum:
Enable this to enable the SDMDEV,
although that could be done better. Maybe just:
Enable the SDMDEV "shared IOMMU Domain Mediated Device"
+ interface for all Hisilicon accelerators if they can. The SDMDEV
probably drop "if they can": accelerators. The SDMDEV interface
+ enable the WarpDrive user space accelerator driver to access the
From: Randy Dunlap <rdunlap@infradead.org> Date: 2018-09-03 02:25:28
On 09/02/2018 05:52 PM, Kenneth Lee wrote:
From: Kenneth Lee <redacted>
This is the sample code to demostrate how WrapDrive user application
should be.
It contains:
1. wd.[ch], the common library to provide WrapDrive interface.
WarpDrive
2. drv/*, the user driver to access the hardware upon spimdev
3. test/*, the test application to use WrapDrive interface to access the
@@ -0,0 +1,32 @@+WD User Land Demonstration+==========================++This directory contains some applications and libraries to demonstrate how a++WrapDrive application can be constructed.
WarpDrive
quoted hunk
+++As a demo, we try to make it simple and clear for understanding. It is not++supposed to be used in business scenario.+++The directory contains the following elements:++wd.[ch]+ A demonstration WrapDrive fundamental library which wraps the basic
WarpDrive
+ operations to the WrapDrive-ed device.
WarpDrive
quoted hunk
++wd_adapter.[ch]+ User driver adaptor for wd.[ch]++wd_utils.[ch]+ Some utitlities function used by WD and its drivers++drv/*+ User drivers. It helps to fulfill the semantic of wd.[ch] for+ particular hardware++test/*+ Test applications to use the wrapdrive library
From: Lu Baolu <baolu.lu@linux.intel.com> Date: 2018-09-03 02:33:23
Hi,
On 09/03/2018 08:51 AM, Kenneth Lee wrote:
From: Kenneth Lee <redacted>
WarpDrive is an accelerator framework to expose the hardware capabilities
directly to the user space. It makes use of the exist vfio and vfio-mdev
facilities. So the user application can send request and DMA to the
hardware without interaction with the kernel. This removes the latency
of syscall.
WarpDrive is the name for the whole framework. The component in kernel
is called SDMDEV, Share Domain Mediated Device. Driver driver exposes its
hardware resource by registering to SDMDEV as a VFIO-Mdev. So the user
library of WarpDrive can access it via VFIO interface.
The patchset contains document for the detail. Please refer to it for more
information.
This patchset is intended to be used with Jean Philippe Brucker's SVA
patch [1], which enables not only IO side page fault, but also PASID
support to IOMMU and VFIO.
With these features, WarpDrive can support non-pinned memory and
multi-process in the same accelerator device. We tested it in our SoC
integrated Accelerator (board ID: D06, Chip ID: HIP08). A reference work
tree can be found here: [2].
But it is not mandatory. This patchset is tested in the latest mainline
kernel without the SVA patches. So it supports only one process for each
accelerator.
We have noticed the IOMMU aware mdev RFC announced recently [3].
The IOMMU aware mdev has similar idea but different intention comparing to
WarpDrive. It intends to dedicate part of the hardware resource to a VM.
And the design is supposed to be used with Scalable I/O Virtualization.
While sdmdev is intended to share the hardware resource with a big amount
of processes. It just requires the hardware supporting address
translation per process (PCIE's PASID or ARM SMMU's substream ID).
But we don't see serious confliction on both design. We believe they can be
normalized as one.
The patch 1 is document of the framework. The patch 2 and 3 add sdmdev
support. The patch 4, 5 and 6 is drivers for Hislicon's ZIP Accelerator
which is registered to both crypto and warpdrive(sdmdev) and can be
used from kernel or user space at the same time. The patch 7 is a user
space sample demonstrating how WarpDrive works.
Change History:
V2 changed from V1:
1. Change kernel framework name from SPIMDEV (Share Parent IOMMU
Mdev) to SDMDEV (Share Domain Mdev).
2. Allocate Hardware Resource when a new mdev is created (While
it is allocated when the mdev is openned)
3. Unmap pages from the shared domain when the sdmdev iommu group is
detached. (This procedure is necessary, but missed in V1)
4. Update document accordingly.
5. Rebase to the latest kernel (4.19.0-rc1)
According the review comment on RFCv1, We did try to use dma-buf
as back end of WarpDrive. It can work properly with the current
solution [4], but it cannot make use of process's
own memory address space directly. This is important to many
acceleration scenario. So dma-buf will be taken as a backup
alternative for noiommu scenario, it will be added in the future
version.
Refernces:
[1] https://www.spinics.net/lists/kernel/msg2651481.html
[2] https://github.com/Kenneth-Lee/linux-kernel-warpdrive/tree/warpdrive-sva-v0.5
[3] https://lkml.org/lkml/2018/7/22/34
From: Lu Baolu <baolu.lu@linux.intel.com> Date: 2018-09-03 02:57:06
Hi,
On 09/03/2018 08:52 AM, Kenneth Lee wrote:
From: Kenneth Lee <redacted>
SDMDEV is "Share Domain Mdev". It is a vfio-mdev. But differ from
the general vfio-mdev, it shares its parent's IOMMU. If Multi-PASID
support is enabled in the IOMMU (not yet in the current kernel HEAD),
multiple process can share the IOMMU by different PASID. If it is not
support, only one process can share the IOMMU with the kernel driver.
If only for share domain purpose, I don't think it's necessary to create
a new device type.
Currently only the vfio type-1 driver is updated to make it to be aware
of.
Signed-off-by: Kenneth Lee <redacted>
Signed-off-by: Zaibo Xu <redacted>
Signed-off-by: Zhou Wang <wangzhou1@hisilicon.com>
---
drivers/vfio/Kconfig | 1 +
drivers/vfio/Makefile | 1 +
drivers/vfio/sdmdev/Kconfig | 10 +
drivers/vfio/sdmdev/Makefile | 3 +
drivers/vfio/sdmdev/vfio_sdmdev.c | 363 ++++++++++++++++++++++++++++++
drivers/vfio/vfio_iommu_type1.c | 151 ++++++++++++-
include/linux/vfio_sdmdev.h | 96 ++++++++
include/uapi/linux/vfio_sdmdev.h | 29 +++
8 files changed, 648 insertions(+), 6 deletions(-)
create mode 100644 drivers/vfio/sdmdev/Kconfig
create mode 100644 drivers/vfio/sdmdev/Makefile
create mode 100644 drivers/vfio/sdmdev/vfio_sdmdev.c
create mode 100644 include/linux/vfio_sdmdev.h
create mode 100644 include/uapi/linux/vfio_sdmdev.h
@@ -1327,6 +1330,109 @@ static bool vfio_iommu_has_sw_msi(struct iommu_group *group, phys_addr_t *base)returnret;}+/* return 0 if the device is not sdmdev.+*return1ifthedeviceissdmdev,thedatawillbeupdatedwithparent+*device'sgroup.+*return-errnoifothererror.+*/+staticintvfio_sdmdev_type(structdevice*dev,void*data)+{+structiommu_group**group=data;+structiommu_group*pgroup;+int(*_is_sdmdev)(structdevice*dev);+structdevice*pdev;+intret=1;++/* vfio_sdmdev module is not configurated */+_is_sdmdev=symbol_get(vfio_sdmdev_is_sdmdev);+if(!_is_sdmdev)+return0;++/* check if it belongs to vfio_sdmdev device */+if(!_is_sdmdev(dev)){+ret=0;+gotoout;+}++pdev=dev->parent;+pgroup=iommu_group_get(pdev);+if(!pgroup){+ret=-ENODEV;+gotoout;+}++if(group){+/* check if all parent devices is the same */+if(*group&&*group!=pgroup)+ret=-ENODEV;+else+*group=pgroup;+}++iommu_group_put(pgroup);++out:+symbol_put(vfio_sdmdev_is_sdmdev);++returnret;+}++/* return 0 or -errno */+staticintvfio_sdmdev_bus(structdevice*dev,void*data)+{+structbus_type**bus=data;++if(!dev->bus)+return-ENODEV;++/* ensure all devices has the same bus_type */+if(*bus&&*bus!=dev->bus)+return-EINVAL;++*bus=dev->bus;+return0;+}++/* return 0 means it is not sd group, 1 means it is, or -EXXX for error */+staticintvfio_iommu_type1_attach_sdgroup(structvfio_domain*domain,+structvfio_group*group,+structiommu_group*iommu_group)+{+intret;+structbus_type*pbus=NULL;+structiommu_group*pgroup=NULL;++ret=iommu_group_for_each_dev(iommu_group,&pgroup,+vfio_sdmdev_type);+if(ret<0)+gotoout;+elseif(ret>0){+domain->domain=iommu_group_share_domain(pgroup);+if(IS_ERR(domain->domain))+gotoout;+ret=iommu_group_for_each_dev(pgroup,&pbus,+vfio_sdmdev_bus);+if(ret<0)+gotoerr_with_share_domain;++if(pbus&&iommu_capable(pbus,IOMMU_CAP_CACHE_COHERENCY))+domain->prot|=IOMMU_CACHE;++group->parent_group=pgroup;+INIT_LIST_HEAD(&domain->group_list);+list_add(&group->next,&domain->group_list);++return1;+}
This doesn't match the function name. It only gets the domain from the
parent device. It hasn't been really attached.
@@ -1373,6 +1479,14 @@ static int vfio_iommu_type1_attach_group(void *iommu_data, if (mdev_bus) { if ((bus == mdev_bus) && !iommu_present(bus)) { symbol_put(mdev_bus_type);++ ret = vfio_iommu_type1_attach_sdgroup(domain, group,+ iommu_group);+ if (ret < 0)+ goto out_free;+ else if (ret > 0)+ goto replay_check;
Here you get the domain from the parent device and save it for later
use. The actual attaching is ignored.
I don't think this follows the philosophy of this function. It actually
make all devices in the group with the same bus type to share a single
domain.
Further more, the parent domain might be a domain of type
IOMMU_DOMAIN_DMA. That will not be able to use as an
IOMMU_DOMAIN_UNMANAGED domain for iommu APIs.
Best regards,
Lu Baolu
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
From: Kenneth Lee <redacted>
WarpDrive is an accelerator framework to expose the hardware capabilities
directly to the user space. It makes use of the exist vfio and vfio-mdev
facilities. So the user application can send request and DMA to the
hardware without interaction with the kernel. This removes the latency
of syscall.
WarpDrive is the name for the whole framework. The component in kernel
is called SDMDEV, Share Domain Mediated Device. Driver driver exposes its
hardware resource by registering to SDMDEV as a VFIO-Mdev. So the user
library of WarpDrive can access it via VFIO interface.
The patchset contains document for the detail. Please refer to it for more
information.
This patchset is intended to be used with Jean Philippe Brucker's SVA
patch [1], which enables not only IO side page fault, but also PASID
support to IOMMU and VFIO.
With these features, WarpDrive can support non-pinned memory and
multi-process in the same accelerator device. We tested it in our SoC
integrated Accelerator (board ID: D06, Chip ID: HIP08). A reference work
tree can be found here: [2].
But it is not mandatory. This patchset is tested in the latest mainline
kernel without the SVA patches. So it supports only one process for each
accelerator.
We have noticed the IOMMU aware mdev RFC announced recently [3].
The IOMMU aware mdev has similar idea but different intention comparing to
WarpDrive. It intends to dedicate part of the hardware resource to a VM.
And the design is supposed to be used with Scalable I/O Virtualization.
While sdmdev is intended to share the hardware resource with a big amount
of processes. It just requires the hardware supporting address
translation per process (PCIE's PASID or ARM SMMU's substream ID).
But we don't see serious confliction on both design. We believe they can be
normalized as one.
So once again i do not understand why you are trying to do things
this way. Kernel already have tons of example of everything you
want to do without a new framework. Moreover i believe you are
confuse by VFIO. To me VFIO is for VM not to create general device
driver frame work.
So here is your use case as i understand it. You have a device
with a limited number of command queues (can be just one) and in
some case it can support SVA/SVM (when hardware support it and it
is not disabled). Final requirement is being able to schedule cmds
from userspace without ioctl. All of this exists already exists
upstream in few device drivers.
So here is how every body else is doing it. Please explain why
this does not work.
1 Userspace open device file driver. Kernel device driver create
a context and associate it with on open. This context can be
uniq to the process and can bind hardware resources (like a
command queue) to the process.
2 Userspace bind/acquire a commands queue and initialize it with
an ioctl on the device file. Through that ioctl userspace can
be inform wether either SVA/SVM works for the device. If SVA/
SVM works then kernel device driver bind the process to the
device as part of this ioctl.
3 If SVM/SVA does not work userspace do an ioctl to create dma
buffer or something that does exactly the same thing.
4 Userspace mmap the command queue (mmap of the device file by
using informations gather at step 2)
5 Userspace can write commands into the queue it mapped
6 When userspace close the device file all resources are release
just like any existing device drivers.
Now if you want to create a device driver framework that expose
a device file with generic API for all of the above steps fine.
But it does not need to be part of VFIO whatsoever or explain
why.
Note that if IOMMU is fully disabled you probably want to block
userspace from being able to directly scheduling commands onto
the hardware as it would allow userspace to DMA anywhere and thus
would open the kernel to easy exploits. In this case you can still
keeps the same API as above and use page fault tricks to valid
commands written by userspace into fake commands ring. This will
be as slow or maybe even slower than ioctl but at least it allows
you to validate commands.
Cheers,
Jérôme
drivers/vfio/sdmdev/vfio_sdmdev.c:106:30: sparse: symbol 'vfio_sdmdev_groups' was not declared. Should it be static?
drivers/vfio/sdmdev/vfio_sdmdev.c: In function 'vfio_sdmdev_mdev_remove':
drivers/vfio/sdmdev/vfio_sdmdev.c:178:2: warning: this 'if' clause does not guard... [-Wmisleading-indentation]
if (sdmdev->ops->put_queue);
^~
drivers/vfio/sdmdev/vfio_sdmdev.c:179:3: note: ...this statement, but the latter is misleadingly indented as if it were guarded by the 'if'
sdmdev->ops->put_queue(q);
^~~~~~
Please review and possibly fold the followup patch.
---
0-DAY kernel test infrastructure Open Source Technology Center
https://lists.01.org/pipermail/kbuild-all Intel Corporation
From: kbuild test robot <hidden> Date: 2018-09-04 15:21:10
Hi Kenneth,
Thank you for the patch! Perhaps something to improve:
[auto build test WARNING on cryptodev/master]
[also build test WARNING on v4.19-rc2 next-20180831]
[if your patch is applied to the wrong git tree, please drop us a note to help improve the system]
url: https://github.com/0day-ci/linux/commits/Kenneth-Lee/A-General-Accelerator-Framework-WarpDrive/20180903-162733
base: https://git.kernel.org/pub/scm/linux/kernel/git/herbert/cryptodev-2.6.git master
config: i386-allmodconfig (attached as .config)
compiler: gcc-7 (Debian 7.3.0-16) 7.3.0
reproduce:
# save the attached .config to linux build tree
make ARCH=i386
:::::: branch date: 2 hours ago
:::::: commit date: 2 hours ago
All warnings (new ones prefixed by >>):
drivers/vfio/sdmdev/vfio_sdmdev.c: In function 'vfio_sdmdev_mdev_remove':
quoted
drivers/vfio/sdmdev/vfio_sdmdev.c:178:2: warning: this 'if' clause does not guard... [-Wmisleading-indentation]
if (sdmdev->ops->put_queue);
^~
drivers/vfio/sdmdev/vfio_sdmdev.c:179:3: note: ...this statement, but the latter is misleadingly indented as if it were guarded by the 'if'
sdmdev->ops->put_queue(q);
^~~~~~
# https://github.com/0day-ci/linux/commit/1e47d5e608652b4a2c813dbeaf5aa6811f6ceaf7
git remote add linux-review https://github.com/0day-ci/linux
git remote update linux-review
git checkout 1e47d5e608652b4a2c813dbeaf5aa6811f6ceaf7
vim +/if +178 drivers/vfio/sdmdev/vfio_sdmdev.c
1e47d5e6 Kenneth Lee 2018-09-03 168
1e47d5e6 Kenneth Lee 2018-09-03 169 static int vfio_sdmdev_mdev_remove(struct mdev_device *mdev)
1e47d5e6 Kenneth Lee 2018-09-03 170 {
1e47d5e6 Kenneth Lee 2018-09-03 171 struct vfio_sdmdev_queue *q =
1e47d5e6 Kenneth Lee 2018-09-03 172 (struct vfio_sdmdev_queue *)mdev_get_drvdata(mdev);
1e47d5e6 Kenneth Lee 2018-09-03 173 struct vfio_sdmdev *sdmdev = q->sdmdev;
1e47d5e6 Kenneth Lee 2018-09-03 174 struct device *pdev = mdev_parent_dev(mdev);
1e47d5e6 Kenneth Lee 2018-09-03 175
1e47d5e6 Kenneth Lee 2018-09-03 176 put_device(pdev);
1e47d5e6 Kenneth Lee 2018-09-03 177
1e47d5e6 Kenneth Lee 2018-09-03 @178 if (sdmdev->ops->put_queue);
1e47d5e6 Kenneth Lee 2018-09-03 179 sdmdev->ops->put_queue(q);
1e47d5e6 Kenneth Lee 2018-09-03 180
1e47d5e6 Kenneth Lee 2018-09-03 181 return 0;
1e47d5e6 Kenneth Lee 2018-09-03 182 }
1e47d5e6 Kenneth Lee 2018-09-03 183
---
0-DAY kernel test infrastructure Open Source Technology Center
https://lists.01.org/pipermail/kbuild-all Intel Corporation
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
quoted
From: Kenneth Lee <redacted>
WarpDrive is an accelerator framework to expose the hardware capabilities
directly to the user space. It makes use of the exist vfio and vfio-mdev
facilities. So the user application can send request and DMA to the
hardware without interaction with the kernel. This removes the latency
of syscall.
WarpDrive is the name for the whole framework. The component in kernel
is called SDMDEV, Share Domain Mediated Device. Driver driver exposes its
hardware resource by registering to SDMDEV as a VFIO-Mdev. So the user
library of WarpDrive can access it via VFIO interface.
The patchset contains document for the detail. Please refer to it for more
information.
This patchset is intended to be used with Jean Philippe Brucker's SVA
patch [1], which enables not only IO side page fault, but also PASID
support to IOMMU and VFIO.
With these features, WarpDrive can support non-pinned memory and
multi-process in the same accelerator device. We tested it in our SoC
integrated Accelerator (board ID: D06, Chip ID: HIP08). A reference work
tree can be found here: [2].
But it is not mandatory. This patchset is tested in the latest mainline
kernel without the SVA patches. So it supports only one process for each
accelerator.
We have noticed the IOMMU aware mdev RFC announced recently [3].
The IOMMU aware mdev has similar idea but different intention comparing to
WarpDrive. It intends to dedicate part of the hardware resource to a VM.
And the design is supposed to be used with Scalable I/O Virtualization.
While sdmdev is intended to share the hardware resource with a big amount
of processes. It just requires the hardware supporting address
translation per process (PCIE's PASID or ARM SMMU's substream ID).
But we don't see serious confliction on both design. We believe they can be
normalized as one.
So once again i do not understand why you are trying to do things
this way. Kernel already have tons of example of everything you
want to do without a new framework. Moreover i believe you are
confuse by VFIO. To me VFIO is for VM not to create general device
driver frame work.
VFIO is a userspace driver framework, the VM use case just happens to
be a rather prolific one. VFIO was never intended to be solely a VM
device interface and has several other userspace users, notably DPDK
and SPDK, an NVMe backend in QEMU, a userspace NVMe driver, a ruby
wrapper, and perhaps others that I'm not aware of. Whether vfio is
appropriate interface here might certainly still be a debatable topic,
but I would strongly disagree with your last sentence above. Thanks,
Alex
So here is your use case as i understand it. You have a device
with a limited number of command queues (can be just one) and in
some case it can support SVA/SVM (when hardware support it and it
is not disabled). Final requirement is being able to schedule cmds
from userspace without ioctl. All of this exists already exists
upstream in few device drivers.
So here is how every body else is doing it. Please explain why
this does not work.
1 Userspace open device file driver. Kernel device driver create
a context and associate it with on open. This context can be
uniq to the process and can bind hardware resources (like a
command queue) to the process.
2 Userspace bind/acquire a commands queue and initialize it with
an ioctl on the device file. Through that ioctl userspace can
be inform wether either SVA/SVM works for the device. If SVA/
SVM works then kernel device driver bind the process to the
device as part of this ioctl.
3 If SVM/SVA does not work userspace do an ioctl to create dma
buffer or something that does exactly the same thing.
4 Userspace mmap the command queue (mmap of the device file by
using informations gather at step 2)
5 Userspace can write commands into the queue it mapped
6 When userspace close the device file all resources are release
just like any existing device drivers.
Now if you want to create a device driver framework that expose
a device file with generic API for all of the above steps fine.
But it does not need to be part of VFIO whatsoever or explain
why.
Note that if IOMMU is fully disabled you probably want to block
userspace from being able to directly scheduling commands onto
the hardware as it would allow userspace to DMA anywhere and thus
would open the kernel to easy exploits. In this case you can still
keeps the same API as above and use page fault tricks to valid
commands written by userspace into fake commands ring. This will
be as slow or maybe even slower than ioctl but at least it allows
you to validate commands.
Cheers,
Jérôme
From: Kenneth Lee <hidden> Date: 2018-09-06 09:03:39
On Mon, Sep 03, 2018 at 10:55:57AM +0800, Lu Baolu wrote:
Date: Mon, 3 Sep 2018 10:55:57 +0800
From: Lu Baolu <baolu.lu@linux.intel.com>
To: Kenneth Lee <redacted>, Jonathan Corbet <corbet@lwn.net>,
Herbert Xu [off-list ref], "David S . Miller"
[off-list ref], Joerg Roedel [off-list ref], Alex Williamson
[off-list ref], Kenneth Lee [off-list ref], Hao
Fang [off-list ref], Zhou Wang [off-list ref], Zaibo Xu
[off-list ref], Philippe Ombredanne [off-list ref], Greg
Kroah-Hartman [off-list ref], Thomas Gleixner
[off-list ref], linux-doc@vger.kernel.org,
linux-kernel@vger.kernel.org, linux-crypto@vger.kernel.org,
iommu@lists.linux-foundation.org, kvm@vger.kernel.org,
linux-accelerators@lists.ozlabs.org, Sanjay Kumar
[off-list ref]
CC: linuxarm@huawei.com, baolu.lu@linux.intel.com
Subject: Re: [PATCH 3/7] vfio: add sdmdev support
User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:52.0) Gecko/20100101
Thunderbird/52.9.1
Message-ID: [off-list ref]
Hi,
On 09/03/2018 08:52 AM, Kenneth Lee wrote:
quoted
From: Kenneth Lee <redacted>
SDMDEV is "Share Domain Mdev". It is a vfio-mdev. But differ from
the general vfio-mdev, it shares its parent's IOMMU. If Multi-PASID
support is enabled in the IOMMU (not yet in the current kernel HEAD),
multiple process can share the IOMMU by different PASID. If it is not
support, only one process can share the IOMMU with the kernel driver.
If only for share domain purpose, I don't think it's necessary to create
a new device type.
Yes, if ONLY for share domain purpose. But we need also to share the interrupt.
quoted
Currently only the vfio type-1 driver is updated to make it to be aware
of.
Signed-off-by: Kenneth Lee <redacted>
Signed-off-by: Zaibo Xu <redacted>
Signed-off-by: Zhou Wang <wangzhou1@hisilicon.com>
---
drivers/vfio/Kconfig | 1 +
drivers/vfio/Makefile | 1 +
drivers/vfio/sdmdev/Kconfig | 10 +
drivers/vfio/sdmdev/Makefile | 3 +
drivers/vfio/sdmdev/vfio_sdmdev.c | 363 ++++++++++++++++++++++++++++++
drivers/vfio/vfio_iommu_type1.c | 151 ++++++++++++-
include/linux/vfio_sdmdev.h | 96 ++++++++
include/uapi/linux/vfio_sdmdev.h | 29 +++
8 files changed, 648 insertions(+), 6 deletions(-)
create mode 100644 drivers/vfio/sdmdev/Kconfig
create mode 100644 drivers/vfio/sdmdev/Makefile
create mode 100644 drivers/vfio/sdmdev/vfio_sdmdev.c
create mode 100644 include/linux/vfio_sdmdev.h
create mode 100644 include/uapi/linux/vfio_sdmdev.h
@@ -1327,6 +1330,109 @@ static bool vfio_iommu_has_sw_msi(struct iommu_group *group, phys_addr_t *base)returnret;}+/* return 0 if the device is not sdmdev.+*return1ifthedeviceissdmdev,thedatawillbeupdatedwithparent+*device'sgroup.+*return-errnoifothererror.+*/+staticintvfio_sdmdev_type(structdevice*dev,void*data)+{+structiommu_group**group=data;+structiommu_group*pgroup;+int(*_is_sdmdev)(structdevice*dev);+structdevice*pdev;+intret=1;++/* vfio_sdmdev module is not configurated */+_is_sdmdev=symbol_get(vfio_sdmdev_is_sdmdev);+if(!_is_sdmdev)+return0;++/* check if it belongs to vfio_sdmdev device */+if(!_is_sdmdev(dev)){+ret=0;+gotoout;+}++pdev=dev->parent;+pgroup=iommu_group_get(pdev);+if(!pgroup){+ret=-ENODEV;+gotoout;+}++if(group){+/* check if all parent devices is the same */+if(*group&&*group!=pgroup)+ret=-ENODEV;+else+*group=pgroup;+}++iommu_group_put(pgroup);++out:+symbol_put(vfio_sdmdev_is_sdmdev);++returnret;+}++/* return 0 or -errno */+staticintvfio_sdmdev_bus(structdevice*dev,void*data)+{+structbus_type**bus=data;++if(!dev->bus)+return-ENODEV;++/* ensure all devices has the same bus_type */+if(*bus&&*bus!=dev->bus)+return-EINVAL;++*bus=dev->bus;+return0;+}++/* return 0 means it is not sd group, 1 means it is, or -EXXX for error */+staticintvfio_iommu_type1_attach_sdgroup(structvfio_domain*domain,+structvfio_group*group,+structiommu_group*iommu_group)+{+intret;+structbus_type*pbus=NULL;+structiommu_group*pgroup=NULL;++ret=iommu_group_for_each_dev(iommu_group,&pgroup,+vfio_sdmdev_type);+if(ret<0)+gotoout;+elseif(ret>0){+domain->domain=iommu_group_share_domain(pgroup);+if(IS_ERR(domain->domain))+gotoout;+ret=iommu_group_for_each_dev(pgroup,&pbus,+vfio_sdmdev_bus);+if(ret<0)+gotoerr_with_share_domain;++if(pbus&&iommu_capable(pbus,IOMMU_CAP_CACHE_COHERENCY))+domain->prot|=IOMMU_CACHE;++group->parent_group=pgroup;+INIT_LIST_HEAD(&domain->group_list);+list_add(&group->next,&domain->group_list);++return1;+}
This doesn't match the function name. It only gets the domain from the
parent device. It hasn't been really attached.
@@ -1373,6 +1479,14 @@ static int vfio_iommu_type1_attach_group(void *iommu_data, if (mdev_bus) { if ((bus == mdev_bus) && !iommu_present(bus)) { symbol_put(mdev_bus_type);++ ret = vfio_iommu_type1_attach_sdgroup(domain, group,+ iommu_group);+ if (ret < 0)+ goto out_free;+ else if (ret > 0)+ goto replay_check;
Here you get the domain from the parent device and save it for later
use. The actual attaching is ignored.
I don't think this follows the philosophy of this function. It actually
make all devices in the group with the same bus type to share a single
domain.
I think the original logic here is:
1. Create a new vfio_domain along with a iommu_domain for the group attached to
the container
2. Try to match the vfio_domain with the domain list in the container. If there
is a match, free the created one and and reuse it, or add the new vfio_domain
to the list.
With this design, the same configuration to the IOMMU(unit) will be applied
only once.
For iommu_group that shares IOMMU with its parent, the configuration will never
be the same (The PASID will be different), so it is not necessary to merge them.
Further more, the parent domain might be a domain of type
IOMMU_DOMAIN_DMA. That will not be able to use as an
IOMMU_DOMAIN_UNMANAGED domain for iommu APIs.
Indeed, it should be checked when the domain is shared. Unmanaged domain should
not be used for sharing. I will update it in the future.
Best regards,
Lu Baolu
--
-Kenneth(Hisilicon)
================================================================================
本邮件及其附件含有华为公司的保密信息,仅限于发送给上面地址中列出的个人或群组。禁
止任何其他人以任何形式使用(包括但不限于全部或部分地泄露、复制、或散发)本邮件中
的信息。如果您错收了本邮件,请您立即电话或邮件通知发件人并删除本邮件!
This e-mail and its attachments contain confidential information from HUAWEI,
which is intended only for the person or entity whose address is listed above.
Any use of the
information contained herein in any way (including, but not limited to, total or
partial disclosure, reproduction, or dissemination) by persons other than the
intended
recipient(s) is prohibited. If you receive this e-mail in error, please notify
the sender by phone or email immediately and delete it!
@@ -1,4 +1,8 @@# SPDX-License-Identifier: GPL-2.0+configCRYPTO_DEV_HISILICON+tristate"Support for HISILICON CRYPTO ACCELERATOR"+help+EnablethistouseHisiliconHardwareAccelerators
Accelerators.
Thanks, will change it in next version.
--
~Randy
--
-Kenneth(Hisilicon)
================================================================================
本邮件及其附件含有华为公司的保密信息,仅限于发送给上面地址中列出的个人或群组。禁
止任何其他人以任何形式使用(包括但不限于全部或部分地泄露、复制、或散发)本邮件中
的信息。如果您错收了本邮件,请您立即电话或邮件通知发件人并删除本邮件!
This e-mail and its attachments contain confidential information from HUAWEI,
which is intended only for the person or entity whose address is listed above.
Any use of the
information contained herein in any way (including, but not limited to, total or
partial disclosure, reproduction, or dissemination) by persons other than the
intended
recipient(s) is prohibited. If you receive this e-mail in error, please notify
the sender by phone or email immediately and delete it!
At a minimum:
Enable this to enable the SDMDEV,
although that could be done better. Maybe just:
Enable the SDMDEV "shared IOMMU Domain Mediated Device"
quoted
+ interface for all Hisilicon accelerators if they can. The SDMDEV
probably drop "if they can": accelerators. The SDMDEV interface
quoted
+ enable the WarpDrive user space accelerator driver to access the
enables
Thank you, will change them all in the coming version.
quoted
+ hardware function directly.
+
--
~Randy
--
-Kenneth(Hisilicon)
================================================================================
本邮件及其附件含有华为公司的保密信息,仅限于发送给上面地址中列出的个人或群组。禁
止任何其他人以任何形式使用(包括但不限于全部或部分地泄露、复制、或散发)本邮件中
的信息。如果您错收了本邮件,请您立即电话或邮件通知发件人并删除本邮件!
This e-mail and its attachments contain confidential information from HUAWEI,
which is intended only for the person or entity whose address is listed above.
Any use of the
information contained herein in any way (including, but not limited to, total or
partial disclosure, reproduction, or dissemination) by persons other than the
intended
recipient(s) is prohibited. If you receive this e-mail in error, please notify
the sender by phone or email immediately and delete it!
From: Kenneth Lee <hidden> Date: 2018-09-06 09:12:34
On Sun, Sep 02, 2018 at 07:25:12PM -0700, Randy Dunlap wrote:
Date: Sun, 2 Sep 2018 19:25:12 -0700
From: Randy Dunlap <rdunlap@infradead.org>
To: Kenneth Lee <redacted>, Jonathan Corbet <corbet@lwn.net>,
Herbert Xu [off-list ref], "David S . Miller"
[off-list ref], Joerg Roedel [off-list ref], Alex Williamson
[off-list ref], Kenneth Lee [off-list ref], Hao
Fang [off-list ref], Zhou Wang [off-list ref], Zaibo Xu
[off-list ref], Philippe Ombredanne [off-list ref], Greg
Kroah-Hartman [off-list ref], Thomas Gleixner
[off-list ref], linux-doc@vger.kernel.org,
linux-kernel@vger.kernel.org, linux-crypto@vger.kernel.org,
iommu@lists.linux-foundation.org, kvm@vger.kernel.org,
linux-accelerators@lists.ozlabs.org, Lu Baolu [off-list ref],
Sanjay Kumar [off-list ref]
CC: linuxarm@huawei.com
Subject: Re: [PATCH 7/7] vfio/sdmdev: add user sample
User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:52.0) Gecko/20100101
Thunderbird/52.9.1
Message-ID: [off-list ref]
On 09/02/2018 05:52 PM, Kenneth Lee wrote:
quoted
From: Kenneth Lee <redacted>
This is the sample code to demostrate how WrapDrive user application
should be.
It contains:
1. wd.[ch], the common library to provide WrapDrive interface.
WarpDrive
quoted
2. drv/*, the user driver to access the hardware upon spimdev
3. test/*, the test application to use WrapDrive interface to access the
@@ -0,0 +1,32 @@+WD User Land Demonstration+==========================++This directory contains some applications and libraries to demonstrate how a++WrapDrive application can be constructed.
WarpDrive
quoted
+++As a demo, we try to make it simple and clear for understanding. It is not++supposed to be used in business scenario.+++The directory contains the following elements:++wd.[ch]+ A demonstration WrapDrive fundamental library which wraps the basic
WarpDrive
quoted
+ operations to the WrapDrive-ed device.
WarpDrive
quoted
++wd_adapter.[ch]+ User driver adaptor for wd.[ch]++wd_utils.[ch]+ Some utitlities function used by WD and its drivers++drv/*+ User drivers. It helps to fulfill the semantic of wd.[ch] for+ particular hardware++test/*+ Test applications to use the wrapdrive library
warpdrive
--
~Randy
Thank you, will change them all in the coming version.
--
-Kenneth(Hisilicon)
================================================================================
本邮件及其附件含有华为公司的保密信息,仅限于发送给上面地址中列出的个人或群组。禁
止任何其他人以任何形式使用(包括但不限于全部或部分地泄露、复制、或散发)本邮件中
的信息。如果您错收了本邮件,请您立即电话或邮件通知发件人并删除本邮件!
This e-mail and its attachments contain confidential information from HUAWEI,
which is intended only for the person or entity whose address is listed above.
Any use of the
information contained herein in any way (including, but not limited to, total or
partial disclosure, reproduction, or dissemination) by persons other than the
intended
recipient(s) is prohibited. If you receive this e-mail in error, please notify
the sender by phone or email immediately and delete it!
From: Kenneth Lee <hidden> Date: 2018-09-06 09:13:34
On Mon, Sep 03, 2018 at 10:32:16AM +0800, Lu Baolu wrote:
Date: Mon, 3 Sep 2018 10:32:16 +0800
From: Lu Baolu <baolu.lu@linux.intel.com>
To: Kenneth Lee <redacted>, Jonathan Corbet <corbet@lwn.net>,
Herbert Xu [off-list ref], "David S . Miller"
[off-list ref], Joerg Roedel [off-list ref], Alex Williamson
[off-list ref], Kenneth Lee [off-list ref], Hao
Fang [off-list ref], Zhou Wang [off-list ref], Zaibo Xu
[off-list ref], Philippe Ombredanne [off-list ref], Greg
Kroah-Hartman [off-list ref], Thomas Gleixner
[off-list ref], linux-doc@vger.kernel.org,
linux-kernel@vger.kernel.org, linux-crypto@vger.kernel.org,
iommu@lists.linux-foundation.org, kvm@vger.kernel.org,
linux-accelerators@lists.ozlabs.org, Sanjay Kumar
[off-list ref]
CC: baolu.lu@linux.intel.com, linuxarm@huawei.com
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:52.0) Gecko/20100101
Thunderbird/52.9.1
Message-ID: [off-list ref]
Hi,
On 09/03/2018 08:51 AM, Kenneth Lee wrote:
quoted
From: Kenneth Lee <redacted>
WarpDrive is an accelerator framework to expose the hardware capabilities
directly to the user space. It makes use of the exist vfio and vfio-mdev
facilities. So the user application can send request and DMA to the
hardware without interaction with the kernel. This removes the latency
of syscall.
WarpDrive is the name for the whole framework. The component in kernel
is called SDMDEV, Share Domain Mediated Device. Driver driver exposes its
hardware resource by registering to SDMDEV as a VFIO-Mdev. So the user
library of WarpDrive can access it via VFIO interface.
The patchset contains document for the detail. Please refer to it for more
information.
This patchset is intended to be used with Jean Philippe Brucker's SVA
patch [1], which enables not only IO side page fault, but also PASID
support to IOMMU and VFIO.
With these features, WarpDrive can support non-pinned memory and
multi-process in the same accelerator device. We tested it in our SoC
integrated Accelerator (board ID: D06, Chip ID: HIP08). A reference work
tree can be found here: [2].
But it is not mandatory. This patchset is tested in the latest mainline
kernel without the SVA patches. So it supports only one process for each
accelerator.
We have noticed the IOMMU aware mdev RFC announced recently [3].
The IOMMU aware mdev has similar idea but different intention comparing to
WarpDrive. It intends to dedicate part of the hardware resource to a VM.
And the design is supposed to be used with Scalable I/O Virtualization.
While sdmdev is intended to share the hardware resource with a big amount
of processes. It just requires the hardware supporting address
translation per process (PCIE's PASID or ARM SMMU's substream ID).
But we don't see serious confliction on both design. We believe they can be
normalized as one.
The patch 1 is document of the framework. The patch 2 and 3 add sdmdev
support. The patch 4, 5 and 6 is drivers for Hislicon's ZIP Accelerator
which is registered to both crypto and warpdrive(sdmdev) and can be
used from kernel or user space at the same time. The patch 7 is a user
space sample demonstrating how WarpDrive works.
Change History:
V2 changed from V1:
1. Change kernel framework name from SPIMDEV (Share Parent IOMMU
Mdev) to SDMDEV (Share Domain Mdev).
2. Allocate Hardware Resource when a new mdev is created (While
it is allocated when the mdev is openned)
3. Unmap pages from the shared domain when the sdmdev iommu group is
detached. (This procedure is necessary, but missed in V1)
4. Update document accordingly.
5. Rebase to the latest kernel (4.19.0-rc1)
According the review comment on RFCv1, We did try to use dma-buf
as back end of WarpDrive. It can work properly with the current
solution [4], but it cannot make use of process's
own memory address space directly. This is important to many
acceleration scenario. So dma-buf will be taken as a backup
alternative for noiommu scenario, it will be added in the future
version.
Refernces:
[1] https://www.spinics.net/lists/kernel/msg2651481.html
[2] https://github.com/Kenneth-Lee/linux-kernel-warpdrive/tree/warpdrive-sva-v0.5
[3] https://lkml.org/lkml/2018/7/22/34
--
-Kenneth(Hisilicon)
================================================================================
本邮件及其附件含有华为公司的保密信息,仅限于发送给上面地址中列出的个人或群组。禁
止任何其他人以任何形式使用(包括但不限于全部或部分地泄露、复制、或散发)本邮件中
的信息。如果您错收了本邮件,请您立即电话或邮件通知发件人并删除本邮件!
This e-mail and its attachments contain confidential information from HUAWEI,
which is intended only for the person or entity whose address is listed above.
Any use of the
information contained herein in any way (including, but not limited to, total or
partial disclosure, reproduction, or dissemination) by persons other than the
intended
recipient(s) is prohibited. If you receive this e-mail in error, please notify
the sender by phone or email immediately and delete it!
From: Kenneth Lee <hidden> Date: 2018-09-06 09:47:33
On Tue, Sep 04, 2018 at 10:15:09AM -0600, Alex Williamson wrote:
Date: Tue, 4 Sep 2018 10:15:09 -0600
From: Alex Williamson <redacted>
To: Jerome Glisse <redacted>
CC: Kenneth Lee <redacted>, Jonathan Corbet <corbet@lwn.net>,
Herbert Xu [off-list ref], "David S . Miller"
[off-list ref], Joerg Roedel [off-list ref], Kenneth Lee
[off-list ref], Hao Fang [off-list ref], Zhou Wang
[off-list ref], Zaibo Xu [off-list ref], Philippe
Ombredanne [off-list ref], Greg Kroah-Hartman
[off-list ref], Thomas Gleixner [off-list ref],
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-crypto@vger.kernel.org, iommu@lists.linux-foundation.org,
kvm@vger.kernel.org, linux-accelerators@lists.ozlabs.org, Lu Baolu
[off-list ref], Sanjay Kumar [off-list ref],
linuxarm@huawei.com
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
Message-ID: [off-list ref]
On Tue, 4 Sep 2018 11:00:19 -0400
Jerome Glisse [off-list ref] wrote:
quoted
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
quoted
From: Kenneth Lee <redacted>
WarpDrive is an accelerator framework to expose the hardware capabilities
directly to the user space. It makes use of the exist vfio and vfio-mdev
facilities. So the user application can send request and DMA to the
hardware without interaction with the kernel. This removes the latency
of syscall.
WarpDrive is the name for the whole framework. The component in kernel
is called SDMDEV, Share Domain Mediated Device. Driver driver exposes its
hardware resource by registering to SDMDEV as a VFIO-Mdev. So the user
library of WarpDrive can access it via VFIO interface.
The patchset contains document for the detail. Please refer to it for more
information.
This patchset is intended to be used with Jean Philippe Brucker's SVA
patch [1], which enables not only IO side page fault, but also PASID
support to IOMMU and VFIO.
With these features, WarpDrive can support non-pinned memory and
multi-process in the same accelerator device. We tested it in our SoC
integrated Accelerator (board ID: D06, Chip ID: HIP08). A reference work
tree can be found here: [2].
But it is not mandatory. This patchset is tested in the latest mainline
kernel without the SVA patches. So it supports only one process for each
accelerator.
We have noticed the IOMMU aware mdev RFC announced recently [3].
The IOMMU aware mdev has similar idea but different intention comparing to
WarpDrive. It intends to dedicate part of the hardware resource to a VM.
And the design is supposed to be used with Scalable I/O Virtualization.
While sdmdev is intended to share the hardware resource with a big amount
of processes. It just requires the hardware supporting address
translation per process (PCIE's PASID or ARM SMMU's substream ID).
But we don't see serious confliction on both design. We believe they can be
normalized as one.
So once again i do not understand why you are trying to do things
this way. Kernel already have tons of example of everything you
want to do without a new framework. Moreover i believe you are
confuse by VFIO. To me VFIO is for VM not to create general device
driver frame work.
VFIO is a userspace driver framework, the VM use case just happens to
be a rather prolific one. VFIO was never intended to be solely a VM
device interface and has several other userspace users, notably DPDK
and SPDK, an NVMe backend in QEMU, a userspace NVMe driver, a ruby
wrapper, and perhaps others that I'm not aware of. Whether vfio is
appropriate interface here might certainly still be a debatable topic,
but I would strongly disagree with your last sentence above. Thanks,
Alex
Yes, that is also my standpoint here.
quoted
So here is your use case as i understand it. You have a device
with a limited number of command queues (can be just one) and in
some case it can support SVA/SVM (when hardware support it and it
is not disabled). Final requirement is being able to schedule cmds
from userspace without ioctl. All of this exists already exists
upstream in few device drivers.
So here is how every body else is doing it. Please explain why
this does not work.
1 Userspace open device file driver. Kernel device driver create
a context and associate it with on open. This context can be
uniq to the process and can bind hardware resources (like a
command queue) to the process.
2 Userspace bind/acquire a commands queue and initialize it with
an ioctl on the device file. Through that ioctl userspace can
be inform wether either SVA/SVM works for the device. If SVA/
SVM works then kernel device driver bind the process to the
device as part of this ioctl.
3 If SVM/SVA does not work userspace do an ioctl to create dma
buffer or something that does exactly the same thing.
4 Userspace mmap the command queue (mmap of the device file by
using informations gather at step 2)
5 Userspace can write commands into the queue it mapped
6 When userspace close the device file all resources are release
just like any existing device drivers.
Hi, Jerome,
Just one thing, as I said in the cover letter, dma-buf requires the application
to use memory created by the driver for DMA. I did try the dma-buf way in
WrapDrive (refer to [4] in the cover letter), it is a good backup for NOIOMMU
mode or we cannot solve the problem in VFIO.
But, in many of my application scenario, the application already has some memory
in hand, maybe allocated by the framework or libraries. Anyway, they don't get
memory from my library, and they pass the poiter for data operation. And they
may also have pointer in the buffer. Those pointer may be used by the
accelerator. So I need hardware fully share the address space with the
application. That is what dmabuf cannot do.
quoted
Now if you want to create a device driver framework that expose
a device file with generic API for all of the above steps fine.
But it does not need to be part of VFIO whatsoever or explain
why.
Note that if IOMMU is fully disabled you probably want to block
userspace from being able to directly scheduling commands onto
the hardware as it would allow userspace to DMA anywhere and thus
would open the kernel to easy exploits. In this case you can still
keeps the same API as above and use page fault tricks to valid
commands written by userspace into fake commands ring. This will
be as slow or maybe even slower than ioctl but at least it allows
you to validate commands.
Cheers,
Jérôme
--
-Kenneth(Hisilicon)
================================================================================
本邮件及其附件含有华为公司的保密信息,仅限于发送给上面地址中列出的个人或群组。禁
止任何其他人以任何形式使用(包括但不限于全部或部分地泄露、复制、或散发)本邮件中
的信息。如果您错收了本邮件,请您立即电话或邮件通知发件人并删除本邮件!
This e-mail and its attachments contain confidential information from HUAWEI,
which is intended only for the person or entity whose address is listed above.
Any use of the
information contained herein in any way (including, but not limited to, total or
partial disclosure, reproduction, or dissemination) by persons other than the
intended
recipient(s) is prohibited. If you receive this e-mail in error, please notify
the sender by phone or email immediately and delete it!
On Thu, Sep 06, 2018 at 05:45:32PM +0800, Kenneth Lee wrote:
On Tue, Sep 04, 2018 at 10:15:09AM -0600, Alex Williamson wrote:
quoted
Date: Tue, 4 Sep 2018 10:15:09 -0600
From: Alex Williamson <redacted>
To: Jerome Glisse <redacted>
CC: Kenneth Lee <redacted>, Jonathan Corbet <corbet@lwn.net>,
Herbert Xu [off-list ref], "David S . Miller"
[off-list ref], Joerg Roedel [off-list ref], Kenneth Lee
[off-list ref], Hao Fang [off-list ref], Zhou Wang
[off-list ref], Zaibo Xu [off-list ref], Philippe
Ombredanne [off-list ref], Greg Kroah-Hartman
[off-list ref], Thomas Gleixner [off-list ref],
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-crypto@vger.kernel.org, iommu@lists.linux-foundation.org,
kvm@vger.kernel.org, linux-accelerators@lists.ozlabs.org, Lu Baolu
[off-list ref], Sanjay Kumar [off-list ref],
linuxarm@huawei.com
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
Message-ID: [off-list ref]
On Tue, 4 Sep 2018 11:00:19 -0400
Jerome Glisse [off-list ref] wrote:
quoted
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
quoted
From: Kenneth Lee <redacted>
WarpDrive is an accelerator framework to expose the hardware capabilities
directly to the user space. It makes use of the exist vfio and vfio-mdev
facilities. So the user application can send request and DMA to the
hardware without interaction with the kernel. This removes the latency
of syscall.
WarpDrive is the name for the whole framework. The component in kernel
is called SDMDEV, Share Domain Mediated Device. Driver driver exposes its
hardware resource by registering to SDMDEV as a VFIO-Mdev. So the user
library of WarpDrive can access it via VFIO interface.
The patchset contains document for the detail. Please refer to it for more
information.
This patchset is intended to be used with Jean Philippe Brucker's SVA
patch [1], which enables not only IO side page fault, but also PASID
support to IOMMU and VFIO.
With these features, WarpDrive can support non-pinned memory and
multi-process in the same accelerator device. We tested it in our SoC
integrated Accelerator (board ID: D06, Chip ID: HIP08). A reference work
tree can be found here: [2].
But it is not mandatory. This patchset is tested in the latest mainline
kernel without the SVA patches. So it supports only one process for each
accelerator.
We have noticed the IOMMU aware mdev RFC announced recently [3].
The IOMMU aware mdev has similar idea but different intention comparing to
WarpDrive. It intends to dedicate part of the hardware resource to a VM.
And the design is supposed to be used with Scalable I/O Virtualization.
While sdmdev is intended to share the hardware resource with a big amount
of processes. It just requires the hardware supporting address
translation per process (PCIE's PASID or ARM SMMU's substream ID).
But we don't see serious confliction on both design. We believe they can be
normalized as one.
So once again i do not understand why you are trying to do things
this way. Kernel already have tons of example of everything you
want to do without a new framework. Moreover i believe you are
confuse by VFIO. To me VFIO is for VM not to create general device
driver frame work.
VFIO is a userspace driver framework, the VM use case just happens to
be a rather prolific one. VFIO was never intended to be solely a VM
device interface and has several other userspace users, notably DPDK
and SPDK, an NVMe backend in QEMU, a userspace NVMe driver, a ruby
wrapper, and perhaps others that I'm not aware of. Whether vfio is
appropriate interface here might certainly still be a debatable topic,
but I would strongly disagree with your last sentence above. Thanks,
Alex
Yes, that is also my standpoint here.
quoted
quoted
So here is your use case as i understand it. You have a device
with a limited number of command queues (can be just one) and in
some case it can support SVA/SVM (when hardware support it and it
is not disabled). Final requirement is being able to schedule cmds
from userspace without ioctl. All of this exists already exists
upstream in few device drivers.
So here is how every body else is doing it. Please explain why
this does not work.
1 Userspace open device file driver. Kernel device driver create
a context and associate it with on open. This context can be
uniq to the process and can bind hardware resources (like a
command queue) to the process.
2 Userspace bind/acquire a commands queue and initialize it with
an ioctl on the device file. Through that ioctl userspace can
be inform wether either SVA/SVM works for the device. If SVA/
SVM works then kernel device driver bind the process to the
device as part of this ioctl.
3 If SVM/SVA does not work userspace do an ioctl to create dma
buffer or something that does exactly the same thing.
4 Userspace mmap the command queue (mmap of the device file by
using informations gather at step 2)
5 Userspace can write commands into the queue it mapped
6 When userspace close the device file all resources are release
just like any existing device drivers.
Hi, Jerome,
Just one thing, as I said in the cover letter, dma-buf requires the application
to use memory created by the driver for DMA. I did try the dma-buf way in
WrapDrive (refer to [4] in the cover letter), it is a good backup for NOIOMMU
mode or we cannot solve the problem in VFIO.
But, in many of my application scenario, the application already has some memory
in hand, maybe allocated by the framework or libraries. Anyway, they don't get
memory from my library, and they pass the poiter for data operation. And they
may also have pointer in the buffer. Those pointer may be used by the
accelerator. So I need hardware fully share the address space with the
application. That is what dmabuf cannot do.
dmabuf can do that ... it is call uptr you can look at i915 for
instance. Still this does not answer my question above, why do
you need to be in VFIO to do any of the above thing ? Kernel has
tons of examples that does all of the above and are not in VFIO
(including usinng existing user pointer with device).
Cheers,
Jérôme
From: Randy Dunlap <rdunlap@infradead.org> Date: 2018-09-06 18:36:50
Hi,
On 09/02/2018 05:51 PM, Kenneth Lee wrote:
From: Kenneth Lee <redacted>
WarpDrive is a common user space accelerator framework. Its main component
in Kernel is called sdmdev, Share Domain Mediated Device. It exposes
the hardware capabilities to the user space via vfio-mdev. So processes in
user land can obtain a "queue" by open the device and direct access the
hardware MMIO space or do DMA operation via VFIO interface.
WarpDrive is intended to be used with Jean Philippe Brucker's SVA
patchset to support multi-process. But This is not a must. Without the
SVA patches, WarpDrive can still work for one process for every hardware
device.
This patch add detail documents for the framework.
Signed-off-by: Kenneth Lee <redacted>
---
Documentation/00-INDEX | 2 +
Documentation/warpdrive/warpdrive.rst | 100 ++++
Documentation/warpdrive/wd-arch.svg | 728 ++++++++++++++++++++++++++
3 files changed, 830 insertions(+)
create mode 100644 Documentation/warpdrive/warpdrive.rst
create mode 100644 Documentation/warpdrive/wd-arch.svg
@@ -0,0 +1,100 @@+Introduction of WarpDrive+=========================++*WarpDrive* is a general accelerator framework for user space. It intends to+provide interface for the user process to send request to hardware+accelerator without heavy user-kernel interaction cost.++The *WarpDrive* user library is supposed to provide a pipe-based API, such as:
Do you say "is supposed to" because it doesn't do that (yet)?
Or you could just change that to say:
The WarpDrive user library provides a pipe-based API, such as:
quoted hunk
+ ::+ int wd_request_queue(struct wd_queue *q);+ void wd_release_queue(struct wd_queue *q);++ int wd_send(struct wd_queue *q, void *req);+ int wd_recv(struct wd_queue *q, void **req);+ int wd_recv_sync(struct wd_queue *q, void **req);+ int wd_flush(struct wd_queue *q);++*wd_request_queue* creates the pipe connection, *queue*, between the+application and the hardware. The application sends request and pulls the+answer back by asynchronized wd_send/wd_recv, which directly interact with the+hardware (by MMIO or share memory) without syscall.++*WarpDrive* maintains a unified application address space among all involved+accelerators. With the following APIs: ::
Seems like an extra '.' there. How about:
accelerators with the following APIs: ::
quoted hunk
++ int wd_mem_share(struct wd_queue *q, const void *addr,+ size_t size, int flags);+ void wd_mem_unshare(struct wd_queue *q, const void *addr, size_t size);++The referred process space shared by these APIs can be directly referred by the+hardware. The process can also dedicate its whole process space with flags,+*WD_SHARE_ALL* (not in this patch yet).++The name *WarpDrive* is simply a cool and general name meaning the framework+makes the application faster. As it will be explained in this text later, the+facility in kernel is called *SDMDEV*, namely "Share Domain Mediated Device".+++How does it work+================++*WarpDrive* is built upon *VFIO-MDEV*. The queue is wrapped as *mdev* in VFIO.+So memory sharing can be done via standard VFIO standard DMA interface.++The architecture is illustrated as follow figure:++.. image:: wd-arch.svg+ :alt: WarpDrive Architecture++Accelerator driver shares its capability via *SDMDEV* API: ::++ vfio_sdmdev_register(struct vfio_sdmdev *sdmdev);+ vfio_sdmdev_unregister(struct vfio_sdmdev *sdmdev);+ vfio_sdmdev_wake_up(struct spimdev_queue *q);++*vfio_sdmdev_register* is a helper function to register the hardware to the+*VFIO_MDEV* framework. The queue creation is done by *mdev* creation interface.++*WarpDrive* User library mmap the mdev to access its mmio space and shared
s/mmio/MMIO/
quoted hunk
+memory. Request can be sent to, or receive from, hardware in this mmap-ed+space until the queue is full or empty.++The user library can wait on the queue by ioctl(VFIO_SDMDEV_CMD_WAIT) the mdev+if the queue is full or empty. If the queue status is changed, the hardware+driver use *vfio_sdmdev_wake_up* to wake up the waiting process.+++Multiple processes support+==========================++In the latest mainline kernel (4.18) when this document is written,+multi-process is not supported in VFIO yet.++Jean Philippe Brucker has a patchset to enable it[1]_. We have tested it+with our hardware (which is known as *D06*). It works well. *WarpDrive* rely+on them to support multiple processes. If it is not enabled, *WarpDrive* can+still work, but it support only one mdev for a process, which will share the+same io map table with kernel. (But it is not going to be a security problem,+since the user application cannot access the kernel address space)++When multiprocess is support, mdev can be created based on how many+hardware resource (queue) is available. Because the VFIO framework accepts only+one open from one mdev iommu_group. Mdev become the smallest unit for process+to use queue. And the mdev will not be released if the user process exist. So+it will need a resource agent to manage the mdev allocation for the user+process. This is not in this document's range.+++Legacy Mode Support+===================+For the hardware on which IOMMU is not support, WarpDrive can run on *NOIOMMU*+mode. That require some update to the mdev driver, which is not included in+this version yet.+++References+==========+.. [1] https://patchwork.kernel.org/patch/10394851/++.. vim: tw=78
From: Kenneth Lee <hidden> Date: 2018-09-07 02:23:18
On Thu, Sep 06, 2018 at 11:36:36AM -0700, Randy Dunlap wrote:
Date: Thu, 6 Sep 2018 11:36:36 -0700
From: Randy Dunlap <rdunlap@infradead.org>
To: Kenneth Lee <redacted>, Jonathan Corbet <corbet@lwn.net>,
Herbert Xu [off-list ref], "David S . Miller"
[off-list ref], Joerg Roedel [off-list ref], Alex Williamson
[off-list ref], Kenneth Lee [off-list ref], Hao
Fang [off-list ref], Zhou Wang [off-list ref], Zaibo Xu
[off-list ref], Philippe Ombredanne [off-list ref], Greg
Kroah-Hartman [off-list ref], Thomas Gleixner
[off-list ref], linux-doc@vger.kernel.org,
linux-kernel@vger.kernel.org, linux-crypto@vger.kernel.org,
iommu@lists.linux-foundation.org, kvm@vger.kernel.org,
linux-accelerators@lists.ozlabs.org, Lu Baolu [off-list ref],
Sanjay Kumar [off-list ref]
CC: linuxarm@huawei.com
Subject: Re: [PATCH 1/7] vfio/sdmdev: Add documents for WarpDrive framework
User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:52.0) Gecko/20100101
Thunderbird/52.9.1
Message-ID: [off-list ref]
Hi,
On 09/02/2018 05:51 PM, Kenneth Lee wrote:
quoted
From: Kenneth Lee <redacted>
WarpDrive is a common user space accelerator framework. Its main component
in Kernel is called sdmdev, Share Domain Mediated Device. It exposes
the hardware capabilities to the user space via vfio-mdev. So processes in
user land can obtain a "queue" by open the device and direct access the
hardware MMIO space or do DMA operation via VFIO interface.
WarpDrive is intended to be used with Jean Philippe Brucker's SVA
patchset to support multi-process. But This is not a must. Without the
SVA patches, WarpDrive can still work for one process for every hardware
device.
This patch add detail documents for the framework.
Signed-off-by: Kenneth Lee <redacted>
---
Documentation/00-INDEX | 2 +
Documentation/warpdrive/warpdrive.rst | 100 ++++
Documentation/warpdrive/wd-arch.svg | 728 ++++++++++++++++++++++++++
3 files changed, 830 insertions(+)
create mode 100644 Documentation/warpdrive/warpdrive.rst
create mode 100644 Documentation/warpdrive/wd-arch.svg
@@ -0,0 +1,100 @@+Introduction of WarpDrive+=========================++*WarpDrive* is a general accelerator framework for user space. It intends to+provide interface for the user process to send request to hardware+accelerator without heavy user-kernel interaction cost.++The *WarpDrive* user library is supposed to provide a pipe-based API, such as:
Do you say "is supposed to" because it doesn't do that (yet)?
Or you could just change that to say:
The WarpDrive user library provides a pipe-based API, such as:
Actually, I tried to say it can be defined like this. But people can choose
other implementation with the same kernel API.
I will say it explicitly in the future version. Thank you.
quoted
+ ::+ int wd_request_queue(struct wd_queue *q);+ void wd_release_queue(struct wd_queue *q);++ int wd_send(struct wd_queue *q, void *req);+ int wd_recv(struct wd_queue *q, void **req);+ int wd_recv_sync(struct wd_queue *q, void **req);+ int wd_flush(struct wd_queue *q);++*wd_request_queue* creates the pipe connection, *queue*, between the+application and the hardware. The application sends request and pulls the+answer back by asynchronized wd_send/wd_recv, which directly interact with the+hardware (by MMIO or share memory) without syscall.++*WarpDrive* maintains a unified application address space among all involved+accelerators. With the following APIs: ::
Seems like an extra '.' there. How about:
accelerators with the following APIs: ::
Err, the "with..." clause belong to the following "The referred process
space...".
quoted
++ int wd_mem_share(struct wd_queue *q, const void *addr,+ size_t size, int flags);+ void wd_mem_unshare(struct wd_queue *q, const void *addr, size_t size);++The referred process space shared by these APIs can be directly referred by the+hardware. The process can also dedicate its whole process space with flags,+*WD_SHARE_ALL* (not in this patch yet).++The name *WarpDrive* is simply a cool and general name meaning the framework+makes the application faster. As it will be explained in this text later, the+facility in kernel is called *SDMDEV*, namely "Share Domain Mediated Device".+++How does it work+================++*WarpDrive* is built upon *VFIO-MDEV*. The queue is wrapped as *mdev* in VFIO.+So memory sharing can be done via standard VFIO standard DMA interface.++The architecture is illustrated as follow figure:++.. image:: wd-arch.svg+ :alt: WarpDrive Architecture++Accelerator driver shares its capability via *SDMDEV* API: ::++ vfio_sdmdev_register(struct vfio_sdmdev *sdmdev);+ vfio_sdmdev_unregister(struct vfio_sdmdev *sdmdev);+ vfio_sdmdev_wake_up(struct spimdev_queue *q);++*vfio_sdmdev_register* is a helper function to register the hardware to the+*VFIO_MDEV* framework. The queue creation is done by *mdev* creation interface.++*WarpDrive* User library mmap the mdev to access its mmio space and shared
s/mmio/MMIO/
quoted
+memory. Request can be sent to, or receive from, hardware in this mmap-ed+space until the queue is full or empty.++The user library can wait on the queue by ioctl(VFIO_SDMDEV_CMD_WAIT) the mdev+if the queue is full or empty. If the queue status is changed, the hardware+driver use *vfio_sdmdev_wake_up* to wake up the waiting process.+++Multiple processes support+==========================++In the latest mainline kernel (4.18) when this document is written,+multi-process is not supported in VFIO yet.++Jean Philippe Brucker has a patchset to enable it[1]_. We have tested it+with our hardware (which is known as *D06*). It works well. *WarpDrive* rely+on them to support multiple processes. If it is not enabled, *WarpDrive* can+still work, but it support only one mdev for a process, which will share the+same io map table with kernel. (But it is not going to be a security problem,+since the user application cannot access the kernel address space)++When multiprocess is support, mdev can be created based on how many+hardware resource (queue) is available. Because the VFIO framework accepts only+one open from one mdev iommu_group. Mdev become the smallest unit for process+to use queue. And the mdev will not be released if the user process exist. So+it will need a resource agent to manage the mdev allocation for the user+process. This is not in this document's range.+++Legacy Mode Support+===================+For the hardware on which IOMMU is not support, WarpDrive can run on *NOIOMMU*+mode. That require some update to the mdev driver, which is not included in+this version yet.+++References+==========+.. [1] https://patchwork.kernel.org/patch/10394851/++.. vim: tw=78
thanks,
--
~Randy
--
-Kenneth(Hisilicon)
================================================================================
本邮件及其附件含有华为公司的保密信息,仅限于发送给上面地址中列出的个人或群组。禁
止任何其他人以任何形式使用(包括但不限于全部或部分地泄露、复制、或散发)本邮件中
的信息。如果您错收了本邮件,请您立即电话或邮件通知发件人并删除本邮件!
This e-mail and its attachments contain confidential information from HUAWEI,
which is intended only for the person or entity whose address is listed above.
Any use of the
information contained herein in any way (including, but not limited to, total or
partial disclosure, reproduction, or dissemination) by persons other than the
intended
recipient(s) is prohibited. If you receive this e-mail in error, please notify
the sender by phone or email immediately and delete it!
From: Kenneth Lee <hidden> Date: 2018-09-07 04:03:38
On Thu, Sep 06, 2018 at 09:31:33AM -0400, Jerome Glisse wrote:
Date: Thu, 6 Sep 2018 09:31:33 -0400
From: Jerome Glisse <redacted>
To: Kenneth Lee <redacted>
CC: Alex Williamson <redacted>, Kenneth Lee
[off-list ref], Jonathan Corbet [off-list ref], Herbert Xu
[off-list ref], "David S . Miller" [off-list ref],
Joerg Roedel [off-list ref], Hao Fang [off-list ref], Zhou Wang
[off-list ref], Zaibo Xu [off-list ref], Philippe
Ombredanne [off-list ref], Greg Kroah-Hartman
[off-list ref], Thomas Gleixner [off-list ref],
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-crypto@vger.kernel.org, iommu@lists.linux-foundation.org,
kvm@vger.kernel.org, linux-accelerators@lists.ozlabs.org, Lu Baolu
[off-list ref], Sanjay Kumar [off-list ref],
linuxarm@huawei.com
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
User-Agent: Mutt/1.10.0 (2018-05-17)
Message-ID: [off-list ref]
On Thu, Sep 06, 2018 at 05:45:32PM +0800, Kenneth Lee wrote:
quoted
On Tue, Sep 04, 2018 at 10:15:09AM -0600, Alex Williamson wrote:
quoted
Date: Tue, 4 Sep 2018 10:15:09 -0600
From: Alex Williamson <redacted>
To: Jerome Glisse <redacted>
CC: Kenneth Lee <redacted>, Jonathan Corbet <corbet@lwn.net>,
Herbert Xu [off-list ref], "David S . Miller"
[off-list ref], Joerg Roedel [off-list ref], Kenneth Lee
[off-list ref], Hao Fang [off-list ref], Zhou Wang
[off-list ref], Zaibo Xu [off-list ref], Philippe
Ombredanne [off-list ref], Greg Kroah-Hartman
[off-list ref], Thomas Gleixner [off-list ref],
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-crypto@vger.kernel.org, iommu@lists.linux-foundation.org,
kvm@vger.kernel.org, linux-accelerators@lists.ozlabs.org, Lu Baolu
[off-list ref], Sanjay Kumar [off-list ref],
linuxarm@huawei.com
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
Message-ID: [off-list ref]
On Tue, 4 Sep 2018 11:00:19 -0400
Jerome Glisse [off-list ref] wrote:
quoted
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
quoted
From: Kenneth Lee <redacted>
WarpDrive is an accelerator framework to expose the hardware capabilities
directly to the user space. It makes use of the exist vfio and vfio-mdev
facilities. So the user application can send request and DMA to the
hardware without interaction with the kernel. This removes the latency
of syscall.
WarpDrive is the name for the whole framework. The component in kernel
is called SDMDEV, Share Domain Mediated Device. Driver driver exposes its
hardware resource by registering to SDMDEV as a VFIO-Mdev. So the user
library of WarpDrive can access it via VFIO interface.
The patchset contains document for the detail. Please refer to it for more
information.
This patchset is intended to be used with Jean Philippe Brucker's SVA
patch [1], which enables not only IO side page fault, but also PASID
support to IOMMU and VFIO.
With these features, WarpDrive can support non-pinned memory and
multi-process in the same accelerator device. We tested it in our SoC
integrated Accelerator (board ID: D06, Chip ID: HIP08). A reference work
tree can be found here: [2].
But it is not mandatory. This patchset is tested in the latest mainline
kernel without the SVA patches. So it supports only one process for each
accelerator.
We have noticed the IOMMU aware mdev RFC announced recently [3].
The IOMMU aware mdev has similar idea but different intention comparing to
WarpDrive. It intends to dedicate part of the hardware resource to a VM.
And the design is supposed to be used with Scalable I/O Virtualization.
While sdmdev is intended to share the hardware resource with a big amount
of processes. It just requires the hardware supporting address
translation per process (PCIE's PASID or ARM SMMU's substream ID).
But we don't see serious confliction on both design. We believe they can be
normalized as one.
So once again i do not understand why you are trying to do things
this way. Kernel already have tons of example of everything you
want to do without a new framework. Moreover i believe you are
confuse by VFIO. To me VFIO is for VM not to create general device
driver frame work.
VFIO is a userspace driver framework, the VM use case just happens to
be a rather prolific one. VFIO was never intended to be solely a VM
device interface and has several other userspace users, notably DPDK
and SPDK, an NVMe backend in QEMU, a userspace NVMe driver, a ruby
wrapper, and perhaps others that I'm not aware of. Whether vfio is
appropriate interface here might certainly still be a debatable topic,
but I would strongly disagree with your last sentence above. Thanks,
Alex
Yes, that is also my standpoint here.
quoted
quoted
So here is your use case as i understand it. You have a device
with a limited number of command queues (can be just one) and in
some case it can support SVA/SVM (when hardware support it and it
is not disabled). Final requirement is being able to schedule cmds
from userspace without ioctl. All of this exists already exists
upstream in few device drivers.
So here is how every body else is doing it. Please explain why
this does not work.
1 Userspace open device file driver. Kernel device driver create
a context and associate it with on open. This context can be
uniq to the process and can bind hardware resources (like a
command queue) to the process.
2 Userspace bind/acquire a commands queue and initialize it with
an ioctl on the device file. Through that ioctl userspace can
be inform wether either SVA/SVM works for the device. If SVA/
SVM works then kernel device driver bind the process to the
device as part of this ioctl.
3 If SVM/SVA does not work userspace do an ioctl to create dma
buffer or something that does exactly the same thing.
4 Userspace mmap the command queue (mmap of the device file by
using informations gather at step 2)
5 Userspace can write commands into the queue it mapped
6 When userspace close the device file all resources are release
just like any existing device drivers.
Hi, Jerome,
Just one thing, as I said in the cover letter, dma-buf requires the application
to use memory created by the driver for DMA. I did try the dma-buf way in
WrapDrive (refer to [4] in the cover letter), it is a good backup for NOIOMMU
mode or we cannot solve the problem in VFIO.
But, in many of my application scenario, the application already has some memory
in hand, maybe allocated by the framework or libraries. Anyway, they don't get
memory from my library, and they pass the poiter for data operation. And they
may also have pointer in the buffer. Those pointer may be used by the
accelerator. So I need hardware fully share the address space with the
application. That is what dmabuf cannot do.
dmabuf can do that ... it is call uptr you can look at i915 for
instance. Still this does not answer my question above, why do
you need to be in VFIO to do any of the above thing ? Kernel has
tons of examples that does all of the above and are not in VFIO
(including usinng existing user pointer with device).
Cheers,
Jérôme
I took a look at i915_gem_execbuffer_ioctl(). It seems it "copy_from_user" the
user memory to the kernel. That is not what we need. What we try to get is: the
user application do something on its data, and push it away to the accelerator,
and says: "I'm tied, it is your turn to do the job...". Then the accelerator has
the memory, referring any portion of it with the same VAs of the application,
even the VAs are stored inside the memory itself.
And I don't understand why I should avoid to use VFIO? As Alex said, VFIO is the
user driver framework. And I need exactly a user driver interface. Why should I
invent another wheel? It has most of stuff I need:
1. Connecting multiple devices to the same application space
2. Pinning and DMA from the application space to the whole set of device
3. Managing hardware resource by device
We just need the last step: make sure multiple applications and the kernel can
share the same IOMMU. Then why shouldn't we use VFIO?
And personally, I believe the maturity and correctness of a framework are driven
by applications. Now the problem in accelerator world is that we don't have a
direction. If we believe the requirement is right, the method itself is not a
big problem in the end. We just need to let people have a unify platform to
share their work together.
Cheers
--
-Kenneth(Hisilicon)
================================================================================
本邮件及其附件含有华为公司的保密信息,仅限于发送给上面地址中列出的个人或群组。禁
止任何其他人以任何形式使用(包括但不限于全部或部分地泄露、复制、或散发)本邮件中
的信息。如果您错收了本邮件,请您立即电话或邮件通知发件人并删除本邮件!
This e-mail and its attachments contain confidential information from HUAWEI,
which is intended only for the person or entity whose address is listed above.
Any use of the
information contained herein in any way (including, but not limited to, total or
partial disclosure, reproduction, or dissemination) by persons other than the
intended
recipient(s) is prohibited. If you receive this e-mail in error, please notify
the sender by phone or email immediately and delete it!
On Fri, Sep 07, 2018 at 12:01:38PM +0800, Kenneth Lee wrote:
On Thu, Sep 06, 2018 at 09:31:33AM -0400, Jerome Glisse wrote:
quoted
Date: Thu, 6 Sep 2018 09:31:33 -0400
From: Jerome Glisse <redacted>
To: Kenneth Lee <redacted>
CC: Alex Williamson <redacted>, Kenneth Lee
[off-list ref], Jonathan Corbet [off-list ref], Herbert Xu
[off-list ref], "David S . Miller" [off-list ref],
Joerg Roedel [off-list ref], Hao Fang [off-list ref], Zhou Wang
[off-list ref], Zaibo Xu [off-list ref], Philippe
Ombredanne [off-list ref], Greg Kroah-Hartman
[off-list ref], Thomas Gleixner [off-list ref],
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-crypto@vger.kernel.org, iommu@lists.linux-foundation.org,
kvm@vger.kernel.org, linux-accelerators@lists.ozlabs.org, Lu Baolu
[off-list ref], Sanjay Kumar [off-list ref],
linuxarm@huawei.com
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
User-Agent: Mutt/1.10.0 (2018-05-17)
Message-ID: [off-list ref]
On Thu, Sep 06, 2018 at 05:45:32PM +0800, Kenneth Lee wrote:
quoted
On Tue, Sep 04, 2018 at 10:15:09AM -0600, Alex Williamson wrote:
quoted
Date: Tue, 4 Sep 2018 10:15:09 -0600
From: Alex Williamson <redacted>
To: Jerome Glisse <redacted>
CC: Kenneth Lee <redacted>, Jonathan Corbet <corbet@lwn.net>,
Herbert Xu [off-list ref], "David S . Miller"
[off-list ref], Joerg Roedel [off-list ref], Kenneth Lee
[off-list ref], Hao Fang [off-list ref], Zhou Wang
[off-list ref], Zaibo Xu [off-list ref], Philippe
Ombredanne [off-list ref], Greg Kroah-Hartman
[off-list ref], Thomas Gleixner [off-list ref],
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-crypto@vger.kernel.org, iommu@lists.linux-foundation.org,
kvm@vger.kernel.org, linux-accelerators@lists.ozlabs.org, Lu Baolu
[off-list ref], Sanjay Kumar [off-list ref],
linuxarm@huawei.com
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
Message-ID: [off-list ref]
On Tue, 4 Sep 2018 11:00:19 -0400
Jerome Glisse [off-list ref] wrote:
quoted
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
quoted
From: Kenneth Lee <redacted>
WarpDrive is an accelerator framework to expose the hardware capabilities
directly to the user space. It makes use of the exist vfio and vfio-mdev
facilities. So the user application can send request and DMA to the
hardware without interaction with the kernel. This removes the latency
of syscall.
WarpDrive is the name for the whole framework. The component in kernel
is called SDMDEV, Share Domain Mediated Device. Driver driver exposes its
hardware resource by registering to SDMDEV as a VFIO-Mdev. So the user
library of WarpDrive can access it via VFIO interface.
The patchset contains document for the detail. Please refer to it for more
information.
This patchset is intended to be used with Jean Philippe Brucker's SVA
patch [1], which enables not only IO side page fault, but also PASID
support to IOMMU and VFIO.
With these features, WarpDrive can support non-pinned memory and
multi-process in the same accelerator device. We tested it in our SoC
integrated Accelerator (board ID: D06, Chip ID: HIP08). A reference work
tree can be found here: [2].
But it is not mandatory. This patchset is tested in the latest mainline
kernel without the SVA patches. So it supports only one process for each
accelerator.
We have noticed the IOMMU aware mdev RFC announced recently [3].
The IOMMU aware mdev has similar idea but different intention comparing to
WarpDrive. It intends to dedicate part of the hardware resource to a VM.
And the design is supposed to be used with Scalable I/O Virtualization.
While sdmdev is intended to share the hardware resource with a big amount
of processes. It just requires the hardware supporting address
translation per process (PCIE's PASID or ARM SMMU's substream ID).
But we don't see serious confliction on both design. We believe they can be
normalized as one.
So once again i do not understand why you are trying to do things
this way. Kernel already have tons of example of everything you
want to do without a new framework. Moreover i believe you are
confuse by VFIO. To me VFIO is for VM not to create general device
driver frame work.
VFIO is a userspace driver framework, the VM use case just happens to
be a rather prolific one. VFIO was never intended to be solely a VM
device interface and has several other userspace users, notably DPDK
and SPDK, an NVMe backend in QEMU, a userspace NVMe driver, a ruby
wrapper, and perhaps others that I'm not aware of. Whether vfio is
appropriate interface here might certainly still be a debatable topic,
but I would strongly disagree with your last sentence above. Thanks,
Alex
Yes, that is also my standpoint here.
quoted
quoted
So here is your use case as i understand it. You have a device
with a limited number of command queues (can be just one) and in
some case it can support SVA/SVM (when hardware support it and it
is not disabled). Final requirement is being able to schedule cmds
from userspace without ioctl. All of this exists already exists
upstream in few device drivers.
So here is how every body else is doing it. Please explain why
this does not work.
1 Userspace open device file driver. Kernel device driver create
a context and associate it with on open. This context can be
uniq to the process and can bind hardware resources (like a
command queue) to the process.
2 Userspace bind/acquire a commands queue and initialize it with
an ioctl on the device file. Through that ioctl userspace can
be inform wether either SVA/SVM works for the device. If SVA/
SVM works then kernel device driver bind the process to the
device as part of this ioctl.
3 If SVM/SVA does not work userspace do an ioctl to create dma
buffer or something that does exactly the same thing.
4 Userspace mmap the command queue (mmap of the device file by
using informations gather at step 2)
5 Userspace can write commands into the queue it mapped
6 When userspace close the device file all resources are release
just like any existing device drivers.
Hi, Jerome,
Just one thing, as I said in the cover letter, dma-buf requires the application
to use memory created by the driver for DMA. I did try the dma-buf way in
WrapDrive (refer to [4] in the cover letter), it is a good backup for NOIOMMU
mode or we cannot solve the problem in VFIO.
But, in many of my application scenario, the application already has some memory
in hand, maybe allocated by the framework or libraries. Anyway, they don't get
memory from my library, and they pass the poiter for data operation. And they
may also have pointer in the buffer. Those pointer may be used by the
accelerator. So I need hardware fully share the address space with the
application. That is what dmabuf cannot do.
dmabuf can do that ... it is call uptr you can look at i915 for
instance. Still this does not answer my question above, why do
you need to be in VFIO to do any of the above thing ? Kernel has
tons of examples that does all of the above and are not in VFIO
(including usinng existing user pointer with device).
Cheers,
Jérôme
I took a look at i915_gem_execbuffer_ioctl(). It seems it "copy_from_user" the
user memory to the kernel. That is not what we need. What we try to get is: the
user application do something on its data, and push it away to the accelerator,
and says: "I'm tied, it is your turn to do the job...". Then the accelerator has
the memory, referring any portion of it with the same VAs of the application,
even the VAs are stored inside the memory itself.
You were not looking at right place see drivers/gpu/drm/i915/i915_gem_userptr.c
It does GUP and create GEM object AFAICR you can wrap that GEM object into a
dma buffer object.
And I don't understand why I should avoid to use VFIO? As Alex said, VFIO is the
user driver framework. And I need exactly a user driver interface. Why should I
invent another wheel? It has most of stuff I need:
1. Connecting multiple devices to the same application space
2. Pinning and DMA from the application space to the whole set of device
3. Managing hardware resource by device
We just need the last step: make sure multiple applications and the kernel can
share the same IOMMU. Then why shouldn't we use VFIO?
Because tons of other drivers already do all of the above outside VFIO. Many
driver have a sizeable userspace side to them (anything with ioctl do) so they
can be construded as userspace driver too.
So there is no reasons to do that under VFIO. Especialy as in your example
it is not a real user space device driver, the userspace portion only knows
about writting command into command buffer AFAICT.
VFIO is for real userspace driver where interrupt, configurations, ... ie
all the driver is handled in userspace. This means that the userspace have
to be trusted as it could program the device to do DMA to anywhere (if
IOMMU is disabled at boot which is still the default configuration in the
kernel).
So i do not see any reasons to do anything you want inside VFIO. All you
want to do can be done outside as easily. Moreover it would be better if
you define clearly each scenario because from where i sit it looks like
you are opening the door wide open to userspace to DMA anywhere when IOMMU
is disabled.
When IOMMU is disabled you can _not_ expose command queue to userspace
unless your device has its own page table and all commands are relative
to that page table and the device page table is populated by kernel driver
in secure way (ie by checking that what is populated can be access).
I do not believe your example device to have such page table nor do i see
a fallback path when IOMMU is disabled that force user to do ioctl for
each commands.
Yes i understand that you target SVA/SVM but still you claim to support
non SVA/SVM. The point is that userspace can not be trusted if you want
to have random program use your device. I am pretty sure that all user
of VFIO are trusted process (like QEMU).
Finaly i am convince that the IOMMU grouping stuff related to VFIO is
useless for your usecase. I really do not see the point of that, it
does complicate things for you for no reasons AFAICT.
And personally, I believe the maturity and correctness of a framework are driven
by applications. Now the problem in accelerator world is that we don't have a
direction. If we believe the requirement is right, the method itself is not a
big problem in the end. We just need to let people have a unify platform to
share their work together.
I am not against that but it seems to me that all you want to do is only
a matter of simplifying discovery of such devices and sharing few common
ioctl (DMA mapping, creating command queue, managing command queue, ...)
and again for all this i do not see the point of doing this under VFIO.
Cheers,
Jérôme
So there is no reasons to do that under VFIO. Especialy as in your example
it is not a real user space device driver, the userspace portion only knows
about writting command into command buffer AFAICT.
VFIO is for real userspace driver where interrupt, configurations, ... ie
all the driver is handled in userspace. This means that the userspace have
to be trusted as it could program the device to do DMA to anywhere (if
IOMMU is disabled at boot which is still the default configuration in the
kernel).
If the IOMMU is disabled (not exactly a kernel default by the way, I
think most IOMMU drivers enable it by default), your userspace driver
can't bypass DMA isolation by accident. It just won't be allowed to
access the device. VFIO requires an IOMMU unless the admin forces the
NOIOMMU mode with the "enable_unsafe_noiommu_mode" module parameter, and
the userspace explicitly asks for it with VFIO_NOIOMMU_IOMMU, which
taints the kernel. Not for production. A normal userspace driver that
uses VFIO can only do DMA to its own memory.
Thanks,
Jean
On Fri, Sep 07, 2018 at 06:55:45PM +0100, Jean-Philippe Brucker wrote:
On 07/09/2018 17:53, Jerome Glisse wrote:
quoted
So there is no reasons to do that under VFIO. Especialy as in your example
it is not a real user space device driver, the userspace portion only knows
about writting command into command buffer AFAICT.
VFIO is for real userspace driver where interrupt, configurations, ... ie
all the driver is handled in userspace. This means that the userspace have
to be trusted as it could program the device to do DMA to anywhere (if
IOMMU is disabled at boot which is still the default configuration in the
kernel).
If the IOMMU is disabled (not exactly a kernel default by the way, I
think most IOMMU drivers enable it by default), your userspace driver
can't bypass DMA isolation by accident. It just won't be allowed to
access the device. VFIO requires an IOMMU unless the admin forces the
NOIOMMU mode with the "enable_unsafe_noiommu_mode" module parameter, and
the userspace explicitly asks for it with VFIO_NOIOMMU_IOMMU, which
taints the kernel. Not for production. A normal userspace driver that
uses VFIO can only do DMA to its own memory.
Didn't know about VFIO check, which is a sane thing. On Intel IOMMU
is disabled by default (see INTEL_IOMMU_DEFAULT_ON Kconfig option).
I am pretty sure it use to be the same for AMD but maybe it is now
enabled by default.
Cheers,
Jérôme
From: Kenneth Lee <hidden> Date: 2018-09-10 03:30:15
On Fri, Sep 07, 2018 at 12:53:06PM -0400, Jerome Glisse wrote:
Date: Fri, 7 Sep 2018 12:53:06 -0400
From: Jerome Glisse <redacted>
To: Kenneth Lee <redacted>
CC: Kenneth Lee <redacted>, Herbert Xu
[off-list ref], kvm@vger.kernel.org, Jonathan Corbet
[off-list ref], Greg Kroah-Hartman [off-list ref], Joerg
Roedel [off-list ref], linux-doc@vger.kernel.org, Sanjay Kumar
[off-list ref], Hao Fang [off-list ref],
iommu@lists.linux-foundation.org, linux-kernel@vger.kernel.org,
linuxarm@huawei.com, Alex Williamson [off-list ref], Thomas
Gleixner [off-list ref], linux-crypto@vger.kernel.org, Zhou Wang
[off-list ref], Philippe Ombredanne [off-list ref],
Zaibo Xu [off-list ref], "David S . Miller" [off-list ref],
linux-accelerators@lists.ozlabs.org, Lu Baolu [off-list ref]
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
User-Agent: Mutt/1.10.0 (2018-05-17)
Message-ID: [off-list ref]
On Fri, Sep 07, 2018 at 12:01:38PM +0800, Kenneth Lee wrote:
quoted
On Thu, Sep 06, 2018 at 09:31:33AM -0400, Jerome Glisse wrote:
quoted
Date: Thu, 6 Sep 2018 09:31:33 -0400
From: Jerome Glisse <redacted>
To: Kenneth Lee <redacted>
CC: Alex Williamson <redacted>, Kenneth Lee
[off-list ref], Jonathan Corbet [off-list ref], Herbert Xu
[off-list ref], "David S . Miller" [off-list ref],
Joerg Roedel [off-list ref], Hao Fang [off-list ref], Zhou Wang
[off-list ref], Zaibo Xu [off-list ref], Philippe
Ombredanne [off-list ref], Greg Kroah-Hartman
[off-list ref], Thomas Gleixner [off-list ref],
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-crypto@vger.kernel.org, iommu@lists.linux-foundation.org,
kvm@vger.kernel.org, linux-accelerators@lists.ozlabs.org, Lu Baolu
[off-list ref], Sanjay Kumar [off-list ref],
linuxarm@huawei.com
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
User-Agent: Mutt/1.10.0 (2018-05-17)
Message-ID: [off-list ref]
On Thu, Sep 06, 2018 at 05:45:32PM +0800, Kenneth Lee wrote:
quoted
On Tue, Sep 04, 2018 at 10:15:09AM -0600, Alex Williamson wrote:
quoted
Date: Tue, 4 Sep 2018 10:15:09 -0600
From: Alex Williamson <redacted>
To: Jerome Glisse <redacted>
CC: Kenneth Lee <redacted>, Jonathan Corbet <corbet@lwn.net>,
Herbert Xu [off-list ref], "David S . Miller"
[off-list ref], Joerg Roedel [off-list ref], Kenneth Lee
[off-list ref], Hao Fang [off-list ref], Zhou Wang
[off-list ref], Zaibo Xu [off-list ref], Philippe
Ombredanne [off-list ref], Greg Kroah-Hartman
[off-list ref], Thomas Gleixner [off-list ref],
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-crypto@vger.kernel.org, iommu@lists.linux-foundation.org,
kvm@vger.kernel.org, linux-accelerators@lists.ozlabs.org, Lu Baolu
[off-list ref], Sanjay Kumar [off-list ref],
linuxarm@huawei.com
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
Message-ID: [off-list ref]
On Tue, 4 Sep 2018 11:00:19 -0400
Jerome Glisse [off-list ref] wrote:
quoted
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
quoted
From: Kenneth Lee <redacted>
WarpDrive is an accelerator framework to expose the hardware capabilities
directly to the user space. It makes use of the exist vfio and vfio-mdev
facilities. So the user application can send request and DMA to the
hardware without interaction with the kernel. This removes the latency
of syscall.
WarpDrive is the name for the whole framework. The component in kernel
is called SDMDEV, Share Domain Mediated Device. Driver driver exposes its
hardware resource by registering to SDMDEV as a VFIO-Mdev. So the user
library of WarpDrive can access it via VFIO interface.
The patchset contains document for the detail. Please refer to it for more
information.
This patchset is intended to be used with Jean Philippe Brucker's SVA
patch [1], which enables not only IO side page fault, but also PASID
support to IOMMU and VFIO.
With these features, WarpDrive can support non-pinned memory and
multi-process in the same accelerator device. We tested it in our SoC
integrated Accelerator (board ID: D06, Chip ID: HIP08). A reference work
tree can be found here: [2].
But it is not mandatory. This patchset is tested in the latest mainline
kernel without the SVA patches. So it supports only one process for each
accelerator.
We have noticed the IOMMU aware mdev RFC announced recently [3].
The IOMMU aware mdev has similar idea but different intention comparing to
WarpDrive. It intends to dedicate part of the hardware resource to a VM.
And the design is supposed to be used with Scalable I/O Virtualization.
While sdmdev is intended to share the hardware resource with a big amount
of processes. It just requires the hardware supporting address
translation per process (PCIE's PASID or ARM SMMU's substream ID).
But we don't see serious confliction on both design. We believe they can be
normalized as one.
So once again i do not understand why you are trying to do things
this way. Kernel already have tons of example of everything you
want to do without a new framework. Moreover i believe you are
confuse by VFIO. To me VFIO is for VM not to create general device
driver frame work.
VFIO is a userspace driver framework, the VM use case just happens to
be a rather prolific one. VFIO was never intended to be solely a VM
device interface and has several other userspace users, notably DPDK
and SPDK, an NVMe backend in QEMU, a userspace NVMe driver, a ruby
wrapper, and perhaps others that I'm not aware of. Whether vfio is
appropriate interface here might certainly still be a debatable topic,
but I would strongly disagree with your last sentence above. Thanks,
Alex
Yes, that is also my standpoint here.
quoted
quoted
So here is your use case as i understand it. You have a device
with a limited number of command queues (can be just one) and in
some case it can support SVA/SVM (when hardware support it and it
is not disabled). Final requirement is being able to schedule cmds
from userspace without ioctl. All of this exists already exists
upstream in few device drivers.
So here is how every body else is doing it. Please explain why
this does not work.
1 Userspace open device file driver. Kernel device driver create
a context and associate it with on open. This context can be
uniq to the process and can bind hardware resources (like a
command queue) to the process.
2 Userspace bind/acquire a commands queue and initialize it with
an ioctl on the device file. Through that ioctl userspace can
be inform wether either SVA/SVM works for the device. If SVA/
SVM works then kernel device driver bind the process to the
device as part of this ioctl.
3 If SVM/SVA does not work userspace do an ioctl to create dma
buffer or something that does exactly the same thing.
4 Userspace mmap the command queue (mmap of the device file by
using informations gather at step 2)
5 Userspace can write commands into the queue it mapped
6 When userspace close the device file all resources are release
just like any existing device drivers.
Hi, Jerome,
Just one thing, as I said in the cover letter, dma-buf requires the application
to use memory created by the driver for DMA. I did try the dma-buf way in
WrapDrive (refer to [4] in the cover letter), it is a good backup for NOIOMMU
mode or we cannot solve the problem in VFIO.
But, in many of my application scenario, the application already has some memory
in hand, maybe allocated by the framework or libraries. Anyway, they don't get
memory from my library, and they pass the poiter for data operation. And they
may also have pointer in the buffer. Those pointer may be used by the
accelerator. So I need hardware fully share the address space with the
application. That is what dmabuf cannot do.
dmabuf can do that ... it is call uptr you can look at i915 for
instance. Still this does not answer my question above, why do
you need to be in VFIO to do any of the above thing ? Kernel has
tons of examples that does all of the above and are not in VFIO
(including usinng existing user pointer with device).
Cheers,
Jérôme
I took a look at i915_gem_execbuffer_ioctl(). It seems it "copy_from_user" the
user memory to the kernel. That is not what we need. What we try to get is: the
user application do something on its data, and push it away to the accelerator,
and says: "I'm tied, it is your turn to do the job...". Then the accelerator has
the memory, referring any portion of it with the same VAs of the application,
even the VAs are stored inside the memory itself.
You were not looking at right place see drivers/gpu/drm/i915/i915_gem_userptr.c
It does GUP and create GEM object AFAICR you can wrap that GEM object into a
dma buffer object.
Thank you for directing me to this implementation. It is interesting:).
But it is not yet solve my problem. If I understand it right, the userptr in
i915 do the following:
1. The user process sets a user pointer with size to the kernel via ioctl.
2. The kernel wraps it as a dma-buf and keeps the process's mm for further
reference.
3. The user pages are allocated, GUPed or DMA mapped to the device. So the data
can be shared between the user space and the hardware.
But my scenario is:
1. The user process has some data in the user space, pointed by a pointer, say
ptr1. And within the memory, there may be some other pointers, let's say one
of them is ptr2.
2. Now I need to assign ptr1 *directly* to the hardware MMIO space. And the
hardware must refer ptr1 and ptr2 *directly* for data.
Userptr lets the hardware and process share the same memory space. But I need
them to share the same *address space*. So IOMMU is a MUST for WarpDrive,
NOIOMMU mode, as Jean said, is just for verifying some of the procedure is OK.
quoted
And I don't understand why I should avoid to use VFIO? As Alex said, VFIO is the
user driver framework. And I need exactly a user driver interface. Why should I
invent another wheel? It has most of stuff I need:
1. Connecting multiple devices to the same application space
2. Pinning and DMA from the application space to the whole set of device
3. Managing hardware resource by device
We just need the last step: make sure multiple applications and the kernel can
share the same IOMMU. Then why shouldn't we use VFIO?
Because tons of other drivers already do all of the above outside VFIO. Many
driver have a sizeable userspace side to them (anything with ioctl do) so they
can be construded as userspace driver too.
Ignoring if there are *tons* of drivers are doing that;), even I do the same as
i915 and solve the address space problem. And if I don't need to with VFIO, why
should I spend so much effort to do it again?
So there is no reasons to do that under VFIO. Especialy as in your example
it is not a real user space device driver, the userspace portion only knows
about writting command into command buffer AFAICT.
VFIO is for real userspace driver where interrupt, configurations, ... ie
all the driver is handled in userspace. This means that the userspace have
to be trusted as it could program the device to do DMA to anywhere (if
IOMMU is disabled at boot which is still the default configuration in the
kernel).
But as Alex explained, VFIO is not simply used by VM. So it need not to have all
stuffs as a driver in host system. And I do need to share the user space as DMA
buffer to the hardware. And I can get it with just a little update, then it can
service me perfectly. I don't understand why I should choose a long route.
So i do not see any reasons to do anything you want inside VFIO. All you
want to do can be done outside as easily. Moreover it would be better if
you define clearly each scenario because from where i sit it looks like
you are opening the door wide open to userspace to DMA anywhere when IOMMU
is disabled.
When IOMMU is disabled you can _not_ expose command queue to userspace
unless your device has its own page table and all commands are relative
to that page table and the device page table is populated by kernel driver
in secure way (ie by checking that what is populated can be access).
I do not believe your example device to have such page table nor do i see
a fallback path when IOMMU is disabled that force user to do ioctl for
each commands.
Yes i understand that you target SVA/SVM but still you claim to support
non SVA/SVM. The point is that userspace can not be trusted if you want
to have random program use your device. I am pretty sure that all user
of VFIO are trusted process (like QEMU).
Finaly i am convince that the IOMMU grouping stuff related to VFIO is
useless for your usecase. I really do not see the point of that, it
does complicate things for you for no reasons AFAICT.
Indeed, I don't like the group thing. I believe VFIO's maintains would not like
it very much either;). But the problem is, the group reflects to the same
IOMMU(unit), which may shared with other devices. It is a security problem. I
cannot ignore it. I have to take it into account event I don't use VFIO.
quoted
And personally, I believe the maturity and correctness of a framework are driven
by applications. Now the problem in accelerator world is that we don't have a
direction. If we believe the requirement is right, the method itself is not a
big problem in the end. We just need to let people have a unify platform to
share their work together.
I am not against that but it seems to me that all you want to do is only
a matter of simplifying discovery of such devices and sharing few common
ioctl (DMA mapping, creating command queue, managing command queue, ...)
and again for all this i do not see the point of doing this under VFIO.
It is not a problem of device management, it is a problem of sharing address
space.
Cheers,
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
[...]
quoted
quoted
I took a look at i915_gem_execbuffer_ioctl(). It seems it "copy_from_user" the
user memory to the kernel. That is not what we need. What we try to get is: the
user application do something on its data, and push it away to the accelerator,
and says: "I'm tied, it is your turn to do the job...". Then the accelerator has
the memory, referring any portion of it with the same VAs of the application,
even the VAs are stored inside the memory itself.
You were not looking at right place see drivers/gpu/drm/i915/i915_gem_userptr.c
It does GUP and create GEM object AFAICR you can wrap that GEM object into a
dma buffer object.
Thank you for directing me to this implementation. It is interesting:).
But it is not yet solve my problem. If I understand it right, the userptr in
i915 do the following:
1. The user process sets a user pointer with size to the kernel via ioctl.
2. The kernel wraps it as a dma-buf and keeps the process's mm for further
reference.
3. The user pages are allocated, GUPed or DMA mapped to the device. So the data
can be shared between the user space and the hardware.
But my scenario is:
1. The user process has some data in the user space, pointed by a pointer, say
ptr1. And within the memory, there may be some other pointers, let's say one
of them is ptr2.
2. Now I need to assign ptr1 *directly* to the hardware MMIO space. And the
hardware must refer ptr1 and ptr2 *directly* for data.
Userptr lets the hardware and process share the same memory space. But I need
them to share the same *address space*. So IOMMU is a MUST for WarpDrive,
NOIOMMU mode, as Jean said, is just for verifying some of the procedure is OK.
So to be 100% clear should we _ignore_ the non SVA/SVM case ?
If so then wait for necessary SVA/SVM to land and do warp drive
without non SVA/SVM path.
If you still want non SVA/SVM path what you want to do only works
if both ptr1 and ptr2 are in a range that is DMA mapped to the
device (moreover you need DMA address to match process address
which is not an easy feat).
Now even if you only want SVA/SVM, i do not see what is the point
of doing this inside VFIO. AMD GPU driver does not and there would
be no benefit for them to be there. Well a AMD VFIO mdev device
driver for QEMU guest might be useful but they have SVIO IIRC.
For SVA/SVM your usage model is:
Setup:
- user space create a warp drive context for the process
- user space create a device specific context for the process
- user space create a user space command queue for the device
- user space bind command queue
At this point the kernel driver has bound the process address
space to the device with a command queue and userspace
Usage:
- user space schedule work and call appropriate flush/update
ioctl from time to time. Might be optional depends on the
hardware, but probably a good idea to enforce so that kernel
can unbind the command queue to bind another process command
queue.
...
Cleanup:
- user space unbind command queue
- user space destroy device specific context
- user space destroy warp drive context
All the above can be implicit when closing the device file.
So again in the above model i do not see anywhere something from
VFIO that would benefit this model.
quoted
quoted
And I don't understand why I should avoid to use VFIO? As Alex said, VFIO is the
user driver framework. And I need exactly a user driver interface. Why should I
invent another wheel? It has most of stuff I need:
1. Connecting multiple devices to the same application space
2. Pinning and DMA from the application space to the whole set of device
3. Managing hardware resource by device
We just need the last step: make sure multiple applications and the kernel can
share the same IOMMU. Then why shouldn't we use VFIO?
Because tons of other drivers already do all of the above outside VFIO. Many
driver have a sizeable userspace side to them (anything with ioctl do) so they
can be construded as userspace driver too.
Ignoring if there are *tons* of drivers are doing that;), even I do the same as
i915 and solve the address space problem. And if I don't need to with VFIO, why
should I spend so much effort to do it again?
Because you do not need any code from VFIO, nor do you need to reinvent
things. If non SVA/SVM matters to you then use dma buffer. If not then
i do not see anything in VFIO that you need.
quoted
So there is no reasons to do that under VFIO. Especialy as in your example
it is not a real user space device driver, the userspace portion only knows
about writting command into command buffer AFAICT.
VFIO is for real userspace driver where interrupt, configurations, ... ie
all the driver is handled in userspace. This means that the userspace have
to be trusted as it could program the device to do DMA to anywhere (if
IOMMU is disabled at boot which is still the default configuration in the
kernel).
But as Alex explained, VFIO is not simply used by VM. So it need not to have all
stuffs as a driver in host system. And I do need to share the user space as DMA
buffer to the hardware. And I can get it with just a little update, then it can
service me perfectly. I don't understand why I should choose a long route.
Again this is not the long route i do not see anything in VFIO that
benefit you in the SVA/SVM case. A basic character device driver can
do that.
quoted
So i do not see any reasons to do anything you want inside VFIO. All you
want to do can be done outside as easily. Moreover it would be better if
you define clearly each scenario because from where i sit it looks like
you are opening the door wide open to userspace to DMA anywhere when IOMMU
is disabled.
When IOMMU is disabled you can _not_ expose command queue to userspace
unless your device has its own page table and all commands are relative
to that page table and the device page table is populated by kernel driver
in secure way (ie by checking that what is populated can be access).
I do not believe your example device to have such page table nor do i see
a fallback path when IOMMU is disabled that force user to do ioctl for
each commands.
Yes i understand that you target SVA/SVM but still you claim to support
non SVA/SVM. The point is that userspace can not be trusted if you want
to have random program use your device. I am pretty sure that all user
of VFIO are trusted process (like QEMU).
Finaly i am convince that the IOMMU grouping stuff related to VFIO is
useless for your usecase. I really do not see the point of that, it
does complicate things for you for no reasons AFAICT.
Indeed, I don't like the group thing. I believe VFIO's maintains would not like
it very much either;). But the problem is, the group reflects to the same
IOMMU(unit), which may shared with other devices. It is a security problem. I
cannot ignore it. I have to take it into account event I don't use VFIO.
To me it seems you are making a policy decission in kernel space ie
wether the device should be isolated in its own group or not is a
decission that is up to the sys admin or something in userspace.
Right now existing user of SVA/SVM don't (at least AFAICT).
Do we really want to force such isolation ?
quoted
quoted
And personally, I believe the maturity and correctness of a framework are driven
by applications. Now the problem in accelerator world is that we don't have a
direction. If we believe the requirement is right, the method itself is not a
big problem in the end. We just need to let people have a unify platform to
share their work together.
I am not against that but it seems to me that all you want to do is only
a matter of simplifying discovery of such devices and sharing few common
ioctl (DMA mapping, creating command queue, managing command queue, ...)
and again for all this i do not see the point of doing this under VFIO.
It is not a problem of device management, it is a problem of sharing address
space.
This ties back to IOMMU SVA/SVM group isolation above.
Jérôme
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
[...]
quoted
quoted
quoted
I took a look at i915_gem_execbuffer_ioctl(). It seems it "copy_from_user" the
user memory to the kernel. That is not what we need. What we try to get is: the
user application do something on its data, and push it away to the accelerator,
and says: "I'm tied, it is your turn to do the job...". Then the accelerator has
the memory, referring any portion of it with the same VAs of the application,
even the VAs are stored inside the memory itself.
You were not looking at right place see drivers/gpu/drm/i915/i915_gem_userptr.c
It does GUP and create GEM object AFAICR you can wrap that GEM object into a
dma buffer object.
Thank you for directing me to this implementation. It is interesting:).
But it is not yet solve my problem. If I understand it right, the userptr in
i915 do the following:
1. The user process sets a user pointer with size to the kernel via ioctl.
2. The kernel wraps it as a dma-buf and keeps the process's mm for further
reference.
3. The user pages are allocated, GUPed or DMA mapped to the device. So the data
can be shared between the user space and the hardware.
But my scenario is:
1. The user process has some data in the user space, pointed by a pointer, say
ptr1. And within the memory, there may be some other pointers, let's say one
of them is ptr2.
2. Now I need to assign ptr1 *directly* to the hardware MMIO space. And the
hardware must refer ptr1 and ptr2 *directly* for data.
Userptr lets the hardware and process share the same memory space. But I need
them to share the same *address space*. So IOMMU is a MUST for WarpDrive,
NOIOMMU mode, as Jean said, is just for verifying some of the procedure is OK.
So to be 100% clear should we _ignore_ the non SVA/SVM case ?
If so then wait for necessary SVA/SVM to land and do warp drive
without non SVA/SVM path.
I think we should clear the concept of SVA/SVM here. As my understanding, Share
Virtual Address/Memory means: any virtual address in a process can be used by
device at the same time. This requires IOMMU device to support PASID. And
optionally, it requires the feature of page-fault-from-device.
But before the feature is settled down, IOMMU can be used immediately in the
current kernel. That make it possible to assign ONE process's virtual addresses
to the device's IOMMU page table with GUP. This make WarpDrive work well for one
process.
Now We are talking about SVA and PASID, just to make sure WarpDrive can benefit
from the feature in the future. It dose not means WarpDrive is useless before
that. And it works for our Zip and RSA accelerators in physical world.
If you still want non SVA/SVM path what you want to do only works
if both ptr1 and ptr2 are in a range that is DMA mapped to the
device (moreover you need DMA address to match process address
which is not an easy feat).
Now even if you only want SVA/SVM, i do not see what is the point
of doing this inside VFIO. AMD GPU driver does not and there would
be no benefit for them to be there. Well a AMD VFIO mdev device
driver for QEMU guest might be useful but they have SVIO IIRC.
For SVA/SVM your usage model is:
Setup:
- user space create a warp drive context for the process
- user space create a device specific context for the process
- user space create a user space command queue for the device
- user space bind command queue
At this point the kernel driver has bound the process address
space to the device with a command queue and userspace
Usage:
- user space schedule work and call appropriate flush/update
ioctl from time to time. Might be optional depends on the
hardware, but probably a good idea to enforce so that kernel
can unbind the command queue to bind another process command
queue.
...
Cleanup:
- user space unbind command queue
- user space destroy device specific context
- user space destroy warp drive context
All the above can be implicit when closing the device file.
So again in the above model i do not see anywhere something from
VFIO that would benefit this model.
Let me show you how the model will be if I use VFIO:
Setup (Kernel part)
- Kernel driver do every as usual to serve the other functionality, NIC
can still be registered to netdev, encryptor can still be registered
to crypto...
- At the same time, the driver can devote some of its hardware resource
and register them as a mdev creator to the VFIO framework. This just
need limited change to the VFIO type1 driver.
Setup (User space)
- System administrator create mdev via the mdev creator interface.
- Following VFIO setup routine, user space open the mdev's group, there is
only one group for one device.
- Without PASID support, you don't need to do anything. With PASID, bind
the PASID to the device via VFIO interface.
- Get the device from the group via VFIO interface and mmap it the user
space for device's MMIO access (for the queue).
- Map whatever memory you need to share with the device with VFIO
interface.
- (opt) Add more devices into the container if you want to share the
same address space with them
Cleanup:
- User space close the group file handler
- There will be a problem to let the other process know the mdev is
freed to be used again. My RFCv1 choose a file handler solution. Alex
dose not like it. But it is not a big problem. We can always have a
scheduler process to manage the state of the mdev or even we can
switch back to the RFCv1 solution without too much effort if we like
in the future.
Except for the minimum update to the type1 driver and use sdmdev to manage the
interrupt sharing, I don't need any extra code to gain the address sharing
capability. And the capability will be strengthen along with the upgrade of VFIO.
quoted
quoted
quoted
And I don't understand why I should avoid to use VFIO? As Alex said, VFIO is the
user driver framework. And I need exactly a user driver interface. Why should I
invent another wheel? It has most of stuff I need:
1. Connecting multiple devices to the same application space
2. Pinning and DMA from the application space to the whole set of device
3. Managing hardware resource by device
We just need the last step: make sure multiple applications and the kernel can
share the same IOMMU. Then why shouldn't we use VFIO?
Because tons of other drivers already do all of the above outside VFIO. Many
driver have a sizeable userspace side to them (anything with ioctl do) so they
can be construded as userspace driver too.
Ignoring if there are *tons* of drivers are doing that;), even I do the same as
i915 and solve the address space problem. And if I don't need to with VFIO, why
should I spend so much effort to do it again?
Because you do not need any code from VFIO, nor do you need to reinvent
things. If non SVA/SVM matters to you then use dma buffer. If not then
i do not see anything in VFIO that you need.
As I have explain, if I don't use VFIO, at lease I have to do all that has been
done in i915 or even more than that.
quoted
quoted
So there is no reasons to do that under VFIO. Especialy as in your example
it is not a real user space device driver, the userspace portion only knows
about writting command into command buffer AFAICT.
VFIO is for real userspace driver where interrupt, configurations, ... ie
all the driver is handled in userspace. This means that the userspace have
to be trusted as it could program the device to do DMA to anywhere (if
IOMMU is disabled at boot which is still the default configuration in the
kernel).
But as Alex explained, VFIO is not simply used by VM. So it need not to have all
stuffs as a driver in host system. And I do need to share the user space as DMA
buffer to the hardware. And I can get it with just a little update, then it can
service me perfectly. I don't understand why I should choose a long route.
Again this is not the long route i do not see anything in VFIO that
benefit you in the SVA/SVM case. A basic character device driver can
do that.
quoted
quoted
So i do not see any reasons to do anything you want inside VFIO. All you
want to do can be done outside as easily. Moreover it would be better if
you define clearly each scenario because from where i sit it looks like
you are opening the door wide open to userspace to DMA anywhere when IOMMU
is disabled.
When IOMMU is disabled you can _not_ expose command queue to userspace
unless your device has its own page table and all commands are relative
to that page table and the device page table is populated by kernel driver
in secure way (ie by checking that what is populated can be access).
I do not believe your example device to have such page table nor do i see
a fallback path when IOMMU is disabled that force user to do ioctl for
each commands.
Yes i understand that you target SVA/SVM but still you claim to support
non SVA/SVM. The point is that userspace can not be trusted if you want
to have random program use your device. I am pretty sure that all user
of VFIO are trusted process (like QEMU).
Finaly i am convince that the IOMMU grouping stuff related to VFIO is
useless for your usecase. I really do not see the point of that, it
does complicate things for you for no reasons AFAICT.
Indeed, I don't like the group thing. I believe VFIO's maintains would not like
it very much either;). But the problem is, the group reflects to the same
IOMMU(unit), which may shared with other devices. It is a security problem. I
cannot ignore it. I have to take it into account event I don't use VFIO.
To me it seems you are making a policy decission in kernel space ie
wether the device should be isolated in its own group or not is a
decission that is up to the sys admin or something in userspace.
Right now existing user of SVA/SVM don't (at least AFAICT).
Do we really want to force such isolation ?
But it is not my decision, that how the iommu subsystem is designed. Personally
I don't like it at all, because all our hardwares have their own stream id
(device id). I don't need the group concept at all. But the iommu subsystem
assume some devices may share the name device ID to a single IOMMU.
quoted
quoted
quoted
And personally, I believe the maturity and correctness of a framework are driven
by applications. Now the problem in accelerator world is that we don't have a
direction. If we believe the requirement is right, the method itself is not a
big problem in the end. We just need to let people have a unify platform to
share their work together.
I am not against that but it seems to me that all you want to do is only
a matter of simplifying discovery of such devices and sharing few common
ioctl (DMA mapping, creating command queue, managing command queue, ...)
and again for all this i do not see the point of doing this under VFIO.
It is not a problem of device management, it is a problem of sharing address
space.
This ties back to IOMMU SVA/SVM group isolation above.
Jérôme
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
[...]
quoted
quoted
quoted
I took a look at i915_gem_execbuffer_ioctl(). It seems it "copy_from_user" the
user memory to the kernel. That is not what we need. What we try to get is: the
user application do something on its data, and push it away to the accelerator,
and says: "I'm tied, it is your turn to do the job...". Then the accelerator has
the memory, referring any portion of it with the same VAs of the application,
even the VAs are stored inside the memory itself.
You were not looking at right place see drivers/gpu/drm/i915/i915_gem_userptr.c
It does GUP and create GEM object AFAICR you can wrap that GEM object into a
dma buffer object.
Thank you for directing me to this implementation. It is interesting:).
But it is not yet solve my problem. If I understand it right, the userptr in
i915 do the following:
1. The user process sets a user pointer with size to the kernel via ioctl.
2. The kernel wraps it as a dma-buf and keeps the process's mm for further
reference.
3. The user pages are allocated, GUPed or DMA mapped to the device. So the data
can be shared between the user space and the hardware.
But my scenario is:
1. The user process has some data in the user space, pointed by a pointer, say
ptr1. And within the memory, there may be some other pointers, let's say one
of them is ptr2.
2. Now I need to assign ptr1 *directly* to the hardware MMIO space. And the
hardware must refer ptr1 and ptr2 *directly* for data.
Userptr lets the hardware and process share the same memory space. But I need
them to share the same *address space*. So IOMMU is a MUST for WarpDrive,
NOIOMMU mode, as Jean said, is just for verifying some of the procedure is OK.
So to be 100% clear should we _ignore_ the non SVA/SVM case ?
If so then wait for necessary SVA/SVM to land and do warp drive
without non SVA/SVM path.
I think we should clear the concept of SVA/SVM here. As my understanding, Share
Virtual Address/Memory means: any virtual address in a process can be used by
device at the same time. This requires IOMMU device to support PASID. And
optionally, it requires the feature of page-fault-from-device.
Yes we agree on what SVA/SVM is. There is a one gotcha thought, access
to range that are MMIO map ie CPU page table pointing to IO memory, IIRC
it is undefined what happens on some platform for a device trying to
access those using SVA/SVM.
But before the feature is settled down, IOMMU can be used immediately in the
current kernel. That make it possible to assign ONE process's virtual addresses
to the device's IOMMU page table with GUP. This make WarpDrive work well for one
process.
UH ? How ? You want to GUP _every_ single valid address in the process
and map it to the device ? How do you handle new vma, page being replace
(despite GUP because of things that utimately calls zap pte) ...
Again here you said that the device must be able to access _any_ valid
pointer. With GUP this is insane.
So i am assuming this is not what you want to do without SVA/SVM ie with
GUP you have a different programming model, one in which the userspace
must first bind _range_ of memory to the device and get a DMA address
for the range.
Again, GUP range of process address space to map it to a device so that
userspace can use the device on the mapped range is something that do
exist in various places in the kernel.
Now We are talking about SVA and PASID, just to make sure WarpDrive can benefit
from the feature in the future. It dose not means WarpDrive is useless before
that. And it works for our Zip and RSA accelerators in physical world.
Just not with random process address ...
quoted
If you still want non SVA/SVM path what you want to do only works
if both ptr1 and ptr2 are in a range that is DMA mapped to the
device (moreover you need DMA address to match process address
which is not an easy feat).
Now even if you only want SVA/SVM, i do not see what is the point
of doing this inside VFIO. AMD GPU driver does not and there would
be no benefit for them to be there. Well a AMD VFIO mdev device
driver for QEMU guest might be useful but they have SVIO IIRC.
For SVA/SVM your usage model is:
Setup:
- user space create a warp drive context for the process
- user space create a device specific context for the process
- user space create a user space command queue for the device
- user space bind command queue
At this point the kernel driver has bound the process address
space to the device with a command queue and userspace
Usage:
- user space schedule work and call appropriate flush/update
ioctl from time to time. Might be optional depends on the
hardware, but probably a good idea to enforce so that kernel
can unbind the command queue to bind another process command
queue.
...
Cleanup:
- user space unbind command queue
- user space destroy device specific context
- user space destroy warp drive context
All the above can be implicit when closing the device file.
So again in the above model i do not see anywhere something from
VFIO that would benefit this model.
Let me show you how the model will be if I use VFIO:
Setup (Kernel part)
- Kernel driver do every as usual to serve the other functionality, NIC
can still be registered to netdev, encryptor can still be registered
to crypto...
- At the same time, the driver can devote some of its hardware resource
and register them as a mdev creator to the VFIO framework. This just
need limited change to the VFIO type1 driver.
In the above VFIO does not help you one bit ... you can do that with
as much code with new common device as front end.
Setup (User space)
- System administrator create mdev via the mdev creator interface.
- Following VFIO setup routine, user space open the mdev's group, there is
only one group for one device.
- Without PASID support, you don't need to do anything. With PASID, bind
the PASID to the device via VFIO interface.
- Get the device from the group via VFIO interface and mmap it the user
space for device's MMIO access (for the queue).
- Map whatever memory you need to share with the device with VFIO
interface.
- (opt) Add more devices into the container if you want to share the
same address space with them
So all VFIO buys you here is boiler plate code that does insert_pfn()
to handle MMIO mapping. Which is just couple hundred lines of boiler
plate code.
Cleanup:
- User space close the group file handler
- There will be a problem to let the other process know the mdev is
freed to be used again. My RFCv1 choose a file handler solution. Alex
dose not like it. But it is not a big problem. We can always have a
scheduler process to manage the state of the mdev or even we can
switch back to the RFCv1 solution without too much effort if we like
in the future.
If you were outside VFIO you would have more freedom on how to do that.
For instance process opening the device file can be placed on queue and
first one in the queue get to use the device until it closes/release the
device. Then next one in queue get the device ...
Except for the minimum update to the type1 driver and use sdmdev to manage the
interrupt sharing, I don't need any extra code to gain the address sharing
capability. And the capability will be strengthen along with the upgrade of VFIO.
quoted
quoted
quoted
quoted
And I don't understand why I should avoid to use VFIO? As Alex said, VFIO is the
user driver framework. And I need exactly a user driver interface. Why should I
invent another wheel? It has most of stuff I need:
1. Connecting multiple devices to the same application space
2. Pinning and DMA from the application space to the whole set of device
3. Managing hardware resource by device
We just need the last step: make sure multiple applications and the kernel can
share the same IOMMU. Then why shouldn't we use VFIO?
Because tons of other drivers already do all of the above outside VFIO. Many
driver have a sizeable userspace side to them (anything with ioctl do) so they
can be construded as userspace driver too.
Ignoring if there are *tons* of drivers are doing that;), even I do the same as
i915 and solve the address space problem. And if I don't need to with VFIO, why
should I spend so much effort to do it again?
Because you do not need any code from VFIO, nor do you need to reinvent
things. If non SVA/SVM matters to you then use dma buffer. If not then
i do not see anything in VFIO that you need.
As I have explain, if I don't use VFIO, at lease I have to do all that has been
done in i915 or even more than that.
So beside the MMIO mmap() handling and dma mapping of range of user space
address space (again all very boiler plate code duplicated accross the
kernel several time in different forms). You do not gain anything being
inside VFIO right ?
quoted
quoted
quoted
So there is no reasons to do that under VFIO. Especialy as in your example
it is not a real user space device driver, the userspace portion only knows
about writting command into command buffer AFAICT.
VFIO is for real userspace driver where interrupt, configurations, ... ie
all the driver is handled in userspace. This means that the userspace have
to be trusted as it could program the device to do DMA to anywhere (if
IOMMU is disabled at boot which is still the default configuration in the
kernel).
But as Alex explained, VFIO is not simply used by VM. So it need not to have all
stuffs as a driver in host system. And I do need to share the user space as DMA
buffer to the hardware. And I can get it with just a little update, then it can
service me perfectly. I don't understand why I should choose a long route.
Again this is not the long route i do not see anything in VFIO that
benefit you in the SVA/SVM case. A basic character device driver can
do that.
quoted
quoted
So i do not see any reasons to do anything you want inside VFIO. All you
want to do can be done outside as easily. Moreover it would be better if
you define clearly each scenario because from where i sit it looks like
you are opening the door wide open to userspace to DMA anywhere when IOMMU
is disabled.
When IOMMU is disabled you can _not_ expose command queue to userspace
unless your device has its own page table and all commands are relative
to that page table and the device page table is populated by kernel driver
in secure way (ie by checking that what is populated can be access).
I do not believe your example device to have such page table nor do i see
a fallback path when IOMMU is disabled that force user to do ioctl for
each commands.
Yes i understand that you target SVA/SVM but still you claim to support
non SVA/SVM. The point is that userspace can not be trusted if you want
to have random program use your device. I am pretty sure that all user
of VFIO are trusted process (like QEMU).
Finaly i am convince that the IOMMU grouping stuff related to VFIO is
useless for your usecase. I really do not see the point of that, it
does complicate things for you for no reasons AFAICT.
Indeed, I don't like the group thing. I believe VFIO's maintains would not like
it very much either;). But the problem is, the group reflects to the same
IOMMU(unit), which may shared with other devices. It is a security problem. I
cannot ignore it. I have to take it into account event I don't use VFIO.
To me it seems you are making a policy decission in kernel space ie
wether the device should be isolated in its own group or not is a
decission that is up to the sys admin or something in userspace.
Right now existing user of SVA/SVM don't (at least AFAICT).
Do we really want to force such isolation ?
But it is not my decision, that how the iommu subsystem is designed. Personally
I don't like it at all, because all our hardwares have their own stream id
(device id). I don't need the group concept at all. But the iommu subsystem
assume some devices may share the name device ID to a single IOMMU.
My question was do you really want to force group isolation for the
device ? Existing SVA/SVM capable driver do not force that, they let
the userspace decide this (sysadm, distributions, ...). Being part of
VFIO (in the way you do, likely ways to avoid this inside VFIO too)
force this decision ie make a policy decision without userspace having
anything to say about it.
The IOMMU group thing as always been doubt full to me, it is advertise
as allowing to share resources (ie IOMMU page table) between devices.
But this assume that all device driver in the group have some way of
communicating with each other to share common DMA address that point
to memory devices care. I believe only VFIO does that and probably
only when use by QEMU.
Anyway my question is:
Is it that much useful to be inside VFIO (to avoid few hundred lines
of boiler plate code) given that it forces you into a model (group
isolation) that so far have never been the prefered way for all
existing device driver that already do what you want to achieve ?
From where i stand i do not see overwhelming reasons to do what you
are doing inside VFIO.
To me it would make more sense to have regular device driver. They
all can have device file under same hierarchy to make devices with
same programming model easy to discover.
Jérôme
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
[...]
quoted
quoted
quoted
I took a look at i915_gem_execbuffer_ioctl(). It seems it "copy_from_user" the
user memory to the kernel. That is not what we need. What we try to get is: the
user application do something on its data, and push it away to the accelerator,
and says: "I'm tied, it is your turn to do the job...". Then the accelerator has
the memory, referring any portion of it with the same VAs of the application,
even the VAs are stored inside the memory itself.
You were not looking at right place see drivers/gpu/drm/i915/i915_gem_userptr.c
It does GUP and create GEM object AFAICR you can wrap that GEM object into a
dma buffer object.
Thank you for directing me to this implementation. It is interesting:).
But it is not yet solve my problem. If I understand it right, the userptr in
i915 do the following:
1. The user process sets a user pointer with size to the kernel via ioctl.
2. The kernel wraps it as a dma-buf and keeps the process's mm for further
reference.
3. The user pages are allocated, GUPed or DMA mapped to the device. So the data
can be shared between the user space and the hardware.
But my scenario is:
1. The user process has some data in the user space, pointed by a pointer, say
ptr1. And within the memory, there may be some other pointers, let's say one
of them is ptr2.
2. Now I need to assign ptr1 *directly* to the hardware MMIO space. And the
hardware must refer ptr1 and ptr2 *directly* for data.
Userptr lets the hardware and process share the same memory space. But I need
them to share the same *address space*. So IOMMU is a MUST for WarpDrive,
NOIOMMU mode, as Jean said, is just for verifying some of the procedure is OK.
So to be 100% clear should we _ignore_ the non SVA/SVM case ?
If so then wait for necessary SVA/SVM to land and do warp drive
without non SVA/SVM path.
I think we should clear the concept of SVA/SVM here. As my understanding, Share
Virtual Address/Memory means: any virtual address in a process can be used by
device at the same time. This requires IOMMU device to support PASID. And
optionally, it requires the feature of page-fault-from-device.
Yes we agree on what SVA/SVM is. There is a one gotcha thought, access
to range that are MMIO map ie CPU page table pointing to IO memory, IIRC
it is undefined what happens on some platform for a device trying to
access those using SVA/SVM.
quoted
But before the feature is settled down, IOMMU can be used immediately in the
current kernel. That make it possible to assign ONE process's virtual addresses
to the device's IOMMU page table with GUP. This make WarpDrive work well for one
process.
UH ? How ? You want to GUP _every_ single valid address in the process
and map it to the device ? How do you handle new vma, page being replace
(despite GUP because of things that utimately calls zap pte) ...
Again here you said that the device must be able to access _any_ valid
pointer. With GUP this is insane.
So i am assuming this is not what you want to do without SVA/SVM ie with
GUP you have a different programming model, one in which the userspace
must first bind _range_ of memory to the device and get a DMA address
for the range.
Again, GUP range of process address space to map it to a device so that
userspace can use the device on the mapped range is something that do
exist in various places in the kernel.
Yes same as your expectation, in WarpDrive, we use the concept of "sharing" to
do so. If some memory is going to be shared among process and devices, we use
wd_share_mem(queue, ptr, size) to share those memory. When the queue is working
in this mode, the point is valid in those memory segments. The wd_share_mem call
vfio dma map syscall which will do GUP.
If SVA/SVM is enabled, user space can set SHARE_ALL flags to the queue. Then
wd_share_mem() is not necessary.
This is really not popular when we started the work on WarpDrive. The GUP
document said it should be put within the scope of mm_sem is locked. Because GUP
simply increase the page refcount, not keep the mapping between the page and the
vma. We keep our work together with VFIO to make sure the problem can be solved
in one deal.
And now we have GUP-longterm and many accounting work in VFIO, we don't want to
do that again.
quoted
Now We are talking about SVA and PASID, just to make sure WarpDrive can benefit
from the feature in the future. It dose not means WarpDrive is useless before
that. And it works for our Zip and RSA accelerators in physical world.
Just not with random process address ...
quoted
quoted
If you still want non SVA/SVM path what you want to do only works
if both ptr1 and ptr2 are in a range that is DMA mapped to the
device (moreover you need DMA address to match process address
which is not an easy feat).
Now even if you only want SVA/SVM, i do not see what is the point
of doing this inside VFIO. AMD GPU driver does not and there would
be no benefit for them to be there. Well a AMD VFIO mdev device
driver for QEMU guest might be useful but they have SVIO IIRC.
For SVA/SVM your usage model is:
Setup:
- user space create a warp drive context for the process
- user space create a device specific context for the process
- user space create a user space command queue for the device
- user space bind command queue
At this point the kernel driver has bound the process address
space to the device with a command queue and userspace
Usage:
- user space schedule work and call appropriate flush/update
ioctl from time to time. Might be optional depends on the
hardware, but probably a good idea to enforce so that kernel
can unbind the command queue to bind another process command
queue.
...
Cleanup:
- user space unbind command queue
- user space destroy device specific context
- user space destroy warp drive context
All the above can be implicit when closing the device file.
So again in the above model i do not see anywhere something from
VFIO that would benefit this model.
Let me show you how the model will be if I use VFIO:
Setup (Kernel part)
- Kernel driver do every as usual to serve the other functionality, NIC
can still be registered to netdev, encryptor can still be registered
to crypto...
- At the same time, the driver can devote some of its hardware resource
and register them as a mdev creator to the VFIO framework. This just
need limited change to the VFIO type1 driver.
In the above VFIO does not help you one bit ... you can do that with
as much code with new common device as front end.
quoted
Setup (User space)
- System administrator create mdev via the mdev creator interface.
- Following VFIO setup routine, user space open the mdev's group, there is
only one group for one device.
- Without PASID support, you don't need to do anything. With PASID, bind
the PASID to the device via VFIO interface.
- Get the device from the group via VFIO interface and mmap it the user
space for device's MMIO access (for the queue).
- Map whatever memory you need to share with the device with VFIO
interface.
- (opt) Add more devices into the container if you want to share the
same address space with them
So all VFIO buys you here is boiler plate code that does insert_pfn()
to handle MMIO mapping. Which is just couple hundred lines of boiler
plate code.
No. With VFIO, I don't need to:
1. GUP and accounting for RLIMIT_MEMLOCK
2. Keep all GUP pages for releasing (VFIO uses the rb_tree to do so)
2. Handle the PASID on SMMU (ARM's IOMMU) myself.
3. Multiple devices menagement (VFIO uses container to manage this)
And even as a boiler plate, it is valueable, the memory thing is sensitive
interface to user space, it can easily become a security problem. If I can
achieve my target within the scope of VFIO, why not? At lease it has been
proved to be safe for the time being.
quoted
Cleanup:
- User space close the group file handler
- There will be a problem to let the other process know the mdev is
freed to be used again. My RFCv1 choose a file handler solution. Alex
dose not like it. But it is not a big problem. We can always have a
scheduler process to manage the state of the mdev or even we can
switch back to the RFCv1 solution without too much effort if we like
in the future.
If you were outside VFIO you would have more freedom on how to do that.
For instance process opening the device file can be placed on queue and
first one in the queue get to use the device until it closes/release the
device. Then next one in queue get the device ...
Yes. I do like the file handle solution. But I hope the solution become mature
as soon as possible. Many of our products, and as I know include some of our
partners, are waiting for a long term solution as direction. If I rely on some
unmature solution, they may choose some deviated, customized solution. That will
be much harmful. Compare to this, the freedom is not so important...
quoted
Except for the minimum update to the type1 driver and use sdmdev to manage the
interrupt sharing, I don't need any extra code to gain the address sharing
capability. And the capability will be strengthen along with the upgrade of VFIO.
quoted
quoted
quoted
quoted
And I don't understand why I should avoid to use VFIO? As Alex said, VFIO is the
user driver framework. And I need exactly a user driver interface. Why should I
invent another wheel? It has most of stuff I need:
1. Connecting multiple devices to the same application space
2. Pinning and DMA from the application space to the whole set of device
3. Managing hardware resource by device
We just need the last step: make sure multiple applications and the kernel can
share the same IOMMU. Then why shouldn't we use VFIO?
Because tons of other drivers already do all of the above outside VFIO. Many
driver have a sizeable userspace side to them (anything with ioctl do) so they
can be construded as userspace driver too.
Ignoring if there are *tons* of drivers are doing that;), even I do the same as
i915 and solve the address space problem. And if I don't need to with VFIO, why
should I spend so much effort to do it again?
Because you do not need any code from VFIO, nor do you need to reinvent
things. If non SVA/SVM matters to you then use dma buffer. If not then
i do not see anything in VFIO that you need.
As I have explain, if I don't use VFIO, at lease I have to do all that has been
done in i915 or even more than that.
So beside the MMIO mmap() handling and dma mapping of range of user space
address space (again all very boiler plate code duplicated accross the
kernel several time in different forms). You do not gain anything being
inside VFIO right ?
As I said, rb-tree for gup, rlimit accounting, cooperation on SMMU, and mature
user interface are our concern.
quoted
quoted
quoted
quoted
So there is no reasons to do that under VFIO. Especialy as in your example
it is not a real user space device driver, the userspace portion only knows
about writting command into command buffer AFAICT.
VFIO is for real userspace driver where interrupt, configurations, ... ie
all the driver is handled in userspace. This means that the userspace have
to be trusted as it could program the device to do DMA to anywhere (if
IOMMU is disabled at boot which is still the default configuration in the
kernel).
But as Alex explained, VFIO is not simply used by VM. So it need not to have all
stuffs as a driver in host system. And I do need to share the user space as DMA
buffer to the hardware. And I can get it with just a little update, then it can
service me perfectly. I don't understand why I should choose a long route.
Again this is not the long route i do not see anything in VFIO that
benefit you in the SVA/SVM case. A basic character device driver can
do that.
quoted
quoted
So i do not see any reasons to do anything you want inside VFIO. All you
want to do can be done outside as easily. Moreover it would be better if
you define clearly each scenario because from where i sit it looks like
you are opening the door wide open to userspace to DMA anywhere when IOMMU
is disabled.
When IOMMU is disabled you can _not_ expose command queue to userspace
unless your device has its own page table and all commands are relative
to that page table and the device page table is populated by kernel driver
in secure way (ie by checking that what is populated can be access).
I do not believe your example device to have such page table nor do i see
a fallback path when IOMMU is disabled that force user to do ioctl for
each commands.
Yes i understand that you target SVA/SVM but still you claim to support
non SVA/SVM. The point is that userspace can not be trusted if you want
to have random program use your device. I am pretty sure that all user
of VFIO are trusted process (like QEMU).
Finaly i am convince that the IOMMU grouping stuff related to VFIO is
useless for your usecase. I really do not see the point of that, it
does complicate things for you for no reasons AFAICT.
Indeed, I don't like the group thing. I believe VFIO's maintains would not like
it very much either;). But the problem is, the group reflects to the same
IOMMU(unit), which may shared with other devices. It is a security problem. I
cannot ignore it. I have to take it into account event I don't use VFIO.
To me it seems you are making a policy decission in kernel space ie
wether the device should be isolated in its own group or not is a
decission that is up to the sys admin or something in userspace.
Right now existing user of SVA/SVM don't (at least AFAICT).
Do we really want to force such isolation ?
But it is not my decision, that how the iommu subsystem is designed. Personally
I don't like it at all, because all our hardwares have their own stream id
(device id). I don't need the group concept at all. But the iommu subsystem
assume some devices may share the name device ID to a single IOMMU.
My question was do you really want to force group isolation for the
device ? Existing SVA/SVM capable driver do not force that, they let
the userspace decide this (sysadm, distributions, ...). Being part of
VFIO (in the way you do, likely ways to avoid this inside VFIO too)
force this decision ie make a policy decision without userspace having
anything to say about it.
The IOMMU group thing as always been doubt full to me, it is advertise
as allowing to share resources (ie IOMMU page table) between devices.
But this assume that all device driver in the group have some way of
communicating with each other to share common DMA address that point
to memory devices care. I believe only VFIO does that and probably
only when use by QEMU.
Anyway my question is:
Is it that much useful to be inside VFIO (to avoid few hundred lines
of boiler plate code) given that it forces you into a model (group
isolation) that so far have never been the prefered way for all
existing device driver that already do what you want to achieve ?
You mean to say I create another framework and copy most of the code from VFIO?
It is hard to believe the mainline kernel will take my code. So how about let me
try the VFIO way first and try that if it won't work? ;)
quoted
From where i stand i do not see overwhelming reasons to do what you
are doing inside VFIO.
To me it would make more sense to have regular device driver. They
all can have device file under same hierarchy to make devices with
same programming model easy to discover.
Jérôme
Cheers
--
-Kenneth(Hisilicon)
================================================================================
本邮件及其附件含有华为公司的保密信息,仅限于发送给上面地址中列出的个人或群组。禁
止任何其他人以任何形式使用(包括但不限于全部或部分地泄露、复制、或散发)本邮件中
的信息。如果您错收了本邮件,请您立即电话或邮件通知发件人并删除本邮件!
This e-mail and its attachments contain confidential information from HUAWEI,
which is intended only for the person or entity whose address is listed above.
Any use of the
information contained herein in any way (including, but not limited to, total or
partial disclosure, reproduction, or dissemination) by persons other than the
intended
recipient(s) is prohibited. If you receive this e-mail in error, please notify
the sender by phone or email immediately and delete it!
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
[...]
quoted
quoted
quoted
I took a look at i915_gem_execbuffer_ioctl(). It seems it "copy_from_user" the
user memory to the kernel. That is not what we need. What we try to get is: the
user application do something on its data, and push it away to the accelerator,
and says: "I'm tied, it is your turn to do the job...". Then the accelerator has
the memory, referring any portion of it with the same VAs of the application,
even the VAs are stored inside the memory itself.
You were not looking at right place see drivers/gpu/drm/i915/i915_gem_userptr.c
It does GUP and create GEM object AFAICR you can wrap that GEM object into a
dma buffer object.
Thank you for directing me to this implementation. It is interesting:).
But it is not yet solve my problem. If I understand it right, the userptr in
i915 do the following:
1. The user process sets a user pointer with size to the kernel via ioctl.
2. The kernel wraps it as a dma-buf and keeps the process's mm for further
reference.
3. The user pages are allocated, GUPed or DMA mapped to the device. So the data
can be shared between the user space and the hardware.
But my scenario is:
1. The user process has some data in the user space, pointed by a pointer, say
ptr1. And within the memory, there may be some other pointers, let's say one
of them is ptr2.
2. Now I need to assign ptr1 *directly* to the hardware MMIO space. And the
hardware must refer ptr1 and ptr2 *directly* for data.
Userptr lets the hardware and process share the same memory space. But I need
them to share the same *address space*. So IOMMU is a MUST for WarpDrive,
NOIOMMU mode, as Jean said, is just for verifying some of the procedure is OK.
So to be 100% clear should we _ignore_ the non SVA/SVM case ?
If so then wait for necessary SVA/SVM to land and do warp drive
without non SVA/SVM path.
I think we should clear the concept of SVA/SVM here. As my understanding, Share
Virtual Address/Memory means: any virtual address in a process can be used by
device at the same time. This requires IOMMU device to support PASID. And
optionally, it requires the feature of page-fault-from-device.
Yes we agree on what SVA/SVM is. There is a one gotcha thought, access
to range that are MMIO map ie CPU page table pointing to IO memory, IIRC
it is undefined what happens on some platform for a device trying to
access those using SVA/SVM.
quoted
But before the feature is settled down, IOMMU can be used immediately in the
current kernel. That make it possible to assign ONE process's virtual addresses
to the device's IOMMU page table with GUP. This make WarpDrive work well for one
process.
UH ? How ? You want to GUP _every_ single valid address in the process
and map it to the device ? How do you handle new vma, page being replace
(despite GUP because of things that utimately calls zap pte) ...
Again here you said that the device must be able to access _any_ valid
pointer. With GUP this is insane.
So i am assuming this is not what you want to do without SVA/SVM ie with
GUP you have a different programming model, one in which the userspace
must first bind _range_ of memory to the device and get a DMA address
for the range.
Again, GUP range of process address space to map it to a device so that
userspace can use the device on the mapped range is something that do
exist in various places in the kernel.
Yes same as your expectation, in WarpDrive, we use the concept of "sharing" to
do so. If some memory is going to be shared among process and devices, we use
wd_share_mem(queue, ptr, size) to share those memory. When the queue is working
in this mode, the point is valid in those memory segments. The wd_share_mem call
vfio dma map syscall which will do GUP.
If SVA/SVM is enabled, user space can set SHARE_ALL flags to the queue. Then
wd_share_mem() is not necessary.
This is really not popular when we started the work on WarpDrive. The GUP
document said it should be put within the scope of mm_sem is locked. Because GUP
simply increase the page refcount, not keep the mapping between the page and the
vma. We keep our work together with VFIO to make sure the problem can be solved
in one deal.
The problem can not be solved in one deal, you can not maintain vaddr
pointing to same page after a fork() this can not be solve without the
use of mmu notifier and device dma mapping invalidation ! So being part
of VFIO will not help you there.
AFAIK VFIO is fine with the way it is as QEMU do not fork() once it
is running a guest and thus the COW that would invalidate vaddr to
physical page assumption is not broken. So i doubt VFIO folks have
any incentive to go down the mmu notifier path and invalidate device
mapping. They also have the replay thing that probably handle some
of fork cases by trusting user space program to do it. In your case
you can not trust the user space program.
In your case AFAICT i do not see any warning or gotcha so the following
scenario is broken (in non SVA/SVM):
1) program setup the device (open container, mdev, setup queue, ...)
2) program map some range of its address space wih VFIO_IOMMU_MAP_DMA
3) program start using the device using map setup in 2)
...
4) program fork()
5) parent trigger COW inside the range setup in 2)
At this point it is the child process that can write to the page that
are access by the device (which was map by the parent in 2)). The
parent can no longer access that memory from the CPU.
There is just no sane way to fix this beside invalidating device mapping
on fork (and you can not rely on userspace to do so) and thus stopping
the device on fork (SVA/SVM case do not have any issue here).
And now we have GUP-longterm and many accounting work in VFIO, we don't want to
do that again.
GUP-longterm does not solve any GUP problem, it just block people to
do GUP on DAX backed vma to avoid pining persistent memory as it is
a nightmare to handle in the block device driver and file system code.
The accounting is the rt limit thing and is litteraly 10 lines of
code so i would not see that as hard to replicate.
quoted
quoted
Now We are talking about SVA and PASID, just to make sure WarpDrive can benefit
from the feature in the future. It dose not means WarpDrive is useless before
that. And it works for our Zip and RSA accelerators in physical world.
Just not with random process address ...
quoted
quoted
If you still want non SVA/SVM path what you want to do only works
if both ptr1 and ptr2 are in a range that is DMA mapped to the
device (moreover you need DMA address to match process address
which is not an easy feat).
Now even if you only want SVA/SVM, i do not see what is the point
of doing this inside VFIO. AMD GPU driver does not and there would
be no benefit for them to be there. Well a AMD VFIO mdev device
driver for QEMU guest might be useful but they have SVIO IIRC.
For SVA/SVM your usage model is:
Setup:
- user space create a warp drive context for the process
- user space create a device specific context for the process
- user space create a user space command queue for the device
- user space bind command queue
At this point the kernel driver has bound the process address
space to the device with a command queue and userspace
Usage:
- user space schedule work and call appropriate flush/update
ioctl from time to time. Might be optional depends on the
hardware, but probably a good idea to enforce so that kernel
can unbind the command queue to bind another process command
queue.
...
Cleanup:
- user space unbind command queue
- user space destroy device specific context
- user space destroy warp drive context
All the above can be implicit when closing the device file.
So again in the above model i do not see anywhere something from
VFIO that would benefit this model.
Let me show you how the model will be if I use VFIO:
Setup (Kernel part)
- Kernel driver do every as usual to serve the other functionality, NIC
can still be registered to netdev, encryptor can still be registered
to crypto...
- At the same time, the driver can devote some of its hardware resource
and register them as a mdev creator to the VFIO framework. This just
need limited change to the VFIO type1 driver.
In the above VFIO does not help you one bit ... you can do that with
as much code with new common device as front end.
quoted
Setup (User space)
- System administrator create mdev via the mdev creator interface.
- Following VFIO setup routine, user space open the mdev's group, there is
only one group for one device.
- Without PASID support, you don't need to do anything. With PASID, bind
the PASID to the device via VFIO interface.
- Get the device from the group via VFIO interface and mmap it the user
space for device's MMIO access (for the queue).
- Map whatever memory you need to share with the device with VFIO
interface.
- (opt) Add more devices into the container if you want to share the
same address space with them
So all VFIO buys you here is boiler plate code that does insert_pfn()
to handle MMIO mapping. Which is just couple hundred lines of boiler
plate code.
No. With VFIO, I don't need to:
1. GUP and accounting for RLIMIT_MEMLOCK
That's 10 line of code ...
2. Keep all GUP pages for releasing (VFIO uses the rb_tree to do so)
GUP pages are not part of rb_tree and what you want to do can be done
in few lines of code here is pseudo code:
warp_dma_map_range(ulong vaddr, ulong npages)
{
struct page *pages = kvzalloc(npages);
for (i = 0; i < npages; ++i, vaddr += PAGE_SIZE) {
GUP(vaddr, &pages[i]);
iommu_map(vaddr, page_to_pfn(pages[i]));
}
kvfree(pages);
}
warp_dma_unmap_range(ulong vaddr, ulong npages)
{
for (i = 0; i < npages; ++i, vaddr += PAGE_SIZE) {
unsigned long pfn;
pfn = iommu_iova_to_phys(vaddr);
iommu_unmap(vaddr);
put_page(pfn_to_page(page)); /* set dirty if mapped write */
}
}
Add locking, error handling, dirtying and comments and you are barely
looking at couple hundred lines of code. You do not need any of the
complexity of VFIO as you do not have the same requirements. Namely
VFIO have to keep track of iova and physical mapping for things like
migration (migrating guest between host) and few others very
virtualization centric requirements.
2. Handle the PASID on SMMU (ARM's IOMMU) myself.
Existing driver do that with 20 lines of with comments and error
handling (see kfd_iommu_bind_process_to_device() for instance) i
doubt you need much more than that.
3. Multiple devices menagement (VFIO uses container to manage this)
All the vfio_group* stuff ? OK that's boiler plate code, note that
hard to replicate thought.
And even as a boiler plate, it is valueable, the memory thing is sensitive
interface to user space, it can easily become a security problem. If I can
achieve my target within the scope of VFIO, why not? At lease it has been
proved to be safe for the time being.
The thing is being part of VFIO impose things on you, things that you
do not need. Like one device per group (maybe it is you imposing this,
i am loosing track here). Or the complex dma mapping tracking ...
quoted
quoted
Cleanup:
- User space close the group file handler
- There will be a problem to let the other process know the mdev is
freed to be used again. My RFCv1 choose a file handler solution. Alex
dose not like it. But it is not a big problem. We can always have a
scheduler process to manage the state of the mdev or even we can
switch back to the RFCv1 solution without too much effort if we like
in the future.
If you were outside VFIO you would have more freedom on how to do that.
For instance process opening the device file can be placed on queue and
first one in the queue get to use the device until it closes/release the
device. Then next one in queue get the device ...
Yes. I do like the file handle solution. But I hope the solution become mature
as soon as possible. Many of our products, and as I know include some of our
partners, are waiting for a long term solution as direction. If I rely on some
unmature solution, they may choose some deviated, customized solution. That will
be much harmful. Compare to this, the freedom is not so important...
I do not see how being part of VFIO protect you from people doing crazy
thing to their kernel ... Time to market being key in this world, i doubt
that being part of VFIO would make anyone think twice before taking a
shortcut.
I have seen horrible things on that front and only players like Google
can impose a minimum level of sanity.
quoted
quoted
Except for the minimum update to the type1 driver and use sdmdev to manage the
interrupt sharing, I don't need any extra code to gain the address sharing
capability. And the capability will be strengthen along with the upgrade of VFIO.
quoted
quoted
quoted
quoted
And I don't understand why I should avoid to use VFIO? As Alex said, VFIO is the
user driver framework. And I need exactly a user driver interface. Why should I
invent another wheel? It has most of stuff I need:
1. Connecting multiple devices to the same application space
2. Pinning and DMA from the application space to the whole set of device
3. Managing hardware resource by device
We just need the last step: make sure multiple applications and the kernel can
share the same IOMMU. Then why shouldn't we use VFIO?
Because tons of other drivers already do all of the above outside VFIO. Many
driver have a sizeable userspace side to them (anything with ioctl do) so they
can be construded as userspace driver too.
Ignoring if there are *tons* of drivers are doing that;), even I do the same as
i915 and solve the address space problem. And if I don't need to with VFIO, why
should I spend so much effort to do it again?
Because you do not need any code from VFIO, nor do you need to reinvent
things. If non SVA/SVM matters to you then use dma buffer. If not then
i do not see anything in VFIO that you need.
As I have explain, if I don't use VFIO, at lease I have to do all that has been
done in i915 or even more than that.
So beside the MMIO mmap() handling and dma mapping of range of user space
address space (again all very boiler plate code duplicated accross the
kernel several time in different forms). You do not gain anything being
inside VFIO right ?
As I said, rb-tree for gup, rlimit accounting, cooperation on SMMU, and mature
user interface are our concern.
quoted
quoted
quoted
quoted
quoted
So there is no reasons to do that under VFIO. Especialy as in your example
it is not a real user space device driver, the userspace portion only knows
about writting command into command buffer AFAICT.
VFIO is for real userspace driver where interrupt, configurations, ... ie
all the driver is handled in userspace. This means that the userspace have
to be trusted as it could program the device to do DMA to anywhere (if
IOMMU is disabled at boot which is still the default configuration in the
kernel).
But as Alex explained, VFIO is not simply used by VM. So it need not to have all
stuffs as a driver in host system. And I do need to share the user space as DMA
buffer to the hardware. And I can get it with just a little update, then it can
service me perfectly. I don't understand why I should choose a long route.
Again this is not the long route i do not see anything in VFIO that
benefit you in the SVA/SVM case. A basic character device driver can
do that.
quoted
quoted
So i do not see any reasons to do anything you want inside VFIO. All you
want to do can be done outside as easily. Moreover it would be better if
you define clearly each scenario because from where i sit it looks like
you are opening the door wide open to userspace to DMA anywhere when IOMMU
is disabled.
When IOMMU is disabled you can _not_ expose command queue to userspace
unless your device has its own page table and all commands are relative
to that page table and the device page table is populated by kernel driver
in secure way (ie by checking that what is populated can be access).
I do not believe your example device to have such page table nor do i see
a fallback path when IOMMU is disabled that force user to do ioctl for
each commands.
Yes i understand that you target SVA/SVM but still you claim to support
non SVA/SVM. The point is that userspace can not be trusted if you want
to have random program use your device. I am pretty sure that all user
of VFIO are trusted process (like QEMU).
Finaly i am convince that the IOMMU grouping stuff related to VFIO is
useless for your usecase. I really do not see the point of that, it
does complicate things for you for no reasons AFAICT.
Indeed, I don't like the group thing. I believe VFIO's maintains would not like
it very much either;). But the problem is, the group reflects to the same
IOMMU(unit), which may shared with other devices. It is a security problem. I
cannot ignore it. I have to take it into account event I don't use VFIO.
To me it seems you are making a policy decission in kernel space ie
wether the device should be isolated in its own group or not is a
decission that is up to the sys admin or something in userspace.
Right now existing user of SVA/SVM don't (at least AFAICT).
Do we really want to force such isolation ?
But it is not my decision, that how the iommu subsystem is designed. Personally
I don't like it at all, because all our hardwares have their own stream id
(device id). I don't need the group concept at all. But the iommu subsystem
assume some devices may share the name device ID to a single IOMMU.
My question was do you really want to force group isolation for the
device ? Existing SVA/SVM capable driver do not force that, they let
the userspace decide this (sysadm, distributions, ...). Being part of
VFIO (in the way you do, likely ways to avoid this inside VFIO too)
force this decision ie make a policy decision without userspace having
anything to say about it.
You still do not answer my question, do you really want to force group
isolation for device in your framework ? Which is a policy decision from
my POV and thus belong to userspace and should not be enforce by kernel.
quoted
The IOMMU group thing as always been doubt full to me, it is advertise
as allowing to share resources (ie IOMMU page table) between devices.
But this assume that all device driver in the group have some way of
communicating with each other to share common DMA address that point
to memory devices care. I believe only VFIO does that and probably
only when use by QEMU.
Anyway my question is:
Is it that much useful to be inside VFIO (to avoid few hundred lines
of boiler plate code) given that it forces you into a model (group
isolation) that so far have never been the prefered way for all
existing device driver that already do what you want to achieve ?
You mean to say I create another framework and copy most of the code from VFIO?
It is hard to believe the mainline kernel will take my code. So how about let me
try the VFIO way first and try that if it won't work? ;)
There is no trying, this is the kernel, once you expose something to
userspace you have to keep supporting it forever ... There is no, hey
let's add this new framework and see how it goes and removing it few
kernel version latter ...
That is why i am being pedantic :) on making sure there is good reasons
to do what you do inside VFIO. I do believe that we want a common frame-
work like the one you are proposing but i do not believe it should be
part of VFIO given the baggages it comes with and that are not relevant
to the use cases for this kind of devices.
Cheers,
Jérôme
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
[...]
quoted
quoted
quoted
I took a look at i915_gem_execbuffer_ioctl(). It seems it "copy_from_user" the
user memory to the kernel. That is not what we need. What we try to get is: the
user application do something on its data, and push it away to the accelerator,
and says: "I'm tied, it is your turn to do the job...". Then the accelerator has
the memory, referring any portion of it with the same VAs of the application,
even the VAs are stored inside the memory itself.
You were not looking at right place see drivers/gpu/drm/i915/i915_gem_userptr.c
It does GUP and create GEM object AFAICR you can wrap that GEM object into a
dma buffer object.
Thank you for directing me to this implementation. It is interesting:).
But it is not yet solve my problem. If I understand it right, the userptr in
i915 do the following:
1. The user process sets a user pointer with size to the kernel via ioctl.
2. The kernel wraps it as a dma-buf and keeps the process's mm for further
reference.
3. The user pages are allocated, GUPed or DMA mapped to the device. So the data
can be shared between the user space and the hardware.
But my scenario is:
1. The user process has some data in the user space, pointed by a pointer, say
ptr1. And within the memory, there may be some other pointers, let's say one
of them is ptr2.
2. Now I need to assign ptr1 *directly* to the hardware MMIO space. And the
hardware must refer ptr1 and ptr2 *directly* for data.
Userptr lets the hardware and process share the same memory space. But I need
them to share the same *address space*. So IOMMU is a MUST for WarpDrive,
NOIOMMU mode, as Jean said, is just for verifying some of the procedure is OK.
So to be 100% clear should we _ignore_ the non SVA/SVM case ?
If so then wait for necessary SVA/SVM to land and do warp drive
without non SVA/SVM path.
I think we should clear the concept of SVA/SVM here. As my understanding, Share
Virtual Address/Memory means: any virtual address in a process can be used by
device at the same time. This requires IOMMU device to support PASID. And
optionally, it requires the feature of page-fault-from-device.
Yes we agree on what SVA/SVM is. There is a one gotcha thought, access
to range that are MMIO map ie CPU page table pointing to IO memory, IIRC
it is undefined what happens on some platform for a device trying to
access those using SVA/SVM.
quoted
But before the feature is settled down, IOMMU can be used immediately in the
current kernel. That make it possible to assign ONE process's virtual addresses
to the device's IOMMU page table with GUP. This make WarpDrive work well for one
process.
UH ? How ? You want to GUP _every_ single valid address in the process
and map it to the device ? How do you handle new vma, page being replace
(despite GUP because of things that utimately calls zap pte) ...
Again here you said that the device must be able to access _any_ valid
pointer. With GUP this is insane.
So i am assuming this is not what you want to do without SVA/SVM ie with
GUP you have a different programming model, one in which the userspace
must first bind _range_ of memory to the device and get a DMA address
for the range.
Again, GUP range of process address space to map it to a device so that
userspace can use the device on the mapped range is something that do
exist in various places in the kernel.
Yes same as your expectation, in WarpDrive, we use the concept of "sharing" to
do so. If some memory is going to be shared among process and devices, we use
wd_share_mem(queue, ptr, size) to share those memory. When the queue is working
in this mode, the point is valid in those memory segments. The wd_share_mem call
vfio dma map syscall which will do GUP.
If SVA/SVM is enabled, user space can set SHARE_ALL flags to the queue. Then
wd_share_mem() is not necessary.
This is really not popular when we started the work on WarpDrive. The GUP
document said it should be put within the scope of mm_sem is locked. Because GUP
simply increase the page refcount, not keep the mapping between the page and the
vma. We keep our work together with VFIO to make sure the problem can be solved
in one deal.
The problem can not be solved in one deal, you can not maintain vaddr
pointing to same page after a fork() this can not be solve without the
use of mmu notifier and device dma mapping invalidation ! So being part
of VFIO will not help you there.
Good point. But sadly, even with mmu notifier and dma mapping invalidation, I
cannot do anything here. If the process fork a sub-process, the sub-process need
a new pasid and hardware resource. The IOMM space mapped should not be used. The
parent process should be aware of this, unmap and close the device file before
the fork. I have the same limitation as VFIO:(
I don't think I can change much here. If I can, VFIO can too:)
AFAIK VFIO is fine with the way it is as QEMU do not fork() once it
is running a guest and thus the COW that would invalidate vaddr to
physical page assumption is not broken. So i doubt VFIO folks have
any incentive to go down the mmu notifier path and invalidate device
mapping. They also have the replay thing that probably handle some
of fork cases by trusting user space program to do it. In your case
you can not trust the user space program.
In your case AFAICT i do not see any warning or gotcha so the following
scenario is broken (in non SVA/SVM):
1) program setup the device (open container, mdev, setup queue, ...)
2) program map some range of its address space wih VFIO_IOMMU_MAP_DMA
3) program start using the device using map setup in 2)
...
4) program fork()
5) parent trigger COW inside the range setup in 2)
At this point it is the child process that can write to the page that
are access by the device (which was map by the parent in 2)). The
parent can no longer access that memory from the CPU.
There is just no sane way to fix this beside invalidating device mapping
on fork (and you can not rely on userspace to do so) and thus stopping
the device on fork (SVA/SVM case do not have any issue here).
Indeed. But as soon as we choose to expose the device space to the user space,
the limitation is already there. If we want to solve the problem, we have to
have a hook in the copy_process() procedure and copy the parent's queue state to
a new queue, assign it to the child's fd and redirect the child's mmap to
it. If I can do so, the same logic can also be applied to VFIO.
The good side is, this is not a security leak. The hardware has been given to
the process. It is the process who choose to share it. If it won't work, it is
the process's problem;)
quoted
And now we have GUP-longterm and many accounting work in VFIO, we don't want to
do that again.
GUP-longterm does not solve any GUP problem, it just block people to
do GUP on DAX backed vma to avoid pining persistent memory as it is
a nightmare to handle in the block device driver and file system code.
The accounting is the rt limit thing and is litteraly 10 lines of
code so i would not see that as hard to replicate.
OK. Agree.
quoted
quoted
quoted
Now We are talking about SVA and PASID, just to make sure WarpDrive can benefit
from the feature in the future. It dose not means WarpDrive is useless before
that. And it works for our Zip and RSA accelerators in physical world.
Just not with random process address ...
quoted
quoted
If you still want non SVA/SVM path what you want to do only works
if both ptr1 and ptr2 are in a range that is DMA mapped to the
device (moreover you need DMA address to match process address
which is not an easy feat).
Now even if you only want SVA/SVM, i do not see what is the point
of doing this inside VFIO. AMD GPU driver does not and there would
be no benefit for them to be there. Well a AMD VFIO mdev device
driver for QEMU guest might be useful but they have SVIO IIRC.
For SVA/SVM your usage model is:
Setup:
- user space create a warp drive context for the process
- user space create a device specific context for the process
- user space create a user space command queue for the device
- user space bind command queue
At this point the kernel driver has bound the process address
space to the device with a command queue and userspace
Usage:
- user space schedule work and call appropriate flush/update
ioctl from time to time. Might be optional depends on the
hardware, but probably a good idea to enforce so that kernel
can unbind the command queue to bind another process command
queue.
...
Cleanup:
- user space unbind command queue
- user space destroy device specific context
- user space destroy warp drive context
All the above can be implicit when closing the device file.
So again in the above model i do not see anywhere something from
VFIO that would benefit this model.
Let me show you how the model will be if I use VFIO:
Setup (Kernel part)
- Kernel driver do every as usual to serve the other functionality, NIC
can still be registered to netdev, encryptor can still be registered
to crypto...
- At the same time, the driver can devote some of its hardware resource
and register them as a mdev creator to the VFIO framework. This just
need limited change to the VFIO type1 driver.
In the above VFIO does not help you one bit ... you can do that with
as much code with new common device as front end.
quoted
Setup (User space)
- System administrator create mdev via the mdev creator interface.
- Following VFIO setup routine, user space open the mdev's group, there is
only one group for one device.
- Without PASID support, you don't need to do anything. With PASID, bind
the PASID to the device via VFIO interface.
- Get the device from the group via VFIO interface and mmap it the user
space for device's MMIO access (for the queue).
- Map whatever memory you need to share with the device with VFIO
interface.
- (opt) Add more devices into the container if you want to share the
same address space with them
So all VFIO buys you here is boiler plate code that does insert_pfn()
to handle MMIO mapping. Which is just couple hundred lines of boiler
plate code.
No. With VFIO, I don't need to:
1. GUP and accounting for RLIMIT_MEMLOCK
That's 10 line of code ...
quoted
2. Keep all GUP pages for releasing (VFIO uses the rb_tree to do so)
GUP pages are not part of rb_tree and what you want to do can be done
in few lines of code here is pseudo code:
warp_dma_map_range(ulong vaddr, ulong npages)
{
struct page *pages = kvzalloc(npages);
for (i = 0; i < npages; ++i, vaddr += PAGE_SIZE) {
GUP(vaddr, &pages[i]);
iommu_map(vaddr, page_to_pfn(pages[i]));
}
kvfree(pages);
}
warp_dma_unmap_range(ulong vaddr, ulong npages)
{
for (i = 0; i < npages; ++i, vaddr += PAGE_SIZE) {
unsigned long pfn;
pfn = iommu_iova_to_phys(vaddr);
iommu_unmap(vaddr);
put_page(pfn_to_page(page)); /* set dirty if mapped write */
}
}
But what if the process exist without unmapping? The pages will be pinned in the
kernel forever.
Add locking, error handling, dirtying and comments and you are barely
looking at couple hundred lines of code. You do not need any of the
complexity of VFIO as you do not have the same requirements. Namely
VFIO have to keep track of iova and physical mapping for things like
migration (migrating guest between host) and few others very
virtualization centric requirements.
quoted
2. Handle the PASID on SMMU (ARM's IOMMU) myself.
Existing driver do that with 20 lines of with comments and error
handling (see kfd_iommu_bind_process_to_device() for instance) i
doubt you need much more than that.
OK, I agree.
quoted
3. Multiple devices menagement (VFIO uses container to manage this)
All the vfio_group* stuff ? OK that's boiler plate code, note that
hard to replicate thought.
No, I meant the container thing. Several devices/group can be assigned to the
same container and the DMA on the container can be assigned to all those
devices. So we can have some devices to share the same name space.
quoted
And even as a boiler plate, it is valueable, the memory thing is sensitive
interface to user space, it can easily become a security problem. If I can
achieve my target within the scope of VFIO, why not? At lease it has been
proved to be safe for the time being.
The thing is being part of VFIO impose things on you, things that you
do not need. Like one device per group (maybe it is you imposing this,
i am loosing track here). Or the complex dma mapping tracking ...
Err... But the one-device-per-group is not VFIO's decision. It is IOMMU's :).
Unless I don't use IOMMU.
quoted
quoted
quoted
Cleanup:
- User space close the group file handler
- There will be a problem to let the other process know the mdev is
freed to be used again. My RFCv1 choose a file handler solution. Alex
dose not like it. But it is not a big problem. We can always have a
scheduler process to manage the state of the mdev or even we can
switch back to the RFCv1 solution without too much effort if we like
in the future.
If you were outside VFIO you would have more freedom on how to do that.
For instance process opening the device file can be placed on queue and
first one in the queue get to use the device until it closes/release the
device. Then next one in queue get the device ...
Yes. I do like the file handle solution. But I hope the solution become mature
as soon as possible. Many of our products, and as I know include some of our
partners, are waiting for a long term solution as direction. If I rely on some
unmature solution, they may choose some deviated, customized solution. That will
be much harmful. Compare to this, the freedom is not so important...
I do not see how being part of VFIO protect you from people doing crazy
thing to their kernel ... Time to market being key in this world, i doubt
that being part of VFIO would make anyone think twice before taking a
shortcut.
I have seen horrible things on that front and only players like Google
can impose a minimum level of sanity.
OK. My fault, to talk about TTM. It has nothing doing with the architecture
decision. But I don't yet see what harm will be brought if I use VFIO when it
can fulfill almost all my requirements.
quoted
quoted
quoted
Except for the minimum update to the type1 driver and use sdmdev to manage the
interrupt sharing, I don't need any extra code to gain the address sharing
capability. And the capability will be strengthen along with the upgrade of VFIO.
quoted
quoted
quoted
quoted
And I don't understand why I should avoid to use VFIO? As Alex said, VFIO is the
user driver framework. And I need exactly a user driver interface. Why should I
invent another wheel? It has most of stuff I need:
1. Connecting multiple devices to the same application space
2. Pinning and DMA from the application space to the whole set of device
3. Managing hardware resource by device
We just need the last step: make sure multiple applications and the kernel can
share the same IOMMU. Then why shouldn't we use VFIO?
Because tons of other drivers already do all of the above outside VFIO. Many
driver have a sizeable userspace side to them (anything with ioctl do) so they
can be construded as userspace driver too.
Ignoring if there are *tons* of drivers are doing that;), even I do the same as
i915 and solve the address space problem. And if I don't need to with VFIO, why
should I spend so much effort to do it again?
Because you do not need any code from VFIO, nor do you need to reinvent
things. If non SVA/SVM matters to you then use dma buffer. If not then
i do not see anything in VFIO that you need.
As I have explain, if I don't use VFIO, at lease I have to do all that has been
done in i915 or even more than that.
So beside the MMIO mmap() handling and dma mapping of range of user space
address space (again all very boiler plate code duplicated accross the
kernel several time in different forms). You do not gain anything being
inside VFIO right ?
As I said, rb-tree for gup, rlimit accounting, cooperation on SMMU, and mature
user interface are our concern.
quoted
quoted
quoted
quoted
quoted
So there is no reasons to do that under VFIO. Especialy as in your example
it is not a real user space device driver, the userspace portion only knows
about writting command into command buffer AFAICT.
VFIO is for real userspace driver where interrupt, configurations, ... ie
all the driver is handled in userspace. This means that the userspace have
to be trusted as it could program the device to do DMA to anywhere (if
IOMMU is disabled at boot which is still the default configuration in the
kernel).
But as Alex explained, VFIO is not simply used by VM. So it need not to have all
stuffs as a driver in host system. And I do need to share the user space as DMA
buffer to the hardware. And I can get it with just a little update, then it can
service me perfectly. I don't understand why I should choose a long route.
Again this is not the long route i do not see anything in VFIO that
benefit you in the SVA/SVM case. A basic character device driver can
do that.
quoted
quoted
So i do not see any reasons to do anything you want inside VFIO. All you
want to do can be done outside as easily. Moreover it would be better if
you define clearly each scenario because from where i sit it looks like
you are opening the door wide open to userspace to DMA anywhere when IOMMU
is disabled.
When IOMMU is disabled you can _not_ expose command queue to userspace
unless your device has its own page table and all commands are relative
to that page table and the device page table is populated by kernel driver
in secure way (ie by checking that what is populated can be access).
I do not believe your example device to have such page table nor do i see
a fallback path when IOMMU is disabled that force user to do ioctl for
each commands.
Yes i understand that you target SVA/SVM but still you claim to support
non SVA/SVM. The point is that userspace can not be trusted if you want
to have random program use your device. I am pretty sure that all user
of VFIO are trusted process (like QEMU).
Finaly i am convince that the IOMMU grouping stuff related to VFIO is
useless for your usecase. I really do not see the point of that, it
does complicate things for you for no reasons AFAICT.
Indeed, I don't like the group thing. I believe VFIO's maintains would not like
it very much either;). But the problem is, the group reflects to the same
IOMMU(unit), which may shared with other devices. It is a security problem. I
cannot ignore it. I have to take it into account event I don't use VFIO.
To me it seems you are making a policy decission in kernel space ie
wether the device should be isolated in its own group or not is a
decission that is up to the sys admin or something in userspace.
Right now existing user of SVA/SVM don't (at least AFAICT).
Do we really want to force such isolation ?
But it is not my decision, that how the iommu subsystem is designed. Personally
I don't like it at all, because all our hardwares have their own stream id
(device id). I don't need the group concept at all. But the iommu subsystem
assume some devices may share the name device ID to a single IOMMU.
My question was do you really want to force group isolation for the
device ? Existing SVA/SVM capable driver do not force that, they let
the userspace decide this (sysadm, distributions, ...). Being part of
VFIO (in the way you do, likely ways to avoid this inside VFIO too)
force this decision ie make a policy decision without userspace having
anything to say about it.
You still do not answer my question, do you really want to force group
isolation for device in your framework ? Which is a policy decision from
my POV and thus belong to userspace and should not be enforce by kernel.
No. But I have to follow the rule defined by IOMMU, haven't I?
quoted
quoted
The IOMMU group thing as always been doubt full to me, it is advertise
as allowing to share resources (ie IOMMU page table) between devices.
But this assume that all device driver in the group have some way of
communicating with each other to share common DMA address that point
to memory devices care. I believe only VFIO does that and probably
only when use by QEMU.
Anyway my question is:
Is it that much useful to be inside VFIO (to avoid few hundred lines
of boiler plate code) given that it forces you into a model (group
isolation) that so far have never been the prefered way for all
existing device driver that already do what you want to achieve ?
You mean to say I create another framework and copy most of the code from VFIO?
It is hard to believe the mainline kernel will take my code. So how about let me
try the VFIO way first and try that if it won't work? ;)
There is no trying, this is the kernel, once you expose something to
userspace you have to keep supporting it forever ... There is no, hey
let's add this new framework and see how it goes and removing it few
kernel version latter ...
No, I don't meant it was unserious when I said "try". I was just not sure if the
community can accept it.
Can Alex say something on this? Is this scenario in the future scope of VFIO? If
it is, we have the season to solve the problem on the way. If it is not, we
should choose other way even we have to copy most of the code.
That is why i am being pedantic :) on making sure there is good reasons
to do what you do inside VFIO. I do believe that we want a common frame-
work like the one you are proposing but i do not believe it should be
part of VFIO given the baggages it comes with and that are not relevant
to the use cases for this kind of devices.
Understood. And I appreciate the discussion and help:)
Cheers
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
[...]
quoted
quoted
quoted
I took a look at i915_gem_execbuffer_ioctl(). It seems it "copy_from_user" the
user memory to the kernel. That is not what we need. What we try to get is: the
user application do something on its data, and push it away to the accelerator,
and says: "I'm tied, it is your turn to do the job...". Then the accelerator has
the memory, referring any portion of it with the same VAs of the application,
even the VAs are stored inside the memory itself.
You were not looking at right place see drivers/gpu/drm/i915/i915_gem_userptr.c
It does GUP and create GEM object AFAICR you can wrap that GEM object into a
dma buffer object.
Thank you for directing me to this implementation. It is interesting:).
But it is not yet solve my problem. If I understand it right, the userptr in
i915 do the following:
1. The user process sets a user pointer with size to the kernel via ioctl.
2. The kernel wraps it as a dma-buf and keeps the process's mm for further
reference.
3. The user pages are allocated, GUPed or DMA mapped to the device. So the data
can be shared between the user space and the hardware.
But my scenario is:
1. The user process has some data in the user space, pointed by a pointer, say
ptr1. And within the memory, there may be some other pointers, let's say one
of them is ptr2.
2. Now I need to assign ptr1 *directly* to the hardware MMIO space. And the
hardware must refer ptr1 and ptr2 *directly* for data.
Userptr lets the hardware and process share the same memory space. But I need
them to share the same *address space*. So IOMMU is a MUST for WarpDrive,
NOIOMMU mode, as Jean said, is just for verifying some of the procedure is OK.
So to be 100% clear should we _ignore_ the non SVA/SVM case ?
If so then wait for necessary SVA/SVM to land and do warp drive
without non SVA/SVM path.
I think we should clear the concept of SVA/SVM here. As my understanding, Share
Virtual Address/Memory means: any virtual address in a process can be used by
device at the same time. This requires IOMMU device to support PASID. And
optionally, it requires the feature of page-fault-from-device.
Yes we agree on what SVA/SVM is. There is a one gotcha thought, access
to range that are MMIO map ie CPU page table pointing to IO memory, IIRC
it is undefined what happens on some platform for a device trying to
access those using SVA/SVM.
quoted
But before the feature is settled down, IOMMU can be used immediately in the
current kernel. That make it possible to assign ONE process's virtual addresses
to the device's IOMMU page table with GUP. This make WarpDrive work well for one
process.
UH ? How ? You want to GUP _every_ single valid address in the process
and map it to the device ? How do you handle new vma, page being replace
(despite GUP because of things that utimately calls zap pte) ...
Again here you said that the device must be able to access _any_ valid
pointer. With GUP this is insane.
So i am assuming this is not what you want to do without SVA/SVM ie with
GUP you have a different programming model, one in which the userspace
must first bind _range_ of memory to the device and get a DMA address
for the range.
Again, GUP range of process address space to map it to a device so that
userspace can use the device on the mapped range is something that do
exist in various places in the kernel.
Yes same as your expectation, in WarpDrive, we use the concept of "sharing" to
do so. If some memory is going to be shared among process and devices, we use
wd_share_mem(queue, ptr, size) to share those memory. When the queue is working
in this mode, the point is valid in those memory segments. The wd_share_mem call
vfio dma map syscall which will do GUP.
If SVA/SVM is enabled, user space can set SHARE_ALL flags to the queue. Then
wd_share_mem() is not necessary.
This is really not popular when we started the work on WarpDrive. The GUP
document said it should be put within the scope of mm_sem is locked. Because GUP
simply increase the page refcount, not keep the mapping between the page and the
vma. We keep our work together with VFIO to make sure the problem can be solved
in one deal.
The problem can not be solved in one deal, you can not maintain vaddr
pointing to same page after a fork() this can not be solve without the
use of mmu notifier and device dma mapping invalidation ! So being part
of VFIO will not help you there.
Good point. But sadly, even with mmu notifier and dma mapping invalidation, I
cannot do anything here. If the process fork a sub-process, the sub-process need
a new pasid and hardware resource. The IOMM space mapped should not be used. The
parent process should be aware of this, unmap and close the device file before
the fork. I have the same limitation as VFIO:(
I don't think I can change much here. If I can, VFIO can too:)
The forbid child to access the device is easy in the kernel whenever
someone open the device file force set the OCLOEXEC flag on the file
some device driver already do that and so should you. With that you
should always have a struct file - mm struct one to one relationship
and thus one PASID per struct file ie per open of the device file.
That does not solve the GUP/fork issue i describe below.
quoted
AFAIK VFIO is fine with the way it is as QEMU do not fork() once it
is running a guest and thus the COW that would invalidate vaddr to
physical page assumption is not broken. So i doubt VFIO folks have
any incentive to go down the mmu notifier path and invalidate device
mapping. They also have the replay thing that probably handle some
of fork cases by trusting user space program to do it. In your case
you can not trust the user space program.
In your case AFAICT i do not see any warning or gotcha so the following
scenario is broken (in non SVA/SVM):
1) program setup the device (open container, mdev, setup queue, ...)
2) program map some range of its address space wih VFIO_IOMMU_MAP_DMA
3) program start using the device using map setup in 2)
...
4) program fork()
5) parent trigger COW inside the range setup in 2)
At this point it is the child process that can write to the page that
are access by the device (which was map by the parent in 2)). The
parent can no longer access that memory from the CPU.
There is just no sane way to fix this beside invalidating device mapping
on fork (and you can not rely on userspace to do so) and thus stopping
the device on fork (SVA/SVM case do not have any issue here).
Indeed. But as soon as we choose to expose the device space to the user space,
the limitation is already there. If we want to solve the problem, we have to
have a hook in the copy_process() procedure and copy the parent's queue state to
a new queue, assign it to the child's fd and redirect the child's mmap to
it. If I can do so, the same logic can also be applied to VFIO.
Except we do not want to do that and this does not solve the COW i describe
above unless you disable COW altogether which is a big no.
The good side is, this is not a security leak. The hardware has been given to
the process. It is the process who choose to share it. If it won't work, it is
the process's problem;)
No this is bad, you can not expect every single userspace program to know and
be aware of that. If some trusted application (say systemd or firefox, ...)
start using your device and is unaware or does not comprehend all side effect
it would allow the child to access/change its parent memory and that is bad.
Device driver need to protect against user doing stupid thing. All existing
device driver that do ATS/PASID already protect themself against child (ie the
child can not use the device through a mix of OCLOEXEC and other checks).
You must set the OCLOEXEC flag but that does not solve the COW problem above.
My motto in life is "do not trust userspace" :) which also translate to "do
not expect userspace will do the right thing".
quoted
quoted
And now we have GUP-longterm and many accounting work in VFIO, we don't want to
do that again.
GUP-longterm does not solve any GUP problem, it just block people to
do GUP on DAX backed vma to avoid pining persistent memory as it is
a nightmare to handle in the block device driver and file system code.
The accounting is the rt limit thing and is litteraly 10 lines of
code so i would not see that as hard to replicate.
OK. Agree.
quoted
quoted
quoted
quoted
Now We are talking about SVA and PASID, just to make sure WarpDrive can benefit
from the feature in the future. It dose not means WarpDrive is useless before
that. And it works for our Zip and RSA accelerators in physical world.
Just not with random process address ...
quoted
quoted
If you still want non SVA/SVM path what you want to do only works
if both ptr1 and ptr2 are in a range that is DMA mapped to the
device (moreover you need DMA address to match process address
which is not an easy feat).
Now even if you only want SVA/SVM, i do not see what is the point
of doing this inside VFIO. AMD GPU driver does not and there would
be no benefit for them to be there. Well a AMD VFIO mdev device
driver for QEMU guest might be useful but they have SVIO IIRC.
For SVA/SVM your usage model is:
Setup:
- user space create a warp drive context for the process
- user space create a device specific context for the process
- user space create a user space command queue for the device
- user space bind command queue
At this point the kernel driver has bound the process address
space to the device with a command queue and userspace
Usage:
- user space schedule work and call appropriate flush/update
ioctl from time to time. Might be optional depends on the
hardware, but probably a good idea to enforce so that kernel
can unbind the command queue to bind another process command
queue.
...
Cleanup:
- user space unbind command queue
- user space destroy device specific context
- user space destroy warp drive context
All the above can be implicit when closing the device file.
So again in the above model i do not see anywhere something from
VFIO that would benefit this model.
Let me show you how the model will be if I use VFIO:
Setup (Kernel part)
- Kernel driver do every as usual to serve the other functionality, NIC
can still be registered to netdev, encryptor can still be registered
to crypto...
- At the same time, the driver can devote some of its hardware resource
and register them as a mdev creator to the VFIO framework. This just
need limited change to the VFIO type1 driver.
In the above VFIO does not help you one bit ... you can do that with
as much code with new common device as front end.
quoted
Setup (User space)
- System administrator create mdev via the mdev creator interface.
- Following VFIO setup routine, user space open the mdev's group, there is
only one group for one device.
- Without PASID support, you don't need to do anything. With PASID, bind
the PASID to the device via VFIO interface.
- Get the device from the group via VFIO interface and mmap it the user
space for device's MMIO access (for the queue).
- Map whatever memory you need to share with the device with VFIO
interface.
- (opt) Add more devices into the container if you want to share the
same address space with them
So all VFIO buys you here is boiler plate code that does insert_pfn()
to handle MMIO mapping. Which is just couple hundred lines of boiler
plate code.
No. With VFIO, I don't need to:
1. GUP and accounting for RLIMIT_MEMLOCK
That's 10 line of code ...
quoted
2. Keep all GUP pages for releasing (VFIO uses the rb_tree to do so)
GUP pages are not part of rb_tree and what you want to do can be done
in few lines of code here is pseudo code:
warp_dma_map_range(ulong vaddr, ulong npages)
{
struct page *pages = kvzalloc(npages);
for (i = 0; i < npages; ++i, vaddr += PAGE_SIZE) {
GUP(vaddr, &pages[i]);
iommu_map(vaddr, page_to_pfn(pages[i]));
}
kvfree(pages);
}
warp_dma_unmap_range(ulong vaddr, ulong npages)
{
for (i = 0; i < npages; ++i, vaddr += PAGE_SIZE) {
unsigned long pfn;
pfn = iommu_iova_to_phys(vaddr);
iommu_unmap(vaddr);
put_page(pfn_to_page(page)); /* set dirty if mapped write */
}
}
But what if the process exist without unmapping? The pages will be pinned in the
kernel forever.
Yeah add a struct warp_map { struct list_head list; unsigned long vaddr,
unsigned long npages; } for every mapping, store the head into the device
file private field of struct file and when the release fs callback is
call you can walk done the list to force unmap any leftover. This is not
that much code. You ca even use interval tree which is 3 lines of code
with interval_tree_generic.h to speed up warp_map lookup on unmap ioctl.
quoted
Add locking, error handling, dirtying and comments and you are barely
looking at couple hundred lines of code. You do not need any of the
complexity of VFIO as you do not have the same requirements. Namely
VFIO have to keep track of iova and physical mapping for things like
migration (migrating guest between host) and few others very
virtualization centric requirements.
quoted
2. Handle the PASID on SMMU (ARM's IOMMU) myself.
Existing driver do that with 20 lines of with comments and error
handling (see kfd_iommu_bind_process_to_device() for instance) i
doubt you need much more than that.
OK, I agree.
quoted
quoted
3. Multiple devices menagement (VFIO uses container to manage this)
All the vfio_group* stuff ? OK that's boiler plate code, note that
hard to replicate thought.
No, I meant the container thing. Several devices/group can be assigned to the
same container and the DMA on the container can be assigned to all those
devices. So we can have some devices to share the same name space.
This was the motivation of my question below, to me this is a policy
decision and it should be left to userspace to decide but not forced
upon userspace because it uses a given device driver.
Maybe i am wrong but i think you can create container and device group
without having a VFIO driver for the devices in the group. It is not
something i do often so i might be wrong here.
quoted
quoted
And even as a boiler plate, it is valueable, the memory thing is sensitive
interface to user space, it can easily become a security problem. If I can
achieve my target within the scope of VFIO, why not? At lease it has been
proved to be safe for the time being.
The thing is being part of VFIO impose things on you, things that you
do not need. Like one device per group (maybe it is you imposing this,
i am loosing track here). Or the complex dma mapping tracking ...
Err... But the one-device-per-group is not VFIO's decision. It is IOMMU's :).
Unless I don't use IOMMU.
AFAIK, on x86 and PPC at least, all PCIE devices are in the same group
by default at boot or at least all devices behind the same bridge.
Maybe they are kernel option to avoid that and userspace init program
can definitly re-arrange that base on sysadmin policy).
quoted
quoted
quoted
quoted
Cleanup:
- User space close the group file handler
- There will be a problem to let the other process know the mdev is
freed to be used again. My RFCv1 choose a file handler solution. Alex
dose not like it. But it is not a big problem. We can always have a
scheduler process to manage the state of the mdev or even we can
switch back to the RFCv1 solution without too much effort if we like
in the future.
If you were outside VFIO you would have more freedom on how to do that.
For instance process opening the device file can be placed on queue and
first one in the queue get to use the device until it closes/release the
device. Then next one in queue get the device ...
Yes. I do like the file handle solution. But I hope the solution become mature
as soon as possible. Many of our products, and as I know include some of our
partners, are waiting for a long term solution as direction. If I rely on some
unmature solution, they may choose some deviated, customized solution. That will
be much harmful. Compare to this, the freedom is not so important...
I do not see how being part of VFIO protect you from people doing crazy
thing to their kernel ... Time to market being key in this world, i doubt
that being part of VFIO would make anyone think twice before taking a
shortcut.
I have seen horrible things on that front and only players like Google
can impose a minimum level of sanity.
OK. My fault, to talk about TTM. It has nothing doing with the architecture
decision. But I don't yet see what harm will be brought if I use VFIO when it
can fulfill almost all my requirements.
The harm is in forcing the device group isolation policy which is not
necessary for all devices like you said so yourself on ARM with the
device stream id so that IOMMU can identify individual devices.
So i would rather see the device isolation as something orthogonal to
what you want to achieve and that should be forced upon user ie sysadmin
should control that distribution can have sane default for each platform.
quoted
quoted
quoted
quoted
Except for the minimum update to the type1 driver and use sdmdev to manage the
interrupt sharing, I don't need any extra code to gain the address sharing
capability. And the capability will be strengthen along with the upgrade of VFIO.
quoted
quoted
quoted
quoted
And I don't understand why I should avoid to use VFIO? As Alex said, VFIO is the
user driver framework. And I need exactly a user driver interface. Why should I
invent another wheel? It has most of stuff I need:
1. Connecting multiple devices to the same application space
2. Pinning and DMA from the application space to the whole set of device
3. Managing hardware resource by device
We just need the last step: make sure multiple applications and the kernel can
share the same IOMMU. Then why shouldn't we use VFIO?
Because tons of other drivers already do all of the above outside VFIO. Many
driver have a sizeable userspace side to them (anything with ioctl do) so they
can be construded as userspace driver too.
Ignoring if there are *tons* of drivers are doing that;), even I do the same as
i915 and solve the address space problem. And if I don't need to with VFIO, why
should I spend so much effort to do it again?
Because you do not need any code from VFIO, nor do you need to reinvent
things. If non SVA/SVM matters to you then use dma buffer. If not then
i do not see anything in VFIO that you need.
As I have explain, if I don't use VFIO, at lease I have to do all that has been
done in i915 or even more than that.
So beside the MMIO mmap() handling and dma mapping of range of user space
address space (again all very boiler plate code duplicated accross the
kernel several time in different forms). You do not gain anything being
inside VFIO right ?
As I said, rb-tree for gup, rlimit accounting, cooperation on SMMU, and mature
user interface are our concern.
quoted
quoted
quoted
quoted
quoted
So there is no reasons to do that under VFIO. Especialy as in your example
it is not a real user space device driver, the userspace portion only knows
about writting command into command buffer AFAICT.
VFIO is for real userspace driver where interrupt, configurations, ... ie
all the driver is handled in userspace. This means that the userspace have
to be trusted as it could program the device to do DMA to anywhere (if
IOMMU is disabled at boot which is still the default configuration in the
kernel).
But as Alex explained, VFIO is not simply used by VM. So it need not to have all
stuffs as a driver in host system. And I do need to share the user space as DMA
buffer to the hardware. And I can get it with just a little update, then it can
service me perfectly. I don't understand why I should choose a long route.
Again this is not the long route i do not see anything in VFIO that
benefit you in the SVA/SVM case. A basic character device driver can
do that.
quoted
quoted
So i do not see any reasons to do anything you want inside VFIO. All you
want to do can be done outside as easily. Moreover it would be better if
you define clearly each scenario because from where i sit it looks like
you are opening the door wide open to userspace to DMA anywhere when IOMMU
is disabled.
When IOMMU is disabled you can _not_ expose command queue to userspace
unless your device has its own page table and all commands are relative
to that page table and the device page table is populated by kernel driver
in secure way (ie by checking that what is populated can be access).
I do not believe your example device to have such page table nor do i see
a fallback path when IOMMU is disabled that force user to do ioctl for
each commands.
Yes i understand that you target SVA/SVM but still you claim to support
non SVA/SVM. The point is that userspace can not be trusted if you want
to have random program use your device. I am pretty sure that all user
of VFIO are trusted process (like QEMU).
Finaly i am convince that the IOMMU grouping stuff related to VFIO is
useless for your usecase. I really do not see the point of that, it
does complicate things for you for no reasons AFAICT.
Indeed, I don't like the group thing. I believe VFIO's maintains would not like
it very much either;). But the problem is, the group reflects to the same
IOMMU(unit), which may shared with other devices. It is a security problem. I
cannot ignore it. I have to take it into account event I don't use VFIO.
To me it seems you are making a policy decission in kernel space ie
wether the device should be isolated in its own group or not is a
decission that is up to the sys admin or something in userspace.
Right now existing user of SVA/SVM don't (at least AFAICT).
Do we really want to force such isolation ?
But it is not my decision, that how the iommu subsystem is designed. Personally
I don't like it at all, because all our hardwares have their own stream id
(device id). I don't need the group concept at all. But the iommu subsystem
assume some devices may share the name device ID to a single IOMMU.
My question was do you really want to force group isolation for the
device ? Existing SVA/SVM capable driver do not force that, they let
the userspace decide this (sysadm, distributions, ...). Being part of
VFIO (in the way you do, likely ways to avoid this inside VFIO too)
force this decision ie make a policy decision without userspace having
anything to say about it.
You still do not answer my question, do you really want to force group
isolation for device in your framework ? Which is a policy decision from
my POV and thus belong to userspace and should not be enforce by kernel.
No. But I have to follow the rule defined by IOMMU, haven't I?
The IOMMU rule does not say that every device _must_ always be in one
group and one domain only. I am pretty sure on x86 by default you get
one domain for all PCIE devices behind same bridge.
My point is that the device grouping into domain/group should be an
orthogonal decision ie it should not be a requirement by the device
driver and should be under control of userspace as it is a policy
decission.
One exception if it is unsafe to have device share a domain in which
case following my motto the driver should refuse to work and return
an error on open (and a kernel explaining why). But this depends on
device and platform.
quoted
quoted
quoted
The IOMMU group thing as always been doubt full to me, it is advertise
as allowing to share resources (ie IOMMU page table) between devices.
But this assume that all device driver in the group have some way of
communicating with each other to share common DMA address that point
to memory devices care. I believe only VFIO does that and probably
only when use by QEMU.
Anyway my question is:
Is it that much useful to be inside VFIO (to avoid few hundred lines
of boiler plate code) given that it forces you into a model (group
isolation) that so far have never been the prefered way for all
existing device driver that already do what you want to achieve ?
You mean to say I create another framework and copy most of the code from VFIO?
It is hard to believe the mainline kernel will take my code. So how about let me
try the VFIO way first and try that if it won't work? ;)
There is no trying, this is the kernel, once you expose something to
userspace you have to keep supporting it forever ... There is no, hey
let's add this new framework and see how it goes and removing it few
kernel version latter ...
No, I don't meant it was unserious when I said "try". I was just not sure if the
community can accept it.
Can Alex say something on this? Is this scenario in the future scope of VFIO? If
it is, we have the season to solve the problem on the way. If it is not, we
should choose other way even we have to copy most of the code.
quoted
That is why i am being pedantic :) on making sure there is good reasons
to do what you do inside VFIO. I do believe that we want a common frame-
work like the one you are proposing but i do not believe it should be
part of VFIO given the baggages it comes with and that are not relevant
to the use cases for this kind of devices.
Understood. And I appreciate the discussion and help:)
Thank you for bearing with in this long discussion :)
Cheers,
Jérôme
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
[...]
quoted
quoted
quoted
I took a look at i915_gem_execbuffer_ioctl(). It seems it "copy_from_user" the
user memory to the kernel. That is not what we need. What we try to get is: the
user application do something on its data, and push it away to the accelerator,
and says: "I'm tied, it is your turn to do the job...". Then the accelerator has
the memory, referring any portion of it with the same VAs of the application,
even the VAs are stored inside the memory itself.
You were not looking at right place see drivers/gpu/drm/i915/i915_gem_userptr.c
It does GUP and create GEM object AFAICR you can wrap that GEM object into a
dma buffer object.
Thank you for directing me to this implementation. It is interesting:).
But it is not yet solve my problem. If I understand it right, the userptr in
i915 do the following:
1. The user process sets a user pointer with size to the kernel via ioctl.
2. The kernel wraps it as a dma-buf and keeps the process's mm for further
reference.
3. The user pages are allocated, GUPed or DMA mapped to the device. So the data
can be shared between the user space and the hardware.
But my scenario is:
1. The user process has some data in the user space, pointed by a pointer, say
ptr1. And within the memory, there may be some other pointers, let's say one
of them is ptr2.
2. Now I need to assign ptr1 *directly* to the hardware MMIO space. And the
hardware must refer ptr1 and ptr2 *directly* for data.
Userptr lets the hardware and process share the same memory space. But I need
them to share the same *address space*. So IOMMU is a MUST for WarpDrive,
NOIOMMU mode, as Jean said, is just for verifying some of the procedure is OK.
So to be 100% clear should we _ignore_ the non SVA/SVM case ?
If so then wait for necessary SVA/SVM to land and do warp drive
without non SVA/SVM path.
I think we should clear the concept of SVA/SVM here. As my understanding, Share
Virtual Address/Memory means: any virtual address in a process can be used by
device at the same time. This requires IOMMU device to support PASID. And
optionally, it requires the feature of page-fault-from-device.
Yes we agree on what SVA/SVM is. There is a one gotcha thought, access
to range that are MMIO map ie CPU page table pointing to IO memory, IIRC
it is undefined what happens on some platform for a device trying to
access those using SVA/SVM.
quoted
But before the feature is settled down, IOMMU can be used immediately in the
current kernel. That make it possible to assign ONE process's virtual addresses
to the device's IOMMU page table with GUP. This make WarpDrive work well for one
process.
UH ? How ? You want to GUP _every_ single valid address in the process
and map it to the device ? How do you handle new vma, page being replace
(despite GUP because of things that utimately calls zap pte) ...
Again here you said that the device must be able to access _any_ valid
pointer. With GUP this is insane.
So i am assuming this is not what you want to do without SVA/SVM ie with
GUP you have a different programming model, one in which the userspace
must first bind _range_ of memory to the device and get a DMA address
for the range.
Again, GUP range of process address space to map it to a device so that
userspace can use the device on the mapped range is something that do
exist in various places in the kernel.
Yes same as your expectation, in WarpDrive, we use the concept of "sharing" to
do so. If some memory is going to be shared among process and devices, we use
wd_share_mem(queue, ptr, size) to share those memory. When the queue is working
in this mode, the point is valid in those memory segments. The wd_share_mem call
vfio dma map syscall which will do GUP.
If SVA/SVM is enabled, user space can set SHARE_ALL flags to the queue. Then
wd_share_mem() is not necessary.
This is really not popular when we started the work on WarpDrive. The GUP
document said it should be put within the scope of mm_sem is locked. Because GUP
simply increase the page refcount, not keep the mapping between the page and the
vma. We keep our work together with VFIO to make sure the problem can be solved
in one deal.
The problem can not be solved in one deal, you can not maintain vaddr
pointing to same page after a fork() this can not be solve without the
use of mmu notifier and device dma mapping invalidation ! So being part
of VFIO will not help you there.
Good point. But sadly, even with mmu notifier and dma mapping invalidation, I
cannot do anything here. If the process fork a sub-process, the sub-process need
a new pasid and hardware resource. The IOMM space mapped should not be used. The
parent process should be aware of this, unmap and close the device file before
the fork. I have the same limitation as VFIO:(
I don't think I can change much here. If I can, VFIO can too:)
The forbid child to access the device is easy in the kernel whenever
someone open the device file force set the OCLOEXEC flag on the file
some device driver already do that and so should you. With that you
should always have a struct file - mm struct one to one relationship
and thus one PASID per struct file ie per open of the device file.
I considerred the OCLOEXEC flag, but it seams it works only for exec, not fork.
That does not solve the GUP/fork issue i describe below.
quoted
quoted
AFAIK VFIO is fine with the way it is as QEMU do not fork() once it
is running a guest and thus the COW that would invalidate vaddr to
physical page assumption is not broken. So i doubt VFIO folks have
any incentive to go down the mmu notifier path and invalidate device
mapping. They also have the replay thing that probably handle some
of fork cases by trusting user space program to do it. In your case
you can not trust the user space program.
In your case AFAICT i do not see any warning or gotcha so the following
scenario is broken (in non SVA/SVM):
1) program setup the device (open container, mdev, setup queue, ...)
2) program map some range of its address space wih VFIO_IOMMU_MAP_DMA
3) program start using the device using map setup in 2)
...
4) program fork()
5) parent trigger COW inside the range setup in 2)
At this point it is the child process that can write to the page that
are access by the device (which was map by the parent in 2)). The
parent can no longer access that memory from the CPU.
There is just no sane way to fix this beside invalidating device mapping
on fork (and you can not rely on userspace to do so) and thus stopping
the device on fork (SVA/SVM case do not have any issue here).
Indeed. But as soon as we choose to expose the device space to the user space,
the limitation is already there. If we want to solve the problem, we have to
have a hook in the copy_process() procedure and copy the parent's queue state to
a new queue, assign it to the child's fd and redirect the child's mmap to
it. If I can do so, the same logic can also be applied to VFIO.
Except we do not want to do that and this does not solve the COW i describe
above unless you disable COW altogether which is a big no.
quoted
The good side is, this is not a security leak. The hardware has been given to
the process. It is the process who choose to share it. If it won't work, it is
the process's problem;)
No this is bad, you can not expect every single userspace program to know and
be aware of that. If some trusted application (say systemd or firefox, ...)
start using your device and is unaware or does not comprehend all side effect
it would allow the child to access/change its parent memory and that is bad.
Device driver need to protect against user doing stupid thing. All existing
device driver that do ATS/PASID already protect themself against child (ie the
child can not use the device through a mix of OCLOEXEC and other checks).
You must set the OCLOEXEC flag but that does not solve the COW problem above.
My motto in life is "do not trust userspace" :) which also translate to "do
not expect userspace will do the right thing".
We don't really trust user space here. We trust that the process cannot do more
than what has been given to it. If an application use WarpDrive, and share its
memory with it, it should know what happen when it fork a sub-process. This is
just like you clone a sub-process with CLONE_FILES or CLONE_VM, the parent
process should know what happen.
quoted
quoted
quoted
And now we have GUP-longterm and many accounting work in VFIO, we don't want to
do that again.
GUP-longterm does not solve any GUP problem, it just block people to
do GUP on DAX backed vma to avoid pining persistent memory as it is
a nightmare to handle in the block device driver and file system code.
The accounting is the rt limit thing and is litteraly 10 lines of
code so i would not see that as hard to replicate.
OK. Agree.
quoted
quoted
quoted
quoted
Now We are talking about SVA and PASID, just to make sure WarpDrive can benefit
from the feature in the future. It dose not means WarpDrive is useless before
that. And it works for our Zip and RSA accelerators in physical world.
Just not with random process address ...
quoted
quoted
If you still want non SVA/SVM path what you want to do only works
if both ptr1 and ptr2 are in a range that is DMA mapped to the
device (moreover you need DMA address to match process address
which is not an easy feat).
Now even if you only want SVA/SVM, i do not see what is the point
of doing this inside VFIO. AMD GPU driver does not and there would
be no benefit for them to be there. Well a AMD VFIO mdev device
driver for QEMU guest might be useful but they have SVIO IIRC.
For SVA/SVM your usage model is:
Setup:
- user space create a warp drive context for the process
- user space create a device specific context for the process
- user space create a user space command queue for the device
- user space bind command queue
At this point the kernel driver has bound the process address
space to the device with a command queue and userspace
Usage:
- user space schedule work and call appropriate flush/update
ioctl from time to time. Might be optional depends on the
hardware, but probably a good idea to enforce so that kernel
can unbind the command queue to bind another process command
queue.
...
Cleanup:
- user space unbind command queue
- user space destroy device specific context
- user space destroy warp drive context
All the above can be implicit when closing the device file.
So again in the above model i do not see anywhere something from
VFIO that would benefit this model.
Let me show you how the model will be if I use VFIO:
Setup (Kernel part)
- Kernel driver do every as usual to serve the other functionality, NIC
can still be registered to netdev, encryptor can still be registered
to crypto...
- At the same time, the driver can devote some of its hardware resource
and register them as a mdev creator to the VFIO framework. This just
need limited change to the VFIO type1 driver.
In the above VFIO does not help you one bit ... you can do that with
as much code with new common device as front end.
quoted
Setup (User space)
- System administrator create mdev via the mdev creator interface.
- Following VFIO setup routine, user space open the mdev's group, there is
only one group for one device.
- Without PASID support, you don't need to do anything. With PASID, bind
the PASID to the device via VFIO interface.
- Get the device from the group via VFIO interface and mmap it the user
space for device's MMIO access (for the queue).
- Map whatever memory you need to share with the device with VFIO
interface.
- (opt) Add more devices into the container if you want to share the
same address space with them
So all VFIO buys you here is boiler plate code that does insert_pfn()
to handle MMIO mapping. Which is just couple hundred lines of boiler
plate code.
No. With VFIO, I don't need to:
1. GUP and accounting for RLIMIT_MEMLOCK
That's 10 line of code ...
quoted
2. Keep all GUP pages for releasing (VFIO uses the rb_tree to do so)
GUP pages are not part of rb_tree and what you want to do can be done
in few lines of code here is pseudo code:
warp_dma_map_range(ulong vaddr, ulong npages)
{
struct page *pages = kvzalloc(npages);
for (i = 0; i < npages; ++i, vaddr += PAGE_SIZE) {
GUP(vaddr, &pages[i]);
iommu_map(vaddr, page_to_pfn(pages[i]));
}
kvfree(pages);
}
warp_dma_unmap_range(ulong vaddr, ulong npages)
{
for (i = 0; i < npages; ++i, vaddr += PAGE_SIZE) {
unsigned long pfn;
pfn = iommu_iova_to_phys(vaddr);
iommu_unmap(vaddr);
put_page(pfn_to_page(page)); /* set dirty if mapped write */
}
}
But what if the process exist without unmapping? The pages will be pinned in the
kernel forever.
Yeah add a struct warp_map { struct list_head list; unsigned long vaddr,
unsigned long npages; } for every mapping, store the head into the device
file private field of struct file and when the release fs callback is
call you can walk done the list to force unmap any leftover. This is not
that much code. You ca even use interval tree which is 3 lines of code
with interval_tree_generic.h to speed up warp_map lookup on unmap ioctl.
Yes, when you add all of them... it is VFIO:)
quoted
quoted
Add locking, error handling, dirtying and comments and you are barely
looking at couple hundred lines of code. You do not need any of the
complexity of VFIO as you do not have the same requirements. Namely
VFIO have to keep track of iova and physical mapping for things like
migration (migrating guest between host) and few others very
virtualization centric requirements.
quoted
2. Handle the PASID on SMMU (ARM's IOMMU) myself.
Existing driver do that with 20 lines of with comments and error
handling (see kfd_iommu_bind_process_to_device() for instance) i
doubt you need much more than that.
OK, I agree.
quoted
quoted
3. Multiple devices menagement (VFIO uses container to manage this)
All the vfio_group* stuff ? OK that's boiler plate code, note that
hard to replicate thought.
No, I meant the container thing. Several devices/group can be assigned to the
same container and the DMA on the container can be assigned to all those
devices. So we can have some devices to share the same name space.
This was the motivation of my question below, to me this is a policy
decision and it should be left to userspace to decide but not forced
upon userspace because it uses a given device driver.
Maybe i am wrong but i think you can create container and device group
without having a VFIO driver for the devices in the group. It is not
something i do often so i might be wrong here.
Container maintains a virtual unify address space for all group/iommu which is
added to it. It simplify the address management. But yes, you can choose to do
it all in the user space.
quoted
quoted
quoted
And even as a boiler plate, it is valueable, the memory thing is sensitive
interface to user space, it can easily become a security problem. If I can
achieve my target within the scope of VFIO, why not? At lease it has been
proved to be safe for the time being.
The thing is being part of VFIO impose things on you, things that you
do not need. Like one device per group (maybe it is you imposing this,
i am loosing track here). Or the complex dma mapping tracking ...
Err... But the one-device-per-group is not VFIO's decision. It is IOMMU's :).
Unless I don't use IOMMU.
AFAIK, on x86 and PPC at least, all PCIE devices are in the same group
by default at boot or at least all devices behind the same bridge.
Maybe they are kernel option to avoid that and userspace init program
can definitly re-arrange that base on sysadmin policy).
But if the IOMMU is enabled, all PCIE devices have their own device IDs. So the
IOMMU can use different page table for every of them.
quoted
quoted
quoted
quoted
quoted
Cleanup:
- User space close the group file handler
- There will be a problem to let the other process know the mdev is
freed to be used again. My RFCv1 choose a file handler solution. Alex
dose not like it. But it is not a big problem. We can always have a
scheduler process to manage the state of the mdev or even we can
switch back to the RFCv1 solution without too much effort if we like
in the future.
If you were outside VFIO you would have more freedom on how to do that.
For instance process opening the device file can be placed on queue and
first one in the queue get to use the device until it closes/release the
device. Then next one in queue get the device ...
Yes. I do like the file handle solution. But I hope the solution become mature
as soon as possible. Many of our products, and as I know include some of our
partners, are waiting for a long term solution as direction. If I rely on some
unmature solution, they may choose some deviated, customized solution. That will
be much harmful. Compare to this, the freedom is not so important...
I do not see how being part of VFIO protect you from people doing crazy
thing to their kernel ... Time to market being key in this world, i doubt
that being part of VFIO would make anyone think twice before taking a
shortcut.
I have seen horrible things on that front and only players like Google
can impose a minimum level of sanity.
OK. My fault, to talk about TTM. It has nothing doing with the architecture
decision. But I don't yet see what harm will be brought if I use VFIO when it
can fulfill almost all my requirements.
The harm is in forcing the device group isolation policy which is not
necessary for all devices like you said so yourself on ARM with the
device stream id so that IOMMU can identify individual devices.
No. Some mini SoC share the same stream id among several small devices. The
iommu has to treat them as the same. That is why IOMMU introduces the concept of
iommu_group. Personally I don't like that. If they share the same stream id,
they should be treated as the same hardware. But it is the decision for the time
being...
So i would rather see the device isolation as something orthogonal to
what you want to achieve and that should be forced upon user ie sysadmin
should control that distribution can have sane default for each platform.
quoted
quoted
quoted
quoted
quoted
Except for the minimum update to the type1 driver and use sdmdev to manage the
interrupt sharing, I don't need any extra code to gain the address sharing
capability. And the capability will be strengthen along with the upgrade of VFIO.
quoted
quoted
quoted
quoted
And I don't understand why I should avoid to use VFIO? As Alex said, VFIO is the
user driver framework. And I need exactly a user driver interface. Why should I
invent another wheel? It has most of stuff I need:
1. Connecting multiple devices to the same application space
2. Pinning and DMA from the application space to the whole set of device
3. Managing hardware resource by device
We just need the last step: make sure multiple applications and the kernel can
share the same IOMMU. Then why shouldn't we use VFIO?
Because tons of other drivers already do all of the above outside VFIO. Many
driver have a sizeable userspace side to them (anything with ioctl do) so they
can be construded as userspace driver too.
Ignoring if there are *tons* of drivers are doing that;), even I do the same as
i915 and solve the address space problem. And if I don't need to with VFIO, why
should I spend so much effort to do it again?
Because you do not need any code from VFIO, nor do you need to reinvent
things. If non SVA/SVM matters to you then use dma buffer. If not then
i do not see anything in VFIO that you need.
As I have explain, if I don't use VFIO, at lease I have to do all that has been
done in i915 or even more than that.
So beside the MMIO mmap() handling and dma mapping of range of user space
address space (again all very boiler plate code duplicated accross the
kernel several time in different forms). You do not gain anything being
inside VFIO right ?
As I said, rb-tree for gup, rlimit accounting, cooperation on SMMU, and mature
user interface are our concern.
quoted
quoted
quoted
quoted
quoted
So there is no reasons to do that under VFIO. Especialy as in your example
it is not a real user space device driver, the userspace portion only knows
about writting command into command buffer AFAICT.
VFIO is for real userspace driver where interrupt, configurations, ... ie
all the driver is handled in userspace. This means that the userspace have
to be trusted as it could program the device to do DMA to anywhere (if
IOMMU is disabled at boot which is still the default configuration in the
kernel).
But as Alex explained, VFIO is not simply used by VM. So it need not to have all
stuffs as a driver in host system. And I do need to share the user space as DMA
buffer to the hardware. And I can get it with just a little update, then it can
service me perfectly. I don't understand why I should choose a long route.
Again this is not the long route i do not see anything in VFIO that
benefit you in the SVA/SVM case. A basic character device driver can
do that.
quoted
quoted
So i do not see any reasons to do anything you want inside VFIO. All you
want to do can be done outside as easily. Moreover it would be better if
you define clearly each scenario because from where i sit it looks like
you are opening the door wide open to userspace to DMA anywhere when IOMMU
is disabled.
When IOMMU is disabled you can _not_ expose command queue to userspace
unless your device has its own page table and all commands are relative
to that page table and the device page table is populated by kernel driver
in secure way (ie by checking that what is populated can be access).
I do not believe your example device to have such page table nor do i see
a fallback path when IOMMU is disabled that force user to do ioctl for
each commands.
Yes i understand that you target SVA/SVM but still you claim to support
non SVA/SVM. The point is that userspace can not be trusted if you want
to have random program use your device. I am pretty sure that all user
of VFIO are trusted process (like QEMU).
Finaly i am convince that the IOMMU grouping stuff related to VFIO is
useless for your usecase. I really do not see the point of that, it
does complicate things for you for no reasons AFAICT.
Indeed, I don't like the group thing. I believe VFIO's maintains would not like
it very much either;). But the problem is, the group reflects to the same
IOMMU(unit), which may shared with other devices. It is a security problem. I
cannot ignore it. I have to take it into account event I don't use VFIO.
To me it seems you are making a policy decission in kernel space ie
wether the device should be isolated in its own group or not is a
decission that is up to the sys admin or something in userspace.
Right now existing user of SVA/SVM don't (at least AFAICT).
Do we really want to force such isolation ?
But it is not my decision, that how the iommu subsystem is designed. Personally
I don't like it at all, because all our hardwares have their own stream id
(device id). I don't need the group concept at all. But the iommu subsystem
assume some devices may share the name device ID to a single IOMMU.
My question was do you really want to force group isolation for the
device ? Existing SVA/SVM capable driver do not force that, they let
the userspace decide this (sysadm, distributions, ...). Being part of
VFIO (in the way you do, likely ways to avoid this inside VFIO too)
force this decision ie make a policy decision without userspace having
anything to say about it.
You still do not answer my question, do you really want to force group
isolation for device in your framework ? Which is a policy decision from
my POV and thus belong to userspace and should not be enforce by kernel.
No. But I have to follow the rule defined by IOMMU, haven't I?
The IOMMU rule does not say that every device _must_ always be in one
group and one domain only. I am pretty sure on x86 by default you get
one domain for all PCIE devices behind same bridge.
Really? Give me some more time. I need to check it out.
My point is that the device grouping into domain/group should be an
orthogonal decision ie it should not be a requirement by the device
driver and should be under control of userspace as it is a policy
decission.
One exception if it is unsafe to have device share a domain in which
case following my motto the driver should refuse to work and return
an error on open (and a kernel explaining why). But this depends on
device and platform.
quoted
quoted
quoted
quoted
The IOMMU group thing as always been doubt full to me, it is advertise
as allowing to share resources (ie IOMMU page table) between devices.
But this assume that all device driver in the group have some way of
communicating with each other to share common DMA address that point
to memory devices care. I believe only VFIO does that and probably
only when use by QEMU.
Anyway my question is:
Is it that much useful to be inside VFIO (to avoid few hundred lines
of boiler plate code) given that it forces you into a model (group
isolation) that so far have never been the prefered way for all
existing device driver that already do what you want to achieve ?
You mean to say I create another framework and copy most of the code from VFIO?
It is hard to believe the mainline kernel will take my code. So how about let me
try the VFIO way first and try that if it won't work? ;)
There is no trying, this is the kernel, once you expose something to
userspace you have to keep supporting it forever ... There is no, hey
let's add this new framework and see how it goes and removing it few
kernel version latter ...
No, I don't meant it was unserious when I said "try". I was just not sure if the
community can accept it.
Can Alex say something on this? Is this scenario in the future scope of VFIO? If
it is, we have the season to solve the problem on the way. If it is not, we
should choose other way even we have to copy most of the code.
quoted
That is why i am being pedantic :) on making sure there is good reasons
to do what you do inside VFIO. I do believe that we want a common frame-
work like the one you are proposing but i do not believe it should be
part of VFIO given the baggages it comes with and that are not relevant
to the use cases for this kind of devices.
Understood. And I appreciate the discussion and help:)
Thank you for bearing with in this long discussion :)
Cheers,
Jérôme
From: Jerome Glisse
Sent: Thursday, September 13, 2018 10:52 PM
[...]
> AFAIK, on x86 and PPC at least, all PCIE devices are in the same group
by default at boot or at least all devices behind the same bridge.
the group thing reflects physical hierarchy limitation, not changed
cross boot. Please note iommu group defines the minimal isolation
boundary - all devices within same group must be attached to the
same iommu domain or address space, because physically IOMMU
cannot differentiate DMAs out of those devices. devices behind
legacy PCI-X bridge is one example. other examples include devices
behind a PCIe switch port which doesn't support ACS thus cannot
route p2p transaction to IOMMU. If talking about typical PCIe
endpoint (with upstreaming ports all supporting ACS), you'll get
one device per group.
One iommu group today is attached to only one iommu domain.
In the future one group may attach to multiple domains, as the
aux domain concept being discussed in another thread.
Maybe they are kernel option to avoid that and userspace init program
can definitly re-arrange that base on sysadmin policy).
I don't think there is such option, as it may break isolation model
enabled by IOMMU.
[...]
quoted
quoted
That is why i am being pedantic :) on making sure there is good reasons
to do what you do inside VFIO. I do believe that we want a common
frame-
quoted
quoted
work like the one you are proposing but i do not believe it should be
part of VFIO given the baggages it comes with and that are not relevant
to the use cases for this kind of devices.
The purpose of VFIO is clear - the kernel portal for granting generic
device resource (mmio, irq, etc.) to user space. VFIO doesn't care
what exactly a resource is used for (queue, cmd reg, etc.). If really
pursuing VFIO path is necessary, maybe such common framework
should lay down in user space, which gets all granted resource from
kernel driver thru VFIO and then provides accelerator services to
other processes?
Thanks
Kevin
From: Kenneth Lee <hidden> Date: 2018-09-14 13:07:58
On Fri, Sep 14, 2018 at 06:50:55AM +0000, Tian, Kevin wrote:
Date: Fri, 14 Sep 2018 06:50:55 +0000
From: "Tian, Kevin" <kevin.tian@intel.com>
To: Jerome Glisse <redacted>, Kenneth Lee <redacted>
CC: Kenneth Lee <redacted>, Alex Williamson
[off-list ref], Herbert Xu [off-list ref],
"kvm@vger.kernel.org" [off-list ref], Jonathan Corbet
[off-list ref], Greg Kroah-Hartman [off-list ref], Zaibo
Xu [off-list ref], "linux-doc@vger.kernel.org"
[off-list ref], "Kumar, Sanjay K" [off-list ref],
Hao Fang [off-list ref], "linux-kernel@vger.kernel.org"
[off-list ref], "linuxarm@huawei.com"
[off-list ref], "iommu@lists.linux-foundation.org"
[off-list ref], "David S . Miller"
[off-list ref], "linux-crypto@vger.kernel.org"
[off-list ref], Zhou Wang [off-list ref],
Philippe Ombredanne [off-list ref], Thomas Gleixner
[off-list ref], Joerg Roedel [off-list ref],
"linux-accelerators@lists.ozlabs.org"
[off-list ref], Lu Baolu [off-list ref]
Subject: RE: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
Message-ID: [off-list ref]
quoted
From: Jerome Glisse
Sent: Thursday, September 13, 2018 10:52 PM
[...]
> AFAIK, on x86 and PPC at least, all PCIE devices are in the same group
quoted
by default at boot or at least all devices behind the same bridge.
the group thing reflects physical hierarchy limitation, not changed
cross boot. Please note iommu group defines the minimal isolation
boundary - all devices within same group must be attached to the
same iommu domain or address space, because physically IOMMU
cannot differentiate DMAs out of those devices. devices behind
legacy PCI-X bridge is one example. other examples include devices
behind a PCIe switch port which doesn't support ACS thus cannot
route p2p transaction to IOMMU. If talking about typical PCIe
endpoint (with upstreaming ports all supporting ACS), you'll get
one device per group.
One iommu group today is attached to only one iommu domain.
In the future one group may attach to multiple domains, as the
aux domain concept being discussed in another thread.
quoted
Maybe they are kernel option to avoid that and userspace init program
can definitly re-arrange that base on sysadmin policy).
I don't think there is such option, as it may break isolation model
enabled by IOMMU.
[...]
quoted
quoted
quoted
That is why i am being pedantic :) on making sure there is good reasons
to do what you do inside VFIO. I do believe that we want a common
frame-
quoted
quoted
work like the one you are proposing but i do not believe it should be
part of VFIO given the baggages it comes with and that are not relevant
to the use cases for this kind of devices.
The purpose of VFIO is clear - the kernel portal for granting generic
device resource (mmio, irq, etc.) to user space. VFIO doesn't care
what exactly a resource is used for (queue, cmd reg, etc.). If really
pursuing VFIO path is necessary, maybe such common framework
should lay down in user space, which gets all granted resource from
kernel driver thru VFIO and then provides accelerator services to
other processes?
Yes. I think this is exactly what WarpDrive is now doing. This patch is just let
the type1 driver use parent IOMMU for mdev.
Thanks
Kevin
--
-Kenneth(Hisilicon)
================================================================================
本邮件及其附件含有华为公司的保密信息,仅限于发送给上面地址中列出的个人或群组。禁
止任何其他人以任何形式使用(包括但不限于全部或部分地泄露、复制、或散发)本邮件中
的信息。如果您错收了本邮件,请您立即电话或邮件通知发件人并删除本邮件!
This e-mail and its attachments contain confidential information from HUAWEI,
which is intended only for the person or entity whose address is listed above.
Any use of the
information contained herein in any way (including, but not limited to, total or
partial disclosure, reproduction, or dissemination) by persons other than the
intended
recipient(s) is prohibited. If you receive this e-mail in error, please notify
the sender by phone or email immediately and delete it!
On Mon, Sep 03, 2018 at 08:51:57AM +0800, Kenneth Lee wrote:
[...]
quoted
quoted
quoted
I took a look at i915_gem_execbuffer_ioctl(). It seems it "copy_from_user" the
user memory to the kernel. That is not what we need. What we try to get is: the
user application do something on its data, and push it away to the accelerator,
and says: "I'm tied, it is your turn to do the job...". Then the accelerator has
the memory, referring any portion of it with the same VAs of the application,
even the VAs are stored inside the memory itself.
You were not looking at right place see drivers/gpu/drm/i915/i915_gem_userptr.c
It does GUP and create GEM object AFAICR you can wrap that GEM object into a
dma buffer object.
Thank you for directing me to this implementation. It is interesting:).
But it is not yet solve my problem. If I understand it right, the userptr in
i915 do the following:
1. The user process sets a user pointer with size to the kernel via ioctl.
2. The kernel wraps it as a dma-buf and keeps the process's mm for further
reference.
3. The user pages are allocated, GUPed or DMA mapped to the device. So the data
can be shared between the user space and the hardware.
But my scenario is:
1. The user process has some data in the user space, pointed by a pointer, say
ptr1. And within the memory, there may be some other pointers, let's say one
of them is ptr2.
2. Now I need to assign ptr1 *directly* to the hardware MMIO space. And the
hardware must refer ptr1 and ptr2 *directly* for data.
Userptr lets the hardware and process share the same memory space. But I need
them to share the same *address space*. So IOMMU is a MUST for WarpDrive,
NOIOMMU mode, as Jean said, is just for verifying some of the procedure is OK.
So to be 100% clear should we _ignore_ the non SVA/SVM case ?
If so then wait for necessary SVA/SVM to land and do warp drive
without non SVA/SVM path.
I think we should clear the concept of SVA/SVM here. As my understanding, Share
Virtual Address/Memory means: any virtual address in a process can be used by
device at the same time. This requires IOMMU device to support PASID. And
optionally, it requires the feature of page-fault-from-device.
Yes we agree on what SVA/SVM is. There is a one gotcha thought, access
to range that are MMIO map ie CPU page table pointing to IO memory, IIRC
it is undefined what happens on some platform for a device trying to
access those using SVA/SVM.
quoted
But before the feature is settled down, IOMMU can be used immediately in the
current kernel. That make it possible to assign ONE process's virtual addresses
to the device's IOMMU page table with GUP. This make WarpDrive work well for one
process.
UH ? How ? You want to GUP _every_ single valid address in the process
and map it to the device ? How do you handle new vma, page being replace
(despite GUP because of things that utimately calls zap pte) ...
Again here you said that the device must be able to access _any_ valid
pointer. With GUP this is insane.
So i am assuming this is not what you want to do without SVA/SVM ie with
GUP you have a different programming model, one in which the userspace
must first bind _range_ of memory to the device and get a DMA address
for the range.
Again, GUP range of process address space to map it to a device so that
userspace can use the device on the mapped range is something that do
exist in various places in the kernel.
Yes same as your expectation, in WarpDrive, we use the concept of "sharing" to
do so. If some memory is going to be shared among process and devices, we use
wd_share_mem(queue, ptr, size) to share those memory. When the queue is working
in this mode, the point is valid in those memory segments. The wd_share_mem call
vfio dma map syscall which will do GUP.
If SVA/SVM is enabled, user space can set SHARE_ALL flags to the queue. Then
wd_share_mem() is not necessary.
This is really not popular when we started the work on WarpDrive. The GUP
document said it should be put within the scope of mm_sem is locked. Because GUP
simply increase the page refcount, not keep the mapping between the page and the
vma. We keep our work together with VFIO to make sure the problem can be solved
in one deal.
The problem can not be solved in one deal, you can not maintain vaddr
pointing to same page after a fork() this can not be solve without the
use of mmu notifier and device dma mapping invalidation ! So being part
of VFIO will not help you there.
Good point. But sadly, even with mmu notifier and dma mapping invalidation, I
cannot do anything here. If the process fork a sub-process, the sub-process need
a new pasid and hardware resource. The IOMM space mapped should not be used. The
parent process should be aware of this, unmap and close the device file before
the fork. I have the same limitation as VFIO:(
I don't think I can change much here. If I can, VFIO can too:)
The forbid child to access the device is easy in the kernel whenever
someone open the device file force set the OCLOEXEC flag on the file
some device driver already do that and so should you. With that you
should always have a struct file - mm struct one to one relationship
and thus one PASID per struct file ie per open of the device file.
I considerred the OCLOEXEC flag, but it seams it works only for exec, not fork.
Every device driver i know protect against that (child having access)
by at least setting VM_DONTCOPY on mmap of device IO (or even all mmap
of device file).
But you right OCLOEXEC does nothing on fork() sometimes i forget posix.
quoted
That does not solve the GUP/fork issue i describe below.
quoted
quoted
AFAIK VFIO is fine with the way it is as QEMU do not fork() once it
is running a guest and thus the COW that would invalidate vaddr to
physical page assumption is not broken. So i doubt VFIO folks have
any incentive to go down the mmu notifier path and invalidate device
mapping. They also have the replay thing that probably handle some
of fork cases by trusting user space program to do it. In your case
you can not trust the user space program.
In your case AFAICT i do not see any warning or gotcha so the following
scenario is broken (in non SVA/SVM):
1) program setup the device (open container, mdev, setup queue, ...)
2) program map some range of its address space wih VFIO_IOMMU_MAP_DMA
3) program start using the device using map setup in 2)
...
4) program fork()
5) parent trigger COW inside the range setup in 2)
At this point it is the child process that can write to the page that
are access by the device (which was map by the parent in 2)). The
parent can no longer access that memory from the CPU.
There is just no sane way to fix this beside invalidating device mapping
on fork (and you can not rely on userspace to do so) and thus stopping
the device on fork (SVA/SVM case do not have any issue here).
Indeed. But as soon as we choose to expose the device space to the user space,
the limitation is already there. If we want to solve the problem, we have to
have a hook in the copy_process() procedure and copy the parent's queue state to
a new queue, assign it to the child's fd and redirect the child's mmap to
it. If I can do so, the same logic can also be applied to VFIO.
Except we do not want to do that and this does not solve the COW i describe
above unless you disable COW altogether which is a big no.
quoted
The good side is, this is not a security leak. The hardware has been given to
the process. It is the process who choose to share it. If it won't work, it is
the process's problem;)
No this is bad, you can not expect every single userspace program to know and
be aware of that. If some trusted application (say systemd or firefox, ...)
start using your device and is unaware or does not comprehend all side effect
it would allow the child to access/change its parent memory and that is bad.
Device driver need to protect against user doing stupid thing. All existing
device driver that do ATS/PASID already protect themself against child (ie the
child can not use the device through a mix of OCLOEXEC and other checks).
You must set the OCLOEXEC flag but that does not solve the COW problem above.
My motto in life is "do not trust userspace" :) which also translate to "do
not expect userspace will do the right thing".
We don't really trust user space here. We trust that the process cannot do more
than what has been given to it. If an application use WarpDrive, and share its
memory with it, it should know what happen when it fork a sub-process. This is
just like you clone a sub-process with CLONE_FILES or CLONE_VM, the parent
process should know what happen.
Except application do not always know, the whole container business is
an example of that. Most application have tons of dependency and reliance
on various library. There is no way they audit all the library all the
times.
So if a library decide to use such device then this trickles down into
application without their knowledge and you can open security issues
through that.
To be clear here, what worry me is the non SVA/SVM case, there is no way
for the device driver to intercept fork and thus noways for it to break
COW for the parent, so child could potentialy write hardware commands for
the device (depends on the device if for instance the cmd queue mmap only
point to regular memory from which it fetches its commands). I guess relying
on IOMMU might be good enough (ie only range mapped by the first user of
the hardware would be accessible).
quoted
quoted
quoted
quoted
And now we have GUP-longterm and many accounting work in VFIO, we don't want to
do that again.
GUP-longterm does not solve any GUP problem, it just block people to
do GUP on DAX backed vma to avoid pining persistent memory as it is
a nightmare to handle in the block device driver and file system code.
The accounting is the rt limit thing and is litteraly 10 lines of
code so i would not see that as hard to replicate.
OK. Agree.
quoted
quoted
quoted
quoted
Now We are talking about SVA and PASID, just to make sure WarpDrive can benefit
from the feature in the future. It dose not means WarpDrive is useless before
that. And it works for our Zip and RSA accelerators in physical world.
Just not with random process address ...
quoted
quoted
If you still want non SVA/SVM path what you want to do only works
if both ptr1 and ptr2 are in a range that is DMA mapped to the
device (moreover you need DMA address to match process address
which is not an easy feat).
Now even if you only want SVA/SVM, i do not see what is the point
of doing this inside VFIO. AMD GPU driver does not and there would
be no benefit for them to be there. Well a AMD VFIO mdev device
driver for QEMU guest might be useful but they have SVIO IIRC.
For SVA/SVM your usage model is:
Setup:
- user space create a warp drive context for the process
- user space create a device specific context for the process
- user space create a user space command queue for the device
- user space bind command queue
At this point the kernel driver has bound the process address
space to the device with a command queue and userspace
Usage:
- user space schedule work and call appropriate flush/update
ioctl from time to time. Might be optional depends on the
hardware, but probably a good idea to enforce so that kernel
can unbind the command queue to bind another process command
queue.
...
Cleanup:
- user space unbind command queue
- user space destroy device specific context
- user space destroy warp drive context
All the above can be implicit when closing the device file.
So again in the above model i do not see anywhere something from
VFIO that would benefit this model.
Let me show you how the model will be if I use VFIO:
Setup (Kernel part)
- Kernel driver do every as usual to serve the other functionality, NIC
can still be registered to netdev, encryptor can still be registered
to crypto...
- At the same time, the driver can devote some of its hardware resource
and register them as a mdev creator to the VFIO framework. This just
need limited change to the VFIO type1 driver.
In the above VFIO does not help you one bit ... you can do that with
as much code with new common device as front end.
quoted
Setup (User space)
- System administrator create mdev via the mdev creator interface.
- Following VFIO setup routine, user space open the mdev's group, there is
only one group for one device.
- Without PASID support, you don't need to do anything. With PASID, bind
the PASID to the device via VFIO interface.
- Get the device from the group via VFIO interface and mmap it the user
space for device's MMIO access (for the queue).
- Map whatever memory you need to share with the device with VFIO
interface.
- (opt) Add more devices into the container if you want to share the
same address space with them
So all VFIO buys you here is boiler plate code that does insert_pfn()
to handle MMIO mapping. Which is just couple hundred lines of boiler
plate code.
No. With VFIO, I don't need to:
1. GUP and accounting for RLIMIT_MEMLOCK
That's 10 line of code ...
quoted
2. Keep all GUP pages for releasing (VFIO uses the rb_tree to do so)
GUP pages are not part of rb_tree and what you want to do can be done
in few lines of code here is pseudo code:
warp_dma_map_range(ulong vaddr, ulong npages)
{
struct page *pages = kvzalloc(npages);
for (i = 0; i < npages; ++i, vaddr += PAGE_SIZE) {
GUP(vaddr, &pages[i]);
iommu_map(vaddr, page_to_pfn(pages[i]));
}
kvfree(pages);
}
warp_dma_unmap_range(ulong vaddr, ulong npages)
{
for (i = 0; i < npages; ++i, vaddr += PAGE_SIZE) {
unsigned long pfn;
pfn = iommu_iova_to_phys(vaddr);
iommu_unmap(vaddr);
put_page(pfn_to_page(page)); /* set dirty if mapped write */
}
}
But what if the process exist without unmapping? The pages will be pinned in the
kernel forever.
Yeah add a struct warp_map { struct list_head list; unsigned long vaddr,
unsigned long npages; } for every mapping, store the head into the device
file private field of struct file and when the release fs callback is
call you can walk done the list to force unmap any leftover. This is not
that much code. You ca even use interval tree which is 3 lines of code
with interval_tree_generic.h to speed up warp_map lookup on unmap ioctl.
Yes, when you add all of them... it is VFIO:)
No, VFIO not only need to track all mapping but it also need to track iova
to vaddr which you do not need. The VFIO code is much more complex because
of that. What you would need is only ~200lines of code.
quoted
quoted
quoted
Add locking, error handling, dirtying and comments and you are barely
looking at couple hundred lines of code. You do not need any of the
complexity of VFIO as you do not have the same requirements. Namely
VFIO have to keep track of iova and physical mapping for things like
migration (migrating guest between host) and few others very
virtualization centric requirements.
quoted
2. Handle the PASID on SMMU (ARM's IOMMU) myself.
Existing driver do that with 20 lines of with comments and error
handling (see kfd_iommu_bind_process_to_device() for instance) i
doubt you need much more than that.
OK, I agree.
quoted
quoted
3. Multiple devices menagement (VFIO uses container to manage this)
All the vfio_group* stuff ? OK that's boiler plate code, note that
hard to replicate thought.
No, I meant the container thing. Several devices/group can be assigned to the
same container and the DMA on the container can be assigned to all those
devices. So we can have some devices to share the same name space.
This was the motivation of my question below, to me this is a policy
decision and it should be left to userspace to decide but not forced
upon userspace because it uses a given device driver.
Maybe i am wrong but i think you can create container and device group
without having a VFIO driver for the devices in the group. It is not
something i do often so i might be wrong here.
Container maintains a virtual unify address space for all group/iommu which is
added to it. It simplify the address management. But yes, you can choose to do
it all in the user space.
quoted
quoted
quoted
quoted
And even as a boiler plate, it is valueable, the memory thing is sensitive
interface to user space, it can easily become a security problem. If I can
achieve my target within the scope of VFIO, why not? At lease it has been
proved to be safe for the time being.
The thing is being part of VFIO impose things on you, things that you
do not need. Like one device per group (maybe it is you imposing this,
i am loosing track here). Or the complex dma mapping tracking ...
Err... But the one-device-per-group is not VFIO's decision. It is IOMMU's :).
Unless I don't use IOMMU.
AFAIK, on x86 and PPC at least, all PCIE devices are in the same group
by default at boot or at least all devices behind the same bridge.
Maybe they are kernel option to avoid that and userspace init program
can definitly re-arrange that base on sysadmin policy).
But if the IOMMU is enabled, all PCIE devices have their own device IDs. So the
IOMMU can use different page table for every of them.
On PCIE AFAIR IOMMU can not differentiate between devices, hence why
by default they end up with the same domain in the same group.
quoted
quoted
quoted
quoted
quoted
quoted
Cleanup:
- User space close the group file handler
- There will be a problem to let the other process know the mdev is
freed to be used again. My RFCv1 choose a file handler solution. Alex
dose not like it. But it is not a big problem. We can always have a
scheduler process to manage the state of the mdev or even we can
switch back to the RFCv1 solution without too much effort if we like
in the future.
If you were outside VFIO you would have more freedom on how to do that.
For instance process opening the device file can be placed on queue and
first one in the queue get to use the device until it closes/release the
device. Then next one in queue get the device ...
Yes. I do like the file handle solution. But I hope the solution become mature
as soon as possible. Many of our products, and as I know include some of our
partners, are waiting for a long term solution as direction. If I rely on some
unmature solution, they may choose some deviated, customized solution. That will
be much harmful. Compare to this, the freedom is not so important...
I do not see how being part of VFIO protect you from people doing crazy
thing to their kernel ... Time to market being key in this world, i doubt
that being part of VFIO would make anyone think twice before taking a
shortcut.
I have seen horrible things on that front and only players like Google
can impose a minimum level of sanity.
OK. My fault, to talk about TTM. It has nothing doing with the architecture
decision. But I don't yet see what harm will be brought if I use VFIO when it
can fulfill almost all my requirements.
The harm is in forcing the device group isolation policy which is not
necessary for all devices like you said so yourself on ARM with the
device stream id so that IOMMU can identify individual devices.
No. Some mini SoC share the same stream id among several small devices. The
iommu has to treat them as the same. That is why IOMMU introduces the concept of
iommu_group. Personally I don't like that. If they share the same stream id,
they should be treated as the same hardware. But it is the decision for the time
being...
Still policy decission ie how to group device together should be some-
thing left to the sysadmin/distribution. When hardware have constraint
it just limits the choice of userspace.
quoted
So i would rather see the device isolation as something orthogonal to
what you want to achieve and that should be forced upon user ie sysadmin
should control that distribution can have sane default for each platform.
quoted
quoted
quoted
quoted
quoted
Except for the minimum update to the type1 driver and use sdmdev to manage the
interrupt sharing, I don't need any extra code to gain the address sharing
capability. And the capability will be strengthen along with the upgrade of VFIO.
quoted
quoted
quoted
quoted
And I don't understand why I should avoid to use VFIO? As Alex said, VFIO is the
user driver framework. And I need exactly a user driver interface. Why should I
invent another wheel? It has most of stuff I need:
1. Connecting multiple devices to the same application space
2. Pinning and DMA from the application space to the whole set of device
3. Managing hardware resource by device
We just need the last step: make sure multiple applications and the kernel can
share the same IOMMU. Then why shouldn't we use VFIO?
Because tons of other drivers already do all of the above outside VFIO. Many
driver have a sizeable userspace side to them (anything with ioctl do) so they
can be construded as userspace driver too.
Ignoring if there are *tons* of drivers are doing that;), even I do the same as
i915 and solve the address space problem. And if I don't need to with VFIO, why
should I spend so much effort to do it again?
Because you do not need any code from VFIO, nor do you need to reinvent
things. If non SVA/SVM matters to you then use dma buffer. If not then
i do not see anything in VFIO that you need.
As I have explain, if I don't use VFIO, at lease I have to do all that has been
done in i915 or even more than that.
So beside the MMIO mmap() handling and dma mapping of range of user space
address space (again all very boiler plate code duplicated accross the
kernel several time in different forms). You do not gain anything being
inside VFIO right ?
As I said, rb-tree for gup, rlimit accounting, cooperation on SMMU, and mature
user interface are our concern.
quoted
quoted
quoted
quoted
quoted
So there is no reasons to do that under VFIO. Especialy as in your example
it is not a real user space device driver, the userspace portion only knows
about writting command into command buffer AFAICT.
VFIO is for real userspace driver where interrupt, configurations, ... ie
all the driver is handled in userspace. This means that the userspace have
to be trusted as it could program the device to do DMA to anywhere (if
IOMMU is disabled at boot which is still the default configuration in the
kernel).
But as Alex explained, VFIO is not simply used by VM. So it need not to have all
stuffs as a driver in host system. And I do need to share the user space as DMA
buffer to the hardware. And I can get it with just a little update, then it can
service me perfectly. I don't understand why I should choose a long route.
Again this is not the long route i do not see anything in VFIO that
benefit you in the SVA/SVM case. A basic character device driver can
do that.
quoted
quoted
So i do not see any reasons to do anything you want inside VFIO. All you
want to do can be done outside as easily. Moreover it would be better if
you define clearly each scenario because from where i sit it looks like
you are opening the door wide open to userspace to DMA anywhere when IOMMU
is disabled.
When IOMMU is disabled you can _not_ expose command queue to userspace
unless your device has its own page table and all commands are relative
to that page table and the device page table is populated by kernel driver
in secure way (ie by checking that what is populated can be access).
I do not believe your example device to have such page table nor do i see
a fallback path when IOMMU is disabled that force user to do ioctl for
each commands.
Yes i understand that you target SVA/SVM but still you claim to support
non SVA/SVM. The point is that userspace can not be trusted if you want
to have random program use your device. I am pretty sure that all user
of VFIO are trusted process (like QEMU).
Finaly i am convince that the IOMMU grouping stuff related to VFIO is
useless for your usecase. I really do not see the point of that, it
does complicate things for you for no reasons AFAICT.
Indeed, I don't like the group thing. I believe VFIO's maintains would not like
it very much either;). But the problem is, the group reflects to the same
IOMMU(unit), which may shared with other devices. It is a security problem. I
cannot ignore it. I have to take it into account event I don't use VFIO.
To me it seems you are making a policy decission in kernel space ie
wether the device should be isolated in its own group or not is a
decission that is up to the sys admin or something in userspace.
Right now existing user of SVA/SVM don't (at least AFAICT).
Do we really want to force such isolation ?
But it is not my decision, that how the iommu subsystem is designed. Personally
I don't like it at all, because all our hardwares have their own stream id
(device id). I don't need the group concept at all. But the iommu subsystem
assume some devices may share the name device ID to a single IOMMU.
My question was do you really want to force group isolation for the
device ? Existing SVA/SVM capable driver do not force that, they let
the userspace decide this (sysadm, distributions, ...). Being part of
VFIO (in the way you do, likely ways to avoid this inside VFIO too)
force this decision ie make a policy decision without userspace having
anything to say about it.
You still do not answer my question, do you really want to force group
isolation for device in your framework ? Which is a policy decision from
my POV and thus belong to userspace and should not be enforce by kernel.
No. But I have to follow the rule defined by IOMMU, haven't I?
The IOMMU rule does not say that every device _must_ always be in one
group and one domain only. I am pretty sure on x86 by default you get
one domain for all PCIE devices behind same bridge.
Really? Give me some more time. I need to check it out.
On my computer right now i have one group per pcie bridge and all
devices behind same bridge in same groups.
ls /sys/kernel/iommu_groups/*/devices
quoted
My point is that the device grouping into domain/group should be an
orthogonal decision ie it should not be a requirement by the device
driver and should be under control of userspace as it is a policy
decission.
One exception if it is unsafe to have device share a domain in which
case following my motto the driver should refuse to work and return
an error on open (and a kernel explaining why). But this depends on
device and platform.
On Fri, Sep 14, 2018 at 06:50:55AM +0000, Tian, Kevin wrote:
quoted
From: Jerome Glisse
Sent: Thursday, September 13, 2018 10:52 PM
[...]
> AFAIK, on x86 and PPC at least, all PCIE devices are in the same group
quoted
by default at boot or at least all devices behind the same bridge.
the group thing reflects physical hierarchy limitation, not changed
cross boot. Please note iommu group defines the minimal isolation
boundary - all devices within same group must be attached to the
same iommu domain or address space, because physically IOMMU
cannot differentiate DMAs out of those devices. devices behind
legacy PCI-X bridge is one example. other examples include devices
behind a PCIe switch port which doesn't support ACS thus cannot
route p2p transaction to IOMMU. If talking about typical PCIe
endpoint (with upstreaming ports all supporting ACS), you'll get
one device per group.
One iommu group today is attached to only one iommu domain.
In the future one group may attach to multiple domains, as the
aux domain concept being discussed in another thread.
Thanks for the info.
quoted
Maybe they are kernel option to avoid that and userspace init program
can definitly re-arrange that base on sysadmin policy).
I don't think there is such option, as it may break isolation model
enabled by IOMMU.
[...]
quoted
quoted
quoted
That is why i am being pedantic :) on making sure there is good reasons
to do what you do inside VFIO. I do believe that we want a common
frame-
quoted
quoted
work like the one you are proposing but i do not believe it should be
part of VFIO given the baggages it comes with and that are not relevant
to the use cases for this kind of devices.
The purpose of VFIO is clear - the kernel portal for granting generic
device resource (mmio, irq, etc.) to user space. VFIO doesn't care
what exactly a resource is used for (queue, cmd reg, etc.). If really
pursuing VFIO path is necessary, maybe such common framework
should lay down in user space, which gets all granted resource from
kernel driver thru VFIO and then provides accelerator services to
other processes?
Except that many existing device driver falls under that description
(ie exposing mmio, command queues, ...) and are not under VFIO.
Up to mdev VFIO was all about handling a full device to userspace AFAIK.
With the introduction of mdev a host kernel driver can "slice" its
device and share it through VFIO to userspace. Note that in that case
it might never end over any mmio, irq, ... the host driver might just
be handling over memory and would be polling from it to schedule on
the real hardware.
The question i am asking about warpdrive is wether being in VFIO is
necessary ? as i do not see the requirement myself.
Cheers,
Jérôme
So i want to summarize issues i have as this threads have dig deep into
details. For this i would like to differentiate two cases first the easy
one when relying on SVA/SVM. Then the second one when there is no SVA/SVM.
In both cases your objectives as i understand them:
[R1]- expose a common user space API that make it easy to share boiler
plate code accross many devices (discovering devices, opening
device, creating context, creating command queue ...).
[R2]- try to share the device as much as possible up to device limits
(number of independant queues the device has)
[R3]- minimize syscall by allowing user space to directly schedule on the
device queue without a round trip to the kernel
I don't think i missed any.
(1) Device with SVA/SVM
For that case it is easy, you do not need to be in VFIO or part of any
thing specific in the kernel. There is no security risk (modulo bug in
the SVA/SVM silicon). Fork/exec is properly handle and binding a process
to a device is just couple dozen lines of code.
(2) Device does not have SVA/SVM (or it is disabled)
You want to still allow device to be part of your framework. However
here i see fundamentals securities issues and you move the burden of
being careful to user space which i think is a bad idea. We should
never trus the userspace from kernel space.
To keep the same API for the user space code you want a 1:1 mapping
between device physical address and process virtual address (ie if
device access device physical address A it is accessing the same
memory as what is backing the virtual address A in the process.
Security issues are on two things:
[I1]- fork/exec, a process who opened any such device and created an
active queue can transfer without its knowledge control of its
commands queue through COW. The parent map some anonymous region
to the device as a command queue buffer but because of COW the
parent can be the first to copy on write and thus the child can
inherit the original pages that are mapped to the hardware.
Here parent lose control and child gain it.
[I2]- Because of [R3] you want to allow userspace to schedule commands
on the device without doing an ioctl and thus here user space
can schedule any commands to the device with any address. What
happens if that address have not been mapped by the user space
is undefined and in fact can not be defined as what each IOMMU
does on invalid address access is different from IOMMU to IOMMU.
In case of a bad IOMMU, or simply an IOMMU improperly setup by
the kernel, this can potentialy allow user space to DMA anywhere.
[I3]- By relying on GUP in VFIO you are not abiding by the implicit
contract (at least i hope it is implicit) that you should not
try to map to the device any file backed vma (private or share).
The VFIO code never check the vma controlling the addresses that
are provided to VFIO_IOMMU_MAP_DMA ioctl. Which means that the
user space can provide file backed range.
I am guessing that the VFIO code never had any issues because its
number one user is QEMU and QEMU never does that (and that's good
as no one should ever do that).
So if process does that you are opening your self to serious file
system corruption (depending on file system this can lead to total
data loss for the filesystem).
Issue is that once you GUP you never abide to file system flushing
which write protect the page before writing to the disk. So
because the page is still map with write permission to the device
(assuming VFIO_IOMMU_MAP_DMA was a write map) then the device can
write to the page while it is in the middle of being written back
to disk. Consult your nearest file system specialist to ask him
how bad that can be.
[I4]- Design issue, mdev design As Far As I Understand It is about
sharing a single device to multiple clients (most obvious case
here is again QEMU guest). But you are going against that model,
in fact AFAIUI you are doing the exect opposite. When there is
no SVA/SVM you want only one mdev device that can not be share.
So this is counter intuitive to the mdev existing design. It is
not about sharing device among multiple users but about giving
exclusive access to the device to one user.
All the reasons above is why i believe a different model would serve
you and your user better. Below is a design that avoids all of the
above issues and still delivers all of your objectives with the
exceptions of the third one [R3] when there is no SVA/SVM.
Create a subsystem (very much boiler plate code) which allow device to
register themself against (very much like what you do in your current
patchset but outside of VFIO).
That subsystem will create a device file for each registered system and
expose a common API (ie set of ioctl) for each of those device files.
When user space create a queue (through an ioctl after opening the device
file) the kernel can return -EBUSY if all the device queue are in use,
or create a device queue and return a flag like SYNC_ONLY for device that
do not have SVA/SVM.
For device with SVA/SVM at the time the process create a queue you bind
the process PASID to the device queue. From there on the userspace can
schedule commands and use the device without going to kernel space.
For device without SVA/SVM you create a fake queue that is just pure
memory is not related to the device. From there on the userspace must
call an ioctl every time it wants the device to consume its queue
(hence why the SYNC_ONLY flag for synchronous operation only). The
kernel portion read the fake queue expose to user space and copy
commands into the real hardware queue but first it properly map any
of the process memory needed for those commands to the device and
adjust the device physical address with the one it gets from dma_map
API.
With that model it is "easy" to listen to mmu_notifier and to abide by
them to avoid issues [I1], [I3] and [I4]. You obviously avoid the [I2]
issue by only mapping a fake device queue to userspace.
So yes with that models it means that every device that wish to support
the non SVA/SVM case will have to do extra work (ie emulate its command
queue in software in the kernel). But by doing so, you support an
unlimited number of process on your device (ie all the process can share
one single hardware command queues or multiple hardware queues).
The big advantages i see here is that the process do not have to worry
about doing something wrong. You are protecting yourself and your user
from stupid mistakes.
I hope this is useful to you.
Cheers,
Jérôme
From: Kenneth Lee <hidden> Date: 2018-09-17 08:41:52
On Sun, Sep 16, 2018 at 09:42:44PM -0400, Jerome Glisse wrote:
Date: Sun, 16 Sep 2018 21:42:44 -0400
From: Jerome Glisse <redacted>
To: Kenneth Lee <redacted>
CC: Jonathan Corbet <corbet@lwn.net>, Herbert Xu
[off-list ref], "David S . Miller" [off-list ref],
Joerg Roedel [off-list ref], Alex Williamson
[off-list ref], Kenneth Lee [off-list ref], Hao
Fang [off-list ref], Zhou Wang [off-list ref], Zaibo Xu
[off-list ref], Philippe Ombredanne [off-list ref], Greg
Kroah-Hartman [off-list ref], Thomas Gleixner
[off-list ref], linux-doc@vger.kernel.org,
linux-kernel@vger.kernel.org, linux-crypto@vger.kernel.org,
iommu@lists.linux-foundation.org, kvm@vger.kernel.org,
linux-accelerators@lists.ozlabs.org, Lu Baolu [off-list ref],
Sanjay Kumar [off-list ref], linuxarm@huawei.com
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
User-Agent: Mutt/1.10.1 (2018-07-13)
Message-ID: [off-list ref]
So i want to summarize issues i have as this threads have dig deep into
details. For this i would like to differentiate two cases first the easy
one when relying on SVA/SVM. Then the second one when there is no SVA/SVM.
Thank you very much for the summary.
In both cases your objectives as i understand them:
[R1]- expose a common user space API that make it easy to share boiler
plate code accross many devices (discovering devices, opening
device, creating context, creating command queue ...).
[R2]- try to share the device as much as possible up to device limits
(number of independant queues the device has)
[R3]- minimize syscall by allowing user space to directly schedule on the
device queue without a round trip to the kernel
I don't think i missed any.
(1) Device with SVA/SVM
For that case it is easy, you do not need to be in VFIO or part of any
thing specific in the kernel. There is no security risk (modulo bug in
the SVA/SVM silicon). Fork/exec is properly handle and binding a process
to a device is just couple dozen lines of code.
This is right...logically. But the kernel has no clear definition about "Device
with SVA/SVM" and no boiler plate for doing so. Then VFIO may become one of the
boiler plate.
VFIO is one of the wrappers for IOMMU for user space. And maybe it is the only
one. If we add that support within VFIO, which solve most of the problem of
SVA/SVM, it will save a lot of work in the future.
I think this is the key confliction between us. So could Alex please say
something here? If the VFIO is going to take this into its scope, we can try
together to solve all the problem on the way. If it it is not, it is also
simple, we can just go to another way to fulfill this part of requirements even
we have to duplicate most of the code.
Another point I need to emphasis here: because we have to replace the hardware
queue when fork, so it won't be very simple even in SVA/SVM case.
(2) Device does not have SVA/SVM (or it is disabled)
You want to still allow device to be part of your framework. However
here i see fundamentals securities issues and you move the burden of
being careful to user space which i think is a bad idea. We should
never trus the userspace from kernel space.
To keep the same API for the user space code you want a 1:1 mapping
between device physical address and process virtual address (ie if
device access device physical address A it is accessing the same
memory as what is backing the virtual address A in the process.
Security issues are on two things:
[I1]- fork/exec, a process who opened any such device and created an
active queue can transfer without its knowledge control of its
commands queue through COW. The parent map some anonymous region
to the device as a command queue buffer but because of COW the
parent can be the first to copy on write and thus the child can
inherit the original pages that are mapped to the hardware.
Here parent lose control and child gain it.
This is indeed an issue. But it remains an issue only if you continue to use the
queue and the memory after fork. We can use at_fork kinds of gadget to fix it in
user space.
From some perspectives, I think the issue can be solved by iommu_notifier. For
example, when the process is fork-ed, we can set the mapped device mmio space as
COW for the child process, so a new queue can be created and set to the same
state as the parent's if the space is accessed. Then we can have two separated
queues for both the parent and the child. The memory part can be done in the
same way.
The thing is, the same strategy can be applied to VFIO without changing its
original feature.
[I2]- Because of [R3] you want to allow userspace to schedule commands
on the device without doing an ioctl and thus here user space
can schedule any commands to the device with any address. What
happens if that address have not been mapped by the user space
is undefined and in fact can not be defined as what each IOMMU
does on invalid address access is different from IOMMU to IOMMU.
In case of a bad IOMMU, or simply an IOMMU improperly setup by
the kernel, this can potentialy allow user space to DMA anywhere.
I don't think this is an issue. If you cannot trust IOMMU and proper setup of
IOMMU in kernel, you cannot trust anything. And the whole VFIO framework is
untrustable.
[I3]- By relying on GUP in VFIO you are not abiding by the implicit
contract (at least i hope it is implicit) that you should not
try to map to the device any file backed vma (private or share).
The VFIO code never check the vma controlling the addresses that
are provided to VFIO_IOMMU_MAP_DMA ioctl. Which means that the
user space can provide file backed range.
I am guessing that the VFIO code never had any issues because its
number one user is QEMU and QEMU never does that (and that's good
as no one should ever do that).
So if process does that you are opening your self to serious file
system corruption (depending on file system this can lead to total
data loss for the filesystem).
Issue is that once you GUP you never abide to file system flushing
which write protect the page before writing to the disk. So
because the page is still map with write permission to the device
(assuming VFIO_IOMMU_MAP_DMA was a write map) then the device can
write to the page while it is in the middle of being written back
to disk. Consult your nearest file system specialist to ask him
how bad that can be.
Same as I2, it is an issue, but the problem can be solved in VFIO if we really
take it in the scope of VFIO.
[I4]- Design issue, mdev design As Far As I Understand It is about
sharing a single device to multiple clients (most obvious case
here is again QEMU guest). But you are going against that model,
in fact AFAIUI you are doing the exect opposite. When there is
no SVA/SVM you want only one mdev device that can not be share.
Wait. It is NOT "I want only one mdev device when there is no SVA/SVM", it is "I
can support only one mdev when there is no PASID support for the IOMMU".
So this is counter intuitive to the mdev existing design. It is
not about sharing device among multiple users but about giving
exclusive access to the device to one user.
All the reasons above is why i believe a different model would serve
you and your user better. Below is a design that avoids all of the
above issues and still delivers all of your objectives with the
exceptions of the third one [R3] when there is no SVA/SVM.
Create a subsystem (very much boiler plate code) which allow device to
register themself against (very much like what you do in your current
patchset but outside of VFIO).
That subsystem will create a device file for each registered system and
expose a common API (ie set of ioctl) for each of those device files.
When user space create a queue (through an ioctl after opening the device
file) the kernel can return -EBUSY if all the device queue are in use,
or create a device queue and return a flag like SYNC_ONLY for device that
do not have SVA/SVM.
For device with SVA/SVM at the time the process create a queue you bind
the process PASID to the device queue. From there on the userspace can
schedule commands and use the device without going to kernel space.
As mentioned previously, this is not enough for fork scenario.
For device without SVA/SVM you create a fake queue that is just pure
memory is not related to the device. From there on the userspace must
call an ioctl every time it wants the device to consume its queue
(hence why the SYNC_ONLY flag for synchronous operation only). The
kernel portion read the fake queue expose to user space and copy
commands into the real hardware queue but first it properly map any
of the process memory needed for those commands to the device and
adjust the device physical address with the one it gets from dma_map
API.
But in this way, we will lost most of the benefit of avoiding syscall.
With that model it is "easy" to listen to mmu_notifier and to abide by
them to avoid issues [I1], [I3] and [I4]. You obviously avoid the [I2]
issue by only mapping a fake device queue to userspace.
So yes with that models it means that every device that wish to support
the non SVA/SVM case will have to do extra work (ie emulate its command
queue in software in the kernel). But by doing so, you support an
unlimited number of process on your device (ie all the process can share
one single hardware command queues or multiple hardware queues).
If I can do this, I will not need WarpDrive at all:(
The big advantages i see here is that the process do not have to worry
about doing something wrong. You are protecting yourself and your user
from stupid mistakes.
I hope this is useful to you.
Anyway, I will try to address the problem you mentioned in next version, and
make both options (on-VFIO or off-VFIO) are available. Thanks.
On Mon, Sep 17, 2018 at 04:39:40PM +0800, Kenneth Lee wrote:
On Sun, Sep 16, 2018 at 09:42:44PM -0400, Jerome Glisse wrote:
quoted
So i want to summarize issues i have as this threads have dig deep into
details. For this i would like to differentiate two cases first the easy
one when relying on SVA/SVM. Then the second one when there is no SVA/SVM.
Thank you very much for the summary.
quoted
In both cases your objectives as i understand them:
[R1]- expose a common user space API that make it easy to share boiler
plate code accross many devices (discovering devices, opening
device, creating context, creating command queue ...).
[R2]- try to share the device as much as possible up to device limits
(number of independant queues the device has)
[R3]- minimize syscall by allowing user space to directly schedule on the
device queue without a round trip to the kernel
I don't think i missed any.
(1) Device with SVA/SVM
For that case it is easy, you do not need to be in VFIO or part of any
thing specific in the kernel. There is no security risk (modulo bug in
the SVA/SVM silicon). Fork/exec is properly handle and binding a process
to a device is just couple dozen lines of code.
This is right...logically. But the kernel has no clear definition about "Device
with SVA/SVM" and no boiler plate for doing so. Then VFIO may become one of the
boiler plate.
VFIO is one of the wrappers for IOMMU for user space. And maybe it is the only
one. If we add that support within VFIO, which solve most of the problem of
SVA/SVM, it will save a lot of work in the future.
You do not need to "wrap" IOMMU for SVA/SVM. Existing upstream SVA/SVM user
all do the SVA/SVM setup in couple dozen lines and i failed to see how it
would require any more than that in your case.
I think this is the key confliction between us. So could Alex please say
something here? If the VFIO is going to take this into its scope, we can try
together to solve all the problem on the way. If it it is not, it is also
simple, we can just go to another way to fulfill this part of requirements even
we have to duplicate most of the code.
Another point I need to emphasis here: because we have to replace the hardware
queue when fork, so it won't be very simple even in SVA/SVM case.
I am assuming hardware queue can only be setup by the kernel and thus
you are totaly safe forkwise as the queue is setup against a PASID and
the child does not bind to any PASID and you use VM_DONTCOPY on the
mmap of the hardware MMIO queue because you should really use that flag
for that.
quoted
(2) Device does not have SVA/SVM (or it is disabled)
You want to still allow device to be part of your framework. However
here i see fundamentals securities issues and you move the burden of
being careful to user space which i think is a bad idea. We should
never trus the userspace from kernel space.
To keep the same API for the user space code you want a 1:1 mapping
between device physical address and process virtual address (ie if
device access device physical address A it is accessing the same
memory as what is backing the virtual address A in the process.
Security issues are on two things:
[I1]- fork/exec, a process who opened any such device and created an
active queue can transfer without its knowledge control of its
commands queue through COW. The parent map some anonymous region
to the device as a command queue buffer but because of COW the
parent can be the first to copy on write and thus the child can
inherit the original pages that are mapped to the hardware.
Here parent lose control and child gain it.
This is indeed an issue. But it remains an issue only if you continue to use the
queue and the memory after fork. We can use at_fork kinds of gadget to fix it in
user space.
Trusting user space is a no go from my point of view.
From some perspectives, I think the issue can be solved by iommu_notifier. For
example, when the process is fork-ed, we can set the mapped device mmio space as
COW for the child process, so a new queue can be created and set to the same
state as the parent's if the space is accessed. Then we can have two separated
queues for both the parent and the child. The memory part can be done in the
same way.
The mmap of mmio space for the queue is not an issue just use VM_DONTCOPY
for it. Issue is with COW and IOMMU mapping of pages and this can not be
solve in your model.
The thing is, the same strategy can be applied to VFIO without changing its
original feature.
No it can not it would break existing VFIO contract (which only should be
use against private anonymous vma).
quoted
[I2]- Because of [R3] you want to allow userspace to schedule commands
on the device without doing an ioctl and thus here user space
can schedule any commands to the device with any address. What
happens if that address have not been mapped by the user space
is undefined and in fact can not be defined as what each IOMMU
does on invalid address access is different from IOMMU to IOMMU.
In case of a bad IOMMU, or simply an IOMMU improperly setup by
the kernel, this can potentialy allow user space to DMA anywhere.
I don't think this is an issue. If you cannot trust IOMMU and proper setup of
IOMMU in kernel, you cannot trust anything. And the whole VFIO framework is
untrustable.
VFIO device is usualy restricted to trusted user and other device that
do DMA do various checks to make sure user space can not abuse them, the
assumption i have always seen so far is to not trust that IOMMU will do
all the work. So exposing user space access to device with DMA capabilities
should be done carefuly IMHO.
To be thorough list of potential bugs i am concern about:
- IOMMU hardware bug
- IOMMU does not isolate device after too many fail attempt
- firmware setup the IOMMU in some unsafe way and linux kernel does
not catch that (like passthrough on error or when there is no
entry for the PA which i am told is a thing for debug)
- bug in the linux IOMMU kernel driver
quoted
[I3]- By relying on GUP in VFIO you are not abiding by the implicit
contract (at least i hope it is implicit) that you should not
try to map to the device any file backed vma (private or share).
The VFIO code never check the vma controlling the addresses that
are provided to VFIO_IOMMU_MAP_DMA ioctl. Which means that the
user space can provide file backed range.
I am guessing that the VFIO code never had any issues because its
number one user is QEMU and QEMU never does that (and that's good
as no one should ever do that).
So if process does that you are opening your self to serious file
system corruption (depending on file system this can lead to total
data loss for the filesystem).
Issue is that once you GUP you never abide to file system flushing
which write protect the page before writing to the disk. So
because the page is still map with write permission to the device
(assuming VFIO_IOMMU_MAP_DMA was a write map) then the device can
write to the page while it is in the middle of being written back
to disk. Consult your nearest file system specialist to ask him
how bad that can be.
Same as I2, it is an issue, but the problem can be solved in VFIO if we really
take it in the scope of VFIO.
Except it can not be solve without breaking VFIO. If you use mmu_notifier
it means that the the IOMMU mapping can vanish at _any_ time and because
you allow user space to directly schedule work on the hardware command
queue than you do not have any synchronization point to use.
When notifier calls you must stop all hardware access and wait for any
pending work that might still dereference affected range. Finaly you can
restore mapping only once the notifier is done. AFAICT this would break
all existing VFIO user.
So solving that for your case it means you would have to:
warp_invalidate_range_callback() {
- take some lock to protect against restore
- unmap the command queue (zap_range) from userspace to stop further
commands
- wait for any pending commands on the hardware to complete
- clear all IOMMU mappings for the range
- put_pages()
- drop restore lock
return to let invalidation complete its works
}
warp_restore() {
// This is call by the page fault handler on the command queue mapped
// to user space.
- take restore lock
- go over all IOMMU mapping and restore them (GUP)
- remap command queue to userspace
- drop restore lock
}
This model does not work for VFIO existing users AFAICT.
quoted
[I4]- Design issue, mdev design As Far As I Understand It is about
sharing a single device to multiple clients (most obvious case
here is again QEMU guest). But you are going against that model,
in fact AFAIUI you are doing the exect opposite. When there is
no SVA/SVM you want only one mdev device that can not be share.
Wait. It is NOT "I want only one mdev device when there is no SVA/SVM", it is "I
can support only one mdev when there is no PASID support for the IOMMU".
Except you can support more than one user when no SVA/SVM with the model
i outlined below.
quoted
So this is counter intuitive to the mdev existing design. It is
not about sharing device among multiple users but about giving
exclusive access to the device to one user.
All the reasons above is why i believe a different model would serve
you and your user better. Below is a design that avoids all of the
above issues and still delivers all of your objectives with the
exceptions of the third one [R3] when there is no SVA/SVM.
Create a subsystem (very much boiler plate code) which allow device to
register themself against (very much like what you do in your current
patchset but outside of VFIO).
That subsystem will create a device file for each registered system and
expose a common API (ie set of ioctl) for each of those device files.
When user space create a queue (through an ioctl after opening the device
file) the kernel can return -EBUSY if all the device queue are in use,
or create a device queue and return a flag like SYNC_ONLY for device that
do not have SVA/SVM.
For device with SVA/SVM at the time the process create a queue you bind
the process PASID to the device queue. From there on the userspace can
schedule commands and use the device without going to kernel space.
As mentioned previously, this is not enough for fork scenario.
It is for every existing user of SVA/SVM so i fail to see why it would
be any different in your case. Note that they all use VM_DONTCOPY flag
on the queue mapped to userspace.
quoted
For device without SVA/SVM you create a fake queue that is just pure
memory is not related to the device. From there on the userspace must
call an ioctl every time it wants the device to consume its queue
(hence why the SYNC_ONLY flag for synchronous operation only). The
kernel portion read the fake queue expose to user space and copy
commands into the real hardware queue but first it properly map any
of the process memory needed for those commands to the device and
adjust the device physical address with the one it gets from dma_map
API.
But in this way, we will lost most of the benefit of avoiding syscall.
Yes but only when there is SVA/SVM. What i am trying to stress is that
there is no sane way to mirror user space address space onto device
without mmu_notifier so short of that this model where you have to
syscall to schedules thing on the hardware is the easiest thing to do.
quoted
With that model it is "easy" to listen to mmu_notifier and to abide by
them to avoid issues [I1], [I3] and [I4]. You obviously avoid the [I2]
issue by only mapping a fake device queue to userspace.
So yes with that models it means that every device that wish to support
the non SVA/SVM case will have to do extra work (ie emulate its command
queue in software in the kernel). But by doing so, you support an
unlimited number of process on your device (ie all the process can share
one single hardware command queues or multiple hardware queues).
If I can do this, I will not need WarpDrive at all:(
This is only needed if you wish to support non SVA/SVM, shifting the
burden to the kernel is always the thing to do especially if they are
legitimate security concerns.
Cheers,
Jérôme
From: Kenneth Lee <hidden> Date: 2018-09-18 06:02:23
On Mon, Sep 17, 2018 at 08:37:45AM -0400, Jerome Glisse wrote:
Date: Mon, 17 Sep 2018 08:37:45 -0400
From: Jerome Glisse <redacted>
To: Kenneth Lee <redacted>
CC: Kenneth Lee <redacted>, Herbert Xu
[off-list ref], kvm@vger.kernel.org, Jonathan Corbet
[off-list ref], Greg Kroah-Hartman [off-list ref], Joerg
Roedel [off-list ref], linux-doc@vger.kernel.org, Sanjay Kumar
[off-list ref], Hao Fang [off-list ref],
iommu@lists.linux-foundation.org, linux-kernel@vger.kernel.org,
linuxarm@huawei.com, Alex Williamson [off-list ref], Thomas
Gleixner [off-list ref], linux-crypto@vger.kernel.org, Zhou Wang
[off-list ref], Philippe Ombredanne [off-list ref],
Zaibo Xu [off-list ref], "David S . Miller" [off-list ref],
linux-accelerators@lists.ozlabs.org, Lu Baolu [off-list ref]
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
User-Agent: Mutt/1.10.1 (2018-07-13)
Message-ID: [off-list ref]
On Mon, Sep 17, 2018 at 04:39:40PM +0800, Kenneth Lee wrote:
quoted
On Sun, Sep 16, 2018 at 09:42:44PM -0400, Jerome Glisse wrote:
quoted
So i want to summarize issues i have as this threads have dig deep into
details. For this i would like to differentiate two cases first the easy
one when relying on SVA/SVM. Then the second one when there is no SVA/SVM.
Thank you very much for the summary.
quoted
In both cases your objectives as i understand them:
[R1]- expose a common user space API that make it easy to share boiler
plate code accross many devices (discovering devices, opening
device, creating context, creating command queue ...).
[R2]- try to share the device as much as possible up to device limits
(number of independant queues the device has)
[R3]- minimize syscall by allowing user space to directly schedule on the
device queue without a round trip to the kernel
I don't think i missed any.
(1) Device with SVA/SVM
For that case it is easy, you do not need to be in VFIO or part of any
thing specific in the kernel. There is no security risk (modulo bug in
the SVA/SVM silicon). Fork/exec is properly handle and binding a process
to a device is just couple dozen lines of code.
This is right...logically. But the kernel has no clear definition about "Device
with SVA/SVM" and no boiler plate for doing so. Then VFIO may become one of the
boiler plate.
VFIO is one of the wrappers for IOMMU for user space. And maybe it is the only
one. If we add that support within VFIO, which solve most of the problem of
SVA/SVM, it will save a lot of work in the future.
You do not need to "wrap" IOMMU for SVA/SVM. Existing upstream SVA/SVM user
all do the SVA/SVM setup in couple dozen lines and i failed to see how it
would require any more than that in your case.
quoted
I think this is the key confliction between us. So could Alex please say
something here? If the VFIO is going to take this into its scope, we can try
together to solve all the problem on the way. If it it is not, it is also
simple, we can just go to another way to fulfill this part of requirements even
we have to duplicate most of the code.
Another point I need to emphasis here: because we have to replace the hardware
queue when fork, so it won't be very simple even in SVA/SVM case.
I am assuming hardware queue can only be setup by the kernel and thus
you are totaly safe forkwise as the queue is setup against a PASID and
the child does not bind to any PASID and you use VM_DONTCOPY on the
mmap of the hardware MMIO queue because you should really use that flag
for that.
quoted
quoted
(2) Device does not have SVA/SVM (or it is disabled)
You want to still allow device to be part of your framework. However
here i see fundamentals securities issues and you move the burden of
being careful to user space which i think is a bad idea. We should
never trus the userspace from kernel space.
To keep the same API for the user space code you want a 1:1 mapping
between device physical address and process virtual address (ie if
device access device physical address A it is accessing the same
memory as what is backing the virtual address A in the process.
Security issues are on two things:
[I1]- fork/exec, a process who opened any such device and created an
active queue can transfer without its knowledge control of its
commands queue through COW. The parent map some anonymous region
to the device as a command queue buffer but because of COW the
parent can be the first to copy on write and thus the child can
inherit the original pages that are mapped to the hardware.
Here parent lose control and child gain it.
This is indeed an issue. But it remains an issue only if you continue to use the
queue and the memory after fork. We can use at_fork kinds of gadget to fix it in
user space.
Trusting user space is a no go from my point of view.
Can we dive deeper on this? Maybe we have different understanding on "Trusting
user space". As my understanding, "trusting user space" means "no matter what
the user process does, it should only hurt itself and anything give to it, no
the kernel and the other process".
In our case, we create a channel between a process and the hardware. The process
can do whateven it like to its own memory the channel itself. It won't hurt the
other process and the kernel. And if the process fork a child and give the
channel to the child, it should the freedom on those resource remain within the
parent and the child. We are not trust another else.
So do you refer to something else here?
quoted
From some perspectives, I think the issue can be solved by iommu_notifier. For
example, when the process is fork-ed, we can set the mapped device mmio space as
COW for the child process, so a new queue can be created and set to the same
state as the parent's if the space is accessed. Then we can have two separated
queues for both the parent and the child. The memory part can be done in the
same way.
The mmap of mmio space for the queue is not an issue just use VM_DONTCOPY
for it. Issue is with COW and IOMMU mapping of pages and this can not be
solve in your model.
quoted
The thing is, the same strategy can be applied to VFIO without changing its
original feature.
No it can not it would break existing VFIO contract (which only should be
use against private anonymous vma).
quoted
quoted
[I2]- Because of [R3] you want to allow userspace to schedule commands
on the device without doing an ioctl and thus here user space
can schedule any commands to the device with any address. What
happens if that address have not been mapped by the user space
is undefined and in fact can not be defined as what each IOMMU
does on invalid address access is different from IOMMU to IOMMU.
In case of a bad IOMMU, or simply an IOMMU improperly setup by
the kernel, this can potentialy allow user space to DMA anywhere.
I don't think this is an issue. If you cannot trust IOMMU and proper setup of
IOMMU in kernel, you cannot trust anything. And the whole VFIO framework is
untrustable.
VFIO device is usualy restricted to trusted user and other device that
do DMA do various checks to make sure user space can not abuse them, the
assumption i have always seen so far is to not trust that IOMMU will do
all the work. So exposing user space access to device with DMA capabilities
should be done carefuly IMHO.
To be thorough list of potential bugs i am concern about:
- IOMMU hardware bug
- IOMMU does not isolate device after too many fail attempt
- firmware setup the IOMMU in some unsafe way and linux kernel does
not catch that (like passthrough on error or when there is no
entry for the PA which i am told is a thing for debug)
- bug in the linux IOMMU kernel driver
quoted
quoted
[I3]- By relying on GUP in VFIO you are not abiding by the implicit
contract (at least i hope it is implicit) that you should not
try to map to the device any file backed vma (private or share).
The VFIO code never check the vma controlling the addresses that
are provided to VFIO_IOMMU_MAP_DMA ioctl. Which means that the
user space can provide file backed range.
I am guessing that the VFIO code never had any issues because its
number one user is QEMU and QEMU never does that (and that's good
as no one should ever do that).
So if process does that you are opening your self to serious file
system corruption (depending on file system this can lead to total
data loss for the filesystem).
Issue is that once you GUP you never abide to file system flushing
which write protect the page before writing to the disk. So
because the page is still map with write permission to the device
(assuming VFIO_IOMMU_MAP_DMA was a write map) then the device can
write to the page while it is in the middle of being written back
to disk. Consult your nearest file system specialist to ask him
how bad that can be.
Same as I2, it is an issue, but the problem can be solved in VFIO if we really
take it in the scope of VFIO.
Except it can not be solve without breaking VFIO. If you use mmu_notifier
it means that the the IOMMU mapping can vanish at _any_ time and because
you allow user space to directly schedule work on the hardware command
queue than you do not have any synchronization point to use.
When notifier calls you must stop all hardware access and wait for any
pending work that might still dereference affected range. Finaly you can
restore mapping only once the notifier is done. AFAICT this would break
all existing VFIO user.
So solving that for your case it means you would have to:
warp_invalidate_range_callback() {
- take some lock to protect against restore
- unmap the command queue (zap_range) from userspace to stop further
commands
- wait for any pending commands on the hardware to complete
- clear all IOMMU mappings for the range
- put_pages()
- drop restore lock
return to let invalidation complete its works
}
warp_restore() {
// This is call by the page fault handler on the command queue mapped
// to user space.
- take restore lock
- go over all IOMMU mapping and restore them (GUP)
- remap command queue to userspace
- drop restore lock
}
This model does not work for VFIO existing users AFAICT.
quoted
quoted
[I4]- Design issue, mdev design As Far As I Understand It is about
sharing a single device to multiple clients (most obvious case
here is again QEMU guest). But you are going against that model,
in fact AFAIUI you are doing the exect opposite. When there is
no SVA/SVM you want only one mdev device that can not be share.
Wait. It is NOT "I want only one mdev device when there is no SVA/SVM", it is "I
can support only one mdev when there is no PASID support for the IOMMU".
Except you can support more than one user when no SVA/SVM with the model
i outlined below.
quoted
quoted
So this is counter intuitive to the mdev existing design. It is
not about sharing device among multiple users but about giving
exclusive access to the device to one user.
All the reasons above is why i believe a different model would serve
you and your user better. Below is a design that avoids all of the
above issues and still delivers all of your objectives with the
exceptions of the third one [R3] when there is no SVA/SVM.
Create a subsystem (very much boiler plate code) which allow device to
register themself against (very much like what you do in your current
patchset but outside of VFIO).
That subsystem will create a device file for each registered system and
expose a common API (ie set of ioctl) for each of those device files.
When user space create a queue (through an ioctl after opening the device
file) the kernel can return -EBUSY if all the device queue are in use,
or create a device queue and return a flag like SYNC_ONLY for device that
do not have SVA/SVM.
For device with SVA/SVM at the time the process create a queue you bind
the process PASID to the device queue. From there on the userspace can
schedule commands and use the device without going to kernel space.
As mentioned previously, this is not enough for fork scenario.
It is for every existing user of SVA/SVM so i fail to see why it would
be any different in your case. Note that they all use VM_DONTCOPY flag
on the queue mapped to userspace.
quoted
quoted
For device without SVA/SVM you create a fake queue that is just pure
memory is not related to the device. From there on the userspace must
call an ioctl every time it wants the device to consume its queue
(hence why the SYNC_ONLY flag for synchronous operation only). The
kernel portion read the fake queue expose to user space and copy
commands into the real hardware queue but first it properly map any
of the process memory needed for those commands to the device and
adjust the device physical address with the one it gets from dma_map
API.
But in this way, we will lost most of the benefit of avoiding syscall.
Yes but only when there is SVA/SVM. What i am trying to stress is that
there is no sane way to mirror user space address space onto device
without mmu_notifier so short of that this model where you have to
syscall to schedules thing on the hardware is the easiest thing to do.
quoted
quoted
With that model it is "easy" to listen to mmu_notifier and to abide by
them to avoid issues [I1], [I3] and [I4]. You obviously avoid the [I2]
issue by only mapping a fake device queue to userspace.
So yes with that models it means that every device that wish to support
the non SVA/SVM case will have to do extra work (ie emulate its command
queue in software in the kernel). But by doing so, you support an
unlimited number of process on your device (ie all the process can share
one single hardware command queues or multiple hardware queues).
If I can do this, I will not need WarpDrive at all:(
This is only needed if you wish to support non SVA/SVM, shifting the
burden to the kernel is always the thing to do especially if they are
legitimate security concerns.
For the other part, please give me some more time to write some test code and
come back to the discussino. Thank you.
Cheers
On Tue, Sep 18, 2018 at 02:00:14PM +0800, Kenneth Lee wrote:
On Mon, Sep 17, 2018 at 08:37:45AM -0400, Jerome Glisse wrote:
quoted
On Mon, Sep 17, 2018 at 04:39:40PM +0800, Kenneth Lee wrote:
quoted
On Sun, Sep 16, 2018 at 09:42:44PM -0400, Jerome Glisse wrote:
quoted
So i want to summarize issues i have as this threads have dig deep into
details. For this i would like to differentiate two cases first the easy
one when relying on SVA/SVM. Then the second one when there is no SVA/SVM.
Thank you very much for the summary.
quoted
In both cases your objectives as i understand them:
[R1]- expose a common user space API that make it easy to share boiler
plate code accross many devices (discovering devices, opening
device, creating context, creating command queue ...).
[R2]- try to share the device as much as possible up to device limits
(number of independant queues the device has)
[R3]- minimize syscall by allowing user space to directly schedule on the
device queue without a round trip to the kernel
I don't think i missed any.
(1) Device with SVA/SVM
For that case it is easy, you do not need to be in VFIO or part of any
thing specific in the kernel. There is no security risk (modulo bug in
the SVA/SVM silicon). Fork/exec is properly handle and binding a process
to a device is just couple dozen lines of code.
This is right...logically. But the kernel has no clear definition about "Device
with SVA/SVM" and no boiler plate for doing so. Then VFIO may become one of the
boiler plate.
VFIO is one of the wrappers for IOMMU for user space. And maybe it is the only
one. If we add that support within VFIO, which solve most of the problem of
SVA/SVM, it will save a lot of work in the future.
You do not need to "wrap" IOMMU for SVA/SVM. Existing upstream SVA/SVM user
all do the SVA/SVM setup in couple dozen lines and i failed to see how it
would require any more than that in your case.
quoted
I think this is the key confliction between us. So could Alex please say
something here? If the VFIO is going to take this into its scope, we can try
together to solve all the problem on the way. If it it is not, it is also
simple, we can just go to another way to fulfill this part of requirements even
we have to duplicate most of the code.
Another point I need to emphasis here: because we have to replace the hardware
queue when fork, so it won't be very simple even in SVA/SVM case.
I am assuming hardware queue can only be setup by the kernel and thus
you are totaly safe forkwise as the queue is setup against a PASID and
the child does not bind to any PASID and you use VM_DONTCOPY on the
mmap of the hardware MMIO queue because you should really use that flag
for that.
quoted
quoted
(2) Device does not have SVA/SVM (or it is disabled)
You want to still allow device to be part of your framework. However
here i see fundamentals securities issues and you move the burden of
being careful to user space which i think is a bad idea. We should
never trus the userspace from kernel space.
To keep the same API for the user space code you want a 1:1 mapping
between device physical address and process virtual address (ie if
device access device physical address A it is accessing the same
memory as what is backing the virtual address A in the process.
Security issues are on two things:
[I1]- fork/exec, a process who opened any such device and created an
active queue can transfer without its knowledge control of its
commands queue through COW. The parent map some anonymous region
to the device as a command queue buffer but because of COW the
parent can be the first to copy on write and thus the child can
inherit the original pages that are mapped to the hardware.
Here parent lose control and child gain it.
This is indeed an issue. But it remains an issue only if you continue to use the
queue and the memory after fork. We can use at_fork kinds of gadget to fix it in
user space.
Trusting user space is a no go from my point of view.
Can we dive deeper on this? Maybe we have different understanding on "Trusting
user space". As my understanding, "trusting user space" means "no matter what
the user process does, it should only hurt itself and anything give to it, no
the kernel and the other process".
In our case, we create a channel between a process and the hardware. The process
can do whateven it like to its own memory the channel itself. It won't hurt the
other process and the kernel. And if the process fork a child and give the
channel to the child, it should the freedom on those resource remain within the
parent and the child. We are not trust another else.
So do you refer to something else here?
I am refering to COW giving control to the child on to what happens
in the parent from device point of view. A process hurting itself is
fine, but if process now has to do special steps to protect from
its child ie make sure that its childs can not hurt it, then i see
that as a kernel bug. We can not ask user space process to know about
all the thousands things that needs to be done to avoid issues with
each device driver that the process may use (process can be totaly
ignorant it is using a device if that device is use by a library it
links to).
Maybe what needs to happen will explain it better. So if userspace
wants to be secure and protect itself from its child taking over the
device through COW:
- parent opened a device and is using it
... when parent wants to fork/exec it must:
- parent _must_ flush device command queue and wait for the
device to finish all pending jobs
- parent _must_ unmap all range mapped to the device
- parent should first close device file (unless you force set
the CLOEXEC flag in the kernel)/it could also just flush
but if you are not mapping the device command queue with
VM_DONTCOPY then you should really be closing the device
- now parent can fork/exec
- parent must force COW ie write at least one byte to _all_
pages in the range it wants to use with the device
- parent re-open the device and re-initialize everything
So this is putting quite a burden on a number of steps the parent
_must_ do in order to keep control of memory exposed to the device.
Not doing so can potentialy lead (it depends on who does the COW
first) to the child taking control of memory use by the device,
memory which was mapped by the parent before the child was created.
Forcing CLOEXEC and VM_DONTCOPY somewhat help to simplify this,
but you still need to stop, flush, unmap, before fork/exec and then
re-init everything after.
This is only when not using SVA/SVM, SVA/SVM is totaly fine from
that point of view, no issues whatsoever.
The solution i outlined in previous email do not have that above
issue either, no need to rely on user space doing that dance.
Cheers,
Jérôme
From: Kenneth Lee <hidden> Date: 2018-09-20 05:57:57
On Tue, Sep 18, 2018 at 09:03:14AM -0400, Jerome Glisse wrote:
Date: Tue, 18 Sep 2018 09:03:14 -0400
From: Jerome Glisse <redacted>
To: Kenneth Lee <redacted>
CC: Kenneth Lee <redacted>, Alex Williamson
[off-list ref], Herbert Xu [off-list ref],
kvm@vger.kernel.org, Jonathan Corbet [off-list ref], Greg Kroah-Hartman
[off-list ref], Joerg Roedel [off-list ref],
linux-doc@vger.kernel.org, Sanjay Kumar [off-list ref], Hao
Fang [off-list ref], linux-kernel@vger.kernel.org,
linuxarm@huawei.com, iommu@lists.linux-foundation.org, "David S . Miller"
[off-list ref], linux-crypto@vger.kernel.org, Zhou Wang
[off-list ref], Philippe Ombredanne [off-list ref],
Thomas Gleixner [off-list ref], Zaibo Xu [off-list ref],
linux-accelerators@lists.ozlabs.org, Lu Baolu [off-list ref]
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
User-Agent: Mutt/1.10.1 (2018-07-13)
Message-ID: [off-list ref]
On Tue, Sep 18, 2018 at 02:00:14PM +0800, Kenneth Lee wrote:
quoted
On Mon, Sep 17, 2018 at 08:37:45AM -0400, Jerome Glisse wrote:
quoted
On Mon, Sep 17, 2018 at 04:39:40PM +0800, Kenneth Lee wrote:
quoted
On Sun, Sep 16, 2018 at 09:42:44PM -0400, Jerome Glisse wrote:
quoted
So i want to summarize issues i have as this threads have dig deep into
details. For this i would like to differentiate two cases first the easy
one when relying on SVA/SVM. Then the second one when there is no SVA/SVM.
Thank you very much for the summary.
quoted
In both cases your objectives as i understand them:
[R1]- expose a common user space API that make it easy to share boiler
plate code accross many devices (discovering devices, opening
device, creating context, creating command queue ...).
[R2]- try to share the device as much as possible up to device limits
(number of independant queues the device has)
[R3]- minimize syscall by allowing user space to directly schedule on the
device queue without a round trip to the kernel
I don't think i missed any.
(1) Device with SVA/SVM
For that case it is easy, you do not need to be in VFIO or part of any
thing specific in the kernel. There is no security risk (modulo bug in
the SVA/SVM silicon). Fork/exec is properly handle and binding a process
to a device is just couple dozen lines of code.
This is right...logically. But the kernel has no clear definition about "Device
with SVA/SVM" and no boiler plate for doing so. Then VFIO may become one of the
boiler plate.
VFIO is one of the wrappers for IOMMU for user space. And maybe it is the only
one. If we add that support within VFIO, which solve most of the problem of
SVA/SVM, it will save a lot of work in the future.
You do not need to "wrap" IOMMU for SVA/SVM. Existing upstream SVA/SVM user
all do the SVA/SVM setup in couple dozen lines and i failed to see how it
would require any more than that in your case.
quoted
I think this is the key confliction between us. So could Alex please say
something here? If the VFIO is going to take this into its scope, we can try
together to solve all the problem on the way. If it it is not, it is also
simple, we can just go to another way to fulfill this part of requirements even
we have to duplicate most of the code.
Another point I need to emphasis here: because we have to replace the hardware
queue when fork, so it won't be very simple even in SVA/SVM case.
I am assuming hardware queue can only be setup by the kernel and thus
you are totaly safe forkwise as the queue is setup against a PASID and
the child does not bind to any PASID and you use VM_DONTCOPY on the
mmap of the hardware MMIO queue because you should really use that flag
for that.
quoted
quoted
(2) Device does not have SVA/SVM (or it is disabled)
You want to still allow device to be part of your framework. However
here i see fundamentals securities issues and you move the burden of
being careful to user space which i think is a bad idea. We should
never trus the userspace from kernel space.
To keep the same API for the user space code you want a 1:1 mapping
between device physical address and process virtual address (ie if
device access device physical address A it is accessing the same
memory as what is backing the virtual address A in the process.
Security issues are on two things:
[I1]- fork/exec, a process who opened any such device and created an
active queue can transfer without its knowledge control of its
commands queue through COW. The parent map some anonymous region
to the device as a command queue buffer but because of COW the
parent can be the first to copy on write and thus the child can
inherit the original pages that are mapped to the hardware.
Here parent lose control and child gain it.
This is indeed an issue. But it remains an issue only if you continue to use the
queue and the memory after fork. We can use at_fork kinds of gadget to fix it in
user space.
Trusting user space is a no go from my point of view.
Can we dive deeper on this? Maybe we have different understanding on "Trusting
user space". As my understanding, "trusting user space" means "no matter what
the user process does, it should only hurt itself and anything give to it, no
the kernel and the other process".
In our case, we create a channel between a process and the hardware. The process
can do whateven it like to its own memory the channel itself. It won't hurt the
other process and the kernel. And if the process fork a child and give the
channel to the child, it should the freedom on those resource remain within the
parent and the child. We are not trust another else.
So do you refer to something else here?
I am refering to COW giving control to the child on to what happens
in the parent from device point of view. A process hurting itself is
fine, but if process now has to do special steps to protect from
its child ie make sure that its childs can not hurt it, then i see
that as a kernel bug. We can not ask user space process to know about
all the thousands things that needs to be done to avoid issues with
each device driver that the process may use (process can be totaly
ignorant it is using a device if that device is use by a library it
links to).
Maybe what needs to happen will explain it better. So if userspace
wants to be secure and protect itself from its child taking over the
device through COW:
- parent opened a device and is using it
... when parent wants to fork/exec it must:
- parent _must_ flush device command queue and wait for the
device to finish all pending jobs
- parent _must_ unmap all range mapped to the device
- parent should first close device file (unless you force set
the CLOEXEC flag in the kernel)/it could also just flush
but if you are not mapping the device command queue with
VM_DONTCOPY then you should really be closing the device
- now parent can fork/exec
- parent must force COW ie write at least one byte to _all_
pages in the range it wants to use with the device
- parent re-open the device and re-initialize everything
So this is putting quite a burden on a number of steps the parent
_must_ do in order to keep control of memory exposed to the device.
Not doing so can potentialy lead (it depends on who does the COW
first) to the child taking control of memory use by the device,
memory which was mapped by the parent before the child was created.
Forcing CLOEXEC and VM_DONTCOPY somewhat help to simplify this,
but you still need to stop, flush, unmap, before fork/exec and then
re-init everything after.
This is only when not using SVA/SVM, SVA/SVM is totaly fine from
that point of view, no issues whatsoever.
The solution i outlined in previous email do not have that above
issue either, no need to rely on user space doing that dance.
Thank you. I get the point. I'm now trying to see if I can solve the problem by
seting the vma to VM_SHARED when the portiong is "shared to the hardware".
Cheers,
Jérôme
--
-Kenneth(Hisilicon)
================================================================================
本邮件及其附件含有华为公司的保密信息,仅限于发送给上面地址中列出的个人或群组。禁
止任何其他人以任何形式使用(包括但不限于全部或部分地泄露、复制、或散发)本邮件中
的信息。如果您错收了本邮件,请您立即电话或邮件通知发件人并删除本邮件!
This e-mail and its attachments contain confidential information from HUAWEI,
which is intended only for the person or entity whose address is listed above.
Any use of the
information contained herein in any way (including, but not limited to, total or
partial disclosure, reproduction, or dissemination) by persons other than the
intended
recipient(s) is prohibited. If you receive this e-mail in error, please notify
the sender by phone or email immediately and delete it!
On Thu, Sep 20, 2018 at 01:55:43PM +0800, Kenneth Lee wrote:
On Tue, Sep 18, 2018 at 09:03:14AM -0400, Jerome Glisse wrote:
quoted
On Tue, Sep 18, 2018 at 02:00:14PM +0800, Kenneth Lee wrote:
quoted
On Mon, Sep 17, 2018 at 08:37:45AM -0400, Jerome Glisse wrote:
quoted
On Mon, Sep 17, 2018 at 04:39:40PM +0800, Kenneth Lee wrote:
quoted
On Sun, Sep 16, 2018 at 09:42:44PM -0400, Jerome Glisse wrote:
quoted
So i want to summarize issues i have as this threads have dig deep into
details. For this i would like to differentiate two cases first the easy
one when relying on SVA/SVM. Then the second one when there is no SVA/SVM.
Thank you very much for the summary.
quoted
In both cases your objectives as i understand them:
[R1]- expose a common user space API that make it easy to share boiler
plate code accross many devices (discovering devices, opening
device, creating context, creating command queue ...).
[R2]- try to share the device as much as possible up to device limits
(number of independant queues the device has)
[R3]- minimize syscall by allowing user space to directly schedule on the
device queue without a round trip to the kernel
I don't think i missed any.
(1) Device with SVA/SVM
For that case it is easy, you do not need to be in VFIO or part of any
thing specific in the kernel. There is no security risk (modulo bug in
the SVA/SVM silicon). Fork/exec is properly handle and binding a process
to a device is just couple dozen lines of code.
This is right...logically. But the kernel has no clear definition about "Device
with SVA/SVM" and no boiler plate for doing so. Then VFIO may become one of the
boiler plate.
VFIO is one of the wrappers for IOMMU for user space. And maybe it is the only
one. If we add that support within VFIO, which solve most of the problem of
SVA/SVM, it will save a lot of work in the future.
You do not need to "wrap" IOMMU for SVA/SVM. Existing upstream SVA/SVM user
all do the SVA/SVM setup in couple dozen lines and i failed to see how it
would require any more than that in your case.
quoted
I think this is the key confliction between us. So could Alex please say
something here? If the VFIO is going to take this into its scope, we can try
together to solve all the problem on the way. If it it is not, it is also
simple, we can just go to another way to fulfill this part of requirements even
we have to duplicate most of the code.
Another point I need to emphasis here: because we have to replace the hardware
queue when fork, so it won't be very simple even in SVA/SVM case.
I am assuming hardware queue can only be setup by the kernel and thus
you are totaly safe forkwise as the queue is setup against a PASID and
the child does not bind to any PASID and you use VM_DONTCOPY on the
mmap of the hardware MMIO queue because you should really use that flag
for that.
quoted
quoted
(2) Device does not have SVA/SVM (or it is disabled)
You want to still allow device to be part of your framework. However
here i see fundamentals securities issues and you move the burden of
being careful to user space which i think is a bad idea. We should
never trus the userspace from kernel space.
To keep the same API for the user space code you want a 1:1 mapping
between device physical address and process virtual address (ie if
device access device physical address A it is accessing the same
memory as what is backing the virtual address A in the process.
Security issues are on two things:
[I1]- fork/exec, a process who opened any such device and created an
active queue can transfer without its knowledge control of its
commands queue through COW. The parent map some anonymous region
to the device as a command queue buffer but because of COW the
parent can be the first to copy on write and thus the child can
inherit the original pages that are mapped to the hardware.
Here parent lose control and child gain it.
This is indeed an issue. But it remains an issue only if you continue to use the
queue and the memory after fork. We can use at_fork kinds of gadget to fix it in
user space.
Trusting user space is a no go from my point of view.
Can we dive deeper on this? Maybe we have different understanding on "Trusting
user space". As my understanding, "trusting user space" means "no matter what
the user process does, it should only hurt itself and anything give to it, no
the kernel and the other process".
In our case, we create a channel between a process and the hardware. The process
can do whateven it like to its own memory the channel itself. It won't hurt the
other process and the kernel. And if the process fork a child and give the
channel to the child, it should the freedom on those resource remain within the
parent and the child. We are not trust another else.
So do you refer to something else here?
I am refering to COW giving control to the child on to what happens
in the parent from device point of view. A process hurting itself is
fine, but if process now has to do special steps to protect from
its child ie make sure that its childs can not hurt it, then i see
that as a kernel bug. We can not ask user space process to know about
all the thousands things that needs to be done to avoid issues with
each device driver that the process may use (process can be totaly
ignorant it is using a device if that device is use by a library it
links to).
Maybe what needs to happen will explain it better. So if userspace
wants to be secure and protect itself from its child taking over the
device through COW:
- parent opened a device and is using it
... when parent wants to fork/exec it must:
- parent _must_ flush device command queue and wait for the
device to finish all pending jobs
- parent _must_ unmap all range mapped to the device
- parent should first close device file (unless you force set
the CLOEXEC flag in the kernel)/it could also just flush
but if you are not mapping the device command queue with
VM_DONTCOPY then you should really be closing the device
- now parent can fork/exec
- parent must force COW ie write at least one byte to _all_
pages in the range it wants to use with the device
- parent re-open the device and re-initialize everything
So this is putting quite a burden on a number of steps the parent
_must_ do in order to keep control of memory exposed to the device.
Not doing so can potentialy lead (it depends on who does the COW
first) to the child taking control of memory use by the device,
memory which was mapped by the parent before the child was created.
Forcing CLOEXEC and VM_DONTCOPY somewhat help to simplify this,
but you still need to stop, flush, unmap, before fork/exec and then
re-init everything after.
This is only when not using SVA/SVM, SVA/SVM is totaly fine from
that point of view, no issues whatsoever.
The solution i outlined in previous email do not have that above
issue either, no need to rely on user space doing that dance.
Thank you. I get the point. I'm now trying to see if I can solve the problem by
seting the vma to VM_SHARED when the portiong is "shared to the hardware".
FYI you can not convert a private anonymous vma to a share one it is
illegal AFAIK at least i never heard of it and i am pretty sure the
mm code would break if that happens. The user space is the one that
decide what flags a vma has, not the kernel. Modulo few flags like
DONTCOPY that can be force set by device driver for their vma ie vma
of an mmap against the device file.
If you don't like my solution here is another one but it is ugly and
i think it is a bad idea. Again this is for the non SVA/SVM case and
it assumes that the command queue is a mmap() of the device file:
(A) register mmu_notifier
(B) on _every_ invalidate range callback (_no matter_ what is the
range) you zap the command queue mapped to user space (this is
because you can't tell if the callback happens for a fork or
something else) wait for the hardware queue to finish and clear
all the iommu/dma mapping and you unpin all the pages ie
put_page()
(C) in device file vma page fault handler (vm_operations_struct.
fault) you redo all the GUP and redo all the iommu/dma mapping
and you remap the command queue to the userspace
In (C) you can remap different command queue if you are in the child
than in the parent (just look at current->mm and compare it to the
one the command queue was created against).
Note that this solution will be much __slower__ than what i described
in my previous email. You will see that mmu notifier callbacks happens
often and for tons of reasons and you will be _constantly_ undoing and
redoing tons of work.
This can be mitigated if you can differentiate reasons behind a mmu
notifier callback. I posted patchset to do that a while ago and i
intend to post it again in the next month or so. But this would still
be a bad idea and solution i described previously is much more sane.
Trying to pretend you can have the same thing as SVA/SVM without SVA
is not a good idea. The non SVA case can still expose same API (like
i described previously) but should go through kernel for _every_
hardware submission (you can batch multiple commands in one submission).
Not doing so is way too risky from my POV.
Cheers,
Jérôme
From: Kenneth Lee <hidden> Date: 2018-09-21 10:05:25
On Sun, Sep 16, 2018 at 09:42:44PM -0400, Jerome Glisse wrote:
Received: from POPSCN.huawei.com [10.3.17.45] by Turing-Arch-b with POP3
(fetchmail-6.3.26) for <kenny@localhost> (single-drop); Mon, 17 Sep 2018
09:45:02 +0800 (CST)
Received: from DGGEMM406-HUB.china.huawei.com (10.3.20.214) by
dggeml421-hub.china.huawei.com (10.1.199.38) with Microsoft SMTP Server
(TLS) id 14.3.399.0; Mon, 17 Sep 2018 09:43:07 +0800
Received: from dggwg01-in.huawei.com (172.30.65.32) by
DGGEMM406-HUB.china.huawei.com (10.3.20.214) with Microsoft SMTP Server id
14.3.399.0; Mon, 17 Sep 2018 09:43:00 +0800
Received: from mx1.redhat.com (unknown [209.132.183.28]) by Forcepoint
Email with ESMTPS id A15E04AB7D1C3; Mon, 17 Sep 2018 09:42:56 +0800 (CST)
Received: from smtp.corp.redhat.com
(int-mx11.intmail.prod.int.phx2.redhat.com [10.5.11.26]) (using
TLSv1.2 with cipher AECDH-AES256-SHA (256/256 bits)) (No client
certificate requested) by mx1.redhat.com (Postfix) with ESMTPS id
EC621308212D; Mon, 17 Sep 2018 01:42:52 +0000 (UTC)
Received: from redhat.com (ovpn-121-3.rdu2.redhat.com [10.10.121.3]) by
smtp.corp.redhat.com (Postfix) with ESMTPS id 8874530912F4; Mon, 17 Sep
2018 01:42:46 +0000 (UTC)
Date: Sun, 16 Sep 2018 21:42:44 -0400
From: Jerome Glisse <redacted>
To: Kenneth Lee <redacted>
CC: Jonathan Corbet <corbet@lwn.net>, Herbert Xu
[off-list ref], "David S . Miller" [off-list ref],
Joerg Roedel [off-list ref], Alex Williamson
[off-list ref], Kenneth Lee [off-list ref], Hao
Fang [off-list ref], Zhou Wang [off-list ref], Zaibo Xu
[off-list ref], Philippe Ombredanne [off-list ref], Greg
Kroah-Hartman [off-list ref], Thomas Gleixner
[off-list ref], linux-doc@vger.kernel.org,
linux-kernel@vger.kernel.org, linux-crypto@vger.kernel.org,
iommu@lists.linux-foundation.org, kvm@vger.kernel.org,
linux-accelerators@lists.ozlabs.org, Lu Baolu [off-list ref],
Sanjay Kumar [off-list ref], linuxarm@huawei.com
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
Message-ID: [off-list ref]
References: [off-list ref]
Content-Type: text/plain; charset="iso-8859-1"
Content-Disposition: inline
Content-Transfer-Encoding: 8bit
In-Reply-To: [off-list ref]
User-Agent: Mutt/1.10.1 (2018-07-13)
X-Scanned-By: MIMEDefang 2.84 on 10.5.11.26
X-Greylist: Sender IP whitelisted, not delayed by milter-greylist-4.5.16
(mx1.redhat.com [10.5.110.42]); Mon, 17 Sep 2018 01:42:53 +0000 (UTC)
Return-Path: jglisse@redhat.com
X-MS-Exchange-Organization-AuthSource: DGGEMM406-HUB.china.huawei.com
X-MS-Exchange-Organization-AuthAs: Anonymous
MIME-Version: 1.0
So i want to summarize issues i have as this threads have dig deep into
details. For this i would like to differentiate two cases first the easy
one when relying on SVA/SVM. Then the second one when there is no SVA/SVM.
In both cases your objectives as i understand them:
[R1]- expose a common user space API that make it easy to share boiler
plate code accross many devices (discovering devices, opening
device, creating context, creating command queue ...).
[R2]- try to share the device as much as possible up to device limits
(number of independant queues the device has)
[R3]- minimize syscall by allowing user space to directly schedule on the
device queue without a round trip to the kernel
I don't think i missed any.
(1) Device with SVA/SVM
For that case it is easy, you do not need to be in VFIO or part of any
thing specific in the kernel. There is no security risk (modulo bug in
the SVA/SVM silicon). Fork/exec is properly handle and binding a process
to a device is just couple dozen lines of code.
(2) Device does not have SVA/SVM (or it is disabled)
You want to still allow device to be part of your framework. However
here i see fundamentals securities issues and you move the burden of
being careful to user space which i think is a bad idea. We should
never trus the userspace from kernel space.
To keep the same API for the user space code you want a 1:1 mapping
between device physical address and process virtual address (ie if
device access device physical address A it is accessing the same
memory as what is backing the virtual address A in the process.
Security issues are on two things:
[I1]- fork/exec, a process who opened any such device and created an
active queue can transfer without its knowledge control of its
commands queue through COW. The parent map some anonymous region
to the device as a command queue buffer but because of COW the
parent can be the first to copy on write and thus the child can
inherit the original pages that are mapped to the hardware.
Here parent lose control and child gain it.
Hi, Jerome,
I reconsider your logic. I think the problem can be solved. Let us separate the
SVA/SVM feature into two: fault-from-device and device-va-awareness. A device
with iommu can support only device-va-awareness or both.
VFIO works on top of iommu, so it will support at least device-va-awareness. For
the COW problem, it can be taken as a mmu synchronization issue. If the mmu page
table is changed, it should be synchronize to iommu (via iommu_notifier). In the
case that the device support fault-from-device, it will work fine. In the case
that it supports only device-va-awareness, we can prefault (handle_mm_fault)
also via iommu_notifier and reset to iommu page table.
So this can be considered as a bug of VFIO, cannot it?
[I2]- Because of [R3] you want to allow userspace to schedule commands
on the device without doing an ioctl and thus here user space
can schedule any commands to the device with any address. What
happens if that address have not been mapped by the user space
is undefined and in fact can not be defined as what each IOMMU
does on invalid address access is different from IOMMU to IOMMU.
In case of a bad IOMMU, or simply an IOMMU improperly setup by
the kernel, this can potentialy allow user space to DMA anywhere.
[I3]- By relying on GUP in VFIO you are not abiding by the implicit
contract (at least i hope it is implicit) that you should not
try to map to the device any file backed vma (private or share).
The VFIO code never check the vma controlling the addresses that
are provided to VFIO_IOMMU_MAP_DMA ioctl. Which means that the
user space can provide file backed range.
I am guessing that the VFIO code never had any issues because its
number one user is QEMU and QEMU never does that (and that's good
as no one should ever do that).
So if process does that you are opening your self to serious file
system corruption (depending on file system this can lead to total
data loss for the filesystem).
Issue is that once you GUP you never abide to file system flushing
which write protect the page before writing to the disk. So
because the page is still map with write permission to the device
(assuming VFIO_IOMMU_MAP_DMA was a write map) then the device can
write to the page while it is in the middle of being written back
to disk. Consult your nearest file system specialist to ask him
how bad that can be.
In the case, we cannot do anything if the device do not support
fault-from-device. But we can reject write map with file-backed mapping.
It seems both issues can be solved under VFIO framework:) (But of cause, I don't
mean it has to)
[I4]- Design issue, mdev design As Far As I Understand It is about
sharing a single device to multiple clients (most obvious case
here is again QEMU guest). But you are going against that model,
in fact AFAIUI you are doing the exect opposite. When there is
no SVA/SVM you want only one mdev device that can not be share.
So this is counter intuitive to the mdev existing design. It is
not about sharing device among multiple users but about giving
exclusive access to the device to one user.
All the reasons above is why i believe a different model would serve
you and your user better. Below is a design that avoids all of the
above issues and still delivers all of your objectives with the
exceptions of the third one [R3] when there is no SVA/SVM.
Create a subsystem (very much boiler plate code) which allow device to
register themself against (very much like what you do in your current
patchset but outside of VFIO).
That subsystem will create a device file for each registered system and
expose a common API (ie set of ioctl) for each of those device files.
When user space create a queue (through an ioctl after opening the device
file) the kernel can return -EBUSY if all the device queue are in use,
or create a device queue and return a flag like SYNC_ONLY for device that
do not have SVA/SVM.
For device with SVA/SVM at the time the process create a queue you bind
the process PASID to the device queue. From there on the userspace can
schedule commands and use the device without going to kernel space.
For device without SVA/SVM you create a fake queue that is just pure
memory is not related to the device. From there on the userspace must
call an ioctl every time it wants the device to consume its queue
(hence why the SYNC_ONLY flag for synchronous operation only). The
kernel portion read the fake queue expose to user space and copy
commands into the real hardware queue but first it properly map any
of the process memory needed for those commands to the device and
adjust the device physical address with the one it gets from dma_map
API.
With that model it is "easy" to listen to mmu_notifier and to abide by
them to avoid issues [I1], [I3] and [I4]. You obviously avoid the [I2]
issue by only mapping a fake device queue to userspace.
So yes with that models it means that every device that wish to support
the non SVA/SVM case will have to do extra work (ie emulate its command
queue in software in the kernel). But by doing so, you support an
unlimited number of process on your device (ie all the process can share
one single hardware command queues or multiple hardware queues).
The big advantages i see here is that the process do not have to worry
about doing something wrong. You are protecting yourself and your user
from stupid mistakes.
I hope this is useful to you.
Cheers,
Jérôme
From: Kenneth Lee <hidden> Date: 2018-09-21 10:07:39
On Thu, Sep 20, 2018 at 10:23:40AM -0400, Jerome Glisse wrote:
Received: from popscn.huawei.com [10.3.17.45] by Turing-Arch-b with POP3
(fetchmail-6.3.26) for <kenny@localhost> (single-drop); Thu, 20 Sep 2018
22:30:01 +0800 (CST)
Received: from DGGEMM401-HUB.china.huawei.com (10.3.20.209) by
dggeml405-hub.china.huawei.com (10.3.17.49) with Microsoft SMTP Server
(TLS) id 14.3.382.0; Thu, 20 Sep 2018 22:23:57 +0800
Received: from dggwg01-in.huawei.com (172.30.65.35) by
DGGEMM401-HUB.china.huawei.com (10.3.20.209) with Microsoft SMTP Server id
14.3.399.0; Thu, 20 Sep 2018 22:23:56 +0800
Received: from mx1.redhat.com (unknown [209.132.183.28]) by Forcepoint
Email with ESMTPS id 18963FA9B6FA9; Thu, 20 Sep 2018 22:23:50 +0800 (CST)
Received: from smtp.corp.redhat.com
(int-mx07.intmail.prod.int.phx2.redhat.com [10.5.11.22]) (using
TLSv1.2 with cipher AECDH-AES256-SHA (256/256 bits)) (No client
certificate requested) by mx1.redhat.com (Postfix) with ESMTPS id
CF1D9C058CBE; Thu, 20 Sep 2018 14:23:47 +0000 (UTC)
Received: from redhat.com (unknown [10.20.6.215]) by
smtp.corp.redhat.com (Postfix) with ESMTPS id 1B030106A780; Thu, 20 Sep
2018 14:23:41 +0000 (UTC)
Date: Thu, 20 Sep 2018 10:23:40 -0400
From: Jerome Glisse <redacted>
To: Kenneth Lee <redacted>
CC: Kenneth Lee <redacted>, Alex Williamson
[off-list ref], Herbert Xu [off-list ref],
kvm@vger.kernel.org, Jonathan Corbet [off-list ref], Greg Kroah-Hartman
[off-list ref], Joerg Roedel [off-list ref],
linux-doc@vger.kernel.org, Sanjay Kumar [off-list ref], Hao
Fang [off-list ref], linux-kernel@vger.kernel.org,
linuxarm@huawei.com, iommu@lists.linux-foundation.org, "David S . Miller"
[off-list ref], linux-crypto@vger.kernel.org, Zhou Wang
[off-list ref], Philippe Ombredanne [off-list ref],
Thomas Gleixner [off-list ref], Zaibo Xu [off-list ref],
linux-accelerators@lists.ozlabs.org, Lu Baolu [off-list ref]
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
Message-ID: [off-list ref]
References: [off-list ref]
[off-list ref]
<20180917083940.GE207969@Turing-Arch-b> [off-list ref]
<20180918060014.GF207969@Turing-Arch-b> [off-list ref]
<20180920055543.GG207969@Turing-Arch-b>
Content-Type: text/plain; charset="iso-8859-1"
Content-Disposition: inline
Content-Transfer-Encoding: 8bit
In-Reply-To: <20180920055543.GG207969@Turing-Arch-b>
User-Agent: Mutt/1.10.0 (2018-05-17)
X-Scanned-By: MIMEDefang 2.84 on 10.5.11.22
X-Greylist: Sender IP whitelisted, not delayed by milter-greylist-4.5.16
(mx1.redhat.com [10.5.110.32]); Thu, 20 Sep 2018 14:23:48 +0000 (UTC)
Return-Path: jglisse@redhat.com
X-MS-Exchange-Organization-AuthSource: DGGEMM401-HUB.china.huawei.com
X-MS-Exchange-Organization-AuthAs: Anonymous
MIME-Version: 1.0
On Thu, Sep 20, 2018 at 01:55:43PM +0800, Kenneth Lee wrote:
quoted
On Tue, Sep 18, 2018 at 09:03:14AM -0400, Jerome Glisse wrote:
quoted
On Tue, Sep 18, 2018 at 02:00:14PM +0800, Kenneth Lee wrote:
quoted
On Mon, Sep 17, 2018 at 08:37:45AM -0400, Jerome Glisse wrote:
quoted
On Mon, Sep 17, 2018 at 04:39:40PM +0800, Kenneth Lee wrote:
quoted
On Sun, Sep 16, 2018 at 09:42:44PM -0400, Jerome Glisse wrote:
quoted
So i want to summarize issues i have as this threads have dig deep into
details. For this i would like to differentiate two cases first the easy
one when relying on SVA/SVM. Then the second one when there is no SVA/SVM.
Thank you very much for the summary.
quoted
In both cases your objectives as i understand them:
[R1]- expose a common user space API that make it easy to share boiler
plate code accross many devices (discovering devices, opening
device, creating context, creating command queue ...).
[R2]- try to share the device as much as possible up to device limits
(number of independant queues the device has)
[R3]- minimize syscall by allowing user space to directly schedule on the
device queue without a round trip to the kernel
I don't think i missed any.
(1) Device with SVA/SVM
For that case it is easy, you do not need to be in VFIO or part of any
thing specific in the kernel. There is no security risk (modulo bug in
the SVA/SVM silicon). Fork/exec is properly handle and binding a process
to a device is just couple dozen lines of code.
This is right...logically. But the kernel has no clear definition about "Device
with SVA/SVM" and no boiler plate for doing so. Then VFIO may become one of the
boiler plate.
VFIO is one of the wrappers for IOMMU for user space. And maybe it is the only
one. If we add that support within VFIO, which solve most of the problem of
SVA/SVM, it will save a lot of work in the future.
You do not need to "wrap" IOMMU for SVA/SVM. Existing upstream SVA/SVM user
all do the SVA/SVM setup in couple dozen lines and i failed to see how it
would require any more than that in your case.
quoted
I think this is the key confliction between us. So could Alex please say
something here? If the VFIO is going to take this into its scope, we can try
together to solve all the problem on the way. If it it is not, it is also
simple, we can just go to another way to fulfill this part of requirements even
we have to duplicate most of the code.
Another point I need to emphasis here: because we have to replace the hardware
queue when fork, so it won't be very simple even in SVA/SVM case.
I am assuming hardware queue can only be setup by the kernel and thus
you are totaly safe forkwise as the queue is setup against a PASID and
the child does not bind to any PASID and you use VM_DONTCOPY on the
mmap of the hardware MMIO queue because you should really use that flag
for that.
quoted
quoted
(2) Device does not have SVA/SVM (or it is disabled)
You want to still allow device to be part of your framework. However
here i see fundamentals securities issues and you move the burden of
being careful to user space which i think is a bad idea. We should
never trus the userspace from kernel space.
To keep the same API for the user space code you want a 1:1 mapping
between device physical address and process virtual address (ie if
device access device physical address A it is accessing the same
memory as what is backing the virtual address A in the process.
Security issues are on two things:
[I1]- fork/exec, a process who opened any such device and created an
active queue can transfer without its knowledge control of its
commands queue through COW. The parent map some anonymous region
to the device as a command queue buffer but because of COW the
parent can be the first to copy on write and thus the child can
inherit the original pages that are mapped to the hardware.
Here parent lose control and child gain it.
This is indeed an issue. But it remains an issue only if you continue to use the
queue and the memory after fork. We can use at_fork kinds of gadget to fix it in
user space.
Trusting user space is a no go from my point of view.
Can we dive deeper on this? Maybe we have different understanding on "Trusting
user space". As my understanding, "trusting user space" means "no matter what
the user process does, it should only hurt itself and anything give to it, no
the kernel and the other process".
In our case, we create a channel between a process and the hardware. The process
can do whateven it like to its own memory the channel itself. It won't hurt the
other process and the kernel. And if the process fork a child and give the
channel to the child, it should the freedom on those resource remain within the
parent and the child. We are not trust another else.
So do you refer to something else here?
I am refering to COW giving control to the child on to what happens
in the parent from device point of view. A process hurting itself is
fine, but if process now has to do special steps to protect from
its child ie make sure that its childs can not hurt it, then i see
that as a kernel bug. We can not ask user space process to know about
all the thousands things that needs to be done to avoid issues with
each device driver that the process may use (process can be totaly
ignorant it is using a device if that device is use by a library it
links to).
Maybe what needs to happen will explain it better. So if userspace
wants to be secure and protect itself from its child taking over the
device through COW:
- parent opened a device and is using it
... when parent wants to fork/exec it must:
- parent _must_ flush device command queue and wait for the
device to finish all pending jobs
- parent _must_ unmap all range mapped to the device
- parent should first close device file (unless you force set
the CLOEXEC flag in the kernel)/it could also just flush
but if you are not mapping the device command queue with
VM_DONTCOPY then you should really be closing the device
- now parent can fork/exec
- parent must force COW ie write at least one byte to _all_
pages in the range it wants to use with the device
- parent re-open the device and re-initialize everything
So this is putting quite a burden on a number of steps the parent
_must_ do in order to keep control of memory exposed to the device.
Not doing so can potentialy lead (it depends on who does the COW
first) to the child taking control of memory use by the device,
memory which was mapped by the parent before the child was created.
Forcing CLOEXEC and VM_DONTCOPY somewhat help to simplify this,
but you still need to stop, flush, unmap, before fork/exec and then
re-init everything after.
This is only when not using SVA/SVM, SVA/SVM is totaly fine from
that point of view, no issues whatsoever.
The solution i outlined in previous email do not have that above
issue either, no need to rely on user space doing that dance.
Thank you. I get the point. I'm now trying to see if I can solve the problem by
seting the vma to VM_SHARED when the portiong is "shared to the hardware".
FYI you can not convert a private anonymous vma to a share one it is
illegal AFAIK at least i never heard of it and i am pretty sure the
mm code would break if that happens. The user space is the one that
decide what flags a vma has, not the kernel. Modulo few flags like
DONTCOPY that can be force set by device driver for their vma ie vma
of an mmap against the device file.
If you don't like my solution here is another one but it is ugly and
i think it is a bad idea. Again this is for the non SVA/SVM case and
it assumes that the command queue is a mmap() of the device file:
(A) register mmu_notifier
(B) on _every_ invalidate range callback (_no matter_ what is the
range) you zap the command queue mapped to user space (this is
because you can't tell if the callback happens for a fork or
something else) wait for the hardware queue to finish and clear
all the iommu/dma mapping and you unpin all the pages ie
put_page()
(C) in device file vma page fault handler (vm_operations_struct.
fault) you redo all the GUP and redo all the iommu/dma mapping
and you remap the command queue to the userspace
In (C) you can remap different command queue if you are in the child
than in the parent (just look at current->mm and compare it to the
one the command queue was created against).
Note that this solution will be much __slower__ than what i described
in my previous email. You will see that mmu notifier callbacks happens
often and for tons of reasons and you will be _constantly_ undoing and
redoing tons of work.
This can be mitigated if you can differentiate reasons behind a mmu
notifier callback. I posted patchset to do that a while ago and i
intend to post it again in the next month or so. But this would still
be a bad idea and solution i described previously is much more sane.
Trying to pretend you can have the same thing as SVA/SVM without SVA
is not a good idea. The non SVA case can still expose same API (like
i described previously) but should go through kernel for _every_
hardware submission (you can batch multiple commands in one submission).
Not doing so is way too risky from my POV.
Cheers,
Jérôme
You are quite right. I tried all the way to find a leak in the mm system and
fail to. I will tried other way or maybe discard the non-SVA scenario.
Cheers,
--
-Kenneth(Hisilicon)
On Fri, Sep 21, 2018 at 06:03:14PM +0800, Kenneth Lee wrote:
On Sun, Sep 16, 2018 at 09:42:44PM -0400, Jerome Glisse wrote:
quoted
So i want to summarize issues i have as this threads have dig deep into
details. For this i would like to differentiate two cases first the easy
one when relying on SVA/SVM. Then the second one when there is no SVA/SVM.
In both cases your objectives as i understand them:
[R1]- expose a common user space API that make it easy to share boiler
plate code accross many devices (discovering devices, opening
device, creating context, creating command queue ...).
[R2]- try to share the device as much as possible up to device limits
(number of independant queues the device has)
[R3]- minimize syscall by allowing user space to directly schedule on the
device queue without a round trip to the kernel
I don't think i missed any.
(1) Device with SVA/SVM
For that case it is easy, you do not need to be in VFIO or part of any
thing specific in the kernel. There is no security risk (modulo bug in
the SVA/SVM silicon). Fork/exec is properly handle and binding a process
to a device is just couple dozen lines of code.
(2) Device does not have SVA/SVM (or it is disabled)
You want to still allow device to be part of your framework. However
here i see fundamentals securities issues and you move the burden of
being careful to user space which i think is a bad idea. We should
never trus the userspace from kernel space.
To keep the same API for the user space code you want a 1:1 mapping
between device physical address and process virtual address (ie if
device access device physical address A it is accessing the same
memory as what is backing the virtual address A in the process.
Security issues are on two things:
[I1]- fork/exec, a process who opened any such device and created an
active queue can transfer without its knowledge control of its
commands queue through COW. The parent map some anonymous region
to the device as a command queue buffer but because of COW the
parent can be the first to copy on write and thus the child can
inherit the original pages that are mapped to the hardware.
Here parent lose control and child gain it.
Hi, Jerome,
I reconsider your logic. I think the problem can be solved. Let us separate the
SVA/SVM feature into two: fault-from-device and device-va-awareness. A device
with iommu can support only device-va-awareness or both.
Not sure i follow, are you also talking about the non SVA/SVM case here ?
Either device has SVA/SVM, either it does not. The fact that device can
use same physical address PA for its access as the process virtual address
VA does not change any of the issues listed here, same issues would apply
if device was using PA != VA ...
VFIO works on top of iommu, so it will support at least device-va-awareness. For
the COW problem, it can be taken as a mmu synchronization issue. If the mmu page
table is changed, it should be synchronize to iommu (via iommu_notifier). In the
case that the device support fault-from-device, it will work fine. In the case
that it supports only device-va-awareness, we can prefault (handle_mm_fault)
also via iommu_notifier and reset to iommu page table.
So this can be considered as a bug of VFIO, cannot it?
So again SVA/SVM is fine because it uses the same CPU page table so anything
done to the process address space reflect automaticly to the device. Nothing
to do for SVA/SVM.
For non SVA/SVM you _must_ unmap device command queues, flush and wait for
all pending commands, unmap all the IOMMU mapping and wait for the process
to fault on the unmapped device command queue before trying to restore any
thing. This would be terribly slow i described it in another email.
Also you can not fault inside a mmu_notifier it is illegal, all you can do
is unmap thing and wait for access to finish.
quoted
[I2]- Because of [R3] you want to allow userspace to schedule commands
on the device without doing an ioctl and thus here user space
can schedule any commands to the device with any address. What
happens if that address have not been mapped by the user space
is undefined and in fact can not be defined as what each IOMMU
does on invalid address access is different from IOMMU to IOMMU.
In case of a bad IOMMU, or simply an IOMMU improperly setup by
the kernel, this can potentialy allow user space to DMA anywhere.
[I3]- By relying on GUP in VFIO you are not abiding by the implicit
contract (at least i hope it is implicit) that you should not
try to map to the device any file backed vma (private or share).
The VFIO code never check the vma controlling the addresses that
are provided to VFIO_IOMMU_MAP_DMA ioctl. Which means that the
user space can provide file backed range.
I am guessing that the VFIO code never had any issues because its
number one user is QEMU and QEMU never does that (and that's good
as no one should ever do that).
So if process does that you are opening your self to serious file
system corruption (depending on file system this can lead to total
data loss for the filesystem).
Issue is that once you GUP you never abide to file system flushing
which write protect the page before writing to the disk. So
because the page is still map with write permission to the device
(assuming VFIO_IOMMU_MAP_DMA was a write map) then the device can
write to the page while it is in the middle of being written back
to disk. Consult your nearest file system specialist to ask him
how bad that can be.
In the case, we cannot do anything if the device do not support
fault-from-device. But we can reject write map with file-backed mapping.
Yes this would avoid most issues with file-backed mapping, truncate would
still cause weird behavior but only for your device. But then your non
SVA/SVM case does not behave as the SVA/SVM case so why not do what i out-
lined below that allow same behavior modulo command flushing ...
It seems both issues can be solved under VFIO framework:) (But of cause, I don't
mean it has to)
The COW can not be solve without either the solution i described in this
previous mail below or the other one with mmu notifier. But the one below
is saner and has better performance.
quoted
[I4]- Design issue, mdev design As Far As I Understand It is about
sharing a single device to multiple clients (most obvious case
here is again QEMU guest). But you are going against that model,
in fact AFAIUI you are doing the exect opposite. When there is
no SVA/SVM you want only one mdev device that can not be share.
So this is counter intuitive to the mdev existing design. It is
not about sharing device among multiple users but about giving
exclusive access to the device to one user.
All the reasons above is why i believe a different model would serve
you and your user better. Below is a design that avoids all of the
above issues and still delivers all of your objectives with the
exceptions of the third one [R3] when there is no SVA/SVM.
Create a subsystem (very much boiler plate code) which allow device to
register themself against (very much like what you do in your current
patchset but outside of VFIO).
That subsystem will create a device file for each registered system and
expose a common API (ie set of ioctl) for each of those device files.
When user space create a queue (through an ioctl after opening the device
file) the kernel can return -EBUSY if all the device queue are in use,
or create a device queue and return a flag like SYNC_ONLY for device that
do not have SVA/SVM.
For device with SVA/SVM at the time the process create a queue you bind
the process PASID to the device queue. From there on the userspace can
schedule commands and use the device without going to kernel space.
For device without SVA/SVM you create a fake queue that is just pure
memory is not related to the device. From there on the userspace must
call an ioctl every time it wants the device to consume its queue
(hence why the SYNC_ONLY flag for synchronous operation only). The
kernel portion read the fake queue expose to user space and copy
commands into the real hardware queue but first it properly map any
of the process memory needed for those commands to the device and
adjust the device physical address with the one it gets from dma_map
API.
With that model it is "easy" to listen to mmu_notifier and to abide by
them to avoid issues [I1], [I3] and [I4]. You obviously avoid the [I2]
issue by only mapping a fake device queue to userspace.
So yes with that models it means that every device that wish to support
the non SVA/SVM case will have to do extra work (ie emulate its command
queue in software in the kernel). But by doing so, you support an
unlimited number of process on your device (ie all the process can share
one single hardware command queues or multiple hardware queues).
The big advantages i see here is that the process do not have to worry
about doing something wrong. You are protecting yourself and your user
from stupid mistakes.
I hope this is useful to you.
Cheers,
Jérôme
From: Kenneth Lee <hidden> Date: 2018-09-25 05:57:27
On Fri, Sep 21, 2018 at 10:52:01AM -0400, Jerome Glisse wrote:
Received: from popscn.huawei.com [10.3.17.45] by Turing-Arch-b with POP3
(fetchmail-6.3.26) for <kenny@localhost> (single-drop); Fri, 21 Sep 2018
23:00:01 +0800 (CST)
Received: from DGGEMM406-HUB.china.huawei.com (10.3.20.214) by
DGGEML403-HUB.china.huawei.com (10.3.17.33) with Microsoft SMTP Server
(TLS) id 14.3.399.0; Fri, 21 Sep 2018 22:52:20 +0800
Received: from dggwg01-in.huawei.com (172.30.65.38) by
DGGEMM406-HUB.china.huawei.com (10.3.20.214) with Microsoft SMTP Server id
14.3.399.0; Fri, 21 Sep 2018 22:52:16 +0800
Received: from mx1.redhat.com (unknown [209.132.183.28]) by Forcepoint
Email with ESMTPS id 912ECA2EC6662; Fri, 21 Sep 2018 22:52:12 +0800 (CST)
Received: from smtp.corp.redhat.com
(int-mx09.intmail.prod.int.phx2.redhat.com [10.5.11.24]) (using
TLSv1.2 with cipher AECDH-AES256-SHA (256/256 bits)) (No client
certificate requested) by mx1.redhat.com (Postfix) with ESMTPS id
25BC0792BB; Fri, 21 Sep 2018 14:52:10 +0000 (UTC)
Received: from redhat.com (ovpn-124-21.rdu2.redhat.com [10.10.124.21])
by smtp.corp.redhat.com (Postfix) with ESMTPS id 67B25308BDA0;
Fri, 21 Sep 2018 14:52:03 +0000 (UTC)
Date: Fri, 21 Sep 2018 10:52:01 -0400
From: Jerome Glisse <redacted>
To: Kenneth Lee <redacted>
CC: Kenneth Lee <redacted>, Herbert Xu
[off-list ref], kvm@vger.kernel.org, Jonathan Corbet
[off-list ref], Greg Kroah-Hartman [off-list ref], Joerg
Roedel [off-list ref], linux-doc@vger.kernel.org, Sanjay Kumar
[off-list ref], Hao Fang [off-list ref],
iommu@lists.linux-foundation.org, linux-kernel@vger.kernel.org,
linuxarm@huawei.com, Alex Williamson [off-list ref], Thomas
Gleixner [off-list ref], linux-crypto@vger.kernel.org, Zhou Wang
[off-list ref], Philippe Ombredanne [off-list ref],
Zaibo Xu [off-list ref], "David S . Miller" [off-list ref],
linux-accelerators@lists.ozlabs.org, Lu Baolu [off-list ref]
Subject: Re: [RFCv2 PATCH 0/7] A General Accelerator Framework, WarpDrive
Message-ID: [off-list ref]
References: [off-list ref]
[off-list ref]
<20180921100314.GH207969@Turing-Arch-b>
Content-Type: text/plain; charset="iso-8859-1"
Content-Disposition: inline
Content-Transfer-Encoding: 8bit
In-Reply-To: <20180921100314.GH207969@Turing-Arch-b>
User-Agent: Mutt/1.10.1 (2018-07-13)
X-Scanned-By: MIMEDefang 2.84 on 10.5.11.24
X-Greylist: Sender IP whitelisted, not delayed by milter-greylist-4.5.16
(mx1.redhat.com [10.5.110.39]); Fri, 21 Sep 2018 14:52:10 +0000 (UTC)
Return-Path: jglisse@redhat.com
X-MS-Exchange-Organization-AuthSource: DGGEMM406-HUB.china.huawei.com
X-MS-Exchange-Organization-AuthAs: Anonymous
MIME-Version: 1.0
On Fri, Sep 21, 2018 at 06:03:14PM +0800, Kenneth Lee wrote:
quoted
On Sun, Sep 16, 2018 at 09:42:44PM -0400, Jerome Glisse wrote:
quoted
So i want to summarize issues i have as this threads have dig deep into
details. For this i would like to differentiate two cases first the easy
one when relying on SVA/SVM. Then the second one when there is no SVA/SVM.
In both cases your objectives as i understand them:
[R1]- expose a common user space API that make it easy to share boiler
plate code accross many devices (discovering devices, opening
device, creating context, creating command queue ...).
[R2]- try to share the device as much as possible up to device limits
(number of independant queues the device has)
[R3]- minimize syscall by allowing user space to directly schedule on the
device queue without a round trip to the kernel
I don't think i missed any.
(1) Device with SVA/SVM
For that case it is easy, you do not need to be in VFIO or part of any
thing specific in the kernel. There is no security risk (modulo bug in
the SVA/SVM silicon). Fork/exec is properly handle and binding a process
to a device is just couple dozen lines of code.
(2) Device does not have SVA/SVM (or it is disabled)
You want to still allow device to be part of your framework. However
here i see fundamentals securities issues and you move the burden of
being careful to user space which i think is a bad idea. We should
never trus the userspace from kernel space.
To keep the same API for the user space code you want a 1:1 mapping
between device physical address and process virtual address (ie if
device access device physical address A it is accessing the same
memory as what is backing the virtual address A in the process.
Security issues are on two things:
[I1]- fork/exec, a process who opened any such device and created an
active queue can transfer without its knowledge control of its
commands queue through COW. The parent map some anonymous region
to the device as a command queue buffer but because of COW the
parent can be the first to copy on write and thus the child can
inherit the original pages that are mapped to the hardware.
Here parent lose control and child gain it.
Hi, Jerome,
I reconsider your logic. I think the problem can be solved. Let us separate the
SVA/SVM feature into two: fault-from-device and device-va-awareness. A device
with iommu can support only device-va-awareness or both.
Not sure i follow, are you also talking about the non SVA/SVM case here ?
Either device has SVA/SVM, either it does not. The fact that device can
use same physical address PA for its access as the process virtual address
VA does not change any of the issues listed here, same issues would apply
if device was using PA != VA ...
quoted
VFIO works on top of iommu, so it will support at least device-va-awareness. For
the COW problem, it can be taken as a mmu synchronization issue. If the mmu page
table is changed, it should be synchronize to iommu (via iommu_notifier). In the
case that the device support fault-from-device, it will work fine. In the case
that it supports only device-va-awareness, we can prefault (handle_mm_fault)
also via iommu_notifier and reset to iommu page table.
So this can be considered as a bug of VFIO, cannot it?
So again SVA/SVM is fine because it uses the same CPU page table so anything
done to the process address space reflect automaticly to the device. Nothing
to do for SVA/SVM.
For non SVA/SVM you _must_ unmap device command queues, flush and wait for
all pending commands, unmap all the IOMMU mapping and wait for the process
to fault on the unmapped device command queue before trying to restore any
thing. This would be terribly slow i described it in another email.
Also you can not fault inside a mmu_notifier it is illegal, all you can do
is unmap thing and wait for access to finish.
quoted
quoted
[I2]- Because of [R3] you want to allow userspace to schedule commands
on the device without doing an ioctl and thus here user space
can schedule any commands to the device with any address. What
happens if that address have not been mapped by the user space
is undefined and in fact can not be defined as what each IOMMU
does on invalid address access is different from IOMMU to IOMMU.
In case of a bad IOMMU, or simply an IOMMU improperly setup by
the kernel, this can potentialy allow user space to DMA anywhere.
[I3]- By relying on GUP in VFIO you are not abiding by the implicit
contract (at least i hope it is implicit) that you should not
try to map to the device any file backed vma (private or share).
The VFIO code never check the vma controlling the addresses that
are provided to VFIO_IOMMU_MAP_DMA ioctl. Which means that the
user space can provide file backed range.
I am guessing that the VFIO code never had any issues because its
number one user is QEMU and QEMU never does that (and that's good
as no one should ever do that).
So if process does that you are opening your self to serious file
system corruption (depending on file system this can lead to total
data loss for the filesystem).
Issue is that once you GUP you never abide to file system flushing
which write protect the page before writing to the disk. So
because the page is still map with write permission to the device
(assuming VFIO_IOMMU_MAP_DMA was a write map) then the device can
write to the page while it is in the middle of being written back
to disk. Consult your nearest file system specialist to ask him
how bad that can be.
In the case, we cannot do anything if the device do not support
fault-from-device. But we can reject write map with file-backed mapping.
Yes this would avoid most issues with file-backed mapping, truncate would
still cause weird behavior but only for your device. But then your non
SVA/SVM case does not behave as the SVA/SVM case so why not do what i out-
lined below that allow same behavior modulo command flushing ...
quoted
It seems both issues can be solved under VFIO framework:) (But of cause, I don't
mean it has to)
The COW can not be solve without either the solution i described in this
previous mail below or the other one with mmu notifier. But the one below
is saner and has better performance.
quoted
quoted
[I4]- Design issue, mdev design As Far As I Understand It is about
sharing a single device to multiple clients (most obvious case
here is again QEMU guest). But you are going against that model,
in fact AFAIUI you are doing the exect opposite. When there is
no SVA/SVM you want only one mdev device that can not be share.
So this is counter intuitive to the mdev existing design. It is
not about sharing device among multiple users but about giving
exclusive access to the device to one user.
All the reasons above is why i believe a different model would serve
you and your user better. Below is a design that avoids all of the
above issues and still delivers all of your objectives with the
exceptions of the third one [R3] when there is no SVA/SVM.
Create a subsystem (very much boiler plate code) which allow device to
register themself against (very much like what you do in your current
patchset but outside of VFIO).
That subsystem will create a device file for each registered system and
expose a common API (ie set of ioctl) for each of those device files.
When user space create a queue (through an ioctl after opening the device
file) the kernel can return -EBUSY if all the device queue are in use,
or create a device queue and return a flag like SYNC_ONLY for device that
do not have SVA/SVM.
For device with SVA/SVM at the time the process create a queue you bind
the process PASID to the device queue. From there on the userspace can
schedule commands and use the device without going to kernel space.
For device without SVA/SVM you create a fake queue that is just pure
memory is not related to the device. From there on the userspace must
call an ioctl every time it wants the device to consume its queue
(hence why the SYNC_ONLY flag for synchronous operation only). The
kernel portion read the fake queue expose to user space and copy
commands into the real hardware queue but first it properly map any
of the process memory needed for those commands to the device and
adjust the device physical address with the one it gets from dma_map
API.
With that model it is "easy" to listen to mmu_notifier and to abide by
them to avoid issues [I1], [I3] and [I4]. You obviously avoid the [I2]
issue by only mapping a fake device queue to userspace.
So yes with that models it means that every device that wish to support
the non SVA/SVM case will have to do extra work (ie emulate its command
queue in software in the kernel). But by doing so, you support an
unlimited number of process on your device (ie all the process can share
one single hardware command queues or multiple hardware queues).
The big advantages i see here is that the process do not have to worry
about doing something wrong. You are protecting yourself and your user
from stupid mistakes.
I hope this is useful to you.
Cheers,
Jérôme
Cheers
--
-Kenneth(Hisilicon)
Thank you very much. Jerome. I will push RFCv3 without VFIO soon.
The basic idea will be:
1. Make a chrdev interface for any WarpDrive device instance. Allocate queue
when it is openned.
2. For device with SVA/SVM, the device (in the queue context) share
application's process page table via iommu.
3. For device without SVA/SVM, the application can mmap any WarpDrive device as
its "shared memory". The memory will be shared with any queue openned by this
process, before or after the queues are openned. The VAs will still be shared
with the hardware (by iommu or by kernel translation). This will fulfill the
requirement of "Sharing VA", but it will need the application to use those
mmapped memory only.
Cheers
--
-Kenneth(Hisilicon)