The skiboot firmware has a hot reset handler which fences the NVIDIA V100
GPU RAM on Witherspoons and makes accesses no-op instead of throwing HMIs:
https://github.com/open-power/skiboot/commit/fca2b2b839a67
Now we are going to pass V100 via VFIO which most certainly involves
KVM guests which are often terminated without getting a chance to offline
GPU RAM so we end up with a running machine with misconfigured memory.
Accessing this memory produces hardware management interrupts (HMI)
which bring the host down.
To suppress HMIs, this wires up this hot reset hook to vfio_pci_disable()
via pci_disable_device() which switches NPU2 to a safe mode and prevents
HMIs.
Signed-off-by: Alexey Kardashevskiy <redacted>
---
Changes:
v2:
* updated the commit log
---
arch/powerpc/platforms/powernv/pci-ioda.c | 10 ++++++++++
1 file changed, 10 insertions(+)
Ping?
On 02/10/2018 13:20, Alexey Kardashevskiy wrote:
quoted hunk
The skiboot firmware has a hot reset handler which fences the NVIDIA V100
GPU RAM on Witherspoons and makes accesses no-op instead of throwing HMIs:
https://github.com/open-power/skiboot/commit/fca2b2b839a67
Now we are going to pass V100 via VFIO which most certainly involves
KVM guests which are often terminated without getting a chance to offline
GPU RAM so we end up with a running machine with misconfigured memory.
Accessing this memory produces hardware management interrupts (HMI)
which bring the host down.
To suppress HMIs, this wires up this hot reset hook to vfio_pci_disable()
via pci_disable_device() which switches NPU2 to a safe mode and prevents
HMIs.
Signed-off-by: Alexey Kardashevskiy <redacted>
---
Changes:
v2:
* updated the commit log
---
arch/powerpc/platforms/powernv/pci-ioda.c | 10 ++++++++++
1 file changed, 10 insertions(+)
Hi Alexey,
Looking at the skiboot side I think we only fence the NVLink bricks as part of a
PCIe function level reset (FLR) rather than a PCI Hot or Fundamental reset which
I believe is what the code here does. So to fence the bricks you would need to
do either a FLR on the given link or alter Skiboot to fence a given link as part
of a hot reset.
- Alistair
On Monday, 15 October 2018 6:17:51 PM AEDT Alexey Kardashevskiy wrote:
Ping?
On 02/10/2018 13:20, Alexey Kardashevskiy wrote:
quoted
The skiboot firmware has a hot reset handler which fences the NVIDIA V100
GPU RAM on Witherspoons and makes accesses no-op instead of throwing HMIs:
https://github.com/open-power/skiboot/commit/fca2b2b839a67
Now we are going to pass V100 via VFIO which most certainly involves
KVM guests which are often terminated without getting a chance to offline
GPU RAM so we end up with a running machine with misconfigured memory.
Accessing this memory produces hardware management interrupts (HMI)
which bring the host down.
To suppress HMIs, this wires up this hot reset hook to vfio_pci_disable()
via pci_disable_device() which switches NPU2 to a safe mode and prevents
HMIs.
Signed-off-by: Alexey Kardashevskiy <redacted>
---
Changes:
v2:
* updated the commit log
---
arch/powerpc/platforms/powernv/pci-ioda.c | 10 ++++++++++
1 file changed, 10 insertions(+)
Hi Alexey,
Looking at the skiboot side I think we only fence the NVLink bricks as part of a
PCIe function level reset (FLR) rather than a PCI Hot or Fundamental reset which
I believe is what the code here does. So to fence the bricks you would need to
do either a FLR on the given link or alter Skiboot to fence a given link as part
of a hot reset.
- Alistair
On Monday, 15 October 2018 6:17:51 PM AEDT Alexey Kardashevskiy wrote:
quoted
Ping?
On 02/10/2018 13:20, Alexey Kardashevskiy wrote:
quoted
The skiboot firmware has a hot reset handler which fences the NVIDIA V100
GPU RAM on Witherspoons and makes accesses no-op instead of throwing HMIs:
https://github.com/open-power/skiboot/commit/fca2b2b839a67
Now we are going to pass V100 via VFIO which most certainly involves
KVM guests which are often terminated without getting a chance to offline
GPU RAM so we end up with a running machine with misconfigured memory.
Accessing this memory produces hardware management interrupts (HMI)
which bring the host down.
To suppress HMIs, this wires up this hot reset hook to vfio_pci_disable()
via pci_disable_device() which switches NPU2 to a safe mode and prevents
HMIs.
Signed-off-by: Alexey Kardashevskiy <redacted>
---
Changes:
v2:
* updated the commit log
---
arch/powerpc/platforms/powernv/pci-ioda.c | 10 ++++++++++
1 file changed, 10 insertions(+)
Hi Alexey,
On Tuesday, 16 October 2018 12:37:49 PM AEDT Alexey Kardashevskiy wrote:
On 16/10/2018 11:38, Alistair Popple wrote:
quoted
Hi Alexey,
Looking at the skiboot side I think we only fence the NVLink bricks as part of a
PCIe function level reset (FLR) rather than a PCI Hot or Fundamental reset which
I believe is what the code here does. So to fence the bricks you would need to
do either a FLR on the given link or alter Skiboot to fence a given link as part
of a hot reset.
- Alistair
On Monday, 15 October 2018 6:17:51 PM AEDT Alexey Kardashevskiy wrote:
quoted
Ping?
On 02/10/2018 13:20, Alexey Kardashevskiy wrote:
quoted
The skiboot firmware has a hot reset handler which fences the NVIDIA V100
GPU RAM on Witherspoons and makes accesses no-op instead of throwing HMIs:
https://github.com/open-power/skiboot/commit/fca2b2b839a67
Now we are going to pass V100 via VFIO which most certainly involves
KVM guests which are often terminated without getting a chance to offline
GPU RAM so we end up with a running machine with misconfigured memory.
Accessing this memory produces hardware management interrupts (HMI)
which bring the host down.
To suppress HMIs, this wires up this hot reset hook to vfio_pci_disable()
via pci_disable_device() which switches NPU2 to a safe mode and prevents
HMIs.
Signed-off-by: Alexey Kardashevskiy <redacted>
---
Changes:
v2:
* updated the commit log
---
arch/powerpc/platforms/powernv/pci-ioda.c | 10 ++++++++++
1 file changed, 10 insertions(+)
Hi Alexey,
On Tuesday, 16 October 2018 12:37:49 PM AEDT Alexey Kardashevskiy wrote:
quoted
On 16/10/2018 11:38, Alistair Popple wrote:
quoted
Hi Alexey,
Looking at the skiboot side I think we only fence the NVLink bricks as part of a
PCIe function level reset (FLR) rather than a PCI Hot or Fundamental reset which
I believe is what the code here does. So to fence the bricks you would need to
do either a FLR on the given link or alter Skiboot to fence a given link as part
of a hot reset.
reset_ntl() does what npu2_dev_procedure_reset() does plus more stuff,
there nothing really in npu2_dev_procedure_reset() which reset_ntl()
does not do already from the hardware standpoint. And it did stop HMIs
for me though.
but ok, what will be sufficient then if not reset_ntl()?
- Alistair
quoted
quoted
- Alistair
On Monday, 15 October 2018 6:17:51 PM AEDT Alexey Kardashevskiy wrote:
quoted
Ping?
On 02/10/2018 13:20, Alexey Kardashevskiy wrote:
quoted
The skiboot firmware has a hot reset handler which fences the NVIDIA V100
GPU RAM on Witherspoons and makes accesses no-op instead of throwing HMIs:
https://github.com/open-power/skiboot/commit/fca2b2b839a67
Now we are going to pass V100 via VFIO which most certainly involves
KVM guests which are often terminated without getting a chance to offline
GPU RAM so we end up with a running machine with misconfigured memory.
Accessing this memory produces hardware management interrupts (HMI)
which bring the host down.
To suppress HMIs, this wires up this hot reset hook to vfio_pci_disable()
via pci_disable_device() which switches NPU2 to a safe mode and prevents
HMIs.
Signed-off-by: Alexey Kardashevskiy <redacted>
---
Changes:
v2:
* updated the commit log
---
arch/powerpc/platforms/powernv/pci-ioda.c | 10 ++++++++++
1 file changed, 10 insertions(+)
reset_ntl() does what npu2_dev_procedure_reset() does plus more stuff,
there nothing really in npu2_dev_procedure_reset() which reset_ntl()
does not do already from the hardware standpoint. And it did stop HMIs
for me though.
but ok, what will be sufficient then if not reset_ntl()?
Argh, yes you are correct. Specifically both npu2_dev_procedure_reset() and
reset_ntl() contain:
/* NTL Reset */
val = npu2_read(ndev->npu, NPU2_NTL_MISC_CFG1(ndev));
val |= PPC_BIT(8) | PPC_BIT(9);
npu2_write(ndev->npu, NPU2_NTL_MISC_CFG1(ndev), val);
Which should fence the brick. However from what I recall there was more to
reliably preventing HMIs than merely fencing the brick. It invovled a sequence
of fencing and flushing the cache with dcbf instructions at the right time which
is why we also have the FLR. Unfortunately I don't know the precise details,
perhaps if we send enough coffee Balbir's way he might be able remind us?
- Alistair
quoted
- Alistair
quoted
quoted
- Alistair
On Monday, 15 October 2018 6:17:51 PM AEDT Alexey Kardashevskiy wrote:
quoted
Ping?
On 02/10/2018 13:20, Alexey Kardashevskiy wrote:
quoted
The skiboot firmware has a hot reset handler which fences the NVIDIA V100
GPU RAM on Witherspoons and makes accesses no-op instead of throwing HMIs:
https://github.com/open-power/skiboot/commit/fca2b2b839a67
Now we are going to pass V100 via VFIO which most certainly involves
KVM guests which are often terminated without getting a chance to offline
GPU RAM so we end up with a running machine with misconfigured memory.
Accessing this memory produces hardware management interrupts (HMI)
which bring the host down.
To suppress HMIs, this wires up this hot reset hook to vfio_pci_disable()
via pci_disable_device() which switches NPU2 to a safe mode and prevents
HMIs.
Signed-off-by: Alexey Kardashevskiy <redacted>
---
Changes:
v2:
* updated the commit log
---
arch/powerpc/platforms/powernv/pci-ioda.c | 10 ++++++++++
1 file changed, 10 insertions(+)
reset_ntl() does what npu2_dev_procedure_reset() does plus more stuff,
there nothing really in npu2_dev_procedure_reset() which reset_ntl()
does not do already from the hardware standpoint. And it did stop HMIs
for me though.
but ok, what will be sufficient then if not reset_ntl()?
Argh, yes you are correct. Specifically both npu2_dev_procedure_reset() and
reset_ntl() contain:
/* NTL Reset */
val = npu2_read(ndev->npu, NPU2_NTL_MISC_CFG1(ndev));
val |= PPC_BIT(8) | PPC_BIT(9);
npu2_write(ndev->npu, NPU2_NTL_MISC_CFG1(ndev), val);
Which should fence the brick. However from what I recall there was more to
reliably preventing HMIs than merely fencing the brick. It invovled a sequence
of fencing and flushing the cache with dcbf instructions at the right time which
is why we also have the FLR. Unfortunately I don't know the precise details,
perhaps if we send enough coffee Balbir's way he might be able remind us?
He suggested and ack'ed that skiboot patch, I can repeat beers^wcoffee
but it won't change much ;)
- Alistair
quoted
quoted
- Alistair
quoted
quoted
- Alistair
On Monday, 15 October 2018 6:17:51 PM AEDT Alexey Kardashevskiy wrote:
quoted
Ping?
On 02/10/2018 13:20, Alexey Kardashevskiy wrote:
quoted
The skiboot firmware has a hot reset handler which fences the NVIDIA V100
GPU RAM on Witherspoons and makes accesses no-op instead of throwing HMIs:
https://github.com/open-power/skiboot/commit/fca2b2b839a67
Now we are going to pass V100 via VFIO which most certainly involves
KVM guests which are often terminated without getting a chance to offline
GPU RAM so we end up with a running machine with misconfigured memory.
Accessing this memory produces hardware management interrupts (HMI)
which bring the host down.
To suppress HMIs, this wires up this hot reset hook to vfio_pci_disable()
via pci_disable_device() which switches NPU2 to a safe mode and prevents
HMIs.
Signed-off-by: Alexey Kardashevskiy <redacted>
---
Changes:
v2:
* updated the commit log
---
arch/powerpc/platforms/powernv/pci-ioda.c | 10 ++++++++++
1 file changed, 10 insertions(+)
On Tuesday, 16 October 2018 1:22:53 PM AEDT Alexey Kardashevskiy wrote:
On 16/10/2018 13:19, Alistair Popple wrote:
quoted
quoted
reset_ntl() does what npu2_dev_procedure_reset() does plus more stuff,
there nothing really in npu2_dev_procedure_reset() which reset_ntl()
does not do already from the hardware standpoint. And it did stop HMIs
for me though.
but ok, what will be sufficient then if not reset_ntl()?
Argh, yes you are correct. Specifically both npu2_dev_procedure_reset() and
reset_ntl() contain:
/* NTL Reset */
val = npu2_read(ndev->npu, NPU2_NTL_MISC_CFG1(ndev));
val |= PPC_BIT(8) | PPC_BIT(9);
npu2_write(ndev->npu, NPU2_NTL_MISC_CFG1(ndev), val);
Which should fence the brick. However from what I recall there was more to
reliably preventing HMIs than merely fencing the brick. It invovled a sequence
of fencing and flushing the cache with dcbf instructions at the right time which
is why we also have the FLR. Unfortunately I don't know the precise details,
perhaps if we send enough coffee Balbir's way he might be able remind us?
He suggested and ack'ed that skiboot patch, I can repeat beers^wcoffee
but it won't change much ;)
Ha. Couldn't hurt ;)
I was pretty sure flushing the caches was an important part of the sequence to
avoid HMI's. I believe you are trying to deal with unexpected guest terminations
which means the driver won't have a chance to flush the caches prior to
termination so wouldn't you also need to do that somewhere? Unless the driver
does it at startup?
- Alistair
quoted
- Alistair
quoted
quoted
- Alistair
quoted
quoted
- Alistair
On Monday, 15 October 2018 6:17:51 PM AEDT Alexey Kardashevskiy wrote:
quoted
Ping?
On 02/10/2018 13:20, Alexey Kardashevskiy wrote:
quoted
The skiboot firmware has a hot reset handler which fences the NVIDIA V100
GPU RAM on Witherspoons and makes accesses no-op instead of throwing HMIs:
https://github.com/open-power/skiboot/commit/fca2b2b839a67
Now we are going to pass V100 via VFIO which most certainly involves
KVM guests which are often terminated without getting a chance to offline
GPU RAM so we end up with a running machine with misconfigured memory.
Accessing this memory produces hardware management interrupts (HMI)
which bring the host down.
To suppress HMIs, this wires up this hot reset hook to vfio_pci_disable()
via pci_disable_device() which switches NPU2 to a safe mode and prevents
HMIs.
Signed-off-by: Alexey Kardashevskiy <redacted>
---
Changes:
v2:
* updated the commit log
---
arch/powerpc/platforms/powernv/pci-ioda.c | 10 ++++++++++
1 file changed, 10 insertions(+)
On Tuesday, 16 October 2018 1:22:53 PM AEDT Alexey Kardashevskiy wrote:
quoted
On 16/10/2018 13:19, Alistair Popple wrote:
quoted
quoted
reset_ntl() does what npu2_dev_procedure_reset() does plus more stuff,
there nothing really in npu2_dev_procedure_reset() which reset_ntl()
does not do already from the hardware standpoint. And it did stop HMIs
for me though.
but ok, what will be sufficient then if not reset_ntl()?
Argh, yes you are correct. Specifically both npu2_dev_procedure_reset() and
reset_ntl() contain:
/* NTL Reset */
val = npu2_read(ndev->npu, NPU2_NTL_MISC_CFG1(ndev));
val |= PPC_BIT(8) | PPC_BIT(9);
npu2_write(ndev->npu, NPU2_NTL_MISC_CFG1(ndev), val);
Which should fence the brick. However from what I recall there was more to
reliably preventing HMIs than merely fencing the brick. It invovled a sequence
of fencing and flushing the cache with dcbf instructions at the right time which
is why we also have the FLR. Unfortunately I don't know the precise details,
perhaps if we send enough coffee Balbir's way he might be able remind us?
He suggested and ack'ed that skiboot patch, I can repeat beers^wcoffee
but it won't change much ;)
Ha. Couldn't hurt ;)
I was pretty sure flushing the caches was an important part of the sequence to
avoid HMI's. I believe you are trying to deal with unexpected guest terminations
which means the driver won't have a chance to flush the caches prior to
termination so
Correct.
wouldn't you also need to do that somewhere? Unless the driver
does it at startup?
VFIO performs GPU reset so I'd expect the GPUs to flush its caches
without any software interactions. Am I hoping for too much here?
- Alistair
quoted
quoted
- Alistair
quoted
quoted
- Alistair
quoted
quoted
- Alistair
On Monday, 15 October 2018 6:17:51 PM AEDT Alexey Kardashevskiy wrote:
quoted
Ping?
On 02/10/2018 13:20, Alexey Kardashevskiy wrote:
quoted
The skiboot firmware has a hot reset handler which fences the NVIDIA V100
GPU RAM on Witherspoons and makes accesses no-op instead of throwing HMIs:
https://github.com/open-power/skiboot/commit/fca2b2b839a67
Now we are going to pass V100 via VFIO which most certainly involves
KVM guests which are often terminated without getting a chance to offline
GPU RAM so we end up with a running machine with misconfigured memory.
Accessing this memory produces hardware management interrupts (HMI)
which bring the host down.
To suppress HMIs, this wires up this hot reset hook to vfio_pci_disable()
via pci_disable_device() which switches NPU2 to a safe mode and prevents
HMIs.
Signed-off-by: Alexey Kardashevskiy <redacted>
---
Changes:
v2:
* updated the commit log
---
arch/powerpc/platforms/powernv/pci-ioda.c | 10 ++++++++++
1 file changed, 10 insertions(+)
wouldn't you also need to do that somewhere? Unless the driver
does it at startup?
VFIO performs GPU reset so I'd expect the GPUs to flush its caches
without any software interactions. Am I hoping for too much here?
Sadly you are. It's not the GPU caches that need flushing, it's the CPU caches.
This needs to happen as part of the reset sequence, so I guess you would need
to add it to the VFIO driver.
- Alistair
quoted
- Alistair
quoted
quoted
- Alistair
quoted
quoted
- Alistair
quoted
quoted
- Alistair
On Monday, 15 October 2018 6:17:51 PM AEDT Alexey Kardashevskiy
wrote:
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
Ping?
On 02/10/2018 13:20, Alexey Kardashevskiy wrote:
quoted
The skiboot firmware has a hot reset handler which fences the
NVIDIA V100
GPU RAM on Witherspoons and makes accesses no-op instead of
throwing HMIs:
https://github.com/open-power/skiboot/commit/fca2b2b839a67
Now we are going to pass V100 via VFIO which most certainly
involves
KVM guests which are often terminated without getting a chance to
offline
GPU RAM so we end up with a running machine with misconfigured
memory.
Accessing this memory produces hardware management interrupts
(HMI)
which bring the host down.
To suppress HMIs, this wires up this hot reset hook to
vfio_pci_disable()
via pci_disable_device() which switches NPU2 to a safe mode and
prevents
HMIs.
Signed-off-by: Alexey Kardashevskiy <redacted>
---
Changes:
v2:
* updated the commit log
---
arch/powerpc/platforms/powernv/pci-ioda.c | 10 ++++++++++
1 file changed, 10 insertions(+)
wouldn't you also need to do that somewhere? Unless the driver
does it at startup?
VFIO performs GPU reset so I'd expect the GPUs to flush its caches
without any software interactions. Am I hoping for too much here?
Sadly you are. It's not the GPU caches that need flushing, it's the CPU caches.
This needs to happen as part of the reset sequence, so I guess you would need
to add it to the VFIO driver.
Well, ok. Caches need flushing, will look into this but this fencing is
still needed, is not it?
- Alistair
quoted
quoted
- Alistair
quoted
quoted
- Alistair
quoted
quoted
- Alistair
quoted
quoted
- Alistair
On Monday, 15 October 2018 6:17:51 PM AEDT Alexey Kardashevskiy
wrote:
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
Ping?
On 02/10/2018 13:20, Alexey Kardashevskiy wrote:
quoted
The skiboot firmware has a hot reset handler which fences the
NVIDIA V100
GPU RAM on Witherspoons and makes accesses no-op instead of
throwing HMIs:
https://github.com/open-power/skiboot/commit/fca2b2b839a67
Now we are going to pass V100 via VFIO which most certainly
involves
KVM guests which are often terminated without getting a chance to
offline
GPU RAM so we end up with a running machine with misconfigured
memory.
Accessing this memory produces hardware management interrupts
(HMI)
which bring the host down.
To suppress HMIs, this wires up this hot reset hook to
vfio_pci_disable()
via pci_disable_device() which switches NPU2 to a safe mode and
prevents
HMIs.
Signed-off-by: Alexey Kardashevskiy <redacted>
---
Changes:
v2:
* updated the commit log
---
arch/powerpc/platforms/powernv/pci-ioda.c | 10 ++++++++++
1 file changed, 10 insertions(+)
wouldn't you also need to do that somewhere? Unless the driver
does it at startup?
VFIO performs GPU reset so I'd expect the GPUs to flush its caches
without any software interactions. Am I hoping for too much here?
Sadly you are. It's not the GPU caches that need flushing, it's the CPU
caches. This needs to happen as part of the reset sequence, so I guess
you would need to add it to the VFIO driver.
Well, ok. Caches need flushing, will look into this but this fencing is
still needed, is not it?
Yes. Although without the flushing I think you may get HMI's on any subsequent
driver loads.
So from the point of view of what happens on the Skiboot/HW side this looks ok
so long as all links on an NPU are assigned to the same guest (as this call
resets every link on a given NPU).
Acked-by: Alistair Popple <redacted>
quoted
- Alistair
quoted
quoted
- Alistair
quoted
quoted
- Alistair
quoted
quoted
- Alistair
quoted
quoted
- Alistair
On Monday, 15 October 2018 6:17:51 PM AEDT Alexey Kardashevskiy
wrote:
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
quoted
Ping?
On 02/10/2018 13:20, Alexey Kardashevskiy wrote:
quoted
The skiboot firmware has a hot reset handler which fences the
NVIDIA V100
GPU RAM on Witherspoons and makes accesses no-op instead of
throwing HMIs:
https://github.com/open-power/skiboot/commit/fca2b2b839a67
Now we are going to pass V100 via VFIO which most certainly
involves
KVM guests which are often terminated without getting a chance
to
offline
GPU RAM so we end up with a running machine with misconfigured
memory.
Accessing this memory produces hardware management interrupts
(HMI)
which bring the host down.
To suppress HMIs, this wires up this hot reset hook to
vfio_pci_disable()
via pci_disable_device() which switches NPU2 to a safe mode and
prevents
HMIs.
Signed-off-by: Alexey Kardashevskiy <redacted>
---
Changes:
v2:
* updated the commit log
---
arch/powerpc/platforms/powernv/pci-ioda.c | 10 ++++++++++
1 file changed, 10 insertions(+)