From: Mauricio Faria de Oliveira <hidden> Date: 2016-06-10 13:03:59
This prevents flooding the logs with 'iommu_alloc failed' messages
while I/O is performed (normally) to very fast devices (e.g. NVMe).
That error is not necessarily a problem; device drivers can retry
later / reschedule the requests for which the allocation failed,
and handle things gracefully for the caller stack on top of them.
This helps at least with NVMe devices without "64-bit"/direct DMA
window scenarios (e.g., systems with more than a few terabytes of
memory, on which DDW cannot be enabled, currently), where just an
'dd' command can trigger errors.
# dd if=/dev/zero of=/dev/nvme0n1p1 bs=64k count=512k
<...>
# echo $?
0
# dmesg
nvme 0000:00:06.0: iommu_alloc failed, tbl c0000001fa67c520 vaddr c000000151c90000 npages 16
nvme 0000:00:06.0: iommu_alloc failed, tbl c0000001fa67c520 vaddr c000000151c90000 npages 16
nvme 0000:00:06.0: iommu_alloc failed, tbl c0000001fa67c520 vaddr c000000151c90000 npages 16
<...>
ppc_iommu_map_sg: 8186 callbacks suppressed
nvme 0000:00:06.0: iommu_alloc failed, tbl c0000001fa67c520 vaddr c0000000fa5c0000 npages 16
nvme 0000:00:06.0: iommu_alloc failed, tbl c0000001fa67c520 vaddr c000000100440000 npages 16
nvme 0000:00:06.0: iommu_alloc failed, tbl c0000001fa67c520 vaddr c000000100440000 npages 16
<...>
ppc_iommu_map_sg: 5707 callbacks suppressed
nvme 0000:00:06.0: iommu_alloc failed, tbl c0000001fa67c520 vaddr c0000000b5f50000 npages 16
nvme 0000:00:06.0: iommu_alloc failed, tbl c0000001fa67c520 vaddr c0000000b5c60000 npages 16
nvme 0000:00:06.0: iommu_alloc failed, tbl c0000001fa67c520 vaddr c0000000b4b30000 npages 16
<...>
Tested on next-20160609.
Signed-off-by: Mauricio Faria de Oliveira <redacted>
---
arch/powerpc/kernel/iommu.c | 15 ++++++---------
1 file changed, 6 insertions(+), 9 deletions(-)
From: Benjamin Herrenschmidt <benh@kernel.crashing.org> Date: 2016-06-11 23:02:36
On Fri, 2016-06-10 at 10:03 -0300, Mauricio Faria de Oliveira wrote:
This prevents flooding the logs with 'iommu_alloc failed' messages
while I/O is performed (normally) to very fast devices (e.g. NVMe).
That error is not necessarily a problem; device drivers can retry
later / reschedule the requests for which the allocation failed,
and handle things gracefully for the caller stack on top of them.
This helps at least with NVMe devices without "64-bit"/direct DMA
window scenarios (e.g., systems with more than a few terabytes of
memory, on which DDW cannot be enabled, currently), where just an
'dd' command can trigger errors.
I'm not fan of this. This is a very useful message to diagnose why,
for example, your network adapter is not working properly.
A lot of drivers don't deal well with IOMMU errors.
The fact that NVME trigger these is a problem that needs to be solved
differently.
Cheers,
Ben.
From: Mauricio Faria de Oliveira <hidden> Date: 2016-06-13 13:28:12
Hi Ben,
On 06/11/2016 08:02 PM, Benjamin Herrenschmidt wrote:
I'm not fan of this. This is a very useful message to diagnose why,
for example, your network adapter is not working properly.
A lot of drivers don't deal well with IOMMU errors.
The fact that NVME trigger these is a problem that needs to be solved
differently.
Ok, I understand your points.
Thanks for the review -- helps in setting another direction.
Kind regards,
--
Mauricio Faria de Oliveira
IBM Linux Technology Center
From: Benjamin Herrenschmidt <benh@kernel.crashing.org> Date: 2016-06-13 21:27:11
On Mon, 2016-06-13 at 10:27 -0300, Mauricio Faria de Oliveira wrote:
Hi Ben,
On 06/11/2016 08:02 PM, Benjamin Herrenschmidt wrote:
quoted
I'm not fan of this. This is a very useful message to diagnose why,
for example, your network adapter is not working properly.
A lot of drivers don't deal well with IOMMU errors.
The fact that NVME trigger these is a problem that needs to be
solved
differently.
Ok, I understand your points.
Thanks for the review -- helps in setting another direction.
I've been thinking about this a bit... it might be worthwhile adding
a dma_* call to query the approximate size of the IOMMU window, as
a way for the device to adjust its requirements dynamically.
Another option would be to use a dma_attr for silencing mapping errors
which NVME could use provided it does handle them gracefully ...
Cheers,
Ben.
From: Mauricio Faria de Oliveira <hidden> Date: 2016-06-13 21:43:52
Hi Ben,
On 06/13/2016 06:26 PM, Benjamin Herrenschmidt wrote:
I've been thinking about this a bit... it might be worthwhile adding
a dma_* call to query the approximate size of the IOMMU window, as
a way for the device to adjust its requirements dynamically.
Ok, cool; something like it was one of the options being discussed here.
What do you mean by 'approximate'? Maybe the size of 'free regions' in
the pools? -- not sure because iiuic the window size is static / 2 gig,
so didn't get why (or of what) to provide an approximation (for).
Another option would be to use a dma_attr for silencing mapping errors
which NVME could use provided it does handle them gracefully ...
Ah, that's new. Interesting. Thanks for suggestion!
--
Mauricio Faria de Oliveira
IBM Linux Technology Center
From: Benjamin Herrenschmidt <benh@kernel.crashing.org> Date: 2016-06-13 21:51:28
On Mon, 2016-06-13 at 18:43 -0300, Mauricio Faria de Oliveira wrote:
Hi Ben,
On 06/13/2016 06:26 PM, Benjamin Herrenschmidt wrote:
quoted
I've been thinking about this a bit... it might be worthwhile adding
a dma_* call to query the approximate size of the IOMMU window, as
a way for the device to adjust its requirements dynamically.
Ok, cool; something like it was one of the options being discussed here.
What do you mean by 'approximate'? Maybe the size of 'free regions' in
the pools? -- not sure because iiuic the window size is static / 2 gig,
so didn't get why (or of what) to provide an approximation (for).
Approximate wasn't a great choice of word but what I meant is:
- The size doesn't mean you can do an allocation that size (pools
layout etc..)
- And it might be shared with another device (though less likely
these days).
quoted
Another option would be to use a dma_attr for silencing mapping errors
which NVME could use provided it does handle them gracefully ...
Ah, that's new. Interesting. Thanks for suggestion!
From: Benjamin Herrenschmidt <benh@kernel.crashing.org> Date: 2016-08-03 21:34:48
On Wed, 2016-08-03 at 16:39 -0300, Mauricio Faria de Oliveira wrote:
Hi Ben,
On 06/13/2016 06:26 PM, Benjamin Herrenschmidt wrote:
quoted
Another option would be to use a dma_attr for silencing mapping
errors
which NVME could use provided it does handle them gracefully ...
I recently submitted patches that implement your suggestion [1].
May you please review/comment if they're OK with you?
I think this is best done by the relevant community maintainer,
I just threw an idea but I'm not that familiar with the details :-)
Did you send them to the lkml list ?