Thread (30 messages) 30 messages, 11 authors, 2026-08-26

Re: [PATCH] swiotlb: avoid double copy with swiotlb on tx socket

From: Mostafa Saleh <smostafa@google.com>
Date: 2026-08-25 16:33:26
Also in: driver-core, linux-iommu, linux-mm, lkml

On Mon, Aug 24, 2026 at 10:59:21AM +0200, Dragos Tatulea wrote:

On 16.06.26 13:06, Mostafa Saleh wrote:
quoted
On Tue, Jun 16, 2026 at 02:33:52AM +0200, Luigi Rizzo wrote:
quoted
On Tue, Jun 16, 2026 at 2:25 AM Jakub Kicinski [off-list ref] wrote:
quoted
On Mon, 15 Jun 2026 23:42:20 +0000 Luigi Rizzo wrote:
quoted
The use of swiotlb causes an extra data copy on I/O.  For tx sockets,
especially with greedy senders, this has a high chance of happening in
the softirq handler for tx network interrupts, creating a significant
performance bottleneck.
What's the use case? I associate swiotlb with debug / testing mostly,
so it'd be useful for people like me to explain why you care.
Ah sorry, I forgot to mention.
swiotlb is used in guest kernels for confidential computing VMs.
Ordinary memory pages are encrypted and the host or devices
have no way to decrypt them, so the kernel must use
unencrypted bounce buffers to exchange data with I/O devices.
I started looking into the same problem recently, to reduce the
bouncing in protected KVM (pKVM) confidential guests.
My first attempt was to update dma_direct_map_phys() to skip
bouncing and do inline memory decryption (for pKVM that is a hypercall
which updates the stage-2 page tables), however, that was really slow
compared to the memcpy in bouncing even for massive pages.
My conclusion was similar that we need to solve this at construction
by making this memory allocated from a pre-decrypted pool (which
does not have to be part of the SWIOTLB)
My initial idea was to teach some of the kernel subsystems (SKB,
BLK, SLAB) about "CoCo allocators" that allocate decrypted memory,
as this is not a net specific problem.
An example of this is Jiri's system_cc_shared heap which is a dma-buf
heap with decrypted memory for userspace.
quoted
I am still looking into this, I was planning to bring this up in the
upcoming LPC.
I will give this patch a try. However, I believe that we need a more
generalised concept for CoCo pre-decrypted allocators in the kernel.
There is a talk at LPC in the networking track about this [2]. This is
exactly the type of discussion that I was hoping to have there.
I see, thanks for point that. I plan to be in LPC, so I will aim to
attend this talk.
Besides the issues mentioned in this thread we've also found that a lot
of overhead can come only from swiotlb allocations when running many queues.

I will add information about this series in my talk. Hopefully I will also
have time to add some numbers for comparison.
I had a quick look, I am not sure how that will shape at the end, but I
think we need a general solution beyond NICs. For example in pKVM the
bouncing is used with virtio devices, so doing this per-driver won't
really work and is not possible in some scenarios where memory is
allocated from the core kernel and then passed to the driver.
So, I was thinking the kernel relying on
CC_ATTR_MEM_ENCRYPT/force_dma_unencrypted() could detect that and
allocate pre-shared memory for those cases.

Thanks,
Mostafa
Sorry for the late reply but I spotted this thread only now by
accident.

[1] https://lore.kernel.org/all/20260325192352.437608-3-jiri@resnulli.us/ (local)
[2] https://lpc.events/event/20/contributions/2464

Thanks,
Dragos
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help