Thread (27 messages) 27 messages, 2 authors, 23h ago

Re: [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support

From: mingming cao <hidden>
Date: 2026-09-26 17:41:27
Also in: netdev

Hi Jacub,

I noticed this v7 series is flagged red on patchwork. The apply failure 
is due to a dependency on my own [PATCH net 0/2] (Message-ID: 
cover.1790357373.git.mmc@linux.ibm.com) sent the same day — kept 
separate per your earlier feedback to peel fixes out of feature series. 
Both touch ibmveth_open() in the same region, and the net pair is not 
yet in net-next.

The v7 MQ series is based on net-next 161ea2d4f2a7 (2026-09-24). The 
fixes are already subsumed by MQ patches 3 and 6.

Shall I wait and rebase MQ once the net pair lands in net-next, or send 
a v8 now with the net pair folded in? Happy to do either.

Thanks, Mingming

On 9/25/26 11:38 AM, Mingming Cao wrote:
Hi,

Power11 PHYP adds Virtual Ethernet multi-queue (MQ) RX: multiple
logical-LAN RX queues, per-queue buffer posting, and completion
delivery. Guest Linux did not use that; ibmveth still registered one
RX queue even when PHYP was MQ-capable.

This series adds the ibmveth MQ client for net-next. When PHYP
advertises IBMVETH_ILLAN_RX_MULTI_QUEUE_SUPPORT via H_ILLAN_ATTRIBUTES,
probe enables MQ with a default RX count of min(num_online_cpus(), 8)
(same cap as TX today); ethtool -L can raise RX up to 16. Packets are
received on per-queue NAPI. Older firmware without the bit is unchanged.
Queue selection remains firmware-defined (PHYP hash). Ethtool RSS hash
get/set for that algorithm is deferred to a follow-up series so this
one stays MQ datapath only.

User-visible bits: ethtool -l/-L (channels); standard per-queue
packets/bytes/drops via netdev_stat_ops (ethtool -S keeps only
driver-specific counters; ndo_get_stats64 is the aggregate, including
retired-queue history); and a read-only debugfs buffer_pools dump
(v3's multi-line sysfs dump moved to debugfs; the historical queue-0
poolN/ sysfs ABI is unchanged).

Background:

ibmveth today uses one logical LAN, one set of buffer pools, and one
NAPI context. PHYP MQ mode gives each RX queue its own handle (post via
H_ADD_LOGICAL_LAN_BUFFERS_QUEUE, subordinate register via
H_REG_LOGICAL_LAN_QUEUE); traffic can land on any active queue. The
driver needs per-queue pools, IRQs, and NAPI to match. Legacy firmware
keeps the original hcall path.

Series layout (15 patches):

   1-2   Hypercall wrappers; MQ adapter layout (MAX_RX_QUEUES stays 1)
   3-9   Queue-aware helpers (still SQ runtime): RX, per-queue pools,
         IRQ, TX, PHYP, buffer submit (open/close 3-8); poll harden (9)
   10    Enable MQ datapath at probe/open (subordinate register helpers
         land here with first use)
   11-13 Per-queue RX/TX stats; get_channels MQ counts; debugfs buffer_pools
   14    Incremental RX resize; live ethtool -L rx
   15    Down-path rollback and mq_fallback max_rx cap

- Helper patches (3-8) reshape ibmveth_open()/close() into
   queue-aware helpers. Patch 9 hardens the SQ poll path with the same
   queue-index helpers; it does not change open/close. MQ stays off
   through 3-9: num_rx_queues stays 1 and multi_queue is false until
   patch 10. The live single-queue path still changes where the review
   required it (open/close unwind, IRQ remask, replenish lock, poll
   harden).
- Patch 10 is the switch: probe sets multi_queue from firmware, raises
   num_rx_queues, registers subordinates, and replenishes every active
   queue.
- Patch 11 moves counters per-queue and exports packets/bytes/drops
   through netdev_stat_ops. The thirteen existing -S keys stay; no
   hcall_* or pool%d_ keys.

Testing:

ppc64le PowerVM LPAR, MQ-capable firmware:
* ethtool -L cycling (16/1/8/11/1/3/16/8/1) with ping - no hangs
* ethtool -L under iperf3; link down/up during traffic
* ifdown/ifup under iperf3 RX+TX (MQ and ethtool -L rx 1)
* Legacy firmware (no MQ bit): open/close/stress on helper path
* Bisect-safe build and boot at every commit; W=1 clean at tip

Changes in v7:

Same 15 patches as v6. We followed up the v6 netdev-bot review
with replies; this v7 is the series after that, plus a few items
from our own re-review.

* Patch 3: update_rx_no_buffer() returns if buffer_list_addr[0] is
   NULL (the per-queue form stays in patch 10).
* Patch 6: synchronize_net() on the late TX-alloc unwind before RX is
   freed.
* Patch 8: replenish failure log names the wrapper from the filled
   count; advance ring on NULL buffer so poll does not spin.
* Patch 9: oversize bound is min(skb_tailroom, pool->buff_size).
* Patch 10: unregister_netdev before cancel_work_sync and gate reset
   on NETREG_REGISTERED (moved from patch 11); wait for pool kobject
   release before free_netdev(). Drop the probe CMO refresh and the two
   CMO follow-ups: CMO (Power9 and earlier) and MQ firmware (Power11+)
   do not coexist.
* Patch 11: replenish_lock on the close harvest (remove-path
   unregister/cancel reorder moved to patch 10 with the reset producer).
* Patch 12: set_channels() returns -EOPNOTSUPP on rx_count changes until
   patch 14 implements live resize.
* Patch 14: key buffer-list unmap on allocation presence because
   DMA address zero is valid; reject an RX count change while down
   with -EOPNOTSUPP. Failed H_FREE skips unmap and restores the
   surviving count; widen real_num before scale-up unmask. Scale-up
   register -EOPNOTSUPP latches mq_fallback. The scale-up /
   scale-down helper split is code motion only.
* Patch 15: publish the down-path RX count, which lifts patch 14's
   temporary rejection. get_channels max_tx is at least the live
   tx_count, and set_channels uses the same ceiling, so CPU offline
   cannot block an RX-only ethtool -L.
* Commit message / kdoc / comment / debug-log updates on 1, 3, 4, 5,
   7, 8, 9, 10, 11, 12, 14 and 15 (patch 1 also names the new hcalls in the perf
   powerpc-hcalls script and documents H_BUSY on the register-queue
   wrapper; patch 10 prints the register-queue failure with %ld).
* Kept: enable_irq on schedule_prep failure; mask PHYP before
   napi_disable; get_channels reports the live rx_count (no clamp).
* Reopen unwind in patches 3 and 6 is pre-existing. No Fixes: tag
   here. The SQ open-fail path is already on the list as
   [PATCH net 0/2] (Message-ID:
   [ref]). This series does
   not depend on it. If both land, keep the helper versions in
   patches 3/4/6; the net pair is the current single-queue path
   only.

Known leftovers (not this series):

* Single-queue: replenish vs free_buffer_pool is not serialized,
   irqsave still covers the whole fill, and close skips
   netpoll_poll_disable. That is a lock-protocol rewrite, not this
   series.
* RX IRQ teardown: teardown masks PHYP, disables NAPI, then masks
   again, but a poll tail that already passed the shutdown checks can
   still re-enable PHYP after that second mask and after free_irq, and
   a mask hcall that failed is never acknowledged. Closing this needs a
   poll/teardown handshake rather than another remask, so the ordering
   is unchanged here.

Changes in v6:

Same 15 patches as v5. Jakub v5 review folded in; per-patch detail is
below --- on each commit.

* Both new registration wrappers use plpar_hcall(), not plpar_hcall9().
* Poll: IPv4 check through skb->data; budget 0 does not complete NAPI.
* Scale-down: publish the surviving count, then synchronize_net(),
   then destroy. num_rx_queues uses smp_store_release / smp_load_acquire.
* packets/bytes/drops through netdev_stat_ops, not private -S strings.
   Thirteen existing -S keys kept. No hcall_* or pool%d_ keys.
   replenish_* are per-queue u64; no atomics. get_base_stats() is the
   retired-queue remainder.
* Reset worker gated on NETREG_REGISTERED (cannot reopen after
   unregister).
* get_channels() keeps the live rx_count; mq_fallback caps max_rx so
   a TX-only ethtool -L is not a silent RX shrink.

Changes in v5:

* Restack mailed v4 (14 patches) to v5 (15):

     v4 1-8  helpers                -> v5 1-8
     (new)   SQ poll harden         -> v5 9   (before MQ enable)
     v4 9    MQ enable              -> v5 10
     v4 10   stats                  -> v5 11
     (new)   get_channels           -> v5 12  (peeled from stats)
     v4 11   debugfs                -> v5 13
     v4 12   resize                 -> v5 14
     v4 13   set_channels           -> v5 15
     v4 14   trailing poll/shutdown -> folded into v5 5/9/10/14
             (mailed "P14" was that trailer, not v5 14)
* Teardown-first resize after aggressive ethtool -L; thin defensive
   poll skip remains; no correlator generation field this series
* opened / rx_irq_setup; set_channels keys on opened (not IFF_UP)
* filter_list_dma=0 on map error; restore default-active 64 KiB pool;
   unwind pools by allocation presence; probe_cleanup clears vio
   drvdata; remove: unregister then cancel_work
* TX quiesce before freeing bounce buffers; guard start_xmit if LTB gone
* MQ H_FUNCTION recovery (reset + SQ fallback); no printk under
   replenish_lock; lock harvest with replenish; resume kicks all queues
* Per-queue update_rx_no_buffer; publish-before-free on resize;
   CMO refresh; IRQ helpers return errno
* Harvest abort (no fake GRO / UAF); poll refuses PHYP re-arm on close;
   wrap-safe skb_put; atomic set_channels; monotonic stats across shrink
* Keep mask -> sync -> napi_disable on teardown; open stays
   request_irq -> napi_enable while PHYP masked; scale-up/recovery keep
   napi_enable before enable_irq
* Pool geometry kept on free; restart_rx_queue after open/scale-down;
   remask after napi_disable; schedule_rx_queue masks only when
   napi_schedule_prep succeeds (STOP + poll no-rearm for storms)

Changes in v4:

Addresses Simon's v3 review and related fixes:
* First-use helpers/includes (irqdomain.h with first dispose); no
   unused statics; dropped orphan open/close pipeline patch
* Open/close unwind (free LAN before RX pools); no double TX teardown
* MQ open: replenish all queues before PHYP unmask; H_FUNCTION on
   subordinate register is a hard open failure
* Resize/set_channels hardenings; stats probe-lifetime + sum-on-read;
   buffer_pools diagnostic on debugfs
* Patch 9: put already-created pool kobjects on probe failure paths
* Patch 14: correlator skip, skb tailroom check, napi_complete_done
   shutdown return < budget
* Bisect-friendly restack (helpers with first use)

Changes in v3:

* Dropped RFC; addressed style / DMA feedback from earlier revisions
* Early MQ enablement iterations (see lore links below)

Comments welcome.

---
v6 lore:
   https://lore.kernel.org/r/cover.1788102125.git.mmc@linux.ibm.com (local)
Sashiko NIPA (v6):
   https://netdev-ai.bots.linux.dev/sashiko/#/patchset/cover.1788102125.git.mmc@linux.ibm.com
v5 lore:
   https://lore.kernel.org/r/20260814073642.24630-1-mmc@linux.ibm.com (local)
v5 review (Jakub Kicinski):
   https://lore.kernel.org/r/20260818014710.3853684-1-kuba@kernel.org (local)
Sashiko Gemini (sashiko.dev):
   https://sashiko.dev/#/patchset/20260814073642.24630-1-mmc@linux.ibm.com
Sashiko NIPA (v5):
   https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260814073642.24630-1-mmc@linux.ibm.com

Previous versions
v6: https://lore.kernel.org/r/cover.1788102125.git.mmc@linux.ibm.com (local)
v5: https://lore.kernel.org/r/20260814073642.24630-1-mmc@linux.ibm.com (local)
v4: https://lore.kernel.org/r/cover.1785457143.git.mmc@linux.ibm.com (local)
v3: https://lore.kernel.org/r/20260706193603.8039-1-mmc@linux.ibm.com (local)
v2: https://lore.kernel.org/r/20260701222327.61325-1-mmc@linux.ibm.com (local)
v1: https://lore.kernel.org/r/cover.1782758799.git.mmc@linux.ibm.com (local)
v4 review (Jakub Kicinski):
   https://lore.kernel.org/r/20260806183614.3171785-1-kuba@kernel.org (local)
v3 review (Simon Horman):
   https://lore.kernel.org/r/20260714124327.GJ1364329@horms.kernel.org (local)

Mingming Cao (15):
   ibmveth: Add MQ RX hypercall wrappers and call definitions
   ibmveth: Prepare MQ RX adapter data structures
   ibmveth: Refactor RX resource allocation for MQ RX bring-up
   ibmveth: Refactor buffer pool management for per-queue MQ RX
   ibmveth: Refactor RX interrupt control for MQ RX queues
   ibmveth: Refactor TX resource allocation in open/close paths
   ibmveth: Add RX queue register helpers for MQ
   ibmveth: Add queue-aware RX buffer submit helper for MQ
   ibmveth: Harden RX poll path with helpers
   ibmveth: Enable multi-queue RX receive path
   ibmveth: Add per-queue RX and TX statistics collection
   ibmveth: Report MQ-aware RX counts in ethtool get_channels
   ibmveth: Expose per-queue buffer pool details via debugfs
   ibmveth: Implement incremental MQ RX queue resize
   ibmveth: Complete set_channels down-path and mq_fallback max_rx cap

  arch/powerpc/include/asm/hvcall.h           |    6 +-
  drivers/net/ethernet/ibm/ibmveth.c          | 4441 +++++++++++++++----
  drivers/net/ethernet/ibm/ibmveth.h          |  232 +-
  tools/perf/scripts/python/powerpc-hcalls.py |    4 +
  4 files changed, 3921 insertions(+), 762 deletions(-)


base-commit: 161ea2d4f2a7e784f14b5b0548fcef3e05fc34f8
  
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help