Thread (16 messages) flat view 16 messages, 1 author, 1d ago
WARM1d REVIEWED: 2 (0M)

Revision v6 of 6 in this series; 2 review trailers.

Revisions (6)
  1. v1 [diff vs current]
  2. v2 [diff vs current]
  3. v3 [diff vs current]
  4. v4 [diff vs current]
  5. v5 [diff vs current]
  6. v6 current

[PATCH net-next v6 11/15] ibmveth: Add per-queue RX and TX statistics collection

From: Mingming Cao <hidden>
Date: 2026-08-31 15:09:32
Also in: linuxppc-dev
Subsystem: ibm power virtual ethernet device driver, linux for powerpc (32-bit and 64-bit), networking drivers, the rest · Maintainers: Nick Child, Madhavan Srinivasan, Andrew Lunn, "David S. Miller", Eric Dumazet, Jakub Kicinski, Paolo Abeni, Linus Torvalds

MQ RX points several queues at the same adapter-wide counters, which
races the updates and leaves no way to attribute a count to a queue.

Move every counter that has more than one writer into per-queue
structs allocated at probe and freed at remove:

  struct ibmveth_rx_queue_stats
  struct ibmveth_tx_queue_stats

Each slot has a single writer — replenish_* under that queue's
replenish_lock, the other RX fields from that queue's NAPI, TX under
the stack's per-queue TX lock — so plain u64 is enough on this
PPC64-only driver. No atomic and no u64_stats_sync.

packets, bytes and drops go through struct netdev_stat_ops
(.get_queue_stats_rx, .get_queue_stats_tx, .get_base_stats).
.ndo_get_stats64() sums the per-queue packet and byte counters into
64-bit device totals. ethtool -S keeps only the driver-specific keys
that have no standard equivalent: interrupts, polls, large_packets,
invalid_buffers and no_buffer_drops per RX queue; large_packets,
send_failures and checksum_offload per TX queue. ETH_SS_STATS becomes
variable-length because that block scales with the live queue count.
Hypercall counters and pool%d_ keys are not added: which hcall a
batch picks is not ABI, and size/active already have sysfs (available
for every queue is patch 13).

The thirteen existing ethtool -S keys keep their exact names, their
order and their adapter-wide values, summed from the per-queue slots
on read. The storage moved; that ABI did not. The four replenish_*
counters get per-queue storage but no per-queue key of their own.

Holding that ABI while the storage moves needs the ethtool -S table
to record where each key lives. IBMVETH_STAT_OFF() could only express
an offset into struct ibmveth_adapter. Tag every entry with an enum
ibmveth_stat_src naming the struct it indexes: adapter-wide keys are
read directly, per-queue keys are summed across the slots by one pair
of offset-keyed helpers. That is what lets the field names change
while the key names do not (rx_invalid_buffer now reads
invalid_buffers, tx_send_failed reads send_failures, and the two
large_packets fields live in different structs). tx_map_failed still
reads from the adapter — it has no writer, here or in mainline —
and the three fw_enabled_* keys are capability flags, not counters.

The per-queue keys come from their own tables with the counts derived
by ARRAY_SIZE(), so get_strings(), get_ethtool_stats() and
get_sset_count() cannot drift apart.

ndo_get_stats64() walks every allocated slot rather than only the live
queues, so device totals cannot go backwards when ethtool -L shrinks
the queue count. get_base_stats() therefore reports the retired-queue
remainder rather than zero; the core sums it with the live queues it
iterates itself. Zeroing would assert that the live-queue sum is
already complete. Every field the per-queue callbacks fill is also
initialised there, because netdev_nl_stats_add() drops a field from
the device total unless both sides set it.

Give every queue a no_buffer_retired carry. PHYP's drop counter is
absolute for the buffer-list page currently mapped, so a reopen or a
queue reuse restarts it near zero. Storing only the newest absolute in
adapter->rx_no_buffer meant whichever queue ran last won, and the
value could go backwards. The carry sits beside the no_buffer_drops it
accumulates from, so both belong to one queue. The rx%d_no_buffer_drops
key reports that live page absolute on its own, so it is the one
exported value that is not monotonic; the adapter-wide rx_no_buffer
sums the two and per-queue rx-hw-drops includes both.

Freeing these arrays in remove() forces the teardown order to be
fixed first. Mainline cancels reset work before unregister_netdev(),
but the RX path stays live until unregister and can re-arm it, so the
worker could run after the cancel and reach memory this patch now
frees. unregister_netdev() therefore moves ahead of cancel_work_sync(),
and ibmveth_reset() returns early unless reg_state is NETREG_REGISTERED.
That reorder is a use-after-free fix in its own right; it is carried
here because this patch depends on it. No Fixes: tag — a stable
backport of a feature patch this size is the wrong vehicle; if the
fix is wanted on its own it should be lifted and tagged separately.

ibmveth_probe_cleanup() also clears the vio drvdata before
free_netdev(). A probe failure never reaches ibmveth_remove(), and
CMO get_desired_dma() reads that pointer on a later rebind.

Readers do not test the arrays for NULL: both exist from before
register_netdev() until after unregister_netdev() and
cancel_work_sync(), and probe fails -ENOMEM if either allocation
does.

Signed-off-by: Mingming Cao <redacted>
Reviewed-by: Dave Marquardt <redacted>
Tested-by: Shaik Abdulla <redacted>
---

Changes in v6:
- wrap the qstats local in replenish (81 cols)
- replenish_* per-queue u64, summed on the existing adapter-wide keys;
  no per-queue replenish key; no atomics
- netdev_stat_ops for packets/bytes/drops, not private -S strings
- enum ibmveth_stat_src so existing -S keys keep names while storage
  moves; ARRAY_SIZE() for the per-queue key counts
- no hcall_* keys (buffer-submit and H_SEND_LOGICAL_LAN)
- drop the three pool%d_ keys (15 ethtool entries)
- drop the fifteen qstats NULL checks outside the allocators
  (five in the RX hot path)
- per-queue no_buffer_retired carry
- get_base_stats() reports the retired-queue remainder
- gate reset on NETREG_REGISTERED
- noted: harvest no_buffer on -L shrink is patch 14

Changes in v5:
- Series renumber: mailed v4 10/14 stats -> tip P11 (P09 peel;
  get_channels -> P12)
- rx_no_buffer_retired + sum MAX_* slots so adapter no-buffer / qstat
  totals stay monotonic across reopen and channel shrink
- probe_cleanup: clear vio drvdata before free_netdev (CMO cannot see a
  freed netdev on rebind)
- remove: unregister_netdev then cancel_work_sync (no UAF reset worker)

Changes in v4:
- Merge v3's separate RX and TX stats commits into one patch.
- Introduce rx_queue_stats / tx_qstats / NUM macros here (first use).
- Allocate/free qstats at probe/remove instead of open/close.
- Report adapter-level ethtool strings by summing per-queue counters on
  read; drop aggregate_* helpers.
- Sum global rx_no_buffer across MQ queues into this statistics patch.
- Cacheline-align per-queue stats; derive field counts with offsetof so
  alignment padding is not counted as a statistic.
- probe_cleanup() cancels reset work, puts pool kobjects via helper from
  the prior patch, and frees qstats on probe failure paths.
- Keep plain u64 qstats like existing ibmveth / ibmvnic (PPC_PSERIES).

 drivers/net/ethernet/ibm/ibmveth.c | 507 +++++++++++++++++++++++++----
 drivers/net/ethernet/ibm/ibmveth.h |  57 +++-
 2 files changed, 492 insertions(+), 72 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index 2e8896ea5af2..f4fddfa56571 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -38,6 +38,7 @@
 #include <asm/firmware.h>
 #include <net/tcp.h>
 #include <net/ip6_checksum.h>
+#include <net/netdev_queues.h>
 
 #include "ibmveth.h"
 
@@ -75,32 +76,101 @@ module_param(old_large_send, bool, 0444);
 MODULE_PARM_DESC(old_large_send,
 	"Use old large send method on firmware that supports the new method");
 
+/**
+ * enum ibmveth_stat_src - where an ethtool -S counter is stored
+ * @IBMVETH_STAT_ADAPTER: plain u64 in struct ibmveth_adapter
+ * @IBMVETH_STAT_RX_QSUM: per-queue u64, summed over rx_qstats[]
+ * @IBMVETH_STAT_TX_QSUM: per-queue u64, summed over tx_qstats[]
+ * @IBMVETH_STAT_RX_NO_BUFFER: rx_qstats[] live-page absolute plus the
+ *	absolutes carried over from pages the queue has already retired
+ *
+ * Counters live per-queue so multi-queue writers never share a field.
+ * The adapter is only ever read from ethtool, so summing there is free.
+ */
+enum ibmveth_stat_src {
+	IBMVETH_STAT_ADAPTER,
+	IBMVETH_STAT_RX_QSUM,
+	IBMVETH_STAT_TX_QSUM,
+	IBMVETH_STAT_RX_NO_BUFFER,
+};
+
 struct ibmveth_stat {
 	char name[ETH_GSTRING_LEN];
-	int offset;
+	enum ibmveth_stat_src src;
+	/* Offset into the struct named by @src. */
+	size_t off;
 };
 
 #define IBMVETH_STAT_OFF(stat) offsetof(struct ibmveth_adapter, stat)
+#define IBMVETH_RXQ_OFF(stat) offsetof(struct ibmveth_rx_queue_stats, stat)
+#define IBMVETH_TXQ_OFF(stat) offsetof(struct ibmveth_tx_queue_stats, stat)
 #define IBMVETH_GET_STAT(a, off) *((u64 *)(((unsigned long)(a)) + off))
 
+#define IBMVETH_ADAPTER_STAT(key, field) \
+	{ key, IBMVETH_STAT_ADAPTER, IBMVETH_STAT_OFF(field) }
+#define IBMVETH_RXQ_STAT(key, field) \
+	{ key, IBMVETH_STAT_RX_QSUM, IBMVETH_RXQ_OFF(field) }
+#define IBMVETH_TXQ_STAT(key, field) \
+	{ key, IBMVETH_STAT_TX_QSUM, IBMVETH_TXQ_OFF(field) }
+
+/*
+ * Key names and their order are ABI. Do not reorder or rename; append
+ * only, and only when the counter is worth a permanent interface.
+ */
 static struct ibmveth_stat ibmveth_stats[] = {
-	{ "replenish_task_cycles", IBMVETH_STAT_OFF(replenish_task_cycles) },
-	{ "replenish_no_mem", IBMVETH_STAT_OFF(replenish_no_mem) },
-	{ "replenish_add_buff_failure",
-			IBMVETH_STAT_OFF(replenish_add_buff_failure) },
-	{ "replenish_add_buff_success",
-			IBMVETH_STAT_OFF(replenish_add_buff_success) },
-	{ "rx_invalid_buffer", IBMVETH_STAT_OFF(rx_invalid_buffer) },
-	{ "rx_no_buffer", IBMVETH_STAT_OFF(rx_no_buffer) },
-	{ "tx_map_failed", IBMVETH_STAT_OFF(tx_map_failed) },
-	{ "tx_send_failed", IBMVETH_STAT_OFF(tx_send_failed) },
-	{ "fw_enabled_ipv4_csum", IBMVETH_STAT_OFF(fw_ipv4_csum_support) },
-	{ "fw_enabled_ipv6_csum", IBMVETH_STAT_OFF(fw_ipv6_csum_support) },
-	{ "tx_large_packets", IBMVETH_STAT_OFF(tx_large_packets) },
-	{ "rx_large_packets", IBMVETH_STAT_OFF(rx_large_packets) },
-	{ "fw_enabled_large_send", IBMVETH_STAT_OFF(fw_large_send_support) }
+	IBMVETH_RXQ_STAT("replenish_task_cycles", replenish_task_cycles),
+	IBMVETH_RXQ_STAT("replenish_no_mem", replenish_no_mem),
+	IBMVETH_RXQ_STAT("replenish_add_buff_failure",
+			 replenish_add_buff_failure),
+	IBMVETH_RXQ_STAT("replenish_add_buff_success",
+			 replenish_add_buff_success),
+	IBMVETH_RXQ_STAT("rx_invalid_buffer", invalid_buffers),
+	{ "rx_no_buffer", IBMVETH_STAT_RX_NO_BUFFER,
+	  IBMVETH_RXQ_OFF(no_buffer_drops) },
+	IBMVETH_ADAPTER_STAT("tx_map_failed", tx_map_failed),
+	IBMVETH_TXQ_STAT("tx_send_failed", send_failures),
+	IBMVETH_ADAPTER_STAT("fw_enabled_ipv4_csum", fw_ipv4_csum_support),
+	IBMVETH_ADAPTER_STAT("fw_enabled_ipv6_csum", fw_ipv6_csum_support),
+	IBMVETH_TXQ_STAT("tx_large_packets", large_packets),
+	IBMVETH_RXQ_STAT("rx_large_packets", large_packets),
+	IBMVETH_ADAPTER_STAT("fw_enabled_large_send", fw_large_send_support),
 };
 
+/**
+ * struct ibmveth_qstat - a per-queue counter exposed through ethtool -S
+ * @fmt: key name, taking the queue index as its only argument
+ * @off: offset into the matching per-queue stats struct
+ *
+ * Driving the strings and the values from one table keeps the two in
+ * step; get_sset_count() derives its length from ARRAY_SIZE() so the
+ * three cannot drift apart.
+ */
+struct ibmveth_qstat {
+	const char *fmt;
+	size_t off;
+};
+
+/*
+ * Only counters with no home in the standard interfaces belong here.
+ * packets, bytes and drops are reported through netdev_stat_ops.
+ */
+static const struct ibmveth_qstat ibmveth_rx_qstat_keys[] = {
+	{ "rx%d_interrupts", IBMVETH_RXQ_OFF(interrupts) },
+	{ "rx%d_polls", IBMVETH_RXQ_OFF(polls) },
+	{ "rx%d_large_packets", IBMVETH_RXQ_OFF(large_packets) },
+	{ "rx%d_invalid_buffers", IBMVETH_RXQ_OFF(invalid_buffers) },
+	{ "rx%d_no_buffer_drops", IBMVETH_RXQ_OFF(no_buffer_drops) },
+};
+
+static const struct ibmveth_qstat ibmveth_tx_qstat_keys[] = {
+	{ "tx%d_large_packets", IBMVETH_TXQ_OFF(large_packets) },
+	{ "tx%d_send_failures", IBMVETH_TXQ_OFF(send_failures) },
+	{ "tx%d_checksum_offload", IBMVETH_TXQ_OFF(checksum_offload) },
+};
+
+#define IBMVETH_NUM_RX_QSTATS ARRAY_SIZE(ibmveth_rx_qstat_keys)
+#define IBMVETH_NUM_TX_QSTATS ARRAY_SIZE(ibmveth_tx_qstat_keys)
+
 /* simple methods of getting data from the current rxq entry */
 static u32 ibmveth_rxq_flags(struct ibmveth_adapter *adapter,
 			     int queue_index)
@@ -241,6 +311,60 @@ ibmveth_free_filter_list(struct ibmveth_adapter *adapter)
 	}
 }
 
+/**
+ * ibmveth_alloc_rx_qstats - Allocate per-queue RX statistics
+ * @adapter: ibmveth adapter structure
+ *
+ * Return: 0 on success, -ENOMEM on failure
+ */
+static int ibmveth_alloc_rx_qstats(struct ibmveth_adapter *adapter)
+{
+	adapter->rx_qstats = kcalloc(IBMVETH_MAX_RX_QUEUES,
+				     sizeof(*adapter->rx_qstats),
+				     GFP_KERNEL);
+	if (!adapter->rx_qstats)
+		return -ENOMEM;
+
+	return 0;
+}
+
+/**
+ * ibmveth_free_rx_qstats - Free per-queue RX statistics
+ * @adapter: ibmveth adapter structure
+ */
+static void ibmveth_free_rx_qstats(struct ibmveth_adapter *adapter)
+{
+	kfree(adapter->rx_qstats);
+	adapter->rx_qstats = NULL;
+}
+
+/**
+ * ibmveth_alloc_tx_qstats - Allocate per-queue TX statistics
+ * @adapter: ibmveth adapter structure
+ *
+ * Return: 0 on success, -ENOMEM on failure
+ */
+static int ibmveth_alloc_tx_qstats(struct ibmveth_adapter *adapter)
+{
+	adapter->tx_qstats = kcalloc(IBMVETH_MAX_QUEUES,
+				     sizeof(*adapter->tx_qstats),
+				     GFP_KERNEL);
+	if (!adapter->tx_qstats)
+		return -ENOMEM;
+
+	return 0;
+}
+
+/**
+ * ibmveth_free_tx_qstats - Free per-queue TX statistics
+ * @adapter: ibmveth adapter structure
+ */
+static void ibmveth_free_tx_qstats(struct ibmveth_adapter *adapter)
+{
+	kfree(adapter->tx_qstats);
+	adapter->tx_qstats = NULL;
+}
+
 /**
  * ibmveth_alloc_rx_queues - Allocate per-queue RX resources
  * @adapter: ibmveth adapter structure
@@ -839,6 +963,8 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
 					 int queue_index,
 					 struct ibmveth_replenish_fail *fail)
 {
+	struct ibmveth_rx_queue_stats *qstats =
+		&adapter->rx_qstats[queue_index];
 	union ibmveth_buf_desc descs[IBMVETH_MAX_RX_PER_HCALL] = {0};
 	u32 remaining = pool->size - atomic_read(&pool->available);
 	u64 correlators[IBMVETH_MAX_RX_PER_HCALL] = {0};
@@ -865,7 +991,7 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
 		for (filled = 0; filled < min(remaining, batch); filled++) {
 			index = pool->free_map[free_index];
 			if (index == IBM_VETH_INVALID_MAP) {
-				adapter->replenish_add_buff_failure++;
+				qstats->replenish_add_buff_failure++;
 				outcome = IBMVETH_REPLENISH_RESET_MAP;
 				break;
 			}
@@ -876,8 +1002,8 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
 				skb = netdev_alloc_skb(adapter->netdev,
 						       pool->buff_size);
 				if (!skb) {
-					adapter->replenish_no_mem++;
-					adapter->replenish_add_buff_failure++;
+					qstats->replenish_no_mem++;
+					qstats->replenish_add_buff_failure++;
 					break;
 				}
 
@@ -892,7 +1018,7 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
 							     DMA_ATTR_NO_WARN);
 				if (dma_mapping_error(dev, dma_addr)) {
 					dev_kfree_skb_any(skb);
-					adapter->replenish_add_buff_failure++;
+					qstats->replenish_add_buff_failure++;
 					break;
 				}
 
@@ -953,7 +1079,7 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
 		}
 
 		buffers_added += filled;
-		adapter->replenish_add_buff_success += filled;
+		qstats->replenish_add_buff_success += filled;
 		remaining -= filled;
 
 		memset(&descs, 0, sizeof(descs));
@@ -976,7 +1102,7 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
 				pool->skbuff[index] = NULL;
 			}
 		}
-		adapter->replenish_add_buff_failure += filled;
+		qstats->replenish_add_buff_failure += filled;
 
 		if (lpar_rc == H_FUNCTION) {
 			if (adapter->multi_queue) {
@@ -1017,6 +1143,7 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
 static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter,
 					int queue_index)
 {
+	struct ibmveth_rx_queue_stats *qstats;
 	__be64 *p;
 	u64 drops;
 
@@ -1028,7 +1155,18 @@ static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter,
 	p = adapter->buffer_list_addr[queue_index] + 4096 - 8;
 	drops = be64_to_cpup(p);
 
-	adapter->rx_no_buffer = drops;
+	/*
+	 * PHYP's buffer-list page counter is absolute for that page. A new
+	 * page (reopen / queue reuse after -L) starts near zero; fold the
+	 * previous absolute into this queue's retired carry so sums stay
+	 * monotonic. Both fields belong to the queue being updated, so this
+	 * stays single-writer under the queue's replenish_lock.
+	 */
+	qstats = &adapter->rx_qstats[queue_index];
+
+	if (drops < qstats->no_buffer_drops)
+		qstats->no_buffer_retired += qstats->no_buffer_drops;
+	qstats->no_buffer_drops = drops;
 }
 
 /* replenish routine */
@@ -1050,10 +1188,10 @@ static void ibmveth_replenish_task(struct ibmveth_adapter *adapter,
 		return;
 	}
 
-	adapter->replenish_task_cycles++;
-
 	spin_lock_irqsave(&rxq->replenish_lock, flags);
 
+	adapter->rx_qstats[queue_index].replenish_task_cycles++;
+
 	for (i = (IBMVETH_NUM_BUFF_POOLS - 1); i >= 0; i--) {
 		struct ibmveth_buff_pool *pool =
 			&adapter->rx_buff_pool[queue_index][i];
@@ -2038,6 +2176,10 @@ static void ibmveth_reset(struct work_struct *w)
 	netdev_dbg(netdev, "reset starting\n");
 
 	rtnl_lock();
+	if (netdev->reg_state != NETREG_REGISTERED) {
+		rtnl_unlock();
+		return;
+	}
 
 	dev_close(adapter->netdev);
 	dev_open(adapter->netdev, NULL);
@@ -2271,22 +2413,96 @@ static int ibmveth_set_features(struct net_device *dev,
 	return rc1 ? rc1 : rc2;
 }
 
-static void ibmveth_get_strings(struct net_device *dev, u32 stringset, u8 *data)
+/*
+ * Sum per-queue counters for rare ethtool reads. The hot paths only ever
+ * touch their own queue's slot, so nothing here needs an atomic; the cost
+ * of aggregation is paid by the reader instead (ibmvnic-style).
+ *
+ * Every slot is summed, not just the live ones, so that shrinking the
+ * queue count with ethtool -L cannot make a counter go backwards.
+ */
+static u64 ibmveth_sum_rx_qstat(struct ibmveth_adapter *adapter, size_t off)
+{
+	u64 total = 0;
+	int i;
+
+	for (i = 0; i < IBMVETH_MAX_RX_QUEUES; i++)
+		total += *(u64 *)((u8 *)&adapter->rx_qstats[i] + off);
+
+	return total;
+}
+
+static u64 ibmveth_sum_tx_qstat(struct ibmveth_adapter *adapter, size_t off)
 {
+	u64 total = 0;
 	int i;
 
+	for (i = 0; i < IBMVETH_MAX_QUEUES; i++)
+		total += *(u64 *)((u8 *)&adapter->tx_qstats[i] + off);
+
+	return total;
+}
+
+static u64 ibmveth_ethtool_adapter_stat(struct ibmveth_adapter *adapter,
+					int index)
+{
+	const struct ibmveth_stat *stat = &ibmveth_stats[index];
+
+	switch (stat->src) {
+	case IBMVETH_STAT_RX_QSUM:
+		return ibmveth_sum_rx_qstat(adapter, stat->off);
+	case IBMVETH_STAT_TX_QSUM:
+		return ibmveth_sum_tx_qstat(adapter, stat->off);
+	case IBMVETH_STAT_RX_NO_BUFFER:
+		/*
+		 * PHYP's page counter is absolute for the page currently
+		 * mapped, so a reopen or queue reuse restarts it near zero.
+		 * ibmveth_update_rx_no_buffer() folds each decrease into the
+		 * queue's retired carry; add both back to stay monotonic.
+		 */
+		return ibmveth_sum_rx_qstat(adapter, stat->off) +
+		       ibmveth_sum_rx_qstat(adapter,
+					    IBMVETH_RXQ_OFF(no_buffer_retired));
+	case IBMVETH_STAT_ADAPTER:
+		break;
+	}
+
+	return IBMVETH_GET_STAT(adapter, stat->off);
+}
+
+static void ibmveth_get_strings(struct net_device *dev, u32 stringset, u8 *data)
+{
+	struct ibmveth_adapter *adapter = netdev_priv(dev);
+	u8 *p = data;
+	int i, j;
+
 	if (stringset != ETH_SS_STATS)
 		return;
 
-	for (i = 0; i < ARRAY_SIZE(ibmveth_stats); i++, data += ETH_GSTRING_LEN)
-		memcpy(data, ibmveth_stats[i].name, ETH_GSTRING_LEN);
+	for (i = 0; i < ARRAY_SIZE(ibmveth_stats); i++) {
+		memcpy(p, ibmveth_stats[i].name, ETH_GSTRING_LEN);
+		p += ETH_GSTRING_LEN;
+	}
+
+	for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++)
+		for (j = 0; j < IBMVETH_NUM_RX_QSTATS; j++)
+			ethtool_sprintf(&p, ibmveth_rx_qstat_keys[j].fmt, i);
+
+	for (i = 0; i < dev->real_num_tx_queues; i++)
+		for (j = 0; j < IBMVETH_NUM_TX_QSTATS; j++)
+			ethtool_sprintf(&p, ibmveth_tx_qstat_keys[j].fmt, i);
 }
 
 static int ibmveth_get_sset_count(struct net_device *dev, int sset)
 {
+	struct ibmveth_adapter *adapter = netdev_priv(dev);
+
 	switch (sset) {
 	case ETH_SS_STATS:
-		return ARRAY_SIZE(ibmveth_stats);
+		return ARRAY_SIZE(ibmveth_stats) +
+		       ibmveth_get_num_rx_queues(adapter) *
+		       IBMVETH_NUM_RX_QSTATS +
+		       dev->real_num_tx_queues * IBMVETH_NUM_TX_QSTATS;
 	default:
 		return -EOPNOTSUPP;
 	}
@@ -2295,11 +2511,27 @@ static int ibmveth_get_sset_count(struct net_device *dev, int sset)
 static void ibmveth_get_ethtool_stats(struct net_device *dev,
 				      struct ethtool_stats *stats, u64 *data)
 {
-	int i;
 	struct ibmveth_adapter *adapter = netdev_priv(dev);
+	int i, j, k;
 
 	for (i = 0; i < ARRAY_SIZE(ibmveth_stats); i++)
-		data[i] = IBMVETH_GET_STAT(adapter, ibmveth_stats[i].offset);
+		data[i] = ibmveth_ethtool_adapter_stat(adapter, i);
+
+	for (j = 0; j < ibmveth_get_num_rx_queues(adapter); j++) {
+		const u8 *q = (const u8 *)&adapter->rx_qstats[j];
+
+		for (k = 0; k < IBMVETH_NUM_RX_QSTATS; k++)
+			data[i++] = *(const u64 *)
+				(q + ibmveth_rx_qstat_keys[k].off);
+	}
+
+	for (j = 0; j < dev->real_num_tx_queues; j++) {
+		const u8 *q = (const u8 *)&adapter->tx_qstats[j];
+
+		for (k = 0; k < IBMVETH_NUM_TX_QSTATS; k++)
+			data[i++] = *(const u64 *)
+				(q + ibmveth_tx_qstat_keys[k].off);
+	}
 }
 
 static void ibmveth_get_channels(struct net_device *netdev,
@@ -2411,8 +2643,10 @@ static int ibmveth_send(struct ibmveth_adapter *adapter,
 }
 
 static int ibmveth_is_packet_unsupported(struct sk_buff *skb,
-					 struct net_device *netdev)
+					 struct ibmveth_adapter *adapter,
+					 int queue_num)
 {
+	struct net_device *netdev = adapter->netdev;
 	struct ethhdr *ether_header;
 	int ret = 0;
 
@@ -2420,7 +2654,7 @@ static int ibmveth_is_packet_unsupported(struct sk_buff *skb,
 
 	if (ether_addr_equal(ether_header->h_dest, netdev->dev_addr)) {
 		netdev_dbg(netdev, "veth doesn't support loopback packets, dropping packet.\n");
-		netdev->stats.tx_dropped++;
+		adapter->tx_qstats[queue_num].dropped_packets++;
 		ret = -EOPNOTSUPP;
 	}
 
@@ -2438,11 +2672,11 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
 
 	/* Close / failed reopen can free LTBs while IFF_UP is still set. */
 	if (unlikely(!adapter->tx_ltb_ptr[queue_num])) {
-		netdev->stats.tx_dropped++;
+		adapter->tx_qstats[queue_num].dropped_packets++;
 		goto out;
 	}
 
-	if (ibmveth_is_packet_unsupported(skb, netdev))
+	if (ibmveth_is_packet_unsupported(skb, adapter, queue_num))
 		goto out;
 	/* veth can't checksum offload UDP */
 	if (skb->ip_summed == CHECKSUM_PARTIAL &&
@@ -2453,7 +2687,7 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
 	    skb_checksum_help(skb)) {
 
 		netdev_err(netdev, "tx: failed to checksum packet\n");
-		netdev->stats.tx_dropped++;
+		adapter->tx_qstats[queue_num].dropped_packets++;
 		goto out;
 	}
 
@@ -2465,6 +2699,8 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
 
 		desc_flags |= (IBMVETH_BUF_NO_CSUM | IBMVETH_BUF_CSUM_GOOD);
 
+		adapter->tx_qstats[queue_num].checksum_offload++;
+
 		/* Need to zero out the checksum */
 		buf[0] = 0;
 		buf[1] = 0;
@@ -2476,7 +2712,7 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
 	if (skb->ip_summed == CHECKSUM_PARTIAL && skb_is_gso(skb)) {
 		if (adapter->fw_large_send_support) {
 			mss = (unsigned long)skb_shinfo(skb)->gso_size;
-			adapter->tx_large_packets++;
+			adapter->tx_qstats[queue_num].large_packets++;
 		} else if (!skb_is_gso_v6(skb)) {
 			/* Put -1 in the IP checksum to tell phyp it
 			 * is a largesend packet. Put the mss in
@@ -2485,7 +2721,7 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
 			ip_hdr(skb)->check = 0xffff;
 			tcp_hdr(skb)->check =
 				cpu_to_be16(skb_shinfo(skb)->gso_size);
-			adapter->tx_large_packets++;
+			adapter->tx_qstats[queue_num].large_packets++;
 		}
 	}
 
@@ -2493,7 +2729,7 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
 	if (unlikely(skb->len > adapter->tx_ltb_size)) {
 		netdev_err(adapter->netdev, "tx: packet size (%u) exceeds ltb (%u)\n",
 			   skb->len, adapter->tx_ltb_size);
-		netdev->stats.tx_dropped++;
+		adapter->tx_qstats[queue_num].dropped_packets++;
 		goto out;
 	}
 	memcpy(adapter->tx_ltb_ptr[queue_num], skb->data, skb_headlen(skb));
@@ -2510,7 +2746,7 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
 	if (unlikely(total_bytes != skb->len)) {
 		netdev_err(adapter->netdev, "tx: incorrect packet len copied into ltb (%u != %u)\n",
 			   skb->len, total_bytes);
-		netdev->stats.tx_dropped++;
+		adapter->tx_qstats[queue_num].dropped_packets++;
 		goto out;
 	}
 	desc.fields.flags_len = desc_flags | skb->len;
@@ -2519,11 +2755,11 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
 	dma_wmb();
 
 	if (ibmveth_send(adapter, desc.desc, mss)) {
-		adapter->tx_send_failed++;
-		netdev->stats.tx_dropped++;
+		adapter->tx_qstats[queue_num].send_failures++;
+		adapter->tx_qstats[queue_num].dropped_packets++;
 	} else {
-		netdev->stats.tx_packets++;
-		netdev->stats.tx_bytes += skb->len;
+		adapter->tx_qstats[queue_num].packets++;
+		adapter->tx_qstats[queue_num].bytes += skb->len;
 	}
 
 out:
@@ -2652,7 +2888,7 @@ static void ibmveth_rx_csum_helper(struct sk_buff *skb,
 static void ibmveth_poll_bump_invalid(struct ibmveth_adapter *adapter,
 				      int queue_index)
 {
-	adapter->rx_invalid_buffer++;
+	adapter->rx_qstats[queue_index].invalid_buffers++;
 }
 
 static bool ibmveth_poll_stopping(struct net_device *netdev,
@@ -2788,7 +3024,7 @@ static int ibmveth_poll_deliver_frame(struct napi_struct *napi,
 	if ((length > netdev->mtu + ETH_HLEN) || lrg_pkt ||
 	    iph_check == 0xffff) {
 		ibmveth_rx_mss_helper(skb, mss, lrg_pkt);
-		adapter->rx_large_packets++;
+		adapter->rx_qstats[queue_index].large_packets++;
 	}
 
 	if (csum_good) {
@@ -2799,8 +3035,8 @@ static int ibmveth_poll_deliver_frame(struct napi_struct *napi,
 	skb_record_rx_queue(skb, queue_index);
 	napi_gro_receive(napi, skb);
 
-	netdev->stats.rx_packets++;
-	netdev->stats.rx_bytes += length;
+	adapter->rx_qstats[queue_index].packets++;
+	adapter->rx_qstats[queue_index].bytes += length;
 
 	return 1;
 }
@@ -2827,6 +3063,8 @@ static int ibmveth_poll(struct napi_struct *napi, int budget)
 		return 0;
 	}
 
+	adapter->rx_qstats[queue_index].polls++;
+
 restart_poll:
 	while (frames_processed < budget) {
 		if (ibmveth_poll_stopping(netdev, napi))
@@ -2915,6 +3153,8 @@ static irqreturn_t ibmveth_interrupt(int irq, void *dev_instance)
 	if (qindex < 0 || qindex >= ibmveth_get_num_rx_queues(adapter))
 		return IRQ_NONE;
 
+	adapter->rx_qstats[qindex].interrupts++;
+
 	ibmveth_schedule_rx_queue(adapter, qindex);
 	return IRQ_HANDLED;
 }
@@ -3132,6 +3372,124 @@ static netdev_features_t ibmveth_features_check(struct sk_buff *skb,
 	return vlan_features_check(skb, features);
 }
 
+/**
+ * ibmveth_get_stats64 - Return aggregated per-queue statistics
+ * @dev: network device
+ * @stats: rtnl link statistics storage
+ *
+ * Sums per-queue rx_qstats and tx_qstats into the rtnl counters.
+ * Walk the full allocated arrays (not the live queue count) so shrinking
+ * channels cannot make the totals go backwards.
+ * Callers use ndo_get_stats64(); avoid updating netdev->stats on the
+ * xmit/poll paths to keep per-queue counters off the hot cache line.
+ */
+static void ibmveth_get_stats64(struct net_device *dev,
+				struct rtnl_link_stats64 *stats)
+{
+	struct ibmveth_adapter *adapter = netdev_priv(dev);
+	int i;
+
+	for (i = 0; i < IBMVETH_MAX_RX_QUEUES; i++) {
+		stats->rx_packets += adapter->rx_qstats[i].packets;
+		stats->rx_bytes += adapter->rx_qstats[i].bytes;
+	}
+
+	for (i = 0; i < IBMVETH_MAX_QUEUES; i++) {
+		stats->tx_packets += adapter->tx_qstats[i].packets;
+		stats->tx_bytes += adapter->tx_qstats[i].bytes;
+		stats->tx_dropped += adapter->tx_qstats[i].dropped_packets;
+	}
+}
+
+static void ibmveth_get_queue_stats_rx(struct net_device *dev, int idx,
+				       struct netdev_queue_stats_rx *stats)
+{
+	struct ibmveth_adapter *adapter = netdev_priv(dev);
+
+	stats->packets = adapter->rx_qstats[idx].packets;
+	stats->bytes = adapter->rx_qstats[idx].bytes;
+	/*
+	 * All three are frames that entered the device and never left it,
+	 * which is what rx-hw-drops is specified to cover: no_buffer_drops
+	 * is PHYP dropping for lack of buffer space on the page mapped now,
+	 * no_buffer_retired the same for pages this queue has already
+	 * released, and invalid_buffers is a processing error.
+	 */
+	stats->hw_drops = adapter->rx_qstats[idx].no_buffer_drops +
+			  adapter->rx_qstats[idx].no_buffer_retired +
+			  adapter->rx_qstats[idx].invalid_buffers;
+	stats->alloc_fail = adapter->rx_qstats[idx].replenish_no_mem;
+}
+
+static void ibmveth_get_queue_stats_tx(struct net_device *dev, int idx,
+				       struct netdev_queue_stats_tx *stats)
+{
+	struct ibmveth_adapter *adapter = netdev_priv(dev);
+
+	stats->packets = adapter->tx_qstats[idx].packets;
+	stats->bytes = adapter->tx_qstats[idx].bytes;
+	stats->hw_drops = adapter->tx_qstats[idx].dropped_packets;
+}
+
+/**
+ * ibmveth_get_base_stats - account for traffic not on a live queue
+ * @dev: network device
+ * @rx: RX base statistics storage
+ * @tx: TX base statistics storage
+ *
+ * get_queue_stats_{rx,tx}() only report queues the core still iterates,
+ * i.e. below real_num_{rx,tx}_queues, while ibmveth_get_stats64() walks
+ * the full arrays so device totals stay monotonic across a shrink.
+ * Report the retired-queue remainder here, otherwise qstats and
+ * rtnl_link_stats64 disagree by a delta that grows with every shrink.
+ * Zeroing would not be neutral: per netdev_stat_ops it asserts the
+ * per-queue sum is already exact.
+ *
+ * Bound the live side with real_num_*_queues rather than the adapter's
+ * own count, so the split lines up with the core's iteration exactly.
+ *
+ * Every field the per-queue callbacks fill must also be initialised
+ * here: netdev_nl_stats_add() starts the sum at NETDEV_STAT_NOT_SET and
+ * only accumulates while both sides are set, so a field left unset here
+ * is dropped from the device total even though the queues report it.
+ */
+static void ibmveth_get_base_stats(struct net_device *dev,
+				   struct netdev_queue_stats_rx *rx,
+				   struct netdev_queue_stats_tx *tx)
+{
+	struct ibmveth_adapter *adapter = netdev_priv(dev);
+	unsigned int i;
+
+	rx->packets = 0;
+	rx->bytes = 0;
+	rx->alloc_fail = 0;
+	rx->hw_drops = 0;
+	tx->packets = 0;
+	tx->bytes = 0;
+	tx->hw_drops = 0;
+
+	for (i = dev->real_num_rx_queues; i < IBMVETH_MAX_RX_QUEUES; i++) {
+		rx->packets += adapter->rx_qstats[i].packets;
+		rx->bytes += adapter->rx_qstats[i].bytes;
+		rx->hw_drops += adapter->rx_qstats[i].no_buffer_drops +
+				adapter->rx_qstats[i].no_buffer_retired +
+				adapter->rx_qstats[i].invalid_buffers;
+		rx->alloc_fail += adapter->rx_qstats[i].replenish_no_mem;
+	}
+
+	for (i = dev->real_num_tx_queues; i < IBMVETH_MAX_QUEUES; i++) {
+		tx->packets += adapter->tx_qstats[i].packets;
+		tx->bytes += adapter->tx_qstats[i].bytes;
+		tx->hw_drops += adapter->tx_qstats[i].dropped_packets;
+	}
+}
+
+static const struct netdev_stat_ops ibmveth_stat_ops = {
+	.get_queue_stats_rx	= ibmveth_get_queue_stats_rx,
+	.get_queue_stats_tx	= ibmveth_get_queue_stats_tx,
+	.get_base_stats		= ibmveth_get_base_stats,
+};
+
 static const struct net_device_ops ibmveth_netdev_ops = {
 	.ndo_open		= ibmveth_open,
 	.ndo_stop		= ibmveth_close,
@@ -3144,6 +3502,7 @@ static const struct net_device_ops ibmveth_netdev_ops = {
 	.ndo_validate_addr	= eth_validate_addr,
 	.ndo_set_mac_address    = ibmveth_set_mac_addr,
 	.ndo_features_check	= ibmveth_features_check,
+	.ndo_get_stats64	= ibmveth_get_stats64,
 #ifdef CONFIG_NET_POLL_CONTROLLER
 	.ndo_poll_controller	= ibmveth_poll_controller,
 #endif
@@ -3158,6 +3517,23 @@ static void ibmveth_put_pool_kobjs(struct ibmveth_adapter *adapter,
 		kobject_put(&adapter->rx_buff_pool[0][i].kobj);
 }
 
+static void ibmveth_probe_cleanup(struct ibmveth_adapter *adapter,
+				  int pools_ready)
+{
+	struct net_device *netdev = adapter->netdev;
+
+	cancel_work_sync(&adapter->work);
+	ibmveth_put_pool_kobjs(adapter, pools_ready);
+
+	ibmveth_free_tx_qstats(adapter);
+	ibmveth_free_rx_qstats(adapter);
+	/* Probe failure never reaches ibmveth_remove(); clear before free so
+	 * CMO get_desired_dma() cannot see a freed netdev on rebind.
+	 */
+	dev_set_drvdata(&adapter->vdev->dev, NULL);
+	free_netdev(netdev);
+}
+
 static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
 {
 	int rc, i, mac_len, pools_ready = 0;
@@ -3223,9 +3599,16 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
 		netif_napi_add_weight(netdev, &adapter->napi[i],
 				      ibmveth_poll, 16);
 
+	if (ibmveth_alloc_rx_qstats(adapter) ||
+	    ibmveth_alloc_tx_qstats(adapter)) {
+		ibmveth_probe_cleanup(adapter, 0);
+		return -ENOMEM;
+	}
+
 	netdev->irq = dev->irq;
 	netdev->netdev_ops = &ibmveth_netdev_ops;
 	netdev->ethtool_ops = &netdev_ethtool_ops;
+	netdev->stat_ops = &ibmveth_stat_ops;
 	SET_NETDEV_DEV(netdev, &dev->dev);
 	netdev->hw_features = NETIF_F_SG;
 	if (vio_get_attribute(dev, "ibm,illan-options", NULL) != NULL) {
@@ -3305,9 +3688,7 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
 				"failed to create pool%d kobject: %d\n", i, rc);
 			/* init_and_add takes a ref even on failure */
 			kobject_put(kobj);
-			ibmveth_put_pool_kobjs(adapter, pools_ready);
-			dev_set_drvdata(&dev->dev, NULL);
-			free_netdev(netdev);
+			ibmveth_probe_cleanup(adapter, pools_ready);
 			return rc;
 		}
 
@@ -3327,9 +3708,7 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
 	if (rc) {
 		netdev_dbg(netdev, "failed to set number of tx queues rc=%d\n",
 			   rc);
-		ibmveth_put_pool_kobjs(adapter, pools_ready);
-		dev_set_drvdata(&dev->dev, NULL);
-		free_netdev(netdev);
+		ibmveth_probe_cleanup(adapter, pools_ready);
 		return rc;
 	}
 
@@ -3344,9 +3723,7 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
 	if (rc) {
 		netdev_dbg(netdev, "failed to set number of rx queues rc=%d\n",
 			   rc);
-		ibmveth_put_pool_kobjs(adapter, pools_ready);
-		dev_set_drvdata(&dev->dev, NULL);
-		free_netdev(netdev);
+		ibmveth_probe_cleanup(adapter, pools_ready);
 		return rc;
 	}
 
@@ -3363,9 +3740,7 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
 
 	if (rc) {
 		netdev_dbg(netdev, "failed to register netdev rc=%d\n", rc);
-		ibmveth_put_pool_kobjs(adapter, pools_ready);
-		dev_set_drvdata(&dev->dev, NULL);
-		free_netdev(netdev);
+		ibmveth_probe_cleanup(adapter, pools_ready);
 		return rc;
 	}
 
@@ -3380,12 +3755,20 @@ static void ibmveth_remove(struct vio_dev *dev)
 	struct ibmveth_adapter *adapter = netdev_priv(netdev);
 	int i;
 
-	cancel_work_sync(&adapter->work);
-
 	for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
 		kobject_put(&adapter->rx_buff_pool[0][i].kobj);
 
+	/*
+	 * Unregister first so NAPI/xmit cannot re-arm reset work after we
+	 * cancel it. cancel_work_sync() before unregister left a window
+	 * where poll could schedule_work() and the worker ran after
+	 * free_netdev().
+	 */
 	unregister_netdev(netdev);
+	cancel_work_sync(&adapter->work);
+
+	ibmveth_free_tx_qstats(adapter);
+	ibmveth_free_rx_qstats(adapter);
 
 	free_netdev(netdev);
 	dev_set_drvdata(&dev->dev, NULL);
diff --git a/drivers/net/ethernet/ibm/ibmveth.h b/drivers/net/ethernet/ibm/ibmveth.h
index cf9e77fc2190..0f2971c8627a 100644
--- a/drivers/net/ethernet/ibm/ibmveth.h
+++ b/drivers/net/ethernet/ibm/ibmveth.h
@@ -275,6 +275,43 @@ static int pool_active[] = { 1, 1, 0, 0, 1};
 
 #define IBM_VETH_INVALID_MAP ((u16)0xffff)
 
+/*
+ * Per-queue RX counters. No field has two concurrent writers:
+ * interrupts is written only from this queue's IRQ handler; polls,
+ * packets, bytes, large_packets and invalid_buffers only from its NAPI
+ * poll; replenish_* only under its replenish_lock; and no_buffer_drops
+ * and no_buffer_retired under that lock or from a teardown path already
+ * quiesced by napi_disable()/synchronize_irq(). Plain u64 is therefore
+ * sufficient and no atomic or u64_stats_sync is needed: the driver is
+ * PPC64-only, so 64-bit loads and stores do not tear.
+ */
+struct ibmveth_rx_queue_stats {
+	u64 packets;
+	u64 bytes;
+	u64 interrupts;
+	u64 polls;
+	u64 large_packets;
+	u64 invalid_buffers;
+	/* PHYP's per-page absolute drop count for the live page. */
+	u64 no_buffer_drops;
+	/* Absolutes from pages this queue has already retired. */
+	u64 no_buffer_retired;
+	u64 replenish_task_cycles;
+	u64 replenish_no_mem;
+	u64 replenish_add_buff_failure;
+	u64 replenish_add_buff_success;
+} ____cacheline_aligned_in_smp;
+
+/* Per-queue TX counters; serialized by the stack's per-queue TX lock. */
+struct ibmveth_tx_queue_stats {
+	u64 packets;
+	u64 bytes;
+	u64 large_packets;
+	u64 dropped_packets;
+	u64 send_failures;
+	u64 checksum_offload;
+} ____cacheline_aligned_in_smp;
+
 struct ibmveth_buff_pool {
     u32 size;
     u32 index;
@@ -333,17 +370,17 @@ struct ibmveth_adapter {
 	u64 fw_ipv6_csum_support;
 	u64 fw_ipv4_csum_support;
 	u64 fw_large_send_support;
-	/* adapter specific stats */
-	u64 replenish_task_cycles;
-	u64 replenish_no_mem;
-	u64 replenish_add_buff_failure;
-	u64 replenish_add_buff_success;
-	u64 rx_invalid_buffer;
-	u64 rx_no_buffer;
+	/*
+	 * Every other ethtool -S counter lives in rx_qstats/tx_qstats and is
+	 * summed on read. tx_map_failed predates multi-queue, has never been
+	 * updated by any code path, and is kept only so the key keeps
+	 * reporting the zero userspace already sees.
+	 */
 	u64 tx_map_failed;
-	u64 tx_send_failed;
-	u64 tx_large_packets;
-	u64 rx_large_packets;
+
+	struct ibmveth_rx_queue_stats *rx_qstats;
+	struct ibmveth_tx_queue_stats *tx_qstats;
+
 	/* Ethtool settings */
 	u8 duplex;
 	u32 speed;
-- 
2.50.1 (Apple Git-155)
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help