From: Alexander Lobakin <hidden> Date: 2021-11-23 16:41:20
This is an almost complete rework of [0].
This series introduces generic XDP statistics infra based on rtnl
xstats (Ethtool standard stats previously), and wires up the drivers
which collect appropriate statistics to this new interface. Finally,
it introduces XDP/XSK statistics to all XDP-capable Intel drivers.
Those counters are:
* packets: number of frames passed to bpf_prog_run_xdp().
* bytes: number of bytes went through bpf_prog_run_xdp().
* errors: number of general XDP errors, if driver has one unified
counter.
* aborted: number of XDP_ABORTED returns.
* drop: number of XDP_DROP returns.
* invalid: number of returns of unallowed values (i.e. not XDP_*).
* pass: number of XDP_PASS returns.
* redirect: number of successfully performed XDP_REDIRECT requests.
* redirect_errors: number of failed XDP_REDIRECT requests.
* tx: number of successfully performed XDP_TX requests.
* tx_errors: number of failed XDP_TX requests.
* xmit_packets: number of successfully transmitted XDP/XSK frames.
* xmit_bytes: number of successfully transmitted XDP/XSK frames.
* xmit_errors: of XDP/XSK frames failed to transmit.
* xmit_full: number of XDP/XSK queue being full at the moment of
transmission.
To provide them, developers need to implement .ndo_get_xdp_stats()
and, if they want to expose stats on a per-channel basis,
.ndo_get_xdp_stats_nch(). include/net/xdp.h contains some helper
structs and functions which might be useful for doing this.
It is up to developers to decide whether to implement XDP stats in
their drivers or not, depending on the needs and so on, but if so,
it is implied that they will be implemented using this new infra
rather than custom Ethtool entries. XDP stats {,type} list can be
expanded if needed as counters are being provided as nested NL attrs.
There's an option to provide XDP and XSK counters separately, and I
used it in mlx5 (has separate RQs and SQs) and Intel (have separate
NAPI poll routines) drivers.
Example output of iproute2's new command:
$ ip link xdpstats dev enp178s0
16: enp178s0:
xdp-channel0-rx_xdp_packets: 0
xdp-channel0-rx_xdp_bytes: 1
xdp-channel0-rx_xdp_errors: 2
xdp-channel0-rx_xdp_aborted: 3
xdp-channel0-rx_xdp_drop: 4
xdp-channel0-rx_xdp_invalid: 5
xdp-channel0-rx_xdp_pass: 6
xdp-channel0-rx_xdp_redirect: 7
xdp-channel0-rx_xdp_redirect_errors: 8
xdp-channel0-rx_xdp_tx: 9
xdp-channel0-rx_xdp_tx_errors: 10
xdp-channel0-tx_xdp_xmit_packets: 11
xdp-channel0-tx_xdp_xmit_bytes: 12
xdp-channel0-tx_xdp_xmit_errors: 13
xdp-channel0-tx_xdp_xmit_full: 14
[ ... ]
This series doesn't touch existing Ethtool-exposed XDP stats due
to the potential breakage of custom{,er} scripts and other stuff
that might be hardwired on their presence. Developers are free to
drop them on their own if possible. In ideal case we would see
Ethtool stats free from anything related to XDP, but it's unlikely
to happen (:
XDP_PASS kpps on an ice NIC with ~50 Gbps line rate:
Frame size 64 | 128 | 256 | 512 | 1024 | 1532
----------------------------------------------------------
net-next 23557 | 23750 | 20731 | 11723 | 6270 | 4377
This series 23484 | 23812 | 20679 | 11720 | 6270 | 4377
The same situation with XDP_DROP and several more common cases:
nothing past stddev (which is a bit surprising, but not complaining
at all).
A brief series breakdown:
* 1-2: introduce new infra and APIs, for rtnetlink and drivers
respectively;
* 3-19: add needed callback to the existing drivers to export
their stats using new infra. Some of the patches are cosmetic
prereqs;
* 20-25: add XDP/XSK stats to all Intel drivers;
* 26: mention generic XDP stats in Documentation.
This set is also available here: [1]
A separate iproute2-next patch will be published in parallel,
for now you can find it here: [2]
From v1 [0]:
- use rtnl xstats instead of Ethtool standard -- XDP stats are
purely software while Ethtool infra was designed for HW stats
(Jakub);
- split xmit into xmit_packets and xmit_bytes, add xmit_full;
- don't touch existing drivers custom XDP stats exposed via
Ethtool (Jakub);
- add a bunch of helper structs and functions to reduce boilerplates
and code duplication in new drivers;
- merge with the series which adds XDP stats to Intel drivers;
- add some perf numbers per Jakub's request.
[0] https://lore.kernel.org/all/20210803163641.3743-1-alexandr.lobakin@intel.com
[1] https://github.com/alobakin/linux/pull/11
[2] https://github.com/alobakin/iproute2/pull/1
Alexander Lobakin (26):
rtnetlink: introduce generic XDP statistics
xdp: provide common driver helpers for implementing XDP stats
ena: implement generic XDP statistics callbacks
dpaa2: implement generic XDP stats callbacks
enetc: implement generic XDP stats callbacks
mvneta: reformat mvneta_netdev_ops
mvneta: add .ndo_get_xdp_stats() callback
mvpp2: provide .ndo_get_xdp_stats() callback
mlx5: don't mix XDP_DROP and Rx XDP error cases
mlx5: provide generic XDP stats callbacks
sf100, sfx: implement generic XDP stats callbacks
veth: don't mix XDP_DROP counter with Rx XDP errors
veth: drop 'xdp_' suffix from packets and bytes stats
veth: reformat veth_netdev_ops
veth: add generic XDP stats callbacks
virtio_net: don't mix XDP_DROP counter with Rx XDP errors
virtio_net: rename xdp_tx{,_drops} SQ stats to xdp_xmit{,_errors}
virtio_net: reformat virtnet_netdev
virtio_net: add callbacks for generic XDP stats
i40e: add XDP and XSK generic per-channel statistics
ice: add XDP and XSK generic per-channel statistics
igb: add XDP generic per-channel statistics
igc: bail out early on XSK xmit if no descs are available
igc: add XDP and XSK generic per-channel statistics
ixgbe: add XDP and XSK generic per-channel statistics
Documentation: reflect generic XDP statistics
Documentation/networking/statistics.rst | 33 +++
drivers/net/ethernet/amazon/ena/ena_netdev.c | 53 ++++
.../net/ethernet/freescale/dpaa2/dpaa2-eth.c | 45 +++
drivers/net/ethernet/freescale/enetc/enetc.c | 48 ++++
drivers/net/ethernet/freescale/enetc/enetc.h | 3 +
.../net/ethernet/freescale/enetc/enetc_pf.c | 2 +
drivers/net/ethernet/intel/i40e/i40e.h | 1 +
drivers/net/ethernet/intel/i40e/i40e_main.c | 38 ++-
drivers/net/ethernet/intel/i40e/i40e_txrx.c | 40 ++-
drivers/net/ethernet/intel/i40e/i40e_txrx.h | 1 +
drivers/net/ethernet/intel/i40e/i40e_xsk.c | 33 ++-
drivers/net/ethernet/intel/ice/ice.h | 2 +
drivers/net/ethernet/intel/ice/ice_lib.c | 21 ++
drivers/net/ethernet/intel/ice/ice_main.c | 17 ++
drivers/net/ethernet/intel/ice/ice_txrx.c | 33 ++-
drivers/net/ethernet/intel/ice/ice_txrx.h | 12 +-
drivers/net/ethernet/intel/ice/ice_txrx_lib.c | 3 +
drivers/net/ethernet/intel/ice/ice_xsk.c | 51 +++-
drivers/net/ethernet/intel/igb/igb.h | 14 +-
drivers/net/ethernet/intel/igb/igb_main.c | 102 ++++++-
drivers/net/ethernet/intel/igc/igc.h | 3 +
drivers/net/ethernet/intel/igc/igc_main.c | 89 +++++-
drivers/net/ethernet/intel/ixgbe/ixgbe.h | 1 +
drivers/net/ethernet/intel/ixgbe/ixgbe_lib.c | 3 +-
drivers/net/ethernet/intel/ixgbe/ixgbe_main.c | 69 ++++-
drivers/net/ethernet/intel/ixgbe/ixgbe_xsk.c | 56 +++-
drivers/net/ethernet/marvell/mvneta.c | 78 +++++-
.../net/ethernet/marvell/mvpp2/mvpp2_main.c | 51 ++++
drivers/net/ethernet/mellanox/mlx5/core/en.h | 5 +
.../net/ethernet/mellanox/mlx5/core/en/xdp.c | 3 +-
.../net/ethernet/mellanox/mlx5/core/en_main.c | 2 +
.../ethernet/mellanox/mlx5/core/en_stats.c | 76 +++++
.../ethernet/mellanox/mlx5/core/en_stats.h | 3 +
drivers/net/ethernet/sfc/ef100_netdev.c | 2 +
drivers/net/ethernet/sfc/efx.c | 2 +
drivers/net/ethernet/sfc/efx_common.c | 42 +++
drivers/net/ethernet/sfc/efx_common.h | 3 +
drivers/net/veth.c | 128 +++++++--
drivers/net/virtio_net.c | 104 +++++--
include/linux/if_link.h | 39 ++-
include/linux/netdevice.h | 13 +
include/net/xdp.h | 162 +++++++++++
include/uapi/linux/if_link.h | 67 +++++
net/core/rtnetlink.c | 264 ++++++++++++++++++
net/core/xdp.c | 124 ++++++++
45 files changed, 1798 insertions(+), 143 deletions(-)
--
2.33.1
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:41:27
Lots of the driver-side XDP enabled drivers provide some statistics
on XDP programs runs and different actions taken (number of passes,
drops, redirects etc.). Regarding that it's quite similar across all
the drivers (which is obvious), we can implement some sort of
generic statistics using rtnetlink xstats infra to provide a way for
exposing XDP/XSK statistics without code and stringsets duplication
inside drivers' Ethtool callbacks.
These 15 fields provided by the standard XDP stats should cover most
stats that might be interesting for collecting and tracking.
Note that most NIC drivers keep XDP statistics on a per-channel
basis, so this also introduces a new callback for getting a number
of channels which a driver will provide stats for. If it's not
implemented, we assume the driver stats are shared across channels.
If it's here, it should return either the number of channels, or 0
if stats for this type (XDP or XSK) is shared, or -EOPNOTSUPP if it
doesn't store such type of statistics, or -ENODATA if it does, but
can't provide them right now for any reason (purely for better code
readability, acts the same as -EOPNOTSUPP).
Stats are provided as nested attrs to be able to expand them later
on.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
Reviewed-by: Michal Swiatkowski <redacted>
Reviewed-by: Maciej Fijalkowski <maciej.fijalkowski@intel.com>
---
include/linux/if_link.h | 39 +++++-
include/linux/netdevice.h | 12 ++
include/uapi/linux/if_link.h | 67 +++++++++
net/core/rtnetlink.c | 264 +++++++++++++++++++++++++++++++++++
4 files changed, 381 insertions(+), 1 deletion(-)
@@ -4,8 +4,8 @@#include<uapi/linux/if_link.h>+/* We don't want these structures exposed to user space */-/* We don't want this structure exposed to user space */structifla_vf_stats{__u64rx_packets;__u64tx_packets;
@@ -1175,6 +1176,72 @@ enum {};#define IFLA_OFFLOAD_XSTATS_MAX (__IFLA_OFFLOAD_XSTATS_MAX - 1)+/* These are embedded into IFLA_STATS_LINK_XDP_XSTATS */+enum{+IFLA_XDP_XSTATS_TYPE_UNSPEC,+/* Stats collected on a "regular" channel(s) */+IFLA_XDP_XSTATS_TYPE_XDP,+/* Stats collected on an XSK channel(s) */+IFLA_XDP_XSTATS_TYPE_XSK,++__IFLA_XDP_XSTATS_TYPE_CNT,+};++#define IFLA_XDP_XSTATS_TYPE_START (IFLA_XDP_XSTATS_TYPE_UNSPEC + 1)+#define IFLA_XDP_XSTATS_TYPE_MAX (__IFLA_XDP_XSTATS_TYPE_CNT - 1)++/* Embedded into IFLA_XDP_XSTATS_TYPE_XDP or IFLA_XDP_XSTATS_TYPE_XSK */+enum{+IFLA_XDP_XSTATS_SCOPE_UNSPEC,+/* netdev-wide stats */+IFLA_XDP_XSTATS_SCOPE_SHARED,+/* Per-channel stats */+IFLA_XDP_XSTATS_SCOPE_CHANNEL,++__IFLA_XDP_XSTATS_SCOPE_CNT,+};++/* Embedded into IFLA_XDP_XSTATS_SCOPE_SHARED/IFLA_XDP_XSTATS_SCOPE_CHANNEL */+enum{+/* Padding for 64-bit alignment */+IFLA_XDP_XSTATS_UNSPEC,+/* Number of frames passed to bpf_prog_run_xdp() */+IFLA_XDP_XSTATS_PACKETS,+/* Number of bytes went through bpf_prog_run_xdp() */+IFLA_XDP_XSTATS_BYTES,+/* Number of general XDP errors if driver counts them together */+IFLA_XDP_XSTATS_ERRORS,+/* Number of %XDP_ABORTED returns */+IFLA_XDP_XSTATS_ABORTED,+/* Number of %XDP_DROP returns */+IFLA_XDP_XSTATS_DROP,+/* Number of returns of unallowed values (i.e. not XDP_*) */+IFLA_XDP_XSTATS_INVALID,+/* Number of %XDP_PASS returns */+IFLA_XDP_XSTATS_PASS,+/* Number of successfully performed %XDP_REDIRECT requests */+IFLA_XDP_XSTATS_REDIRECT,+/* Number of failed %XDP_REDIRECT requests */+IFLA_XDP_XSTATS_REDIRECT_ERRORS,+/* Number of successfully performed %XDP_TX requests */+IFLA_XDP_XSTATS_TX,+/* Number of failed %XDP_TX requests */+IFLA_XDP_XSTATS_TX_ERRORS,+/* Number of successfully transmitted XDP/XSK frames */+IFLA_XDP_XSTATS_XMIT_PACKETS,+/* Number of successfully transmitted XDP/XSK bytes */+IFLA_XDP_XSTATS_XMIT_BYTES,+/* Number of XDP/XSK frames failed to transmit */+IFLA_XDP_XSTATS_XMIT_ERRORS,+/* Number of XDP/XSK queue being full at the moment of transmission */+IFLA_XDP_XSTATS_XMIT_FULL,++__IFLA_XDP_XSTATS_CNT,+};++#define IFLA_XDP_XSTATS_START (IFLA_XDP_XSTATS_UNSPEC + 1)+#define IFLA_XDP_XSTATS_MAX (__IFLA_XDP_XSTATS_CNT - 1)+/* XDP section */#define XDP_FLAGS_UPDATE_IF_NOEXIST (1U << 0)
@@ -5107,6 +5107,262 @@ static int rtnl_get_offload_stats_size(const struct net_device *dev)returnnla_size;}+#define IFLA_XDP_XSTATS_NUM (__IFLA_XDP_XSTATS_CNT - \+IFLA_XDP_XSTATS_START)++static_assert(sizeof(structifla_xdp_stats)/sizeof(__u64)==+IFLA_XDP_XSTATS_NUM);++staticu32rtnl_get_xdp_stats_num(u32attr_id)+{+switch(attr_id){+caseIFLA_XDP_XSTATS_TYPE_XDP:+caseIFLA_XDP_XSTATS_TYPE_XSK:+returnIFLA_XDP_XSTATS_NUM;+default:+return0;+}+}++staticboolrtnl_get_xdp_stats_xdpxsk(structsk_buff*skb,u32ch,+constvoid*attr_data)+{+conststructifla_xdp_stats*xstats=attr_data;++xstats+=ch;++if(nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_PACKETS,xstats->packets,+IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_BYTES,xstats->bytes,+IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_ERRORS,xstats->errors,+IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_ABORTED,xstats->aborted,+IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_DROP,xstats->drop,+IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_INVALID,xstats->invalid,+IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_PASS,xstats->pass,+IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_REDIRECT,xstats->redirect,+IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_REDIRECT_ERRORS,+xstats->redirect_errors,+IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_TX,xstats->tx,+IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_TX_ERRORS,+xstats->tx_errors,IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_XMIT_PACKETS,+xstats->xmit_packets,IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_XMIT_BYTES,+xstats->xmit_bytes,IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_XMIT_ERRORS,+xstats->xmit_errors,IFLA_XDP_XSTATS_UNSPEC)||+nla_put_u64_64bit(skb,IFLA_XDP_XSTATS_XMIT_FULL,+xstats->xmit_full,IFLA_XDP_XSTATS_UNSPEC))+returnfalse;++returntrue;+}++staticboolrtnl_get_xdp_stats_one(structsk_buff*skb,u32attr_id,+u32scope_id,u32ch,constvoid*attr_data)+{+structnlattr*scope;++scope=nla_nest_start_noflag(skb,scope_id);+if(!scope)+returnfalse;++switch(attr_id){+caseIFLA_XDP_XSTATS_TYPE_XDP:+caseIFLA_XDP_XSTATS_TYPE_XSK:+if(!rtnl_get_xdp_stats_xdpxsk(skb,ch,attr_data))+gotofail;++break;+default:+fail:+nla_nest_cancel(skb,scope);++returnfalse;+}++nla_nest_end(skb,scope);++returntrue;+}++staticboolrtnl_get_xdp_stats(structsk_buff*skb,+conststructnet_device*dev,+int*idxattr,int*prividx)+{+conststructnet_device_ops*ops=dev->netdev_ops;+structnlattr*xstats,*type=NULL;+u32saved_ch=*prividx&U16_MAX;+u32saved_attr=*prividx>>16;+boolnuke_xstats=true;+u32attr_id,ch=0;+intret;++if(!ops||!ops->ndo_get_xdp_stats)+gotonodata;++*idxattr=IFLA_STATS_LINK_XDP_XSTATS;++xstats=nla_nest_start_noflag(skb,IFLA_STATS_LINK_XDP_XSTATS);+if(!xstats)+returnfalse;++for(attr_id=IFLA_XDP_XSTATS_TYPE_START;+attr_id<__IFLA_XDP_XSTATS_TYPE_CNT;+attr_id++){+u32nstat,scope_id,nch;+boolnuke_type=true;+void*attr_data;+size_tsize;++if(attr_id>saved_attr)+saved_ch=0;+if(attr_id<saved_attr)+continue;++nstat=rtnl_get_xdp_stats_num(attr_id);+if(!nstat)+continue;++scope_id=IFLA_XDP_XSTATS_SCOPE_SHARED;+nch=1;++if(!ops->ndo_get_xdp_stats_nch)+gotoshared;++ret=ops->ndo_get_xdp_stats_nch(dev,attr_id);+if(ret==-EOPNOTSUPP||ret==-ENODATA)+continue;+if(ret<0)+gotoout;+if(!ret)+gotoshared;++scope_id=IFLA_XDP_XSTATS_SCOPE_CHANNEL;+nch=ret;++shared:+size=array3_size(nch,nstat,sizeof(__u64));+if(unlikely(size==SIZE_MAX)){+ret=-EOVERFLOW;+gotoout;+}++attr_data=kzalloc(size,GFP_KERNEL);+if(!attr_data){+ret=-ENOMEM;+gotoout;+}++ret=ops->ndo_get_xdp_stats(dev,attr_id,attr_data);+if(ret==-EOPNOTSUPP||ret==-ENODATA)+gotokfree_cont;+if(ret){+kfree_out:+kfree(attr_data);+gotoout;+}++ret=-EMSGSIZE;++type=nla_nest_start_noflag(skb,attr_id);+if(!type)+gotokfree_out;++for(ch=saved_ch;ch<nch;ch++)+if(!rtnl_get_xdp_stats_one(skb,attr_id,scope_id,+ch,attr_data)){+if(nuke_type)+nla_nest_cancel(skb,type);+else+nla_nest_end(skb,type);++gotokfree_out;+}else{+nuke_xstats=false;+nuke_type=false;+}++nla_nest_end(skb,type);+kfree_cont:+kfree(attr_data);+}++ret=0;++out:+if(nuke_xstats)+nla_nest_cancel(skb,xstats);+else+nla_nest_end(skb,xstats);++if(ret&&ret!=-EOPNOTSUPP&&ret!=-ENODATA){+/* If the driver has 60+ queues, we can run out of skb+*tailroomevenwhenputtingstatsforonetype.Save+*channelnumberinprividxtoresumefromitnexttime+*ratherthanrestaringthewholetypeandrunninginto+*thesameproblemagain.+*/+*prividx=(attr_id<<16)|ch;+returnfalse;+}++*prividx=0;+nodata:+*idxattr=0;++returntrue;+}++staticsize_trtnl_get_xdp_stats_size(conststructnet_device*dev)+{+conststructnet_device_ops*ops=dev->netdev_ops;+size_tsize=0;+u32attr_id;++if(!ops||!ops->ndo_get_xdp_stats)+return0;++for(attr_id=IFLA_XDP_XSTATS_TYPE_START;+attr_id<__IFLA_XDP_XSTATS_TYPE_CNT;+attr_id++){+u32nstat=rtnl_get_xdp_stats_num(attr_id);+u32nch=1;+intret;++if(!nstat)+continue;++if(!ops->ndo_get_xdp_stats_nch)+gotoshared;++ret=ops->ndo_get_xdp_stats_nch(dev,attr_id);+if(ret<0)+continue;+if(ret>0)+nch=ret;++shared:+size+=nla_total_size(0)+/* IFLA_XDP_XSTATS_TYPE_* */+(nla_total_size(0)+/* IFLA_XDP_XSTATS_SCOPE_* */+nla_total_size_64bit(sizeof(__u64))*nstat)*nch;+}++if(size)+size+=nla_total_size(0);/* IFLA_STATS_LINK_XDP_XSTATS */++returnsize;+}+staticintrtnl_fill_statsinfo(structsk_buff*skb,structnet_device*dev,inttype,u32pid,u32seq,u32change,unsignedintflags,unsignedintfilter_mask,
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:41:34
ena driver has 6 XDP counters collected per-channel. Add callbacks
for getting the number of channels and those counters using generic
XDP stats infra.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
drivers/net/ethernet/amazon/ena/ena_netdev.c | 53 ++++++++++++++++++++
1 file changed, 53 insertions(+)
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:41:40
Add several shorthands to reduce driver boilerplates and unify
storing and accessing generic XDP statistics in the drivers.
If the driver has one of xdp_{rx,tx}_drv_stats embedded into
a ring structure, it can reuse pretty much everything, but needs
to implement its own .ndo_xdp_stats() and .ndo_xdp_stats_nch()
if needed. If the driver stores a separate array of xdp_drv_stats,
it can then export it as net_device::xstats, implement only
.ndo_xdp_stats_nch() and wire up xdp_get_drv_stats_generic()
as .ndo_xdp_stats().
Both XDP and XSK blocks of xdp_drv_stats are cacheline-aligned
to avoid false-sharing, only extremely unlikely 'aborted' and
'invalid' falls out of a 64-byte CL. xdp_rx_drv_stats_local is
provided to put it on stack and collect the stats on hotpath,
with accessing a real container and its atomic/seqcount sync
points just once when exiting Rx NAPI polling.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
Reviewed-by: Michal Swiatkowski <redacted>
Reviewed-by: Maciej Fijalkowski <maciej.fijalkowski@intel.com>
---
include/linux/netdevice.h | 1 +
include/net/xdp.h | 162 ++++++++++++++++++++++++++++++++++++++
net/core/xdp.c | 124 +++++++++++++++++++++++++++++
3 files changed, 287 insertions(+)
@@ -611,3 +611,127 @@ struct xdp_frame *xdpf_clone(struct xdp_frame *xdpf)returnnxdpf;}++/**+*xdp_fetch_rx_drv_stats-helperforimplementing.ndo_get_xdp_stats()+*@if_stats:targetcontainerpassedfromrtnetlinkcore+*@rstats:drivercontainerifitusesgenericxdp_rx_drv_stats+*+*FetchesRxpathXDPstatisticsfromasuggesteddriverstructureto+*theoneusedbyrtnetlink,respectingatomic/seqcountsynchronization.+*/+voidxdp_fetch_rx_drv_stats(structifla_xdp_stats*if_stats,+conststructxdp_rx_drv_stats*rstats)+{+u32start;++do{+start=u64_stats_fetch_begin_irq(&rstats->syncp);++if_stats->packets=u64_stats_read(&rstats->packets);+if_stats->bytes=u64_stats_read(&rstats->bytes);+if_stats->pass=u64_stats_read(&rstats->pass);+if_stats->drop=u64_stats_read(&rstats->drop);+if_stats->tx=u64_stats_read(&rstats->tx);+if_stats->tx_errors=u64_stats_read(&rstats->tx_errors);+if_stats->redirect=u64_stats_read(&rstats->redirect);+if_stats->redirect_errors=+u64_stats_read(&rstats->redirect_errors);+if_stats->aborted=u64_stats_read(&rstats->aborted);+if_stats->invalid=u64_stats_read(&rstats->invalid);+}while(u64_stats_fetch_retry_irq(&rstats->syncp,start));+}+EXPORT_SYMBOL_GPL(xdp_fetch_rx_drv_stats);++/**+*xdp_fetch_tx_drv_stats-helperforimplementing.ndo_get_xdp_stats()+*@if_stats:targetcontainerpassedfromrtnetlinkcore+*@tstats:drivercontainerifitusesgenericxdp_tx_drv_stats+*+*FetchesTxpathXDPstatisticsfromasuggesteddriverstructureto+*theoneusedbyrtnetlink,respectingatomic/seqcountsynchronization.+*/+voidxdp_fetch_tx_drv_stats(structifla_xdp_stats*if_stats,+conststructxdp_tx_drv_stats*tstats)+{+u32start;++do{+start=u64_stats_fetch_begin_irq(&tstats->syncp);++if_stats->xmit_packets=u64_stats_read(&tstats->packets);+if_stats->xmit_bytes=u64_stats_read(&tstats->bytes);+if_stats->xmit_errors=u64_stats_read(&tstats->errors);+if_stats->xmit_full=u64_stats_read(&tstats->full);+}while(u64_stats_fetch_retry_irq(&tstats->syncp,start));+}+EXPORT_SYMBOL_GPL(xdp_fetch_tx_drv_stats);++/**+*xdp_get_drv_stats_generic-genericimplementationof.ndo_get_xdp_stats()+*@dev:networkinterfacedevicestructure+*@attr_id:typeofstatistics(XDP,XSK,...)+*@attr_data:targetstatscontainer+*+*Returns0onsuccess,-%EOPNOTSUPPifeitherdriverorthisfunctiondoesn't+*supportthisattr_id,-%ENODATAifthedriversupportsattr_id,butcan't+*provideanythingrightnow,and-%EINVALifdriverconfigurationisinvalid.+*/+intxdp_get_drv_stats_generic(conststructnet_device*dev,u32attr_id,+void*attr_data)+{+constboolxsk=attr_id==IFLA_XDP_XSTATS_TYPE_XSK;+conststructxdp_drv_stats*drv_iter=dev->xstats;+conststructnet_device_ops*ops=dev->netdev_ops;+structifla_xdp_stats*iter=attr_data;+intnch;+u32i;++switch(attr_id){+caseIFLA_XDP_XSTATS_TYPE_XDP:+if(unlikely(!ops->ndo_bpf))+return-EINVAL;++break;+caseIFLA_XDP_XSTATS_TYPE_XSK:+if(!ops->ndo_xsk_wakeup)+return-EOPNOTSUPP;++break;+default:+return-EOPNOTSUPP;+}++if(unlikely(!drv_iter||!ops->ndo_get_xdp_stats_nch))+return-EINVAL;++nch=ops->ndo_get_xdp_stats_nch(dev,attr_id);+switch(nch){+case0:+/* Stats are shared across the netdev */+nch=1;+break;+case1...INT_MAX:+/* Stats are per-channel */+break;+default:+returnnch;+}++for(i=0;i<nch;i++){+conststructxdp_rx_drv_stats*rstats;+conststructxdp_tx_drv_stats*tstats;++rstats=xsk?&drv_iter->xsk_rx:&drv_iter->xdp_rx;+xdp_fetch_rx_drv_stats(iter,rstats);++tstats=xsk?&drv_iter->xsk_tx:&drv_iter->xdp_tx;+xdp_fetch_tx_drv_stats(iter,tstats);++drv_iter++;+iter++;+}++return0;+}+EXPORT_SYMBOL_GPL(xdp_get_drv_stats_generic);--
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:41:53
Some of the initializers are aligned with spaces, others with tabs.
Reindent it using tabs only.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
drivers/net/ethernet/marvell/mvneta.c | 24 ++++++++++++------------
1 file changed, 12 insertions(+), 12 deletions(-)
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:41:59
mvneta driver implements 7 per-cpu counters which means we can
only provide them as a global sum across CPUs.
Implement a callback for querying them using generic XDP stats
infra.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
drivers/net/ethernet/marvell/mvneta.c | 54 +++++++++++++++++++++++++++
1 file changed, 54 insertions(+)
@@ -802,6 +802,59 @@ mvneta_get_stats64(struct net_device *dev,stats->tx_dropped=dev->stats.tx_dropped;}+staticintmvneta_get_xdp_stats(conststructnet_device*dev,u32attr_id,+void*attr_data)+{+conststructmvneta_port*pp=netdev_priv(dev);+structifla_xdp_stats*xdp_stats=attr_data;+u32cpu;++switch(attr_id){+caseIFLA_XDP_XSTATS_TYPE_XDP:+break;+default:+return-EOPNOTSUPP;+}++for_each_possible_cpu(cpu){+conststructmvneta_pcpu_stats*stats;+conststructmvneta_stats*ps;+u64xdp_xmit_err;+u64xdp_redirect;+u64xdp_tx_err;+u64xdp_pass;+u64xdp_drop;+u64xdp_xmit;+u64xdp_tx;+u32start;++stats=per_cpu_ptr(pp->stats,cpu);+ps=&stats->es.ps;++do{+start=u64_stats_fetch_begin_irq(&stats->syncp);++xdp_drop=ps->xdp_drop;+xdp_pass=ps->xdp_pass;+xdp_redirect=ps->xdp_redirect;+xdp_tx=ps->xdp_tx;+xdp_tx_err=ps->xdp_tx_err;+xdp_xmit=ps->xdp_xmit;+xdp_xmit_err=ps->xdp_xmit_err;+}while(u64_stats_fetch_retry_irq(&stats->syncp,start));++xdp_stats->drop+=xdp_drop;+xdp_stats->pass+=xdp_pass;+xdp_stats->redirect+=xdp_redirect;+xdp_stats->tx+=xdp_tx;+xdp_stats->tx_errors+=xdp_tx_err;+xdp_stats->xmit_packets+=xdp_xmit;+xdp_stats->xmit_errors+=xdp_xmit_err;+}++return0;+}+/* Rx descriptors helper methods *//* Checks whether the RX descriptor having this status is both the first
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:42:04
They get updated not only on XDP path. Moreover, packet counter
stores the total number of frames, not only the ones passed to
bpf_prog_run_xdp(), so it's rather confusing.
Drop the xdp_ suffix from both of them to not mix XDP-only stats
with the general ones.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
drivers/net/veth.c | 36 ++++++++++++++++++------------------
1 file changed, 18 insertions(+), 18 deletions(-)
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:42:09
mlx5 driver has a bunch of per-channel stats for XDP. 7 and 5 of
them can be exported through generic XDP stats infra for XDP and XSK
correspondingly.
Add necessary calbacks for that. Note that the driver doesn't expose
XSK stats if XSK setup has never been requested.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
drivers/net/ethernet/mellanox/mlx5/core/en.h | 5 ++
.../net/ethernet/mellanox/mlx5/core/en_main.c | 2 +
.../ethernet/mellanox/mlx5/core/en_stats.c | 69 +++++++++++++++++++
3 files changed, 76 insertions(+)
@@ -1212,4 +1212,9 @@ int mlx5e_set_vf_rate(struct net_device *dev, int vf, int min_tx_rate, int max_tintmlx5e_get_vf_config(structnet_device*dev,intvf,structifla_vf_info*ivi);intmlx5e_get_vf_stats(structnet_device*dev,intvf,structifla_vf_stats*vf_stats);#endif++intmlx5e_get_xdp_stats_nch(conststructnet_device*dev,u32attr_id);+intmlx5e_get_xdp_stats(conststructnet_device*dev,u32attr_id,+void*attr_data);+#endif /* __MLX5_EN_H__ */
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:42:18
Provide a separate counter [rx_]xdp_errors for drops related to
XDP_ABORTED and other errors/exceptions and leave [rx_]xdp_drop
only for XDP_DROP case.
This will align the driver better with generic XDP stats.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
drivers/net/ethernet/mellanox/mlx5/core/en/xdp.c | 3 ++-
drivers/net/ethernet/mellanox/mlx5/core/en_stats.c | 7 +++++++
drivers/net/ethernet/mellanox/mlx5/core/en_stats.h | 3 +++
3 files changed, 12 insertions(+), 1 deletion(-)
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:42:22
Dedicate a separate counter for tracking XDP_ABORTED and other XDP
errors and to leave xdp_drop for XDP_DROP case solely.
Needed to better align with generic XDP stats.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
drivers/net/virtio_net.c | 18 ++++++++++++++----
1 file changed, 14 insertions(+), 4 deletions(-)
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:42:28
To align better with other drivers and generic XDP stats, rename
xdp_tx{,_drops} to xdp_xmit{,_errors} as they're used on
.ndo_xdp_xmit() path.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
drivers/net/virtio_net.c | 12 ++++++------
1 file changed, 6 insertions(+), 6 deletions(-)
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:42:30
Some of the initializers are aligned with spaces, others with tabs.
Reindent it using tabs only.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
drivers/net/virtio_net.c | 18 +++++++++---------
1 file changed, 9 insertions(+), 9 deletions(-)
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:42:34
Some of the initializers are aligned with spaces, others with tabs.¬
Reindent it using tabs only.¬
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
drivers/net/veth.c | 14 +++++++-------
1 file changed, 7 insertions(+), 7 deletions(-)
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:42:51
There's no need to fetch an XSK pool desc in case our ring is full,
we can rather quit under unlikely branch.
Can't skip taking a lock here unfortunately since igc_desc_unused()
assumes we call it being locked.
This was designed to track xsk_tx::full counter, but won't hurt
either way.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
Reviewed-by: Michal Swiatkowski <redacted>
---
drivers/net/ethernet/intel/igc/igc_main.c | 3 +++
1 file changed, 3 insertions(+)
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:42:56
Make ixgbe driver collect and provide all generic XDP/XSK counters.
Unfortunately, XDP rings have a lifetime of an XDP prog, and all
ring stats structures get wiped on xsk_pool attach/detach, so
store them in a separate array with a lifetime of a netdev.
Reuse all previously introduced helpers and
xdp_get_drv_stats_generic(). Performance wavering from incrementing
a bunch of counters on hotpath is around stddev at [64 ... 1532]
frame sizes.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
Reviewed-by: Michal Swiatkowski <redacted>
---
drivers/net/ethernet/intel/ixgbe/ixgbe.h | 1 +
drivers/net/ethernet/intel/ixgbe/ixgbe_lib.c | 3 +-
drivers/net/ethernet/intel/ixgbe/ixgbe_main.c | 69 ++++++++++++++++---
drivers/net/ethernet/intel/ixgbe/ixgbe_xsk.c | 56 +++++++++++----
4 files changed, 106 insertions(+), 23 deletions(-)
@@ -951,6 +951,7 @@ static int ixgbe_alloc_q_vector(struct ixgbe_adapter *adapter,ring->queue_index=xdp_idx;set_ring_xdp(ring);spin_lock_init(&ring->tx_lock);+ring->xdp_stats=adapter->netdev->xstats+xdp_idx;/* assign ring to adapter */WRITE_ONCE(adapter->xdp_ring[xdp_idx],ring);
@@ -994,6 +995,7 @@ static int ixgbe_alloc_q_vector(struct ixgbe_adapter *adapter,/* apply Rx specific ring traits */ring->count=adapter->rx_ring_count;ring->queue_index=rxr_idx;+ring->xdp_stats=adapter->netdev->xstats+rxr_idx;/* assign ring to adapter */WRITE_ONCE(adapter->rx_ring[rxr_idx],ring);
@@ -2348,7 +2370,7 @@ static int ixgbe_clean_rx_irq(struct ixgbe_q_vector *q_vector,/* At larger PAGE_SIZE, frame_sz depend on len size */xdp.frame_sz=ixgbe_rx_frame_truesize(rx_ring,size);#endif-skb=ixgbe_run_xdp(adapter,rx_ring,&xdp);+skb=ixgbe_run_xdp(adapter,rx_ring,&xdp,&lrstats);}if(IS_ERR(skb)){
@@ -2440,6 +2462,7 @@ static int ixgbe_clean_rx_irq(struct ixgbe_q_vector *q_vector,rx_ring->stats.packets+=total_rx_packets;rx_ring->stats.bytes+=total_rx_bytes;u64_stats_update_end(&rx_ring->syncp);+xdp_update_rx_drv_stats(&rx_ring->xdp_stats->xdp_rx,&lrstats);q_vector->rx.total_packets+=total_rx_packets;q_vector->rx.total_bytes+=total_rx_bytes;
@@ -8552,8 +8575,10 @@ int ixgbe_xmit_xdp_ring(struct ixgbe_ring *ring,len=xdpf->len;-if(unlikely(!ixgbe_desc_unused(ring)))+if(unlikely(!ixgbe_desc_unused(ring))){+xdp_update_tx_drv_full(&ring->xdp_stats->xdp_tx);returnIXGBE_XDP_CONSUMED;+}dma=dma_map_single(ring->dev,xdpf->data,len,DMA_TO_DEVICE);if(dma_mapping_error(ring->dev,dma))
@@ -10257,12 +10282,26 @@ static int ixgbe_xdp_xmit(struct net_device *dev, int n,if(unlikely(flags&XDP_XMIT_FLUSH))ixgbe_xdp_ring_update_tail(ring);+if(unlikely(nxmit<n))+xdp_update_tx_drv_err(&ring->xdp_stats->xdp_tx,n-nxmit);+if(static_branch_unlikely(&ixgbe_xdp_locking_key))spin_unlock(&ring->tx_lock);returnnxmit;}+staticintixgbe_get_xdp_stats_nch(conststructnet_device*dev,u32attr_id)+{+switch(attr_id){+caseIFLA_XDP_XSTATS_TYPE_XDP:+caseIFLA_XDP_XSTATS_TYPE_XSK:+returnIXGBE_MAX_XDP_QS;+default:+return-EOPNOTSUPP;+}+}+staticconststructnet_device_opsixgbe_netdev_ops={.ndo_open=ixgbe_open,.ndo_stop=ixgbe_close,
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:43:00
Make i40e driver collect and provide all generic XDP/XSK counters.
Unfortunately, XDP rings have a lifetime of an XDP prog, and all
ring stats structures get wiped on xsk_pool attach/detach, so
store them in a separate array with a lifetime of a VSI.
Reuse all previously introduced helpers and
xdp_get_drv_stats_generic(). Performance wavering from incrementing
a bunch of counters on hotpath is around stddev at [64 ... 1532]
frame sizes.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
Reviewed-by: Michal Swiatkowski <redacted>
---
drivers/net/ethernet/intel/i40e/i40e.h | 1 +
drivers/net/ethernet/intel/i40e/i40e_main.c | 38 +++++++++++++++++++-
drivers/net/ethernet/intel/i40e/i40e_txrx.c | 40 +++++++++++++++++----
drivers/net/ethernet/intel/i40e/i40e_txrx.h | 1 +
drivers/net/ethernet/intel/i40e/i40e_xsk.c | 33 +++++++++++++----
5 files changed, 99 insertions(+), 14 deletions(-)
@@ -11087,7 +11087,7 @@ static int i40e_set_num_rings_in_vsi(struct i40e_vsi *vsi)staticinti40e_vsi_alloc_arrays(structi40e_vsi*vsi,boolalloc_qvectors){structi40e_ring**next_rings;-intsize;+intsize,i;intret=0;/* allocate memory for both Tx, XDP Tx and Rx ring pointers */
@@ -11103,6 +11103,15 @@ static int i40e_vsi_alloc_arrays(struct i40e_vsi *vsi, bool alloc_qvectors)}vsi->rx_rings=next_rings;+vsi->xdp_stats=kcalloc(vsi->alloc_queue_pairs,+sizeof(*vsi->xdp_stats),+GFP_KERNEL);+if(!vsi->xdp_stats)+gotoerr_xdp_stats;++for(i=0;i<vsi->alloc_queue_pairs;i++)+xdp_init_drv_stats(vsi->xdp_stats+i);+if(alloc_qvectors){/* allocate memory for q_vector pointers */size=sizeof(structi40e_q_vector*)*vsi->num_q_vectors;
@@ -2303,33 +2308,48 @@ static int i40e_run_xdp(struct i40e_ring *rx_ring, struct xdp_buff *xdp)if(!xdp_prog)gotoxdp_out;+lrstats->bytes+=xdp->data_end-xdp->data;+lrstats->packets++;+prefetchw(xdp->data_hard_start);/* xdp_frame write */act=bpf_prog_run_xdp(xdp_prog,xdp);switch(act){caseXDP_PASS:+lrstats->pass++;break;caseXDP_TX:xdp_ring=rx_ring->vsi->xdp_rings[rx_ring->queue_index];result=i40e_xmit_xdp_tx_ring(xdp,xdp_ring);-if(result==I40E_XDP_CONSUMED)+if(result==I40E_XDP_CONSUMED){+lrstats->tx_errors++;gotoout_failure;+}+lrstats->tx++;break;caseXDP_REDIRECT:err=xdp_do_redirect(rx_ring->netdev,xdp,xdp_prog);-if(err)+if(err){+lrstats->redirect_errors++;gotoout_failure;+}result=I40E_XDP_REDIR;+lrstats->redirect++;break;default:bpf_warn_invalid_xdp_action(act);-fallthrough;+lrstats->invalid++;+gotoout_failure;caseXDP_ABORTED:+lrstats->aborted++;out_failure:trace_xdp_exception(rx_ring->netdev,xdp_prog,act);-fallthrough;/* handle aborts by dropping packet */+/* handle aborts by dropping packet */+result=I40E_XDP_CONSUMED;+break;caseXDP_DROP:result=I40E_XDP_CONSUMED;+lrstats->drop++;break;}xdp_out:
@@ -2441,6 +2461,7 @@ static int i40e_clean_rx_irq(struct i40e_ring *rx_ring, int budget){unsignedinttotal_rx_bytes=0,total_rx_packets=0,frame_sz=0;u16cleaned_count=I40E_DESC_UNUSED(rx_ring);+structxdp_rx_drv_stats_locallrstats={};unsignedintoffset=rx_ring->rx_offset;structsk_buff*skb=rx_ring->skb;unsignedintxdp_xmit=0;
@@ -2512,7 +2533,7 @@ static int i40e_clean_rx_irq(struct i40e_ring *rx_ring, int budget)/* At larger PAGE_SIZE, frame_sz depend on len size */xdp.frame_sz=i40e_rx_frame_truesize(rx_ring,size);#endif-xdp_res=i40e_run_xdp(rx_ring,&xdp);+xdp_res=i40e_run_xdp(rx_ring,&xdp,&lrstats);}if(xdp_res){
@@ -2569,6 +2590,7 @@ static int i40e_clean_rx_irq(struct i40e_ring *rx_ring, int budget)rx_ring->skb=skb;i40e_update_rx_stats(rx_ring,total_rx_bytes,total_rx_packets);+xdp_update_rx_drv_stats(&rx_ring->xdp_stats->xdp_rx,&lrstats);/* guarantee a trip back through this routine if there was a failure */returnfailure?budget:(int)total_rx_packets;
@@ -3696,6 +3718,7 @@ static int i40e_xmit_xdp_ring(struct xdp_frame *xdpf,dma_addr_tdma;if(!unlikely(I40E_DESC_UNUSED(xdp_ring))){+xdp_update_tx_drv_full(&xdp_ring->xdp_stats->xdp_tx);xdp_ring->tx_stats.tx_busy++;returnI40E_XDP_CONSUMED;}
@@ -3923,5 +3946,8 @@ int i40e_xdp_xmit(struct net_device *dev, int n, struct xdp_frame **frames,if(unlikely(flags&XDP_XMIT_FLUSH))i40e_xdp_ring_update_tail(xdp_ring);+if(unlikely(nxmit<n))+xdp_update_tx_drv_err(&xdp_ring->xdp_stats->xdp_tx,n-nxmit);+returnnxmit;}
@@ -368,6 +368,7 @@ struct i40e_ring {structi40e_tx_queue_statstx_stats;structi40e_rx_queue_statsrx_stats;};+structxdp_drv_stats*xdp_stats;unsignedintsize;/* length of descriptor ring in bytes */dma_addr_tdma;/* physical address of ring */
@@ -143,16 +143,21 @@ int i40e_xsk_pool_setup(struct i40e_vsi *vsi, struct xsk_buff_pool *pool,*i40e_run_xdp_zc-ExecutesanXDPprogramonanxdp_buff*@rx_ring:Rxring*@xdp:xdp_buffusedasinputtotheXDPprogram+*@lrstats:onstackRxXDPstatsstructure**ReturnsanyofI40E_XDP_{PASS,CONSUMED,TX,REDIR}**/-staticinti40e_run_xdp_zc(structi40e_ring*rx_ring,structxdp_buff*xdp)+staticinti40e_run_xdp_zc(structi40e_ring*rx_ring,structxdp_buff*xdp,+structxdp_rx_drv_stats_local*lrstats){interr,result=I40E_XDP_PASS;structi40e_ring*xdp_ring;structbpf_prog*xdp_prog;u32act;+lrstats->bytes+=xdp->data_end-xdp->data;+lrstats->packets++;+/* NB! xdp_prog will always be !NULL, due to the fact that*thispathisenabledbysettinganXDPprogram.*/
@@ -161,29 +166,41 @@ static int i40e_run_xdp_zc(struct i40e_ring *rx_ring, struct xdp_buff *xdp)if(likely(act==XDP_REDIRECT)){err=xdp_do_redirect(rx_ring->netdev,xdp,xdp_prog);-if(err)+if(err){+lrstats->redirect_errors++;gotoout_failure;+}+lrstats->redirect++;returnI40E_XDP_REDIR;}switch(act){caseXDP_PASS:+lrstats->pass++;break;caseXDP_TX:xdp_ring=rx_ring->vsi->xdp_rings[rx_ring->queue_index];result=i40e_xmit_xdp_tx_ring(xdp,xdp_ring);-if(result==I40E_XDP_CONSUMED)+if(result==I40E_XDP_CONSUMED){+lrstats->tx_errors++;gotoout_failure;+}+lrstats->tx++;break;default:bpf_warn_invalid_xdp_action(act);-fallthrough;+lrstats->invalid++;+gotoout_failure;caseXDP_ABORTED:+lrstats->aborted++;out_failure:trace_xdp_exception(rx_ring->netdev,xdp_prog,act);-fallthrough;/* handle aborts by dropping packet */+/* handle aborts by dropping packet */+result=I40E_XDP_CONSUMED;+break;caseXDP_DROP:result=I40E_XDP_CONSUMED;+lrstats->drop++;break;}returnresult;
@@ -325,6 +342,7 @@ int i40e_clean_rx_irq_zc(struct i40e_ring *rx_ring, int budget){unsignedinttotal_rx_bytes=0,total_rx_packets=0;u16cleaned_count=I40E_DESC_UNUSED(rx_ring);+structxdp_rx_drv_stats_locallrstats={};u16next_to_clean=rx_ring->next_to_clean;u16count_mask=rx_ring->count-1;unsignedintxdp_res,xdp_xmit=0;
@@ -366,7 +384,7 @@ int i40e_clean_rx_irq_zc(struct i40e_ring *rx_ring, int budget)xsk_buff_set_size(bi,size);xsk_buff_dma_sync_for_cpu(bi,rx_ring->xsk_pool);-xdp_res=i40e_run_xdp_zc(rx_ring,bi);+xdp_res=i40e_run_xdp_zc(rx_ring,bi,&lrstats);i40e_handle_xdp_result_zc(rx_ring,bi,rx_desc,&rx_packets,&rx_bytes,size,xdp_res);total_rx_packets+=rx_packets;
@@ -383,6 +401,7 @@ int i40e_clean_rx_irq_zc(struct i40e_ring *rx_ring, int budget)i40e_finalize_xdp_rx(rx_ring,xdp_xmit);i40e_update_rx_stats(rx_ring,total_rx_bytes,total_rx_packets);+xdp_update_rx_drv_stats(&rx_ring->xdp_stats->xsk_rx,&lrstats);if(xsk_uses_need_wakeup(rx_ring->xsk_pool)){if(failure||next_to_clean==rx_ring->next_to_use)
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:43:02
Add a couple of hints on how to retrieve and implement generic XDP
statistics for drivers/interfaces. Mention that it's unwanted to
include related XDP counters in driver-defined Ethtool stats.
Signed-off-by: Alexander Lobakin <redacted>
---
Documentation/networking/statistics.rst | 33 +++++++++++++++++++++++++
1 file changed, 33 insertions(+)
@@ -41,6 +41,29 @@ If `-s` is specified once the detailed errors won't be shown.`ip` supports JSON formatting via the `-j` option.+For some interfaces, standard XDP statistics are available.+It can be accessed the same ways, e.g. `ip`::++ $ ip link xdpstats dev enp178s0+ 16: enp178s0:+ xdp-channel0-rx_xdp_packets: 0+ xdp-channel0-rx_xdp_bytes: 1+ xdp-channel0-rx_xdp_errors: 2+ xdp-channel0-rx_xdp_aborted: 3+ xdp-channel0-rx_xdp_drop: 4+ xdp-channel0-rx_xdp_invalid: 5+ xdp-channel0-rx_xdp_pass: 6+ xdp-channel0-rx_xdp_redirect: 7+ xdp-channel0-rx_xdp_redirect_errors: 8+ xdp-channel0-rx_xdp_tx: 9+ xdp-channel0-rx_xdp_tx_errors: 10+ xdp-channel0-tx_xdp_xmit_packets: 11+ xdp-channel0-tx_xdp_xmit_bytes: 12+ xdp-channel0-tx_xdp_xmit_errors: 13+ xdp-channel0-tx_xdp_xmit_full: 14++Those are usually per-channel. JSON is also supported via the `-j` opt.+ Protocol-specific statistics ----------------------------
@@ -147,6 +170,8 @@ Statistics are reported both in the responses to link information requests (`RTM_GETLINK`) and statistic requests (`RTM_GETSTATS`, when `IFLA_STATS_LINK_64` bit is set in the `.filter_mask` of the request).+`IFLA_STATS_LINK_XDP_XSTATS` bit is used to retrieve standard XDP statstics.+ ethtool -------
@@ -206,6 +231,14 @@ Retrieving ethtool statistics is a multi-syscall process, drivers are advised to keep the number of statistics constant to avoid race conditions with user space trying to read them.+It is up to the developers whether to implement XDP statistics or not due to+possible performance hits. If so, it is encouraged to export it using generic+XDP statistics infrastructure, not driver-defined Ethtool stats.+It can be achieve by implementing `.ndo_get_xdp_stats` and, optionally but+preferred, `.ndo_get_xdp_stats_nch`. There are several common helper structures+and functions in `include/net/xdp.h` to make this simpler and keep the code+compact.+ Statistics must persist across routine operations like bringing the interface down and up.
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:43:28
Make ice driver collect and provide all generic XDP/XSK counters.
Unfortunately, XDP rings have a lifetime of an XDP prog, and all
ring stats structures get wiped on xsk_pool attach/detach, so
store them in a separate array with a lifetime of a VSI. New
alloc_xdp_stats field is used to calculate the maximum possible
number of XDP-enabled queues just once and refer to it later.
Reuse all previously introduced helpers and
xdp_get_drv_stats_generic(). Performance wavering from incrementing
a bunch of counters on hotpath is around stddev at [64 ... 1532]
frame sizes.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
Reviewed-by: Michal Swiatkowski <redacted>
Reviewed-by: Maciej Fijalkowski <maciej.fijalkowski@intel.com>
---
drivers/net/ethernet/intel/ice/ice.h | 2 +
drivers/net/ethernet/intel/ice/ice_lib.c | 21 ++++++++
drivers/net/ethernet/intel/ice/ice_main.c | 17 +++++++
drivers/net/ethernet/intel/ice/ice_txrx.c | 33 +++++++++---
drivers/net/ethernet/intel/ice/ice_txrx.h | 12 +++--
drivers/net/ethernet/intel/ice/ice_txrx_lib.c | 3 ++
drivers/net/ethernet/intel/ice/ice_xsk.c | 51 ++++++++++++++-----
7 files changed, 118 insertions(+), 21 deletions(-)
@@ -627,6 +642,9 @@ ice_xdp_xmit(struct net_device *dev, int n, struct xdp_frame **frames,if(static_branch_unlikely(&ice_xdp_locking_key))spin_unlock(&xdp_ring->tx_lock);+if(unlikely(nxmit<n))+xdp_update_tx_drv_err(&xdp_ring->xdp_stats->xdp_tx,n-nxmit);+returnnxmit;}
@@ -1089,6 +1107,7 @@ int ice_clean_rx_irq(struct ice_rx_ring *rx_ring, int budget){unsignedinttotal_rx_bytes=0,total_rx_pkts=0,frame_sz=0;u16cleaned_count=ICE_DESC_UNUSED(rx_ring);+structxdp_rx_drv_stats_locallrstats={};unsignedintoffset=rx_ring->rx_offset;structice_tx_ring*xdp_ring=NULL;unsignedintxdp_res,xdp_xmit=0;
@@ -1173,7 +1192,8 @@ int ice_clean_rx_irq(struct ice_rx_ring *rx_ring, int budget)if(!xdp_prog)gotoconstruct_skb;-xdp_res=ice_run_xdp(rx_ring,&xdp,xdp_prog,xdp_ring);+xdp_res=ice_run_xdp(rx_ring,&xdp,xdp_prog,xdp_ring,+&lrstats);if(!xdp_res)gotoconstruct_skb;if(xdp_res&(ICE_XDP_TX|ICE_XDP_REDIR)){
@@ -1254,6 +1274,7 @@ int ice_clean_rx_irq(struct ice_rx_ring *rx_ring, int budget)rx_ring->skb=skb;ice_update_rx_ring_stats(rx_ring,total_rx_pkts,total_rx_bytes);+xdp_update_rx_drv_stats(&rx_ring->xdp_stats->xdp_rx,&lrstats);/* guarantee a trip back through this routine if there was a failure */returnfailure?budget:(int)total_rx_pkts;
@@ -284,9 +284,9 @@ struct ice_rx_ring {structice_rxq_statsrx_stats;structice_q_statsstats;structu64_stats_syncsyncp;+structxdp_drv_stats*xdp_stats;-structrcu_headrcu;/* to avoid race on free */-/* CL4 - 3rd cacheline starts here */+/* CL4 - 4rd cacheline starts here */structice_channel*ch;structbpf_prog*xdp_prog;structice_tx_ring*xdp_ring;
@@ -298,6 +298,9 @@ struct ice_rx_ring {u8dcb_tc;/* Traffic class of ring */u8ptp_rx;u8flags;++/* CL5 - 5th cacheline starts here */+structrcu_headrcu;/* to avoid race on free */}____cacheline_internodealigned_in_smp;structice_tx_ring{
@@ -324,13 +327,16 @@ struct ice_tx_ring {/* stats structs */structice_q_statsstats;structu64_stats_syncsyncp;-structice_txq_statstx_stats;+structxdp_drv_stats*xdp_stats;/* CL3 - 3rd cacheline starts here */+structice_txq_statstx_stats;structrcu_headrcu;/* to avoid race on free */DECLARE_BITMAP(xps_state,ICE_TX_NBITS);/* XPS Config State */structice_channel*ch;structice_ptp_tx*tx_tstamps;++/* CL4 - 4th cacheline starts here */spinlock_ttx_lock;u32txq_teid;/* Added Tx queue TEID */#define ICE_TX_FLAGS_RING_XDP BIT(0)
@@ -507,6 +523,7 @@ int ice_clean_rx_irq_zc(struct ice_rx_ring *rx_ring, int budget){unsignedinttotal_rx_bytes=0,total_rx_packets=0;u16cleaned_count=ICE_DESC_UNUSED(rx_ring);+structxdp_rx_drv_stats_locallrstats={};structice_tx_ring*xdp_ring;unsignedintxdp_xmit=0;structbpf_prog*xdp_prog;
@@ -548,7 +565,8 @@ int ice_clean_rx_irq_zc(struct ice_rx_ring *rx_ring, int budget)xsk_buff_set_size(*xdp,size);xsk_buff_dma_sync_for_cpu(*xdp,rx_ring->xsk_pool);-xdp_res=ice_run_xdp_zc(rx_ring,*xdp,xdp_prog,xdp_ring);+xdp_res=ice_run_xdp_zc(rx_ring,*xdp,xdp_prog,xdp_ring,+&lrstats);if(xdp_res){if(xdp_res&(ICE_XDP_TX|ICE_XDP_REDIR))xdp_xmit|=xdp_res;
@@ -598,6 +616,7 @@ int ice_clean_rx_irq_zc(struct ice_rx_ring *rx_ring, int budget)ice_finalize_xdp_rx(xdp_ring,xdp_xmit);ice_update_rx_ring_stats(rx_ring,total_rx_packets,total_rx_bytes);+xdp_update_rx_drv_stats(&rx_ring->xdp_stats->xsk_rx,&lrstats);if(xsk_uses_need_wakeup(rx_ring->xsk_pool)){if(failure||rx_ring->next_to_clean==rx_ring->next_to_use)
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:43:37
Make igc driver collect and provide all generic XDP/XSK counters.
Unfortunately, igc has an unified ice_ring structure for both Rx
and Tx, so embedding xdp_drv_stats would bloat it for no good.
Store them in a separate array with a lifetime of an igc_adapter.
IGC_MAX_QUEUES is introduced purely for convenience to not hardcode
max(RX, TX) all the time.
Reuse all previously introduced helpers and
xdp_get_drv_stats_generic(). Performance wavering from incrementing
a bunch of counters on hotpath is around stddev at [64 ... 1532]
frame sizes.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
Reviewed-by: Michal Swiatkowski <redacted>
---
drivers/net/ethernet/intel/igc/igc.h | 3 +
drivers/net/ethernet/intel/igc/igc_main.c | 88 +++++++++++++++++++----
2 files changed, 77 insertions(+), 14 deletions(-)
@@ -2148,8 +2148,10 @@ static int igc_xdp_init_tx_descriptor(struct igc_ring *ring,u32cmd_type,olinfo_status;interr;-if(!igc_desc_unused(ring))+if(!igc_desc_unused(ring)){+xdp_update_tx_drv_full(&ring->xdp_stats->xdp_tx);return-EBUSY;+}buffer=&ring->tx_buffer_info[ring->next_to_use];err=igc_xdp_init_tx_buffer(buffer,xdpf,ring);
@@ -2214,36 +2216,51 @@ static int igc_xdp_xmit_back(struct igc_adapter *adapter, struct xdp_buff *xdp)/* This function assumes rcu_read_lock() is held by the caller. */staticint__igc_xdp_run_prog(structigc_adapter*adapter,structbpf_prog*prog,-structxdp_buff*xdp)+structxdp_buff*xdp,+structxdp_rx_drv_stats_local*lrstats){-u32act=bpf_prog_run_xdp(prog,xdp);+u32act;++lrstats->bytes+=xdp->data_end-xdp->data;+lrstats->packets++;+act=bpf_prog_run_xdp(prog,xdp);switch(act){caseXDP_PASS:+lrstats->pass++;returnIGC_XDP_PASS;caseXDP_TX:-if(igc_xdp_xmit_back(adapter,xdp)<0)+if(igc_xdp_xmit_back(adapter,xdp)<0){+lrstats->tx_errors++;gotoout_failure;+}+lrstats->tx++;returnIGC_XDP_TX;caseXDP_REDIRECT:-if(xdp_do_redirect(adapter->netdev,xdp,prog)<0)+if(xdp_do_redirect(adapter->netdev,xdp,prog)<0){+lrstats->redirect_errors++;gotoout_failure;+}+lrstats->redirect++;returnIGC_XDP_REDIRECT;-break;default:bpf_warn_invalid_xdp_action(act);-fallthrough;+lrstats->invalid++;+gotoout_failure;caseXDP_ABORTED:+lrstats->aborted++;out_failure:trace_xdp_exception(adapter->netdev,prog,act);-fallthrough;+returnIGC_XDP_CONSUMED;caseXDP_DROP:+lrstats->drop++;returnIGC_XDP_CONSUMED;}}staticstructsk_buff*igc_xdp_run_prog(structigc_adapter*adapter,-structxdp_buff*xdp)+structxdp_buff*xdp,+structxdp_rx_drv_stats_local*lrstats){structbpf_prog*prog;intres;
@@ -4385,6 +4416,8 @@ static int igc_alloc_q_vector(struct igc_adapter *adapter,ring->count=adapter->tx_ring_count;ring->queue_index=txr_idx;+ring->xdp_stats=adapter->netdev->xstats+txr_idx;+/* assign ring to adapter */adapter->tx_ring[txr_idx]=ring;
@@ -4407,6 +4440,8 @@ static int igc_alloc_q_vector(struct igc_adapter *adapter,ring->count=adapter->rx_ring_count;ring->queue_index=rxr_idx;+ring->xdp_stats=adapter->netdev->xstats+rxr_idx;+/* assign ring to adapter */adapter->rx_ring[rxr_idx]=ring;}
@@ -4515,6 +4550,7 @@ static int igc_sw_init(struct igc_adapter *adapter)structnet_device*netdev=adapter->netdev;structpci_dev*pdev=adapter->pdev;structigc_hw*hw=&adapter->hw;+u32i;pci_read_config_word(pdev,PCI_COMMAND,&hw->bus.pci_cmd_word);
@@ -4544,6 +4580,14 @@ static int igc_sw_init(struct igc_adapter *adapter)igc_init_queue_configuration(adapter);+netdev->xstats=kcalloc(IGC_MAX_QUEUES,sizeof(*netdev->xstats),+GFP_KERNEL);+if(!netdev->xstats)+return-ENOMEM;++for(i=0;i<IGC_MAX_QUEUES;i++)+xdp_init_drv_stats(netdev->xstats+i);+/* This call may decrease the number of queues */if(igc_init_interrupt_scheme(adapter,true)){netdev_err(netdev,"Unable to allocate memory for queues\n");
@@ -6046,11 +6090,25 @@ static int igc_xdp_xmit(struct net_device *dev, int num_frames,if(flags&XDP_XMIT_FLUSH)igc_flush_tx_descriptors(ring);+if(unlikely(drops))+xdp_update_tx_drv_err(&ring->xdp_stats->xdp_tx,drops);+__netif_tx_unlock(nq);returnnum_frames-drops;}+staticintigc_get_xdp_stats_nch(conststructnet_device*dev,u32attr_id)+{+switch(attr_id){+caseIFLA_XDP_XSTATS_TYPE_XDP:+caseIFLA_XDP_XSTATS_TYPE_XSK:+returnIGC_MAX_QUEUES;+default:+return-EOPNOTSUPP;+}+}+staticvoidigc_trigger_rxtxq_interrupt(structigc_adapter*adapter,structigc_q_vector*q_vector){
From: Alexander Lobakin <hidden> Date: 2021-11-23 16:44:37
Make igb driver collect and provide all generic XDP counters.
Unfortunately, igb has an unified ice_ring structure for both Rx
and Tx, so embedding xdp_drv_rx_stats would bloat it for no good.
Store XDP stats in a separate array with a lifetime of a netdev.
Unlike other Intel drivers, igb has no support for XSK, so we can't
use full xdp_drv_stats here. IGB_MAX_ALLOC_QUEUES is introduced
purely for convenience to not hardcode 16 twice more.
Reuse previously introduced helpers where possible. Performance
wavering from incrementing a bunch of counters on hotpath is around
stddev at [64 ... 1532] frame sizes.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
Reviewed-by: Michal Swiatkowski <redacted>
---
drivers/net/ethernet/intel/igb/igb.h | 14 ++-
drivers/net/ethernet/intel/igb/igb_main.c | 102 ++++++++++++++++++++--
2 files changed, 105 insertions(+), 11 deletions(-)
@@ -303,6 +303,11 @@ struct igb_rx_queue_stats {u64alloc_failed;};+structigb_xdp_stats{+structxdp_rx_drv_statsrx;+structxdp_tx_drv_statstx;+}____cacheline_aligned;+structigb_ring_container{structigb_ring*ring;/* pointer to linked list of rings */unsignedinttotal_bytes;/* total bytes processed this int */
@@ -1266,6 +1266,7 @@ static int igb_alloc_q_vector(struct igb_adapter *adapter,u64_stats_init(&ring->tx_syncp);u64_stats_init(&ring->tx_syncp2);+ring->xdp_stats=adapter->xdp_stats+txr_idx;/* assign ring to adapter */adapter->tx_ring[txr_idx]=ring;
@@ -1300,6 +1301,7 @@ static int igb_alloc_q_vector(struct igb_adapter *adapter,ring->queue_index=rxr_idx;u64_stats_init(&ring->rx_syncp);+ring->xdp_stats=adapter->xdp_stats+rxr_idx;/* assign ring to adapter */adapter->rx_ring[rxr_idx]=ring;
@@ -2973,6 +2975,9 @@ static int igb_xdp_xmit(struct net_device *dev, int n,nxmit++;}+if(unlikely(nxmit<n))+xdp_update_tx_drv_err(&tx_ring->xdp_stats->tx,n-nxmit);+__netif_tx_unlock(nq);if(unlikely(flags&XDP_XMIT_FLUSH))
@@ -2981,6 +2986,42 @@ static int igb_xdp_xmit(struct net_device *dev, int n,returnnxmit;}+staticintigb_get_xdp_stats_nch(conststructnet_device*dev,u32attr_id)+{+switch(attr_id){+caseIFLA_XDP_XSTATS_TYPE_XDP:+returnIGB_MAX_ALLOC_QUEUES;+default:+return-EOPNOTSUPP;+}+}++staticintigb_get_xdp_stats(conststructnet_device*dev,u32attr_id,+void*attr_data)+{+conststructigb_adapter*adapter=netdev_priv(dev);+conststructigb_xdp_stats*drv_iter=adapter->xdp_stats;+structifla_xdp_stats*iter=attr_data;+u32i;++switch(attr_id){+caseIFLA_XDP_XSTATS_TYPE_XDP:+break;+default:+return-EOPNOTSUPP;+}++for(i=0;i<IGB_MAX_ALLOC_QUEUES;i++){+xdp_fetch_rx_drv_stats(iter,&drv_iter->rx);+xdp_fetch_tx_drv_stats(iter,&drv_iter->tx);++drv_iter++;+iter++;+}++return0;+}+staticconststructnet_device_opsigb_netdev_ops={.ndo_open=igb_open,.ndo_stop=igb_close,
@@ -3962,6 +4007,7 @@ static int igb_sw_init(struct igb_adapter *adapter)structe1000_hw*hw=&adapter->hw;structnet_device*netdev=adapter->netdev;structpci_dev*pdev=adapter->pdev;+u32i;pci_read_config_word(pdev,PCI_COMMAND,&hw->bus.pci_cmd_word);
@@ -4019,6 +4065,19 @@ static int igb_sw_init(struct igb_adapter *adapter)if(!adapter->shadow_vfta)return-ENOMEM;+adapter->xdp_stats=kcalloc(IGB_MAX_ALLOC_QUEUES,+sizeof(*adapter->xdp_stats),+GFP_KERNEL);+if(!adapter->xdp_stats)+return-ENOMEM;++for(i=0;i<IGB_MAX_ALLOC_QUEUES;i++){+structigb_xdp_stats*xdp_stats=adapter->xdp_stats+i;++xdp_init_rx_drv_stats(&xdp_stats->rx);+xdp_init_tx_drv_stats(&xdp_stats->tx);+}+/* This call may decrease the number of queues */if(igb_init_interrupt_scheme(adapter,true)){dev_err(&pdev->dev,"Unable to allocate memory for queues\n");
@@ -6264,8 +6323,10 @@ int igb_xmit_xdp_ring(struct igb_adapter *adapter,len=xdpf->len;-if(unlikely(!igb_desc_unused(tx_ring)))+if(unlikely(!igb_desc_unused(tx_ring))){+xdp_update_tx_drv_full(&tx_ring->xdp_stats->tx);returnIGB_XDP_CONSUMED;+}dma=dma_map_single(tx_ring->dev,xdpf->data,len,DMA_TO_DEVICE);if(dma_mapping_error(tx_ring->dev,dma))
@@ -8677,6 +8759,7 @@ static int igb_clean_rx_irq(struct igb_q_vector *q_vector, const int budget){structigb_adapter*adapter=q_vector->adapter;structigb_ring*rx_ring=q_vector->rx.ring;+structxdp_rx_drv_stats_locallrstats={};structsk_buff*skb=rx_ring->skb;unsignedinttotal_bytes=0,total_packets=0;u16cleaned_count=igb_desc_unused(rx_ring);
@@ -8740,7 +8823,7 @@ static int igb_clean_rx_irq(struct igb_q_vector *q_vector, const int budget)/* At larger PAGE_SIZE, frame_sz depend on len size */xdp.frame_sz=igb_rx_frame_truesize(rx_ring,size);#endif-skb=igb_run_xdp(adapter,rx_ring,&xdp);+skb=igb_run_xdp(adapter,rx_ring,&xdp,&lrstats);}if(IS_ERR(skb)){
@@ -8814,6 +8897,7 @@ static int igb_clean_rx_irq(struct igb_q_vector *q_vector, const int budget)rx_ring->rx_stats.packets+=total_packets;rx_ring->rx_stats.bytes+=total_bytes;u64_stats_update_end(&rx_ring->rx_syncp);+xdp_update_rx_drv_stats(&rx_ring->xdp_stats->rx,&lrstats);q_vector->rx.total_packets+=total_packets;q_vector->rx.total_bytes+=total_bytes;--
From: Vladimir Oltean <vladimir.oltean@nxp.com> Date: 2021-11-23 17:09:27
On Tue, Nov 23, 2021 at 05:39:34PM +0100, Alexander Lobakin wrote:
Similarly to dpaa2, enetc stores 5 per-channel counters for XDP.
Add necessary callbacks to be able to access them using new generic
XDP stats infra.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
Reviewed-by: Vladimir Oltean <vladimir.oltean@nxp.com>
These counters can be dropped from ethtool, nobody depends on having
them there.
Side question: what does "nch" stand for?
Imho, the overall approach is way too bloated. I can see the packets/bytes but now we
have 3 counter updates with return codes included and then the additional sync of the
on-stack counters into the ring counters via xdp_update_rx_drv_stats(). So we now need
ice_update_rx_ring_stats() as well as xdp_update_rx_drv_stats() which syncs 10 different
stat counters via u64_stats_add() into the per ring ones. :/
I'm just taking our XDP L4LB in Cilium as an example: there we already count errors and
export them via per-cpu map that eventually lead to XDP_DROP cases including the /reason/
which caused the XDP_DROP (e.g. Prometheus can then scrape these insights from all the
nodes in the cluster). Given the different action codes are very often application specific,
there's not much debugging that you can do when /only/ looking at `ip link xdpstats` to
gather insight on *why* some of these actions were triggered (e.g. fib lookup failure, etc).
If really of interest, then maybe libxdp could have such per-action counters as opt-in in
its call chain..
In the case of ice_run_xdp() today, we already bump total_rx_bytes/total_rx_pkts under
XDP and update ice_update_rx_ring_stats(). I do see the case for XDP_TX and XDP_REDIRECT
where we run into driver-specific errors that are /outside of the reach/ of the BPF prog.
For example, we've been running into errors from XDP_TX in ice_xmit_xdp_ring() in the
past during testing, and were able to pinpoint the location as xdp_ring->tx_stats.tx_busy
was increasing. These things are useful and would make sense to standardize for XDP context.
But then it also seems like above in ice_xmit_xdp_ring() we now need to bump counters
twice just for sake of ethtool vs xdp counters which sucks a bit, would be nice to only
having to do it once:
> if (!unlikely(ICE_DESC_UNUSED(xdp_ring))) {
> + xdp_update_tx_drv_full(&xdp_ring->xdp_stats->xdp_tx);
> xdp_ring->tx_stats.tx_busy++;
> return ICE_XDP_CONSUMED;
> }
Anyway, but just to reiterate, for troubleshooting I do care about anomalous events that
led to drops in the driver e.g. due to no space in ring or DMA errors (XDP_TX), or more
detailed insights in xdp_do_redirect() when errors occur (XDP_REDIRECT), very much less
about the action code given the prog has the full error context here already.
One more comment/question on the last doc update patch (I presume you only have dummy
numbers in there from testing?):
+For some interfaces, standard XDP statistics are available.
+It can be accessed the same ways, e.g. `ip`::
+
+ $ ip link xdpstats dev enp178s0
+ 16: enp178s0:
+ xdp-channel0-rx_xdp_packets: 0
+ xdp-channel0-rx_xdp_bytes: 1
+ xdp-channel0-rx_xdp_errors: 2
What are the semantics on xdp_errors? Summary of xdp_redirect_errors, xdp_tx_errors and
xdp_xmit_errors? Or driver specific defined?
+ xdp-channel0-rx_xdp_aborted: 3
+ xdp-channel0-rx_xdp_drop: 4
+ xdp-channel0-rx_xdp_invalid: 5
+ xdp-channel0-rx_xdp_pass: 6
[...]
+ xdp-channel0-rx_xdp_redirect: 7
+ xdp-channel0-rx_xdp_redirect_errors: 8
+ xdp-channel0-rx_xdp_tx: 9
+ xdp-channel0-rx_xdp_tx_errors: 10
+ xdp-channel0-tx_xdp_xmit_packets: 11
+ xdp-channel0-tx_xdp_xmit_bytes: 12
+ xdp-channel0-tx_xdp_xmit_errors: 13
+ xdp-channel0-tx_xdp_xmit_full: 14
From a user PoV to avoid confusion, maybe should be made more clear that the latter refers
to xsk.
quoted hunk
@@ -507,6 +523,7 @@ int ice_clean_rx_irq_zc(struct ice_rx_ring *rx_ring, int budget) { unsigned int total_rx_bytes = 0, total_rx_packets = 0; u16 cleaned_count = ICE_DESC_UNUSED(rx_ring);+ struct xdp_rx_drv_stats_local lrstats = { }; struct ice_tx_ring *xdp_ring; unsigned int xdp_xmit = 0; struct bpf_prog *xdp_prog;
@@ -548,7 +565,8 @@ int ice_clean_rx_irq_zc(struct ice_rx_ring *rx_ring, int budget) xsk_buff_set_size(*xdp, size); xsk_buff_dma_sync_for_cpu(*xdp, rx_ring->xsk_pool);- xdp_res = ice_run_xdp_zc(rx_ring, *xdp, xdp_prog, xdp_ring);+ xdp_res = ice_run_xdp_zc(rx_ring, *xdp, xdp_prog, xdp_ring,+ &lrstats); if (xdp_res) { if (xdp_res & (ICE_XDP_TX | ICE_XDP_REDIR)) xdp_xmit |= xdp_res;
@@ -598,6 +616,7 @@ int ice_clean_rx_irq_zc(struct ice_rx_ring *rx_ring, int budget) ice_finalize_xdp_rx(xdp_ring, xdp_xmit); ice_update_rx_ring_stats(rx_ring, total_rx_packets, total_rx_bytes);+ xdp_update_rx_drv_stats(&rx_ring->xdp_stats->xsk_rx, &lrstats); if (xsk_uses_need_wakeup(rx_ring->xsk_pool)) { if (failure || rx_ring->next_to_clean == rx_ring->next_to_use)
From: Edward Cree <ecree.xilinx@gmail.com> Date: 2021-11-24 10:00:09
On 23/11/2021 16:39, Alexander Lobakin wrote:
Export 4 per-channel XDP counters for both sf100 and sfx drivers
using generic XDP stats infra.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
The usual Subject: prefix for these drivers is sfc:
(or occasionally sfc_ef100: for ef100-specific stuff).
From: "Russell King (Oracle)" <linux@armlinux.org.uk> Date: 2021-11-24 11:34:30
On Tue, Nov 23, 2021 at 05:39:37PM +0100, Alexander Lobakin wrote:
Same as mvneta, mvpp2 stores 7 XDP counters in per-cpu containers.
Expose them via generic XDP stats infra.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
Reviewed-by: Russell King (Oracle) <redacted>
Thanks!
--
RMK's Patch system: https://www.armlinux.org.uk/developer/patches/
FTTP is here! 40Mbps down 10Mbps up. Decent connectivity at last!
Actually, the only concern I have here is the duplication between this
function and mvpp2_get_xdp_stats(). It looks to me like these two
functions could share a lot of their code. Please submit a patch to
make that happen. Thanks.
--
RMK's Patch system: https://www.armlinux.org.uk/developer/patches/
FTTP is here! 40Mbps down 10Mbps up. Decent connectivity at last!
From: Alexander Lobakin <hidden> Date: 2021-11-24 11:38:48
From: Vladimir Oltean <vladimir.oltean@nxp.com>
Date: Tue, 23 Nov 2021 17:09:20 +0000
On Tue, Nov 23, 2021 at 05:39:34PM +0100, Alexander Lobakin wrote:
quoted
Similarly to dpaa2, enetc stores 5 per-channel counters for XDP.
Add necessary callbacks to be able to access them using new generic
XDP stats infra.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
Reviewed-by: Vladimir Oltean <vladimir.oltean@nxp.com>
Thanks!
These counters can be dropped from ethtool, nobody depends on having
them there.
Got it, thanks. I'll remove them in v3 or, in case v2 gets accepted,
will send a follow-up patch(es) for removing redundant Ethtool
stats.
Side question: what does "nch" stand for?
"The number of channels". I was thinking of an intuitial, but short
term, as get_xdp_stats_channels is too long and breaks Tab aligment
of tons of net_device_ops across the tree.
It was "nqs" /number of queues/ previously, but we usually use term
"queue" referring to one-direction ring, in case of these stats and
XDP in general "queue pair" or simply "channel" is more correct.
Thanks,
Al
Same comment as for mvpp2 - this could share a lot of code from
mvneta_ethtool_update_pcpu_stats() (although it means we end up
calculating a little more for the alloc error and refill error
that this API doesn't need) but I think sharing that code would be
a good idea.
--
RMK's Patch system: https://www.armlinux.org.uk/developer/patches/
FTTP is here! 40Mbps down 10Mbps up. Decent connectivity at last!
Daniel asked me to share my opinion, as Cloudflare has an XDP load
balancer as well.
On Wed, 24 Nov 2021 at 00:53, Daniel Borkmann [off-list ref] wrote:
I'm just taking our XDP L4LB in Cilium as an example: there we already count errors and
export them via per-cpu map that eventually lead to XDP_DROP cases including the /reason/
which caused the XDP_DROP (e.g. Prometheus can then scrape these insights from all the
nodes in the cluster). Given the different action codes are very often application specific,
there's not much debugging that you can do when /only/ looking at `ip link xdpstats` to
gather insight on *why* some of these actions were triggered (e.g. fib lookup failure, etc).
Agreed. For our purpose we often want to know whether a specific
program has been invoked. Per-channel or per device stats don't help
us much since we have a chain of programs (not using libxdp though).
My colleague Arthur has written xdpcap [1], which gives per-action,
per-program counters. This way we can correlate an action with a
packet and a program.
If really of interest, then maybe libxdp could have such per-action counters as opt-in in
its call chain..
We could also make it part of BPF_ENABLE_STATS, it's kind of coarse
grained though.
In the case of ice_run_xdp() today, we already bump total_rx_bytes/total_rx_pkts under
XDP and update ice_update_rx_ring_stats(). I do see the case for XDP_TX and XDP_REDIRECT
where we run into driver-specific errors that are /outside of the reach/ of the BPF prog.
For example, we've been running into errors from XDP_TX in ice_xmit_xdp_ring() in the
past during testing, and were able to pinpoint the location as xdp_ring->tx_stats.tx_busy
was increasing. These things are useful and would make sense to standardize for XDP context.
I'd like to see more tracepoints like trace_xdp_exception, personally.
We can use things like bpftrace for exploration and ebpf_exporter [2]
to generate alerts much more easily than something wired into
iproute2.
Best
Lorenz
1: https://github.com/cloudflare/xdpcap
2: https://github.com/cloudflare/ebpf_exporter
--
Lorenz Bauer | Systems Engineer
6th Floor, County Hall/The Riverside Building, SE1 7PB, UK
www.cloudflare.com
Imho, the overall approach is way too bloated. I can see the
packets/bytes but now we have 3 counter updates with return codes
included and then the additional sync of the on-stack counters into
the ring counters via xdp_update_rx_drv_stats(). So we now need
ice_update_rx_ring_stats() as well as xdp_update_rx_drv_stats() which
syncs 10 different stat counters via u64_stats_add() into the per ring
ones. :/
I'm just taking our XDP L4LB in Cilium as an example: there we already
count errors and export them via per-cpu map that eventually lead to
XDP_DROP cases including the /reason/ which caused the XDP_DROP (e.g.
Prometheus can then scrape these insights from all the nodes in the
cluster). Given the different action codes are very often application
specific, there's not much debugging that you can do when /only/
looking at `ip link xdpstats` to gather insight on *why* some of these
actions were triggered (e.g. fib lookup failure, etc). If really of
interest, then maybe libxdp could have such per-action counters as
opt-in in its call chain..
To me, standardising these counters is less about helping people debug
their XDP programs (as you say, you can put your own telemetry into
those), and more about making XDP less "mystical" to the system
administrator (who may not be the same person who wrote the XDP
programs). So at the very least, they need to indicate "where are the
packets going", which means at least counters for DROP, REDIRECT and TX
(+ errors for tx/redirect) in addition to the "processed by XDP" initial
counter. Which in the above means 'pass', 'invalid' and 'aborted' could
be dropped, I guess; but I don't mind terribly keeping them either given
that there's no measurable performance impact.
But then it also seems like above in ice_xmit_xdp_ring() we now need
to bump counters twice just for sake of ethtool vs xdp counters which
sucks a bit, would be nice to only having to do it once:
This I agree with, and while I can see the layering argument for putting
them into 'ip' and rtnetlink instead of ethtool, I also worry that these
counters will simply be lost in obscurity, so I do wonder if it wouldn't
be better to accept the "layering violation" and keeping them all in the
'ethtool -S' output?
[...]
+ xdp-channel0-rx_xdp_redirect: 7
+ xdp-channel0-rx_xdp_redirect_errors: 8
+ xdp-channel0-rx_xdp_tx: 9
+ xdp-channel0-rx_xdp_tx_errors: 10
+ xdp-channel0-tx_xdp_xmit_packets: 11
+ xdp-channel0-tx_xdp_xmit_bytes: 12
+ xdp-channel0-tx_xdp_xmit_errors: 13
+ xdp-channel0-tx_xdp_xmit_full: 14
From a user PoV to avoid confusion, maybe should be made more clear that the latter refers
to xsk.
+1, these should probably be xdp-channel0-tx_xsk_* or something like
that...
-Toke
Imho, the overall approach is way too bloated. I can see the
packets/bytes but now we have 3 counter updates with return codes
included and then the additional sync of the on-stack counters into
the ring counters via xdp_update_rx_drv_stats(). So we now need
ice_update_rx_ring_stats() as well as xdp_update_rx_drv_stats() which
syncs 10 different stat counters via u64_stats_add() into the per ring
ones. :/
I'm just taking our XDP L4LB in Cilium as an example: there we already
count errors and export them via per-cpu map that eventually lead to
XDP_DROP cases including the /reason/ which caused the XDP_DROP (e.g.
Prometheus can then scrape these insights from all the nodes in the
cluster). Given the different action codes are very often application
specific, there's not much debugging that you can do when /only/
looking at `ip link xdpstats` to gather insight on *why* some of these
actions were triggered (e.g. fib lookup failure, etc). If really of
interest, then maybe libxdp could have such per-action counters as
opt-in in its call chain..
To me, standardising these counters is less about helping people debug
their XDP programs (as you say, you can put your own telemetry into
those), and more about making XDP less "mystical" to the system
administrator (who may not be the same person who wrote the XDP
programs). So at the very least, they need to indicate "where are the
packets going", which means at least counters for DROP, REDIRECT and TX
(+ errors for tx/redirect) in addition to the "processed by XDP" initial
counter. Which in the above means 'pass', 'invalid' and 'aborted' could
be dropped, I guess; but I don't mind terribly keeping them either given
that there's no measurable performance impact.
Right.
The other reason is that I want to continue the effort of
standardizing widely-implemented statistics. Ethtool private stats
approach is neither scalable (you can't rely on any fields which may
be not exposed in other drivers) nor good for code hygiene (code
duplication, differences in naming and logics etc.).
Let's say if only mlx5 driver has 'cache_waive' stats, then it's
okay to export it using private stats, but if 10 drivers has
'xdp_drop' field it's better to uniform it, isn't it?
quoted
But then it also seems like above in ice_xmit_xdp_ring() we now need
to bump counters twice just for sake of ethtool vs xdp counters which
sucks a bit, would be nice to only having to do it once:
We'll remove such duplication in the nearest future, as well as some
of duplications between Ethtool private and this XDP stats. I wanted
this series to be as harmless as possible.
This I agree with, and while I can see the layering argument for putting
them into 'ip' and rtnetlink instead of ethtool, I also worry that these
counters will simply be lost in obscurity, so I do wonder if it wouldn't
be better to accept the "layering violation" and keeping them all in the
'ethtool -S' output?
I don't think we should harm the code and the logics in favor of
'some of the users can face something'. We don't control anything
related to XDP using Ethtool at all, but there is some XDP-related
stuff inside iproute2 code, so for me it's even more intuitive to
have them there.
Jakub, may be you'd like to add something at this point?
[...]
quoted
+ xdp-channel0-rx_xdp_redirect: 7
+ xdp-channel0-rx_xdp_redirect_errors: 8
+ xdp-channel0-rx_xdp_tx: 9
+ xdp-channel0-rx_xdp_tx_errors: 10
+ xdp-channel0-tx_xdp_xmit_packets: 11
+ xdp-channel0-tx_xdp_xmit_bytes: 12
+ xdp-channel0-tx_xdp_xmit_errors: 13
+ xdp-channel0-tx_xdp_xmit_full: 14
From a user PoV to avoid confusion, maybe should be made more clear that the latter refers
to xsk.
+1, these should probably be xdp-channel0-tx_xsk_* or something like
that...
I think I should expand this example in Docs a bit. For XSK, there's
a separate set of the same counters, and they differ as follows:
xdp-channel0-rx_xdp_packets: 0
xdp-channel0-rx_xdp_bytes: 1
xdp-channel0-rx_xdp_errors: 2
[ ... ]
xsk-channel0-rx_xdp_packets: 256
xsk-channel0-rx_xdp_bytes: 257
xsk-channel0-rx_xdp_errors: 258
[ ... ]
The only semantic difference is that 'tx_xdp_xmit' for XDP is a
counter for the packets gone through .ndo_xdp_xmit(), and in
case of XSK it's a counter for the packets gone through XSK
ZC xmit.
Same comment as for mvpp2 - this could share a lot of code from
mvneta_ethtool_update_pcpu_stats() (although it means we end up
calculating a little more for the alloc error and refill error
that this API doesn't need) but I think sharing that code would be
a good idea.
Ah, I didn't do that because in my first series I was removing
Ethtool counters at all. In this one, I left them as-is due to
some of folks hinted me that those counters (not specifically
on mvpp2 or mvneta, let's say on virtio-net or so) could have
already been used in some admin scripts somewhere in the world
(but with a TODO to figure out which driver I could remove them
in and do that).
It would be great if you know and would hint me if I could remove
those XDP-related Ethtool counters from Marvell drivers or not.
If so, I'll wipe them, otherwise just factor out common parts to
wipe out code duplication.
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-11-25 17:46:46
On Thu, 25 Nov 2021 18:07:08 +0100 Alexander Lobakin wrote:
quoted
This I agree with, and while I can see the layering argument for putting
them into 'ip' and rtnetlink instead of ethtool, I also worry that these
counters will simply be lost in obscurity, so I do wonder if it wouldn't
be better to accept the "layering violation" and keeping them all in the
'ethtool -S' output?
I don't think we should harm the code and the logics in favor of
'some of the users can face something'. We don't control anything
related to XDP using Ethtool at all, but there is some XDP-related
stuff inside iproute2 code, so for me it's even more intuitive to
have them there.
Jakub, may be you'd like to add something at this point?
TBH I wasn't following this thread too closely since I saw Daniel
nacked it already. I do prefer rtnl xstats, I'd just report them
in -s if they are non-zero. But doesn't sound like we have an agreement
whether they should exist or not.
Can we think of an approach which would make cloudflare and cilium
happy? Feels like we're trying to make the slightly hypothetical
admin happy while ignoring objections of very real users.
Please leave the per-channel stats out. They make a precedent for
channel stats which should be an attribute of a channel. Working for
a large XDP user for a couple of years now I can tell you from my own
experience I've not once found them useful. In fact per-queue stats are
a major PITA as they crowd the output.
From: Alexander Lobakin <hidden> Date: 2021-11-25 20:44:31
From: Jakub Kicinski <kuba@kernel.org>
Date: Thu, 25 Nov 2021 09:44:40 -0800
On Thu, 25 Nov 2021 18:07:08 +0100 Alexander Lobakin wrote:
quoted
quoted
This I agree with, and while I can see the layering argument for putting
them into 'ip' and rtnetlink instead of ethtool, I also worry that these
counters will simply be lost in obscurity, so I do wonder if it wouldn't
be better to accept the "layering violation" and keeping them all in the
'ethtool -S' output?
I don't think we should harm the code and the logics in favor of
'some of the users can face something'. We don't control anything
related to XDP using Ethtool at all, but there is some XDP-related
stuff inside iproute2 code, so for me it's even more intuitive to
have them there.
Jakub, may be you'd like to add something at this point?
TBH I wasn't following this thread too closely since I saw Daniel
nacked it already. I do prefer rtnl xstats, I'd just report them
in -s if they are non-zero. But doesn't sound like we have an agreement
whether they should exist or not.
Right, just -s is fine, if we drop the per-channel approach.
Can we think of an approach which would make cloudflare and cilium
happy? Feels like we're trying to make the slightly hypothetical
admin happy while ignoring objections of very real users.
The initial idea was to only uniform the drivers. But in general
you are right, 10 drivers having something doesn't mean it's
something good.
Maciej, I think you were talking about Cilium asking for those stats
in Intel drivers? Could you maybe provide their exact usecases/needs
so I'll orient myself? I certainly remember about XSK Tx packets and
bytes.
And speaking of XSK Tx, we have per-socket stats, isn't that enough?
Please leave the per-channel stats out. They make a precedent for
channel stats which should be an attribute of a channel. Working for
a large XDP user for a couple of years now I can tell you from my own
experience I've not once found them useful. In fact per-queue stats are
a major PITA as they crowd the output.
Oh okay. My very first iterations were without this, but then I
found most of the drivers expose their XDP stats per-channel. Since
I didn't plan to degrade the functionality, they went that way.
Al
From: Jakub Kicinski <kuba@kernel.org>
Date: Thu, 25 Nov 2021 09:44:40 -0800
quoted
On Thu, 25 Nov 2021 18:07:08 +0100 Alexander Lobakin wrote:
quoted
quoted
This I agree with, and while I can see the layering argument for putting
them into 'ip' and rtnetlink instead of ethtool, I also worry that these
counters will simply be lost in obscurity, so I do wonder if it wouldn't
be better to accept the "layering violation" and keeping them all in the
'ethtool -S' output?
I don't think we should harm the code and the logics in favor of
'some of the users can face something'. We don't control anything
related to XDP using Ethtool at all, but there is some XDP-related
stuff inside iproute2 code, so for me it's even more intuitive to
have them there.
Jakub, may be you'd like to add something at this point?
TBH I wasn't following this thread too closely since I saw Daniel
nacked it already. I do prefer rtnl xstats, I'd just report them
in -s if they are non-zero. But doesn't sound like we have an agreement
whether they should exist or not.
Right, just -s is fine, if we drop the per-channel approach.
I agree that adding them to -s is fine (and that resolves my "no one
will find them" complain as well). If it crowds the output we could also
default to only output'ing a subset, and have the more detailed
statistics hidden behind a verbose switch (or even just in the JSON
output)?
quoted
Can we think of an approach which would make cloudflare and cilium
happy? Feels like we're trying to make the slightly hypothetical
admin happy while ignoring objections of very real users.
The initial idea was to only uniform the drivers. But in general
you are right, 10 drivers having something doesn't mean it's
something good.
I don't think it's accurate to call the admin use case "hypothetical".
We're expending a significant effort explaining to people that XDP can
"eat" your packets, and not having any standard statistics makes this
way harder. We should absolutely cater to our "early adopters", but if
we want XDP to see wider adoption, making it "less weird" is critical!
Maciej, I think you were talking about Cilium asking for those stats
in Intel drivers? Could you maybe provide their exact usecases/needs
so I'll orient myself? I certainly remember about XSK Tx packets and
bytes.
And speaking of XSK Tx, we have per-socket stats, isn't that enough?
IMO, as long as the packets are accounted for in the regular XDP stats,
having a whole separate set of stats only for XSK is less important.
quoted
Please leave the per-channel stats out. They make a precedent for
channel stats which should be an attribute of a channel. Working for
a large XDP user for a couple of years now I can tell you from my own
experience I've not once found them useful. In fact per-queue stats are
a major PITA as they crowd the output.
Oh okay. My very first iterations were without this, but then I
found most of the drivers expose their XDP stats per-channel. Since
I didn't plan to degrade the functionality, they went that way.
I personally find the per-channel stats quite useful. One of the primary
reasons for not achieving full performance with XDP is broken
configuration of packet steering to CPUs, and having per-channel stats
is a nice way of seeing this. I can see the point about them being way
too verbose in the default output, though, and I do generally filter the
output as well when viewing them. But see my point above about only
printing a subset of the stats by default; per-channel stats could be
JSON-only, for instance?
-Toke
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-11-26 18:13:40
On Fri, 26 Nov 2021 13:30:16 +0100 Toke Høiland-Jørgensen wrote:
quoted
quoted
TBH I wasn't following this thread too closely since I saw Daniel
nacked it already. I do prefer rtnl xstats, I'd just report them
in -s if they are non-zero. But doesn't sound like we have an agreement
whether they should exist or not.
Right, just -s is fine, if we drop the per-channel approach.
I agree that adding them to -s is fine (and that resolves my "no one
will find them" complain as well). If it crowds the output we could also
default to only output'ing a subset, and have the more detailed
statistics hidden behind a verbose switch (or even just in the JSON
output)?
quoted
quoted
Can we think of an approach which would make cloudflare and cilium
happy? Feels like we're trying to make the slightly hypothetical
admin happy while ignoring objections of very real users.
The initial idea was to only uniform the drivers. But in general
you are right, 10 drivers having something doesn't mean it's
something good.
I don't think it's accurate to call the admin use case "hypothetical".
We're expending a significant effort explaining to people that XDP can
"eat" your packets, and not having any standard statistics makes this
way harder. We should absolutely cater to our "early adopters", but if
we want XDP to see wider adoption, making it "less weird" is critical!
Fair. In all honesty I said that hoping to push for a more flexible
approach hidden entirely in BPF, and not involving driver changes.
Assuming the XDP program has more fine grained stats we should be able
to extract those instead of double-counting. Hence my vague "let's work
with apps" comment.
For example to a person familiar with the workload it'd be useful to
know if program returned XDP_DROP because of configured policy or
failure to parse a packet. I don't think that sort distinction is
achievable at the level of standard stats.
The information required by the admin is higher level. As you say the
primary concern there is "how many packets did XDP eat".
Speaking of which, one thing that badly needs clarification is our
expectation around XDP packets getting counted towards the interface
stats.
quoted
Maciej, I think you were talking about Cilium asking for those stats
in Intel drivers? Could you maybe provide their exact usecases/needs
so I'll orient myself? I certainly remember about XSK Tx packets and
bytes.
And speaking of XSK Tx, we have per-socket stats, isn't that enough?
IMO, as long as the packets are accounted for in the regular XDP stats,
having a whole separate set of stats only for XSK is less important.
quoted
quoted
Please leave the per-channel stats out. They make a precedent for
channel stats which should be an attribute of a channel. Working for
a large XDP user for a couple of years now I can tell you from my own
experience I've not once found them useful. In fact per-queue stats are
a major PITA as they crowd the output.
Oh okay. My very first iterations were without this, but then I
found most of the drivers expose their XDP stats per-channel. Since
I didn't plan to degrade the functionality, they went that way.
I personally find the per-channel stats quite useful. One of the primary
reasons for not achieving full performance with XDP is broken
configuration of packet steering to CPUs, and having per-channel stats
is a nice way of seeing this.
Right, that's about the only thing I use it for as well. "Is the load
evenly distributed?" But that's not XDP specific and not worth
standardizing for, yet, IMO, because..
I can see the point about them being way too verbose in the default
output, though, and I do generally filter the output as well when
viewing them. But see my point above about only printing a subset of
the stats by default; per-channel stats could be JSON-only, for
instance?
we don't even know what constitutes a channel today. And that will
become increasingly problematic as importance of application specific
queues increases (zctap etc). IMO until the ontological gaps around
queues are filled we should leave per-queue stats in ethtool -S.
On Fri, 26 Nov 2021 13:30:16 +0100 Toke Høiland-Jørgensen wrote:
quoted
quoted
quoted
TBH I wasn't following this thread too closely since I saw Daniel
nacked it already. I do prefer rtnl xstats, I'd just report them
in -s if they are non-zero. But doesn't sound like we have an agreement
whether they should exist or not.
Right, just -s is fine, if we drop the per-channel approach.
I agree that adding them to -s is fine (and that resolves my "no one
will find them" complain as well). If it crowds the output we could also
default to only output'ing a subset, and have the more detailed
statistics hidden behind a verbose switch (or even just in the JSON
output)?
quoted
quoted
Can we think of an approach which would make cloudflare and cilium
happy? Feels like we're trying to make the slightly hypothetical
admin happy while ignoring objections of very real users.
The initial idea was to only uniform the drivers. But in general
you are right, 10 drivers having something doesn't mean it's
something good.
I don't think it's accurate to call the admin use case "hypothetical".
We're expending a significant effort explaining to people that XDP can
"eat" your packets, and not having any standard statistics makes this
way harder. We should absolutely cater to our "early adopters", but if
we want XDP to see wider adoption, making it "less weird" is critical!
Fair. In all honesty I said that hoping to push for a more flexible
approach hidden entirely in BPF, and not involving driver changes.
Assuming the XDP program has more fine grained stats we should be able
to extract those instead of double-counting. Hence my vague "let's work
with apps" comment.
For example to a person familiar with the workload it'd be useful to
know if program returned XDP_DROP because of configured policy or
failure to parse a packet. I don't think that sort distinction is
achievable at the level of standard stats.
The information required by the admin is higher level. As you say the
primary concern there is "how many packets did XDP eat".
Right, sure, I am also totally fine with having only a somewhat
restricted subset of stats available at the interface level and make
everything else be BPF-based. I'm hoping we can converge of a common
understanding of what this "minimal set" should be :)
Speaking of which, one thing that badly needs clarification is our
expectation around XDP packets getting counted towards the interface
stats.
Agreed. My immediate thought is that "XDP packets are interface packets"
but that is certainly not what we do today, so not sure if changing it
at this point would break things?
quoted
quoted
Maciej, I think you were talking about Cilium asking for those stats
in Intel drivers? Could you maybe provide their exact usecases/needs
so I'll orient myself? I certainly remember about XSK Tx packets and
bytes.
And speaking of XSK Tx, we have per-socket stats, isn't that enough?
IMO, as long as the packets are accounted for in the regular XDP stats,
having a whole separate set of stats only for XSK is less important.
quoted
quoted
Please leave the per-channel stats out. They make a precedent for
channel stats which should be an attribute of a channel. Working for
a large XDP user for a couple of years now I can tell you from my own
experience I've not once found them useful. In fact per-queue stats are
a major PITA as they crowd the output.
Oh okay. My very first iterations were without this, but then I
found most of the drivers expose their XDP stats per-channel. Since
I didn't plan to degrade the functionality, they went that way.
I personally find the per-channel stats quite useful. One of the primary
reasons for not achieving full performance with XDP is broken
configuration of packet steering to CPUs, and having per-channel stats
is a nice way of seeing this.
Right, that's about the only thing I use it for as well. "Is the load
evenly distributed?" But that's not XDP specific and not worth
standardizing for, yet, IMO, because..
quoted
I can see the point about them being way too verbose in the default
output, though, and I do generally filter the output as well when
viewing them. But see my point above about only printing a subset of
the stats by default; per-channel stats could be JSON-only, for
instance?
we don't even know what constitutes a channel today. And that will
become increasingly problematic as importance of application specific
queues increases (zctap etc). IMO until the ontological gaps around
queues are filled we should leave per-queue stats in ethtool -S.
Hmm, right, I see. I suppose that as long as the XDP packets show up in
one of the interface counters in ethtool -S, it's possible to answer the
load distribution issue, and any further debugging (say, XDP drops on a
certain queue due to CPU-based queue indexing on TX) can be delegated to
BPF-based tools...
-Toke
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-11-26 19:39:01
On Fri, 26 Nov 2021 19:47:17 +0100 Toke Høiland-Jørgensen wrote:
quoted
Fair. In all honesty I said that hoping to push for a more flexible
approach hidden entirely in BPF, and not involving driver changes.
Assuming the XDP program has more fine grained stats we should be able
to extract those instead of double-counting. Hence my vague "let's work
with apps" comment.
For example to a person familiar with the workload it'd be useful to
know if program returned XDP_DROP because of configured policy or
failure to parse a packet. I don't think that sort distinction is
achievable at the level of standard stats.
The information required by the admin is higher level. As you say the
primary concern there is "how many packets did XDP eat".
Right, sure, I am also totally fine with having only a somewhat
restricted subset of stats available at the interface level and make
everything else be BPF-based. I'm hoping we can converge of a common
understanding of what this "minimal set" should be :)
quoted
Speaking of which, one thing that badly needs clarification is our
expectation around XDP packets getting counted towards the interface
stats.
Agreed. My immediate thought is that "XDP packets are interface packets"
but that is certainly not what we do today, so not sure if changing it
at this point would break things?
I'd vote for taking the risk and trying to align all the drivers.
From: Daniel Borkmann <daniel@iogearbox.net> Date: 2021-11-26 22:32:06
On 11/26/21 7:06 PM, Jakub Kicinski wrote:
On Fri, 26 Nov 2021 13:30:16 +0100 Toke Høiland-Jørgensen wrote:
quoted
quoted
quoted
TBH I wasn't following this thread too closely since I saw Daniel
nacked it already. I do prefer rtnl xstats, I'd just report them
in -s if they are non-zero. But doesn't sound like we have an agreement
whether they should exist or not.
Right, just -s is fine, if we drop the per-channel approach.
I agree that adding them to -s is fine (and that resolves my "no one
will find them" complain as well). If it crowds the output we could also
default to only output'ing a subset, and have the more detailed
statistics hidden behind a verbose switch (or even just in the JSON
output)?
quoted
quoted
Can we think of an approach which would make cloudflare and cilium
happy? Feels like we're trying to make the slightly hypothetical
admin happy while ignoring objections of very real users.
The initial idea was to only uniform the drivers. But in general
you are right, 10 drivers having something doesn't mean it's
something good.
I don't think it's accurate to call the admin use case "hypothetical".
We're expending a significant effort explaining to people that XDP can
"eat" your packets, and not having any standard statistics makes this
way harder. We should absolutely cater to our "early adopters", but if
we want XDP to see wider adoption, making it "less weird" is critical!
Fair. In all honesty I said that hoping to push for a more flexible
approach hidden entirely in BPF, and not involving driver changes.
Assuming the XDP program has more fine grained stats we should be able
to extract those instead of double-counting. Hence my vague "let's work
with apps" comment.
For example to a person familiar with the workload it'd be useful to
know if program returned XDP_DROP because of configured policy or
failure to parse a packet. I don't think that sort distinction is
achievable at the level of standard stats.
Agree on the additional context. How often have you looked at tc clsact
/dropped/ stats specifically when you debug a more complex BPF program
there?
# tc -s qdisc show clsact dev foo
qdisc clsact ffff: parent ffff:fff1
Sent 6800 bytes 120 pkt (dropped 0, overlimits 0 requeues 0)
backlog 0b 0p requeues 0
Similarly, XDP_PASS counters may be of limited use as well for same reason
(and I think we might not even have a tc counter equivalent for it).
The information required by the admin is higher level. As you say the
primary concern there is "how many packets did XDP eat".
Agree. Above said, for XDP_DROP I would see one use case where you compare
different drivers or bond vs no bond as we did in the past in [0] when
testing against a packet generator (although I don't see bond driver covered
in this series here yet where it aggregates the XDP stats from all bond slave
devs).
On a higher-level wrt "how many packets did XDP eat", it would make sense
to have the stats for successful XDP_{TX,REDIRECT} given these are out
of reach from a BPF prog PoV - we can only count there how many times we
returned with XDP_TX but not whether the pkt /successfully made it/.
In terms of error cases, could we just standardize all drivers on the behavior
of e.g. mlx5e_xdp_handle(), meaning, a failure from XDP_{TX,REDIRECT} will
hit the trace_xdp_exception() and then fallthrough to bump a drop counter
(same as we bump in XDP_DROP then). So the drop counter will account for
program drops but also driver-related drops.
At some later point the trace_xdp_exception() could be extended with an error
code that the driver would propagate (given some of them look quite similar
across drivers, fwiw), and then whoever wants to do further processing with
them can do so via bpftrace or other tooling.
So overall wrt this series: from the lrstats we'd be /dropping/ the pass,
tx_errors, redirect_errors, invalid, aborted counters. And we'd be /keeping/
bytes & packets counters that XDP sees, (driver-)successful tx & redirect
counters as well as drop counter. Also, XDP bytes & packets counters should
not be counted twice wrt ethtool stats.
[0] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=9e2ee5c7e7c35d195e2aa0692a7241d47a433d1e
Thanks,
Daniel
From: Daniel Borkmann <daniel@iogearbox.net> Date: 2021-11-26 23:03:37
On 11/26/21 11:27 PM, Daniel Borkmann wrote:
On 11/26/21 7:06 PM, Jakub Kicinski wrote:
[...]
quoted
The information required by the admin is higher level. As you say the
primary concern there is "how many packets did XDP eat".
Agree. Above said, for XDP_DROP I would see one use case where you compare
different drivers or bond vs no bond as we did in the past in [0] when
testing against a packet generator (although I don't see bond driver covered
in this series here yet where it aggregates the XDP stats from all bond slave
devs).
On a higher-level wrt "how many packets did XDP eat", it would make sense
to have the stats for successful XDP_{TX,REDIRECT} given these are out
of reach from a BPF prog PoV - we can only count there how many times we
returned with XDP_TX but not whether the pkt /successfully made it/.
In terms of error cases, could we just standardize all drivers on the behavior
of e.g. mlx5e_xdp_handle(), meaning, a failure from XDP_{TX,REDIRECT} will
hit the trace_xdp_exception() and then fallthrough to bump a drop counter
(same as we bump in XDP_DROP then). So the drop counter will account for
program drops but also driver-related drops.
At some later point the trace_xdp_exception() could be extended with an error
code that the driver would propagate (given some of them look quite similar
across drivers, fwiw), and then whoever wants to do further processing with
them can do so via bpftrace or other tooling.
Just thinking out loud, one straight forward example we could start out with
that is also related to Paolo's series [1] ...
enum xdp_error {
XDP_UNKNOWN,
XDP_ACTION_INVALID,
XDP_ACTION_UNSUPPORTED,
};
... and then bpf_warn_invalid_xdp_action() returns one of the latter two
which we pass to trace_xdp_exception(). Later there could be XDP_DRIVER_*
cases e.g. propagated from XDP_TX error exceptions.
[...]
default:
err = bpf_warn_invalid_xdp_action(act);
fallthrough;
case XDP_ABORTED:
xdp_abort:
trace_xdp_exception(rq->netdev, prog, act, err);
fallthrough;
case XDP_DROP:
lrstats->xdp_drop++;
break;
}
[...]
[1] https://lore.kernel.org/netdev/cover.1637924200.git.pabeni@redhat.com/
+Petr, Nik
On Fri, Nov 26, 2021 at 11:14:31AM -0800, Jakub Kicinski wrote:
On Fri, 26 Nov 2021 19:47:17 +0100 Toke Høiland-Jørgensen wrote:
quoted
quoted
Fair. In all honesty I said that hoping to push for a more flexible
approach hidden entirely in BPF, and not involving driver changes.
Assuming the XDP program has more fine grained stats we should be able
to extract those instead of double-counting. Hence my vague "let's work
with apps" comment.
For example to a person familiar with the workload it'd be useful to
know if program returned XDP_DROP because of configured policy or
failure to parse a packet. I don't think that sort distinction is
achievable at the level of standard stats.
The information required by the admin is higher level. As you say the
primary concern there is "how many packets did XDP eat".
Right, sure, I am also totally fine with having only a somewhat
restricted subset of stats available at the interface level and make
everything else be BPF-based. I'm hoping we can converge of a common
understanding of what this "minimal set" should be :)
quoted
Speaking of which, one thing that badly needs clarification is our
expectation around XDP packets getting counted towards the interface
stats.
Agreed. My immediate thought is that "XDP packets are interface packets"
but that is certainly not what we do today, so not sure if changing it
at this point would break things?
I'd vote for taking the risk and trying to align all the drivers.
I agree. I think IFLA_STATS64 in RTM_NEWLINK should contain statistics
of all the packets seen by the netdev. The breakdown into software /
hardware / XDP should be reported via RTM_NEWSTATS.
Currently, for soft devices such as VLANs, bridges and GRE, user space
only sees statistics of packets forwarded by software, which is quite
useless when forwarding is offloaded from the kernel to hardware.
Petr is working on exposing hardware statistics for such devices via
rtnetlink. Unlike XDP (?), we need to be able to let user space enable /
disable hardware statistics as we have a limited number of hardware
counters and they can also reduce the bandwidth when enabled. We are
thinking of adding a new RTM_SETSTATS for that:
# ip stats set dev swp1 hw_stats on
For query, something like (under discussion):
# ip stats show dev swp1 // all groups
# ip stats show dev swp1 group link
# ip stats show dev swp1 group offload // all sub-groups
# ip stats show dev swp1 group offload sub-group cpu
# ip stats show dev swp1 group offload sub-group hw
Like other iproute2 commands, these follow the nesting of the
RTM_{NEW,GET}STATS uAPI.
Looking at patch #1 [1], I think that whatever you decide to expose for
XDP can be queried via:
# ip stats show dev swp1 group xdp
# ip stats show dev swp1 group xdp sub-group regular
# ip stats show dev swp1 group xdp sub-group xsk
Regardless, the following command should show statistics of all the
packets seen by the netdev:
# ip -s link show dev swp1
There is a PR [2] for node_exporter to use rtnetlink to fetch netdev
statistics instead of the old proc interface. It should be possible to
extend it to use RTM_*STATS for more fine-grained statistics.
[1] https://lore.kernel.org/netdev/20211123163955.154512-2-alexandr.lobakin@intel.com/
[2] https://github.com/prometheus/node_exporter/pull/2074
From: David Ahern <hidden> Date: 2021-11-28 22:26:39
On 11/23/21 9:39 AM, Alexander Lobakin wrote:
This is an almost complete rework of [0].
This series introduces generic XDP statistics infra based on rtnl
xstats (Ethtool standard stats previously), and wires up the drivers
which collect appropriate statistics to this new interface. Finally,
it introduces XDP/XSK statistics to all XDP-capable Intel drivers.
Those counters are:
* packets: number of frames passed to bpf_prog_run_xdp().
* bytes: number of bytes went through bpf_prog_run_xdp().
* errors: number of general XDP errors, if driver has one unified
counter.
* aborted: number of XDP_ABORTED returns.
* drop: number of XDP_DROP returns.
* invalid: number of returns of unallowed values (i.e. not XDP_*).
* pass: number of XDP_PASS returns.
* redirect: number of successfully performed XDP_REDIRECT requests.
* redirect_errors: number of failed XDP_REDIRECT requests.
* tx: number of successfully performed XDP_TX requests.
* tx_errors: number of failed XDP_TX requests.
* xmit_packets: number of successfully transmitted XDP/XSK frames.
* xmit_bytes: number of successfully transmitted XDP/XSK frames.
* xmit_errors: of XDP/XSK frames failed to transmit.
* xmit_full: number of XDP/XSK queue being full at the moment of
transmission.
To provide them, developers need to implement .ndo_get_xdp_stats()
and, if they want to expose stats on a per-channel basis,
.ndo_get_xdp_stats_nch(). include/net/xdp.h contains some helper
Why the tie to a channel? There are Rx queues and Tx queues and no
requirement to link them into a channel. It would be better (more
flexible) to allow them to be independent. Rather than ask the driver
"how many channels", ask 'how many Rx queues' and 'how many Tx queues'
for which xdp stats are reported.
From there, allow queue numbers or queue id's to be non-consecutive and
add a queue id or number as an attribute. e.g.,
[XDP stats]
[ Rx queue N]
counters
[ Tx queue N]
counters
This would allow a follow on patch set to do something like "Give me XDP
stats for Rx queue N" instead of doing a full dump.
On Fri, 26 Nov 2021 13:30:16 +0100 Toke Høiland-Jørgensen wrote:
quoted
quoted
quoted
TBH I wasn't following this thread too closely since I saw Daniel
nacked it already. I do prefer rtnl xstats, I'd just report them
in -s if they are non-zero. But doesn't sound like we have an agreement
whether they should exist or not.
Right, just -s is fine, if we drop the per-channel approach.
I agree that adding them to -s is fine (and that resolves my "no one
will find them" complain as well). If it crowds the output we could also
default to only output'ing a subset, and have the more detailed
statistics hidden behind a verbose switch (or even just in the JSON
output)?
quoted
quoted
Can we think of an approach which would make cloudflare and cilium
happy? Feels like we're trying to make the slightly hypothetical
admin happy while ignoring objections of very real users.
The initial idea was to only uniform the drivers. But in general
you are right, 10 drivers having something doesn't mean it's
something good.
I don't think it's accurate to call the admin use case "hypothetical".
We're expending a significant effort explaining to people that XDP can
"eat" your packets, and not having any standard statistics makes this
way harder. We should absolutely cater to our "early adopters", but if
we want XDP to see wider adoption, making it "less weird" is critical!
Fair. In all honesty I said that hoping to push for a more flexible
approach hidden entirely in BPF, and not involving driver changes.
Assuming the XDP program has more fine grained stats we should be able
to extract those instead of double-counting. Hence my vague "let's work
with apps" comment.
For example to a person familiar with the workload it'd be useful to
know if program returned XDP_DROP because of configured policy or
failure to parse a packet. I don't think that sort distinction is
achievable at the level of standard stats.
Agree on the additional context. How often have you looked at tc clsact
/dropped/ stats specifically when you debug a more complex BPF program
there?
# tc -s qdisc show clsact dev foo
qdisc clsact ffff: parent ffff:fff1
Sent 6800 bytes 120 pkt (dropped 0, overlimits 0 requeues 0)
backlog 0b 0p requeues 0
Similarly, XDP_PASS counters may be of limited use as well for same reason
(and I think we might not even have a tc counter equivalent for it).
quoted
The information required by the admin is higher level. As you say the
primary concern there is "how many packets did XDP eat".
Agree. Above said, for XDP_DROP I would see one use case where you compare
different drivers or bond vs no bond as we did in the past in [0] when
testing against a packet generator (although I don't see bond driver covered
in this series here yet where it aggregates the XDP stats from all bond slave
devs).
On a higher-level wrt "how many packets did XDP eat", it would make sense
to have the stats for successful XDP_{TX,REDIRECT} given these are out
of reach from a BPF prog PoV - we can only count there how many times we
returned with XDP_TX but not whether the pkt /successfully made it/.
In terms of error cases, could we just standardize all drivers on the behavior
of e.g. mlx5e_xdp_handle(), meaning, a failure from XDP_{TX,REDIRECT} will
hit the trace_xdp_exception() and then fallthrough to bump a drop counter
(same as we bump in XDP_DROP then). So the drop counter will account for
program drops but also driver-related drops.
At some later point the trace_xdp_exception() could be extended with an error
code that the driver would propagate (given some of them look quite similar
across drivers, fwiw), and then whoever wants to do further processing with
them can do so via bpftrace or other tooling.
So overall wrt this series: from the lrstats we'd be /dropping/ the pass,
tx_errors, redirect_errors, invalid, aborted counters. And we'd be /keeping/
bytes & packets counters that XDP sees, (driver-)successful tx & redirect
counters as well as drop counter. Also, XDP bytes & packets counters should
not be counted twice wrt ethtool stats.
This sounds reasonable to me, and I also like the error code to
tracepoint idea :)
-Toke
ena driver has 6 XDP counters collected per-channel. Add
callbacks
for getting the number of channels and those counters using
generic
XDP stats infra.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
drivers/net/ethernet/amazon/ena/ena_netdev.c | 53
++++++++++++++++++++
1 file changed, 53 insertions(+)
Hi,
thank you for the time you took in adding ENA support, this code
doesn't update the XDP TX queues (which only available when an XDP
program is loaded).
In theory the following patch should fix it, but I was unable to
compile your version of iproute2 and test the patch properly. Can
you please let me know if I need to do anything special to bring
up your version of iproute2 and test this patch?
The information required by the admin is higher level. As you say the
primary concern there is "how many packets did XDP eat".
Agree. Above said, for XDP_DROP I would see one use case where you
compare
different drivers or bond vs no bond as we did in the past in [0] when
testing against a packet generator (although I don't see bond driver
covered
in this series here yet where it aggregates the XDP stats from all
bond slave devs).
On a higher-level wrt "how many packets did XDP eat", it would make sense
to have the stats for successful XDP_{TX,REDIRECT} given these are out
of reach from a BPF prog PoV - we can only count there how many times we
returned with XDP_TX but not whether the pkt /successfully made it/.
Exactly.
quoted
In terms of error cases, could we just standardize all drivers on the
behavior
of e.g. mlx5e_xdp_handle(), meaning, a failure from XDP_{TX,REDIRECT} will
hit the trace_xdp_exception() and then fallthrough to bump a drop counter
(same as we bump in XDP_DROP then). So the drop counter will account for
program drops but also driver-related drops.
Hmm... I don't agree here. IMHO the BPF-program's *choice* to drop (via
XDP_DROP) should NOT share the counter with the driver-related drops.
The driver-related drops must be accounted separate.
For the record, I think mlx5e_xdp_handle() does the wrong thing, of
accounting everything as XDP_DROP in (rq->stats->xdp_drop++).
Current mlx5 driver stats are highly problematic actually.
Please don't model stats behavior after this driver.
E.g. if BPF-prog takes the *choice* XDP_TX or XDP_REDIRECT or XDP_DROP,
then the packet is invisible to "ifconfig" stats. It is like the driver
never received these packets (which is wrong IMHO). (The stats are only
avail via ethtool -S).
quoted
At some later point the trace_xdp_exception() could be extended with
an error
code that the driver would propagate (given some of them look quite
similar
across drivers, fwiw), and then whoever wants to do further processing
with them can do so via bpftrace or other tooling.
I do like trace_xdp_exception() is invoked in mlx5e_xdp_handle(), but do
notice that xdp_do_redirect() also have a tracepoint that can be used
for troubleshooting. (I usually use xdp_monitor for troubleshooting
which catch both).
I like the stats XDP handling better in mvneta_run_xdp().
Just thinking out loud, one straight forward example we could start out
with that is also related to Paolo's series [1] ...
enum xdp_error {
XDP_UNKNOWN,
XDP_ACTION_INVALID,
XDP_ACTION_UNSUPPORTED,
};
... and then bpf_warn_invalid_xdp_action() returns one of the latter two
which we pass to trace_xdp_exception(). Later there could be XDP_DRIVER_*
cases e.g. propagated from XDP_TX error exceptions.
[...]
default:
err = bpf_warn_invalid_xdp_action(act);
fallthrough;
case XDP_ABORTED:
xdp_abort:
trace_xdp_exception(rq->netdev, prog, act, err);
fallthrough;
case XDP_DROP:
lrstats->xdp_drop++;
break;
}
[...]
[1]
https://lore.kernel.org/netdev/cover.1637924200.git.pabeni@redhat.com/
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-11-29 14:50:19
On Sun, 28 Nov 2021 19:54:53 +0200 Ido Schimmel wrote:
quoted
quoted
Right, sure, I am also totally fine with having only a somewhat
restricted subset of stats available at the interface level and make
everything else be BPF-based. I'm hoping we can converge of a common
understanding of what this "minimal set" should be :)
Agreed. My immediate thought is that "XDP packets are interface packets"
but that is certainly not what we do today, so not sure if changing it
at this point would break things?
I'd vote for taking the risk and trying to align all the drivers.
I agree. I think IFLA_STATS64 in RTM_NEWLINK should contain statistics
of all the packets seen by the netdev. The breakdown into software /
hardware / XDP should be reported via RTM_NEWSTATS.
Hm, in the offload case "seen by the netdev" may be unclear. For
the offload case I believe our recommendation was phrased more like
"all packets which would be seen by the netdev if there was no
routing/tc offload", right?
Currently, for soft devices such as VLANs, bridges and GRE, user space
only sees statistics of packets forwarded by software, which is quite
useless when forwarding is offloaded from the kernel to hardware.
Petr is working on exposing hardware statistics for such devices via
rtnetlink. Unlike XDP (?), we need to be able to let user space enable /
disable hardware statistics as we have a limited number of hardware
counters and they can also reduce the bandwidth when enabled. We are
thinking of adding a new RTM_SETSTATS for that:
# ip stats set dev swp1 hw_stats on
Does it belong on the switch port? Not the netdev we want to track?
For query, something like (under discussion):
# ip stats show dev swp1 // all groups
# ip stats show dev swp1 group link
# ip stats show dev swp1 group offload // all sub-groups
# ip stats show dev swp1 group offload sub-group cpu
# ip stats show dev swp1 group offload sub-group hw
Like other iproute2 commands, these follow the nesting of the
RTM_{NEW,GET}STATS uAPI.
But we do have IFLA_STATS_LINK_OFFLOAD_XSTATS, isn't it effectively
the same use case?
Looking at patch #1 [1], I think that whatever you decide to expose for
XDP can be queried via:
# ip stats show dev swp1 group xdp
# ip stats show dev swp1 group xdp sub-group regular
# ip stats show dev swp1 group xdp sub-group xsk
Regardless, the following command should show statistics of all the
packets seen by the netdev:
# ip -s link show dev swp1
There is a PR [2] for node_exporter to use rtnetlink to fetch netdev
statistics instead of the old proc interface. It should be possible to
extend it to use RTM_*STATS for more fine-grained statistics.
[1] https://lore.kernel.org/netdev/20211123163955.154512-2-alexandr.lobakin@intel.com/
[2] https://github.com/prometheus/node_exporter/pull/2074
From: Petr Machata <petrm@nvidia.com> Date: 2021-11-29 15:53:29
Jakub Kicinski [off-list ref] writes:
On Sun, 28 Nov 2021 19:54:53 +0200 Ido Schimmel wrote:
quoted
quoted
quoted
Right, sure, I am also totally fine with having only a somewhat
restricted subset of stats available at the interface level and make
everything else be BPF-based. I'm hoping we can converge of a common
understanding of what this "minimal set" should be :)
Agreed. My immediate thought is that "XDP packets are interface packets"
but that is certainly not what we do today, so not sure if changing it
at this point would break things?
I'd vote for taking the risk and trying to align all the drivers.
I agree. I think IFLA_STATS64 in RTM_NEWLINK should contain statistics
of all the packets seen by the netdev. The breakdown into software /
hardware / XDP should be reported via RTM_NEWSTATS.
Hm, in the offload case "seen by the netdev" may be unclear. For
the offload case I believe our recommendation was phrased more like
"all packets which would be seen by the netdev if there was no
routing/tc offload", right?
Yes. The idea is to expose to Linux stats about traffic at conceptually
corresponding objects in the HW.
quoted
Currently, for soft devices such as VLANs, bridges and GRE, user space
only sees statistics of packets forwarded by software, which is quite
useless when forwarding is offloaded from the kernel to hardware.
Petr is working on exposing hardware statistics for such devices via
rtnetlink. Unlike XDP (?), we need to be able to let user space enable /
disable hardware statistics as we have a limited number of hardware
counters and they can also reduce the bandwidth when enabled. We are
thinking of adding a new RTM_SETSTATS for that:
# ip stats set dev swp1 hw_stats on
Does it belong on the switch port? Not the netdev we want to track?
Yes, it does, and is designed that way. That was just muscle memory
typing that "swp1" above :)
You would do e.g. "ip stats set dev swp1.200 hw_stats on" or, "dev br1",
or something like that.
quoted
For query, something like (under discussion):
# ip stats show dev swp1 // all groups
# ip stats show dev swp1 group link
# ip stats show dev swp1 group offload // all sub-groups
# ip stats show dev swp1 group offload sub-group cpu
# ip stats show dev swp1 group offload sub-group hw
Like other iproute2 commands, these follow the nesting of the
RTM_{NEW,GET}STATS uAPI.
But we do have IFLA_STATS_LINK_OFFLOAD_XSTATS, isn't it effectively
the same use case?
IFLA_STATS_LINK_OFFLOAD_XSTATS is a nest. Currently it carries just
CPU_HIT stats. The idea is to carry HW stats as well in that group.
quoted
Looking at patch #1 [1], I think that whatever you decide to expose for
XDP can be queried via:
# ip stats show dev swp1 group xdp
# ip stats show dev swp1 group xdp sub-group regular
# ip stats show dev swp1 group xdp sub-group xsk
Regardless, the following command should show statistics of all the
packets seen by the netdev:
# ip -s link show dev swp1
There is a PR [2] for node_exporter to use rtnetlink to fetch netdev
statistics instead of the old proc interface. It should be possible to
extend it to use RTM_*STATS for more fine-grained statistics.
[1] https://lore.kernel.org/netdev/20211123163955.154512-2-alexandr.lobakin@intel.com/
[2] https://github.com/prometheus/node_exporter/pull/2074
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-11-29 16:07:14
On Mon, 29 Nov 2021 16:51:02 +0100 Petr Machata wrote:
Jakub Kicinski [off-list ref] writes:
quoted
On Sun, 28 Nov 2021 19:54:53 +0200 Ido Schimmel wrote:
quoted
I agree. I think IFLA_STATS64 in RTM_NEWLINK should contain statistics
of all the packets seen by the netdev. The breakdown into software /
hardware / XDP should be reported via RTM_NEWSTATS.
Hm, in the offload case "seen by the netdev" may be unclear. For
the offload case I believe our recommendation was phrased more like
"all packets which would be seen by the netdev if there was no
routing/tc offload", right?
Yes. The idea is to expose to Linux stats about traffic at conceptually
corresponding objects in the HW.
Great.
quoted
quoted
Currently, for soft devices such as VLANs, bridges and GRE, user space
only sees statistics of packets forwarded by software, which is quite
useless when forwarding is offloaded from the kernel to hardware.
Petr is working on exposing hardware statistics for such devices via
rtnetlink. Unlike XDP (?), we need to be able to let user space enable /
disable hardware statistics as we have a limited number of hardware
counters and they can also reduce the bandwidth when enabled. We are
thinking of adding a new RTM_SETSTATS for that:
# ip stats set dev swp1 hw_stats on
Does it belong on the switch port? Not the netdev we want to track?
Yes, it does, and is designed that way. That was just muscle memory
typing that "swp1" above :)
You would do e.g. "ip stats set dev swp1.200 hw_stats on" or, "dev br1",
or something like that.
I see :)
quoted
quoted
For query, something like (under discussion):
# ip stats show dev swp1 // all groups
# ip stats show dev swp1 group link
# ip stats show dev swp1 group offload // all sub-groups
# ip stats show dev swp1 group offload sub-group cpu
# ip stats show dev swp1 group offload sub-group hw
Like other iproute2 commands, these follow the nesting of the
RTM_{NEW,GET}STATS uAPI.
But we do have IFLA_STATS_LINK_OFFLOAD_XSTATS, isn't it effectively
the same use case?
IFLA_STATS_LINK_OFFLOAD_XSTATS is a nest. Currently it carries just
CPU_HIT stats. The idea is to carry HW stats as well in that group.
Hm, the expectation was that the HW stats == total - SW. I believe that
still holds true for you, even if HW stats are not "complete" (e.g.
user enabled them after device was already forwarding for a while).
Is the concern about backward compat or such?
From: Petr Machata <petrm@nvidia.com> Date: 2021-11-29 17:10:34
Jakub Kicinski [off-list ref] writes:
On Mon, 29 Nov 2021 16:51:02 +0100 Petr Machata wrote:
quoted
Jakub Kicinski [off-list ref] writes:
quoted
On Sun, 28 Nov 2021 19:54:53 +0200 Ido Schimmel wrote:
quoted
For query, something like (under discussion):
# ip stats show dev swp1 // all groups
# ip stats show dev swp1 group link
# ip stats show dev swp1 group offload // all sub-groups
# ip stats show dev swp1 group offload sub-group cpu
# ip stats show dev swp1 group offload sub-group hw
Like other iproute2 commands, these follow the nesting of the
RTM_{NEW,GET}STATS uAPI.
But we do have IFLA_STATS_LINK_OFFLOAD_XSTATS, isn't it effectively
the same use case?
IFLA_STATS_LINK_OFFLOAD_XSTATS is a nest. Currently it carries just
CPU_HIT stats. The idea is to carry HW stats as well in that group.
Hm, the expectation was that the HW stats == total - SW. I believe that
still holds true for you, even if HW stats are not "complete" (e.g.
user enabled them after device was already forwarding for a while).
Is the concern about backward compat or such?
I guess you could call it backward compat. But not only. I think a
typical user doing "ip -s l sh", including various scripts, wants to see
the full picture and not worry what's going on where. Physical
netdevices already do that, and by extension bond and team of physical
netdevices. It also makes sense from the point of view of an offloaded
datapath as an implementation detail that you would ideally not notice.
For those who care to know about the offloaded datapath, it would be
nice to have the option to request either just the SW stats, or just the
HW stats. A logical place to put these would be under the OFFLOAD_XSTATS
nest of the RTM_GETSTATS message, but maybe the SW ones should be up
there next to IFLA_STATS_LINK_64. (After all it's going to be
independent from not only offload datapath, but also XDP.)
This way you get the intuitive default behavior, but still have a way to
e.g. request just the SW stats without hitting the HW, or just request
the HW stats if that's what you care about.
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-11-29 18:48:28
On Mon, 29 Nov 2021 14:59:53 +0100 Jesper Dangaard Brouer wrote:
Hmm... I don't agree here. IMHO the BPF-program's *choice* to drop (via
XDP_DROP) should NOT share the counter with the driver-related drops.
The driver-related drops must be accounted separate.
+1 FWIW. The Tx stat is a little misleading because it differs from the
definition of our other tx stats which mean _successfully_ transmitted
(and are accounted on the completion path in many drivers).
In the past I've used act_*, e.g. act_tx, to indicate the stat counts
returned actions, not whether the packet made it.
I still wonder whether it makes sense to count the stats per-action or
just have one "XDP consumed it" stat and that's it. The semantics of the
action are not of interest to the admin. A firewall can drop or tx
depending if it wants to send an ICMP reject or TCP RST message in
response. I need to know what the application does to understand the
difference, and if I do I can as well look at app stats. But I'm aware
I'm not going to find much support for this position, so just saying...
;)
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-11-29 20:40:50
On Mon, 29 Nov 2021 18:08:12 +0100 Petr Machata wrote:
Jakub Kicinski [off-list ref] writes:
quoted
On Mon, 29 Nov 2021 16:51:02 +0100 Petr Machata wrote:
quoted
IFLA_STATS_LINK_OFFLOAD_XSTATS is a nest. Currently it carries just
CPU_HIT stats. The idea is to carry HW stats as well in that group.
Hm, the expectation was that the HW stats == total - SW. I believe that
still holds true for you, even if HW stats are not "complete" (e.g.
user enabled them after device was already forwarding for a while).
Is the concern about backward compat or such?
I guess you could call it backward compat. But not only. I think a
typical user doing "ip -s l sh", including various scripts, wants to see
the full picture and not worry what's going on where. Physical
netdevices already do that, and by extension bond and team of physical
netdevices. It also makes sense from the point of view of an offloaded
datapath as an implementation detail that you would ideally not notice.
Agreed.
For those who care to know about the offloaded datapath, it would be
nice to have the option to request either just the SW stats, or just the
HW stats. A logical place to put these would be under the OFFLOAD_XSTATS
nest of the RTM_GETSTATS message, but maybe the SW ones should be up
there next to IFLA_STATS_LINK_64. (After all it's going to be
independent from not only offload datapath, but also XDP.)
What I'm getting at is that I thought IFLA_OFFLOAD_XSTATS_CPU_HIT
should be sufficient from uAPI perspective in terms of reporting.
User space can do the simple math to calculate the "SW stats" if
it wants to. We may well be talking about the same thing, so maybe
let's wait for the code?
This way you get the intuitive default behavior, but still have a way to
e.g. request just the SW stats without hitting the HW, or just request
the HW stats if that's what you care about.
Another thought on this patch: with individual attributes you could save
some overhead by not sending 0 counters to userspace. e.g., define a
helper that does:
static inline int nla_put_u64_if_set(struct sk_buff *skb, int attrtype,
u64 value, int padattr)
{
if (value)
return nla_put_u64_64bit(skb, attrtype, value, padattr);
return 0;
}
From: Petr Machata <petrm@nvidia.com> Date: 2021-11-30 11:57:24
Jakub Kicinski [off-list ref] writes:
On Mon, 29 Nov 2021 18:08:12 +0100 Petr Machata wrote:
quoted
For those who care to know about the offloaded datapath, it would be
nice to have the option to request either just the SW stats, or just the
HW stats. A logical place to put these would be under the OFFLOAD_XSTATS
nest of the RTM_GETSTATS message, but maybe the SW ones should be up
there next to IFLA_STATS_LINK_64. (After all it's going to be
independent from not only offload datapath, but also XDP.)
What I'm getting at is that I thought IFLA_OFFLOAD_XSTATS_CPU_HIT
should be sufficient from uAPI perspective in terms of reporting.
User space can do the simple math to calculate the "SW stats" if
it wants to. We may well be talking about the same thing, so maybe
let's wait for the code?
Ha, OK, now I understand. Yeah, CPU_HIT actually does fit the bill for
the traffic that took place in SW. We can reuse it.
I still think it would be better to report HW_STATS explicitly as well
though. One reason is simply convenience. The other is that OK, now we
have SW stats, and XDP stats, and total stats, and I (as a client) don't
necessarily know how it all fits together. But the contract for HW_STATS
is very clear.
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-11-30 15:14:21
On Tue, 30 Nov 2021 12:55:47 +0100 Petr Machata wrote:
I still think it would be better to report HW_STATS explicitly as well
though. One reason is simply convenience. The other is that OK, now we
have SW stats, and XDP stats, and total stats, and I (as a client) don't
necessarily know how it all fits together. But the contract for HW_STATS
is very clear.
Would be good to check with Jiri, my recollection is that this argument
was brought up when CPU_HIT stats were added. I don't recall the
reasoning.
<insert xkcd standards>
From: Alexander Lobakin <hidden> Date: 2021-11-30 15:56:57
From: Alexander Lobakin <redacted>
Date: Tue, 23 Nov 2021 17:39:29 +0100
Ok, open questions:
1. Channels vs queues vs global.
Jakub: no per-channel.
David (Ahern): it's worth it to separate as Rx/Tx.
Toke is fine with globals at the end I think?
My point was that for most of the systems we have 1:1 Rx:Tx
(usually num_online_cpus()), so asking drivers separately for
the number of RQs and then SQs would end up asking for the same
number twice.
But the main reason TBH was that most of the drivers store stats
on a per-channel basis and I didn't want them to regress in
functionality. I'm fine with reporting only netdev-wide if
everyone are.
In case if we keep per-channel: report per-channel only by request
and cumulative globals by default to not flood the output?
2. Count all errors as "drops" vs separately.
Daniel: account everything as drops, plus errors should be
reported as exceptions for tracing sub.
Jesper: we shouldn't mix drops and errors.
My point: we shouldn't, that's why there are patches for 2 drivers
to give errors a separate counter.
I provided an option either to report all errors together ('errors'
in stats structure) or to provide individual counters for each of
them (sonamed ctrs), but personally prefer detailed errors. However,
they might "go detailed" under trace_xdp_exception() only, sound
fine (OTOH in RTNL stats we have both "general" errors and detailed
error counters).
3. XDP and XSK ctrs separately or not.
My PoV is that those are two quite different worlds.
However, stats for actions on XSK really make a little sense since
99% of time we have xskmap redirect. So I think it'd be fine to just
expand stats structure with xsk_{rx,tx}_{packets,bytes} and count
the rest (actions, errors) together with XDP.
Rest:
- don't create a separate `ip` command and report under `-s`;
- save some RTNL skb space by skipping zeroed counters.
Also, regarding that I count all on the stack and then add to the
storage once in a polling cycle -- most drivers don't do that and
just increment the values in the storage directly, but this can be
less performant for frequently updated stats (or it's just my
embedded past).
Re u64 vs u64_stats_t -- the latter is more universal and
architecture-friendly, the former is used directly in most of the
drivers primarily because those drivers and the corresponding HW
are being run on 64-bit systems in the vast majority of cases, and
Ethtools stats themselves are not so critical to guard them with
anti-tearing. Anyways, local64_t is cheap on ARM64/x86_64 I guess?
Thanks,
Al
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-11-30 16:12:16
On Tue, 30 Nov 2021 16:56:12 +0100 Alexander Lobakin wrote:
3. XDP and XSK ctrs separately or not.
My PoV is that those are two quite different worlds.
However, stats for actions on XSK really make a little sense since
99% of time we have xskmap redirect. So I think it'd be fine to just
expand stats structure with xsk_{rx,tx}_{packets,bytes} and count
the rest (actions, errors) together with XDP.
Rest:
- don't create a separate `ip` command and report under `-s`;
- save some RTNL skb space by skipping zeroed counters.
Let me ruin this point of clarity for you. I think that stats should
be skipped when they are not collected (see ETHTOOL_STAT_NOT_SET).
If messages get large user should use the GETSTATS call and avoid
the problem more effectively.
Also, regarding that I count all on the stack and then add to the
storage once in a polling cycle -- most drivers don't do that and
just increment the values in the storage directly, but this can be
less performant for frequently updated stats (or it's just my
embedded past).
Re u64 vs u64_stats_t -- the latter is more universal and
architecture-friendly, the former is used directly in most of the
drivers primarily because those drivers and the corresponding HW
are being run on 64-bit systems in the vast majority of cases, and
Ethtools stats themselves are not so critical to guard them with
anti-tearing. Anyways, local64_t is cheap on ARM64/x86_64 I guess?
From: Alexander Lobakin <redacted>
Date: Tue, 23 Nov 2021 17:39:29 +0100
Ok, open questions:
1. Channels vs queues vs global.
Jakub: no per-channel.
David (Ahern): it's worth it to separate as Rx/Tx.
Toke is fine with globals at the end I think?
Well, I don't like throwing data away, so in that sense I do like
per-queue stats, but it's not a very strong preference (i.e., I can live
with either)...
My point was that for most of the systems we have 1:1 Rx:Tx
(usually num_online_cpus()), so asking drivers separately for
the number of RQs and then SQs would end up asking for the same
number twice.
But the main reason TBH was that most of the drivers store stats
on a per-channel basis and I didn't want them to regress in
functionality. I'm fine with reporting only netdev-wide if
everyone are.
In case if we keep per-channel: report per-channel only by request
and cumulative globals by default to not flood the output?
... however if we do go with per-channel stats I do agree that they
shouldn't be in the default output. I guess netlink could still split
them out and iproute2 could just sum them before display?
2. Count all errors as "drops" vs separately.
Daniel: account everything as drops, plus errors should be
reported as exceptions for tracing sub.
Jesper: we shouldn't mix drops and errors.
My point: we shouldn't, that's why there are patches for 2 drivers
to give errors a separate counter.
I provided an option either to report all errors together ('errors'
in stats structure) or to provide individual counters for each of
them (sonamed ctrs), but personally prefer detailed errors. However,
they might "go detailed" under trace_xdp_exception() only, sound
fine (OTOH in RTNL stats we have both "general" errors and detailed
error counters).
I agree it would be nice to have a separate error counter, but a single
counter is enough when combined with the tracepoints.
3. XDP and XSK ctrs separately or not.
My PoV is that those are two quite different worlds.
However, stats for actions on XSK really make a little sense since
99% of time we have xskmap redirect. So I think it'd be fine to just
expand stats structure with xsk_{rx,tx}_{packets,bytes} and count
the rest (actions, errors) together with XDP.
A whole set of separate counters for XSK is certainly overkill. No
strong preference as to whether they need a separate counter at all...
Rest:
- don't create a separate `ip` command and report under `-s`;
- save some RTNL skb space by skipping zeroed counters.
Also, regarding that I count all on the stack and then add to the
storage once in a polling cycle -- most drivers don't do that and
just increment the values in the storage directly, but this can be
less performant for frequently updated stats (or it's just my
embedded past).
Re u64 vs u64_stats_t -- the latter is more universal and
architecture-friendly, the former is used directly in most of the
drivers primarily because those drivers and the corresponding HW
are being run on 64-bit systems in the vast majority of cases, and
Ethtools stats themselves are not so critical to guard them with
anti-tearing. Anyways, local64_t is cheap on ARM64/x86_64 I guess?
I'm generally a fan of correctness first, so since you're touching all
the drivers anyway why I'd say go for u64_stats_t :)
-Toke
From: Alexander Lobakin <hidden> Date: 2021-11-30 16:35:23
From: Jakub Kicinski <kuba@kernel.org>
Date: Tue, 30 Nov 2021 08:12:07 -0800
On Tue, 30 Nov 2021 16:56:12 +0100 Alexander Lobakin wrote:
quoted
3. XDP and XSK ctrs separately or not.
My PoV is that those are two quite different worlds.
However, stats for actions on XSK really make a little sense since
99% of time we have xskmap redirect. So I think it'd be fine to just
expand stats structure with xsk_{rx,tx}_{packets,bytes} and count
the rest (actions, errors) together with XDP.
Rest:
- don't create a separate `ip` command and report under `-s`;
- save some RTNL skb space by skipping zeroed counters.
Let me ruin this point of clarity for you. I think that stats should
be skipped when they are not collected (see ETHTOOL_STAT_NOT_SET).
If messages get large user should use the GETSTATS call and avoid
the problem more effectively.
Well, it was Dave's thought here: [0]
Another thought on this patch: with individual attributes you could save
some overhead by not sending 0 counters to userspace. e.g., define a
helper that does:
I know about ETHTOOL_STAT_NOT_SET, but RTNL xstats doesn't use this,
does it?
GETSTATS is another thing, and I'll use it, thanks.
quoted
Also, regarding that I count all on the stack and then add to the
storage once in a polling cycle -- most drivers don't do that and
just increment the values in the storage directly, but this can be
less performant for frequently updated stats (or it's just my
embedded past).
Re u64 vs u64_stats_t -- the latter is more universal and
architecture-friendly, the former is used directly in most of the
drivers primarily because those drivers and the corresponding HW
are being run on 64-bit systems in the vast majority of cases, and
Ethtools stats themselves are not so critical to guard them with
anti-tearing. Anyways, local64_t is cheap on ARM64/x86_64 I guess?
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-11-30 17:05:07
On Tue, 30 Nov 2021 17:34:54 +0100 Alexander Lobakin wrote:
quoted
Another thought on this patch: with individual attributes you could save
some overhead by not sending 0 counters to userspace. e.g., define a
helper that does:
I know about ETHTOOL_STAT_NOT_SET, but RTNL xstats doesn't use this,
does it?
Not sure if you're asking me or Dave but no, to my knowledge RTNL does
not use such semantics today. But the reason is mostly because there
weren't many driver stats added there. Knowing if an error counter is
not supported or supporter and 0 is important for monitoring. Even if
XDP stats don't have a counter which may not be supported today it's
not a good precedent to make IMO.
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-11-30 17:07:25
On Tue, 30 Nov 2021 17:17:24 +0100 Toke Høiland-Jørgensen wrote:
quoted
1. Channels vs queues vs global.
Jakub: no per-channel.
David (Ahern): it's worth it to separate as Rx/Tx.
Toke is fine with globals at the end I think?
Well, I don't like throwing data away, so in that sense I do like
per-queue stats, but it's not a very strong preference (i.e., I can live
with either)...
We don't even have a clear definition of a queue in Linux.
As I said, adding this API today without a strong user and letting
drivers diverge in behavior would be a mistake.
From: David Ahern <hidden> Date: 2021-11-30 17:38:28
On 11/30/21 10:04 AM, Jakub Kicinski wrote:
On Tue, 30 Nov 2021 17:34:54 +0100 Alexander Lobakin wrote:
quoted
quoted
Another thought on this patch: with individual attributes you could save
some overhead by not sending 0 counters to userspace. e.g., define a
helper that does:
I know about ETHTOOL_STAT_NOT_SET, but RTNL xstats doesn't use this,
does it?
Not sure if you're asking me or Dave but no, to my knowledge RTNL does
not use such semantics today. But the reason is mostly because there
weren't many driver stats added there. Knowing if an error counter is
not supported or supporter and 0 is important for monitoring. Even if
XDP stats don't have a counter which may not be supported today it's
not a good precedent to make IMO.
Today, stats are sent as a struct so skipping stats whose value is 0 is
not an option. When using individual attributes for the counters this
becomes an option. Given there is no value in sending '0' why do it?
Is your pushback that there should be a uapi to opt-in to this behavior?
From: David Ahern <hidden> Date: 2021-11-30 17:46:08
On 11/30/21 8:56 AM, Alexander Lobakin wrote:
Rest:
- don't create a separate `ip` command and report under `-s`;
Reporting XDP stats under 'ip -s' is not going to be scalable from a
readability perspective.
ifstat (misc/ifstat.c) has support for extended stats which is where you
are adding these.
From: David Ahern <hidden> Date: 2021-11-30 17:56:33
On 11/30/21 10:07 AM, Jakub Kicinski wrote:
On Tue, 30 Nov 2021 17:17:24 +0100 Toke Høiland-Jørgensen wrote:
quoted
quoted
1. Channels vs queues vs global.
Jakub: no per-channel.
David (Ahern): it's worth it to separate as Rx/Tx.
Toke is fine with globals at the end I think?
Well, I don't like throwing data away, so in that sense I do like
per-queue stats, but it's not a very strong preference (i.e., I can live
with either)...
We don't even have a clear definition of a queue in Linux.
The summary above says "Jakub: no per-channel", and then you have this
comment about a clear definition of a queue. What is your preference
here, Jakub? I think I have gotten lost in all of the coments.
My request was just to not lump Rx and Tx together under a 'channel'
definition as a new API. Proposals like zctap and 'queues as a first
class citizen' are examples of intentions / desires to move towards Rx
and Tx queues beyond what exists today.
ena driver has 6 XDP counters collected per-channel. Add
callbacks
for getting the number of channels and those counters using
generic
XDP stats infra.
Signed-off-by: Alexander Lobakin <redacted>
Reviewed-by: Jesse Brandeburg <redacted>
---
drivers/net/ethernet/amazon/ena/ena_netdev.c | 53
++++++++++++++++++++
1 file changed, 53 insertions(+)
Hi,
thank you for the time you took in adding ENA support, this code
doesn't update the XDP TX queues (which only available when an XDP
program is loaded).
In theory the following patch should fix it, but I was unable to
compile your version of iproute2 and test the patch properly. Can
you please let me know if I need to do anything special to bring
up your version of iproute2 and test this patch?
Did you clone 'xdp_stats' branch? I've just rechecked on a freshly
cloned copy, works for me.
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-11-30 19:46:33
On Tue, 30 Nov 2021 10:38:14 -0700 David Ahern wrote:
On 11/30/21 10:04 AM, Jakub Kicinski wrote:
quoted
On Tue, 30 Nov 2021 17:34:54 +0100 Alexander Lobakin wrote:
quoted
I know about ETHTOOL_STAT_NOT_SET, but RTNL xstats doesn't use this,
does it?
Not sure if you're asking me or Dave but no, to my knowledge RTNL does
not use such semantics today. But the reason is mostly because there
weren't many driver stats added there. Knowing if an error counter is
not supported or supporter and 0 is important for monitoring. Even if
XDP stats don't have a counter which may not be supported today it's
not a good precedent to make IMO.
Today, stats are sent as a struct so skipping stats whose value is 0 is
not an option. When using individual attributes for the counters this
becomes an option. Given there is no value in sending '0' why do it?
To establish semantics of what it means that the statistic is not
reported. If we need to save space we'd need an extra attr with
a bitmap of "these stats were skipped because they were zero".
Or conversely some way of querying supported stats.
Is your pushback that there should be a uapi to opt-in to this behavior?
Not where I was going with it, but it is an option. If skipping 0s was
controlled by a flag a dump without such flag set would basically serve
as a way to query supported stats.
From: Jakub Kicinski <kuba@kernel.org> Date: 2021-11-30 19:53:25
On Tue, 30 Nov 2021 10:56:26 -0700 David Ahern wrote:
On 11/30/21 10:07 AM, Jakub Kicinski wrote:
quoted
On Tue, 30 Nov 2021 17:17:24 +0100 Toke Høiland-Jørgensen wrote:
quoted
Well, I don't like throwing data away, so in that sense I do like
per-queue stats, but it's not a very strong preference (i.e., I can live
with either)...
We don't even have a clear definition of a queue in Linux.
The summary above says "Jakub: no per-channel", and then you have this
comment about a clear definition of a queue. What is your preference
here, Jakub? I think I have gotten lost in all of the coments.
I'm against per-channel and against per-queue stats. I'm not saying "do
one instead of the other". Hope that makes it clear.
My request was just to not lump Rx and Tx together under a 'channel'
definition as a new API. Proposals like zctap and 'queues as a first
class citizen' are examples of intentions / desires to move towards Rx
and Tx queues beyond what exists today.
Right, and when we have the objects to control those we'll hang the
stats off them. Right now half of the NICs will destroy queue stats
on random reconfiguration requests, others will mix the stats between
queue instantiations.. mlx5 does it's shadow queue thing. It's a mess.
uAPI which is not portable and not usable in production is pure burden.