Re: [PATCH bpf RFC 1/4] xdp: rss hash types representation
From: Stanislav Fomichev <hidden>
Date: 2023-03-30 17:11:42
Also in:
bpf, intel-wired-lan
On 03/30, Jesper Dangaard Brouer wrote:
On 30/03/2023 01.19, Stanislav Fomichev wrote:quoted
On 03/29, Jesper Dangaard Brouer wrote:quoted
On 29/03/2023 19.18, Stanislav Fomichev wrote:quoted
On 03/29, Jesper Dangaard Brouer wrote:quoted
On 28/03/2023 23.58, Stanislav Fomichev wrote:quoted
On 03/28, Jesper Dangaard Brouer wrote:quoted
The RSS hash type specifies what portion of packet data NIChardware usedquoted
quoted
quoted
quoted
quoted
quoted
when calculating RSS hash value. The RSS types are focused onInternetquoted
quoted
quoted
quoted
quoted
quoted
traffic protocols at OSI layers L3 and L4. L2 (e.g. ARP)often get hashquoted
quoted
quoted
quoted
quoted
quoted
value zero and no RSS type. For L3 focused on IPv4 vs. IPv6,and L4quoted
quoted
quoted
quoted
quoted
quoted
primarily TCP vs UDP, but some hardware supports SCTP.quoted
Hardware RSS types are differently encoded for each hardwareNIC. Mostquoted
quoted
quoted
quoted
quoted
quoted
hardware represent RSS hash type as a number. Determining L3vs L4 oftenquoted
quoted
quoted
quoted
quoted
quoted
requires a mapping table as there often isn't a pattern orsortingquoted
quoted
quoted
quoted
quoted
quoted
according to ISO layer.quoted
The patch introduce a XDP RSS hash type (xdp_rss_hash_type)that can bothquoted
quoted
quoted
quoted
quoted
quoted
be seen as a number that is ordered according by ISO layer,and can be bitquoted
quoted
quoted
quoted
quoted
quoted
masked to separate IPv4 and IPv6 types for L4 protocols. Roomis availablequoted
quoted
quoted
quoted
quoted
quoted
for extending later while keeping these properties. This mapsand unifiesquoted
quoted
quoted
quoted
quoted
quoted
difference to hardware specific hashes.Looks good overall. Any reason we're making this specificlayout?quoted
quoted
quoted
quoted
One important goal is to have a simple/fast way to determining L3vs L4,quoted
quoted
quoted
quoted
because a L4 hash can be used for flow handling (e.g.load-balancing).quoted
quoted
quoted
quoted
We below layout you can:quoted
if (rss_type & XDP_RSS_TYPE_L4_MASK) bool hw_hash_do_LB = true;quoted
Or using it as a number:quoted
if (rss_type > XDP_RSS_TYPE_L4) bool hw_hash_do_LB = true;Why is it strictly better then the following? if (rss_type & (TYPE_UDP | TYPE_TCP | TYPE_SCTP)) {}quoted
See V2 I dropped the idea of this being a number (that idea was not a good idea).👍quoted
quoted
If we add some new L4 format, the bpf programs can be updated tosupportquoted
quoted
quoted
it?quoted
I'm very open to changes to my "specific" layout. I am in doubtifquoted
quoted
quoted
quoted
using it as a number is the right approach and worth the trouble.quoted
quoted
Why not simply the following? enum { ����XDP_RSS_TYPE_NONE = 0, ����XDP_RSS_TYPE_IPV4 = BIT(0), ����XDP_RSS_TYPE_IPV6 = BIT(1), ����/* IPv6 with extension header. */ ����/* let's note ^^^ it in the UAPI? */ ����XDP_RSS_TYPE_IPV6_EX = BIT(2), ����XDP_RSS_TYPE_UDP = BIT(3), ����XDP_RSS_TYPE_TCP = BIT(4), ����XDP_RSS_TYPE_SCTP = BIT(5),quoted
We know these bits for UDP, TCP, SCTP (and IPSEC) are exclusive,theyquoted
quoted
quoted
quoted
cannot be set at the same time, e.g. as a packet cannot both beUDP andquoted
quoted
quoted
quoted
TCP. Thus, using these bits as a number make sense to me, and ismorequoted
quoted
quoted
quoted
compact.
See below, why I'm wrong (in storing this as numbers).
quoted
quoted
quoted
[..]quoted
This BIT() approach also have the issue of extending it later(forwardquoted
quoted
quoted
quoted
compatibility). As mentioned a common task will be to check if hash-type is a L4 type. See mlx5 [patch 4/4] needed to extendwithquoted
quoted
quoted
quoted
IPSEC. Notice how my XDP_RSS_TYPE_L4_MASK covers all the bitsthat thisquoted
quoted
quoted
quoted
can be extended with new L4 types, such that existing progs willstillquoted
quoted
quoted
quoted
work checking for L4 check. It can of-cause be solved in thesame wayquoted
quoted
quoted
quoted
for this BIT() approach by reserving some bits upfront in a mask.We're using 6 bits out of 64, we should be good for awhile? If there is ever a forward compatibility issue, we can always come up with a new kfunc.quoted
I want/need store the RSS-type in the xdp_frame, for XDP_REDIRECT and SKB use-cases. Thus, I don't want to use 64-bit/8-bytes, as xdp_frame size is limited (given it reduces headroom expansion).quoted
quoted
One other related question I have is: should we export the type over some additional new kfunc argument? (instead of abusing thereturnquoted
quoted
quoted
type)quoted
Good question. I was also wondering if it wouldn't be better to add another kfunc argument with the rss_hash_type?quoted
That will change the call signature, so that will not be easy tohandlequoted
quoted
between kernel releases.Agree with Toke on a separate thread; might not be too late to fit it into an rc..quoted
quoted
Maybe that will let us drop the explicit BTF_TYPE_EMIT as well?quoted
Sure, if we define it as an argument, then it will automatically exported as BTF.quoted
quoted
quoted
quoted
} And then using XDP_RSS_TYPE_IPV4|XDP_RSS_TYPE_UDP vs XDP_RSS_TYPE_IPV6|XXX ?quoted
Do notice, that I already does some level of or'ing ("|") in this proposal. The main difference is that I hide this from thedriver, andquoted
quoted
quoted
quoted
kind of pre-combine the valid combination (enum's) drivers canselectquoted
quoted
quoted
quoted
from. I do get the point, and I think I will come up with acombinedquoted
quoted
quoted
quoted
solution based on your input.quoted
The RSS hashing types and combinations comes from M$ standards: [1]https://learn.microsoft.com/en-us/windows-hardware/drivers/network/rss-hashing-types#ipv4-hash-type-combinationsquoted
quoted
quoted
My main concern here is that we're over-complicating it with themasksquoted
quoted
quoted
and the format. With the explicit bits we can easily map to that spec you mention.quoted
See if you like my RFC-V2 proposal better. It should go more in your direction.Yeah, I like it better. Btw, why have a separate bit for XDP_RSS_BIT_EX?
Yes, we can rename the EX bit define (which is in V2). I reduced the name-length, because it allowed to keep code on-one-line when OR'ing.
quoted
Any reason it's not a XDP_RSS_L3_IPV6_EX within XDP_RSS_L3_MASK?
Hmm... I guess it belongs with L3.
Do notice that both IPv4 and IPv6 have a flexible header called either options/extensions headers, after their fixed header. (Mlx4 HW contains this info for IPv4, but I didn't extend xdp_rss_hash_type in that patch). Thus, we could have a single BIT that is valid for both IPv4 and IPv6. (This can help speedup packet parsing having this info).
A separate bit for both v4/v6 sounds good. But thinking more about it, not sure what the users are supposed to do with it. Whether the flow is hashed over the extension header should a config option, not a per-packet signal?
[...]quoted
quoted
quoted
For example, for forward compat, I'm not sure we can assume thatthe peoplequoted
quoted
quoted
will do: "rss_type & XDP_RSS_TYPE_L4_MASK" instead of something like: "rss_type & (XDP_RSS_TYPE_L4_IPV4_TCP|XDP_RSS_TYPE_L4_IPV4_UDP)"quoted
quoted
quoted
quoted
This code is allowed in V2 and should be. It is a choice of BPF-programmer in line-2 to not be forward compatible with newer L4 types.
The above code made me realize, I was wrong and you are right, we should represent the L4 types as BITs (and not as numbers). Even-though a single packet cannot be both UDP and TCP at the same time, then it is reasonable to have a code path that want to match both UDP and TCP. If L4 types are BITs then code can do a single compare (via ORing), while if they are numbers then we need more compares. Thus, I'll change scheme in V3 to use BITs.
So you are saying that the following: if (rss_type & (TCP|UDP) is much faster than the following: proto = rss_type & L4_MASK; if (proto == TCP || proto == UDP) ? idk, as long as we have enough bits to represent everything, I'm fine with either way, up to you. (not sure how much you want to constrain the data to fit it into xdp_frame; assuming u16 is fine?)
quoted
quoted
quoted
quoted
quoted
quoted
This proposal change the kfunc APIbpf_xdp_metadata_rx_hash() > > > > to return this RSS hash type on success.quoted
This is the real question (as also raised above)... Should we use return value or add an argument for type?Let's fix the prototype while it's still early in the rc?
Okay, in V3 I will propose adding an argument for the type then.
SG, thx!
quoted
Maybe also extend the tests to drop/decode/verify the mask?
Yes, I/we obviously need to update the selftests.
One problem with selftests is that it's using veth SKB-based mode, and SKB's have lost the RSS hash info and converted this into a single BIT telling us if this was L4 based. Thus, its hard to do some e.g. UDP type verification, but I guess we can check if expected UDP packet is RSS type L4.
Yeah, sounds fair.
In xdp_hw_metadata, I will add something that uses the RSS type bits. I was thinking to match against L4-UDP RSS type as program only AF_XDP redirect UDP packets, so we can verify it was a UDP packet by HW info.
Or maybe just dump it, idk.
--Jesper