[PATCH net-next v3 3/6] vxlan: vnifilter: bound the number of VNIs one request may touch
From: Ali Firas <hidden>
Date: 2026-09-27 21:54:01
Also in:
linux-kselftest, lkml
Subsystem:
networking drivers, the rest · Maintainers:
Andrew Lunn, "David S. Miller", Eric Dumazet, Jakub Kicinski, Paolo Abeni, Linus Torvalds
A single RTM_NEWTUNNEL or RTM_DELTUNNEL message can ask for the whole
24-bit space in one range. vxlan_vni_add_del() then walks it one VNI at
a time under rtnl_lock, holding the lock for the length of the walk.
That walk is the cost being bounded, not the memory: an add allocates a
VNI node and a per-CPU stats block per VNI, a delete creates neither,
but both walk the same span and hold rtnl the same way, so both are
bounded.
The span of one VXLAN_VNIFILTER_ENTRY is not the quantity to bound. An
entry carrying just START and END is 20 bytes on the wire and a message
may carry many of them, vxlan_process_vni_filter() being called once per
entry, so bounding each entry alone would still let one message ask for
thousands of times the limit. Sum the spans of every entry and reject
the message as a whole in vxlan_vnifilter_check_msg(), before the
dispatch loop: entries are applied and notified one at a time, so a
limit enforced during dispatch would return an error only after every
preceding entry had been acted on. The range extraction is factored
into vxlan_vni_filter_entry_range() so the count and the range later
acted on cannot drift.
The limit is 4096, following the usable VLAN ID space, since vnifilter
is mainly used on bridged devices where the VNI is derived from the
VLAN. It bounds one request, not how many VNIs a device may hold: the
whole space can still be installed in more than one request, and a
device holding it is still torn down in one step by
vxlan_vnigroup_uninit(), which is device teardown, not a netlink
message, and is not subject to this cap.
This is a policy narrowing of what a single message may ask for, not a
fix for a crash or corruption, and is not a backport candidate.
Suggested-by: Ido Schimmel <idosch@nvidia.com>
Assisted-by: LLM
Signed-off-by: Ali Firas <redacted>
---
Notes:
v3: was 2/5. Symmetric cap kept (bounds add and delete); changelog reframed around the rtnl walk, and the dump-range asymmetry moved to patch 4.
drivers/net/vxlan/vxlan_vnifilter.c | 89 ++++++++++++++++++++++++++---
1 file changed, 81 insertions(+), 8 deletions(-)
diff --git a/drivers/net/vxlan/vxlan_vnifilter.c b/drivers/net/vxlan/vxlan_vnifilter.c
index 9cffaf4998a9..13f4e115701a 100644
--- a/drivers/net/vxlan/vxlan_vnifilter.c
+++ b/drivers/net/vxlan/vxlan_vnifilter.c@@ -17,6 +17,15 @@ #include "vxlan_private.h" +/* Maximum number of VNIs one RTM_NEWTUNNEL or RTM_DELTUNNEL message may add or + * delete, summed over all of its VXLAN_VNIFILTER_ENTRY attributes. VNI + * filtering is mainly used on bridged VXLAN devices where the VNI is derived + * from the VLAN, so one request touching more VNIs than the VLAN ID space has + * no practical use, while an unbounded request walks the 24-bit space one VNI + * at a time under rtnl_lock. + */ +#define VXLAN_VNI_FILTER_MSG_MAX 4096 + static inline int vxlan_vni_cmp(struct rhashtable_compare_arg *arg, const void *ptr) {
@@ -846,12 +855,78 @@ static int vxlan_vni_add_del(struct vxlan_dev *vxlan, __u32 start_vni, return err; } +/* Derive the VNI range one VXLAN_VNIFILTER_ENTRY selects. Shared so that the + * count taken by vxlan_vnifilter_check_msg() cannot drift from the range + * vxlan_process_vni_filter() then acts on. + */ +static void vxlan_vni_filter_entry_range(struct nlattr **vattrs, u32 *vni_start, + u32 *vni_end) +{ + *vni_start = 0; + *vni_end = 0; + + if (vattrs[VXLAN_VNIFILTER_ENTRY_START]) { + *vni_start = nla_get_u32(vattrs[VXLAN_VNIFILTER_ENTRY_START]); + *vni_end = *vni_start; + } + + if (vattrs[VXLAN_VNIFILTER_ENTRY_END]) + *vni_end = nla_get_u32(vattrs[VXLAN_VNIFILTER_ENTRY_END]); +} + +/* Reject a request touching more than VXLAN_VNI_FILTER_MSG_MAX VNIs before any + * of its entries is acted on. Both add and delete walk the span one VNI at a + * time under rtnl_lock, so both are bounded. Entries are applied one at a time + * and each one notifies as it goes, so a limit checked inside the dispatch loop + * would leave the entries ahead of the offending one already applied. + */ +static int vxlan_vnifilter_check_msg(const struct nlmsghdr *nlh, + struct netlink_ext_ack *extack) +{ + struct nlattr *vattrs[VXLAN_VNIFILTER_ENTRY_MAX + 1]; + struct nlattr *attr; + u32 vnis = 0; + int err, rem; + + nlmsg_for_each_attr_type(attr, VXLAN_VNIFILTER_ENTRY, nlh, + sizeof(struct tunnel_msg), rem) { + u32 vni_start, vni_end; + + err = nla_parse_nested(vattrs, VXLAN_VNIFILTER_ENTRY_MAX, attr, + vni_filter_entry_policy, extack); + if (err) + return err; + + vxlan_vni_filter_entry_range(vattrs, &vni_start, &vni_end); + + /* A start above the end selects no VNI and costs nothing; + * leave it behaving as it does today. + */ + if (vni_end < vni_start) + continue; + + /* vni_filter_entry_policy has already bounded both endpoints + * below VXLAN_N_VID, so one entry adds at most VXLAN_N_VID and + * vnis cannot wrap before the test below rejects it. + */ + vnis += vni_end - vni_start + 1; + if (vnis > VXLAN_VNI_FILTER_MSG_MAX) { + NL_SET_ERR_MSG_ATTR_FMT(extack, attr, + "Request asks for more than %u VNIs", + VXLAN_VNI_FILTER_MSG_MAX); + return -EINVAL; + } + } + + return 0; +} + static int vxlan_process_vni_filter(struct vxlan_dev *vxlan, struct nlattr *nlvnifilter, int cmd, struct netlink_ext_ack *extack) { struct nlattr *vattrs[VXLAN_VNIFILTER_ENTRY_MAX + 1]; - u32 vni_start = 0, vni_end = 0; + u32 vni_start, vni_end; union vxlan_addr group; int err;
@@ -862,13 +937,7 @@ static int vxlan_process_vni_filter(struct vxlan_dev *vxlan, if (err) return err; - if (vattrs[VXLAN_VNIFILTER_ENTRY_START]) { - vni_start = nla_get_u32(vattrs[VXLAN_VNIFILTER_ENTRY_START]); - vni_end = vni_start; - } - - if (vattrs[VXLAN_VNIFILTER_ENTRY_END]) - vni_end = nla_get_u32(vattrs[VXLAN_VNIFILTER_ENTRY_END]); + vxlan_vni_filter_entry_range(vattrs, &vni_start, &vni_end); if (!vni_start && !vni_end) { NL_SET_ERR_MSG_ATTR(extack, nlvnifilter,
@@ -975,6 +1044,10 @@ static int vxlan_vnifilter_process(struct sk_buff *skb, struct nlmsghdr *nlh, if (!(vxlan->cfg.flags & VXLAN_F_VNIFILTER)) return -EOPNOTSUPP; + err = vxlan_vnifilter_check_msg(nlh, extack); + if (err) + return err; + nlmsg_for_each_attr_type(attr, VXLAN_VNIFILTER_ENTRY, nlh, sizeof(*tmsg), rem) { err = vxlan_process_vni_filter(vxlan, attr, nlh->nlmsg_type,
--
2.53.0