"ip maddr show" prints link-layer, IPv4 and IPv6 entries. The IPv4 and
IPv6 ones come over netlink today, with IFA_MC_USERS for the user count.
The link-layer list, dev->mc, is only exported via /proc/net/dev_mcast,
so iproute2 and other users still parse procfs for it.
This series adds the AF_PACKET family to RTM_GETMULTICAST so that the
link-layer multicast addresses are reported the same way. Patch 2 dumps
dev->mc in the ifaddrmsg format of the IPv4 and IPv6 dumps: IFA_MULTICAST
carries the address, IFA_MC_USERS the reference count and a new
IFA_F_GLOBAL flag marks entries added explicitly, the "static" column of
/proc/net/dev_mcast. ifa_index limits the dump to one device and
IFA_TARGET_NETNSID selects another netns. Patches 1 and 3 update the
rt-addr spec and patch 4 adds a selftest.
AF_PACKET dumps returned -EOPNOTSUPP before, so iproute2 can keep the
procfs fallback for older kernels. The iproute2 side is ready and will
be posted once this is in.
Changes in v6:
- Move the dump to net/core/dev_addr_lists.c next to the dev->mc
helpers
- Stamp cb->seq from dev_base_seq and check dump consistency
- Let an unexpected dump error propagate instead of skipping
- Check for an address only present in the peer netns in the
target-netnsid selftest, the peer and local ifindex can be equal
- Document the ifa-index filter and the zero header fields in the spec
Changes in v5:
- Filter the target-netnsid selftest dump by ifa-index, a new netns
also contains the fallback tunnel devices
Changes in v4:
- Reset the resume offset when the device the dump stopped at is gone
- Use a tracked netns reference (put_net_track)
- Say ifa-family must be set in the spec doc, drop the AF_UNSPEC remark
- Close the netlink socket and guard the checks in the selftest
- Drop the Fixes tag
Changes in v3:
- Report the static bit as a new IFA_F_GLOBAL flag in IFA_FLAGS
instead of IFA_F_PERMANENT
- Support IFA_TARGET_NETNSID and test it
- Describe global_use accurately, it is also set by dev_mc_add_excl()
- Fix the target-netnsid type in the rt-addr spec, as its own patch
Changes in v2:
- Always validate the request header, not only with strict checking
- Use a single "with" statement in the selftest (ruff)
Yuyang Huang (4):
netlink: specs: rt-addr: fix the type of target-netnsid
net: add AF_PACKET multicast dumps
netlink: specs: rt-addr: document AF_PACKET multicast dumps
selftests: net: test AF_PACKET multicast dumps
Documentation/netlink/specs/rt-addr.yaml | 19 ++-
include/linux/netdevice.h | 1 +
include/uapi/linux/if_addr.h | 1 +
net/core/dev_addr_lists.c | 177 +++++++++++++++++++++++
net/core/rtnetlink.c | 2 +
tools/testing/selftests/net/rtnetlink.py | 73 +++++++++-
6 files changed, 268 insertions(+), 5 deletions(-)
--
2.43.0
The kernel parses IFA_TARGET_NETNSID as NLA_S32 and rt-link.yaml
declares its target-netnsid as s32, but rt-addr.yaml has it as binary.
Signed-off-by: Yuyang Huang <redacted>
Reviewed-by: Nicolas Dichtel <redacted>
---
Documentation/netlink/specs/rt-addr.yaml | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
RTM_GETMULTICAST dumps IPv4 and IPv6 multicast group memberships, but
the device multicast list (dev->mc) is only available through
/proc/net/dev_mcast, so "ip maddr show" still has to parse procfs for
its link-layer entries.
Handle RTM_GETMULTICAST dumps with ifa_family set to AF_PACKET next to
the dev->mc helpers in dev_addr_lists.c and report every entry of
dev->mc in the existing ifaddrmsg format:
- IFA_MULTICAST carries the raw link-layer address
- IFA_MC_USERS carries the entry reference count
- IFA_F_GLOBAL in IFA_FLAGS reports netdev_hw_addr::global_use, set
by dev_mc_add_global() (SIOCADDMULTI) and dev_mc_add_excl()
("bridge fdb add ... self"), i.e. entries added explicitly rather
than by a protocol join. This is the static column of
/proc/net/dev_mcast
- ifa_scope is RT_SCOPE_LINK
This covers every column of /proc/net/dev_mcast. AF_PACKET is the
family iproute2 already uses for link-layer addresses ("ip -0").
The default FDB dump also walks dev->mc, but only for Ethernet devices
without an ndo_fdb_dump of their own, so bridge, vxlan or macvlan
devices never show their multicast filter there, and it has no users
count or global_use bit. Extending it would change "bridge fdb show"
output and add NDA_* attributes.
Requests are always validated, there are no legacy users: prefixlen,
flags and scope must be zero, ifa_index selects one device and
IFA_TARGET_NETNSID is the only attribute accepted. The dump runs under
RCU and netif_addr_lock_bh() without RTNL, and stamps cb->seq from
dev_base_seq so a device added or removed between dump rounds sets
NLM_F_DUMP_INTR.
Signed-off-by: Yuyang Huang <redacted>
Reviewed-by: Nicolas Dichtel <redacted>
---
include/linux/netdevice.h | 1 +
include/uapi/linux/if_addr.h | 1 +
net/core/dev_addr_lists.c | 177 +++++++++++++++++++++++++++++++++++
net/core/rtnetlink.c | 2 +
4 files changed, 181 insertions(+)
@@ -1180,6 +1184,179 @@ void dev_mc_init(struct net_device *dev)}EXPORT_SYMBOL(dev_mc_init);+staticintdev_mc_fill_addr(structsk_buff*skb,conststructnet_device*dev,+conststructnetdev_hw_addr*ha,u32portid,+u32seq,unsignedintflags,intnetnsid)+{+u32ifa_flags=ha->global_use?IFA_F_GLOBAL:0;+structifaddrmsg*ifm;+structnlmsghdr*nlh;++nlh=nlmsg_put(skb,portid,seq,RTM_GETMULTICAST,sizeof(*ifm),+flags);+if(!nlh)+return-EMSGSIZE;++ifm=nlmsg_data(nlh);+ifm->ifa_family=AF_PACKET;+ifm->ifa_prefixlen=0;+/* ifm->ifa_flags holds 8 bits, the full value is in IFA_FLAGS */+ifm->ifa_flags=(__u8)ifa_flags;+ifm->ifa_scope=RT_SCOPE_LINK;+ifm->ifa_index=dev->ifindex;++if((netnsid>=0&&+nla_put_s32(skb,IFA_TARGET_NETNSID,netnsid))||+nla_put(skb,IFA_MULTICAST,dev->addr_len,ha->addr)||+nla_put_u32(skb,IFA_MC_USERS,ha->refcount)||+nla_put_u32(skb,IFA_FLAGS,ifa_flags)){+nlmsg_cancel(skb,nlh);+return-EMSGSIZE;+}++nlmsg_end(skb,nlh);+return0;+}++staticintdev_mc_dump_dev(structnet_device*dev,structsk_buff*skb,+structnetlink_callback*cb,int*s_addr_idx,+unsignedintflags,intnetnsid)+{+structnetdev_hw_addr*ha;+intaddr_idx=0;+interr=0;++netif_addr_lock_bh(dev);+netdev_for_each_mc_addr(ha,dev){+if(addr_idx<*s_addr_idx){+addr_idx++;+continue;+}+err=dev_mc_fill_addr(skb,dev,ha,NETLINK_CB(cb->skb).portid,+cb->nlh->nlmsg_seq,flags,netnsid);+if(err<0)+break;+nl_dump_check_consistent(cb,nlmsg_hdr(skb));+addr_idx++;+}+netif_addr_unlock_bh(dev);++*s_addr_idx=err<0?addr_idx:0;++returnerr;+}++structdev_mc_dump_filter{+structnet*tgt_net;+netns_trackerns_tracker;+intnetnsid;+intifindex;+};++staticconststructnla_policydev_mc_dump_policy[IFA_MAX+1]={+[IFA_TARGET_NETNSID]={.type=NLA_S32},+};++staticintdev_mc_valid_dump_req(conststructnlmsghdr*nlh,structsock*sk,+structdev_mc_dump_filter*filter,+structnetlink_ext_ack*extack)+{+structnlattr*tb[IFA_MAX+1];+structifaddrmsg*ifm;+interr;++ifm=nlmsg_payload(nlh,sizeof(*ifm));+if(!ifm){+NL_SET_ERR_MSG(extack,+"Invalid header for multicast dump request");+return-EINVAL;+}++if(ifm->ifa_prefixlen||ifm->ifa_flags||ifm->ifa_scope){+NL_SET_ERR_MSG(extack,+"Invalid values in multicast dump header");+return-EINVAL;+}++err=nlmsg_parse(nlh,sizeof(*ifm),tb,IFA_MAX,+dev_mc_dump_policy,extack);+if(err<0)+returnerr;++if(tb[IFA_TARGET_NETNSID]){+structnet*net;++filter->netnsid=nla_get_s32(tb[IFA_TARGET_NETNSID]);+net=rtnl_get_net_ns_capable(sk,filter->netnsid);+if(IS_ERR(net)){+NL_SET_ERR_MSG(extack,+"Invalid target network namespace id");+returnPTR_ERR(net);+}+netns_tracker_alloc(net,&filter->ns_tracker,GFP_KERNEL);+filter->tgt_net=net;+}++filter->ifindex=ifm->ifa_index;++return0;+}++intdev_mc_dump(structsk_buff*skb,structnetlink_callback*cb)+{+structdev_mc_dump_filterfilter={+.tgt_net=sock_net(skb->sk),+.netnsid=-1,+};+unsignedintflags=NLM_F_MULTI;+struct{+unsignedlongifindex;+intaddr_idx;+}*ctx=(void*)cb->ctx;+unsignedlongs_ifindex;+structnet_device*dev;+interr;++err=dev_mc_valid_dump_req(cb->nlh,skb->sk,&filter,cb->extack);+if(err<0)+returnerr;++cb->seq=READ_ONCE(filter.tgt_net->dev_base_seq);++rcu_read_lock();++if(filter.ifindex){+cb->answer_flags|=NLM_F_DUMP_FILTERED;+flags|=NLM_F_DUMP_FILTERED;+dev=dev_get_by_index_rcu(filter.tgt_net,filter.ifindex);+if(!dev){+err=-ENODEV;+gotoout;+}+err=dev_mc_dump_dev(dev,skb,cb,&ctx->addr_idx,flags,+filter.netnsid);+gotoout;+}++s_ifindex=ctx->ifindex;+for_each_netdev_dump(filter.tgt_net,dev,ctx->ifindex){+/* The device the dump stopped at is gone, do not skip+*entriesofthenextone.+*/+if(dev->ifindex!=s_ifindex)+ctx->addr_idx=0;+err=dev_mc_dump_dev(dev,skb,cb,&ctx->addr_idx,flags,+filter.netnsid);+if(err<0)+break;+}+out:+rcu_read_unlock();+if(filter.netnsid>=0)+put_net_track(filter.tgt_net,&filter.ns_tracker);+returnerr;+}+staticintnetif_addr_lists_snapshot(structnet_device*dev,structnetdev_hw_addr_list*uc_snap,structnetdev_hw_addr_list*mc_snap,
Add the global flag, list the attributes the AF_PACKET dump uses and
describe how ifa-family selects IPv4, IPv6 or link-layer output for
RTM_GETMULTICAST.
Signed-off-by: Yuyang Huang <redacted>
Reviewed-by: Nicolas Dichtel <redacted>
---
Documentation/netlink/specs/rt-addr.yaml | 17 +++++++++++++++--
1 file changed, 15 insertions(+), 2 deletions(-)
@@ -168,7 +170,15 @@ operations:attributes:*ifaddr-all-name:getmulticast-doc:Get / dump IPv4/IPv6 multicast addresses.+doc:|+Get / dump multicast addresses. ifa-family must select the address+family:AF_INET or AF_INET6 for the IP multicast groups joined on+a device, AF_PACKET for the link-layer multicast addresses in the+device filter. Link-layer entries added explicitly, e.g. with+SIOCADDMULTI or "bridge fdb add ... self", rather than by a+protocol join are reported with the global flag set. A non-zero+ifa-index restricts the dump to that device, ifa-prefixlen,+ifa-flags and ifa-scope must be zero.attribute-set:addr-attrsfixed-header:ifaddrmsgdo:
Dump the link-layer multicast addresses of a dummy device and verify
that ifa_index restricts the dump to that device, that the all-hosts
address joined on link up is listed without IFA_F_GLOBAL, that an
address added with SIOCADDMULTI is listed with IFA_F_GLOBAL and
IFA_MC_USERS, and that IFA_TARGET_NETNSID dumps another netns.
Signed-off-by: Yuyang Huang <redacted>
Reviewed-by: Nicolas Dichtel <redacted>
---
tools/testing/selftests/net/rtnetlink.py | 73 +++++++++++++++++++++++-
1 file changed, 71 insertions(+), 2 deletions(-)
@@ -105,6 +110,69 @@ def dump_mcaddr6_check() -> None:s2.close()+defdump_mcaddr_l2_check()->None:+"""+Verifylink-layermulticastaddressesinanAF_PACKETRTM_GETMULTICAST+dump:theifa-indexfilter,mc-users,theglobalflagand+target-netnsid.+"""++withNetNS()asns,NetNSEnter(str(ns)):+forifnamein("dummy1","dummy2"):+ip(f"link add name {ifname} type dummy")+ip(f"link set {ifname} up")+dev_idx=socket.if_nametoindex("dummy1")+ip(f"maddr add {ETH_TEST_MULTICAST_STR} dev dummy1")++rtnl=RtnlAddrFamily()+defer(rtnl.close)+addresses=rtnl.getmulticast(+{"ifa-family":socket.AF_PACKET,"ifa-index":dev_idx},+dump=True)++# dummy2 has entries as well, only dummy1 may be listed+ksft_eq({addr['ifa-index']foraddrinaddresses},{dev_idx},+"AF_PACKET multicast dump ignored ifa-index filter")++entries={addr['multicast']:addrforaddrinaddresses}++# Bringing an Ethernet device up joins 224.0.0.1, which maps+# to 01:00:5e:00:00:01 in the device multicast list.+all_hosts=entries.get(ETH_ALL_HOSTS_MULTICAST)+ksft_not_none(all_hosts,+"dummy1 does not have the all-hosts link-layer address")+ifall_hostsisnotNone:+ksft_not_in('global',all_hosts['flags'],+"protocol entry is global")++static=entries.get(ETH_TEST_MULTICAST)+ksft_not_none(static,"dummy1 does not have the SIOCADDMULTI address")+ifstaticisnotNone:+ksft_eq(static['mc-users'],1,+"unexpected mc-users for the SIOCADDMULTI address")+ksft_in('global',static['flags'],+"SIOCADDMULTI entry is not global")++# target-netnsid dumps another netns, ifa-index is relative to it+withNetNS()aspeer:+ip(f"netns set {peer} 5")+ip("link add name dummy3 type dummy",ns=peer)+ip("link set dummy3 up",ns=peer)+ip(f"maddr add {ETH_PEER_MULTICAST_STR} dev dummy3",ns=peer)+peer_idx=ip("link show dummy3",json=True,ns=peer)[0]['ifindex']++addresses=rtnl.getmulticast(+{"ifa-family":socket.AF_PACKET,"target-netnsid":5,+"ifa-index":peer_idx},dump=True)+ksft_eq({(addr['ifa-index'],addr['target-netnsid'])+foraddrinaddresses},{(peer_idx,5)},+"target-netnsid did not dump the peer netns")+# dummy1 in this netns can have the same ifindex as dummy3+ksft_in(ETH_PEER_MULTICAST,+{addr['multicast']foraddrinaddresses},+"target-netnsid did not dump the peer device")++defipv4_devconf_notify()->None:"""Configureaninterfaceandsetipv4-devconfvaluesthroughnetlink
From: Nicolas Dichtel <hidden> Date: 2026-09-22 07:20:52
Le 22/09/2026 à 01:59, Yuyang Huang a écrit :
RTM_GETMULTICAST dumps IPv4 and IPv6 multicast group memberships, but
the device multicast list (dev->mc) is only available through
/proc/net/dev_mcast, so "ip maddr show" still has to parse procfs for
its link-layer entries.
Handle RTM_GETMULTICAST dumps with ifa_family set to AF_PACKET next to
the dev->mc helpers in dev_addr_lists.c and report every entry of
dev->mc in the existing ifaddrmsg format:
- IFA_MULTICAST carries the raw link-layer address
- IFA_MC_USERS carries the entry reference count
- IFA_F_GLOBAL in IFA_FLAGS reports netdev_hw_addr::global_use, set
by dev_mc_add_global() (SIOCADDMULTI) and dev_mc_add_excl()
("bridge fdb add ... self"), i.e. entries added explicitly rather
than by a protocol join. This is the static column of
/proc/net/dev_mcast
- ifa_scope is RT_SCOPE_LINK
This covers every column of /proc/net/dev_mcast. AF_PACKET is the
family iproute2 already uses for link-layer addresses ("ip -0").
The default FDB dump also walks dev->mc, but only for Ethernet devices
without an ndo_fdb_dump of their own, so bridge, vxlan or macvlan
devices never show their multicast filter there, and it has no users
count or global_use bit. Extending it would change "bridge fdb show"
output and add NDA_* attributes.
Requests are always validated, there are no legacy users: prefixlen,
flags and scope must be zero, ifa_index selects one device and
IFA_TARGET_NETNSID is the only attribute accepted. The dump runs under
RCU and netif_addr_lock_bh() without RTNL, and stamps cb->seq from
dev_base_seq so a device added or removed between dump rounds sets
NLM_F_DUMP_INTR.
Signed-off-by: Yuyang Huang <redacted>
Reviewed-by: Nicolas Dichtel <redacted>
---
[snip]
quoted hunk
+static int dev_mc_dump_dev(struct net_device *dev, struct sk_buff *skb,+ struct netlink_callback *cb, int *s_addr_idx,+ unsigned int flags, int netnsid)+{+ struct netdev_hw_addr *ha;+ int addr_idx = 0;+ int err = 0;++ netif_addr_lock_bh(dev);+ netdev_for_each_mc_addr(ha, dev) {+ if (addr_idx < *s_addr_idx) {+ addr_idx++;+ continue;+ }+ err = dev_mc_fill_addr(skb, dev, ha, NETLINK_CB(cb->skb).portid,+ cb->nlh->nlmsg_seq, flags, netnsid);+ if (err < 0)+ break;+ nl_dump_check_consistent(cb, nlmsg_hdr(skb));+ addr_idx++;+ }+ netif_addr_unlock_bh(dev);++ *s_addr_idx = err < 0 ? addr_idx : 0;++ return err;+}
[snip]
quoted hunk
+int dev_mc_dump(struct sk_buff *skb, struct netlink_callback *cb)+{+ struct dev_mc_dump_filter filter = {+ .tgt_net = sock_net(skb->sk),+ .netnsid = -1,+ };+ unsigned int flags = NLM_F_MULTI;+ struct {+ unsigned long ifindex;+ int addr_idx;+ } *ctx = (void *)cb->ctx;+ unsigned long s_ifindex;+ struct net_device *dev;+ int err;++ err = dev_mc_valid_dump_req(cb->nlh, skb->sk, &filter, cb->extack);+ if (err < 0)+ return err;++ cb->seq = READ_ONCE(filter.tgt_net->dev_base_seq);
dev_base_seq only indicates whether the netdevice list has changed, not whether
the addr list has changed. It's only a partial consistency check.
See inet_base_seq()/inet6_base_seq().
I don't see an equivalent of dev_addr_genid for the L2 mcaddr list; it should
probably be added.
On Tue, Sep 22, 2026 at 4:20 PM Nicolas Dichtel
[off-list ref] wrote:
dev_base_seq only indicates whether the netdevice list has changed, not whether
the addr list has changed. It's only a partial consistency check.
See inet_base_seq()/inet6_base_seq().
I don't see an equivalent of dev_addr_genid for the L2 mcaddr list; it should
probably be added.
Thanks for the suggestion, I will fix this issue in v7: I think I need
to add a per netns counter bumped on dev->mc changes and combine it
with dev_base_seq like inet_base_seq() does.
Thanks,
Yuyang
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] The patch fixes a real, user-visible defect in a shipped netlink
spec — `Documentation/netlink/specs/rt-addr.yaml` declared `addr-…
--
netlink: specs: rt-addr: fix the type of target-netnsid
This changes the target-netnsid attribute in the addr-attrs set of
rt-addr.yaml from binary to s32, matching what the kernel policy
declares. No Link: tags are present in the commit header.
[Severity: Low]
This isn't a bug in the change itself, but should this carry a Fixes: tag?
The spec has been shipped with the wrong type, and the generated ynl code and
the Python decoder have been treating the netns ID as an opaque byte blob,
with the request path rejecting an integer for a binary attribute. The
kernel side has always declared it signed 32-bit:
net/ipv4/devinet.c:ifa_ipv4_policy:
[IFA_TARGET_NETNSID] = { .type = NLA_S32 },
with nla_put_s32() on the emit side and nla_get_s32() on the parse side in
net/ipv6/addrconf.c.
git blame on the pre-patch tree attributes the type: binary line in
Documentation/netlink/specs/rt-addr.yaml to:
dfb0f7d9d97956 ("doc/netlink: Add spec for rt addr messages")
so would adding
Fixes: dfb0f7d9d979 ("doc/netlink: Add spec for rt addr messages")
help backport tooling and users identify which releases carry the wrong
type?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260921235951.3214-1-sigefriedhyy%40gmail.com
Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 1
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] dev_mc_dump()/dev_mc_dump_dev() resume a paused dump with a
bare positional index into dev->mc (ctx->addr_idx / *s_addr_idx) and…
- [Low] uAPI vs netlink YAML spec out of sync: the patch adds
IFA_F_GLOBAL (0x1000), emits IFA_FLAGS/IFA_TARGET_NETNSID in…
--
net: add AF_PACKET multicast dumps
This adds an RTM_GETMULTICAST dump handler for ifa_family AF_PACKET that
reports every entry of dev->mc in the ifaddrmsg format, with the raw
link-layer address in IFA_MULTICAST, the reference count in IFA_MC_USERS
and a new IFA_F_GLOBAL bit mirroring netdev_hw_addr::global_use. The
dump runs under RCU and netif_addr_lock_bh() without RTNL.
[Severity: Low]
This isn't a bug, just an observation about the intermediate state of the
series: this commit adds the new bit 12 and starts emitting IFA_FLAGS and
IFA_TARGET_NETNSID for AF_PACKET RTM_GETMULTICAST, but
Documentation/netlink/specs/rt-addr.yaml is not touched here. At this
commit the ifa-flags enum still ends at stable-privacy, the getmulticast
dump request lists no attributes, the reply set is {multicast, mc-users,
cacheinfo} and the op doc mentions IPv4/IPv6 only, so a spec driven
decoder such as ynl would see bit 12 as unknown and would not describe
target-netnsid or flags.
The next patch in the series, "netlink: specs: rt-addr: document AF_PACKET
multicast dumps", adds "- name: global" to ifa-flags, adds target-netnsid
to the dump request, adds flags and target-netnsid to the shared
mcaddr-attrs reply set and updates the doc, so the gap is closed within
the series. Would it be worth folding the spec update into this commit so
that no bisect point emits attributes the spec does not describe?
+static int dev_mc_dump_dev(struct net_device *dev, struct sk_buff *skb,+ struct netlink_callback *cb, int *s_addr_idx,+ unsigned int flags, int netnsid)+{+ struct netdev_hw_addr *ha;+ int addr_idx = 0;+ int err = 0;++ netif_addr_lock_bh(dev);+ netdev_for_each_mc_addr(ha, dev) {+ if (addr_idx < *s_addr_idx) {+ addr_idx++;+ continue;+ }+ err = dev_mc_fill_addr(skb, dev, ha, NETLINK_CB(cb->skb).portid,+ cb->nlh->nlmsg_seq, flags, netnsid);+ if (err < 0)+ break;+ nl_dump_check_consistent(cb, nlmsg_hdr(skb));+ addr_idx++;+ }+ netif_addr_unlock_bh(dev);++ *s_addr_idx = err < 0 ? addr_idx : 0;++ return err;+}
[Severity: Medium]
Can this dump silently skip or repeat a dev->mc entry without setting
NLM_F_DUMP_INTR?
The resume cursor here is a bare ordinal into dev->mc, and
netif_addr_lock_bh(dev) is dropped before returning, so the lock is not
held across a dump round boundary:
*s_addr_idx = err < 0 ? addr_idx : 0;
If round N stops at ordinal k+1 and an entry before k is removed while the
dump is paused, every later entry shifts down by one and round N+1 starts
at k+1, so the entry that moved from k+1 to k is never reported. An
insertion before the cursor reports one entry twice.
The only consistency stamp is dev_base_seq:
and dev_base_seq_inc() is called only from list_netdevice() and
unlist_netdevice() in net/core/dev.c, so it tracks device register and
unregister, not dev->mc mutations. nl_dump_check_consistent() therefore
cannot raise NLM_F_DUMP_INTR for the mutation that breaks the resume
index, and user space sees what looks like a complete, consistent dump.
The writers are reachable without any privilege, for instance
IP_DROP_MEMBERSHIP: __ip_mc_dec_group()->ip_mc_filter_del()->dev_mc_del(),
or IP_ADD_MEMBERSHIP: ip_mc_inc_group()->dev_mc_add(), and multiple dump
rounds happen as soon as the reply does not fit in one skb.
Since RTNL_FLAG_DUMP_UNLOCKED is used, even the dev->mc writers that do
hold RTNL (SIOCADDMULTI via dev_mc_add_global(), "bridge fdb add ... self"
via dev_mc_add_excl(), dev_mc_flush()) are free to run between two rounds
of this dump. Would a generation counter bumped on dev->mc changes, and
folded into cb->seq, be needed here, or could the resume point be keyed on
the address itself rather than on its position in the list?
The commit message says:
stamps cb->seq from dev_base_seq so a device added or removed between
dump rounds sets NLM_F_DUMP_INTR
That statement is accurate for device add and remove, but should it also
note that changes to the dumped objects themselves, the dev->mc entries,
are not covered by that counter?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260921235951.3214-1-sigefriedhyy%40gmail.com
Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.
Critical: 0 · High: 1 · Medium: 1 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [High] dev_mc_fill_addr() (net/core/dev_addr_lists.c:1210) emits
IFA_MULTICAST with `nla_put(skb, IFA_MULTICAST, dev->addr_len,…
- [Medium] dump_mcaddr_l2_check()
(tools/testing/selftests/net/rtnetlink.py:113) issues the new…
--
selftests: net: test AF_PACKET multicast dumps
Adds dump_mcaddr_l2_check() to tools/testing/selftests/net/rtnetlink.py,
which dumps the link-layer multicast addresses of dummy devices over
AF_PACKET RTM_GETMULTICAST and checks the ifa-index filter, the global
flag, mc-users and target-netnsid handling.
@@ -105,6 +110,69 @@ def dump_mcaddr6_check() -> None: s2.close()+def dump_mcaddr_l2_check() -> None:+ """+ Verify link-layer multicast addresses in an AF_PACKET RTM_GETMULTICAST+ dump: the ifa-index filter, mc-users, the global flag and+ target-netnsid.+ """++ with NetNS() as ns, NetNSEnter(str(ns)):+ for ifname in ("dummy1", "dummy2"):+ ip(f"link add name {ifname} type dummy")+ ip(f"link set {ifname} up")+ dev_idx = socket.if_nametoindex("dummy1")+ ip(f"maddr add {ETH_TEST_MULTICAST_STR} dev dummy1")++ rtnl = RtnlAddrFamily()+ defer(rtnl.close)+ addresses = rtnl.getmulticast(+ {"ifa-family": socket.AF_PACKET, "ifa-index": dev_idx},+ dump=True)
[Severity: Medium]
Should this dump be guarded so that an older kernel produces a skip
instead of a failure?
On a kernel without the PF_PACKET RTM_GETMULTICAST registration,
rtnetlink_rcv_msg() looks up the handler for the requested family, then
retries with PF_UNSPEC:
net/core/rtnetlink.c:rtnetlink_rcv_msg() {
...
link = rtnl_get_link(family, type);
if (!link || !link->dumpit) {
family = PF_UNSPEC;
link = rtnl_get_link(family, type);
if (!link || !link->dumpit)
goto err_unlock;
...
}
There is no PF_UNSPEC RTM_GETMULTICAST dumpit (only PF_INET in
net/ipv4/devinet.c and PF_INET6 in net/ipv6/addrconf.c, plus the
PF_PACKET entry added by this series), so err stays at its -EOPNOTSUPP
initialisation and an error ack comes back. ynl then turns that into an
exception:
tools/net/ynl/pyynl/lib/ynl.py:YnlFamily._op() {
...
if nl_msg.error:
raise NlError(nl_msg)
...
}
which ksft_run() reports as "not ok" plus a traceback.
The two neighbouring tests in this same file take the other route:
_users_for() returns None when mc-users is absent, and
dump_mcaddr_check()/dump_mcaddr6_check() convert that into
raise KsftSkipEx("kernel does not expose IFA_MC_USERS"). Would wrapping
the first getmulticast() call and raising KsftSkipEx on NlError keep this
new case consistent with the rest of the file when the selftest tree is
newer than the kernel under test?
quoted hunk
++ # dummy2 has entries as well, only dummy1 may be listed+ ksft_eq({addr['ifa-index'] for addr in addresses}, {dev_idx},+ "AF_PACKET multicast dump ignored ifa-index filter")++ entries = {addr['multicast']: addr for addr in addresses}
[Severity: High]
The IFA_MULTICAST payload consumed here is emitted by dev_mc_fill_addr()
in net/core/dev_addr_lists.c, added by this series, using the device's
current addr_len for every entry of dev->mc:
net/core/dev_addr_lists.c:dev_mc_fill_addr() {
...
nla_put(skb, IFA_MULTICAST, dev->addr_len, ha->addr) ||
...
}
Each entry was created with an unzeroed allocation where only addr_len
bytes were written:
net/core/dev_addr_lists.c:__hw_addr_create() {
...
ha = kmalloc(alloc_size, GFP_ATOMIC);
if (!ha)
return NULL;
memcpy(ha->addr, addr, addr_len);
...
}
ha->addr is MAX_ADDR_LEN (32) bytes, so bytes addr_len..31 hold stale
slab contents. Can this dump then copy uninitialised heap bytes to user
space if dev->addr_len grows after the entry was inserted?
The sequence I end up with is:
1. create a tap device (ARPHRD_ETHER, addr_len 6; tun_net_mclist is
provided as .ndo_set_rx_mode, which SIOCADDMULTI requires per
net/core/dev_ioctl.c)
2. ip maddr add 01:00:5e:01:01:01 dev tap0, so __hw_addr_create()
copies 6 bytes
3. with the device down, TUNSETLINK to ARPHRD_TUNNEL6:
drivers/net/tun.c:__tun_chr_ioctl() {
...
tun->dev->type = (int) arg;
tun->dev->addr_len = tun_get_addr_len(tun->dev->type);
...
}
addr_len becomes sizeof(struct in6_addr), and dev->mc is not
flushed; the NETDEV_PRE_TYPE_CHANGE handlers only drop the
protocol-joined mappings via ip_mc_unmap()/ipv6_mc_unmap(), not
global_use entries
4. an AF_PACKET RTM_GETMULTICAST dump (RTNL_KIND_GET, no capability
required) now returns 16 bytes for that entry, 10 of which were
never initialised
The pre-existing dev->mc dump path does not reach this because it is
gated on the device type:
net/core/rtnetlink.c:ndo_dflt_fdb_dump() {
...
if (dev->type != ARPHRD_ETHER)
return -EINVAL;
...
}
The new dump has no such restriction. Would bounding the copy by what was
actually initialised - recording or clamping a per-entry length, or
zeroing ha->addr in __hw_addr_create() - be preferable to trusting
dev->addr_len here?
For the record, the bonding path is not a trigger: bond_enslave() calls
dev_uc_flush()/dev_mc_flush() immediately before bond_setup_by_slave()
changes addr_len, so only the tun/tap TUNSETLINK path above applies.
quoted hunk
++ # Bringing an Ethernet device up joins 224.0.0.1, which maps+ # to 01:00:5e:00:00:01 in the device multicast list.+ all_hosts = entries.get(ETH_ALL_HOSTS_MULTICAST)+ ksft_not_none(all_hosts,+ "dummy1 does not have the all-hosts link-layer address")+ if all_hosts is not None:+ ksft_not_in('global', all_hosts['flags'],+ "protocol entry is global")++ static = entries.get(ETH_TEST_MULTICAST)+ ksft_not_none(static, "dummy1 does not have the SIOCADDMULTI address")+ if static is not None:+ ksft_eq(static['mc-users'], 1,+ "unexpected mc-users for the SIOCADDMULTI address")+ ksft_in('global', static['flags'],+ "SIOCADDMULTI entry is not global")
On Wed, Sep 23, 2026 at 9:19 AM [off-list ref] wrote:
[Severity: Low]
This isn't a bug in the change itself, but should this carry a Fixes: tag?
The spec has been shipped with the wrong type, and the generated ynl code and
the Python decoder have been treating the netns ID as an opaque byte blob,
with the request path rejecting an integer for a binary attribute. The
kernel side has always declared it signed 32-bit:
net/ipv4/devinet.c:ifa_ipv4_policy:
[IFA_TARGET_NETNSID] = { .type = NLA_S32 },
with nla_put_s32() on the emit side and nla_get_s32() on the parse side in
net/ipv6/addrconf.c.
git blame on the pre-patch tree attributes the type: binary line in
Documentation/netlink/specs/rt-addr.yaml to:
dfb0f7d9d97956 ("doc/netlink: Add spec for rt addr messages")
so would adding
Fixes: dfb0f7d9d979 ("doc/netlink: Add spec for rt addr messages")
help backport tooling and users identify which releases carry the wrong
type?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260921235951.3214-1-sigefriedhyy%40gmail.com
This issue has been discussed in a previous review comment before, we
want to keep this patch target net-next without the Fixes tag.