Re: [PATCH net 8/8] netfilter: nf_tables: skip expired catchall elements on insert and delete
From: Pablo Neira Ayuso <pablo@netfilter.org>
Date: 2026-09-19 14:57:24
Also in:
netfilter-devel
Hi Jakub, This comments below look like brainstorming, nothing really that should prevent this patch from being merged. On Sat, Sep 19, 2026 at 01:34:51PM +0000, netdev-bot+sashiko@kernel.org wrote:
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider: Critical: 0 · High: 0 · Medium: 1 · Low: 0 - [Medium] The relaxed duplicate predicate in nft_setelem_catchall_insert() (nf_tables_api.c:7003-7016) removes the… -- netfilter: nf_tables: skip expired catchall elements on insert and delete This aligns nft_setelem_catchall_insert() and nft_setelem_catchall_deactivate() with the lookup and dump predicates, so an expired catchall element no longer counts as a duplicate and no longer gets picked as the delete candidate. Both walks now also test __nft_set_elem_expired() with the per-netns transaction timestamp and nft_set_elem_is_dead().quoted
diff --git a/net/netfilter/nf_tables_api.c b/net/netfilter/nf_tables_api.c index c0b754a2d45b..b59628e6240c 100644 --- a/net/netfilter/nf_tables_api.c +++ b/net/netfilter/nf_tables_api.c@@ -6995,11 +6995,14 @@ static int nft_setelem_catchall_insert(const struct net *net, { struct nft_set_elem_catchall *catchall; u8 genmask = nft_genmask_next(net); + u64 tstamp = nft_net_tstamp(net); struct nft_set_ext *ext; list_for_each_entry(catchall, &set->catchall_list, list) { ext = nft_set_elem_ext(set, catchall->elem); - if (nft_set_elem_active(ext, genmask)) { + if (nft_set_elem_active(ext, genmask) && + !__nft_set_elem_expired(ext, tstamp) && + !nft_set_elem_is_dead(ext)) { *priv = catchall->elem; return -EEXIST; } }[Severity: Medium] Can repeated expire-then-readd cycles grow set->catchall_list without bound here? With the relaxed predicate, the expired entry is skipped but stays linked and genmask-active until GC runs, so the tail of nft_setelem_catchall_insert() appends a second entry alongside it: catchall = kmalloc_obj(*catchall, GFP_KERNEL_ACCOUNT); if (!catchall) return -ENOMEM; catchall->elem = elem->priv; list_add_tail_rcu(&catchall->list, &set->catchall_list); Unlike ordinary elements, catchall elements are excluded from the size accounting in nft_add_set_elem():
Set size is specified by user... User can create a set without a set size... This comment makes no sense.
if (!(flags & NFT_SET_ELEM_CATCHALL)) {
unsigned int max = nft_set_maxsize(set), nelems;
nelems = atomic_inc_return(&set->nelems);
if (nelems > max)
set_full = true;
}
so the "return set_full ? -ENFILE : 0;" limiter that caps accumulation of
ordinary elements does not apply to this path. Is there anything else that
bounds the number of stale catchall entries a set can hold?We could sets a cap for maximum number of items in a transaction, but noone requested this so far. There is also memcg which should provide a cap.
Reclamation looks like it depends entirely on GC, and the interval is
user-tunable with no upper bound (NFTA_SET_GC_INTERVAL is accepted as a u32
in milliseconds and nft_set_gc_interval() just converts it):
include/net/netfilter/nf_tables.h:nft_set_gc_interval() {
u32 gc_int = READ_ONCE(set->gc_int);
return gc_int ? msecs_to_jiffies(gc_int) : HZ;
}
For nft_set_rhash and nft_set_rbtree the catchall reclaim runs only from
the delayed gc work requeued with that value, so with a large gc-interval
the window for accumulation is correspondingly large.
There is also a packet-path effect: nft_set_catchall_lookup() walks the
whole list, and new elements are appended at the tail, so every accumulated
stale entry is dereferenced before the live one is found:
list_for_each_entry_rcu(catchall, &set->catchall_list, list) {
ext = nft_set_elem_ext(set, catchall->elem);
if (nft_set_elem_active(ext, genmask) &&
!nft_set_elem_expired(ext) &&
!nft_set_elem_is_dead(ext))
return ext;
}
Would it be preferable to reuse or force-reclaim the stale catchall element
on insert, rather than only ignoring it?Maybe, but that is unrelated to this patch.
quoted
@@ -7092,11 +7095,14 @@ static int nft_setelem_catchall_deactivate(const struct net *net, struct nft_set_elem *elem) { struct nft_set_elem_catchall *catchall; + u64 tstamp = nft_net_tstamp(net); struct nft_set_ext *ext; list_for_each_entry(catchall, &set->catchall_list, list) { ext = nft_set_elem_ext(set, catchall->elem); - if (!nft_is_active_next(net, ext)) + if (!nft_is_active_next(net, ext) || + __nft_set_elem_expired(ext, tstamp) || + nft_set_elem_is_dead(ext)) continue; kfree(elem->priv);[Severity: Medium] Related to the above: with stale entries now skipped and -ENOENT returned, per-element DELSETELEM can no longer retire the accumulated entries, so a full set flush or set deletion appears to be the only way to drop them before GC fires. Is that the intended behaviour? -- Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260918112844.194503-2-pablo%40netfilter.org