Thread (16 messages) 16 messages, 3 authors, 8d ago

Re: [PATCH net 8/8] netfilter: nf_tables: skip expired catchall elements on insert and delete

From: Pablo Neira Ayuso <pablo@netfilter.org>
Date: 2026-09-19 14:57:24
Also in: netfilter-devel

Hi Jakub,

This comments below look like brainstorming, nothing really that
should prevent this patch from being merged.

On Sat, Sep 19, 2026 at 01:34:51PM +0000, netdev-bot+sashiko@kernel.org wrote:
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider:

Critical: 0 · High: 0 · Medium: 1 · Low: 0

- [Medium] The relaxed duplicate predicate in
  nft_setelem_catchall_insert() (nf_tables_api.c:7003-7016) removes the…

--

netfilter: nf_tables: skip expired catchall elements on insert and delete

This aligns nft_setelem_catchall_insert() and
nft_setelem_catchall_deactivate() with the lookup and dump predicates, so
an expired catchall element no longer counts as a duplicate and no longer
gets picked as the delete candidate. Both walks now also test
__nft_set_elem_expired() with the per-netns transaction timestamp and
nft_set_elem_is_dead().
quoted
diff --git a/net/netfilter/nf_tables_api.c b/net/netfilter/nf_tables_api.c
index c0b754a2d45b..b59628e6240c 100644
--- a/net/netfilter/nf_tables_api.c
+++ b/net/netfilter/nf_tables_api.c
@@ -6995,11 +6995,14 @@ static int nft_setelem_catchall_insert(const struct net *net,
 {
 	struct nft_set_elem_catchall *catchall;
 	u8 genmask = nft_genmask_next(net);
+	u64 tstamp = nft_net_tstamp(net);
 	struct nft_set_ext *ext;
 
 	list_for_each_entry(catchall, &set->catchall_list, list) {
 		ext = nft_set_elem_ext(set, catchall->elem);
-		if (nft_set_elem_active(ext, genmask)) {
+		if (nft_set_elem_active(ext, genmask) &&
+		    !__nft_set_elem_expired(ext, tstamp) &&
+		    !nft_set_elem_is_dead(ext)) {
 			*priv = catchall->elem;
 			return -EEXIST;
 		}
 	}
[Severity: Medium]
Can repeated expire-then-readd cycles grow set->catchall_list without
bound here?

With the relaxed predicate, the expired entry is skipped but stays linked
and genmask-active until GC runs, so the tail of
nft_setelem_catchall_insert() appends a second entry alongside it:

	catchall = kmalloc_obj(*catchall, GFP_KERNEL_ACCOUNT);
	if (!catchall)
		return -ENOMEM;

	catchall->elem = elem->priv;
	list_add_tail_rcu(&catchall->list, &set->catchall_list);

Unlike ordinary elements, catchall elements are excluded from the size
accounting in nft_add_set_elem():
Set size is specified by user... User can create a set without a set
size... This comment makes no sense.
	if (!(flags & NFT_SET_ELEM_CATCHALL)) {
		unsigned int max = nft_set_maxsize(set), nelems;

		nelems = atomic_inc_return(&set->nelems);
		if (nelems > max)
			set_full = true;
	}

so the "return set_full ? -ENFILE : 0;" limiter that caps accumulation of
ordinary elements does not apply to this path. Is there anything else that
bounds the number of stale catchall entries a set can hold?
We could sets a cap for maximum number of items in a transaction, but
noone requested this so far. There is also memcg which should provide
a cap.
Reclamation looks like it depends entirely on GC, and the interval is
user-tunable with no upper bound (NFTA_SET_GC_INTERVAL is accepted as a u32
in milliseconds and nft_set_gc_interval() just converts it):

include/net/netfilter/nf_tables.h:nft_set_gc_interval() {
	u32 gc_int = READ_ONCE(set->gc_int);

	return gc_int ? msecs_to_jiffies(gc_int) : HZ;
}

For nft_set_rhash and nft_set_rbtree the catchall reclaim runs only from
the delayed gc work requeued with that value, so with a large gc-interval
the window for accumulation is correspondingly large.

There is also a packet-path effect: nft_set_catchall_lookup() walks the
whole list, and new elements are appended at the tail, so every accumulated
stale entry is dereferenced before the live one is found:

	list_for_each_entry_rcu(catchall, &set->catchall_list, list) {
		ext = nft_set_elem_ext(set, catchall->elem);
		if (nft_set_elem_active(ext, genmask) &&
		    !nft_set_elem_expired(ext) &&
		    !nft_set_elem_is_dead(ext))
			return ext;
	}

Would it be preferable to reuse or force-reclaim the stale catchall element
on insert, rather than only ignoring it?
Maybe, but that is unrelated to this patch.
quoted
@@ -7092,11 +7095,14 @@ static int nft_setelem_catchall_deactivate(const struct net *net,
 					   struct nft_set_elem *elem)
 {
 	struct nft_set_elem_catchall *catchall;
+	u64 tstamp = nft_net_tstamp(net);
 	struct nft_set_ext *ext;
 
 	list_for_each_entry(catchall, &set->catchall_list, list) {
 		ext = nft_set_elem_ext(set, catchall->elem);
-		if (!nft_is_active_next(net, ext))
+		if (!nft_is_active_next(net, ext) ||
+		    __nft_set_elem_expired(ext, tstamp) ||
+		    nft_set_elem_is_dead(ext))
 			continue;
 
 		kfree(elem->priv);
[Severity: Medium]
Related to the above: with stale entries now skipped and -ENOENT returned,
per-element DELSETELEM can no longer retire the accumulated entries, so a
full set flush or set deletion appears to be the only way to drop them
before GC fires. Is that the intended behaviour?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260918112844.194503-2-pablo%40netfilter.org
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help