Re: [PATCH net] bridge: mrp: reject zero test interval to avoid OOM panic
From: Xiang Mei <hidden>
Date: 2026-03-28 06:19:37
Also in:
bridge
On Fri, Mar 27, 2026 at 01:46:39PM +0200, Nikolay Aleksandrov wrote:
On 27/03/2026 13:34, Simon Horman wrote:quoted
On Wed, Mar 25, 2026 at 08:24:38PM -0700, Xiang Mei wrote:quoted
br_mrp_start_test() and br_mrp_start_in_test() accept the user-supplied interval value from netlink without validation. When interval is 0, usecs_to_jiffies(0) yields 0, causing the delayed work (br_mrp_test_work_expired / br_mrp_in_test_work_expired) to reschedule itself with zero delay. This creates a tight loop on system_percpu_wq that allocates and transmits MRP test frames at maximum rate, exhausting all system memory and causing a kernel panic via OOM deadlock.I would suspect the primary outcome of this problem is high CPU consumption rather than memory exhaustion. Is there a reason to expect that the transmitted fames can't be consumed as fast as they are created?+1 More so with CAP_NET_ADMIN you can cause all sorts of OOM and high-cpu usage conditions. This is a configuration error and OOM doesn't lead to panic unless instructed to. I don't think this is worth changing at all.
Thanks for your review. This path is reachable from an unprivileged user namespace. The capability check goes through rtnetlink_rcv_msg() -> netlink_net_capable() -> netlink_ns_capable(), which checks CAP_NET_ADMIN against the network namespace's user_ns, not init_user_ns. An unprivileged user can create a user+net namespace, get CAP_NET_ADMIN within it, set up a bridge with MRP, and trigger the zero-interval loop. This is not a privileged misconfiguration scenario. Also, the PoC can crash a kernel without "oops=panic" with this bug.
quoted
quoted
The same zero-interval issue applies to br_mrp_start_in_test_parse() for interconnect test frames. Use NLA_POLICY_MIN(NLA_U32, 1) in the nla_policy tables for both IFLA_BRIDGE_MRP_START_TEST_INTERVAL and IFLA_BRIDGE_MRP_START_IN_TEST_INTERVAL, so zero is rejected at the netlink attribute parsing layer before the value ever reaches the workqueue scheduling code. This is consistent with how other bridge subsystems (br_fdb, br_mst) enforce range constraints on netlink attributes. Fixes: 7ab1748e4ce6 ("bridge: mrp: Extend MRP netlink interface for configuring MRP interconnect")I think you also want Fixes: 20f6a05ef635 ("bridge: mrp: Rework the MRP netlink interface") As highlighted by AI review.quoted
Reported-by: Weiming Shi <redacted> Signed-off-by: Xiang Mei <redacted>...