From: Antoine Tenart <atenart@kernel.org> Date: 2021-03-22 15:44:15
xps_queue_show is mostly made of an RCU read-side critical section and
calls bitmap_zalloc with GFP_KERNEL in the middle of it. That is not
allowed as this call may sleep and such behaviours aren't allowed in RCU
read-side critical sections. Fix this by using GFP_NOWAIT instead.
Fixes: 5478fcd0f483 ("net: embed nr_ids in the xps maps")
Reported-by: kernel test robot <redacted>
Suggested-by: Matthew Wilcox <willy@infradead.org>
Signed-off-by: Antoine Tenart <atenart@kernel.org>
---
Fix sent to net-next as it fixes an issue only in net-next.
net/core/net-sysfs.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Matthew Wilcox <willy@infradead.org> Date: 2021-03-22 16:55:54
On Mon, Mar 22, 2021 at 04:43:29PM +0100, Antoine Tenart wrote:
xps_queue_show is mostly made of an RCU read-side critical section and
calls bitmap_zalloc with GFP_KERNEL in the middle of it. That is not
allowed as this call may sleep and such behaviours aren't allowed in RCU
read-side critical sections. Fix this by using GFP_NOWAIT instead.
This would be another way of fixing the problem that is slightly less
complex than my initial proposal, but does allow for using GFP_KERNEL
for fewer failures:
@@ -1366,11 +1366,10 @@ static ssize_t xps_queue_show(struct net_device *dev, unsigned int index, { struct xps_dev_maps *dev_maps; unsigned long *mask;- unsigned int nr_ids;+ unsigned int nr_ids, new_nr_ids; int j, len;- rcu_read_lock();- dev_maps = rcu_dereference(dev->xps_maps[type]);+ dev_maps = READ_ONCE(dev->xps_maps[type]); /* Default to nr_cpu_ids/dev->num_rx_queues and do not just return 0 * when dev_maps hasn't been allocated yet, to be backward compatible.
@@ -1379,10 +1378,18 @@ static ssize_t xps_queue_show(struct net_device *dev, unsigned int index, (type == XPS_CPUS ? nr_cpu_ids : dev->num_rx_queues); mask = bitmap_zalloc(nr_ids, GFP_KERNEL);- if (!mask) {- rcu_read_unlock();+ if (!mask) return -ENOMEM;- }++ rcu_read_lock();+ dev_maps = rcu_dereference(dev->xps_maps[type]);+ /* if nr_ids shrank in the meantime, do not overrun array.+ * if it increased, we just won't show the new ones+ */+ new_nr_ids = dev_maps ? dev_maps->nr_ids :+ (type == XPS_CPUS ? nr_cpu_ids : dev->num_rx_queues);+ if (new_nr_ids < nr_ids)+ nr_ids = new_nr_ids; if (!dev_maps || tc >= dev_maps->num_tc) goto out_no_maps;
(or do we need the rcu read lock to read dev->num_rcx_queues? i'm assuming
we only need it to read the xps_maps array)
From: Antoine Tenart <atenart@kernel.org> Date: 2021-03-22 17:42:13
Quoting Matthew Wilcox (2021-03-22 17:54:39)
quoted hunk
On Mon, Mar 22, 2021 at 04:43:29PM +0100, Antoine Tenart wrote:
quoted
xps_queue_show is mostly made of an RCU read-side critical section and
calls bitmap_zalloc with GFP_KERNEL in the middle of it. That is not
allowed as this call may sleep and such behaviours aren't allowed in RCU
read-side critical sections. Fix this by using GFP_NOWAIT instead.
This would be another way of fixing the problem that is slightly less
complex than my initial proposal, but does allow for using GFP_KERNEL
for fewer failures:
@@ -1366,11 +1366,10 @@ static ssize_t xps_queue_show(struct net_device *dev, unsigned int index, { struct xps_dev_maps *dev_maps; unsigned long *mask;- unsigned int nr_ids;+ unsigned int nr_ids, new_nr_ids; int j, len;- rcu_read_lock();- dev_maps = rcu_dereference(dev->xps_maps[type]);+ dev_maps = READ_ONCE(dev->xps_maps[type]);
Couldn't dev_maps be freed between here and the read of dev_maps->nr_ids
as we're not in an RCU read-side critical section?
quoted hunk
/* Default to nr_cpu_ids/dev->num_rx_queues and do not just return 0
* when dev_maps hasn't been allocated yet, to be backward compatible.
@@ -1379,10 +1378,18 @@ static ssize_t xps_queue_show(struct net_device *dev, unsigned int index, (type == XPS_CPUS ? nr_cpu_ids : dev->num_rx_queues); mask = bitmap_zalloc(nr_ids, GFP_KERNEL);- if (!mask) {- rcu_read_unlock();+ if (!mask) return -ENOMEM;- }++ rcu_read_lock();+ dev_maps = rcu_dereference(dev->xps_maps[type]);+ /* if nr_ids shrank in the meantime, do not overrun array.+ * if it increased, we just won't show the new ones+ */+ new_nr_ids = dev_maps ? dev_maps->nr_ids :+ (type == XPS_CPUS ? nr_cpu_ids : dev->num_rx_queues);+ if (new_nr_ids < nr_ids)+ nr_ids = new_nr_ids; if (!dev_maps || tc >= dev_maps->num_tc) goto out_no_maps;
My feeling is there is not much value in having a tricky allocation
logic for reads from xps_cpus and xps_rxqs. While we could come up with
something, returning -ENOMEM on memory pressure should be fine.
Antoine
Couldn't dev_maps be freed between here and the read of dev_maps->nr_ids
as we're not in an RCU read-side critical section?
Oh, good point. Never mind, then.
My feeling is there is not much value in having a tricky allocation
logic for reads from xps_cpus and xps_rxqs. While we could come up with
something, returning -ENOMEM on memory pressure should be fine.
That's fine. It's your code, and this is probably a small allocation
anyway.
From: Antoine Tenart <atenart@kernel.org> Date: 2021-03-22 17:47:01
Quoting Antoine Tenart (2021-03-22 18:41:30)
Quoting Matthew Wilcox (2021-03-22 17:54:39)
quoted
On Mon, Mar 22, 2021 at 04:43:29PM +0100, Antoine Tenart wrote:
quoted
xps_queue_show is mostly made of an RCU read-side critical section and
calls bitmap_zalloc with GFP_KERNEL in the middle of it. That is not
allowed as this call may sleep and such behaviours aren't allowed in RCU
read-side critical sections. Fix this by using GFP_NOWAIT instead.
This would be another way of fixing the problem that is slightly less
complex than my initial proposal, but does allow for using GFP_KERNEL
for fewer failures:
@@ -1366,11 +1366,10 @@ static ssize_t xps_queue_show(struct net_device *dev, unsigned int index, { struct xps_dev_maps *dev_maps; unsigned long *mask;- unsigned int nr_ids;+ unsigned int nr_ids, new_nr_ids; int j, len;- rcu_read_lock();- dev_maps = rcu_dereference(dev->xps_maps[type]);+ dev_maps = READ_ONCE(dev->xps_maps[type]);
Couldn't dev_maps be freed between here and the read of dev_maps->nr_ids
as we're not in an RCU read-side critical section?
* The first read of dev_maps->nr_ids, happening before rcu_read_lock,
not the one shown below.
quoted
/* Default to nr_cpu_ids/dev->num_rx_queues and do not just return 0
* when dev_maps hasn't been allocated yet, to be backward compatible.
@@ -1379,10 +1378,18 @@ static ssize_t xps_queue_show(struct net_device *dev, unsigned int index, (type == XPS_CPUS ? nr_cpu_ids : dev->num_rx_queues); mask = bitmap_zalloc(nr_ids, GFP_KERNEL);- if (!mask) {- rcu_read_unlock();+ if (!mask) return -ENOMEM;- }++ rcu_read_lock();+ dev_maps = rcu_dereference(dev->xps_maps[type]);+ /* if nr_ids shrank in the meantime, do not overrun array.+ * if it increased, we just won't show the new ones+ */+ new_nr_ids = dev_maps ? dev_maps->nr_ids :+ (type == XPS_CPUS ? nr_cpu_ids : dev->num_rx_queues);+ if (new_nr_ids < nr_ids)+ nr_ids = new_nr_ids;
Couldn't dev_maps be freed between here and the read of dev_maps->nr_ids
as we're not in an RCU read-side critical section?
Oh, good point. Never mind, then.
quoted
My feeling is there is not much value in having a tricky allocation
logic for reads from xps_cpus and xps_rxqs. While we could come up with
something, returning -ENOMEM on memory pressure should be fine.
That's fine. It's your code, and this is probably a small allocation
anyway.
All right. Thanks for the suggestions anyway!
Antoine
Hello:
This patch was applied to netdev/net-next.git (refs/heads/master):
On Mon, 22 Mar 2021 16:43:29 +0100 you wrote:
xps_queue_show is mostly made of an RCU read-side critical section and
calls bitmap_zalloc with GFP_KERNEL in the middle of it. That is not
allowed as this call may sleep and such behaviours aren't allowed in RCU
read-side critical sections. Fix this by using GFP_NOWAIT instead.
Fixes: 5478fcd0f483 ("net: embed nr_ids in the xps maps")
Reported-by: kernel test robot <redacted>
Suggested-by: Matthew Wilcox <willy@infradead.org>
Signed-off-by: Antoine Tenart <atenart@kernel.org>
[...]