From: Marco Crivellari <hidden> Date: 2026-07-20 10:09:15
Hello,
Currently the code uses the per-cpu workqueue system_long_wq to schedule
long running works.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Another good reason to have this unbound,
is the "queue_delayed_work()" function, used to enqueue the work item.
More details on this will follow in the next section.
Recently, a new unbound workqueue specific for long running work has been
added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
~~~ Details about queue_delayed_work ~~~
system_long_wq is a per-cpu workqueue and it is used as a parameter of
queue_delayed_work(). This function schedule an item that it will later
be enqueued (once the timer will fire). __queue_delayed_work() does the job
receiving as "cpu" WORK_CPU_UNBOUND:
if (housekeeping_enabled(HK_TYPE_TIMER)) {
// [....]
} else {
if (likely(cpu == WORK_CPU_UNBOUND))
add_timer_global(timer);
else
add_timer_on(timer, cpu);
}
The timer is global, so can fire everywhere, and the work item will be
enqueued where the timer fired.
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change the
workqueue with the new system_dfl_long_wq, so that the used workqueue is
now unbound and can benefit from scheduler task placement.
Thanks!
---
Changes in v3:
- rebased on v7.2-rc4
- added "net: ti: icssg-prueth: Move long delayed work on system_dfl_long_wq"
- changed also the workqueue in ibmvnic_reset().
Link to v2: https://lore.kernel.org/all/20260706134033.244295-1-marco.crivellari@suse.com/
Changes in v2:
- rebased on v7.2-rc2
- dropped the RFC prefix, kept Ack and review tags
Link to v1: https://lore.kernel.org/all/20260511092846.120141-1-marco.crivellari@suse.com/
Marco Crivellari (6):
ibmvnic: Move long delayed work on system_dfl_long_wq
net: ti: icssg-stats: Move long delayed work on system_dfl_long_wq
net: ti: icssg-prueth: Move long delayed work on system_dfl_long_wq
net: thunderbolt: Move long delayed work on system_dfl_long_wq
net: usb: pegasus: Move long delayed work on system_dfl_long_wq
net: usb: r8152: Move long delayed work on system_dfl_long_wq
drivers/net/ethernet/ibm/ibmvnic.c | 6 +++---
drivers/net/ethernet/ti/icssg/icssg_prueth.c | 2 +-
drivers/net/ethernet/ti/icssg/icssg_stats.c | 2 +-
drivers/net/thunderbolt/main.c | 7 ++++---
drivers/net/usb/pegasus.c | 9 +++++----
drivers/net/usb/r8152.c | 7 ++++---
6 files changed, 18 insertions(+), 15 deletions(-)
--
2.54.0
From: Marco Crivellari <hidden> Date: 2026-07-20 10:09:18
Currently the code enqueue work items using {queue|mod}_delayed_work(),
using system_long_wq. This workqueue should be used when long works are
expected and it is a per-cpu workqueue.
The function(s) end up calling __queue_delayed_work(), which set a global
timer that could fire anywhere, enqueuing the work where the timer fired.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Long work shouldn't stick to a single
CPU.
Recently, a new unbound workqueue specific for long running work has
been added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change
system_long_wq with system_dfl_long_wq so that the work may benefit from
scheduler task placement.
Cc: Mika Westerberg <westeri@kernel.org>
Cc: Yehezkel Bernat <YehezkelShB@gmail.com>
Signed-off-by: Marco Crivellari <redacted>
Acked-by: Mika Westerberg <westeri@kernel.org>
---
drivers/net/thunderbolt/main.c | 7 ++++---
1 file changed, 4 insertions(+), 3 deletions(-)
From: Marco Crivellari <hidden> Date: 2026-07-20 10:09:20
Currently the code enqueue work items using {queue|mod}_delayed_work(),
using system_long_wq. This workqueue should be used when long works are
expected and it is a per-cpu workqueue.
The function(s) end up calling __queue_delayed_work(), which set a global
timer that could fire anywhere, enqueuing the work where the timer fired.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Long work shouldn't stick to a single
CPU.
Recently, a new unbound workqueue specific for long running work has
been added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change
system_long_wq with system_dfl_long_wq so that the work may benefit from
scheduler task placement.
Cc: Ethan Nelson-Moore <redacted>
Cc: linux-usb@vger.kernel.org
Signed-off-by: Marco Crivellari <redacted>
---
drivers/net/usb/r8152.c | 7 ++++---
1 file changed, 4 insertions(+), 3 deletions(-)
@@ -7072,7 +7072,8 @@ static void rtl_hw_phy_work_func_t(struct work_struct *work)/* Delay execution in case request_firmware() is not ready yet.*/-queue_delayed_work(system_long_wq,&tp->hw_phy_work,HZ*10);+queue_delayed_work(system_dfl_long_wq,&tp->hw_phy_work,+HZ*10);gotoignore_once;}
@@ -8840,7 +8841,7 @@ static int rtl8152_reset_resume(struct usb_interface *intf)clear_bit(SELECTIVE_SUSPEND,&tp->flags);rtl_reset_ocp_base(tp);tp->rtl_ops.init(tp);-queue_delayed_work(system_long_wq,&tp->hw_phy_work,0);+queue_delayed_work(system_dfl_long_wq,&tp->hw_phy_work,0);set_ethernet_addr(tp,true);returnrtl8152_resume(intf);}
@@ -10295,7 +10296,7 @@ static int rtl8152_probe_once(struct usb_interface *intf,/* Retry in case request_firmware() is not ready yet. */tp->rtl_fw.retry=true;#endif-queue_delayed_work(system_long_wq,&tp->hw_phy_work,0);+queue_delayed_work(system_dfl_long_wq,&tp->hw_phy_work,0);set_ethernet_addr(tp,false);usb_set_intfdata(intf,tp);
From: Marco Crivellari <hidden> Date: 2026-07-20 10:09:20
Currently the code enqueue work items using {queue|mod}_delayed_work(),
using system_long_wq. This workqueue should be used when long works are
expected and it is a per-cpu workqueue.
The function(s) end up calling __queue_delayed_work(), which set a global
timer that could fire anywhere, enqueuing the work where the timer fired.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Long work shouldn't stick to a single
CPU.
Recently, a new unbound workqueue specific for long running work has
been added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change
system_long_wq with system_dfl_long_wq so that the work may benefit from
scheduler task placement.
Cc: Petko Manolov <petkan@nucleusys.com>
Cc: linux-usb@vger.kernel.org
Signed-off-by: Marco Crivellari <redacted>
---
drivers/net/usb/pegasus.c | 9 +++++----
1 file changed, 5 insertions(+), 4 deletions(-)
From: Marco Crivellari <hidden> Date: 2026-07-20 10:09:20
Currently the code enqueue work items using {queue|mod}_delayed_work(),
using system_long_wq. This workqueue should be used when long works are
expected and it is a per-cpu workqueue.
The function(s) end up calling __queue_delayed_work(), which set a global
timer that could fire anywhere, enqueuing the work where the timer fired.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Long work shouldn't stick to a single
CPU.
Recently, a new unbound workqueue specific for long running work has
been added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change
system_long_wq with system_dfl_long_wq so that the work may benefit from
scheduler task placement.
Cc: Haren Myneni <haren@linux.ibm.com>
Cc: Rick Lindsley <ricklind@linux.ibm.com>
Cc: Nick Child <nnac123@linux.ibm.com>
Cc: Madhavan Srinivasan <maddy@linux.ibm.com>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Nicholas Piggin <npiggin@gmail.com>
Cc: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
Cc: linuxppc-dev@lists.ozlabs.org
Signed-off-by: Marco Crivellari <redacted>
---
drivers/net/ethernet/ibm/ibmvnic.c | 6 +++---
1 file changed, 3 insertions(+), 3 deletions(-)
From: Marco Crivellari <hidden> Date: 2026-07-20 10:09:20
Currently the code enqueue work items using {queue|mod}_delayed_work(),
using system_long_wq. This workqueue should be used when long works are
expected and it is a per-cpu workqueue.
The function(s) end up calling __queue_delayed_work(), which set a global
timer that could fire anywhere, enqueuing the work where the timer fired.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Long work shouldn't stick to a single
CPU.
Recently, a new unbound workqueue specific for long running work has
been added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change
system_long_wq with system_dfl_long_wq so that the work may benefit from
scheduler task placement.
Cc: MD Danish Anwar <danishanwar@ti.com>
Cc: Roger Quadros <rogerq@kernel.org>
Signed-off-by: Marco Crivellari <redacted>
---
drivers/net/ethernet/ti/icssg/icssg_prueth.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Marco Crivellari <hidden> Date: 2026-07-20 10:09:24
Currently the code enqueue work items using {queue|mod}_delayed_work(),
using system_long_wq. This workqueue should be used when long works are
expected and it is a per-cpu workqueue.
The function(s) end up calling __queue_delayed_work(), which set a global
timer that could fire anywhere, enqueuing the work where the timer fired.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Long work shouldn't stick to a single
CPU.
Recently, a new unbound workqueue specific for long running work has
been added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change
system_long_wq with system_dfl_long_wq so that the work may benefit from
scheduler task placement.
Cc: MD Danish Anwar <danishanwar@ti.com>
Cc: Roger Quadros <rogerq@kernel.org>
Cc: linux-arm-kernel@lists.infradead.org
Signed-off-by: Marco Crivellari <redacted>
Reviewed-by: Richard Cheng <redacted>
---
drivers/net/ethernet/ti/icssg/icssg_stats.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Jacob Keller <jacob.e.keller@intel.com> Date: 2026-07-20 22:35:18
On 7/20/2026 3:08 AM, Marco Crivellari wrote:
Hello,
Currently the code uses the per-cpu workqueue system_long_wq to schedule
long running works.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Another good reason to have this unbound,
is the "queue_delayed_work()" function, used to enqueue the work item.
More details on this will follow in the next section.
Recently, a new unbound workqueue specific for long running work has been
added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
~~~ Details about queue_delayed_work ~~~
system_long_wq is a per-cpu workqueue and it is used as a parameter of
queue_delayed_work(). This function schedule an item that it will later
be enqueued (once the timer will fire). __queue_delayed_work() does the job
receiving as "cpu" WORK_CPU_UNBOUND:
if (housekeeping_enabled(HK_TYPE_TIMER)) {
// [....]
} else {
if (likely(cpu == WORK_CPU_UNBOUND))
add_timer_global(timer);
else
add_timer_on(timer, cpu);
}
The timer is global, so can fire everywhere, and the work item will be
enqueued where the timer fired.
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change the
workqueue with the new system_dfl_long_wq, so that the used workqueue is
now unbound and can benefit from scheduler task placement.
Ok. So if I am understanding this correctly, the current code uses
system_long_wq which is per-CPU, but is fired using an unbound timer. As
a result, whichever CPU the timer triggers on will be the one which
selects the work queue. From there, the work item will be enqueued to
that work queue and remain on that work queue until resolving with no
way for scheduler to adjust it?
With the new change, we schedule on the system_dfl_long_wq which *isn't*
per CPU, so the scheduler is free to move the task around and
reschedule. As a result we get better overall behavior with more input
from the scheduler, instead of effective randomness from the timer which
is then forced so that such long running task cannot migrate?
That sounds like a pretty good improvement for the cases where the
queued work doesn't depend on any per-cpu behavior. Nice!
I am not sure I can speak to any of the individual drivers here since I
wouldn't know whether moving that particular work item would be
affected.. so feel free to take this review with a grain of salt :)
Reviewed-by: Jacob Keller <jacob.e.keller@intel.com>
From: Marco Crivellari <hidden> Date: 2026-07-21 08:20:28
Hi,
On Tue, Jul 21, 2026 at 12:35 AM Jacob Keller [off-list ref] wrote:
[...]
quoted
system_long_wq is a per-cpu workqueue and it is used as a parameter of
queue_delayed_work(). This function schedule an item that it will later
be enqueued (once the timer will fire). __queue_delayed_work() does the job
receiving as "cpu" WORK_CPU_UNBOUND:
if (housekeeping_enabled(HK_TYPE_TIMER)) {
// [....]
} else {
if (likely(cpu == WORK_CPU_UNBOUND))
add_timer_global(timer);
else
add_timer_on(timer, cpu);
}
The timer is global, so can fire everywhere, and the work item will be
enqueued where the timer fired.
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change the
workqueue with the new system_dfl_long_wq, so that the used workqueue is
now unbound and can benefit from scheduler task placement.
Ok. So if I am understanding this correctly, the current code uses
system_long_wq which is per-CPU, but is fired using an unbound timer. As
a result, whichever CPU the timer triggers on will be the one which
selects the work queue. From there, the work item will be enqueued to
that work queue and remain on that work queue until resolving with no
way for scheduler to adjust it?
With the new change, we schedule on the system_dfl_long_wq which *isn't*
per CPU, so the scheduler is free to move the task around and
reschedule. As a result we get better overall behavior with more input
from the scheduler, instead of effective randomness from the timer which
is then forced so that such long running task cannot migrate?
That sounds like a pretty good improvement for the cases where the
queued work doesn't depend on any per-cpu behavior. Nice!
Yes, that's pretty much it!
I am not sure I can speak to any of the individual drivers here since I
wouldn't know whether moving that particular work item would be
affected.. so feel free to take this review with a grain of salt :)
Reviewed-by: Jacob Keller <jacob.e.keller@intel.com>
From: Oliver Neukum <oneukum@suse.com> Date: 2026-07-22 08:29:54
On 20.07.26 12:08, Marco Crivellari wrote:
Hi,
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change
system_long_wq with system_dfl_long_wq so that the work may benefit from
scheduler task placement.
these changes are problematic, although they look like a good cleanup
in first place. But the test you are using to determine whether USB
devices need their own work queue is incomplete because you are not
considering the reason they allocate their own work queues.
These drivers have their own work queues because they are part of the block layer.
USB devices can share a device with a block device (storage & UAS) and
USB devices have common, per device operations, in particular reset
and runtime power management and disconnect handling. Because these operations
can be necessary to complete block IO neither they nor anything
they depend on can use IO to allocate memory. That is they need to
perform any memory allocation with GFP_NOIO or GFP_ATOMIC.
That is also true for any operation on a work queue they need to wait
for to make progress. That means you cannot limit your check to per-cpu
variables. You also need to check for such dependencies. In particular
any usage of flush_work() on such queues can deadlock, if you use
common queues.
Please refrain from making such changes unless you have fully analyzed
the dependencies.
Regards
Oliver