From: Sean Anderson <hidden> Date: 2021-08-02 17:09:32
This puts tasks submitted to the SDIO workqueue at the head of the queue
and runs them immediately. This gets higher RX throughput with the SDIO
bus.
This was originally submitted as [1]. The original author Wright Feng
reports
throughput result with 43455(11ac) on 1 core 1.6 Ghz platform is
Without WQ_HIGGPRI TX/RX: 293/301 (mbps)
With WQ_HIGHPRI TX/RX: 293/321 (mbps)
I tested this with a 43364(11bgn) on a 1 core 800 MHz platform and got
Without WQ_HIGHPRI TX/RX: 16/19 (Mbits/sec)
With WQ_HIGHPRI TX/RX: 24/20 (MBits/sec)
[1] https://lore.kernel.org/linux-wireless/1584604406-15452-4-git-send-email-wright.feng@cypress.com/
Signed-off-by: Sean Anderson <redacted>
---
drivers/net/wireless/broadcom/brcm80211/brcmfmac/sdio.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Sean Anderson <hidden> Date: 2021-08-17 16:48:47
ping?
On 8/2/21 1:09 PM, Sean Anderson wrote:
quoted hunk
This puts tasks submitted to the SDIO workqueue at the head of the queue
and runs them immediately. This gets higher RX throughput with the SDIO
bus.
This was originally submitted as [1]. The original author Wright Feng
reports
quoted
throughput result with 43455(11ac) on 1 core 1.6 Ghz platform is
Without WQ_HIGGPRI TX/RX: 293/301 (mbps)
With WQ_HIGHPRI TX/RX: 293/321 (mbps)
I tested this with a 43364(11bgn) on a 1 core 800 MHz platform and got
Without WQ_HIGHPRI TX/RX: 16/19 (Mbits/sec)
With WQ_HIGHPRI TX/RX: 24/20 (MBits/sec)
[1] https://lore.kernel.org/linux-wireless/1584604406-15452-4-git-send-email-wright.feng@cypress.com/
Signed-off-by: Sean Anderson <redacted>
---
drivers/net/wireless/broadcom/brcm80211/brcmfmac/sdio.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Arend van Spriel <hidden> Date: 2021-08-17 17:17:15
On August 17, 2021 6:50:50 PM Sean Anderson [off-list ref] wrote:
ping?
Good idea to ping with a top-level post :-p
On 8/2/21 1:09 PM, Sean Anderson wrote:
quoted
This puts tasks submitted to the SDIO workqueue at the head of the queue
and runs them immediately. This gets higher RX throughput with the SDIO
bus.
This was originally submitted as [1]. The original author Wright Feng
reports
quoted
throughput result with 43455(11ac) on 1 core 1.6 Ghz platform is
Without WQ_HIGGPRI TX/RX: 293/301 (mbps)
With WQ_HIGHPRI TX/RX: 293/321 (mbps)
While I understand the obvious gain it seems like a wrong move to me. What
if all workqueues in the kernel would start using this flag? I bet the gain
above would be negated and all are equal in the eyes of .. the kernel
Regards,
Arend
From: Sean Anderson <hidden> Date: 2021-08-17 18:22:05
On 8/17/21 1:17 PM, Arend van Spriel wrote:
On August 17, 2021 6:50:50 PM Sean Anderson [off-list ref] wrote:
quoted
ping?
Good idea to ping with a top-level post :-p
Sorry, I'm not subscribed to this list; did I mess up the list/CC for the original?
quoted
On 8/2/21 1:09 PM, Sean Anderson wrote:
quoted
This puts tasks submitted to the SDIO workqueue at the head of the queue
and runs them immediately. This gets higher RX throughput with the SDIO
bus.
This was originally submitted as [1]. The original author Wright Feng
reports
quoted
throughput result with 43455(11ac) on 1 core 1.6 Ghz platform is
Without WQ_HIGGPRI TX/RX: 293/301 (mbps)
With WQ_HIGHPRI TX/RX: 293/321 (mbps)
While I understand the obvious gain it seems like a wrong move to me. What if all workqueues in the kernel would start using this flag? I bet the gain above would be negated and all are equal in the eyes of .. the kernel
Is there an official policy on what counts as high-priority? Using some
very-scientific methodology [1], it seems like most high-priority
workqueues are in drivers/net and fs. Making these queues high-priority
seems to be commonplace. For example, in fe101716c7c9 ("rtw88: replace
tx tasklet with work queue"), Po-Hao Huang remarks:
Since throughput is delay-sensitive in most cases, we allocate a
dedicated, high priority wq for our needs.
which is effectively the same rationale as this patch. At least for my
application, network transfer speed is one of the most important
performance metrics.
The original patch got the following feedback [2] from Kalle Valo:
Why would someone want to disable this? Like in patch 2, please avoid
adding new module parameters as much as possible.
Hello,
On Tue, Aug 17, 2021 at 02:21:55PM -0400, Sean Anderson wrote:
quoted
While I understand the obvious gain it seems like a wrong move to me. What if all workqueues in the kernel would start using this flag? I bet the gain above would be negated and all are equal in the eyes of .. the kernel
Is there an official policy on what counts as high-priority? Using some
very-scientific methodology [1], it seems like most high-priority
workqueues are in drivers/net and fs. Making these queues high-priority
seems to be commonplace. For example, in fe101716c7c9 ("rtw88: replace
tx tasklet with work queue"), Po-Hao Huang remarks:
I think this is actually a good candidate for HIGHPRI. As you noted, stuff
which interacts with hardware in latency sensitive manner with impact on
observable performance is one of the common use cases. The alternatives
would be doing it from hard/softirqs which are higher priorities anyway.
Thanks.
--
tejun
From: Arend van Spriel <hidden> Date: 2021-08-17 21:14:53
On August 17, 2021 8:28:23 PM Tejun Heo [off-list ref] wrote:
Hello,
On Tue, Aug 17, 2021 at 02:21:55PM -0400, Sean Anderson wrote:
quoted
quoted
While I understand the obvious gain it seems like a wrong move to me. What
if all workqueues in the kernel would start using this flag? I bet the gain
above would be negated and all are equal in the eyes of .. the kernel
Is there an official policy on what counts as high-priority? Using some
very-scientific methodology [1], it seems like most high-priority
workqueues are in drivers/net and fs. Making these queues high-priority
seems to be commonplace. For example, in fe101716c7c9 ("rtw88: replace
tx tasklet with work queue"), Po-Hao Huang remarks:
I think this is actually a good candidate for HIGHPRI. As you noted, stuff
which interacts with hardware in latency sensitive manner with impact on
observable performance is one of the common use cases. The alternatives
would be doing it from hard/softirqs which are higher priorities anyway.
Hi, Tejun
Thanks for the explanation.
Regards,
Arend
From: Arend van Spriel <hidden> Date: 2021-08-18 04:53:43
On August 2, 2021 7:11:12 PM Sean Anderson [off-list ref] wrote:
This puts tasks submitted to the SDIO workqueue at the head of the queue
and runs them immediately. This gets higher RX throughput with the SDIO
bus.
This was originally submitted as [1]. The original author Wright Feng
reports
quoted
throughput result with 43455(11ac) on 1 core 1.6 Ghz platform is
Without WQ_HIGGPRI TX/RX: 293/301 (mbps)
With WQ_HIGHPRI TX/RX: 293/321 (mbps)
Not sure if Wright Feng needs to be attributed as you clearly had a good
look at his patch and the discussion related to it. You can add my ...
Reviewed-by: Arend van Spriel <redacted>
Signed-off-by: Sean Anderson <redacted>
---
drivers/net/wireless/broadcom/brcm80211/brcmfmac/sdio.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Kalle Valo <hidden> Date: 2021-08-21 16:59:29
Sean Anderson [off-list ref] wrote:
This puts tasks submitted to the SDIO workqueue at the head of the queue
and runs them immediately. This gets higher RX throughput with the SDIO
bus.
This was originally submitted as [1]. The original author Wright Feng
reports
quoted
throughput result with 43455(11ac) on 1 core 1.6 Ghz platform is
Without WQ_HIGGPRI TX/RX: 293/301 (mbps)
With WQ_HIGHPRI TX/RX: 293/321 (mbps)