Thread (15 messages) 15 messages, 4 authors, 2021-05-20

Re: [PATCH] blk-mq: plug request for shared sbitmap

flat view

From: Ming Lei <hidden>
Date: 2021-05-20 01:23:57

On Wed, May 19, 2021 at 09:41:01AM +0100, John Garry wrote:
On 19/05/2021 01:21, Ming Lei wrote:
quoted
quoted
The 'after' results are similar to without shared sbitmap, i.e using
reply-map:

reply-map:
450K (read), 430K IOPs (randread)
OK, that is expected result. After shared sbitmap, IO merge gets improved
when batching submission is bypassed, meantime IOPS of random IO drops
because cpu utilization is increased.

So that isn't a regression, let's live with this awkward situation,:-(
Well at least we have ~ parity with non-shared sbitmap now. And also know
higher performance is possible for "read" (vs "randread") scenario, FWIW.
NVMe too, but never see it becomes true, :-)
BTW, recently we have seen 2x optimisation/improvement for shared sbitmap
which were under/related to nr_hw_queues == 1 check - this patch and the
changing of the default IO sched.
You mean you saw 2X improvement in your hisilicon SAS compared with
non-shared sbitmap? In Yanhui's virt test, we just bring back the perf
to non-shared sbitmap's level.
I am wondering how you detected/analyzed this issue, and whether we need to
audit other nr_hw_queues == 1 checks? I did a quick scan, and the only
possible thing I see is the other q->nr_hw_queues > 1 check for direct issue
in blk_mq_subit_bio() - I suspect you know more about that topic.
IMO, the direct issue code path is fine.

For slow devices, we won't use none, so the code path can't be reached.

For fast device, direct issue is often a win.

Also looks your test on non has shown that perf is fine.

Thanks,
Ming
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help