From: Shaohua Li <shli@kernel.org> Date: 2017-08-24 19:24:51
From: Shaohua Li <redacted>
two small patches to improve performance for loop in directio mode. The goal is
to increase IO size sending to underlayer disks. As Omar pointed out, the
patches have slight conflict with his, but should be easy to fix.
Thanks,
Shaohua
Shaohua Li (2):
block/loop: set hw_sectors
block/loop: allow request merge for directio mode
drivers/block/loop.c | 67 ++++++++++++++++++++++++++++++++++++++++------------
drivers/block/loop.h | 1 +
2 files changed, 53 insertions(+), 15 deletions(-)
--
2.9.5
From: Shaohua Li <shli@kernel.org> Date: 2017-08-24 19:24:52
From: Shaohua Li <redacted>
Loop can handle any size of request. Limiting it to 255 sectors just
burns the CPU for bio split and request merge for underlayer disk and
also cause bad fs block allocation in directio mode.
Reviewed-by: Omar Sandoval <redacted>
Signed-off-by: Shaohua Li <redacted>
---
drivers/block/loop.c | 1 +
1 file changed, 1 insertion(+)
From: Shaohua Li <shli@kernel.org> Date: 2017-08-24 19:24:53
From: Shaohua Li <redacted>
Currently loop disables merge. While it makes sense for buffer IO mode,
directio mode can benefit from request merge. Without merge, loop could
send small size IO to underlayer disk and harm performance.
Reviewed-by: Omar Sandoval <redacted>
Signed-off-by: Shaohua Li <redacted>
---
drivers/block/loop.c | 66 ++++++++++++++++++++++++++++++++++++++++------------
drivers/block/loop.h | 1 +
2 files changed, 52 insertions(+), 15 deletions(-)
@@ -72,6 +72,7 @@ struct loop_cmd {booluse_aio;/* use AIO interface to handle I/O */longret;structkiocbiocb;+structbio_vec*bvec;};/* Support for loadable transfer modules */
On Thu, Aug 24, 2017 at 12:24:52PM -0700, Shaohua Li wrote:
quoted hunk
From: Shaohua Li <redacted>
Loop can handle any size of request. Limiting it to 255 sectors just
burns the CPU for bio split and request merge for underlayer disk and
also cause bad fs block allocation in directio mode.
Reviewed-by: Omar Sandoval <redacted>
Signed-off-by: Shaohua Li <redacted>
---
drivers/block/loop.c | 1 +
1 file changed, 1 insertion(+)
On Thu, Aug 24, 2017 at 12:24:53PM -0700, Shaohua Li wrote:
From: Shaohua Li <redacted>
Currently loop disables merge. While it makes sense for buffer IO mode,
directio mode can benefit from request merge. Without merge, loop could
send small size IO to underlayer disk and harm performance.
Hi Shaohua,
IMO no matter if merge is used, loop always sends page by page
to VFS in both dio or buffer I/O.
Also if merge is enabled on loop, that means merge is run
on both loop and low level block driver, and not sure if we
can benefit from that.
So Could you provide some performance data about this patch?
From: Shaohua Li <shli@kernel.org> Date: 2017-08-29 15:13:39
On Tue, Aug 29, 2017 at 05:56:05PM +0800, Ming Lei wrote:
On Thu, Aug 24, 2017 at 12:24:53PM -0700, Shaohua Li wrote:
quoted
From: Shaohua Li <redacted>
Currently loop disables merge. While it makes sense for buffer IO mode,
directio mode can benefit from request merge. Without merge, loop could
send small size IO to underlayer disk and harm performance.
Hi Shaohua,
IMO no matter if merge is used, loop always sends page by page
to VFS in both dio or buffer I/O.
Why do you think so?
Also if merge is enabled on loop, that means merge is run
on both loop and low level block driver, and not sure if we
can benefit from that.
why does merge still happen in low level block driver?
So Could you provide some performance data about this patch?
In my virtual machine, a workload improves from ~20M/s to ~50M/s. And I clearly
see the request size becomes bigger.
@@ -464,6 +467,8 @@ static void lo_rw_aio_complete(struct kiocb *iocb, long ret, long ret2){structloop_cmd*cmd=container_of(iocb,structloop_cmd,iocb);+kfree(cmd->bvec);+cmd->bvec=NULL;cmd->ret=ret;blk_mq_complete_request(cmd->rq);}
@@ -473,22 +478,50 @@ static int lo_rw_aio(struct loop_device *lo, struct loop_cmd *cmd,{structiov_iteriter;structbio_vec*bvec;-structbio*bio=cmd->rq->bio;+structrequest*rq=cmd->rq;+structbio*bio=rq->bio;structfile*file=lo->lo_backing_file;+unsignedintoffset;+intsegments=0;intret;-/* nomerge for loop request queue */-WARN_ON(cmd->rq->bio!=cmd->rq->biotail);+if(rq->bio!=rq->biotail){+structreq_iteratoriter;+structbio_vectmp;++__rq_for_each_bio(bio,rq)+segments+=bio_segments(bio);+bvec=kmalloc(sizeof(structbio_vec)*segments,GFP_KERNEL);
The allocation should have been GFP_NOIO.
Sounds good. To make this completely correct isn't easy though, we are calling
into the underlayer fs operations which can do allocation.
Thanks,
Shaohua
On Tue, Aug 29, 2017 at 08:13:39AM -0700, Shaohua Li wrote:
On Tue, Aug 29, 2017 at 05:56:05PM +0800, Ming Lei wrote:
quoted
On Thu, Aug 24, 2017 at 12:24:53PM -0700, Shaohua Li wrote:
quoted
From: Shaohua Li <redacted>
Currently loop disables merge. While it makes sense for buffer IO mode,
directio mode can benefit from request merge. Without merge, loop could
send small size IO to underlayer disk and harm performance.
Hi Shaohua,
IMO no matter if merge is used, loop always sends page by page
to VFS in both dio or buffer I/O.
Why do you think so?
do_blockdev_direct_IO() still handles page by page from iov_iter, and
with bigger request, I guess it might be the plug merge working.
quoted
Also if merge is enabled on loop, that means merge is run
on both loop and low level block driver, and not sure if we
can benefit from that.
why does merge still happen in low level block driver?
Because scheduler is still working on low level disk. My question
is that why the scheduler in low level disk doesn't work now
if scheduler on loop can merge?
quoted
So Could you provide some performance data about this patch?
In my virtual machine, a workload improves from ~20M/s to ~50M/s. And I clearly
see the request size becomes bigger.
Could you share us what the low level disk is?
--
Ming
From: Shaohua Li <shli@kernel.org> Date: 2017-08-30 04:43:20
On Wed, Aug 30, 2017 at 10:51:21AM +0800, Ming Lei wrote:
On Tue, Aug 29, 2017 at 08:13:39AM -0700, Shaohua Li wrote:
quoted
On Tue, Aug 29, 2017 at 05:56:05PM +0800, Ming Lei wrote:
quoted
On Thu, Aug 24, 2017 at 12:24:53PM -0700, Shaohua Li wrote:
quoted
From: Shaohua Li <redacted>
Currently loop disables merge. While it makes sense for buffer IO mode,
directio mode can benefit from request merge. Without merge, loop could
send small size IO to underlayer disk and harm performance.
Hi Shaohua,
IMO no matter if merge is used, loop always sends page by page
to VFS in both dio or buffer I/O.
Why do you think so?
do_blockdev_direct_IO() still handles page by page from iov_iter, and
with bigger request, I guess it might be the plug merge working.
This is not true. directio sends big size bio directly, not because of plug
merge. Please at least check the code before you complain.
quoted
quoted
Also if merge is enabled on loop, that means merge is run
on both loop and low level block driver, and not sure if we
can benefit from that.
why does merge still happen in low level block driver?
Because scheduler is still working on low level disk. My question
is that why the scheduler in low level disk doesn't work now
if scheduler on loop can merge?
The low level disk can still do merge, but since this is directio, the upper
layer already dispatches request as big as possible. There is very little
chance the requests can be merged again.
quoted
quoted
So Could you provide some performance data about this patch?
In my virtual machine, a workload improves from ~20M/s to ~50M/s. And I clearly
see the request size becomes bigger.
On Tue, Aug 29, 2017 at 09:43:20PM -0700, Shaohua Li wrote:
On Wed, Aug 30, 2017 at 10:51:21AM +0800, Ming Lei wrote:
quoted
On Tue, Aug 29, 2017 at 08:13:39AM -0700, Shaohua Li wrote:
quoted
On Tue, Aug 29, 2017 at 05:56:05PM +0800, Ming Lei wrote:
quoted
On Thu, Aug 24, 2017 at 12:24:53PM -0700, Shaohua Li wrote:
quoted
From: Shaohua Li <redacted>
Currently loop disables merge. While it makes sense for buffer IO mode,
directio mode can benefit from request merge. Without merge, loop could
send small size IO to underlayer disk and harm performance.
Hi Shaohua,
IMO no matter if merge is used, loop always sends page by page
to VFS in both dio or buffer I/O.
Why do you think so?
do_blockdev_direct_IO() still handles page by page from iov_iter, and
with bigger request, I guess it might be the plug merge working.
This is not true. directio sends big size bio directly, not because of plug
merge. Please at least check the code before you complain.
I complain nothing, just try to understand the idea behind,
never mind, :-)
quoted
quoted
quoted
Also if merge is enabled on loop, that means merge is run
on both loop and low level block driver, and not sure if we
can benefit from that.
why does merge still happen in low level block driver?
Because scheduler is still working on low level disk. My question
is that why the scheduler in low level disk doesn't work now
if scheduler on loop can merge?
The low level disk can still do merge, but since this is directio, the upper
layer already dispatches request as big as possible. There is very little
chance the requests can be merged again.
That is true, but these requests need to enter scheduler queue and
be tried to merge again, even though it is less possible to succeed.
Double merge may take extra CPU utilization.
Looks it doesn't answer my question.
Without this patch, the requests dispatched to loop won't be merged,
so they may be small and their sectors may be continuous, my question
is why dio bios converted from these small loop requests can't be
merged in block layer when queuing these dio bios to low level device?
quoted
quoted
quoted
So Could you provide some performance data about this patch?
In my virtual machine, a workload improves from ~20M/s to ~50M/s. And I clearly
see the request size becomes bigger.
Could you share us what the low level disk is?
It's a SATA ssd.
For sata, it is pretty easy to trigger I/O merge.
--
Ming
From: Shaohua Li <shli@kernel.org> Date: 2017-08-30 22:06:16
On Wed, Aug 30, 2017 at 02:43:40PM +0800, Ming Lei wrote:
On Tue, Aug 29, 2017 at 09:43:20PM -0700, Shaohua Li wrote:
quoted
On Wed, Aug 30, 2017 at 10:51:21AM +0800, Ming Lei wrote:
quoted
On Tue, Aug 29, 2017 at 08:13:39AM -0700, Shaohua Li wrote:
quoted
On Tue, Aug 29, 2017 at 05:56:05PM +0800, Ming Lei wrote:
quoted
On Thu, Aug 24, 2017 at 12:24:53PM -0700, Shaohua Li wrote:
quoted
From: Shaohua Li <redacted>
Currently loop disables merge. While it makes sense for buffer IO mode,
directio mode can benefit from request merge. Without merge, loop could
send small size IO to underlayer disk and harm performance.
Hi Shaohua,
IMO no matter if merge is used, loop always sends page by page
to VFS in both dio or buffer I/O.
Why do you think so?
do_blockdev_direct_IO() still handles page by page from iov_iter, and
with bigger request, I guess it might be the plug merge working.
This is not true. directio sends big size bio directly, not because of plug
merge. Please at least check the code before you complain.
I complain nothing, just try to understand the idea behind,
never mind, :-)
quoted
quoted
quoted
quoted
Also if merge is enabled on loop, that means merge is run
on both loop and low level block driver, and not sure if we
can benefit from that.
why does merge still happen in low level block driver?
Because scheduler is still working on low level disk. My question
is that why the scheduler in low level disk doesn't work now
if scheduler on loop can merge?
The low level disk can still do merge, but since this is directio, the upper
layer already dispatches request as big as possible. There is very little
chance the requests can be merged again.
That is true, but these requests need to enter scheduler queue and
be tried to merge again, even though it is less possible to succeed.
Double merge may take extra CPU utilization.
Looks it doesn't answer my question.
Without this patch, the requests dispatched to loop won't be merged,
so they may be small and their sectors may be continuous, my question
is why dio bios converted from these small loop requests can't be
merged in block layer when queuing these dio bios to low level device?
loop thread doesn't have plug there. Even we have plug there, it's still a bad
idea to do the merge in low level layer. If we run direct_IO for every 4k, the
overhead is much much higher than bio merge. The direct_IO will call into fs
code, take different mutexes, metadata update for write and so on.
On Wed, Aug 30, 2017 at 03:06:16PM -0700, Shaohua Li wrote:
On Wed, Aug 30, 2017 at 02:43:40PM +0800, Ming Lei wrote:
quoted
On Tue, Aug 29, 2017 at 09:43:20PM -0700, Shaohua Li wrote:
quoted
On Wed, Aug 30, 2017 at 10:51:21AM +0800, Ming Lei wrote:
quoted
On Tue, Aug 29, 2017 at 08:13:39AM -0700, Shaohua Li wrote:
quoted
On Tue, Aug 29, 2017 at 05:56:05PM +0800, Ming Lei wrote:
quoted
On Thu, Aug 24, 2017 at 12:24:53PM -0700, Shaohua Li wrote:
quoted
From: Shaohua Li <redacted>
Currently loop disables merge. While it makes sense for buffer IO mode,
directio mode can benefit from request merge. Without merge, loop could
send small size IO to underlayer disk and harm performance.
Hi Shaohua,
IMO no matter if merge is used, loop always sends page by page
to VFS in both dio or buffer I/O.
Why do you think so?
do_blockdev_direct_IO() still handles page by page from iov_iter, and
with bigger request, I guess it might be the plug merge working.
This is not true. directio sends big size bio directly, not because of plug
merge. Please at least check the code before you complain.
I complain nothing, just try to understand the idea behind,
never mind, :-)
quoted
quoted
quoted
quoted
Also if merge is enabled on loop, that means merge is run
on both loop and low level block driver, and not sure if we
can benefit from that.
why does merge still happen in low level block driver?
Because scheduler is still working on low level disk. My question
is that why the scheduler in low level disk doesn't work now
if scheduler on loop can merge?
The low level disk can still do merge, but since this is directio, the upper
layer already dispatches request as big as possible. There is very little
chance the requests can be merged again.
That is true, but these requests need to enter scheduler queue and
be tried to merge again, even though it is less possible to succeed.
Double merge may take extra CPU utilization.
Looks it doesn't answer my question.
Without this patch, the requests dispatched to loop won't be merged,
so they may be small and their sectors may be continuous, my question
is why dio bios converted from these small loop requests can't be
merged in block layer when queuing these dio bios to low level device?
loop thread doesn't have plug there. Even we have plug there, it's still a bad
idea to do the merge in low level layer. If we run direct_IO for every 4k, the
overhead is much much higher than bio merge. The direct_IO will call into fs
code, take different mutexes, metadata update for write and so on.