From: Jeff Moyer <hidden> Date: 2012-01-23 18:28:46
Andrea Arcangeli [off-list ref] writes:
On Mon, Jan 23, 2012 at 05:18:57PM +0100, Jan Kara wrote:
quoted
requst granularity. Sure, big requests will take longer to complete but
maximum request size is relatively low (512k by default) so writing maximum
sized request isn't that much slower than writing 4k. So it works OK in
practice.
Totally unrelated to the writeback, but the merged big 512k requests
actually adds up some measurable I/O scheduler latencies and they in
turn slightly diminish the fairness that cfq could provide with
smaller max request size. Probably even more measurable with SSDs (but
then SSDs are even faster).
Are you speaking from experience? If so, what workloads were negatively
affected by merging, and how did you measure that?
Cheers,
Jeff
From: Andrea Arcangeli <hidden> Date: 2012-01-23 18:56:50
On Mon, Jan 23, 2012 at 01:28:08PM -0500, Jeff Moyer wrote:
Are you speaking from experience? If so, what workloads were negatively
affected by merging, and how did you measure that?
Any workload where two processes compete for accessing the same disk
and one process writes big requests (usually async writes), the other
small (usually sync reads). The one with the small 4k requests
(usually reads) gets some artificial latency if the big requests are
512k. Vivek did a recent measurement to verify the issue is still
there, and it's basically an hardware issue. Software can't do much
other than possibly reducing the max request size when we notice such
an I/O pattern coming in cfq. I did old measurements that's how I knew
it, but they were so ancient they're worthless by now, this is why
Vivek had to repeat it to verify before we could assume it still
existed on recent hardware.
These days with cgroups it may be a bit more relevant as max write
bandwidth may be secondary to latency/QoS.
From: Chris Mason <hidden> Date: 2012-01-24 16:51:33
On Mon, Jan 23, 2012 at 01:28:08PM -0500, Jeff Moyer wrote:
Andrea Arcangeli [off-list ref] writes:
quoted
On Mon, Jan 23, 2012 at 05:18:57PM +0100, Jan Kara wrote:
quoted
requst granularity. Sure, big requests will take longer to complete but
maximum request size is relatively low (512k by default) so writing maximum
sized request isn't that much slower than writing 4k. So it works OK in
practice.
Totally unrelated to the writeback, but the merged big 512k requests
actually adds up some measurable I/O scheduler latencies and they in
turn slightly diminish the fairness that cfq could provide with
smaller max request size. Probably even more measurable with SSDs (but
then SSDs are even faster).
Are you speaking from experience? If so, what workloads were negatively
affected by merging, and how did you measure that?
https://lkml.org/lkml/2011/12/13/326
This patch is another example, although for a slight different reason.
I really have no idea yet what the right answer is in a generic sense,
but you don't need a 512K request to see higher latencies from merging.
-chris
From: Christoph Hellwig <hch@infradead.org> Date: 2012-01-24 16:56:40
On Tue, Jan 24, 2012 at 10:15:04AM -0500, Chris Mason wrote:
https://lkml.org/lkml/2011/12/13/326
This patch is another example, although for a slight different reason.
I really have no idea yet what the right answer is in a generic sense,
but you don't need a 512K request to see higher latencies from merging.
That assumes the 512k requests is created by merging. We have enough
workloads that create large I/O from the get go, and not splitting them
and eventually merging them again would be a big win. E.g. I'm
currently looking at a distributed block device which uses internal 4MB
chunks, and increasing the maximum request size to that dramatically
increases the read performance.
From: Andreas Dilger <hidden> Date: 2012-01-24 17:00:50
Cheers, Andreas
On 2012-01-24, at 9:56, Christoph Hellwig [off-list ref] wrote:
On Tue, Jan 24, 2012 at 10:15:04AM -0500, Chris Mason wrote:
quoted
https://lkml.org/lkml/2011/12/13/326
This patch is another example, although for a slight different reason.
I really have no idea yet what the right answer is in a generic sense,
but you don't need a 512K request to see higher latencies from merging.
That assumes the 512k requests is created by merging. We have enough
workloads that create large I/O from the get go, and not splitting them
and eventually merging them again would be a big win. E.g. I'm
currently looking at a distributed block device which uses internal 4MB
chunks, and increasing the maximum request size to that dramatically
increases the read performance.
--
To unsubscribe from this list: send the line "unsubscribe linux-fsdevel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
From: Andrea Arcangeli <hidden> Date: 2012-01-24 17:06:36
On Tue, Jan 24, 2012 at 11:56:31AM -0500, Christoph Hellwig wrote:
That assumes the 512k requests is created by merging. We have enough
workloads that create large I/O from the get go, and not splitting them
and eventually merging them again would be a big win. E.g. I'm
currently looking at a distributed block device which uses internal 4MB
chunks, and increasing the maximum request size to that dramatically
increases the read performance.
Depends on the device though, if it's a normal disk, it likely only
reduces the number of dma ops without increasing performance too
much. Most disks should reach platter speed at 64KB, so larger request
only saves a bit of cpu in interrutps and stuff.
But I think nobody here was suggesting to reduce the request size by
default. cfq should easily notice when there are multiple queues that
are being submitted in the same time range. A device in addition to
specifying the max request dma size it can handle it could specify the
minimum it runs at platter speed and cfq could degrade to it when
there's multiple queues running in parallel over the same millisecond
or so. Reads will return in the I/O queue almost immediately but
they'll be out for a little while until the data is copied to
userland. So it'd need to keep it down to the min request size the
device allows to reach platter speed, for a little while. Then if no
other queue presents itself it double up the request size for each
unit of time until it reaches the max again. Maybe that could work, maybe
not :). Waiting only once for 4MB sounds better than waiting every
time 4MB for each 4k metadata seeking read.
From: Chris Mason <hidden> Date: 2012-01-24 17:08:18
On Tue, Jan 24, 2012 at 11:56:31AM -0500, Christoph Hellwig wrote:
On Tue, Jan 24, 2012 at 10:15:04AM -0500, Chris Mason wrote:
quoted
https://lkml.org/lkml/2011/12/13/326
This patch is another example, although for a slight different reason.
I really have no idea yet what the right answer is in a generic sense,
but you don't need a 512K request to see higher latencies from merging.
That assumes the 512k requests is created by merging. We have enough
workloads that create large I/O from the get go, and not splitting them
and eventually merging them again would be a big win. E.g. I'm
currently looking at a distributed block device which uses internal 4MB
chunks, and increasing the maximum request size to that dramatically
increases the read performance.
Is this read latency or read tput? If you're waiting on the whole 4MB
anyway, I'd expect one request to be better for both. But Andrea's
original question was on the impact of the big request on other requests
being serviced by the drive....there's really not much we can do about
that outside of more knobs for the admin.
-chris
From: Andreas Dilger <hidden> Date: 2012-01-24 17:08:47
On 2012-01-24, at 9:56, Christoph Hellwig [off-list ref] wrote:
On Tue, Jan 24, 2012 at 10:15:04AM -0500, Chris Mason wrote:
quoted
https://lkml.org/lkml/2011/12/13/326
This patch is another example, although for a slight different reason.
I really have no idea yet what the right answer is in a generic sense,
but you don't need a 512K request to see higher latencies from merging.
That assumes the 512k requests is created by merging. We have enough
workloads that create large I/O from the get go, and not splitting them
and eventually merging them again would be a big win. E.g. I'm
currently looking at a distributed block device which uses internal 4MB
chunks, and increasing the maximum request size to that dramatically
increases the read performance.
(sorry about last email, hit send by accident)
I don't think we can have a "one size fits all" policy here. In most RAID devices the IO size needs to be at least 1MB, and with newer devices 4MB gives better performance.
One of the reasons that Lustre used to hack so much around the VFS and VM APIs is exactly to avoid the splitting of read/write requests into pages and then depend on the elevator to reconstruct a good-sized IO out of it.
Things have gotten better with newer kernels, but there is still a ways to go w.r.t. allowing large IO requests to pass unhindered through to disk (or at least as far as enduring that the IO is aligned to the underlying disk geometry).
Cheers, Andreas