From: Pavel Emelyanov <hidden> Date: 2012-07-04 07:11:23
On 07/04/2012 07:01 AM, Nikolaus Rath wrote:
Hi Pavel,
I think it's great that you're working on this! I've been waiting for
FUSE being able to supply write data in bigger chunks for a long time,
and I'm very excited to see some progress on this. I'm not a kernel
developer, but I'll be happy to try the patches.
Just to make it clear. I didn't increase the 32 pages per request limit. What
I did is made FUSE submit more than one request at a time while serving massive
writes. So yes, bigger chunks can be now seen by the daemon, but it should read
several requests for that.
While I try to get this to compile,
That's good news! Thanks a lot!
a few more questions:
Pavel Emelyanov [off-list ref] writes:
quoted
Hi everyone.
One of the problems with the existing FUSE implementation is that it uses the
write-through cache policy which results in performance problems on certain
workloads. E.g. when copying a big file into a FUSE file the cp pushes every
128k to the userspace synchronously. This becomes a problem when the userspace
back-end uses networking for storing the data.
A good solution of this is switching the FUSE page cache into a write-back policy.
With this file data are pushed to the userspace with big chunks (depending on the
dirty memory limits, but this is much more than 128k) which lets the FUSE daemons
handle the size updates in a more efficient manner.
The writeback feature is per-connection and is explicitly configurable at the
init stage (is it worth making it CAP_SOMETHING protected?)
From your description it sounds as if the only effect of write-back is
to increase the chunk size. Why is the a need to require special
privileges for this?
Provided I understand the code correctly: if FUSE daemon turns writeback on and sets
per-bdi dirty limit too high it can cause a deadlock on the box. Thus then daemon
should be trusted by the kernel, i.e. -- privileged.
quoted
When the writeback is turned ON:
* still copy writeback pages to temporary buffer when sending a writeback request
and finish the page writeback immediately
Could you elaborate? I don't understand what you're saying here.
This is an implementation detail. To avoid a deadlock in the memory reclaim code
existing FUSE copies a mmapped dirty page contents into a temporary buffer before
sending it to the user space. What I wanted to say here is that I did use the same
trick in the introduced writeback paths. This doesn't affect kernel-to-user API
at all.
Best,
-Nikolaus
Thanks,
Pavel
------------------------------------------------------------------------------
Live Security Virtual Conference
Exclusive live event will cover all the ways today's security and
threat landscape has changed and how IT managers can respond. Discussions
will include endpoint security, mobile security and the latest in malware
threats. http://www.accelacomm.com/jaw/sfrnl04242012/114/50122263/
Hi Pavel,
I think it's great that you're working on this! I've been waiting for
FUSE being able to supply write data in bigger chunks for a long time,
and I'm very excited to see some progress on this. I'm not a kernel
developer, but I'll be happy to try the patches.
Just to make it clear. I didn't increase the 32 pages per request limit. What
I did is made FUSE submit more than one request at a time while serving massive
writes. So yes, bigger chunks can be now seen by the daemon, but it should read
several requests for that.
Ah, I thought that your patch would do both. So with the patch an
userspace client can now writes data in say 4 kb chunks, and the FUSE
daemon will still receive it from the kernel in 128 kb chunks? But if
the client writes a say 1 MB chunk, the FUSE daemon will still see 8
128kb write requests?
Would it be very hard to raise the 32 pages per request limit at the
same time?
quoted
quoted
A good solution of this is switching the FUSE page cache into a write-back policy.
With this file data are pushed to the userspace with big chunks (depending on the
dirty memory limits, but this is much more than 128k) which lets the FUSE daemons
handle the size updates in a more efficient manner.
The writeback feature is per-connection and is explicitly configurable at the
init stage (is it worth making it CAP_SOMETHING protected?)
From your description it sounds as if the only effect of write-back is
to increase the chunk size. Why the need to require special
privileges for this?
Provided I understand the code correctly: if FUSE daemon turns writeback on and sets
per-bdi dirty limit too high it can cause a deadlock on the box. Thus then daemon
should be trusted by the kernel, i.e. -- privileged.
Wouldn't it be more reasonable to enforce that the bdi dirty limit is
not set too high then?
Thanks,
-Nikolaus
--
»Time flies like an arrow, fruit flies like a Banana.«
PGP fingerprint: 5B93 61F8 4EA2 E279 ABF6 02CF A9AD B7F8 AE4E 425C
--
To unsubscribe from this list: send the line "unsubscribe linux-fsdevel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Ok, I got it to compile and tested it by copying a large file into an
S3QL FUSE file system using different block sizes. At first glance,
things look fantastic.
With kernel 3.2:
Blocksize: 4k
131072000 bytes (131 MB) copied, 7.76037 s, 16.9 MB/s
Blocksize: 8k
131072000 bytes (131 MB) copied, 2.48469 s, 52.8 MB/s
Blocksize: 16k
131072000 bytes (131 MB) copied, 2.21868 s, 59.1 MB/s
Blocksize: 32k
131072000 bytes (131 MB) copied, 2.0455 s, 64.1 MB/s
Blocksize: 64k
131072000 bytes (131 MB) copied, 1.83897 s, 71.3 MB/s
Blocksize: 96k
131072000 bytes (131 MB) copied, 1.72298 s, 76.1 MB/s
Blocksize: 128k
131072000 bytes (131 MB) copied, 2.10194 s, 62.4 MB/s
Blocksize: 256k
131072000 bytes (131 MB) copied, 1.96788 s, 66.6 MB/s
Blocksize: 512k
131072000 bytes (131 MB) copied, 2.42342 s, 54.1 MB/s
Blocksize: 1024k
131072000 bytes (131 MB) copied, 1.6651 s, 78.7 MB/s
With 3.5-pre and write-back patches:
Blocksize: 4k
131072000 bytes (131 MB) copied, 5.50489 s, 23.8 MB/s
Blocksize: 8k
131072000 bytes (131 MB) copied, 1.43694 s, 91.2 MB/s
Blocksize: 16k
131072000 bytes (131 MB) copied, 0.967118 s, 136 MB/s
Blocksize: 32k
131072000 bytes (131 MB) copied, 0.707767 s, 185 MB/s
Blocksize: 64k
131072000 bytes (131 MB) copied, 0.605011 s, 217 MB/s
Blocksize: 96k
131072000 bytes (131 MB) copied, 0.573445 s, 229 MB/s
Blocksize: 128k
131072000 bytes (131 MB) copied, 0.51943 s, 252 MB/s
Blocksize: 256k
131072000 bytes (131 MB) copied, 0.458335 s, 286 MB/s
Blocksize: 512k
131072000 bytes (131 MB) copied, 0.452351 s, 290 MB/s
Blocksize: 1024k
131072000 bytes (131 MB) copied, 0.446027 s, 294 MB/s
However, I suspect that most of the gain is really because with the
patch most of the data is still in the kernel cache when dd finishes and
hasn't yet been received by the FUSE client.
Is there a way to force flushing of the fuse cache, so that I can
measure the time for dd + final flush?
Best,
-Nikolaus
--
»Time flies like an arrow, fruit flies like a Banana.«
PGP fingerprint: 5B93 61F8 4EA2 E279 ABF6 02CF A9AD B7F8 AE4E 425C
--
To unsubscribe from this list: send the line "unsubscribe linux-fsdevel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
From: Pavel Emelyanov <hidden> Date: 2012-07-05 14:09:28
On 07/05/2012 05:07 PM, Nikolaus Rath wrote:
Pavel Emelyanov [off-list ref] writes:
quoted
quoted
While I try to get this to compile,
That's good news! Thanks a lot!
The fuse and fsdevel mainling lists were corrupted in you original message,
so looping the lists back in.
Ok, I got it to compile and tested it by copying a large file into an
S3QL FUSE file system using different block sizes. At first glance,
things look fantastic.
However, I suspect that most of the gain is really because with the
patch most of the data is still in the kernel cache when dd finishes and
hasn't yet been received by the FUSE client.
Is there a way to force flushing of the fuse cache, so that I can
measure the time for dd + final flush?
Actually when dd closes an output file the whole page cache which relates to
it gets flushed to the userspace! This is due to how FUSE write requests are
served.
That said -- you've already measured writeback with flush :)
The fuse and fsdevel mainling lists were corrupted in you original message,
so looping the lists back in.
quoted
Ok, I got it to compile and tested it by copying a large file into an
S3QL FUSE file system using different block sizes. At first glance,
things look fantastic.
However, I suspect that most of the gain is really because with the
patch most of the data is still in the kernel cache when dd finishes and
hasn't yet been received by the FUSE client.
Is there a way to force flushing of the fuse cache, so that I can
measure the time for dd + final flush?
Actually when dd closes an output file the whole page cache which relates to
it gets flushed to the userspace! This is due to how FUSE write requests are
served.
That said -- you've already measured writeback with flush :)
But then how is it possible that there is such a big difference even
when using 128kb blocks?
Kernel 3.2:
131072000 bytes (131 MB) copied, 2.10194 s, 62.4 MB/s
Kernel 3.5-pre:
Blocksize: 128k
131072000 bytes (131 MB) copied, 0.51943 s, 252 MB/s
I would think that it both cases the FUSE daemon gets the data in 128 kb
packets, so why the difference? Or is this due to some other change
between 3.2 and 3.5 that significantly increased FUSE performance?
I can try the same test with the unpatched 3.5 tonight if that would be
of interest.
Best,
-Nikolaus
--
»Time flies like an arrow, fruit flies like a Banana.«
PGP fingerprint: 5B93 61F8 4EA2 E279 ABF6 02CF A9AD B7F8 AE4E 425C
--
To unsubscribe from this list: send the line "unsubscribe linux-fsdevel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
From: Pavel Emelyanov <hidden> Date: 2012-07-05 14:34:59
On 07/05/2012 06:29 PM, Nikolaus Rath wrote:
On 07/05/2012 10:08 AM, Pavel Emelyanov wrote:
quoted
On 07/05/2012 05:07 PM, Nikolaus Rath wrote:
quoted
Pavel Emelyanov [off-list ref] writes:
quoted
quoted
While I try to get this to compile,
That's good news! Thanks a lot!
The fuse and fsdevel mainling lists were corrupted in you original message,
so looping the lists back in.
quoted
Ok, I got it to compile and tested it by copying a large file into an
S3QL FUSE file system using different block sizes. At first glance,
things look fantastic.
However, I suspect that most of the gain is really because with the
patch most of the data is still in the kernel cache when dd finishes and
hasn't yet been received by the FUSE client.
Is there a way to force flushing of the fuse cache, so that I can
measure the time for dd + final flush?
Actually when dd closes an output file the whole page cache which relates to
it gets flushed to the userspace! This is due to how FUSE write requests are
served.
That said -- you've already measured writeback with flush :)
But then how is it possible that there is such a big difference even
when using 128kb blocks?
Kernel 3.2:
131072000 bytes (131 MB) copied, 2.10194 s, 62.4 MB/s
Kernel 3.5-pre:
Blocksize: 128k
131072000 bytes (131 MB) copied, 0.51943 s, 252 MB/s
I would think that it both cases the FUSE daemon gets the data in 128 kb
packets, so why the difference? Or is this due to some other change
between 3.2 and 3.5 that significantly increased FUSE performance?
I don't know for sure. Need to read an analyze the FUSE userspace side
logs (if any).
I can try the same test with the unpatched 3.5 tonight if that would be
of interest.