Re: [PATCH v10 09/21] Replace the XIP page fault handler with the DAX page... | linux-mm

[PATCH v10 00/21] Support ext4 on NV-DIMMs · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 01/21] axonram: Fix bug in direct_access · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 02/21] Change direct_access calling convention · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 03/21] Fix XIP fault vs truncate race · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 04/21] Allow page fault handlers to perform the COW · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 05/21] Introduce IS_DAX(inode) · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 06/21] Add copy_to_iter(), copy_from_iter() and iov_iter_zero() · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 07/21] Replace XIP read and write with DAX I/O · Matthew Wilcox <hidden> · 2014-08-27
Re: [PATCH v10 07/21] Replace XIP read and write with DAX I/O · Boaz Harrosh <hidden> · 2014-09-14
[PATCH v10 08/21] Replace ext2_clear_xip_target with dax_clear_blocks · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 09/21] Replace the XIP page fault handler with the DAX page fault handler · Matthew Wilcox <hidden> · 2014-08-27
Re: [PATCH v10 09/21] Replace the XIP page fault handler with the DAX page fault handler · Dave Chinner <david@fromorbit.com> · 2014-09-03
Re: [PATCH v10 09/21] Replace the XIP page fault handler with the DAX page fault handler · Matthew Wilcox <hidden> · 2014-09-10
Re: [PATCH v10 09/21] Replace the XIP page fault handler with the DAX page fault handler · Dave Chinner <david@fromorbit.com> · 2014-09-11
Re: [PATCH v10 09/21] Replace the XIP page fault handler with the DAX page fault handler · Matthew Wilcox <hidden> · 2014-09-24
Re: [PATCH v10 09/21] Replace the XIP page fault handler with the DAX page fault handler · Dave Chinner <david@fromorbit.com> · 2014-09-25
[PATCH v10 10/21] Replace xip_truncate_page with dax_truncate_page · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 11/21] Replace XIP documentation with DAX documentation · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 12/21] Remove get_xip_mem · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 13/21] ext2: Remove ext2_xip_verify_sb() · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 14/21] ext2: Remove ext2_use_xip · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 15/21] ext2: Remove xip.c and xip.h · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 16/21] Remove CONFIG_EXT2_FS_XIP and rename CONFIG_FS_XIP to CONFIG_FS_DAX · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 17/21] ext2: Remove ext2_aops_xip · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 18/21] Get rid of most mentions of XIP in ext2 · Matthew Wilcox <hidden> · 2014-08-27
[PATCH v10 19/21] xip: Add xip_zero_page_range · Matthew Wilcox <hidden> · 2014-08-27
Re: [PATCH v10 19/21] xip: Add xip_zero_page_range · Dave Chinner <david@fromorbit.com> · 2014-09-03
Re: [PATCH v10 19/21] xip: Add xip_zero_page_range · Matthew Wilcox <hidden> · 2014-09-04
Re: [PATCH v10 19/21] xip: Add xip_zero_page_range · Theodore Ts'o <tytso@mit.edu> · 2014-09-04
Re: [PATCH v10 19/21] xip: Add xip_zero_page_range · Matthew Wilcox <hidden> · 2014-09-08
[PATCH v10 20/21] ext4: Add DAX functionality · Matthew Wilcox <hidden> · 2014-08-27
Re: [PATCH v10 20/21] ext4: Add DAX functionality · Dave Chinner <david@fromorbit.com> · 2014-09-03
Re: [PATCH v10 20/21] ext4: Add DAX functionality · Boaz Harrosh <hidden> · 2014-09-10
Re: [PATCH v10 20/21] ext4: Add DAX functionality · Dave Chinner <david@fromorbit.com> · 2014-09-11
Re: [PATCH v10 20/21] ext4: Add DAX functionality · Boaz Harrosh <hidden> · 2014-09-14
Re: [PATCH v10 20/21] ext4: Add DAX functionality · Dave Chinner <david@fromorbit.com> · 2014-09-15
Re: [PATCH v10 20/21] ext4: Add DAX functionality · Boaz Harrosh <hidden> · 2014-09-15
[PATCH v10 21/21] brd: Rename XIP to DAX · Matthew Wilcox <hidden> · 2014-08-27
Re: [PATCH v10 00/21] Support ext4 on NV-DIMMs · Andrew Morton <akpm@linux-foundation.org> · 2014-08-27
Re: [PATCH v10 00/21] Support ext4 on NV-DIMMs · Matthew Wilcox <hidden> · 2014-08-27
Re: [PATCH v10 00/21] Support ext4 on NV-DIMMs · Andrew Morton <akpm@linux-foundation.org> · 2014-08-27
Re: [PATCH v10 00/21] Support ext4 on NV-DIMMs · Andy Lutomirski <luto@amacapital.net> · 2014-08-28
Re: [PATCH v10 00/21] Support ext4 on NV-DIMMs · Matthew Wilcox <hidden> · 2014-08-28
Re: [PATCH v10 00/21] Support ext4 on NV-DIMMs · Matthew Wilcox <hidden> · 2014-08-28
Re: [PATCH v10 00/21] Support ext4 on NV-DIMMs · Christoph Lameter <hidden> · 2014-08-27
Re: [PATCH v10 00/21] Support ext4 on NV-DIMMs · Andrew Morton <akpm@linux-foundation.org> · 2014-08-27
Re: [PATCH v10 00/21] Support ext4 on NV-DIMMs · One Thousand Gnomes <hidden> · 2014-08-27
Re: [PATCH v10 00/21] Support ext4 on NV-DIMMs · Dave Chinner <david@fromorbit.com> · 2014-08-28
Re: [PATCH v10 00/21] Support ext4 on NV-DIMMs · Christian Stroetmann <hidden> · 2014-08-30
Re: [PATCH v10 00/21] Support ext4 on NV-DIMMs · Boaz Harrosh <hidden> · 2014-08-28
Re: [PATCH v10 00/21] Support ext4 on NV-DIMMs · Zwisler, Ross <hidden> · 2014-08-28
[PATCH 1/1] xfs: add DAX support · Dave Chinner <david@fromorbit.com> · 2014-09-03

Re: [PATCH v10 09/21] Replace the XIP page fault handler with the DAX page fault handler

From: Dave Chinner <david@fromorbit.com>
Date: 2014-09-11 03:09:26
Also in: linux-fsdevel, lkml

On Wed, Sep 10, 2014 at 11:23:37AM -0400, Matthew Wilcox wrote:

On Wed, Sep 03, 2014 at 05:47:24PM +1000, Dave Chinner wrote:

quoted

+	error = get_block(inode, block, &bh, 0);
+	if (!error && (bh.b_size < PAGE_SIZE))
+		error = -EIO;
+	if (error)
+		goto unlock_page;

page fault into unwritten region, returns buffer_unwritten(bh) ==
true. Hence buffer_written(bh) is false, and we take this branch:

quoted

+	if (!buffer_written(&bh) && !vmf->cow_page) {
+		if (vmf->flags & FAULT_FLAG_WRITE) {
+			error = get_block(inode, block, &bh, 1);

Exactly what are you expecting to happen here? We don't do
allocation because there are already unwritten blocks over this
extent, and so bh will be unchanged when returning. i.e. it will
still be mapping an unwritten extent.

I was expecting calling get_block() on an unwritten extent to convert it
to a written extent.  Your suggestion below of using b_end_io() to do that
is a better idea.

So this should be:

	if (!buffer_mapped(&bh) && !vmf->cow_page) {

... right?

Yes, that is the conclusion I reached as well. ;)

quoted

dax: add IO completion callback for page faults

From: Dave Chinner <redacted>

When a page fault drops into a hole, it needs to allocate an extent.
Filesystems may allocate unwritten extents so that the underlying
contents are not exposed until data is written to the extent. In
that case, we need an io completion callback to run once the blocks
have been zeroed to indicate that it is safe for the filesystem to
mark those blocks written without exposing stale data in the event
of a crash.

Signed-off-by: Dave Chinner <redacted>
---
 fs/dax.c | 7 ++++++-
 1 file changed, 6 insertions(+), 1 deletion(-)

diff --git a/fs/dax.c b/fs/dax.c
index 96c4fed..387ca78 100644
--- a/fs/dax.c
+++ b/fs/dax.c

@@ -306,6 +306,7 @@ static int do_dax_fault(struct vm_area_struct *vma, struct vm_fault *vmf,
 	memset(&bh, 0, sizeof(bh));
 	block = (sector_t)vmf->pgoff << (PAGE_SHIFT - blkbits);
 	bh.b_size = PAGE_SIZE;
+	bh.b_end_io = NULL;

Given the above memset, I don't think we need to explicitly set b_end_io
to NULL.

I missed that ;)

quoted

  repeat:
 	page = find_get_page(mapping, vmf->pgoff);

@@ -364,8 +365,12 @@ static int do_dax_fault(struct vm_area_struct *vma, struct vm_fault *vmf,
 		return VM_FAULT_LOCKED;
 	}
 
-	if (buffer_unwritten(&bh) || buffer_new(&bh))
+	if (buffer_unwritten(&bh) || buffer_new(&bh)) {
+		/* XXX: errors zeroing the blocks are propagated how? */
 		dax_clear_blocks(inode, bh.b_blocknr, bh.b_size);

That's a great question.  I think we need to segfault here.

I suspect there are other cases where we need to do similar "trigger
segv" error handling rather than ignoring errors altogether...

quoted

+		if (bh.b_end_io)
+			bh.b_end_io(&bh, 1);
+	}

I think ext4 is going to need to set b_end_io too.  Right now, it uses the
dio_iodone_t to convert unwritten extents to written extents, but we don't
have (and I don't think we should have) a kiocb for page faults.

Yes, ext4 is going to need this as well. After I got XFS running
without problems, I then went back and ran xfstests on ext4 and it
failed many of the tests that do operations into unwritten regions.

So, if it's OK with you, I'm going to fold this patch into version 11 and
add your Reviewed-by to it.

Fold it in, I'll review the result ;)

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org.  For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>

`h`	back out one level
`j`	next message in thread
`k`	previous message in thread
`l`	drill in
`Esc`	close help / fold thread tree
`?`	toggle this help