Re: [PATCH v10 09/25] mm: protect VMA modifications using VMA sequence count

[PATCH v10 00/25] Speculative page faults · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 03/25] powerpc/mm: set ARCH_SUPPORTS_SPECULATIVE_PAGE_FAULT · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 02/25] x86/mm: define ARCH_SUPPORTS_SPECULATIVE_PAGE_FAULT · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 04/25] mm: prepare for FAULT_FLAG_SPECULATIVE · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 06/25] mm: make pte_unmap_same compatible with SPF · Laurent Dufour <hidden> · 2018-04-17
Re: [PATCH v10 06/25] mm: make pte_unmap_same compatible with SPF · Minchan Kim <minchan@kernel.org> · 2018-04-23
Re: [PATCH v10 06/25] mm: make pte_unmap_same compatible with SPF · Laurent Dufour <hidden> · 2018-04-30
Re: [PATCH v10 06/25] mm: make pte_unmap_same compatible with SPF · Minchan Kim <minchan@kernel.org> · 2018-05-01
Re: [PATCH v10 06/25] mm: make pte_unmap_same compatible with SPF · vinayak menon <hidden> · 2018-05-10
Re: [PATCH v10 06/25] mm: make pte_unmap_same compatible with SPF · Laurent Dufour <hidden> · 2018-05-14
[PATCH v10 07/25] mm: introduce INIT_VMA() · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 08/25] mm: VMA sequence count · Laurent Dufour <hidden> · 2018-04-17
Re: [PATCH v10 08/25] mm: VMA sequence count · Minchan Kim <minchan@kernel.org> · 2018-04-23
Re: [PATCH v10 08/25] mm: VMA sequence count · Laurent Dufour <hidden> · 2018-04-30
Re: [PATCH v10 08/25] mm: VMA sequence count · Minchan Kim <minchan@kernel.org> · 2018-05-01
Re: [PATCH v10 08/25] mm: VMA sequence count · Laurent Dufour <hidden> · 2018-05-03
[PATCH v10 10/25] mm: protect mremap() against SPF hanlder · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 09/25] mm: protect VMA modifications using VMA sequence count · Laurent Dufour <hidden> · 2018-04-17
Re: [PATCH v10 09/25] mm: protect VMA modifications using VMA sequence count · Minchan Kim <minchan@kernel.org> · 2018-04-23
Re: [PATCH v10 09/25] mm: protect VMA modifications using VMA sequence count · Laurent Dufour <hidden> · 2018-05-14
[PATCH v10 13/25] mm/migrate: Pass vm_fault pointer to migrate_misplaced_page() · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 15/25] mm: introduce __vm_normal_page() · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 17/25] mm: protect mm_rb tree with a rwlock · Laurent Dufour <hidden> · 2018-04-17
Re: [PATCH v10 17/25] mm: protect mm_rb tree with a rwlock · Punit Agrawal <hidden> · 2018-04-30
Re: [PATCH v10 17/25] mm: protect mm_rb tree with a rwlock · Laurent Dufour <hidden> · 2018-05-02
[PATCH v10 19/25] mm: adding speculative page fault failure trace events · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 20/25] perf: add a speculative page fault sw event · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 23/25] mm: add speculative page fault vmstats · Laurent Dufour <hidden> · 2018-04-17
Re: [PATCH v10 23/25] mm: add speculative page fault vmstats · Ganesh Mahendran <hidden> · 2018-05-16
Re: [PATCH v10 23/25] mm: add speculative page fault vmstats · Laurent Dufour <hidden> · 2018-05-16
[PATCH v10 22/25] mm: speculative page fault handler return VMA · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 25/25] powerpc/mm: add speculative page fault · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 24/25] x86/mm: add speculative pagefault handling · Laurent Dufour <hidden> · 2018-04-17
Re: [PATCH v10 24/25] x86/mm: add speculative pagefault handling · Punit Agrawal <hidden> · 2018-04-30
Re: [PATCH v10 24/25] x86/mm: add speculative pagefault handling · Laurent Dufour <hidden> · 2018-05-03
[PATCH v10 21/25] perf tools: add support for the SPF perf event · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 18/25] mm: provide speculative fault infrastructure · Laurent Dufour <hidden> · 2018-04-17
Re: [PATCH v10 18/25] mm: provide speculative fault infrastructure · vinayak menon <hidden> · 2018-05-15
Re: [PATCH v10 18/25] mm: provide speculative fault infrastructure · Laurent Dufour <hidden> · 2018-05-15
[PATCH v10 16/25] mm: introduce __page_add_new_anon_rmap() · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 14/25] mm: introduce __lru_cache_add_active_or_unevictable · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 12/25] mm: cache some VMA fields in the vm_fault structure · Laurent Dufour <hidden> · 2018-04-17
Re: [PATCH v10 12/25] mm: cache some VMA fields in the vm_fault structure · Minchan Kim <minchan@kernel.org> · 2018-04-23
Re: [PATCH v10 12/25] mm: cache some VMA fields in the vm_fault structure · Laurent Dufour <hidden> · 2018-05-03
Re: [PATCH v10 12/25] mm: cache some VMA fields in the vm_fault structure · Minchan Kim <minchan@kernel.org> · 2018-05-03
Re: [PATCH v10 12/25] mm: cache some VMA fields in the vm_fault structure · Laurent Dufour <hidden> · 2018-05-04
Re: [PATCH v10 12/25] mm: cache some VMA fields in the vm_fault structure · Minchan Kim <minchan@kernel.org> · 2018-05-08
[PATCH v10 11/25] mm: protect SPF handler against anon_vma changes · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 05/25] mm: introduce pte_spinlock for FAULT_FLAG_SPECULATIVE · Laurent Dufour <hidden> · 2018-04-17
[PATCH v10 01/25] mm: introduce CONFIG_SPECULATIVE_PAGE_FAULT · Laurent Dufour <hidden> · 2018-04-17
Re: [PATCH v10 01/25] mm: introduce CONFIG_SPECULATIVE_PAGE_FAULT · Minchan Kim <minchan@kernel.org> · 2018-04-23
Re: [PATCH v10 01/25] mm: introduce CONFIG_SPECULATIVE_PAGE_FAULT · Laurent Dufour <hidden> · 2018-04-23

From: Minchan Kim <minchan@kernel.org>
Date: 2018-04-23 07:19:54
Also in: linux-mm, lkml

On Tue, Apr 17, 2018 at 04:33:15PM +0200, Laurent Dufour wrote:

quoted hunk ↗ jump to hunk

The VMA sequence count has been introduced to allow fast detection of
VMA modification when running a page fault handler without holding
the mmap_sem.

This patch provides protection against the VMA modification done in :
	- madvise()
	- mpol_rebind_policy()
	- vma_replace_policy()
	- change_prot_numa()
	- mlock(), munlock()
	- mprotect()
	- mmap_region()
	- collapse_huge_page()
	- userfaultd registering services

In addition, VMA fields which will be read during the speculative fault
path needs to be written using WRITE_ONCE to prevent write to be split
and intermediate values to be pushed to other CPUs.

Signed-off-by: Laurent Dufour <redacted>
---
 fs/proc/task_mmu.c |  5 ++++-
 fs/userfaultfd.c   | 17 +++++++++++++----
 mm/khugepaged.c    |  3 +++
 mm/madvise.c       |  6 +++++-
 mm/mempolicy.c     | 51 ++++++++++++++++++++++++++++++++++-----------------
 mm/mlock.c         | 13 ++++++++-----
 mm/mmap.c          | 22 +++++++++++++---------
 mm/mprotect.c      |  4 +++-
 mm/swap_state.c    |  8 ++++++--
 9 files changed, 89 insertions(+), 40 deletions(-)

diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
index c486ad4b43f0..aeb417f28839 100644
--- a/fs/proc/task_mmu.c
+++ b/fs/proc/task_mmu.c

@@ -1136,8 +1136,11 @@ static ssize_t clear_refs_write(struct file *file, const char __user *buf,
 					goto out_mm;
 				}
 				for (vma = mm->mmap; vma; vma = vma->vm_next) {
-					vma->vm_flags &= ~VM_SOFTDIRTY;
+					vm_write_begin(vma);
+					WRITE_ONCE(vma->vm_flags,
+						   vma->vm_flags & ~VM_SOFTDIRTY);
 					vma_set_page_prot(vma);
+					vm_write_end(vma);

trivial:

I think It's tricky to maintain that VMA fields to be read during SPF should be
(READ|WRITE_ONCE). I think we need some accessor to read/write them rather than
raw accessing like like vma_set_page_prot. Maybe spf prefix would be helpful. 

	vma_spf_set_value(vma, vm_flags, val);

We also add some markers in vm_area_struct's fileds to indicate that
people shouldn't access those fields directly.

Just a thought.

 				}
 				downgrade_write(&mm->mmap_sem);

quoted hunk ↗ jump to hunk

diff --git a/mm/swap_state.c b/mm/swap_state.c
index fe079756bb18..8a8a402ed59f 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c

@@ -575,6 +575,10 @@ static unsigned long swapin_nr_pages(unsigned long offset)
  * the readahead.
  *
  * Caller must hold down_read on the vma->vm_mm if vmf->vma is not NULL.
+ * This is needed to ensure the VMA will not be freed in our back. In the case
+ * of the speculative page fault handler, this cannot happen, even if we don't
+ * hold the mmap_sem. Callees are assumed to take care of reading VMA's fields

I guess reader would be curious on *why* is safe with SPF.
Comment about the why could be helpful for reviewer.

quoted hunk ↗ jump to hunk

+ * using READ_ONCE() to read consistent values.
  */
 struct page *swap_cluster_readahead(swp_entry_t entry, gfp_t gfp_mask,
 				struct vm_fault *vmf)

@@ -668,9 +672,9 @@ static inline void swap_ra_clamp_pfn(struct vm_area_struct *vma,
 				     unsigned long *start,
 				     unsigned long *end)
 {
-	*start = max3(lpfn, PFN_DOWN(vma->vm_start),
+	*start = max3(lpfn, PFN_DOWN(READ_ONCE(vma->vm_start)),
 		      PFN_DOWN(faddr & PMD_MASK));
-	*end = min3(rpfn, PFN_DOWN(vma->vm_end),
+	*end = min3(rpfn, PFN_DOWN(READ_ONCE(vma->vm_end)),
 		    PFN_DOWN((faddr & PMD_MASK) + PMD_SIZE));
 }

-- 
2.7.4

`h`	back out one level
`j`	next message in thread
`k`	previous message in thread
`l`	drill in
`Esc`	close help / fold thread tree
`?`	toggle this help