Re: [PATCH v11 10/26] mm: protect VMA modifications using VMA sequence count

[PATCH v11 00/26] Speculative page faults · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 01/26] mm: introduce CONFIG_SPECULATIVE_PAGE_FAULT · Laurent Dufour <hidden> · 2018-05-17
Re: [PATCH v11 01/26] mm: introduce CONFIG_SPECULATIVE_PAGE_FAULT · Randy Dunlap <hidden> · 2018-05-17
Re: [PATCH v11 01/26] mm: introduce CONFIG_SPECULATIVE_PAGE_FAULT · Matthew Wilcox <willy@infradead.org> · 2018-05-17
Re: [PATCH v11 01/26] mm: introduce CONFIG_SPECULATIVE_PAGE_FAULT · Randy Dunlap <hidden> · 2018-05-17
[FIX PATCH v11 01/26] mm: introduce CONFIG_SPECULATIVE_PAGE_FAULT · Laurent Dufour <hidden> · 2018-05-22
Re: [PATCH v11 01/26] mm: introduce CONFIG_SPECULATIVE_PAGE_FAULT · Laurent Dufour <hidden> · 2018-05-22
Re: [PATCH v11 01/26] mm: introduce CONFIG_SPECULATIVE_PAGE_FAULT · Laurent Dufour <hidden> · 2018-05-22
[PATCH v11 04/26] arm64/mm: define ARCH_SUPPORTS_SPECULATIVE_PAGE_FAULT · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 07/26] mm: make pte_unmap_same compatible with SPF · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 09/26] mm: VMA sequence count · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 06/26] mm: introduce pte_spinlock for FAULT_FLAG_SPECULATIVE · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 08/26] mm: introduce INIT_VMA() · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 13/26] mm: cache some VMA fields in the vm_fault structure · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 14/26] mm/migrate: Pass vm_fault pointer to migrate_misplaced_page() · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 17/26] mm: introduce __page_add_new_anon_rmap() · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 19/26] mm: provide speculative fault infrastructure · Laurent Dufour <hidden> · 2018-05-17
Re: [PATCH v11 19/26] mm: provide speculative fault infrastructure · zhong jiang <hidden> · 2018-07-24
Re: [PATCH v11 19/26] mm: provide speculative fault infrastructure · Laurent Dufour <hidden> · 2018-07-24
Re: [PATCH v11 19/26] mm: provide speculative fault infrastructure · zhong jiang <hidden> · 2018-07-25
Re: [PATCH v11 19/26] mm: provide speculative fault infrastructure · Laurent Dufour <hidden> · 2018-07-25
Re: [PATCH v11 19/26] mm: provide speculative fault infrastructure · zhong jiang <hidden> · 2018-07-25
[PATCH v11 22/26] perf tools: add support for the SPF perf event · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 26/26] arm64/mm: add speculative page fault · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 25/26] powerpc/mm: add speculative page fault · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 24/26] x86/mm: add speculative pagefault handling · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 10/26] mm: protect VMA modifications using VMA sequence count · Laurent Dufour <hidden> · 2018-05-17
Re: [PATCH v11 10/26] mm: protect VMA modifications using VMA sequence count · vinayak menon <hidden> · 2018-11-05
Re: [PATCH v11 10/26] mm: protect VMA modifications using VMA sequence count · Laurent Dufour <hidden> · 2018-11-05
Re: [PATCH v11 10/26] mm: protect VMA modifications using VMA sequence count · Vinayak Menon <hidden> · 2018-11-06
[PATCH v11 23/26] mm: add speculative page fault vmstats · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 20/26] mm: adding speculative page fault failure trace events · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 21/26] perf: add a speculative page fault sw event · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 18/26] mm: protect mm_rb tree with a rwlock · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 16/26] mm: introduce __vm_normal_page() · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 15/26] mm: introduce __lru_cache_add_active_or_unevictable · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 12/26] mm: protect SPF handler against anon_vma changes · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 11/26] mm: protect mremap() against SPF hanlder · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 05/26] mm: prepare for FAULT_FLAG_SPECULATIVE · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 03/26] powerpc/mm: set ARCH_SUPPORTS_SPECULATIVE_PAGE_FAULT · Laurent Dufour <hidden> · 2018-05-17
[PATCH v11 02/26] x86/mm: define ARCH_SUPPORTS_SPECULATIVE_PAGE_FAULT · Laurent Dufour <hidden> · 2018-05-17
RE: [PATCH v11 00/26] Speculative page faults · Song, HaiyanX <hidden> · 2018-05-28
Re: [PATCH v11 00/26] Speculative page faults · Laurent Dufour <hidden> · 2018-05-28
Re: [PATCH v11 00/26] Speculative page faults · Haiyan Song <hidden> · 2018-05-28
Re: [PATCH v11 00/26] Speculative page faults · Laurent Dufour <hidden> · 2018-05-28
RE: [PATCH v11 00/26] Speculative page faults · Wang, Kemi <hidden> · 2018-05-28
RE: [PATCH v11 00/26] Speculative page faults · Song, HaiyanX <hidden> · 2018-06-11
Re: [PATCH v11 00/26] Speculative page faults · Laurent Dufour <hidden> · 2018-06-11
Re: [PATCH v11 00/26] Speculative page faults · Haiyan Song <hidden> · 2018-06-19
Re: [PATCH v11 00/26] Speculative page faults · Laurent Dufour <hidden> · 2018-07-02
RE: [PATCH v11 00/26] Speculative page faults · Song, HaiyanX <hidden> · 2018-07-04
Re: [PATCH v11 00/26] Speculative page faults · Laurent Dufour <hidden> · 2018-07-04
Re: [PATCH v11 00/26] Speculative page faults · Laurent Dufour <hidden> · 2018-07-11
RE: [PATCH v11 00/26] Speculative page faults · Song, HaiyanX <hidden> · 2018-07-13
Re: [PATCH v11 00/26] Speculative page faults · Laurent Dufour <hidden> · 2018-07-17
RE: [PATCH v11 00/26] Speculative page faults · Song, HaiyanX <hidden> · 2018-08-03
RE: [PATCH v11 00/26] Speculative page faults · Song, HaiyanX <hidden> · 2018-08-03
Re: [PATCH v11 00/26] Speculative page faults · Laurent Dufour <hidden> · 2018-08-22
RE: [PATCH v11 00/26] Speculative page faults · Song, HaiyanX <hidden> · 2018-09-18
Re: [PATCH v11 00/26] Speculative page faults · Balbir Singh <bsingharora@gmail.com> · 2018-11-05
Re: [PATCH v11 00/26] Speculative page faults · Laurent Dufour <hidden> · 2018-11-05

From: Vinayak Menon <hidden>
Date: 2018-11-06 09:29:03
Also in: linux-mm, lkml

On 11/5/2018 11:52 PM, Laurent Dufour wrote:

Le 05/11/2018 à 08:04, vinayak menon a écrit :

quoted

Hi Laurent,

On Thu, May 17, 2018 at 4:37 PM Laurent Dufour
[off-list ref] wrote:

quoted

The VMA sequence count has been introduced to allow fast detection of
VMA modification when running a page fault handler without holding
the mmap_sem.

This patch provides protection against the VMA modification done in :
         - madvise()
         - mpol_rebind_policy()
         - vma_replace_policy()
         - change_prot_numa()
         - mlock(), munlock()
         - mprotect()
         - mmap_region()
         - collapse_huge_page()
         - userfaultd registering services

In addition, VMA fields which will be read during the speculative fault
path needs to be written using WRITE_ONCE to prevent write to be split
and intermediate values to be pushed to other CPUs.

Signed-off-by: Laurent Dufour <redacted>
---
  fs/proc/task_mmu.c |  5 ++++-
  fs/userfaultfd.c   | 17 +++++++++++++----
  mm/khugepaged.c    |  3 +++
  mm/madvise.c       |  6 +++++-
  mm/mempolicy.c     | 51 ++++++++++++++++++++++++++++++++++-----------------
  mm/mlock.c         | 13 ++++++++-----
  mm/mmap.c          | 22 +++++++++++++---------
  mm/mprotect.c      |  4 +++-
  mm/swap_state.c    |  8 ++++++--
  9 files changed, 89 insertions(+), 40 deletions(-)

  struct page *swap_cluster_readahead(swp_entry_t entry, gfp_t gfp_mask,
                                 struct vm_fault *vmf)

@@ -665,9 +669,9 @@ static inline void swap_ra_clamp_pfn(struct vm_area_struct *vma,

                                      unsigned long *start,
                                      unsigned long *end)
  {
-       *start = max3(lpfn, PFN_DOWN(vma->vm_start),
+       *start = max3(lpfn, PFN_DOWN(READ_ONCE(vma->vm_start)),
                       PFN_DOWN(faddr & PMD_MASK));
-       *end = min3(rpfn, PFN_DOWN(vma->vm_end),
+       *end = min3(rpfn, PFN_DOWN(READ_ONCE(vma->vm_end)),
                     PFN_DOWN((faddr & PMD_MASK) + PMD_SIZE));
  }

-- 
2.7.4

I have got a crash on 4.14 kernel with speculative page faults enabled
and here is my analysis of the problem.
The issue was reported only once.

Hi Vinayak,

Thanks for reporting this.

quoted

[23409.303395]  el1_da+0x24/0x84
[23409.303400]  __radix_tree_lookup+0x8/0x90
[23409.303407]  find_get_entry+0x64/0x14c
[23409.303410]  pagecache_get_page+0x5c/0x27c
[23409.303416]  __read_swap_cache_async+0x80/0x260
[23409.303420]  swap_vma_readahead+0x264/0x37c
[23409.303423]  swapin_readahead+0x5c/0x6c
[23409.303428]  do_swap_page+0x128/0x6e4
[23409.303431]  handle_pte_fault+0x230/0xca4
[23409.303435]  __handle_speculative_fault+0x57c/0x7c8
[23409.303438]  do_page_fault+0x228/0x3e8
[23409.303442]  do_translation_fault+0x50/0x6c
[23409.303445]  do_mem_abort+0x5c/0xe0
[23409.303447]  el0_da+0x20/0x24

Process A accesses address ADDR (part of VMA A) and that results in a
translation fault.
Kernel enters __handle_speculative_fault to fix the fault.
Process A enters do_swap_page->swapin_readahead->swap_vma_readahead
from speculative path.
During this time, another process B which shares the same mm, does a
mprotect from another CPU which follows
mprotect_fixup->__split_vma, and it splits VMA A into VMAs A and B.
After the split, ADDR falls into VMA B, but process A is still using
VMA A.
Now ADDR is greater than VMA_A->vm_start and VMA_A->vm_end.
swap_vma_readahead->swap_ra_info uses start and end of vma to
calculate ptes and nr_pte, which goes wrong due to this and finally
resulting in wrong "entry" passed to
swap_vma_readahead->__read_swap_cache_async, and in turn causing
invalid swapper_space
being passed to __read_swap_cache_async->find_get_page, causing an abort.

The fix I have tried is to cache vm_start and vm_end also in vmf and
use it in swap_ra_clamp_pfn. Let me know your thoughts on this. I can
send
the patch I am a using if you feel that is the right thing to do.

I think the best would be to don't do swap readahead during the speculatvive page fault. If the page is found in the swap cache, that's fine, but otherwise, we should f    allback to the regular page fault.

The attached -untested- patch is doing this, if you want to give it a try. I'll review that for the next series.

Thanks Laurent. I and going to try this patch.

With this patch, since all non-SWP_SYNCHRONOUS_IO swapins result in non-speculative fault
and a retry, wouldn't this have an impact on some perf numbers ? If so, would caching start
and end be a better option ?

Also, would it make sense to move the FAULT_FLAG_SPECULATIVE check inside swapin_readahead,
in a way that  swap_cluster_readahead can take the speculative path ? swap_cluster_readahead
doesn't seem to use vma values.

Thanks,
Vinayak

`h`	back out one level
`j`	next message in thread
`k`	previous message in thread
`l`	drill in
`Esc`	close help / fold thread tree
`?`	toggle this help