Thread (26 messages) 26 messages, 5 authors, 1d ago

Re: [PATCH v7 1/3] mm: make persistent huge zero folio read-only

From: Xueyuan Chen <hidden>
Date: 2026-09-07 01:11:47
Also in: linux-mm, lkml

On Sun, Sep 6, 2026 at 6:00 PM Mike Rapoport [off-list ref] wrote:
On Thu, Sep 03, 2026 at 08:49:52PM +0800, Xueyuan Chen wrote:
quoted
On Thu, Sep 3, 2026 at 5:22 PM Mike Rapoport [off-list ref] wrote:
quoted
On Tue, Sep 01, 2026 at 11:18:16PM +0800, Xueyuan Chen wrote:
quoted
The persistent huge zero folio is shared globally and should stay zero
after initialization. As Jann Horn pointed out [1], kernel bugs have
ended up writing to pages that were meant to be read-only, including in
security-sensitive cases. Making the folio read-only in the direct map
turns such writes into faults instead of silent zero-page corruption.

Add a page-based helper consistent with the existing direct-map interfaces.
Handle TLB invalidation in the architecture implementation; unsupported
architectures retain their current behavior.

Protect the folio after initialization. Skip highmem folios, which have no
permanent direct-map mapping.

Inspired by Jann Horn's read-only zero page work [1] and follow-up
discussion [3] with Yang Shi.

Link: https://lore.kernel.org/r/20260508-ro-zeropage-v1-1-9808abc20b49@google.com (local) [1]
Link: https://lore.kernel.org/r/0e5b23a6-4895-454a-9dfa-6dc21adc2991@kernel.org (local) [2]
Link: https://lore.kernel.org/r/CAHbLzkrXXe7r3n3jXgDKtwZhRqj=jDx9E6dLOULohnhBguvi9A@mail.gmail.com (local) [3]

Suggested-by: David Hildenbrand <david@kernel.org>
Suggested-by: Usama Arif <usama.arif@linux.dev>
Co-developed-by: Lance Yang <lance.yang@linux.dev>
Signed-off-by: Lance Yang <lance.yang@linux.dev>
Signed-off-by: Xueyuan Chen <redacted>
---
 include/linux/set_memory.h | 17 +++++++++++++++++
 mm/huge_memory.c           | 13 ++++++++++---
 2 files changed, 27 insertions(+), 3 deletions(-)
diff --git a/include/linux/set_memory.h b/include/linux/set_memory.h
index 3fe293cfed8c..ed9ce04b18a1 100644
--- a/include/linux/set_memory.h
+++ b/include/linux/set_memory.h
@@ -54,6 +54,23 @@ static inline bool can_set_direct_map(void)
 #endif
 #endif /* CONFIG_ARCH_HAS_SET_DIRECT_MAP */

+#ifndef set_direct_map_ro
+/**
+ * set_direct_map_ro - make a direct-map range read-only
+ * @page: first page in the direct-map range
+ * @nr: number of pages in the range
+ *
+ * Make the direct-map range starting at @page read-only and invalidate stale
+ * writable translations before returning.
+ *
+ * Return: 0 on success, or a negative error code on failure.
+ */
+static inline int set_direct_map_ro(struct page *page, unsigned int nr)
+{
+     return 0;
+}
+#endif
+
 #ifdef CONFIG_X86_64
 int set_mce_nospec(unsigned long pfn);
 int clear_mce_nospec(unsigned long pfn);
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 54494c3fa983..742283b36d74 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -42,6 +42,7 @@
 #include <linux/pgalloc_tag.h>
 #include <linux/pagewalk.h>
 #include <linux/cleanup.h>
+#include <linux/set_memory.h>

 #include <asm/tlb.h>
 #include "internal.h"
@@ -291,10 +292,16 @@ static int __init huge_zero_init(void)
      huge_zero_folio = alloc_huge_zero_folio();
      if (!huge_zero_folio) {
              pr_warn("Allocating persistent huge zero folio failed\n");
-     } else {
-             huge_zero_pfn = folio_pfn(huge_zero_folio);
-             count_vm_event(THP_ZERO_PAGE_ALLOC);
+             return 0;
      }
+
+     huge_zero_pfn = folio_pfn(huge_zero_folio);
+     count_vm_event(THP_ZERO_PAGE_ALLOC);
+
+     /* Highmem folios have no permanent direct-map mapping to protect. */
+     if (!folio_test_highmem(huge_zero_folio))
+             set_direct_map_ro(folio_page(huge_zero_folio, 0), HPAGE_PMD_NR);
Sorry, I don't remember if it was discussed previously, but why can't we
use the existing set_memory_ro() here?
Hi Mike,

We want to change the linear map here, but arm64 set_memory_ro() only
works on vmalloc addresses.
I believe this is an historical artifact. change_memory_common() already
updates the linear map when a vmalloc mapping switches to RO and the system
supports it.
On x86, set_memory_ro() does work on direct-map addresses.
I believe arm64::set_memory_ro() can change the linear map in the general
case as well as long as can_set_direct_map() is true.
On arm64, I'm reading arch/arm64/mm/pageattr.c, and the linear-map
update in change_memory_common() is only reachable for vmalloc
addresses:

  set_memory_ro
    change_memory_common
       area = find_vm_area((void *)addr);
       if (!area || ...)
          return -EINVAL;

So set_memory_ro() on a linear-map address returns -EINVAL on arm64.

Am I missing something here?
quoted
So a new helper is needed on arm64, and x86 implements the same
helper to keep the two architectures consistent. On x86 it is
basically set_memory_ro() minus the alias check.

Thanks,
Xueyuan
quoted
quoted
+
      return 0;
 }

--
2.47.3
--
Sincerely yours,
Mike.
--
Sincerely yours,
Mike.
  
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help