[PATCH 0/6] mm: Unify Device DAX and HugeTLB vmemmap population paths

COOLING14d

Revision v1 of 3 in this series.

18 messages, 3 authors, 14d ago · open the first message on its own page

[PATCH 0/6] mm: Unify Device DAX and HugeTLB vmemmap population paths

From: Muchun Song <hidden>
Date: 2026-09-13 08:40:24

This series is split out from the earlier, larger series "mm: Generalize
HVO for HugeTLB and device DAX" [1]. While the parent series generalizes
vmemmap optimization across HugeTLB and device DAX, this subset addresses
a single, self-contained step: unifying their vmemmap population paths.

After the preceding Device DAX conversion, both HugeTLB and Device DAX
describe optimized vmemmap mappings through memory-section metadata and
use per-zone shared tail vmemmap pages. The generic code, however, still
carries a Device DAX-specific population flag and compound-page population
path, along with arguments and helpers needed only by that path.

This series first removes VMEMMAP_POPULATE_DAX and moves selection and
reference handling for the shared tail page into the common vmemmap
population path. It then removes the generic Device DAX-specific
compound-page population path and routes section vmemmap population
through vmemmap_populate(). The powerpc radix implementation retains its
architecture-specific compound-page population and selects it directly
for optimizable sections.

The remaining patches remove the unused ptpfn argument, make the
powerpc compound-page helper local to radix_pgtable.c, and open-code
vmemmap_populate_address() now that no caller needs its returned PTE.

This is the fourth smaller step toward the broader HVO generalization.
After this series, HugeTLB and Device DAX use the same population model
instead of parallel generic paths, while powerpc keeps its
architecture-specific implementation.

[1] https://lore.kernel.org/all/20260513130542.35604-1-songmuchun@bytedance.com/

Muchun Song (6):
  mm/sparse-vmemmap: drop VMEMMAP_POPULATE_DAX
  mm/sparse-vmemmap: support device DAX in common vmemmap path
  mm/sparse-vmemmap: drop Device DAX-specific population path
  mm/sparse-vmemmap: remove the unused ptpfn argument
  powerpc/mm: make vmemmap_populate_compound_pages() static
  mm/sparse-vmemmap: open-code vmemmap_populate_address()

 arch/powerpc/include/asm/book3s/64/radix.h |   6 -
 arch/powerpc/mm/book3s64/radix_pgtable.c   |  14 +-
 mm/mm_init.c                               |   2 +-
 mm/sparse-vmemmap.c                        | 188 +++++----------------
 4 files changed, 50 insertions(+), 160 deletions(-)


base-commit: 383fc05d4650b021f3c17e36a145106dcc61a294
-- 
2.54.0

[PATCH 1/6] mm/sparse-vmemmap: drop VMEMMAP_POPULATE_DAX

From: Muchun Song <hidden>
Date: 2026-09-13 08:40:31

VMEMMAP_POPULATE_DAX currently distinguishes DAX vmemmap population in two
places: it keeps allocations on the normal path and takes a reference when
a backing page is supplied for reuse.

After Device DAX switched to the common per-zone shared tail page, both
conditions can be determined locally. DAX supplies ptpfn for every shared
tail mapping and requests an allocation only for compound head mappings,
whose PFNs are not optimizable. Therefore, vmemmap_optimizable_pfn() alone
selects the correct allocation path.

When ptpfn is supplied, the caller is reusing an existing backing page.
Once the slab allocator is available, take a reference for each reused
mapping to balance the release performed by vmemmap_free(). Early mappings
are backed by memblock/reserved memory and do not need page reference
accounting.

Remove VMEMMAP_POPULATE_DAX and the flags argument from the vmemmap
population helpers.

Signed-off-by: Muchun Song <redacted>
---
 mm/sparse-vmemmap.c | 36 +++++++++++++-----------------------
 1 file changed, 13 insertions(+), 23 deletions(-)
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 96506f594924..878d29a4e862 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -32,12 +32,6 @@
 #include <asm/dma.h>
 #include <asm/tlbflush.h>
 
-/*
- * Flags for vmemmap_populate_range and friends.
- */
-/* Vmemmap population for ZONE_DEVICE compound pages */
-#define VMEMMAP_POPULATE_DAX		0x0001
-
 #include "internal.h"
 #include "mm_init.h"
 #include "sparse.h"
@@ -208,7 +202,7 @@ struct page __ref *vmemmap_shared_tail_page(unsigned int order, struct zone *zon
 }
 
 static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
-		struct vmem_altmap *altmap, unsigned long flags)
+		struct vmem_altmap *altmap)
 {
 	struct zone *zone;
 	struct page *page;
@@ -218,7 +212,7 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
 	 * Device DAX still relies on vmemmap_populate_compound_pages() for
 	 * head/first-tail allocation and tail-page reuse.
 	 */
-	if (!vmemmap_optimizable_pfn(pfn) || flags & VMEMMAP_POPULATE_DAX)
+	if (!vmemmap_optimizable_pfn(pfn))
 		return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
 
 	zone = pfn_to_zone(pfn, node);
@@ -230,8 +224,7 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
 }
 
 static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, int node,
-				       struct vmem_altmap *altmap,
-				       unsigned long ptpfn, unsigned long flags)
+		struct vmem_altmap *altmap, unsigned long ptpfn)
 {
 	pte_t *pte = pte_offset_kernel(pmd, addr);
 	unsigned long pfn = page_to_pfn((struct page *)addr);
@@ -240,7 +233,7 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
 		pte_t entry;
 
 		if (ptpfn == (unsigned long)-1) {
-			void *p = vmemmap_alloc_pte(pfn, node, altmap, flags);
+			void *p = vmemmap_alloc_pte(pfn, node, altmap);
 
 			if (!p)
 				return NULL;
@@ -255,7 +248,7 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
 			 * and through vmemmap_populate_compound_pages() when
 			 * slab is available.
 			 */
-			if (flags & VMEMMAP_POPULATE_DAX)
+			if (slab_is_available())
 				get_page(pfn_to_page(ptpfn));
 		}
 		entry = pfn_pte(ptpfn, PAGE_KERNEL);
@@ -318,8 +311,7 @@ static pgd_t * __meminit vmemmap_pgd_populate(unsigned long addr, int node)
 
 static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
 					      struct vmem_altmap *altmap,
-					      unsigned long ptpfn,
-					      unsigned long flags)
+					      unsigned long ptpfn)
 {
 	pgd_t *pgd;
 	p4d_t *p4d;
@@ -339,7 +331,7 @@ static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
 	pmd = vmemmap_pmd_populate(pud, addr, node);
 	if (!pmd)
 		return NULL;
-	pte = vmemmap_pte_populate(pmd, addr, node, altmap, ptpfn, flags);
+	pte = vmemmap_pte_populate(pmd, addr, node, altmap, ptpfn);
 	if (!pte)
 		return NULL;
 	vmemmap_verify(pte, node, addr, addr + PAGE_SIZE);
@@ -350,15 +342,14 @@ static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
 static int __meminit vmemmap_populate_range(unsigned long start,
 					    unsigned long end, int node,
 					    struct vmem_altmap *altmap,
-					    unsigned long ptpfn,
-					    unsigned long flags)
+					    unsigned long ptpfn)
 {
 	unsigned long addr = start;
 	pte_t *pte;
 
 	for (; addr < end; addr += PAGE_SIZE) {
 		pte = vmemmap_populate_address(addr, node, altmap,
-					       ptpfn, flags);
+					       ptpfn);
 		if (!pte)
 			return -ENOMEM;
 	}
@@ -369,7 +360,7 @@ static int __meminit vmemmap_populate_range(unsigned long start,
 int __meminit vmemmap_populate_basepages(unsigned long start, unsigned long end,
 					 int node, struct vmem_altmap *altmap)
 {
-	return vmemmap_populate_range(start, end, node, altmap, -1, 0);
+	return vmemmap_populate_range(start, end, node, altmap, -1);
 }
 
 /*
@@ -498,7 +489,6 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
 	unsigned long size, addr;
 	pte_t *pte;
 	int rc;
-	unsigned long flags = VMEMMAP_POPULATE_DAX;
 	struct page *page;
 	unsigned int order = pfn_to_section_compound_order(start_pfn);
 
@@ -508,14 +498,14 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
 
 	if (reuse_compound_section(start_pfn, pgmap))
 		return vmemmap_populate_range(start, end, node, NULL,
-					      page_to_pfn(page), flags);
+					      page_to_pfn(page));
 
 	size = min(end - start, (1UL << order) * sizeof(struct page));
 	for (addr = start; addr < end; addr += size) {
 		unsigned long next, last = addr + size;
 
 		/* Populate the head page vmemmap page */
-		pte = vmemmap_populate_address(addr, node, NULL, -1, flags);
+		pte = vmemmap_populate_address(addr, node, NULL, -1);
 		if (!pte)
 			return -ENOMEM;
 
@@ -525,7 +515,7 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
 		 */
 		next = addr + PAGE_SIZE;
 		rc = vmemmap_populate_range(next, last, node, NULL,
-					    page_to_pfn(page), flags);
+					    page_to_pfn(page));
 		if (rc)
 			return -ENOMEM;
 	}
-- 
2.54.0

[PATCH 2/6] mm/sparse-vmemmap: support device DAX in common vmemmap path

From: Muchun Song <hidden>
Date: 2026-09-13 08:40:41

The common vmemmap population path cannot yet handle optimized Device DAX
mappings on its own. It uses pfn_to_zone() to find the shared tail page,
but Device DAX populates its vmemmap at runtime before the ZONE_DEVICE span
is initialized.

Teach the common path to use device_zone() for runtime optimized vmemmap
population while retaining pfn_to_zone() for early boot. This allows the
same path to support both early boot mappings and Device DAX.

The backing PFN supplied by the Device DAX-specific population path is no
longer used, allowing the redundant lookup and population code to be
removed later.

Signed-off-by: Muchun Song <redacted>
---
 mm/sparse-vmemmap.c | 44 +++++++++++++++++++-------------------------
 1 file changed, 19 insertions(+), 25 deletions(-)
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 878d29a4e862..e83821768c12 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -208,18 +208,27 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
 	struct page *page;
 	const unsigned int order = pfn_to_section_compound_order(pfn);
 
-	/*
-	 * Device DAX still relies on vmemmap_populate_compound_pages() for
-	 * head/first-tail allocation and tail-page reuse.
-	 */
 	if (!vmemmap_optimizable_pfn(pfn))
 		return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
 
-	zone = pfn_to_zone(pfn, node);
+	/*
+	 * At runtime (slab available), only ZONE_DEVICE pages trigger vmemmap
+	 * optimization, so device_zone() suffices. Note that pfn_to_zone()
+	 * cannot be used at runtime because the zone span is not set up now.
+	 */
+	zone = slab_is_available() ? device_zone(node) : pfn_to_zone(pfn, node);
 	page = vmemmap_shared_tail_page(order, zone);
 	if (!page)
 		return NULL;
 
+	/*
+	 * When a PTE entry is freed, a free_pages() call occurs. This get_page()
+	 * pairs with put_page_testzero() on the freeing path. This can only occur
+	 * when slab is available.
+	 */
+	if (slab_is_available())
+		get_page(page);
+
 	return page_address(page);
 }
 
@@ -231,27 +240,12 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
 
 	if (pte_none(ptep_get(pte))) {
 		pte_t entry;
+		void *p = vmemmap_alloc_pte(pfn, node, altmap);
 
-		if (ptpfn == (unsigned long)-1) {
-			void *p = vmemmap_alloc_pte(pfn, node, altmap);
-
-			if (!p)
-				return NULL;
-			ptpfn = PHYS_PFN(__pa(p));
-		} else {
-			/*
-			 * When a PTE/PMD entry is freed from the init_mm
-			 * there's a free_pages() call to this page allocated
-			 * above. Thus this get_page() is paired with the
-			 * put_page_testzero() on the freeing path.
-			 * This can only called by certain ZONE_DEVICE path,
-			 * and through vmemmap_populate_compound_pages() when
-			 * slab is available.
-			 */
-			if (slab_is_available())
-				get_page(pfn_to_page(ptpfn));
-		}
-		entry = pfn_pte(ptpfn, PAGE_KERNEL);
+		if (!p)
+			return NULL;
+
+		entry = pfn_pte(PHYS_PFN(__pa(p)), PAGE_KERNEL);
 		set_pte_at(&init_mm, addr, pte, entry);
 	} else if (WARN_ON_ONCE(vmemmap_optimizable_pfn(pfn)))
 		return NULL;
-- 
2.54.0

[PATCH 3/6] mm/sparse-vmemmap: drop Device DAX-specific population path

From: Muchun Song <hidden>
Date: 2026-09-13 08:40:49

The common vmemmap path selects the shared page for optimized mappings
itself, so Device DAX no longer needs vmemmap_populate_compound_pages()
to find a shared tail page and pass its backing PFN through the generic
population helpers.

Remove the Device DAX-specific population path and let section memmap
population always use vmemmap_populate(). The powerpc retains an
architecture-specific compound-page implementation, so select it directly
from radix__vmemmap_populate() for optimizable sections.

Signed-off-by: Muchun Song <redacted>
---
 arch/powerpc/mm/book3s64/radix_pgtable.c |  3 +
 mm/mm_init.c                             |  2 +-
 mm/sparse-vmemmap.c                      | 71 +-----------------------
 3 files changed, 5 insertions(+), 71 deletions(-)
diff --git a/arch/powerpc/mm/book3s64/radix_pgtable.c b/arch/powerpc/mm/book3s64/radix_pgtable.c
index 9ca28e4a610a..cb72d9ccf747 100644
--- a/arch/powerpc/mm/book3s64/radix_pgtable.c
+++ b/arch/powerpc/mm/book3s64/radix_pgtable.c
@@ -1122,7 +1122,10 @@ int __meminit radix__vmemmap_populate(unsigned long start, unsigned long end, in
 	pud_t *pud;
 	pmd_t *pmd;
 	pte_t *pte;
+	unsigned long pfn = page_to_pfn((struct page *)start);
 
+	if (section_vmemmap_optimizable(__pfn_to_section(pfn)))
+		return vmemmap_populate_compound_pages(pfn, start, end, node, NULL);
 	/*
 	 * If altmap is present, Make sure we align the start vmemmap addr
 	 * to PAGE_SIZE so that we calculate the correct start_pfn in
diff --git a/mm/mm_init.c b/mm/mm_init.c
index 56bb4567a494..1650d6bc1211 100644
--- a/mm/mm_init.c
+++ b/mm/mm_init.c
@@ -1046,7 +1046,7 @@ static void zone_device_page_init_from_template(struct page *page,
  * initialize is a lot smaller that the total amount of struct pages being
  * mapped. This is a paired / mild layering violation with explicit knowledge
  * of how the sparse_vmemmap internals handle compound pages in the lack
- * of an altmap. See vmemmap_populate_compound_pages().
+ * of an altmap.
  */
 static inline unsigned long compound_nr_pages(unsigned long pfn,
 					      struct dev_pagemap *pgmap)
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index e83821768c12..028f844c90f5 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -454,71 +454,6 @@ int __meminit vmemmap_populate_hugepages(unsigned long start, unsigned long end,
 	return 0;
 }
 
-#ifndef vmemmap_populate_compound_pages
-/*
- * For compound pages bigger than section size (e.g. x86 1G compound
- * pages with 2M subsection size) fill the rest of sections as tail
- * pages.
- *
- * Note that memremap_pages() resets @nr_range value and will increment
- * it after each range successful onlining. Thus the value or @nr_range
- * at section memmap populate corresponds to the in-progress range
- * being onlined here.
- */
-static bool __meminit reuse_compound_section(unsigned long start_pfn,
-					     struct dev_pagemap *pgmap)
-{
-	unsigned long nr_pages = pgmap_vmemmap_nr(pgmap);
-	unsigned long offset = start_pfn -
-		PHYS_PFN(pgmap->ranges[pgmap->nr_range].start);
-
-	return !IS_ALIGNED(offset, nr_pages) && nr_pages > PAGES_PER_SUBSECTION;
-}
-
-static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
-						     unsigned long start,
-						     unsigned long end, int node,
-						     struct dev_pagemap *pgmap)
-{
-	unsigned long size, addr;
-	pte_t *pte;
-	int rc;
-	struct page *page;
-	unsigned int order = pfn_to_section_compound_order(start_pfn);
-
-	page = vmemmap_shared_tail_page(order, device_zone(node));
-	if (!page)
-		return -ENOMEM;
-
-	if (reuse_compound_section(start_pfn, pgmap))
-		return vmemmap_populate_range(start, end, node, NULL,
-					      page_to_pfn(page));
-
-	size = min(end - start, (1UL << order) * sizeof(struct page));
-	for (addr = start; addr < end; addr += size) {
-		unsigned long next, last = addr + size;
-
-		/* Populate the head page vmemmap page */
-		pte = vmemmap_populate_address(addr, node, NULL, -1);
-		if (!pte)
-			return -ENOMEM;
-
-		/*
-		 * Reuse the shared page for the rest of tail pages
-		 * See layout diagram in Documentation/mm/vmemmap_dedup.rst
-		 */
-		next = addr + PAGE_SIZE;
-		rc = vmemmap_populate_range(next, last, node, NULL,
-					    page_to_pfn(page));
-		if (rc)
-			return -ENOMEM;
-	}
-
-	return 0;
-}
-
-#endif
-
 struct page * __meminit __populate_section_memmap(unsigned long pfn,
 		unsigned long nr_pages, int nid, struct vmem_altmap *altmap,
 		struct dev_pagemap *pgmap)
@@ -531,11 +466,7 @@ struct page * __meminit __populate_section_memmap(unsigned long pfn,
 		!IS_ALIGNED(nr_pages, PAGES_PER_SUBSECTION)))
 		return NULL;
 
-	if (pgmap && section_vmemmap_optimizable(__pfn_to_section(pfn)))
-		r = vmemmap_populate_compound_pages(pfn, start, end, nid, pgmap);
-	else
-		r = vmemmap_populate(start, end, nid, altmap);
-
+	r = vmemmap_populate(start, end, nid, altmap);
 	if (r < 0)
 		return NULL;
 
-- 
2.54.0

[PATCH 4/6] mm/sparse-vmemmap: remove the unused ptpfn argument

From: Muchun Song <hidden>
Date: 2026-09-13 08:41:03

vmemmap_pte_populate() no longer uses ptpfn as an input. Drop the
argument to simplify the code.

Signed-off-by: Muchun Song <redacted>
---
 mm/sparse-vmemmap.c | 15 ++++++---------
 1 file changed, 6 insertions(+), 9 deletions(-)
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 028f844c90f5..8b4904d9c65e 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -233,7 +233,7 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
 }
 
 static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, int node,
-		struct vmem_altmap *altmap, unsigned long ptpfn)
+		struct vmem_altmap *altmap)
 {
 	pte_t *pte = pte_offset_kernel(pmd, addr);
 	unsigned long pfn = page_to_pfn((struct page *)addr);
@@ -304,8 +304,7 @@ static pgd_t * __meminit vmemmap_pgd_populate(unsigned long addr, int node)
 }
 
 static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
-					      struct vmem_altmap *altmap,
-					      unsigned long ptpfn)
+						  struct vmem_altmap *altmap)
 {
 	pgd_t *pgd;
 	p4d_t *p4d;
@@ -325,7 +324,7 @@ static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
 	pmd = vmemmap_pmd_populate(pud, addr, node);
 	if (!pmd)
 		return NULL;
-	pte = vmemmap_pte_populate(pmd, addr, node, altmap, ptpfn);
+	pte = vmemmap_pte_populate(pmd, addr, node, altmap);
 	if (!pte)
 		return NULL;
 	vmemmap_verify(pte, node, addr, addr + PAGE_SIZE);
@@ -335,15 +334,13 @@ static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
 
 static int __meminit vmemmap_populate_range(unsigned long start,
 					    unsigned long end, int node,
-					    struct vmem_altmap *altmap,
-					    unsigned long ptpfn)
+					    struct vmem_altmap *altmap)
 {
 	unsigned long addr = start;
 	pte_t *pte;
 
 	for (; addr < end; addr += PAGE_SIZE) {
-		pte = vmemmap_populate_address(addr, node, altmap,
-					       ptpfn);
+		pte = vmemmap_populate_address(addr, node, altmap);
 		if (!pte)
 			return -ENOMEM;
 	}
@@ -354,7 +351,7 @@ static int __meminit vmemmap_populate_range(unsigned long start,
 int __meminit vmemmap_populate_basepages(unsigned long start, unsigned long end,
 					 int node, struct vmem_altmap *altmap)
 {
-	return vmemmap_populate_range(start, end, node, altmap, -1);
+	return vmemmap_populate_range(start, end, node, altmap);
 }
 
 /*
-- 
2.54.0

[PATCH 5/6] powerpc/mm: make vmemmap_populate_compound_pages() static

From: Muchun Song <hidden>
Date: 2026-09-13 08:41:08

vmemmap_populate_compound_pages() is no longer used outside
radix_pgtable.c.

Make it static and drop the unused dev_pagemap argument from
its only remaining caller to simplify the code.

Signed-off-by: Muchun Song <redacted>
---
 arch/powerpc/include/asm/book3s/64/radix.h |  6 ------
 arch/powerpc/mm/book3s64/radix_pgtable.c   | 13 +++++++------
 2 files changed, 7 insertions(+), 12 deletions(-)
diff --git a/arch/powerpc/include/asm/book3s/64/radix.h b/arch/powerpc/include/asm/book3s/64/radix.h
index da954e779744..8452a2714cb1 100644
--- a/arch/powerpc/include/asm/book3s/64/radix.h
+++ b/arch/powerpc/include/asm/book3s/64/radix.h
@@ -356,11 +356,5 @@ int radix__remove_section_mapping(unsigned long start, unsigned long end);
 #define vmemmap_can_optimize vmemmap_can_optimize
 bool vmemmap_can_optimize(struct vmem_altmap *altmap, struct dev_pagemap *pgmap);
 #endif
-
-#define vmemmap_populate_compound_pages vmemmap_populate_compound_pages
-int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
-					      unsigned long start,
-					      unsigned long end, int node,
-					      struct dev_pagemap *pgmap);
 #endif /* __ASSEMBLER__ */
 #endif
diff --git a/arch/powerpc/mm/book3s64/radix_pgtable.c b/arch/powerpc/mm/book3s64/radix_pgtable.c
index cb72d9ccf747..ee121cb4fc12 100644
--- a/arch/powerpc/mm/book3s64/radix_pgtable.c
+++ b/arch/powerpc/mm/book3s64/radix_pgtable.c
@@ -1110,7 +1110,9 @@ static inline pte_t *vmemmap_pte_alloc(pmd_t *pmdp, int node,
 	return pte_offset_kernel(pmdp, address);
 }
 
-
+static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
+						     unsigned long start,
+						     unsigned long end, int node);
 
 int __meminit radix__vmemmap_populate(unsigned long start, unsigned long end, int node,
 				      struct vmem_altmap *altmap)
@@ -1125,7 +1127,7 @@ int __meminit radix__vmemmap_populate(unsigned long start, unsigned long end, in
 	unsigned long pfn = page_to_pfn((struct page *)start);
 
 	if (section_vmemmap_optimizable(__pfn_to_section(pfn)))
-		return vmemmap_populate_compound_pages(pfn, start, end, node, NULL);
+		return vmemmap_populate_compound_pages(pfn, start, end, node);
 	/*
 	 * If altmap is present, Make sure we align the start vmemmap addr
 	 * to PAGE_SIZE so that we calculate the correct start_pfn in
@@ -1221,10 +1223,9 @@ int __meminit radix__vmemmap_populate(unsigned long start, unsigned long end, in
 	return 0;
 }
 
-int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
-					      unsigned long start,
-					      unsigned long end, int node,
-					      struct dev_pagemap *pgmap)
+static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
+						     unsigned long start,
+						     unsigned long end, int node)
 {
 	/*
 	 * we want to map things as base page size mapping so that
-- 
2.54.0

[PATCH 6/6] mm/sparse-vmemmap: open-code vmemmap_populate_address()

From: Muchun Song <hidden>
Date: 2026-09-13 08:41:18

vmemmap_populate_address() no longer has any callers that need the
returned PTE.  Its only remaining user, vmemmap_populate_range(), only
checks whether population succeeded.

Open-code vmemmap_populate_address() directly in
vmemmap_populate_basepages(), remove the now-redundant range helper,
and return -ENOMEM directly on failure.

Signed-off-by: Muchun Song <redacted>
---
 mm/sparse-vmemmap.c | 54 ++++++++++++++-------------------------------
 1 file changed, 17 insertions(+), 37 deletions(-)
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 8b4904d9c65e..5e0e30c77431 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -303,8 +303,8 @@ static pgd_t * __meminit vmemmap_pgd_populate(unsigned long addr, int node)
 	return pgd;
 }
 
-static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
-						  struct vmem_altmap *altmap)
+int __meminit vmemmap_populate_basepages(unsigned long start, unsigned long end,
+					 int node, struct vmem_altmap *altmap)
 {
 	pgd_t *pgd;
 	p4d_t *p4d;
@@ -312,48 +312,28 @@ static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
 	pmd_t *pmd;
 	pte_t *pte;
 
-	pgd = vmemmap_pgd_populate(addr, node);
-	if (!pgd)
-		return NULL;
-	p4d = vmemmap_p4d_populate(pgd, addr, node);
-	if (!p4d)
-		return NULL;
-	pud = vmemmap_pud_populate(p4d, addr, node);
-	if (!pud)
-		return NULL;
-	pmd = vmemmap_pmd_populate(pud, addr, node);
-	if (!pmd)
-		return NULL;
-	pte = vmemmap_pte_populate(pmd, addr, node, altmap);
-	if (!pte)
-		return NULL;
-	vmemmap_verify(pte, node, addr, addr + PAGE_SIZE);
-
-	return pte;
-}
-
-static int __meminit vmemmap_populate_range(unsigned long start,
-					    unsigned long end, int node,
-					    struct vmem_altmap *altmap)
-{
-	unsigned long addr = start;
-	pte_t *pte;
-
-	for (; addr < end; addr += PAGE_SIZE) {
-		pte = vmemmap_populate_address(addr, node, altmap);
+	for (unsigned long addr = start; addr < end; addr += PAGE_SIZE) {
+		pgd = vmemmap_pgd_populate(addr, node);
+		if (!pgd)
+			return -ENOMEM;
+		p4d = vmemmap_p4d_populate(pgd, addr, node);
+		if (!p4d)
+			return -ENOMEM;
+		pud = vmemmap_pud_populate(p4d, addr, node);
+		if (!pud)
+			return -ENOMEM;
+		pmd = vmemmap_pmd_populate(pud, addr, node);
+		if (!pmd)
+			return -ENOMEM;
+		pte = vmemmap_pte_populate(pmd, addr, node, altmap);
 		if (!pte)
 			return -ENOMEM;
+		vmemmap_verify(pte, node, addr, addr + PAGE_SIZE);
 	}
 
 	return 0;
 }
 
-int __meminit vmemmap_populate_basepages(unsigned long start, unsigned long end,
-					 int node, struct vmem_altmap *altmap)
-{
-	return vmemmap_populate_range(start, end, node, altmap);
-}
-
 /*
  * Write protect the mirrored tail page structs for HVO. This will be
  * called from the hugetlb code when gathering and initializing the
-- 
2.54.0

Re: [PATCH 1/6] mm/sparse-vmemmap: drop VMEMMAP_POPULATE_DAX

From: Qi Zheng <qi.zheng@linux.dev>
Date: 2026-09-19 14:01:24


On 9/13/26 4:37 PM, Muchun Song wrote:
VMEMMAP_POPULATE_DAX currently distinguishes DAX vmemmap population in two
places: it keeps allocations on the normal path and takes a reference when
a backing page is supplied for reuse.

After Device DAX switched to the common per-zone shared tail page, both
conditions can be determined locally. DAX supplies ptpfn for every shared
tail mapping and requests an allocation only for compound head mappings,
whose PFNs are not optimizable. Therefore, vmemmap_optimizable_pfn() alone
selects the correct allocation path.

When ptpfn is supplied, the caller is reusing an existing backing page.
Once the slab allocator is available, take a reference for each reused
Does the availability of slab mean the buddy allocator is already being
used? Could there be a window where the buddy allocator is functional
but slab hasn't become available yet?
quoted hunk
mapping to balance the release performed by vmemmap_free(). Early mappings
are backed by memblock/reserved memory and do not need page reference
accounting.

Remove VMEMMAP_POPULATE_DAX and the flags argument from the vmemmap
population helpers.

Signed-off-by: Muchun Song <redacted>
---
  mm/sparse-vmemmap.c | 36 +++++++++++++-----------------------
  1 file changed, 13 insertions(+), 23 deletions(-)
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 96506f594924..878d29a4e862 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -32,12 +32,6 @@
  #include <asm/dma.h>
  #include <asm/tlbflush.h>
  
-/*
- * Flags for vmemmap_populate_range and friends.
- */
-/* Vmemmap population for ZONE_DEVICE compound pages */
-#define VMEMMAP_POPULATE_DAX		0x0001
-
  #include "internal.h"
  #include "mm_init.h"
  #include "sparse.h"
@@ -208,7 +202,7 @@ struct page __ref *vmemmap_shared_tail_page(unsigned int order, struct zone *zon
  }
  
  static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
-		struct vmem_altmap *altmap, unsigned long flags)
+		struct vmem_altmap *altmap)
  {
  	struct zone *zone;
  	struct page *page;
@@ -218,7 +212,7 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
  	 * Device DAX still relies on vmemmap_populate_compound_pages() for
  	 * head/first-tail allocation and tail-page reuse.
  	 */
-	if (!vmemmap_optimizable_pfn(pfn) || flags & VMEMMAP_POPULATE_DAX)
+	if (!vmemmap_optimizable_pfn(pfn))
  		return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
  
  	zone = pfn_to_zone(pfn, node);
@@ -230,8 +224,7 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
  }
  
  static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, int node,
-				       struct vmem_altmap *altmap,
-				       unsigned long ptpfn, unsigned long flags)
+		struct vmem_altmap *altmap, unsigned long ptpfn)
  {
  	pte_t *pte = pte_offset_kernel(pmd, addr);
  	unsigned long pfn = page_to_pfn((struct page *)addr);
@@ -240,7 +233,7 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
  		pte_t entry;
  
  		if (ptpfn == (unsigned long)-1) {
-			void *p = vmemmap_alloc_pte(pfn, node, altmap, flags);
+			void *p = vmemmap_alloc_pte(pfn, node, altmap);
  
  			if (!p)
  				return NULL;
@@ -255,7 +248,7 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
  			 * and through vmemmap_populate_compound_pages() when
  			 * slab is available.
  			 */
-			if (flags & VMEMMAP_POPULATE_DAX)
+			if (slab_is_available())
  				get_page(pfn_to_page(ptpfn));
  		}
  		entry = pfn_pte(ptpfn, PAGE_KERNEL);
@@ -318,8 +311,7 @@ static pgd_t * __meminit vmemmap_pgd_populate(unsigned long addr, int node)
  
  static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
  					      struct vmem_altmap *altmap,
-					      unsigned long ptpfn,
-					      unsigned long flags)
+					      unsigned long ptpfn)
  {
  	pgd_t *pgd;
  	p4d_t *p4d;
@@ -339,7 +331,7 @@ static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
  	pmd = vmemmap_pmd_populate(pud, addr, node);
  	if (!pmd)
  		return NULL;
-	pte = vmemmap_pte_populate(pmd, addr, node, altmap, ptpfn, flags);
+	pte = vmemmap_pte_populate(pmd, addr, node, altmap, ptpfn);
  	if (!pte)
  		return NULL;
  	vmemmap_verify(pte, node, addr, addr + PAGE_SIZE);
@@ -350,15 +342,14 @@ static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
  static int __meminit vmemmap_populate_range(unsigned long start,
  					    unsigned long end, int node,
  					    struct vmem_altmap *altmap,
-					    unsigned long ptpfn,
-					    unsigned long flags)
+					    unsigned long ptpfn)
  {
  	unsigned long addr = start;
  	pte_t *pte;
  
  	for (; addr < end; addr += PAGE_SIZE) {
  		pte = vmemmap_populate_address(addr, node, altmap,
-					       ptpfn, flags);
+					       ptpfn);
  		if (!pte)
  			return -ENOMEM;
  	}
@@ -369,7 +360,7 @@ static int __meminit vmemmap_populate_range(unsigned long start,
  int __meminit vmemmap_populate_basepages(unsigned long start, unsigned long end,
  					 int node, struct vmem_altmap *altmap)
  {
-	return vmemmap_populate_range(start, end, node, altmap, -1, 0);
+	return vmemmap_populate_range(start, end, node, altmap, -1);
  }
  
  /*
@@ -498,7 +489,6 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
  	unsigned long size, addr;
  	pte_t *pte;
  	int rc;
-	unsigned long flags = VMEMMAP_POPULATE_DAX;
  	struct page *page;
  	unsigned int order = pfn_to_section_compound_order(start_pfn);
  
@@ -508,14 +498,14 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
  
  	if (reuse_compound_section(start_pfn, pgmap))
  		return vmemmap_populate_range(start, end, node, NULL,
-					      page_to_pfn(page), flags);
+					      page_to_pfn(page));
  
  	size = min(end - start, (1UL << order) * sizeof(struct page));
  	for (addr = start; addr < end; addr += size) {
  		unsigned long next, last = addr + size;
  
  		/* Populate the head page vmemmap page */
-		pte = vmemmap_populate_address(addr, node, NULL, -1, flags);
+		pte = vmemmap_populate_address(addr, node, NULL, -1);
  		if (!pte)
  			return -ENOMEM;
  
@@ -525,7 +515,7 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
  		 */
  		next = addr + PAGE_SIZE;
  		rc = vmemmap_populate_range(next, last, node, NULL,
-					    page_to_pfn(page), flags);
+					    page_to_pfn(page));
  		if (rc)
  			return -ENOMEM;
  	}

Re: [PATCH 1/6] mm/sparse-vmemmap: drop VMEMMAP_POPULATE_DAX

From: Muchun Song <muchun.song@linux.dev>
Date: 2026-09-19 14:07:45

On Sep 19, 2026, at 22:01, Qi Zheng [off-list ref] wrote:
On 9/13/26 4:37 PM, Muchun Song wrote:
quoted
VMEMMAP_POPULATE_DAX currently distinguishes DAX vmemmap population in two
places: it keeps allocations on the normal path and takes a reference when
a backing page is supplied for reuse.
After Device DAX switched to the common per-zone shared tail page, both
conditions can be determined locally. DAX supplies ptpfn for every shared
tail mapping and requests an allocation only for compound head mappings,
whose PFNs are not optimizable. Therefore, vmemmap_optimizable_pfn() alone
selects the correct allocation path.
When ptpfn is supplied, the caller is reusing an existing backing page.
Once the slab allocator is available, take a reference for each reused
Does the availability of slab mean the buddy allocator is already being
used? Could there be a window where the buddy allocator is functional
but slab hasn't become available yet?
Yes, because the slab allocator is based on buddy allocator. But I want to know
what's your concern here?

Re: [PATCH 1/6] mm/sparse-vmemmap: drop VMEMMAP_POPULATE_DAX

From: Qi Zheng <qi.zheng@linux.dev>
Date: 2026-09-21 03:44:07


On 9/19/26 10:07 PM, Muchun Song wrote:
quoted
On Sep 19, 2026, at 22:01, Qi Zheng [off-list ref] wrote:
On 9/13/26 4:37 PM, Muchun Song wrote:
quoted
VMEMMAP_POPULATE_DAX currently distinguishes DAX vmemmap population in two
places: it keeps allocations on the normal path and takes a reference when
a backing page is supplied for reuse.
After Device DAX switched to the common per-zone shared tail page, both
conditions can be determined locally. DAX supplies ptpfn for every shared
tail mapping and requests an allocation only for compound head mappings,
whose PFNs are not optimizable. Therefore, vmemmap_optimizable_pfn() alone
selects the correct allocation path.
When ptpfn is supplied, the caller is reusing an existing backing page.
Once the slab allocator is available, take a reference for each reused
Does the availability of slab mean the buddy allocator is already being
used? Could there be a window where the buddy allocator is functional
but slab hasn't become available yet?
Yes, because the slab allocator is based on buddy allocator. But I want to know
what's your concern here?
The goal here is to check if the buddy allocator is ready, but the code
actually checks for slab.

I'm concerned there could be a gap between these two:

buddy is ready <-- gap --> slab is ready

Thanks,
Qi


Re: [PATCH 1/6] mm/sparse-vmemmap: drop VMEMMAP_POPULATE_DAX

From: Muchun Song <muchun.song@linux.dev>
Date: 2026-09-21 03:49:48

On Sep 21, 2026, at 11:43, Qi Zheng [off-list ref] wrote:
On 9/19/26 10:07 PM, Muchun Song wrote:
quoted
quoted
On Sep 19, 2026, at 22:01, Qi Zheng [off-list ref] wrote:
On 9/13/26 4:37 PM, Muchun Song wrote:
quoted
VMEMMAP_POPULATE_DAX currently distinguishes DAX vmemmap population in two
places: it keeps allocations on the normal path and takes a reference when
a backing page is supplied for reuse.
After Device DAX switched to the common per-zone shared tail page, both
conditions can be determined locally. DAX supplies ptpfn for every shared
tail mapping and requests an allocation only for compound head mappings,
whose PFNs are not optimizable. Therefore, vmemmap_optimizable_pfn() alone
selects the correct allocation path.
When ptpfn is supplied, the caller is reusing an existing backing page.
Once the slab allocator is available, take a reference for each reused
Does the availability of slab mean the buddy allocator is already being
used? Could there be a window where the buddy allocator is functional
but slab hasn't become available yet?
Yes, because the slab allocator is based on buddy allocator. But I want to know
what's your concern here?
The goal here is to check if the buddy allocator is ready, but the code
actually checks for slab.

I'm concerned there could be a gap between these two:

buddy is ready <-- gap --> slab is ready
There is no vmemmap population during the gap. We don't need to concern the gap.

Thanks.
Thanks,
Qi

Re: [PATCH 1/6] mm/sparse-vmemmap: drop VMEMMAP_POPULATE_DAX

From: Qi Zheng <qi.zheng@linux.dev>
Date: 2026-09-22 07:54:49


On 9/21/26 11:49 AM, Muchun Song wrote:
quoted
On Sep 21, 2026, at 11:43, Qi Zheng [off-list ref] wrote:
On 9/19/26 10:07 PM, Muchun Song wrote:
quoted
quoted
On Sep 19, 2026, at 22:01, Qi Zheng [off-list ref] wrote:
On 9/13/26 4:37 PM, Muchun Song wrote:
quoted
VMEMMAP_POPULATE_DAX currently distinguishes DAX vmemmap population in two
places: it keeps allocations on the normal path and takes a reference when
a backing page is supplied for reuse.
After Device DAX switched to the common per-zone shared tail page, both
conditions can be determined locally. DAX supplies ptpfn for every shared
tail mapping and requests an allocation only for compound head mappings,
whose PFNs are not optimizable. Therefore, vmemmap_optimizable_pfn() alone
selects the correct allocation path.
When ptpfn is supplied, the caller is reusing an existing backing page.
Once the slab allocator is available, take a reference for each reused
Does the availability of slab mean the buddy allocator is already being
used? Could there be a window where the buddy allocator is functional
but slab hasn't become available yet?
Yes, because the slab allocator is based on buddy allocator. But I want to know
what's your concern here?
The goal here is to check if the buddy allocator is ready, but the code
actually checks for slab.

I'm concerned there could be a gap between these two:

buddy is ready <-- gap --> slab is ready
There is no vmemmap population during the gap. We don't need to concern the gap.
Okay, perhaps we could add more explanation in the commit message.

With this, LGTM, so:

Acked-by: Qi Zheng <qi.zheng@linux.dev>

Thanks,
Qi
Thanks.
quoted
Thanks,
Qi

Re: [PATCH 2/6] mm/sparse-vmemmap: support device DAX in common vmemmap path

From: Qi Zheng <qi.zheng@linux.dev>
Date: 2026-09-22 08:03:04


On 9/13/26 4:37 PM, Muchun Song wrote:
quoted hunk
The common vmemmap population path cannot yet handle optimized Device DAX
mappings on its own. It uses pfn_to_zone() to find the shared tail page,
but Device DAX populates its vmemmap at runtime before the ZONE_DEVICE span
is initialized.

Teach the common path to use device_zone() for runtime optimized vmemmap
population while retaining pfn_to_zone() for early boot. This allows the
same path to support both early boot mappings and Device DAX.

The backing PFN supplied by the Device DAX-specific population path is no
longer used, allowing the redundant lookup and population code to be
removed later.

Signed-off-by: Muchun Song <redacted>
---
  mm/sparse-vmemmap.c | 44 +++++++++++++++++++-------------------------
  1 file changed, 19 insertions(+), 25 deletions(-)
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 878d29a4e862..e83821768c12 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -208,18 +208,27 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
  	struct page *page;
  	const unsigned int order = pfn_to_section_compound_order(pfn);
  
-	/*
-	 * Device DAX still relies on vmemmap_populate_compound_pages() for
-	 * head/first-tail allocation and tail-page reuse.
-	 */
  	if (!vmemmap_optimizable_pfn(pfn))
  		return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
  
-	zone = pfn_to_zone(pfn, node);
+	/*
+	 * At runtime (slab available), only ZONE_DEVICE pages trigger vmemmap
+	 * optimization, so device_zone() suffices. Note that pfn_to_zone()
+	 * cannot be used at runtime because the zone span is not set up now.
+	 */
+	zone = slab_is_available() ? device_zone(node) : pfn_to_zone(pfn, node);
  	page = vmemmap_shared_tail_page(order, zone);
  	if (!page)
  		return NULL;
  
+	/*
+	 * When a PTE entry is freed, a free_pages() call occurs. This get_page()
+	 * pairs with put_page_testzero() on the freeing path. This can only occur
+	 * when slab is available.
+	 */
+	if (slab_is_available())
Would it make sense to introduce a helper function that wraps
slab_is_available() for better readability? Also, it might be worth
adding a comment to the helper function as well.

At least for me, encountering this always causes a moment of confusion.
quoted hunk
+		get_page(page);
+
  	return page_address(page);
  }
  
@@ -231,27 +240,12 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
  
  	if (pte_none(ptep_get(pte))) {
  		pte_t entry;
+		void *p = vmemmap_alloc_pte(pfn, node, altmap);
  
-		if (ptpfn == (unsigned long)-1) {
-			void *p = vmemmap_alloc_pte(pfn, node, altmap);
-
-			if (!p)
-				return NULL;
-			ptpfn = PHYS_PFN(__pa(p));
-		} else {
-			/*
-			 * When a PTE/PMD entry is freed from the init_mm
-			 * there's a free_pages() call to this page allocated
-			 * above. Thus this get_page() is paired with the
-			 * put_page_testzero() on the freeing path.
-			 * This can only called by certain ZONE_DEVICE path,
-			 * and through vmemmap_populate_compound_pages() when
-			 * slab is available.
-			 */
-			if (slab_is_available())
-				get_page(pfn_to_page(ptpfn));
-		}
-		entry = pfn_pte(ptpfn, PAGE_KERNEL);
+		if (!p)
+			return NULL;
+
+		entry = pfn_pte(PHYS_PFN(__pa(p)), PAGE_KERNEL);
  		set_pte_at(&init_mm, addr, pte, entry);
  	} else if (WARN_ON_ONCE(vmemmap_optimizable_pfn(pfn)))
  		return NULL;

Re: [PATCH 3/6] mm/sparse-vmemmap: drop Device DAX-specific population path

From: Qi Zheng <qi.zheng@linux.dev>
Date: 2026-09-22 08:07:20


On 9/13/26 4:37 PM, Muchun Song wrote:
The common vmemmap path selects the shared page for optimized mappings
itself, so Device DAX no longer needs vmemmap_populate_compound_pages()
to find a shared tail page and pass its backing PFN through the generic
population helpers.

Remove the Device DAX-specific population path and let section memmap
population always use vmemmap_populate(). The powerpc retains an
architecture-specific compound-page implementation, so select it directly
from radix__vmemmap_populate() for optimizable sections.

Signed-off-by: Muchun Song <redacted>
---
  arch/powerpc/mm/book3s64/radix_pgtable.c |  3 +
  mm/mm_init.c                             |  2 +-
  mm/sparse-vmemmap.c                      | 71 +-----------------------
  3 files changed, 5 insertions(+), 71 deletions(-)
Acked-by: Qi Zheng <qi.zheng@linux.dev>

Thanks,
Qi

Re: [PATCH 4/6] mm/sparse-vmemmap: remove the unused ptpfn argument

From: Qi Zheng <qi.zheng@linux.dev>
Date: 2026-09-22 08:13:04


On 9/13/26 4:37 PM, Muchun Song wrote:
vmemmap_pte_populate() no longer uses ptpfn as an input. Drop the
argument to simplify the code.

Signed-off-by: Muchun Song <redacted>
---
  mm/sparse-vmemmap.c | 15 ++++++---------
  1 file changed, 6 insertions(+), 9 deletions(-)
Acked-by: Qi Zheng <qi.zheng@linux.dev>

Thanks,
Qi

Re: [PATCH 5/6] powerpc/mm: make vmemmap_populate_compound_pages() static

From: Qi Zheng <qi.zheng@linux.dev>
Date: 2026-09-22 08:14:38


On 9/13/26 4:37 PM, Muchun Song wrote:
vmemmap_populate_compound_pages() is no longer used outside
radix_pgtable.c.

Make it static and drop the unused dev_pagemap argument from
its only remaining caller to simplify the code.

Signed-off-by: Muchun Song <redacted>
---
  arch/powerpc/include/asm/book3s/64/radix.h |  6 ------
  arch/powerpc/mm/book3s64/radix_pgtable.c   | 13 +++++++------
  2 files changed, 7 insertions(+), 12 deletions(-)
Acked-by: Qi Zheng <qi.zheng@linux.dev>

Thanks,
Qi

Re: [PATCH 6/6] mm/sparse-vmemmap: open-code vmemmap_populate_address()

From: Qi Zheng <qi.zheng@linux.dev>
Date: 2026-09-22 08:18:36


On 9/13/26 4:37 PM, Muchun Song wrote:
vmemmap_populate_address() no longer has any callers that need the
returned PTE.  Its only remaining user, vmemmap_populate_range(), only
checks whether population succeeded.

Open-code vmemmap_populate_address() directly in
vmemmap_populate_basepages(), remove the now-redundant range helper,
and return -ENOMEM directly on failure.

Signed-off-by: Muchun Song <redacted>
---
  mm/sparse-vmemmap.c | 54 ++++++++++++++-------------------------------
  1 file changed, 17 insertions(+), 37 deletions(-)
Acked-by: Qi Zheng <qi.zheng@linux.dev>

Thanks,
Qi

Re: [PATCH 2/6] mm/sparse-vmemmap: support device DAX in common vmemmap path

From: Muchun Song <muchun.song@linux.dev>
Date: 2026-09-22 11:34:45


On 2026/9/22 16:02, Qi Zheng wrote:

On 9/13/26 4:37 PM, Muchun Song wrote:
quoted
The common vmemmap population path cannot yet handle optimized Device DAX
mappings on its own. It uses pfn_to_zone() to find the shared tail page,
but Device DAX populates its vmemmap at runtime before the ZONE_DEVICE span
is initialized.

Teach the common path to use device_zone() for runtime optimized vmemmap
population while retaining pfn_to_zone() for early boot. This allows the
same path to support both early boot mappings and Device DAX.

The backing PFN supplied by the Device DAX-specific population path is no
longer used, allowing the redundant lookup and population code to be
removed later.

Signed-off-by: Muchun Song <redacted>
---
  mm/sparse-vmemmap.c | 44 +++++++++++++++++++-------------------------
  1 file changed, 19 insertions(+), 25 deletions(-)
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 878d29a4e862..e83821768c12 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -208,18 +208,27 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
      struct page *page;
      const unsigned int order = pfn_to_section_compound_order(pfn);
  -    /*
-     * Device DAX still relies on vmemmap_populate_compound_pages() for
-     * head/first-tail allocation and tail-page reuse.
-     */
      if (!vmemmap_optimizable_pfn(pfn))
          return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
  -    zone = pfn_to_zone(pfn, node);
+    /*
+     * At runtime (slab available), only ZONE_DEVICE pages trigger vmemmap
+     * optimization, so device_zone() suffices. Note that pfn_to_zone()
+     * cannot be used at runtime because the zone span is not set up now.
+     */
+    zone = slab_is_available() ? device_zone(node) : pfn_to_zone(pfn, node);
      page = vmemmap_shared_tail_page(order, zone);
      if (!page)
          return NULL;
  +    /*
+     * When a PTE entry is freed, a free_pages() call occurs. This get_page()
+     * pairs with put_page_testzero() on the freeing path. This can only occur
+     * when slab is available.
+     */
+    if (slab_is_available())
Would it make sense to introduce a helper function that wraps
slab_is_available() for better readability? Also, it might be worth
adding a comment to the helper function as well.
Thanks for the suggestion. I considered introducing a helper, but the
two uses of slab_is_available() depend on different properties of the
same initialization boundary. One determines how the zone is obtained,
while the other determines whether the shared vmemmap backing page can
participate in page refcounting. Any helper name would therefore be
either too generic or specific to only one of these properties.

I would prefer to keep slab_is_available() explicit and expand the
comments at both call sites. The first comment explains why early
system RAM can use pfn_to_zone(), whereas ZONE_DEVICE population after
slab becomes available must use device_zone(). The second explains why
early shared backing pages, which are allocated from memblock, cannot
yet be refcounted. Once slab becomes available, shared backing pages
are allocated from the buddy allocator and can hold one reference for
each shared PTE mapping.

I will adopt the following revisions. Does this make sense to you?
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 5e0e30c77431..406d6f7918f7 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -212,9 +212,13 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
                return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);

        /*
-        * At runtime (slab available), only ZONE_DEVICE pages trigger vmemmap
-        * optimization, so device_zone() suffices. Note that pfn_to_zone()
-        * cannot be used at runtime because the zone span is not set up now.
+        * Before slab is available, vmemmap optimization is used for early
+        * system RAM, whose zone can be determined from the PFN.
+        *
+        * Once slab is available, only ZONE_DEVICE memory reaches this
+        * optimized population path. Its zone span has not been initialized
+        * while its vmemmap is being populated, so pfn_to_zone() cannot be
+        * used. Obtain ZONE_DEVICE directly from the node instead.
         */
        zone = slab_is_available() ? device_zone(node) : pfn_to_zone(pfn, node);
        page = vmemmap_shared_tail_page(order, zone);
@@ -222,9 +226,17 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
                return NULL;

        /*
-        * When a PTE entry is freed, a free_pages() call occurs. This get_page()
-        * pairs with put_page_testzero() on the freeing path. This can only occur
-        * when slab is available.
+        * During early vmemmap population, the shared tail vmemmap backing
+        * page is allocated from memblock before its struct page can safely
+        * participate in page refcounting. Therefore, no reference can be
+        * held for each shared PTE mapping, and the mappings must be unshared
+        * before the vmemmap is depopulated.
+        *
+        * Once slab is available, the shared backing page is allocated from
+        * the buddy allocator and can be refcounted. Hold one reference for
+        * each shared PTE mapping. The architecture vmemmap teardown drops
+        * the reference through __free_pages() when removing the mapping,
+        * preventing the backing page from being freed while it is shared.
         */
        if (slab_is_available())
                get_page(page);
At least for me, encountering this always causes a moment of confusion.
quoted
+        get_page(page);
+
      return page_address(page);
  }
  @@ -231,27 +240,12 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
        if (pte_none(ptep_get(pte))) {
          pte_t entry;
+        void *p = vmemmap_alloc_pte(pfn, node, altmap);
  -        if (ptpfn == (unsigned long)-1) {
-            void *p = vmemmap_alloc_pte(pfn, node, altmap);
-
-            if (!p)
-                return NULL;
-            ptpfn = PHYS_PFN(__pa(p));
-        } else {
-            /*
-             * When a PTE/PMD entry is freed from the init_mm
-             * there's a free_pages() call to this page allocated
-             * above. Thus this get_page() is paired with the
-             * put_page_testzero() on the freeing path.
-             * This can only called by certain ZONE_DEVICE path,
-             * and through vmemmap_populate_compound_pages() when
-             * slab is available.
-             */
-            if (slab_is_available())
-                get_page(pfn_to_page(ptpfn));
-        }
-        entry = pfn_pte(ptpfn, PAGE_KERNEL);
+        if (!p)
+            return NULL;
+
+        entry = pfn_pte(PHYS_PFN(__pa(p)), PAGE_KERNEL);
          set_pte_at(&init_mm, addr, pte, entry);
      } else if (WARN_ON_ONCE(vmemmap_optimizable_pfn(pfn)))
          return NULL;

Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help