From: Muchun Song <hidden> Date: 2026-09-11 05:03:29
This series is split out from the earlier, larger series "mm: Generalize
HVO for HugeTLB and device DAX" [1]. While the parent series generalizes
vmemmap optimization across HugeTLB and device DAX, this subset addresses
a single, self-contained step: switching device DAX to the section-based
sparse-vmemmap optimization infrastructure introduced for HugeTLB.
After the HugeTLB conversion, optimized vmemmap state is described by
the memory section and the sparse-vmemmap population path can allocate or
reuse shared tail vmemmap pages based on that metadata. Device DAX still
uses the older DAX-specific population model, including a separate tail
vmemmap page reservation and architecture-specific logic to locate or
populate reusable tail pages.
This series makes device DAX use the same section-based model. Device DAX
records the compound page order from pgmap->vmemmap_shift in section
metadata before vmemmap population, uses the common per-zone shared tail
vmemmap page, and drops the extra reserved tail page. The powerpc radix
path is updated to use the same shared tail-page helper, so the generic
and powerpc DAX paths follow the same reservation model.
The first patches prepare the shared infrastructure by introducing a
generic CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION symbol, factoring out
shared tail-page allocation, and keeping the special shared-tail struct
page initialization local to sparse-vmemmap.
The middle patches move device DAX onto that infrastructure by recording
the device DAX compound page order in memory-section metadata, using that
metadata to back generic device DAX mappings with the common per-zone
shared tail page, exposing the shared helpers so the powerpc radix path
can use the same model, and dropping the extra DAX-only tail page
reservation and the now-unused section accounting arguments.
The final patch updates the documentation for the new DAX layout.
This is intended to be the third smaller step toward the broader HVO
generalization. The wider HVO consolidation between HugeTLB and device
DAX is left for follow-up series.
[1] https://lore.kernel.org/all/20260513130542.35604-1-songmuchun@bytedance.com/
v3:
- Use EOPNOTSUPP for partial additions to sections that already use
optimized vmemmap mappings
- Move device_zone() after NODE_DATA() to fix non-NUMA builds
- Collect Acked-by tags from David Hildenbrand and Qi Zheng
- Rebase onto mm/mm-new
v2: https://lore.kernel.org/all/20260908030335.96549-1-songmuchun@bytedance.com/
- Add a missing SPARSEMEM_VMEMMAP dependency (suggested by Qi Zheng,
reported by Sashiko)
- Add an explicit ZONE_DEVICE dependency for DEV_DAX
- Add missing dependencies to the new public header
- Explain why optimized and ordinary layouts cannot share a section
(suggested by Qi Zheng)
- Explain why sharing tail vmemmap pages is safe for DEV-DAX (suggested
by Qi Zheng)
- Clarify the removal of duplicated 4K PUD calculations from the docs
(reported by Sashiko)
- Collect Acked-by tags from Qi Zheng
v1: https://lore.kernel.org/all/20260831075342.57563-1-songmuchun@bytedance.com/
Muchun Song (11):
mm/sparse-vmemmap: introduce CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION
mm/sparse-vmemmap: factor out shared vmemmap tail page allocation
mm/sparse-vmemmap: open-code init_compound_tail()
mm/sparse-vmemmap: prepare DAX vmemmap population for compound page
orders
mm/sparse-vmemmap: set compound page order for device DAX
mm/sparse-vmemmap: switch device DAX to shared tail vmemmap pages
mm/sparse-vmemmap: move vmemmap optimization helpers to a public
header
powerpc/mm: switch device DAX to shared tail vmemmap pages
mm/sparse-vmemmap: drop the extra tail page from device DAX
reservation
mm/sparse-vmemmap: drop unused section_nr_vmemmap_pages() arguments
Documentation/mm: update DAX vmemmap deduplication docs
Documentation/arch/powerpc/vmemmap_dedup.rst | 90 ++-------
Documentation/mm/vmemmap_dedup.rst | 32 +--
MAINTAINERS | 1 +
arch/powerpc/mm/book3s64/radix_pgtable.c | 124 +-----------
arch/x86/entry/vdso/vdso32/fake_32bit_build.h | 2 +-
drivers/dax/Kconfig | 2 +
fs/Kconfig | 1 +
include/linux/mm.h | 6 +-
include/linux/mmzone.h | 23 ++-
include/linux/page-flags.h | 5 +-
include/linux/vmemmap-optimization.h | 94 +++++++++
mm/Kconfig | 4 +
mm/hugetlb.c | 2 +-
mm/hugetlb_vmemmap.c | 30 +--
mm/internal.h | 9 -
mm/memory_hotplug.c | 6 +-
mm/mm_init.c | 17 +-
mm/sparse-vmemmap.c | 184 ++++++++----------
mm/sparse.c | 3 +-
mm/sparse.h | 81 +-------
20 files changed, 253 insertions(+), 463 deletions(-)
create mode 100644 include/linux/vmemmap-optimization.h
base-commit: c46036ca3ab7a7628fabe5790aefd917d0858db6
--
2.54.0
From: Muchun Song <hidden> Date: 2026-09-11 05:03:32
The section-based vmemmap optimization infrastructure is still guarded by
CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP, but it also can be used by device
DAX. Introduce CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION as a common config
for the shared infrastructure.
Select the new option from HUGETLB_PAGE_OPTIMIZE_VMEMMAP and from
DEV_DAX when the architecture opts in to DAX vmemmap optimization, and
use it to guard the generic sparse-vmemmap state and helpers.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v2:
- Fix SPARSEMEM_VMEMMAP_OPTIMIZATION being selected without SPARSEMEM_VMEMMAP
reported by Sashiko.
- Add an explicit DEV_DAX dependency on ZONE_DEVICE
- Collect Acked-by from Qi Zheng
---
arch/x86/entry/vdso/vdso32/fake_32bit_build.h | 2 +-
drivers/dax/Kconfig | 2 ++
fs/Kconfig | 1 +
include/linux/mm.h | 3 +++
include/linux/mmzone.h | 13 +++++++------
include/linux/page-flags.h | 5 ++---
mm/Kconfig | 4 ++++
mm/sparse.h | 4 ++--
8 files changed, 22 insertions(+), 12 deletions(-)
@@ -102,9 +102,9 @@**HVOwhichisonlyactiveifthesizeofstructpageisapowerof2.*/-#define MAX_FOLIO_VMEMMAP_ALIGN \-(IS_ENABLED(CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP)&&\-is_power_of_2(sizeof(structpage))?\+#define MAX_FOLIO_VMEMMAP_ALIGN \+(IS_ENABLED(CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION)&&\+is_power_of_2(sizeof(structpage))?\MAX_FOLIO_NR_PAGES*sizeof(structpage):0)/* The number of retained vmemmap pages with HVO enabled. */
@@ -461,6 +461,10 @@ config SPARSEMEM_VMEMMAPpfn_to_pageandpage_to_pfnoperations.Thisisthemostefficientoptionwhensufficientkernelresourcesareavailable.+configSPARSEMEM_VMEMMAP_OPTIMIZATION+bool+depends onSPARSEMEM_VMEMMAP+## Select this config option from the architecture Kconfig, if it is preferred# to enable the feature of HugeTLB/dev_dax vmemmap optimization.
From: Muchun Song <hidden> Date: 2026-09-11 05:03:37
HugeTLB and sparse-vmemmap each have their own helper to allocate the
shared vmemmap tail page used by vmemmap optimization.
Factor that logic into a common vmemmap_shared_tail_page() helper. It
allocates the page through vmemmap_alloc_block(), and uses cmpxchg()
to install the per-zone shared page.
Expose zone->vmemmap_tails under CONFIG_SPARSEMEM_VMEMMAP to match the
shared helper's build condition. This avoids a
!CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION stub; when optimization is
disabled, the array has no entries and the compiler folds away the unused
paths, so no storage or runtime overhead is added.
This removes duplicate allocation logic while still handling both the
early boot and runtime paths through the same helper.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v2:
- Collect Acked-by from Qi Zheng
---
include/linux/mmzone.h | 2 +-
mm/hugetlb_vmemmap.c | 28 +---------------
mm/sparse-vmemmap.c | 74 +++++++++++++++++-------------------------
mm/sparse.h | 1 +
4 files changed, 33 insertions(+), 72 deletions(-)
@@ -42,27 +42,13 @@#include"mm_init.h"#include"sparse.h"-/*-*Allocateablockofmemorytobeusedtobackthevirtualmemorymap-*ortobackthepagetablesthatareusedtocreatethemapping.-*Usesthemainallocatorsiftheyareavailable,elsebootmem.-*/--staticvoid*__ref__earlyonly_bootmem_alloc(intnode,-unsignedlongsize,-unsignedlongalign,-unsignedlonggoal)-{-returnmemmap_alloc(size,align,goal,node,false);-}--void*__meminitvmemmap_alloc_block(unsignedlongsize,intnode)+void__ref*vmemmap_alloc_block(unsignedlongsize,intnode){/* If the main allocator is up use that, fallback to bootmem. */if(slab_is_available()){gfp_tgfp_mask=GFP_KERNEL|__GFP_RETRY_MAYFAIL|__GFP_NOWARN;intorder=get_order(size);-staticboolwarned__meminitdata;+staticboolwarned;structpage*page;page=alloc_pages_node(node,gfp_mask,order);
@@ -76,8 +62,7 @@ void * __meminit vmemmap_alloc_block(unsigned long size, int node)}returnNULL;}else-return__earlyonly_bootmem_alloc(node,size,size,-__pa(MAX_DMA_ADDRESS));+returnmemmap_alloc(size,size,__pa(MAX_DMA_ADDRESS),node,false);}staticvoid*__meminitaltmap_alloc_block_buf(unsignedlongsize,
@@ -184,39 +169,40 @@ static void * __meminit vmemmap_alloc_block_zero(unsigned long size, int node)returnp;}-#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP-static__meminitstructpage*vmemmap_get_tail(unsignedintorder,structzone*zone)+structpage__ref*vmemmap_shared_tail_page(unsignedintorder,structzone*zone){-structpage*p,*tail;-unsignedintidx;-intnode=zone_to_nid(zone);+void*addr;+structpage*page;+constunsignedintidx=order-VMEMMAP_OPTIMIZATION_MIN_ORDER;-if(WARN_ON_ONCE(order<VMEMMAP_OPTIMIZATION_MIN_ORDER))-returnNULL;-if(WARN_ON_ONCE(order>MAX_FOLIO_ORDER))+if(WARN_ON_ONCE(idx>=ARRAY_SIZE(zone->vmemmap_tails)))returnNULL;-idx=order-VMEMMAP_OPTIMIZATION_MIN_ORDER;-tail=zone->vmemmap_tails[idx];-if(tail)-returntail;-p=vmemmap_alloc_block_zero(PAGE_SIZE,node);-if(!p)+page=READ_ONCE(zone->vmemmap_tails[idx]);+if(likely(page))+returnpage;++addr=vmemmap_alloc_block(PAGE_SIZE,zone_to_nid(zone));+if(!addr)returnNULL;-for(inti=0;i<PAGE_SIZE/sizeof(structpage);i++)-init_compound_tail(p+i,NULL,order,zone);-tail=virt_to_page(p);-zone->vmemmap_tails[idx]=tail;+for(inti=0;i<PAGE_SIZE/sizeof(structpage);i++){+page=(structpage*)addr+i;+mm_zero_struct_page(page);+init_compound_tail(page,NULL,order,zone);+}-returntail;-}-#else-staticinlinestructpage*vmemmap_get_tail(unsignedintorder,structzone*zone)-{-returnNULL;+page=virt_to_page(addr);+if(cmpxchg(&zone->vmemmap_tails[idx],NULL,page)!=NULL){+if(slab_is_available())+__free_page(page);+else+memblock_free(addr,PAGE_SIZE);+page=READ_ONCE(zone->vmemmap_tails[idx]);+}++returnpage;}-#endifstatic__meminitvoid*vmemmap_alloc_pte(unsignedlongpfn,intnode,structvmem_altmap*altmap)
@@ -229,7 +215,7 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,returnvmemmap_alloc_block_buf(PAGE_SIZE,node,altmap);zone=pfn_to_zone(pfn,node);-page=vmemmap_get_tail(order,zone);+page=vmemmap_shared_tail_page(order,zone);if(!page)returnNULL;
From: Muchun Song <hidden> Date: 2026-09-11 05:03:42
init_compound_tail() is only used by vmemmap_shared_tail_page(), where
the shared tail page setup intentionally passes NULL as the compound head.
Keeping this helper in mm/internal.h exposes that special case to the rest
of the MM code and can make the NULL head argument look generally valid.
Open-code the initialization at the only call site so the special-case use
stays local to sparse vmemmap optimization.
No functional change intended.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
---
v3:
- Collect Acked-by from David Hildenbrand
v2:
- Collect Acked-by from Qi Zheng
---
mm/internal.h | 9 ---------
mm/sparse-vmemmap.c | 5 ++++-
2 files changed, 4 insertions(+), 10 deletions(-)
From: Muchun Song <hidden> Date: 2026-09-11 05:03:46
Device DAX still uses vmemmap_populate_compound_pages() to populate its
compound-page vmemmap mappings. That helper allocates the head and first
tail vmemmap pages explicitly, then reuses the first tail page for the
remaining tail page mappings.
Device DAX is being moved to the section-based vmemmap optimization
infrastructure, but it cannot switch to the generic section-based
population path yet. Once a later patch records the DAX compound page
order in section metadata, DAX head and first-tail PFNs can look
optimizable to the generic helpers as well.
Add a DAX-specific population flag for this transition. It keeps DAX
head/first-tail allocations on the normal vmemmap allocation path, while
preserving the existing page reference for reused DAX tail mappings.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Update the subject and commit message to use compound page order
terminology
v2:
- Collect Acked-by from Qi Zheng
---
mm/sparse-vmemmap.c | 27 +++++++++++++++------------
1 file changed, 15 insertions(+), 12 deletions(-)
@@ -35,8 +35,8 @@/**Flagsforvmemmap_populate_rangeandfriends.*/-/* Get a ref on the head page struct page, for ZONE_DEVICE compound pages */-#define VMEMMAP_POPULATE_PAGEREF 0x0001+/* Vmemmap population for ZONE_DEVICE compound pages */+#define VMEMMAP_POPULATE_DAX 0x0001#include"internal.h"#include"mm_init.h"
@@ -208,13 +208,17 @@ struct page __ref *vmemmap_shared_tail_page(unsigned int order, struct zone *zon}static__meminitvoid*vmemmap_alloc_pte(unsignedlongpfn,intnode,-structvmem_altmap*altmap)+structvmem_altmap*altmap,unsignedlongflags){structzone*zone;structpage*page;constunsignedintorder=pfn_to_section_compound_order(pfn);-if(!vmemmap_optimizable_pfn(pfn))+/*+*DeviceDAXstillreliesonvmemmap_populate_compound_pages()for+*head/first-tailallocationandtail-pagereuse.+*/+if(!vmemmap_optimizable_pfn(pfn)||flags&VMEMMAP_POPULATE_DAX)returnvmemmap_alloc_block_buf(PAGE_SIZE,node,altmap);zone=pfn_to_zone(pfn,node);
@@ -511,6 +515,7 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,unsignedlongsize,addr;pte_t*pte;intrc;+unsignedlongflags=VMEMMAP_POPULATE_DAX;if(reuse_compound_section(start_pfn,pgmap)){pte=compound_section_tail_page(start);
@@ -522,8 +527,7 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,*withjusttailstructpages.*/returnvmemmap_populate_range(start,end,node,NULL,-pte_pfn(ptep_get(pte)),-VMEMMAP_POPULATE_PAGEREF);+pte_pfn(ptep_get(pte)),flags);}size=min(end-start,pgmap_vmemmap_nr(pgmap)*sizeof(structpage));
@@ -531,13 +535,13 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,unsignedlongnext,last=addr+size;/* Populate the head page vmemmap page */-pte=vmemmap_populate_address(addr,node,NULL,-1,0);+pte=vmemmap_populate_address(addr,node,NULL,-1,flags);if(!pte)return-ENOMEM;/* Populate the tail pages vmemmap page */next=addr+PAGE_SIZE;-pte=vmemmap_populate_address(next,node,NULL,-1,0);+pte=vmemmap_populate_address(next,node,NULL,-1,flags);if(!pte)return-ENOMEM;
@@ -547,8 +551,7 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,*/next+=PAGE_SIZE;rc=vmemmap_populate_range(next,last,node,NULL,-pte_pfn(ptep_get(pte)),-VMEMMAP_POPULATE_PAGEREF);+pte_pfn(ptep_get(pte)),flags);if(rc)return-ENOMEM;}
From: Muchun Song <hidden> Date: 2026-09-11 05:03:52
Device DAX can use vmemmap optimization only when a full section is
populated with a compound-page geometry. Record that geometry as the
compound page order in section metadata before populating the section, so
later vmemmap accounting and population decisions can use the section state
directly.
Clear the compound page order when the section becomes empty again. Also
reject partial additions to a section that already has optimized vmemmap
mappings. compound_nr_pages() determines how many struct pages to
initialize with a section as the smallest granularity. A section therefore
cannot safely mix optimized and ordinary vmemmap layouts.
Partial additions continue to use ordinary vmemmap population, so they do
not save vmemmap memory. Such additions are uncommon, and the lost saving
is negligible.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Update the subject and commit message to use compound page order
terminology
- Use EOPNOTSUPP instead of ENOTSUPP
v2:
- Explain why optimized and ordinary layouts cannot share a section
(suggested by Qi Zheng)
- Collect Acked-by from Qi Zheng
---
mm/mm_init.c | 15 +++++----------
mm/sparse-vmemmap.c | 16 ++++++++++++----
2 files changed, 17 insertions(+), 14 deletions(-)
@@ -1144,7 +1139,7 @@ void __ref memmap_init_zone_device(struct zone *zone,memcpy(&template,page,sizeof(*page));if(pfns_per_compound!=1)memmap_init_compound(page,pfn,zone_idx,nid,pgmap,-compound_nr_pages(pfn,altmap,pgmap));+compound_nr_pages(pfn,pgmap));pfn+=pfns_per_compound;/* Initialize the remaining head pages from template. */
@@ -1160,7 +1155,7 @@ void __ref memmap_init_zone_device(struct zone *zone,continue;memmap_init_compound(page,pfn,zone_idx,nid,pgmap,-compound_nr_pages(pfn,altmap,pgmap));+compound_nr_pages(pfn,pgmap));}pageblock_migratetype_init_range(start_pfn,nr_pages,MIGRATE_MOVABLE,
@@ -135,14 +135,14 @@ int __meminit section_nr_vmemmap_pages(unsigned long pfn, unsigned long nr_pagesstructvmem_altmap*altmap,structdev_pagemap*pgmap){conststructmem_section*ms=__pfn_to_section(pfn);-constintorder=pgmap?pgmap->vmemmap_shift:section_compound_order(ms);+constintorder=section_compound_order(ms);constintvmemmap_pages=pgmap?VMEMMAP_RESERVE_NR:VMEMMAP_OPTIMIZATION_PAGES;constunsignedlongpages_per_compound=1UL<<order;VM_WARN_ON_ONCE(!IS_ALIGNED(pfn|nr_pages,PAGES_PER_SUBSECTION));VM_WARN_ON_ONCE(nr_pages>PAGES_PER_SECTION);-if(!vmemmap_can_optimize(altmap,pgmap)&&!section_vmemmap_optimizable(ms))+if(!section_vmemmap_optimizable(ms))returnDIV_ROUND_UP(nr_pages*sizeof(structpage),PAGE_SIZE);if(order<PFN_SECTION_SHIFT){
From: Muchun Song <hidden> Date: 2026-09-11 05:03:56
HugeTLB vmemmap optimization now uses per-zone shared tail vmemmap pages.
Device DAX has not been switched to that mechanism yet.
Switch device DAX to vmemmap_shared_tail_page() as well. This aligns DAX
with HugeTLB by using the common per-zone shared tail vmemmap page.
The optimization is enabled only for DEV-DAX through pgmap->vmemmap_shift,
which supplies the compound page order recorded in section metadata before
vmemmap population. Unlike FS-DAX, DEV-DAX does not modify tail struct
pages, so sharing them is safe.
Since the shared tail page can now back ZONE_DEVICE vmemmap mappings,
initialize its entries with PG_reserved for device zones. Also skip
poisoning vmemmap-optimizable sections while their struct pages may be
shared.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Move device_zone() after the definition of NODE_DATA() to fix
non-NUMA builds.
- Update the commit message to describe the compound page order stored
in section metadata
- Collect Acked-by from Qi Zheng
v2:
- Explain why sharing tail vmemmap pages is safe for DEV-DAX
(suggested by Qi Zheng)
---
include/linux/mmzone.h | 10 +++++++++
mm/memory_hotplug.c | 6 ++++--
mm/sparse-vmemmap.c | 47 ++++++++++++++----------------------------
3 files changed, 29 insertions(+), 34 deletions(-)
@@ -554,8 +555,9 @@ void remove_pfn_range_from_zone(struct zone *zone,/* Select all remaining pages up to the next section boundary */cur_nr_pages=min(end_pfn-pfn,SECTION_ALIGN_UP(pfn+1)-pfn);-page_init_poison(pfn_to_page(pfn),-sizeof(structpage)*cur_nr_pages);+if(!section_vmemmap_optimizable(__pfn_to_section(pfn)))+page_init_poison(pfn_to_page(pfn),+sizeof(structpage)*cur_nr_pages);}/*
@@ -193,6 +193,8 @@ struct page __ref *vmemmap_shared_tail_page(unsigned int order, struct zone *zonset_page_node(page,zone_to_nid(zone));set_page_zone(page,zone_idx(zone));prep_compound_tail(page,NULL,order);+if(zone_is_zone_device(zone))+__SetPageReserved(page);}page=virt_to_page(addr);
@@ -490,23 +492,6 @@ static bool __meminit reuse_compound_section(unsigned long start_pfn,return!IS_ALIGNED(offset,nr_pages)&&nr_pages>PAGES_PER_SUBSECTION;}-staticpte_t*__meminitcompound_section_tail_page(unsignedlongaddr)-{-pte_t*pte;--addr-=PAGE_SIZE;--/*-*Assumingsectionsarepopulatedsequentially,theprevioussection's-*pagedatacanbereused.-*/-pte=pte_offset_kernel(pmd_off_k(addr),addr);-if(!pte)-returnNULL;--returnpte;-}-staticint__meminitvmemmap_populate_compound_pages(unsignedlongstart_pfn,unsignedlongstart,unsignedlongend,intnode,
@@ -516,21 +501,18 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,pte_t*pte;intrc;unsignedlongflags=VMEMMAP_POPULATE_DAX;+structpage*page;+unsignedintorder=pfn_to_section_compound_order(start_pfn);-if(reuse_compound_section(start_pfn,pgmap)){-pte=compound_section_tail_page(start);-if(!pte)-return-ENOMEM;+page=vmemmap_shared_tail_page(order,device_zone(node));+if(!page)+return-ENOMEM;-/*-*Reusethepagethatwaspopulatedintheprioriteration-*withjusttailstructpages.-*/+if(reuse_compound_section(start_pfn,pgmap))returnvmemmap_populate_range(start,end,node,NULL,-pte_pfn(ptep_get(pte)),flags);-}+page_to_pfn(page),flags);-size=min(end-start,pgmap_vmemmap_nr(pgmap)*sizeof(structpage));+size=min(end-start,(1UL<<order)*sizeof(structpage));for(addr=start;addr<end;addr+=size){unsignedlongnext,last=addr+size;
@@ -546,12 +528,12 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,return-ENOMEM;/*-*Reusethepreviouspagefortherestoftailpages+*Reusethesharedpagefortherestoftailpages*SeelayoutdiagraminDocumentation/mm/vmemmap_dedup.rst*/next+=PAGE_SIZE;rc=vmemmap_populate_range(next,last,node,NULL,-pte_pfn(ptep_get(pte)),flags);+page_to_pfn(page),flags);if(rc)return-ENOMEM;}
@@ -883,13 +865,14 @@ int __meminit sparse_add_section(int nid, unsigned long start_pfn,if(IS_ERR(memmap))returnPTR_ERR(memmap);+ms=__nr_to_section(section_nr);/**Poisonuninitializedstructpagesinordertocatchinvalidflags*combinations.*/-page_init_poison(memmap,sizeof(structpage)*nr_pages);+if(!section_vmemmap_optimizable(ms))+page_init_poison(memmap,sizeof(structpage)*nr_pages);-ms=__nr_to_section(section_nr);__section_mark_present(ms,section_nr);/* Align memmap to section boundary in the subsection case */
From: Muchun Song <hidden> Date: 2026-09-11 05:04:01
The vmemmap optimization helpers currently live in mm/sparse.h,
which is an internal MM header. That works for MM code, but
prevents powerpc from using the same interfaces without including a
private header.
Move the declarations and inline helpers to
include/linux/vmemmap-optimization.h. This is a preparatory change for
powerpc, which has its own vmemmap optimization implementation and needs
to use the common vmemmap optimization interfaces from architecture code.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Update the subject and commit message to describe common vmemmap
optimization helpers
- Collect Acked-by from Qi Zheng
v2:
- Fix missing header dependencies.
---
MAINTAINERS | 1 +
include/linux/vmemmap-optimization.h | 94 ++++++++++++++++++++++++++++
mm/hugetlb.c | 2 +-
mm/hugetlb_vmemmap.c | 2 +-
mm/sparse.h | 76 +---------------------
5 files changed, 98 insertions(+), 77 deletions(-)
create mode 100644 include/linux/vmemmap-optimization.h
From: Muchun Song <hidden> Date: 2026-09-11 05:04:11
The powerpc radix compound vmemmap population path still finds a reusable
tail page by walking the vmemmap page tables.
Switch it to the common vmemmap_shared_tail_page() helper instead, so it
can use the shared vmemmap page directly to simplify the code.
This removes the powerpc-specific tail-page lookup and its fallback path
and aligns the device DAX vmemmap optimization path with HugeTLB.
Signed-off-by: Muchun Song <redacted>
---
arch/powerpc/mm/book3s64/radix_pgtable.c | 80 +++---------------------
1 file changed, 9 insertions(+), 71 deletions(-)
@@ -1250,59 +1251,6 @@ static pte_t * __meminit radix__vmemmap_populate_address(unsigned long addr, intreturnpte;}-staticpte_t*__meminitvmemmap_compound_tail_page(unsignedlongaddr,-unsignedlongpfn_offset,intnode)-{-pgd_t*pgd;-p4d_t*p4d;-pud_t*pud;-pmd_t*pmd;-pte_t*pte;-unsignedlongmap_addr;--/* the second vmemmap page which we use for duplication */-map_addr=addr-pfn_offset*sizeof(structpage)+PAGE_SIZE;-pgd=pgd_offset_k(map_addr);-p4d=p4d_offset(pgd,map_addr);-pud=vmemmap_pud_alloc(p4d,node,map_addr);-if(!pud)-returnNULL;-pmd=vmemmap_pmd_alloc(pud,node,map_addr);-if(!pmd)-returnNULL;-if(pmd_leaf(*pmd))-/*-*Thesecondpageismappedasahugepageduetoanearbyrequest.-*Forceourmappingtopagesizewithoutdeduplication-*/-returnNULL;-pte=vmemmap_pte_alloc(pmd,node,map_addr);-if(!pte)-returnNULL;-/*-*Checkifthereexistamappingtotheleft-*/-if(pte_none(*pte)){-/*-*Populatetheheadpagevmemmappage.-*Itcanfallindifferentpmd,hence-*vmemmap_populate_address()-*/-pte=radix__vmemmap_populate_address(map_addr-PAGE_SIZE,node,NULL,NULL);-if(!pte)-returnNULL;-/*-*Populatethetailpagesvmemmappage-*/-pte=radix__vmemmap_pte_populate(pmd,map_addr,node,NULL,NULL);-if(!pte)-returnNULL;-vmemmap_verify(pte,node,map_addr,map_addr+PAGE_SIZE);-returnpte;-}-returnpte;-}-int__meminitvmemmap_populate_compound_pages(unsignedlongstart_pfn,unsignedlongstart,unsignedlongend,intnode,
@@ -1320,6 +1268,12 @@ int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,pud_t*pud;pmd_t*pmd;pte_t*pte;+structpage*tail_page;+unsignedintorder=pfn_to_section_compound_order(start_pfn);++tail_page=vmemmap_shared_tail_page(order,device_zone(node));+if(!tail_page)+return-ENOMEM;for(addr=start;addr<end;addr=next){
@@ -1349,10 +1303,9 @@ int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,next=addr+PAGE_SIZE;continue;}else{-unsignedlongnr_pages=pgmap_vmemmap_nr(pgmap);+unsignedlongnr_pages=1UL<<order;unsignedlongaddr_pfn=page_to_pfn((structpage*)addr);unsignedlongpfn_offset=addr_pfn-ALIGN_DOWN(addr_pfn,nr_pages);-pte_t*tail_page_pte;/**iftheaddressisalignedtohugepagesizeitisthe
@@ -1377,23 +1330,8 @@ int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,next=addr+2*PAGE_SIZE;continue;}-/*-*getthe2ndmappingdetails-*Alsocreateitifthatdoesn'texist-*/-tail_page_pte=vmemmap_compound_tail_page(addr,pfn_offset,node);-if(!tail_page_pte){--pte=radix__vmemmap_pte_populate(pmd,addr,node,NULL,NULL);-if(!pte)-return-ENOMEM;-vmemmap_verify(pte,node,addr,addr+PAGE_SIZE);--next=addr+PAGE_SIZE;-continue;-}-pte=radix__vmemmap_pte_populate(pmd,addr,node,NULL,pte_page(*tail_page_pte));+pte=radix__vmemmap_pte_populate(pmd,addr,node,NULL,tail_page);if(!pte)return-ENOMEM;vmemmap_verify(pte,node,addr,addr+PAGE_SIZE);
From: Muchun Song <hidden> Date: 2026-09-11 05:04:14
The device DAX vmemmap population still reserves one extra tail vmemmap
page after the head page.
Drop that extra reservation and let the shared tail page cover all tail
vmemmap pages after the head page, so DAX follows the same reservation
model as HugeTLB.
This reduces the reserved vmemmap pages for optimized DAX mappings to
one and removes the now-unneeded first-tail population from the generic
and powerpc paths to simplify the code as well.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Collect Acked-by from Qi Zheng
---
arch/powerpc/mm/book3s64/radix_pgtable.c | 46 ++----------------------
include/linux/mm.h | 3 +-
mm/mm_init.c | 2 +-
mm/sparse-vmemmap.c | 13 ++-----
4 files changed, 7 insertions(+), 57 deletions(-)
@@ -1218,39 +1218,6 @@ int __meminit radix__vmemmap_populate(unsigned long start, unsigned long end, inreturn0;}-staticpte_t*__meminitradix__vmemmap_populate_address(unsignedlongaddr,intnode,-structvmem_altmap*altmap,-structpage*reuse)-{-pgd_t*pgd;-p4d_t*p4d;-pud_t*pud;-pmd_t*pmd;-pte_t*pte;--pgd=pgd_offset_k(addr);-p4d=p4d_offset(pgd,addr);-pud=vmemmap_pud_alloc(p4d,node,addr);-if(!pud)-returnNULL;-pmd=vmemmap_pmd_alloc(pud,node,addr);-if(!pmd)-returnNULL;-if(pmd_leaf(*pmd))-/*-*Thesecondpageismappedasahugepageduetoanearbyrequest.-*Forceourmappingtopagesizewithoutdeduplication-*/-returnNULL;-pte=vmemmap_pte_alloc(pmd,node,addr);-if(!pte)-returnNULL;-radix__vmemmap_pte_populate(pmd,addr,node,NULL,NULL);-vmemmap_verify(pte,node,addr,addr+PAGE_SIZE);--returnpte;-}-int__meminitvmemmap_populate_compound_pages(unsignedlongstart_pfn,unsignedlongstart,unsignedlongend,intnode,
@@ -1297,7 +1264,7 @@ int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,if(!pte_none(*pte)){/**Thiscouldbebecausewealreadyhaveacompound-*pagewhoseVMEMMAP_RESERVE_NRpagesweremappedand+*pagewhoseretainedvmemmappagewasmappedand*thisrequestfallinthosepages.*/next=addr+PAGE_SIZE;
@@ -1318,16 +1285,7 @@ int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,return-ENOMEM;vmemmap_verify(pte,node,addr,addr+PAGE_SIZE);-/*-*Populatethetailpagesvmemmappage-*Itcanfallindifferentpmd,hence-*vmemmap_populate_address()-*/-pte=radix__vmemmap_populate_address(addr+PAGE_SIZE,node,NULL,NULL);-if(!pte)-return-ENOMEM;--next=addr+2*PAGE_SIZE;+next=addr+PAGE_SIZE;continue;}
@@ -136,7 +136,6 @@ int __meminit section_nr_vmemmap_pages(unsigned long pfn, unsigned long nr_pages{conststructmem_section*ms=__pfn_to_section(pfn);constintorder=section_compound_order(ms);-constintvmemmap_pages=pgmap?VMEMMAP_RESERVE_NR:VMEMMAP_OPTIMIZATION_PAGES;constunsignedlongpages_per_compound=1UL<<order;VM_WARN_ON_ONCE(!IS_ALIGNED(pfn|nr_pages,PAGES_PER_SUBSECTION));
@@ -147,13 +146,13 @@ int __meminit section_nr_vmemmap_pages(unsigned long pfn, unsigned long nr_pagesif(order<PFN_SECTION_SHIFT){VM_WARN_ON_ONCE(!IS_ALIGNED(pfn|nr_pages,pages_per_compound));-returnvmemmap_pages*nr_pages/pages_per_compound;+returnVMEMMAP_OPTIMIZATION_PAGES*nr_pages/pages_per_compound;}VM_WARN_ON_ONCE(!IS_ALIGNED(pfn|nr_pages,PAGES_PER_SECTION));if(IS_ALIGNED(pfn,pages_per_compound))-returnvmemmap_pages;+returnVMEMMAP_OPTIMIZATION_PAGES;return0;}
@@ -521,17 +520,11 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,if(!pte)return-ENOMEM;-/* Populate the tail pages vmemmap page */-next=addr+PAGE_SIZE;-pte=vmemmap_populate_address(next,node,NULL,-1,flags);-if(!pte)-return-ENOMEM;-/**Reusethesharedpagefortherestoftailpages*SeelayoutdiagraminDocumentation/mm/vmemmap_dedup.rst*/-next+=PAGE_SIZE;+next=addr+PAGE_SIZE;rc=vmemmap_populate_range(next,last,node,NULL,page_to_pfn(page),flags);if(rc)
From: Muchun Song <hidden> Date: 2026-09-11 05:04:17
section_nr_vmemmap_pages() no longer uses the altmap or pgmap
arguments, so drop them from the helper and its callers.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Collect Acked-by from Qi Zheng
---
mm/sparse-vmemmap.c | 10 ++++------
mm/sparse.c | 3 +--
mm/sparse.h | 6 ++----
3 files changed, 7 insertions(+), 12 deletions(-)
From: Muchun Song <hidden> Date: 2026-09-11 05:04:23
Device DAX now uses the common per-zone shared tail page for vmemmap
deduplication. The old documentation still described a DAX-specific
layout with a separately populated tail vmemmap page and half the HugeTLB
savings.
Update the generic and powerpc documentation to describe the shared layout.
In the powerpc document, keep the radix and 64K-specific details, drop the
duplicated 4K PUD arithmetic, and replace the repeated device-dax diagrams
with a single parameterized PMD/PUD diagram.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Collect Acked-by from Qi Zheng
v2:
- Clarify the commit message to state that the 4K PUD arithmetic is
intentionally dropped reported by Sashiko.
---
Documentation/arch/powerpc/vmemmap_dedup.rst | 90 ++++----------------
Documentation/mm/vmemmap_dedup.rst | 32 +------
2 files changed, 21 insertions(+), 101 deletions(-)
@@ -192,32 +191,7 @@ to 4 on HugeTLB pages. There's no remapping of vmemmap given that device-dax memory is not part of System RAM ranges initialized at boot. Thus the tail page deduplication-happens at a later stage when we populate the sections. HugeTLB reuses the-the head vmemmap page representing, whereas device-dax reuses the tail-vmemmap page. This results in only half of the savings compared to HugeTLB.--Deduplicated tail pages are not mapped read-only.+happens at a later stage when we populate the sections.-Here's how things look like on device-dax after the sections are populated::-- +-----------+ ---virt_to_page---> +-----------+ mapping to +-----------+-| | | 0 | -------------> | 0 |-| | +-----------+ +-----------+-| | | 1 | -------------> | 1 |-| | +-----------+ +-----------+-| | | 2 | ----------------^ ^ ^ ^ ^ ^-| | +-----------+ | | | | |-| | | 3 | ------------------+ | | | |-| | +-----------+ | | | |-| | | 4 | --------------------+ | | |-| PMD | +-----------+ | | |-| level | | 5 | ----------------------+ | |-| mapping | +-----------+ | |-| | | 6 | ------------------------+ |-| | +-----------+ |-| | | 7 | --------------------------+-| | +-----------+-| |-| |-| |- +-----------++Deduplicated tail pages are not mapped read-only. The mapping layout is the same+as HugeTLB.
From: Mike Rapoport <rppt@kernel.org> Date: 2026-09-14 07:48:22
Hi,
The section-based vmemmap optimization infrastructure is still guarded by
CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP, but it also can be used by device
DAX. Introduce CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION as a common config
CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION is a bit mouthful :)
I think that dropping _SPARSEMEM won't hurt readability.
quoted hunk
for the shared infrastructure.
Select the new option from HUGETLB_PAGE_OPTIMIZE_VMEMMAP and from
DEV_DAX when the architecture opts in to DAX vmemmap optimization, and
use it to guard the generic sparse-vmemmap state and helpers.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
@@ -102,9 +102,9 @@**HVOwhichisonlyactiveifthesizeofstructpageisapowerof2.*/-#define MAX_FOLIO_VMEMMAP_ALIGN \-(IS_ENABLED(CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP)&&\-is_power_of_2(sizeof(structpage))?\+#define MAX_FOLIO_VMEMMAP_ALIGN \+(IS_ENABLED(CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION)&&\+is_power_of_2(sizeof(structpage))?\MAX_FOLIO_NR_PAGES*sizeof(structpage):0)/* The number of retained vmemmap pages with HVO enabled. */
@@ -461,6 +461,10 @@ config SPARSEMEM_VMEMMAPpfn_to_pageandpage_to_pfnoperations.Thisisthemostefficientoptionwhensufficientkernelresourcesareavailable.+configSPARSEMEM_VMEMMAP_OPTIMIZATION+bool+depends onSPARSEMEM_VMEMMAP+## Select this config option from the architecture Kconfig, if it is preferred# to enable the feature of HugeTLB/dev_dax vmemmap optimization.
From: Mike Rapoport <rppt@kernel.org> Date: 2026-09-14 07:48:26
On Fri, 11 Sep 2026 13:02:19 +0800, Muchun Song [off-list ref] wrote:
HugeTLB and sparse-vmemmap each have their own helper to allocate the
shared vmemmap tail page used by vmemmap optimization.
Factor that logic into a common vmemmap_shared_tail_page() helper. It
allocates the page through vmemmap_alloc_block(), and uses cmpxchg()
to install the per-zone shared page.
[...]
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
--
Sincerely yours,
Mike.
From: Mike Rapoport <rppt@kernel.org> Date: 2026-09-14 07:48:30
On Fri, 11 Sep 2026 13:02:20 +0800, Muchun Song [off-list ref] wrote:
init_compound_tail() is only used by vmemmap_shared_tail_page(), where
the shared tail page setup intentionally passes NULL as the compound head.
Keeping this helper in mm/internal.h exposes that special case to the rest
of the MM code and can make the NULL head argument look generally valid.
Open-code the initialization at the only call site so the special-case use
stays local to sparse vmemmap optimization.
[...]
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
--
Sincerely yours,
Mike.
From: Mike Rapoport <rppt@kernel.org> Date: 2026-09-14 07:48:34
On Fri, 11 Sep 2026 13:02:24 +0800, Muchun Song [off-list ref] wrote:
The vmemmap optimization helpers currently live in mm/sparse.h,
which is an internal MM header. That works for MM code, but
prevents powerpc from using the same interfaces without including a
private header.
Move the declarations and inline helpers to
include/linux/vmemmap-optimization.h. This is a preparatory change for
powerpc, which has its own vmemmap optimization implementation and needs
to use the common vmemmap optimization interfaces from architecture code.
[...]
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
--
Sincerely yours,
Mike.
From: Muchun Song <muchun.song@linux.dev> Date: 2026-09-14 07:51:59
On Sep 14, 2026, at 15:48, Mike Rapoport [off-list ref] wrote:
Hi,
Hi,
quoted
The section-based vmemmap optimization infrastructure is still guarded by
CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP, but it also can be used by device
DAX. Introduce CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION as a common config
CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION is a bit mouthful :)
I think that dropping _SPARSEMEM won't hurt readability.
Yes, I'll simplify it in v4.
quoted
for the shared infrastructure.
Select the new option from HUGETLB_PAGE_OPTIMIZE_VMEMMAP and from
DEV_DAX when the architecture opts in to DAX vmemmap optimization, and
use it to guard the generic sparse-vmemmap state and helpers.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
config DEV_DAX
tristate "Device DAX: direct access mapping device"
depends on TRANSPARENT_HUGEPAGE
+ depends on ZONE_DEVICE
+ select SPARSEMEM_VMEMMAP_OPTIMIZATION if ARCH_WANT_OPTIMIZE_DAX_VMEMMAP
help
Support raw access to differentiated (persistence, bandwidth,
latency...) memory via an mmap(2) capable character
unsigned long nr_pages;
unsigned long nr_vmemmap_pages;
+ if (!IS_ENABLED(CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION))
+ return false;
+
if (!pgmap || !is_power_of_2(sizeof(struct page)))
return false;
unsigned long section_mem_map;
struct mem_section_usage *usage;
-#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP
+#ifdef CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION
/*
* Normally, sections hold regular (order-0) pages. However, for
* sections with HVO enabled, this tracks the compound page order
static __always_inline bool compound_info_has_mask(void)
{
/*
- * Limit mask usage to HugeTLB vmemmap optimization (HVO) where it
- * makes a difference.
+ * Limit mask usage to HVO where it makes a difference.
*
* The approach with mask would work in the wider set of conditions,
* but it requires validating that struct pages are naturally aligned
* for all orders up to the MAX_FOLIO_ORDER, which can be tricky.
*/
- if (!IS_ENABLED(CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP))
+ if (!IS_ENABLED(CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION))
return false;
return is_power_of_2(sizeof(struct page));
@@ -75,7 +75,7 @@ static inline bool vmemmap_optimizable_pfn(unsigned long pfn)
static inline bool vmemmap_optimizable_order(unsigned int order)
{
- if (!IS_ENABLED(CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP))
+ if (!IS_ENABLED(CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION))
return false;
if (!is_power_of_2(sizeof(struct page)))
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>