From: Muchun Song <hidden> Date: 2026-09-27 02:54:57
This series is split out from the earlier, larger series "mm: Generalize
HVO for HugeTLB and device DAX" [1]. While the parent series generalizes
vmemmap optimization across HugeTLB and device DAX, this subset addresses
a single, self-contained step: switching device DAX to the section-based
sparse-vmemmap optimization infrastructure introduced for HugeTLB.
After the HugeTLB conversion, optimized vmemmap state is described by
the memory section and the sparse-vmemmap population path can allocate or
reuse shared tail vmemmap pages based on that metadata. Device DAX still
uses the older DAX-specific population model, including a separate tail
vmemmap page reservation and architecture-specific logic to locate or
populate reusable tail pages.
This series makes device DAX use the same section-based model. Device DAX
records the compound page order from pgmap->vmemmap_shift in section
metadata before vmemmap population, uses the common per-zone shared tail
vmemmap page, and drops the extra reserved tail page. The powerpc radix
path is updated to use the same shared tail-page helper, so the generic
and powerpc DAX paths follow the same reservation model.
The first patches prepare the shared infrastructure by factoring out
shared tail-page allocation, allocating the per-zone shared tail-page
array dynamically, and introducing a generic
CONFIG_VMEMMAP_OPTIMIZATION symbol.
The middle patches move device DAX onto that infrastructure by recording
the device DAX compound page order in memory-section metadata, using that
metadata to back generic device DAX mappings with the common per-zone
shared tail page, exposing the shared helpers so the powerpc radix path
can use the same model, and dropping the extra DAX-only tail page
reservation and the now-unused section accounting arguments.
The final patch updates the documentation for the new DAX layout.
This is intended to be the third smaller step toward the broader HVO
generalization. The wider HVO consolidation between HugeTLB and device
DAX is left for follow-up series.
[1] https://lore.kernel.org/all/20260513130542.35604-1-songmuchun@bytedance.com/
v5:
- Move the shared tail-page factoring before introducing
CONFIG_VMEMMAP_OPTIMIZATION
- Add a new patch to allocate the per-zone shared tail-page array
dynamically and fix the RISC-V build failure reported by the kernel
test robot
- Select VMEMMAP_OPTIMIZATION from ZONE_DEVICE instead of DEV_DAX so
MSHV_VTL cannot set vmemmap_shift while leaving the optimization
disabled (reported by Sashiko)
- Move the vmemmap optimization macros and MAX_FOLIO_VMEMMAP_ALIGN from
mmzone.h to vmemmap-optimization.h
v4: https://lore.kernel.org/all/20260916064341.1825793-1-songmuchun@bytedance.com/
- Rename CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION to
CONFIG_VMEMMAP_OPTIMIZATION (suggested by Mike Rapoport)
- Collect Acked-by tags from Mike Rapoport
v3: https://lore.kernel.org/all/20260911050228.58884-1-songmuchun@bytedance.com/
- Use EOPNOTSUPP for partial additions to sections that already use
optimized vmemmap mappings
- Move device_zone() after NODE_DATA() to fix non-NUMA builds
- Collect Acked-by tags from David Hildenbrand and Qi Zheng
- Rebase onto mm/mm-new
v2: https://lore.kernel.org/all/20260908030335.96549-1-songmuchun@bytedance.com/
- Add a missing SPARSEMEM_VMEMMAP dependency (suggested by Qi Zheng,
reported by Sashiko)
- Add an explicit ZONE_DEVICE dependency for DEV_DAX
- Add missing dependencies to the new public header
- Explain why optimized and ordinary layouts cannot share a section
(suggested by Qi Zheng)
- Explain why sharing tail vmemmap pages is safe for DEV-DAX (suggested
by Qi Zheng)
- Clarify the removal of duplicated 4K PUD calculations from the docs
(reported by Sashiko)
- Collect Acked-by tags from Qi Zheng
v1: https://lore.kernel.org/all/20260831075342.57563-1-songmuchun@bytedance.com/
Muchun Song (12):
mm/sparse-vmemmap: factor out shared vmemmap tail page allocation
mm/sparse-vmemmap: allocate shared tail page array dynamically
mm/sparse-vmemmap: introduce CONFIG_VMEMMAP_OPTIMIZATION
mm/sparse-vmemmap: open-code init_compound_tail()
mm/sparse-vmemmap: prepare DAX vmemmap population for compound page
orders
mm/sparse-vmemmap: set compound page order for device DAX
mm/sparse-vmemmap: switch device DAX to shared tail vmemmap pages
mm/sparse-vmemmap: move vmemmap optimization helpers to a public
header
powerpc/mm: switch device DAX to shared tail vmemmap pages
mm/sparse-vmemmap: drop the extra tail page from device DAX
reservation
mm/sparse-vmemmap: drop unused section_nr_vmemmap_pages() arguments
Documentation/mm: update DAX vmemmap deduplication docs
Documentation/arch/powerpc/vmemmap_dedup.rst | 90 ++------
Documentation/mm/vmemmap_dedup.rst | 32 +--
MAINTAINERS | 1 +
arch/loongarch/include/asm/pgtable.h | 1 +
arch/powerpc/mm/book3s64/radix_pgtable.c | 124 +----------
arch/riscv/mm/init.c | 1 +
arch/x86/entry/vdso/vdso32/fake_32bit_build.h | 2 +-
fs/Kconfig | 1 +
include/linux/mm.h | 7 +-
include/linux/mmzone.h | 38 ++--
include/linux/page-flags.h | 5 +-
include/linux/vmemmap-optimization.h | 115 ++++++++++
mm/Kconfig | 5 +
mm/hugetlb.c | 2 +-
mm/hugetlb_vmemmap.c | 31 +--
mm/internal.h | 9 -
mm/memory_hotplug.c | 6 +-
mm/mm_init.c | 17 +-
mm/sparse-vmemmap.c | 209 +++++++++---------
mm/sparse.c | 3 +-
mm/sparse.h | 81 +------
21 files changed, 299 insertions(+), 481 deletions(-)
create mode 100644 include/linux/vmemmap-optimization.h
base-commit: 92068d3f6a4274d952441ba8f46221c3c03787dd
--
2.54.0
From: Muchun Song <hidden> Date: 2026-09-27 02:55:00
HugeTLB and sparse-vmemmap each have their own helper to allocate the
shared vmemmap tail page used by vmemmap optimization.
Factor that logic into a common vmemmap_shared_tail_page() helper. It
allocates the page through vmemmap_alloc_block(), initializes the tail
struct pages, and uses cmpxchg() to install the per-zone shared page.
This removes duplicate allocation logic while handling both early boot
and runtime allocation through the same helper.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
---
v5:
- Move this patch before CONFIG_VMEMMAP_OPTIMIZATION is introduced
v4:
- Update the commit message for the renamed VMEMMAP_OPTIMIZATION config
- Collect Acked-by from Mike Rapoport
v2:
- Collect Acked-by from Qi Zheng
---
mm/hugetlb_vmemmap.c | 29 +-----------------
mm/sparse-vmemmap.c | 70 ++++++++++++++++++++------------------------
mm/sparse.h | 3 ++
3 files changed, 36 insertions(+), 66 deletions(-)
@@ -42,27 +42,13 @@#include"mm_init.h"#include"sparse.h"-/*-*Allocateablockofmemorytobeusedtobackthevirtualmemorymap-*ortobackthepagetablesthatareusedtocreatethemapping.-*Usesthemainallocatorsiftheyareavailable,elsebootmem.-*/--staticvoid*__ref__earlyonly_bootmem_alloc(intnode,-unsignedlongsize,-unsignedlongalign,-unsignedlonggoal)-{-returnmemmap_alloc(size,align,goal,node,false);-}--void*__meminitvmemmap_alloc_block(unsignedlongsize,intnode)+void__ref*vmemmap_alloc_block(unsignedlongsize,intnode){/* If the main allocator is up use that, fallback to bootmem. */if(slab_is_available()){gfp_tgfp_mask=GFP_KERNEL|__GFP_RETRY_MAYFAIL|__GFP_NOWARN;intorder=get_order(size);-staticboolwarned__meminitdata;+staticboolwarned;structpage*page;page=alloc_pages_node(node,gfp_mask,order);
@@ -76,8 +62,7 @@ void * __meminit vmemmap_alloc_block(unsigned long size, int node)}returnNULL;}else-return__earlyonly_bootmem_alloc(node,size,size,-__pa(MAX_DMA_ADDRESS));+returnmemmap_alloc(size,size,__pa(MAX_DMA_ADDRESS),node,false);}staticvoid*__meminitaltmap_alloc_block_buf(unsignedlongsize,
@@ -185,34 +170,43 @@ static void * __meminit vmemmap_alloc_block_zero(unsigned long size, int node)}#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP-static__meminitstructpage*vmemmap_get_tail(unsignedintorder,structzone*zone)+structpage__ref*vmemmap_shared_tail_page(unsignedintorder,structzone*zone){-structpage*p,*tail;-unsignedintidx;-intnode=zone_to_nid(zone);+void*addr;+structpage*page;+constunsignedintidx=order-VMEMMAP_OPTIMIZATION_MIN_ORDER;-if(WARN_ON_ONCE(order<VMEMMAP_OPTIMIZATION_MIN_ORDER))-returnNULL;-if(WARN_ON_ONCE(order>MAX_FOLIO_ORDER))+if(WARN_ON_ONCE(idx>=VMEMMAP_OPTIMIZATION_NR_ORDERS))returnNULL;-idx=order-VMEMMAP_OPTIMIZATION_MIN_ORDER;-tail=zone->vmemmap_tails[idx];-if(tail)-returntail;-p=vmemmap_alloc_block_zero(PAGE_SIZE,node);-if(!p)+page=READ_ONCE(zone->vmemmap_tails[idx]);+if(likely(page))+returnpage;++addr=vmemmap_alloc_block(PAGE_SIZE,zone_to_nid(zone));+if(!addr)returnNULL;-for(inti=0;i<PAGE_SIZE/sizeof(structpage);i++)-init_compound_tail(p+i,NULL,order,zone);-tail=virt_to_page(p);-zone->vmemmap_tails[idx]=tail;+for(inti=0;i<PAGE_SIZE/sizeof(structpage);i++){+page=(structpage*)addr+i;+mm_zero_struct_page(page);+init_compound_tail(page,NULL,order,zone);+}-returntail;+page=virt_to_page(addr);+if(cmpxchg(&zone->vmemmap_tails[idx],NULL,page)!=NULL){+if(slab_is_available())+__free_page(page);+else+memblock_free(addr,PAGE_SIZE);+page=READ_ONCE(zone->vmemmap_tails[idx]);+}++returnpage;}#else-staticinlinestructpage*vmemmap_get_tail(unsignedintorder,structzone*zone)+staticinlinestructpage*vmemmap_shared_tail_page(unsignedintorder,+structzone*zone){returnNULL;}
@@ -229,7 +223,7 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,returnvmemmap_alloc_block_buf(PAGE_SIZE,node,altmap);zone=pfn_to_zone(pfn,node);-page=vmemmap_get_tail(order,zone);+page=vmemmap_shared_tail_page(order,zone);if(!page)returnNULL;
From: Muchun Song <hidden> Date: 2026-09-27 02:55:06
Commit 622026e87c40 ("mm/hugetlb: remove fake head pages") added the
per-zone vmemmap_tails array. Its size depends on MAX_FOLIO_ORDER, which
had been moved to mmzone.h in preparation for the array.
PUD_ORDER is defined by linux/pgtable.h, which cannot be included from
mmzone.h without creating an include cycle. It was therefore open-coded
as PUD_SHIFT - PAGE_SHIFT.
This removed the dependency on PUD_ORDER, but not the underlying
dependency on architecture page-table definitions. PUD_SHIFT is
generally provided by architecture page-table headers, which are not
guaranteed to have been included when mmzone.h is parsed.
The dependency remained hidden because vmemmap_tails was originally
guarded by CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP. Under that condition,
MAX_FOLIO_ORDER resolves to either MAX_PAGE_ORDER or the fixed HugeTLB
limit, rather than the PUD_SHIFT-based definition.
Device DAX, however, does not require CONFIG_HUGETLB_PAGE. When it is
converted to use section-based vmemmap optimization, MAX_FOLIO_ORDER can
resolve to PUD_SHIFT - PAGE_SHIFT while it is being used to size
vmemmap_tails. This would make struct zone depend on architecture
page-table definitions being available when mmzone.h is parsed.
Replace the embedded array with a pointer and allocate it on first use.
This moves the order-count evaluation into sparse-vmemmap.c, after the
architecture page-table definitions are available, and removes the
dependency from mmzone.h.
Removing the compile-time array also removes the original reason for
keeping MAX_FOLIO_ORDER and the vmemmap optimization sizing definitions
in mmzone.h. Follow-up cleanups can place each definition in the header
owned by its respective subsystem.
Signed-off-by: Muchun Song <redacted>
---
v5:
- Add this patch to fix the RISC-V build failure under the configuration
reported by the kernel test robot
---
include/linux/mmzone.h | 7 +------
mm/sparse-vmemmap.c | 35 +++++++++++++++++++++++++++++++----
2 files changed, 32 insertions(+), 10 deletions(-)
From: Muchun Song <hidden> Date: 2026-09-27 02:55:22
The section-based vmemmap optimization infrastructure is guarded by
CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP, but it can also be used by
ZONE_DEVICE users that set dev_pagemap::vmemmap_shift. Introduce
CONFIG_VMEMMAP_OPTIMIZATION as a common config for the shared
infrastructure.
Select the new option from HUGETLB_PAGE_OPTIMIZE_VMEMMAP and from
ZONE_DEVICE when the architecture opts in to DAX vmemmap optimization,
and use it to guard the generic sparse-vmemmap state and helpers.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
---
v5:
- Move this patch after the shared tail-page factoring.
- Select VMEMMAP_OPTIMIZATION from ZONE_DEVICE instead of DEV_DAX,
covering all users of dev_pagemap::vmemmap_shift, reported by
Sashiko.
v4:
- Rename SPARSEMEM_VMEMMAP_OPTIMIZATION to VMEMMAP_OPTIMIZATION
(suggested by Mike Rapoport)
- Collect Acked-by from Mike Rapoport
v2:
- Fix SPARSEMEM_VMEMMAP_OPTIMIZATION being selected without SPARSEMEM_VMEMMAP
reported by Sashiko.
- Add an explicit DEV_DAX dependency on ZONE_DEVICE
- Collect Acked-by from Qi Zheng
---
arch/x86/entry/vdso/vdso32/fake_32bit_build.h | 2 +-
fs/Kconfig | 1 +
include/linux/mm.h | 3 +++
include/linux/mmzone.h | 10 +++++-----
include/linux/page-flags.h | 5 ++---
mm/Kconfig | 5 +++++
mm/sparse-vmemmap.c | 2 +-
mm/sparse.h | 6 +++---
8 files changed, 21 insertions(+), 13 deletions(-)
@@ -102,9 +102,9 @@**HVOwhichisonlyactiveifthesizeofstructpageisapowerof2.*/-#define MAX_FOLIO_VMEMMAP_ALIGN \-(IS_ENABLED(CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP)&&\-is_power_of_2(sizeof(structpage))?\+#define MAX_FOLIO_VMEMMAP_ALIGN \+(IS_ENABLED(CONFIG_VMEMMAP_OPTIMIZATION)&&\+is_power_of_2(sizeof(structpage))?\MAX_FOLIO_NR_PAGES*sizeof(structpage):0)/* The number of retained vmemmap pages with HVO enabled. */
@@ -1150,7 +1150,7 @@ struct zone {/* Zone statistics */atomic_long_tvm_stat[NR_VM_ZONE_STAT_ITEMS];atomic_long_tvm_numa_event[NR_VM_NUMA_EVENT_ITEMS];-#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP+#ifdef CONFIG_VMEMMAP_OPTIMIZATIONstructpage**vmemmap_tails;#endif}____cacheline_internodealigned_in_smp;
@@ -461,6 +461,10 @@ config SPARSEMEM_VMEMMAPpfn_to_pageandpage_to_pfnoperations.Thisisthemostefficientoptionwhensufficientkernelresourcesareavailable.+configVMEMMAP_OPTIMIZATION+bool+depends onSPARSEMEM_VMEMMAP+## Select this config option from the architecture Kconfig, if it is preferred# to enable the feature of HugeTLB/dev_dax vmemmap optimization.
From: Muchun Song <hidden> Date: 2026-09-27 02:55:30
init_compound_tail() is only used by vmemmap_shared_tail_page(), where
the shared tail page setup intentionally passes NULL as the compound head.
Keeping this helper in mm/internal.h exposes that special case to the rest
of the MM code and can make the NULL head argument look generally valid.
Open-code the initialization at the only call site so the special-case use
stays local to sparse vmemmap optimization.
No functional change intended.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
---
v4:
- Collect Acked-by from Mike Rapoport
v3:
- Collect Acked-by from David Hildenbrand
v2:
- Collect Acked-by from Qi Zheng
---
mm/internal.h | 9 ---------
mm/sparse-vmemmap.c | 5 ++++-
2 files changed, 4 insertions(+), 10 deletions(-)
From: Muchun Song <hidden> Date: 2026-09-27 02:55:37
Device DAX still uses vmemmap_populate_compound_pages() to populate its
compound-page vmemmap mappings. That helper allocates the head and first
tail vmemmap pages explicitly, then reuses the first tail page for the
remaining tail page mappings.
Device DAX is being moved to the section-based vmemmap optimization
infrastructure, but it cannot switch to the generic section-based
population path yet. Once a later patch records the DAX compound page
order in section metadata, DAX head and first-tail PFNs can look
optimizable to the generic helpers as well.
Add a DAX-specific population flag for this transition. It keeps DAX
head/first-tail allocations on the normal vmemmap allocation path, while
preserving the existing page reference for reused DAX tail mappings.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Update the subject and commit message to use compound page order
terminology
v2:
- Collect Acked-by from Qi Zheng
---
mm/sparse-vmemmap.c | 27 +++++++++++++++------------
1 file changed, 15 insertions(+), 12 deletions(-)
@@ -35,8 +35,8 @@/**Flagsforvmemmap_populate_rangeandfriends.*/-/* Get a ref on the head page struct page, for ZONE_DEVICE compound pages */-#define VMEMMAP_POPULATE_PAGEREF 0x0001+/* Vmemmap population for ZONE_DEVICE compound pages */+#define VMEMMAP_POPULATE_DAX 0x0001#include"internal.h"#include"mm_init.h"
@@ -546,6 +550,7 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,unsignedlongsize,addr;pte_t*pte;intrc;+unsignedlongflags=VMEMMAP_POPULATE_DAX;if(reuse_compound_section(start_pfn,pgmap)){pte=compound_section_tail_page(start);
@@ -557,8 +562,7 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,*withjusttailstructpages.*/returnvmemmap_populate_range(start,end,node,NULL,-pte_pfn(ptep_get(pte)),-VMEMMAP_POPULATE_PAGEREF);+pte_pfn(ptep_get(pte)),flags);}size=min(end-start,pgmap_vmemmap_nr(pgmap)*sizeof(structpage));
@@ -566,13 +570,13 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,unsignedlongnext,last=addr+size;/* Populate the head page vmemmap page */-pte=vmemmap_populate_address(addr,node,NULL,-1,0);+pte=vmemmap_populate_address(addr,node,NULL,-1,flags);if(!pte)return-ENOMEM;/* Populate the tail pages vmemmap page */next=addr+PAGE_SIZE;-pte=vmemmap_populate_address(next,node,NULL,-1,0);+pte=vmemmap_populate_address(next,node,NULL,-1,flags);if(!pte)return-ENOMEM;
@@ -582,8 +586,7 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,*/next+=PAGE_SIZE;rc=vmemmap_populate_range(next,last,node,NULL,-pte_pfn(ptep_get(pte)),-VMEMMAP_POPULATE_PAGEREF);+pte_pfn(ptep_get(pte)),flags);if(rc)return-ENOMEM;}
From: Muchun Song <hidden> Date: 2026-09-27 02:55:40
The device DAX vmemmap population still reserves one extra tail vmemmap
page after the head page.
Drop that extra reservation and let the shared tail page cover all tail
vmemmap pages after the head page, so DAX follows the same reservation
model as HugeTLB.
This reduces the reserved vmemmap pages for optimized DAX mappings to
one and removes the now-unneeded first-tail population from the generic
and powerpc paths to simplify the code as well.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Collect Acked-by from Qi Zheng
---
arch/powerpc/mm/book3s64/radix_pgtable.c | 46 ++----------------------
include/linux/mm.h | 4 +--
mm/mm_init.c | 2 +-
mm/sparse-vmemmap.c | 13 ++-----
4 files changed, 8 insertions(+), 57 deletions(-)
@@ -1218,39 +1218,6 @@ int __meminit radix__vmemmap_populate(unsigned long start, unsigned long end, inreturn0;}-staticpte_t*__meminitradix__vmemmap_populate_address(unsignedlongaddr,intnode,-structvmem_altmap*altmap,-structpage*reuse)-{-pgd_t*pgd;-p4d_t*p4d;-pud_t*pud;-pmd_t*pmd;-pte_t*pte;--pgd=pgd_offset_k(addr);-p4d=p4d_offset(pgd,addr);-pud=vmemmap_pud_alloc(p4d,node,addr);-if(!pud)-returnNULL;-pmd=vmemmap_pmd_alloc(pud,node,addr);-if(!pmd)-returnNULL;-if(pmd_leaf(*pmd))-/*-*Thesecondpageismappedasahugepageduetoanearbyrequest.-*Forceourmappingtopagesizewithoutdeduplication-*/-returnNULL;-pte=vmemmap_pte_alloc(pmd,node,addr);-if(!pte)-returnNULL;-radix__vmemmap_pte_populate(pmd,addr,node,NULL,NULL);-vmemmap_verify(pte,node,addr,addr+PAGE_SIZE);--returnpte;-}-int__meminitvmemmap_populate_compound_pages(unsignedlongstart_pfn,unsignedlongstart,unsignedlongend,intnode,
@@ -1297,7 +1264,7 @@ int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,if(!pte_none(*pte)){/**Thiscouldbebecausewealreadyhaveacompound-*pagewhoseVMEMMAP_RESERVE_NRpagesweremappedand+*pagewhoseretainedvmemmappagewasmappedand*thisrequestfallinthosepages.*/next=addr+PAGE_SIZE;
@@ -1318,16 +1285,7 @@ int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,return-ENOMEM;vmemmap_verify(pte,node,addr,addr+PAGE_SIZE);-/*-*Populatethetailpagesvmemmappage-*Itcanfallindifferentpmd,hence-*vmemmap_populate_address()-*/-pte=radix__vmemmap_populate_address(addr+PAGE_SIZE,node,NULL,NULL);-if(!pte)-return-ENOMEM;--next=addr+2*PAGE_SIZE;+next=addr+PAGE_SIZE;continue;}
@@ -136,7 +136,6 @@ int __meminit section_nr_vmemmap_pages(unsigned long pfn, unsigned long nr_pages{conststructmem_section*ms=__pfn_to_section(pfn);constintorder=section_compound_order(ms);-constintvmemmap_pages=pgmap?VMEMMAP_RESERVE_NR:VMEMMAP_OPTIMIZATION_PAGES;constunsignedlongpages_per_compound=1UL<<order;VM_WARN_ON_ONCE(!IS_ALIGNED(pfn|nr_pages,PAGES_PER_SUBSECTION));
@@ -147,13 +146,13 @@ int __meminit section_nr_vmemmap_pages(unsigned long pfn, unsigned long nr_pagesif(order<PFN_SECTION_SHIFT){VM_WARN_ON_ONCE(!IS_ALIGNED(pfn|nr_pages,pages_per_compound));-returnvmemmap_pages*nr_pages/pages_per_compound;+returnVMEMMAP_OPTIMIZATION_PAGES*nr_pages/pages_per_compound;}VM_WARN_ON_ONCE(!IS_ALIGNED(pfn|nr_pages,PAGES_PER_SECTION));if(IS_ALIGNED(pfn,pages_per_compound))-returnvmemmap_pages;+returnVMEMMAP_OPTIMIZATION_PAGES;return0;}
@@ -550,17 +549,11 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,if(!pte)return-ENOMEM;-/* Populate the tail pages vmemmap page */-next=addr+PAGE_SIZE;-pte=vmemmap_populate_address(next,node,NULL,-1,flags);-if(!pte)-return-ENOMEM;-/**Reusethesharedpagefortherestoftailpages*SeelayoutdiagraminDocumentation/mm/vmemmap_dedup.rst*/-next+=PAGE_SIZE;+next=addr+PAGE_SIZE;rc=vmemmap_populate_range(next,last,node,NULL,page_to_pfn(page),flags);if(rc)
From: Muchun Song <hidden> Date: 2026-09-27 02:55:44
Device DAX can use vmemmap optimization only when a full section is
populated with a compound-page geometry. Record that geometry as the
compound page order in section metadata before populating the section, so
later vmemmap accounting and population decisions can use the section state
directly.
Clear the compound page order when the section becomes empty again. Also
reject partial additions to a section that already has optimized vmemmap
mappings. compound_nr_pages() determines how many struct pages to
initialize with a section as the smallest granularity. A section therefore
cannot safely mix optimized and ordinary vmemmap layouts.
Partial additions continue to use ordinary vmemmap population, so they do
not save vmemmap memory. Such additions are uncommon, and the lost saving
is negligible.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Update the subject and commit message to use compound page order
terminology
- Use EOPNOTSUPP instead of ENOTSUPP
v2:
- Explain why optimized and ordinary layouts cannot share a section
(suggested by Qi Zheng)
- Collect Acked-by from Qi Zheng
---
mm/mm_init.c | 15 +++++----------
mm/sparse-vmemmap.c | 16 ++++++++++++----
2 files changed, 17 insertions(+), 14 deletions(-)
@@ -1144,7 +1139,7 @@ void __ref memmap_init_zone_device(struct zone *zone,memcpy(&template,page,sizeof(*page));if(pfns_per_compound!=1)memmap_init_compound(page,pfn,zone_idx,nid,pgmap,-compound_nr_pages(pfn,altmap,pgmap));+compound_nr_pages(pfn,pgmap));pfn+=pfns_per_compound;/* Initialize the remaining head pages from template. */
@@ -1160,7 +1155,7 @@ void __ref memmap_init_zone_device(struct zone *zone,continue;memmap_init_compound(page,pfn,zone_idx,nid,pgmap,-compound_nr_pages(pfn,altmap,pgmap));+compound_nr_pages(pfn,pgmap));}pageblock_migratetype_init_range(start_pfn,nr_pages,MIGRATE_MOVABLE,
@@ -135,14 +135,14 @@ int __meminit section_nr_vmemmap_pages(unsigned long pfn, unsigned long nr_pagesstructvmem_altmap*altmap,structdev_pagemap*pgmap){conststructmem_section*ms=__pfn_to_section(pfn);-constintorder=pgmap?pgmap->vmemmap_shift:section_compound_order(ms);+constintorder=section_compound_order(ms);constintvmemmap_pages=pgmap?VMEMMAP_RESERVE_NR:VMEMMAP_OPTIMIZATION_PAGES;constunsignedlongpages_per_compound=1UL<<order;VM_WARN_ON_ONCE(!IS_ALIGNED(pfn|nr_pages,PAGES_PER_SUBSECTION));VM_WARN_ON_ONCE(nr_pages>PAGES_PER_SECTION);-if(!vmemmap_can_optimize(altmap,pgmap)&&!section_vmemmap_optimizable(ms))+if(!section_vmemmap_optimizable(ms))returnDIV_ROUND_UP(nr_pages*sizeof(structpage),PAGE_SIZE);if(order<PFN_SECTION_SHIFT){
From: Muchun Song <hidden> Date: 2026-09-27 02:55:46
section_nr_vmemmap_pages() no longer uses the altmap or pgmap
arguments, so drop them from the helper and its callers.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Collect Acked-by from Qi Zheng
---
mm/sparse-vmemmap.c | 10 ++++------
mm/sparse.c | 3 +--
mm/sparse.h | 6 ++----
3 files changed, 7 insertions(+), 12 deletions(-)
From: Muchun Song <hidden> Date: 2026-09-27 02:55:51
HugeTLB vmemmap optimization now uses per-zone shared tail vmemmap pages.
Device DAX has not been switched to that mechanism yet.
Switch device DAX to vmemmap_shared_tail_page() as well. This aligns DAX
with HugeTLB by using the common per-zone shared tail vmemmap page.
The optimization is enabled only for DEV-DAX through pgmap->vmemmap_shift,
which supplies the compound page order recorded in section metadata before
vmemmap population. Unlike FS-DAX, DEV-DAX does not modify tail struct
pages, so sharing them is safe.
Since the shared tail page can now back ZONE_DEVICE vmemmap mappings,
initialize its entries with PG_reserved for device zones. Also skip
poisoning vmemmap-optimizable sections while their struct pages may be
shared.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Move device_zone() after the definition of NODE_DATA() to fix
non-NUMA builds.
- Update the commit message to describe the compound page order stored
in section metadata
- Collect Acked-by from Qi Zheng
v2:
- Explain why sharing tail vmemmap pages is safe for DEV-DAX
(suggested by Qi Zheng)
---
include/linux/mmzone.h | 10 +++++++++
mm/memory_hotplug.c | 6 ++++--
mm/sparse-vmemmap.c | 47 ++++++++++++++----------------------------
3 files changed, 29 insertions(+), 34 deletions(-)
@@ -554,8 +555,9 @@ void remove_pfn_range_from_zone(struct zone *zone,/* Select all remaining pages up to the next section boundary */cur_nr_pages=min(end_pfn-pfn,SECTION_ALIGN_UP(pfn+1)-pfn);-page_init_poison(pfn_to_page(pfn),-sizeof(structpage)*cur_nr_pages);+if(!section_vmemmap_optimizable(__pfn_to_section(pfn)))+page_init_poison(pfn_to_page(pfn),+sizeof(structpage)*cur_nr_pages);}/*
@@ -221,6 +221,8 @@ struct page __ref *vmemmap_shared_tail_page(unsigned int order, struct zone *zonset_page_node(page,zone_to_nid(zone));set_page_zone(page,zone_idx(zone));prep_compound_tail(page,NULL,order);+if(zone_is_zone_device(zone))+__SetPageReserved(page);}page=virt_to_page(addr);
@@ -525,23 +527,6 @@ static bool __meminit reuse_compound_section(unsigned long start_pfn,return!IS_ALIGNED(offset,nr_pages)&&nr_pages>PAGES_PER_SUBSECTION;}-staticpte_t*__meminitcompound_section_tail_page(unsignedlongaddr)-{-pte_t*pte;--addr-=PAGE_SIZE;--/*-*Assumingsectionsarepopulatedsequentially,theprevioussection's-*pagedatacanbereused.-*/-pte=pte_offset_kernel(pmd_off_k(addr),addr);-if(!pte)-returnNULL;--returnpte;-}-staticint__meminitvmemmap_populate_compound_pages(unsignedlongstart_pfn,unsignedlongstart,unsignedlongend,intnode,
@@ -551,21 +536,18 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,pte_t*pte;intrc;unsignedlongflags=VMEMMAP_POPULATE_DAX;+structpage*page;+unsignedintorder=pfn_to_section_compound_order(start_pfn);-if(reuse_compound_section(start_pfn,pgmap)){-pte=compound_section_tail_page(start);-if(!pte)-return-ENOMEM;+page=vmemmap_shared_tail_page(order,device_zone(node));+if(!page)+return-ENOMEM;-/*-*Reusethepagethatwaspopulatedintheprioriteration-*withjusttailstructpages.-*/+if(reuse_compound_section(start_pfn,pgmap))returnvmemmap_populate_range(start,end,node,NULL,-pte_pfn(ptep_get(pte)),flags);-}+page_to_pfn(page),flags);-size=min(end-start,pgmap_vmemmap_nr(pgmap)*sizeof(structpage));+size=min(end-start,(1UL<<order)*sizeof(structpage));for(addr=start;addr<end;addr+=size){unsignedlongnext,last=addr+size;
@@ -581,12 +563,12 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,return-ENOMEM;/*-*Reusethepreviouspagefortherestoftailpages+*Reusethesharedpagefortherestoftailpages*SeelayoutdiagraminDocumentation/mm/vmemmap_dedup.rst*/next+=PAGE_SIZE;rc=vmemmap_populate_range(next,last,node,NULL,-pte_pfn(ptep_get(pte)),flags);+page_to_pfn(page),flags);if(rc)return-ENOMEM;}
@@ -918,13 +900,14 @@ int __meminit sparse_add_section(int nid, unsigned long start_pfn,if(IS_ERR(memmap))returnPTR_ERR(memmap);+ms=__nr_to_section(section_nr);/**Poisonuninitializedstructpagesinordertocatchinvalidflags*combinations.*/-page_init_poison(memmap,sizeof(structpage)*nr_pages);+if(!section_vmemmap_optimizable(ms))+page_init_poison(memmap,sizeof(structpage)*nr_pages);-ms=__nr_to_section(section_nr);__section_mark_present(ms,section_nr);/* Align memmap to section boundary in the subsection case */
From: Muchun Song <hidden> Date: 2026-09-27 02:55:52
Device DAX now uses the common per-zone shared tail page for vmemmap
deduplication. The old documentation still described a DAX-specific
layout with a separately populated tail vmemmap page and half the HugeTLB
savings.
Update the generic and powerpc documentation to describe the shared layout.
In the powerpc document, keep the radix and 64K-specific details, drop the
duplicated 4K PUD arithmetic, and replace the repeated device-dax diagrams
with a single parameterized PMD/PUD diagram.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Collect Acked-by from Qi Zheng
v2:
- Clarify the commit message to state that the 4K PUD arithmetic is
intentionally dropped reported by Sashiko.
---
Documentation/arch/powerpc/vmemmap_dedup.rst | 90 ++++----------------
Documentation/mm/vmemmap_dedup.rst | 32 +------
2 files changed, 21 insertions(+), 101 deletions(-)
@@ -192,32 +191,7 @@ to 4 on HugeTLB pages. There's no remapping of vmemmap given that device-dax memory is not part of System RAM ranges initialized at boot. Thus the tail page deduplication-happens at a later stage when we populate the sections. HugeTLB reuses the-the head vmemmap page representing, whereas device-dax reuses the tail-vmemmap page. This results in only half of the savings compared to HugeTLB.--Deduplicated tail pages are not mapped read-only.+happens at a later stage when we populate the sections.-Here's how things look like on device-dax after the sections are populated::-- +-----------+ ---virt_to_page---> +-----------+ mapping to +-----------+-| | | 0 | -------------> | 0 |-| | +-----------+ +-----------+-| | | 1 | -------------> | 1 |-| | +-----------+ +-----------+-| | | 2 | ----------------^ ^ ^ ^ ^ ^-| | +-----------+ | | | | |-| | | 3 | ------------------+ | | | |-| | +-----------+ | | | |-| | | 4 | --------------------+ | | |-| PMD | +-----------+ | | |-| level | | 5 | ----------------------+ | |-| mapping | +-----------+ | |-| | | 6 | ------------------------+ |-| | +-----------+ |-| | | 7 | --------------------------+-| | +-----------+-| |-| |-| |- +-----------++Deduplicated tail pages are not mapped read-only. The mapping layout is the same+as HugeTLB.
From: Muchun Song <hidden> Date: 2026-09-27 02:55:59
The vmemmap optimization helpers currently live in mm/sparse.h,
which is an internal MM header. That works for MM code, but
prevents powerpc from using the same interfaces without including a
private header.
Move the declarations and inline helpers to vmemmap-optimization.h.
This is a preparatory change for powerpc, which has its own vmemmap
optimization implementation and needs to use the common vmemmap
optimization interfaces from architecture code.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
---
v5:
- Move the VMEMMAP_OPTIMIZATION_* macros and MAX_FOLIO_VMEMMAP_ALIGN to
vmemmap-optimization.h
v4:
- Use the renamed VMEMMAP_OPTIMIZATION config in the public header
- Collect Acked-by from Mike Rapoport
v3:
- Update the subject and commit message to describe common vmemmap
optimization helpers
- Collect Acked-by from Qi Zheng
v2:
- Fix missing header dependencies.
---
MAINTAINERS | 1 +
arch/loongarch/include/asm/pgtable.h | 1 +
arch/riscv/mm/init.c | 1 +
include/linux/mmzone.h | 17 -----
include/linux/vmemmap-optimization.h | 109 +++++++++++++++++++++++++++
mm/hugetlb.c | 2 +-
mm/hugetlb_vmemmap.c | 2 +-
mm/sparse.h | 78 +------------------
8 files changed, 115 insertions(+), 96 deletions(-)
create mode 100644 include/linux/vmemmap-optimization.h
From: Muchun Song <hidden> Date: 2026-09-27 02:56:06
The powerpc radix compound vmemmap population path still finds a reusable
tail page by walking the vmemmap page tables.
Switch it to the common vmemmap_shared_tail_page() helper instead, so it
can use the shared vmemmap page directly to simplify the code.
This removes the powerpc-specific tail-page lookup and its fallback path
and aligns the device DAX vmemmap optimization path with HugeTLB.
Signed-off-by: Muchun Song <redacted>
---
arch/powerpc/mm/book3s64/radix_pgtable.c | 80 +++---------------------
include/linux/vmemmap-optimization.h | 6 ++
mm/sparse-vmemmap.c | 6 --
3 files changed, 15 insertions(+), 77 deletions(-)
@@ -1250,59 +1251,6 @@ static pte_t * __meminit radix__vmemmap_populate_address(unsigned long addr, intreturnpte;}-staticpte_t*__meminitvmemmap_compound_tail_page(unsignedlongaddr,-unsignedlongpfn_offset,intnode)-{-pgd_t*pgd;-p4d_t*p4d;-pud_t*pud;-pmd_t*pmd;-pte_t*pte;-unsignedlongmap_addr;--/* the second vmemmap page which we use for duplication */-map_addr=addr-pfn_offset*sizeof(structpage)+PAGE_SIZE;-pgd=pgd_offset_k(map_addr);-p4d=p4d_offset(pgd,map_addr);-pud=vmemmap_pud_alloc(p4d,node,map_addr);-if(!pud)-returnNULL;-pmd=vmemmap_pmd_alloc(pud,node,map_addr);-if(!pmd)-returnNULL;-if(pmd_leaf(*pmd))-/*-*Thesecondpageismappedasahugepageduetoanearbyrequest.-*Forceourmappingtopagesizewithoutdeduplication-*/-returnNULL;-pte=vmemmap_pte_alloc(pmd,node,map_addr);-if(!pte)-returnNULL;-/*-*Checkifthereexistamappingtotheleft-*/-if(pte_none(*pte)){-/*-*Populatetheheadpagevmemmappage.-*Itcanfallindifferentpmd,hence-*vmemmap_populate_address()-*/-pte=radix__vmemmap_populate_address(map_addr-PAGE_SIZE,node,NULL,NULL);-if(!pte)-returnNULL;-/*-*Populatethetailpagesvmemmappage-*/-pte=radix__vmemmap_pte_populate(pmd,map_addr,node,NULL,NULL);-if(!pte)-returnNULL;-vmemmap_verify(pte,node,map_addr,map_addr+PAGE_SIZE);-returnpte;-}-returnpte;-}-int__meminitvmemmap_populate_compound_pages(unsignedlongstart_pfn,unsignedlongstart,unsignedlongend,intnode,
@@ -1320,6 +1268,12 @@ int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,pud_t*pud;pmd_t*pmd;pte_t*pte;+structpage*tail_page;+unsignedintorder=pfn_to_section_compound_order(start_pfn);++tail_page=vmemmap_shared_tail_page(order,device_zone(node));+if(!tail_page)+return-ENOMEM;for(addr=start;addr<end;addr=next){
@@ -1349,10 +1303,9 @@ int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,next=addr+PAGE_SIZE;continue;}else{-unsignedlongnr_pages=pgmap_vmemmap_nr(pgmap);+unsignedlongnr_pages=1UL<<order;unsignedlongaddr_pfn=page_to_pfn((structpage*)addr);unsignedlongpfn_offset=addr_pfn-ALIGN_DOWN(addr_pfn,nr_pages);-pte_t*tail_page_pte;/**iftheaddressisalignedtohugepagesizeitisthe
@@ -1377,23 +1330,8 @@ int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,next=addr+2*PAGE_SIZE;continue;}-/*-*getthe2ndmappingdetails-*Alsocreateitifthatdoesn'texist-*/-tail_page_pte=vmemmap_compound_tail_page(addr,pfn_offset,node);-if(!tail_page_pte){--pte=radix__vmemmap_pte_populate(pmd,addr,node,NULL,NULL);-if(!pte)-return-ENOMEM;-vmemmap_verify(pte,node,addr,addr+PAGE_SIZE);--next=addr+PAGE_SIZE;-continue;-}-pte=radix__vmemmap_pte_populate(pmd,addr,node,NULL,pte_page(*tail_page_pte));+pte=radix__vmemmap_pte_populate(pmd,addr,node,NULL,tail_page);if(!pte)return-ENOMEM;vmemmap_verify(pte,node,addr,addr+PAGE_SIZE);
From: Andrew Morton <akpm@linux-foundation.org> Date: 2026-09-27 05:51:06
On Sun, 27 Sep 2026 10:54:29 +0800 Muchun Song [off-list ref] wrote:
After the HugeTLB conversion, optimized vmemmap state is described by
the memory section and the sparse-vmemmap population path can allocate or
reuse shared tail vmemmap pages based on that metadata. Device DAX still
uses the older DAX-specific population model, including a separate tail
vmemmap page reservation and architecture-specific logic to locate or
populate reusable tail pages.
This series makes device DAX use the same section-based model. Device DAX
records the compound page order from pgmap->vmemmap_shift in section
metadata before vmemmap population, uses the common per-zone shared tail
vmemmap page, and drops the extra reserved tail page. The powerpc radix
path is updated to use the same shared tail-page helper, so the generic
and powerpc DAX paths follow the same reservation model.
v5:
- Move the shared tail-page factoring before introducing
CONFIG_VMEMMAP_OPTIMIZATION
- Add a new patch to allocate the per-zone shared tail-page array
dynamically and fix the RISC-V build failure reported by the kernel
test robot
- Select VMEMMAP_OPTIMIZATION from ZONE_DEVICE instead of DEV_DAX so
MSHV_VTL cannot set vmemmap_shift while leaving the optimization
disabled (reported by Sashiko)
- Move the vmemmap optimization macros and MAX_FOLIO_VMEMMAP_ALIGN from
mmzone.h to vmemmap-optimization.h
From: Muchun Song <muchun.song@linux.dev> Date: 2026-09-27 10:51:37
On Sep 27, 2026, at 13:51, Andrew Morton [off-list ref] wrote:
On Sun, 27 Sep 2026 10:54:29 +0800 Muchun Song [off-list ref] wrote:
quoted
After the HugeTLB conversion, optimized vmemmap state is described by
the memory section and the sparse-vmemmap population path can allocate or
reuse shared tail vmemmap pages based on that metadata. Device DAX still
uses the older DAX-specific population model, including a separate tail
vmemmap page reservation and architecture-specific logic to locate or
populate reusable tail pages.
This series makes device DAX use the same section-based model. Device DAX
records the compound page order from pgmap->vmemmap_shift in section
metadata before vmemmap population, uses the common per-zone shared tail
vmemmap page, and drops the extra reserved tail page. The powerpc radix
path is updated to use the same shared tail-page helper, so the generic
and powerpc DAX paths follow the same reservation model.
Sashiko said page->refcount can overflow by incrementing it over 2.14 billion
times when mapping more than **524 TB** of DEV-DAX memory on a single NUMA
node, where the pages share the same node, order, and zone.
I am not aware of any practical hardware configuration approaching this
topology today.
Handling that theoretical limit would add non-trivial lifetime or
architecture-specific teardown complexity. Without a concrete hardware
requirement, I prefer not to over-engineer the current series. We can revisit
it when such a system or use case becomes realistic.
Thanks.
From: Andrew Morton <akpm@linux-foundation.org> Date: 2026-09-27 19:54:53
On Sun, 27 Sep 2026 18:51:15 +0800 Muchun Song [off-list ref] wrote:
quoted
On Sep 27, 2026, at 13:51, Andrew Morton [off-list ref] wrote:
On Sun, 27 Sep 2026 10:54:29 +0800 Muchun Song [off-list ref] wrote:
quoted
After the HugeTLB conversion, optimized vmemmap state is described by
the memory section and the sparse-vmemmap population path can allocate or
reuse shared tail vmemmap pages based on that metadata. Device DAX still
uses the older DAX-specific population model, including a separate tail
vmemmap page reservation and architecture-specific logic to locate or
populate reusable tail pages.
This series makes device DAX use the same section-based model. Device DAX
records the compound page order from pgmap->vmemmap_shift in section
metadata before vmemmap population, uses the common per-zone shared tail
vmemmap page, and drops the extra reserved tail page. The powerpc radix
path is updated to use the same shared tail-page helper, so the generic
and powerpc DAX paths follow the same reservation model.
Sashiko said page->refcount can overflow by incrementing it over 2.14 billion
times when mapping more than **524 TB** of DEV-DAX memory on a single NUMA
node, where the pages share the same node, order, and zone.
I am not aware of any practical hardware configuration approaching this
topology today.
Handling that theoretical limit would add non-trivial lifetime or
architecture-specific teardown complexity. Without a concrete hardware
requirement, I prefer not to over-engineer the current series. We can revisit
it when such a system or use case becomes realistic.
OK. Presumably it would be cheap to add a check for this craziness and
return ENOSOMETHING?
From: Muchun Song <muchun.song@linux.dev> Date: 2026-09-28 04:26:30
On Sep 28, 2026, at 03:54, Andrew Morton [off-list ref] wrote:
On Sun, 27 Sep 2026 18:51:15 +0800 Muchun Song [off-list ref] wrote:
quoted
quoted
On Sep 27, 2026, at 13:51, Andrew Morton [off-list ref] wrote:
On Sun, 27 Sep 2026 10:54:29 +0800 Muchun Song [off-list ref] wrote:
quoted
After the HugeTLB conversion, optimized vmemmap state is described by
the memory section and the sparse-vmemmap population path can allocate or
reuse shared tail vmemmap pages based on that metadata. Device DAX still
uses the older DAX-specific population model, including a separate tail
vmemmap page reservation and architecture-specific logic to locate or
populate reusable tail pages.
This series makes device DAX use the same section-based model. Device DAX
records the compound page order from pgmap->vmemmap_shift in section
metadata before vmemmap population, uses the common per-zone shared tail
vmemmap page, and drops the extra reserved tail page. The powerpc radix
path is updated to use the same shared tail-page helper, so the generic
and powerpc DAX paths follow the same reservation model.
Sashiko said page->refcount can overflow by incrementing it over 2.14 billion
times when mapping more than **524 TB** of DEV-DAX memory on a single NUMA
node, where the pages share the same node, order, and zone.
I am not aware of any practical hardware configuration approaching this
topology today.
Handling that theoretical limit would add non-trivial lifetime or
architecture-specific teardown complexity. Without a concrete hardware
requirement, I prefer not to over-engineer the current series. We can revisit
it when such a system or use case becomes realistic.
OK. Presumably it would be cheap to add a check for this craziness and
return ENOSOMETHING?
Sounds right — I'll send a follow-up fixup patch shortly.
Muchun,
Thanks.
From: Muchun Song <hidden> Date: 2026-09-28 04:41:58
Each PTE mapping the shared device DAX tail page takes a page reference.
A sufficiently large range could therefore cycle the reference count back
to zero if population were allowed to continue after it became
non-positive.
Use try_get_page() so further mappings fail once the reference count is no
longer positive. The section population error path tears down mappings
created for the failed section, while the warning makes this currently
impractical limit visible.
Signed-off-by: Muchun Song <redacted>
---
mm/sparse-vmemmap.c | 10 +++++++---
1 file changed, 7 insertions(+), 3 deletions(-)
The section-based vmemmap optimization infrastructure is guarded by
CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP, but it can also be used by
ZONE_DEVICE users that set dev_pagemap::vmemmap_shift. Introduce
CONFIG_VMEMMAP_OPTIMIZATION as a common config for the shared
infrastructure.
Select the new option from HUGETLB_PAGE_OPTIMIZE_VMEMMAP and from
ZONE_DEVICE when the architecture opts in to DAX vmemmap optimization,
and use it to guard the generic sparse-vmemmap state and helpers.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
---
v5:
- Move this patch after the shared tail-page factoring.
- Select VMEMMAP_OPTIMIZATION from ZONE_DEVICE instead of DEV_DAX,
covering all users of dev_pagemap::vmemmap_shift, reported by
Sashiko.
v4:
- Rename SPARSEMEM_VMEMMAP_OPTIMIZATION to VMEMMAP_OPTIMIZATION
(suggested by Mike Rapoport)
- Collect Acked-by from Mike Rapoport
v2:
- Fix SPARSEMEM_VMEMMAP_OPTIMIZATION being selected without SPARSEMEM_VMEMMAP
reported by Sashiko.
- Add an explicit DEV_DAX dependency on ZONE_DEVICE
- Collect Acked-by from Qi Zheng
---
arch/x86/entry/vdso/vdso32/fake_32bit_build.h | 2 +-
fs/Kconfig | 1 +
include/linux/mm.h | 3 +++
include/linux/mmzone.h | 10 +++++-----
include/linux/page-flags.h | 5 ++---
mm/Kconfig | 5 +++++
mm/sparse-vmemmap.c | 2 +-
mm/sparse.h | 6 +++---
8 files changed, 21 insertions(+), 13 deletions(-)
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Is there a path to remove HUGETLB_PAGE_OPTIMIZE_VMEMMAP, and to merge
ARCH_WANT_OPTIMIZE_DAX_VMEMMAP+ARCH_WANT_OPTIMIZE_HUGETLB_VMEMMAP into a
ARCH_SUPPORTS_VMEMMAP_OPTIMIZATION?
--
Cheers,
David
HugeTLB and sparse-vmemmap each have their own helper to allocate the
shared vmemmap tail page used by vmemmap optimization.
Factor that logic into a common vmemmap_shared_tail_page() helper. It
allocates the page through vmemmap_alloc_block(), initializes the tail
struct pages, and uses cmpxchg() to install the per-zone shared page.
This removes duplicate allocation logic while handling both early boot
and runtime allocation through the same helper.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
---
v5:
- Move this patch before CONFIG_VMEMMAP_OPTIMIZATION is introduced
v4:
- Update the commit message for the renamed VMEMMAP_OPTIMIZATION config
- Collect Acked-by from Mike Rapoport
v2:
- Collect Acked-by from Qi Zheng
---
[...]
quoted hunk
#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP-static __meminit struct page *vmemmap_get_tail(unsigned int order, struct zone *zone)+struct page __ref *vmemmap_shared_tail_page(unsigned int order, struct zone *zone) {- struct page *p, *tail;- unsigned int idx;- int node = zone_to_nid(zone);+ void *addr;+ struct page *page;+ const unsigned int idx = order - VMEMMAP_OPTIMIZATION_MIN_ORDER;
Nit: constants read much nicer all the way at the top.
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
--
Cheers,
David
Commit 622026e87c40 ("mm/hugetlb: remove fake head pages") added the
per-zone vmemmap_tails array. Its size depends on MAX_FOLIO_ORDER, which
had been moved to mmzone.h in preparation for the array.
PUD_ORDER is defined by linux/pgtable.h, which cannot be included from
mmzone.h without creating an include cycle. It was therefore open-coded
as PUD_SHIFT - PAGE_SHIFT.
This removed the dependency on PUD_ORDER, but not the underlying
dependency on architecture page-table definitions. PUD_SHIFT is
generally provided by architecture page-table headers, which are not
guaranteed to have been included when mmzone.h is parsed.
The dependency remained hidden because vmemmap_tails was originally
guarded by CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP. Under that condition,
MAX_FOLIO_ORDER resolves to either MAX_PAGE_ORDER or the fixed HugeTLB
limit, rather than the PUD_SHIFT-based definition.
Device DAX, however, does not require CONFIG_HUGETLB_PAGE. When it is
converted to use section-based vmemmap optimization, MAX_FOLIO_ORDER can
resolve to PUD_SHIFT - PAGE_SHIFT while it is being used to size
vmemmap_tails. This would make struct zone depend on architecture
page-table definitions being available when mmzone.h is parsed.
Replace the embedded array with a pointer and allocate it on first use.
This moves the order-count evaluation into sparse-vmemmap.c, after the
architecture page-table definitions are available, and removes the
dependency from mmzone.h.
Removing the compile-time array also removes the original reason for
keeping MAX_FOLIO_ORDER and the vmemmap optimization sizing definitions
in mmzone.h. Follow-up cleanups can place each definition in the header
owned by its respective subsystem.
Signed-off-by: Muchun Song <redacted>
---
v5:
- Add this patch to fix the RISC-V build failure under the configuration
reported by the kernel test robot
---
include/linux/mmzone.h | 7 +------
mm/sparse-vmemmap.c | 35 +++++++++++++++++++++++++++++++----
2 files changed, 32 insertions(+), 10 deletions(-)
@@ -1156,7 +1151,7 @@ struct zone {atomic_long_tvm_stat[NR_VM_ZONE_STAT_ITEMS];atomic_long_tvm_numa_event[NR_VM_NUMA_EVENT_ITEMS];#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP-structpage*vmemmap_tails[VMEMMAP_OPTIMIZATION_NR_ORDERS];+structpage**vmemmap_tails;#endif}____cacheline_internodealigned_in_smp;
[...]
quoted hunk
struct page __ref *vmemmap_shared_tail_page(unsigned int order, struct zone *zone) { void *addr;- struct page *page;+ struct page *page, **pages; const unsigned int idx = order - VMEMMAP_OPTIMIZATION_MIN_ORDER; if (WARN_ON_ONCE(idx >= VMEMMAP_OPTIMIZATION_NR_ORDERS)) return NULL;- page = READ_ONCE(zone->vmemmap_tails[idx]);+ pages = READ_ONCE(zone->vmemmap_tails) ? : vmemmap_tails_alloc(zone);
This reads much nicer if you handle the READ_ONCE(zone->vmemmap_tails) inside
the function.
pages = vmemmap_tails(zone);
So just place the entire logic of obtaining the array in there.
Apart from that LGTM.
--
Cheers,
David
Device DAX still uses vmemmap_populate_compound_pages() to populate its
compound-page vmemmap mappings. That helper allocates the head and first
tail vmemmap pages explicitly, then reuses the first tail page for the
remaining tail page mappings.
Device DAX is being moved to the section-based vmemmap optimization
infrastructure, but it cannot switch to the generic section-based
population path yet. Once a later patch records the DAX compound page
order in section metadata, DAX head and first-tail PFNs can look
optimizable to the generic helpers as well.
Add a DAX-specific population flag for this transition. It keeps DAX
Well, you're not adding flag, your reusing an existing one and renaming it?
And then you're specifying it on more paths.
[...]
@@ -35,8 +35,8 @@/**Flagsforvmemmap_populate_rangeandfriends.*/-/* Get a ref on the head page struct page, for ZONE_DEVICE compound pages */-#define VMEMMAP_POPULATE_PAGEREF 0x0001+/* Vmemmap population for ZONE_DEVICE compound pages */+#define VMEMMAP_POPULATE_DAX 0x0001
Cleaner.
quoted hunk
#include "internal.h"
#include "mm_init.h"
@@ -243,13 +243,17 @@ static inline struct page *vmemmap_shared_tail_page(unsigned int order, #endif static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,- struct vmem_altmap *altmap)+ struct vmem_altmap *altmap, unsigned long flags) { struct zone *zone; struct page *page; const unsigned int order = pfn_to_section_compound_order(pfn);- if (!vmemmap_optimizable_pfn(pfn))+ /*+ * Device DAX still relies on vmemmap_populate_compound_pages() for+ * head/first-tail allocation and tail-page reuse.+ */+ if (!vmemmap_optimizable_pfn(pfn) || flags & VMEMMAP_POPULATE_DAX) return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap); zone = pfn_to_zone(pfn, node);
@@ -271,7 +275,7 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in pte_t entry; if (ptpfn == (unsigned long)-1) {- void *p = vmemmap_alloc_pte(pfn, node, altmap);+ void *p = vmemmap_alloc_pte(pfn, node, altmap, flags); if (!p) return NULL;
@@ -286,7 +290,7 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in * and through vmemmap_populate_compound_pages() when * slab is available. */- if (flags & VMEMMAP_POPULATE_PAGEREF)+ if (flags & VMEMMAP_POPULATE_DAX) get_page(pfn_to_page(ptpfn)); } entry = pfn_pte(ptpfn, PAGE_KERNEL);
@@ -546,6 +550,7 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn, unsigned long size, addr; pte_t *pte; int rc;+ unsigned long flags = VMEMMAP_POPULATE_DAX;
Device DAX can use vmemmap optimization only when a full section is
populated with a compound-page geometry. Record that geometry as the
compound page order in section metadata before populating the section, so
later vmemmap accounting and population decisions can use the section state
directly.
Clear the compound page order when the section becomes empty again. Also
reject partial additions to a section that already has optimized vmemmap
mappings. compound_nr_pages() determines how many struct pages to
initialize with a section as the smallest granularity. A section therefore
cannot safely mix optimized and ordinary vmemmap layouts.
Partial additions continue to use ordinary vmemmap population, so they do
not save vmemmap memory. Such additions are uncommon, and the lost saving
is negligible.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Update the subject and commit message to use compound page order
terminology
- Use EOPNOTSUPP instead of ENOTSUPP
v2:
- Explain why optimized and ordinary layouts cannot share a section
(suggested by Qi Zheng)
- Collect Acked-by from Qi Zheng
---
[...]>
quoted hunk
static struct page * __meminit section_activate(int nid, unsigned long pfn,
@@ -838,8 +840,13 @@ static struct page * __meminit section_activate(int nid, unsigned long pfn, struct mem_section *ms = __pfn_to_section(pfn); struct mem_section_usage *usage = NULL; struct page *memmap;+ unsigned int order; int rc;+ order = vmemmap_can_optimize(altmap, pgmap) ? pgmap->vmemmap_shift : 0;+ if (nr_pages < PAGES_PER_SECTION && section_compound_order(ms))+ return ERR_PTR(-EOPNOTSUPP);
Hm. Why should we support optimizing the vmemmap in case we fall into the same
memory section as boot memory?
In that case, there already is a memmap allocated during boot for the entire
section. IOW, we really shouldn't mess with the vmemmap in case we have an early
section.
But maybe I am missing something and this is already disallowed?
--
Cheers,
David
HugeTLB vmemmap optimization now uses per-zone shared tail vmemmap pages.
Device DAX has not been switched to that mechanism yet.
Switch device DAX to vmemmap_shared_tail_page() as well. This aligns DAX
with HugeTLB by using the common per-zone shared tail vmemmap page.
The optimization is enabled only for DEV-DAX through pgmap->vmemmap_shift,
which supplies the compound page order recorded in section metadata before
vmemmap population. Unlike FS-DAX, DEV-DAX does not modify tail struct
pages, so sharing them is safe.
Since the shared tail page can now back ZONE_DEVICE vmemmap mappings,
initialize its entries with PG_reserved for device zones. Also skip
poisoning vmemmap-optimizable sections while their struct pages may be
shared.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Move device_zone() after the definition of NODE_DATA() to fix
non-NUMA builds.
- Update the commit message to describe the compound page order stored
in section metadata
- Collect Acked-by from Qi Zheng
v2:
- Explain why sharing tail vmemmap pages is safe for DEV-DAX
(suggested by Qi Zheng)
---
[...]
quoted hunk
- static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn, unsigned long start, unsigned long end, int node,
@@ -551,21 +536,18 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn, pte_t *pte; int rc; unsigned long flags = VMEMMAP_POPULATE_DAX;+ struct page *page;+ unsigned int order = pfn_to_section_compound_order(start_pfn);
const and all the way to the top.
I did wonder about the poisoning change ... because the memmap usually gets
initialized once the memory section gets moved to a zone.
SO now I'm a bit confused about the ordering of events :)
--
Cheers,
David
The vmemmap optimization helpers currently live in mm/sparse.h,
which is an internal MM header. That works for MM code, but
prevents powerpc from using the same interfaces without including a
private header.
Move the declarations and inline helpers to vmemmap-optimization.h.
This is a preparatory change for powerpc, which has its own vmemmap
optimization implementation and needs to use the common vmemmap
optimization interfaces from architecture code.
Which raises the question why powerpc was special and will remain special. Wha's
the big problem here that powerpc must do special things?
Change itself looks good.
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
--
Cheers,
David
section_nr_vmemmap_pages() no longer uses the altmap or pgmap
arguments, so drop them from the helper and its callers.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Collect Acked-by from Qi Zheng
---
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
--
Cheers,
David
Device DAX now uses the common per-zone shared tail page for vmemmap
deduplication. The old documentation still described a DAX-specific
layout with a separately populated tail vmemmap page and half the HugeTLB
savings.
Update the generic and powerpc documentation to describe the shared layout.
In the powerpc document, keep the radix and 64K-specific details, drop the
duplicated 4K PUD arithmetic, and replace the repeated device-dax diagrams
with a single parameterized PMD/PUD diagram.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
--
Cheers,
David
The device DAX vmemmap population still reserves one extra tail vmemmap
page after the head page.
Drop that extra reservation and let the shared tail page cover all tail
vmemmap pages after the head page, so DAX follows the same reservation
model as HugeTLB.
This reduces the reserved vmemmap pages for optimized DAX mappings to
one and removes the now-unneeded first-tail population from the generic
and powerpc paths to simplify the code as well.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
--
Cheers,
David
From: Muchun Song <muchun.song@linux.dev> Date: 2026-09-29 07:53:32
On Sep 29, 2026, at 15:11, David Hildenbrand (Arm) [off-list ref] wrote:
On 9/27/26 04:54, Muchun Song wrote:
quoted
The section-based vmemmap optimization infrastructure is guarded by
CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP, but it can also be used by
ZONE_DEVICE users that set dev_pagemap::vmemmap_shift. Introduce
CONFIG_VMEMMAP_OPTIMIZATION as a common config for the shared
infrastructure.
Select the new option from HUGETLB_PAGE_OPTIMIZE_VMEMMAP and from
ZONE_DEVICE when the architecture opts in to DAX vmemmap optimization,
and use it to guard the generic sparse-vmemmap state and helpers.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
---
v5:
- Move this patch after the shared tail-page factoring.
- Select VMEMMAP_OPTIMIZATION from ZONE_DEVICE instead of DEV_DAX,
covering all users of dev_pagemap::vmemmap_shift, reported by
Sashiko.
v4:
- Rename SPARSEMEM_VMEMMAP_OPTIMIZATION to VMEMMAP_OPTIMIZATION
(suggested by Mike Rapoport)
- Collect Acked-by from Mike Rapoport
v2:
- Fix SPARSEMEM_VMEMMAP_OPTIMIZATION being selected without SPARSEMEM_VMEMMAP
reported by Sashiko.
- Add an explicit DEV_DAX dependency on ZONE_DEVICE
- Collect Acked-by from Qi Zheng
---
arch/x86/entry/vdso/vdso32/fake_32bit_build.h | 2 +-
fs/Kconfig | 1 +
include/linux/mm.h | 3 +++
include/linux/mmzone.h | 10 +++++-----
include/linux/page-flags.h | 5 ++---
mm/Kconfig | 5 +++++
mm/sparse-vmemmap.c | 2 +-
mm/sparse.h | 6 +++---
8 files changed, 21 insertions(+), 13 deletions(-)
def_bool HUGETLB_PAGE
depends on ARCH_WANT_OPTIMIZE_HUGETLB_VMEMMAP
depends on SPARSEMEM_VMEMMAP
+ select VMEMMAP_OPTIMIZATION
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Thanks.
Is there a path to remove HUGETLB_PAGE_OPTIMIZE_VMEMMAP, and to merge
ARCH_WANT_OPTIMIZE_DAX_VMEMMAP+ARCH_WANT_OPTIMIZE_HUGETLB_VMEMMAP into a
ARCH_SUPPORTS_VMEMMAP_OPTIMIZATION?
These are actually two completely different capabilities.
ARCH_WANT_OPTIMIZE_HUGETLB_VMEMMAP requires the architecture
to support dynamic updates to vmemmap page tables, meaning a
PTE entry can be changed from one valid entry to another
valid entry. This does not meet the requirements on arm64,
because arm64 requires page table operations to satisfy BBM
(there is, of course, a series [1] attempting to do this).
However, for ARCH_WANT_OPTIMIZE_DAX_VMEMMAP, the vmemmap page
tables do not involve dynamic updates, so the BBM requirement
can be satisfied. Therefore, arm64 can enable
ARCH_WANT_OPTIMIZE_DAX_VMEMMAP, but cannot enable
ARCH_WANT_OPTIMIZE_HUGETLB_VMEMMAP. To make the naming clearer,
I have another patch [2] that renames it for greater clarity.
As for ARCH_WANT_OPTIMIZE_DAX_VMEMMAP, I plan to remove it
entirely in the future, because architectures that do not support
it can simply choose to disable it, as can be seen in patch [3].
So in my plan, ultimately only one config will remain:
ARCH_SUPPORTS_VMEMMAP_REMAP.
I hope this clarifies the plan. Let me know what you think.
[1] https://lore.kernel.org/20260708031129.3503195-1-jthoughton@google.com/
[2] https://lore.kernel.org/20260903122128.12264-2-songmuchun@bytedance.com/
[3] https://lore.kernel.org/20260513132044.41690-6-songmuchun@bytedance.com/
Thanks,
Muchun
From: Muchun Song <muchun.song@linux.dev> Date: 2026-09-29 07:55:49
On Sep 29, 2026, at 15:16, David Hildenbrand (Arm) [off-list ref] wrote:
On 9/27/26 04:54, Muchun Song wrote:
quoted
HugeTLB and sparse-vmemmap each have their own helper to allocate the
shared vmemmap tail page used by vmemmap optimization.
Factor that logic into a common vmemmap_shared_tail_page() helper. It
allocates the page through vmemmap_alloc_block(), initializes the tail
struct pages, and uses cmpxchg() to install the per-zone shared page.
This removes duplicate allocation logic while handling both early boot
and runtime allocation through the same helper.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
---
v5:
- Move this patch before CONFIG_VMEMMAP_OPTIMIZATION is introduced
v4:
- Update the commit message for the renamed VMEMMAP_OPTIMIZATION config
- Collect Acked-by from Mike Rapoport
v2:
- Collect Acked-by from Qi Zheng
---
[...]
quoted
#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP
-static __meminit struct page *vmemmap_get_tail(unsigned int order, struct zone *zone)
+struct page __ref *vmemmap_shared_tail_page(unsigned int order, struct zone *zone)
{
- struct page *p, *tail;
- unsigned int idx;
- int node = zone_to_nid(zone);
+ void *addr;
+ struct page *page;
+ const unsigned int idx = order - VMEMMAP_OPTIMIZATION_MIN_ORDER;
Nit: constants read much nicer all the way at the top.
I can update to this next version.
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
From: Muchun Song <muchun.song@linux.dev> Date: 2026-09-29 08:01:08
On Sep 29, 2026, at 15:21, David Hildenbrand (Arm) [off-list ref] wrote:
On 9/27/26 04:54, Muchun Song wrote:
quoted
Commit 622026e87c40 ("mm/hugetlb: remove fake head pages") added the
per-zone vmemmap_tails array. Its size depends on MAX_FOLIO_ORDER, which
had been moved to mmzone.h in preparation for the array.
PUD_ORDER is defined by linux/pgtable.h, which cannot be included from
mmzone.h without creating an include cycle. It was therefore open-coded
as PUD_SHIFT - PAGE_SHIFT.
This removed the dependency on PUD_ORDER, but not the underlying
dependency on architecture page-table definitions. PUD_SHIFT is
generally provided by architecture page-table headers, which are not
guaranteed to have been included when mmzone.h is parsed.
The dependency remained hidden because vmemmap_tails was originally
guarded by CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP. Under that condition,
MAX_FOLIO_ORDER resolves to either MAX_PAGE_ORDER or the fixed HugeTLB
limit, rather than the PUD_SHIFT-based definition.
Device DAX, however, does not require CONFIG_HUGETLB_PAGE. When it is
converted to use section-based vmemmap optimization, MAX_FOLIO_ORDER can
resolve to PUD_SHIFT - PAGE_SHIFT while it is being used to size
vmemmap_tails. This would make struct zone depend on architecture
page-table definitions being available when mmzone.h is parsed.
Replace the embedded array with a pointer and allocate it on first use.
This moves the order-count evaluation into sparse-vmemmap.c, after the
architecture page-table definitions are available, and removes the
dependency from mmzone.h.
Removing the compile-time array also removes the original reason for
keeping MAX_FOLIO_ORDER and the vmemmap optimization sizing definitions
in mmzone.h. Follow-up cleanups can place each definition in the header
owned by its respective subsystem.
Signed-off-by: Muchun Song <redacted>
---
v5:
- Add this patch to fix the RISC-V build failure under the configuration
reported by the kernel test robot
---
include/linux/mmzone.h | 7 +------
mm/sparse-vmemmap.c | 35 +++++++++++++++++++++++++++++++----
2 files changed, 32 insertions(+), 10 deletions(-)
From: Muchun Song <muchun.song@linux.dev> Date: 2026-09-29 08:05:18
On Sep 29, 2026, at 15:24, David Hildenbrand (Arm) [off-list ref] wrote:
On 9/27/26 04:54, Muchun Song wrote:
quoted
Device DAX still uses vmemmap_populate_compound_pages() to populate its
compound-page vmemmap mappings. That helper allocates the head and first
tail vmemmap pages explicitly, then reuses the first tail page for the
remaining tail page mappings.
Device DAX is being moved to the section-based vmemmap optimization
infrastructure, but it cannot switch to the generic section-based
population path yet. Once a later patch records the DAX compound page
order in section metadata, DAX head and first-tail PFNs can look
optimizable to the generic helpers as well.
Add a DAX-specific population flag for this transition. It keeps DAX
Well, you're not adding flag, your reusing an existing one and renaming it?
And then you're specifying it on more paths.
You're right. The commit message need to be more precise.
/*
* Flags for vmemmap_populate_range and friends.
*/
-/* Get a ref on the head page struct page, for ZONE_DEVICE compound pages */
-#define VMEMMAP_POPULATE_PAGEREF 0x0001
+/* Vmemmap population for ZONE_DEVICE compound pages */
+#define VMEMMAP_POPULATE_DAX 0x0001
@@ -286,7 +290,7 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
* and through vmemmap_populate_compound_pages() when
* slab is available.
*/
- if (flags & VMEMMAP_POPULATE_PAGEREF)
+ if (flags & VMEMMAP_POPULATE_DAX)
get_page(pfn_to_page(ptpfn));
}
entry = pfn_pte(ptpfn, PAGE_KERNEL);
@@ -546,6 +550,7 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
unsigned long size, addr;
pte_t *pte;
int rc;
+ unsigned long flags = VMEMMAP_POPULATE_DAX;
From: Muchun Song <muchun.song@linux.dev> Date: 2026-09-29 08:22:52
On Sep 29, 2026, at 15:30, David Hildenbrand (Arm) [off-list ref] wrote:
On 9/27/26 04:54, Muchun Song wrote:
quoted
Device DAX can use vmemmap optimization only when a full section is
populated with a compound-page geometry. Record that geometry as the
compound page order in section metadata before populating the section, so
later vmemmap accounting and population decisions can use the section state
directly.
Clear the compound page order when the section becomes empty again. Also
reject partial additions to a section that already has optimized vmemmap
mappings. compound_nr_pages() determines how many struct pages to
initialize with a section as the smallest granularity. A section therefore
cannot safely mix optimized and ordinary vmemmap layouts.
Partial additions continue to use ordinary vmemmap population, so they do
not save vmemmap memory. Such additions are uncommon, and the lost saving
is negligible.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Update the subject and commit message to use compound page order
terminology
- Use EOPNOTSUPP instead of ENOTSUPP
v2:
- Explain why optimized and ordinary layouts cannot share a section
(suggested by Qi Zheng)
- Collect Acked-by from Qi Zheng
---
[...]>
quoted
static struct page * __meminit section_activate(int nid, unsigned long pfn,
struct mem_section *ms = __pfn_to_section(pfn);
struct mem_section_usage *usage = NULL;
struct page *memmap;
+ unsigned int order;
int rc;
+ order = vmemmap_can_optimize(altmap, pgmap) ? pgmap->vmemmap_shift : 0;
+ if (nr_pages < PAGES_PER_SECTION && section_compound_order(ms))
+ return ERR_PTR(-EOPNOTSUPP);
Hm. Why should we support optimizing the vmemmap in case we fall into the same
memory section as boot memory?
In that case, there already is a memmap allocated during boot for the entire
section. IOW, we really shouldn't mess with the vmemmap in case we have an early
section.
But maybe I am missing something and this is already disallowed?
Yes, this is already handled.
For a partial addition to a normal early section, after updating the
subsection map we return the existing boot-time memmap here:
if (nr_pages < PAGES_PER_SECTION && early_section(ms))
return pfn_to_page(pfn);
Therefore, neither section_set_compound_order_range() nor
populate_section_memmap() is called. The fully populated boot memmap is
simply reused, and no vmemmap optimization is attempted.
The check above handles the other case: if the section already has an
optimized vmemmap layout, as indicated by section_compound_order(ms),
a partial addition is rejected because we cannot mix optimized and
ordinary vmemmap layouts within one section. This also covers an early
section whose vmemmap was already optimized during boot.
Thanks,
Muchun
From: Muchun Song <muchun.song@linux.dev> Date: 2026-09-29 08:36:32
On Sep 29, 2026, at 15:36, David Hildenbrand (Arm) [off-list ref] wrote:
On 9/27/26 04:54, Muchun Song wrote:
quoted
HugeTLB vmemmap optimization now uses per-zone shared tail vmemmap pages.
Device DAX has not been switched to that mechanism yet.
Switch device DAX to vmemmap_shared_tail_page() as well. This aligns DAX
with HugeTLB by using the common per-zone shared tail vmemmap page.
The optimization is enabled only for DEV-DAX through pgmap->vmemmap_shift,
which supplies the compound page order recorded in section metadata before
vmemmap population. Unlike FS-DAX, DEV-DAX does not modify tail struct
pages, so sharing them is safe.
Since the shared tail page can now back ZONE_DEVICE vmemmap mappings,
initialize its entries with PG_reserved for device zones. Also skip
poisoning vmemmap-optimizable sections while their struct pages may be
shared.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Move device_zone() after the definition of NODE_DATA() to fix
non-NUMA builds.
- Update the commit message to describe the compound page order stored
in section metadata
- Collect Acked-by from Qi Zheng
v2:
- Explain why sharing tail vmemmap pages is safe for DEV-DAX
(suggested by Qi Zheng)
---
[...]
quoted
-
static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
unsigned long start,
unsigned long end, int node,
@@ -551,21 +536,18 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
pte_t *pte;
int rc;
unsigned long flags = VMEMMAP_POPULATE_DAX;
+ struct page *page;
+ unsigned int order = pfn_to_section_compound_order(start_pfn);
const and all the way to the top.
OK.
I did wonder about the poisoning change ... because the memmap usually gets
initialized once the memory section gets moved to a zone.
SO now I'm a bit confused about the ordering of events :)
Yes, the normal memmap entries are initialized later when the range is
moved into the zone. The shared tail entries are the exception, though.
The ordering for device DAX is:
1. sparse_add_section() sets the section compound order and populates
the vmemmap.
2. During vmemmap population, optimizable tail entries are mapped to
the per-zone shared tail page, which is initialized by
vmemmap_shared_tail_page().
3. page_init_poison() is reached after that population.
4. Later, move_pfn_range_to_zone() calls memmap_init_range(), but the
latter deliberately skips vmemmap_optimizable_pfn() because those
entries have already been initialized.
Therefore, an unconditional poison here would overwrite the initialized
shared tail page, and the later zone initialization would not restore it.
The non-shared entries are still initialized later as usual.
I agree that this ordering is not obvious. I will update the comment to
make it clearer, for example:
/*
* Poison uninitialized struct pages to catch invalid flag combinations.
*
* Tail struct pages in a vmemmap-optimized section are initialized and
* shared during vmemmap population, so they must not be overwritten here.
*/
Thanks,
Muchun
On Sep 29, 2026, at 15:11, David Hildenbrand (Arm) [off-list ref] wrote:
On 9/27/26 04:54, Muchun Song wrote:
quoted
The section-based vmemmap optimization infrastructure is guarded by
CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP, but it can also be used by
ZONE_DEVICE users that set dev_pagemap::vmemmap_shift. Introduce
CONFIG_VMEMMAP_OPTIMIZATION as a common config for the shared
infrastructure.
Select the new option from HUGETLB_PAGE_OPTIMIZE_VMEMMAP and from
ZONE_DEVICE when the architecture opts in to DAX vmemmap optimization,
and use it to guard the generic sparse-vmemmap state and helpers.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
---
v5:
- Move this patch after the shared tail-page factoring.
- Select VMEMMAP_OPTIMIZATION from ZONE_DEVICE instead of DEV_DAX,
covering all users of dev_pagemap::vmemmap_shift, reported by
Sashiko.
v4:
- Rename SPARSEMEM_VMEMMAP_OPTIMIZATION to VMEMMAP_OPTIMIZATION
(suggested by Mike Rapoport)
- Collect Acked-by from Mike Rapoport
v2:
- Fix SPARSEMEM_VMEMMAP_OPTIMIZATION being selected without SPARSEMEM_VMEMMAP
reported by Sashiko.
- Add an explicit DEV_DAX dependency on ZONE_DEVICE
- Collect Acked-by from Qi Zheng
---
arch/x86/entry/vdso/vdso32/fake_32bit_build.h | 2 +-
fs/Kconfig | 1 +
include/linux/mm.h | 3 +++
include/linux/mmzone.h | 10 +++++-----
include/linux/page-flags.h | 5 ++---
mm/Kconfig | 5 +++++
mm/sparse-vmemmap.c | 2 +-
mm/sparse.h | 6 +++---
8 files changed, 21 insertions(+), 13 deletions(-)
def_bool HUGETLB_PAGE
depends on ARCH_WANT_OPTIMIZE_HUGETLB_VMEMMAP
depends on SPARSEMEM_VMEMMAP
+ select VMEMMAP_OPTIMIZATION
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Thanks.
quoted
Is there a path to remove HUGETLB_PAGE_OPTIMIZE_VMEMMAP, and to merge
ARCH_WANT_OPTIMIZE_DAX_VMEMMAP+ARCH_WANT_OPTIMIZE_HUGETLB_VMEMMAP into a
ARCH_SUPPORTS_VMEMMAP_OPTIMIZATION?
These are actually two completely different capabilities.
ARCH_WANT_OPTIMIZE_HUGETLB_VMEMMAP requires the architecture
to support dynamic updates to vmemmap page tables, meaning a
PTE entry can be changed from one valid entry to another
valid entry. This does not meet the requirements on arm64,
because arm64 requires page table operations to satisfy BBM
(there is, of course, a series [1] attempting to do this).
However, for ARCH_WANT_OPTIMIZE_DAX_VMEMMAP, the vmemmap page
tables do not involve dynamic updates, so the BBM requirement
can be satisfied. Therefore, arm64 can enable
ARCH_WANT_OPTIMIZE_DAX_VMEMMAP, but cannot enable
ARCH_WANT_OPTIMIZE_HUGETLB_VMEMMAP. To make the naming clearer,
I have another patch [2] that renames it for greater clarity.
Ah, perfect. Too many patches floating around :)
I thought there is a patch set to avoid the BBM requirement on arm64 from James,
though. So not sure if both things will always stay separate.
As for ARCH_WANT_OPTIMIZE_DAX_VMEMMAP, I plan to remove it
entirely in the future, because architectures that do not support
it can simply choose to disable it, as can be seen in patch [3].
So in my plan, ultimately only one config will remain:
ARCH_SUPPORTS_VMEMMAP_REMAP.
I hope this clarifies the plan. Let me know what you think.
On Sep 29, 2026, at 15:30, David Hildenbrand (Arm) [off-list ref] wrote:
On 9/27/26 04:54, Muchun Song wrote:
quoted
Device DAX can use vmemmap optimization only when a full section is
populated with a compound-page geometry. Record that geometry as the
compound page order in section metadata before populating the section, so
later vmemmap accounting and population decisions can use the section state
directly.
Clear the compound page order when the section becomes empty again. Also
reject partial additions to a section that already has optimized vmemmap
mappings. compound_nr_pages() determines how many struct pages to
initialize with a section as the smallest granularity. A section therefore
cannot safely mix optimized and ordinary vmemmap layouts.
Partial additions continue to use ordinary vmemmap population, so they do
not save vmemmap memory. Such additions are uncommon, and the lost saving
is negligible.
Signed-off-by: Muchun Song <redacted>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Update the subject and commit message to use compound page order
terminology
- Use EOPNOTSUPP instead of ENOTSUPP
v2:
- Explain why optimized and ordinary layouts cannot share a section
(suggested by Qi Zheng)
- Collect Acked-by from Qi Zheng
---
[...]>
quoted
static struct page * __meminit section_activate(int nid, unsigned long pfn,
struct mem_section *ms = __pfn_to_section(pfn);
struct mem_section_usage *usage = NULL;
struct page *memmap;
+ unsigned int order;
int rc;
+ order = vmemmap_can_optimize(altmap, pgmap) ? pgmap->vmemmap_shift : 0;
+ if (nr_pages < PAGES_PER_SECTION && section_compound_order(ms))
+ return ERR_PTR(-EOPNOTSUPP);
Hm. Why should we support optimizing the vmemmap in case we fall into the same
memory section as boot memory?
In that case, there already is a memmap allocated during boot for the entire
section. IOW, we really shouldn't mess with the vmemmap in case we have an early
section.
But maybe I am missing something and this is already disallowed?
Yes, this is already handled.
For a partial addition to a normal early section, after updating the
subsection map we return the existing boot-time memmap here:
if (nr_pages < PAGES_PER_SECTION && early_section(ms))
return pfn_to_page(pfn);
Therefore, neither section_set_compound_order_range() nor
populate_section_memmap() is called. The fully populated boot memmap is
simply reused, and no vmemmap optimization is attempted.
Perfect, thanks
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
--
Cheers,
David
From: Muchun Song <muchun.song@linux.dev> Date: 2026-09-29 08:44:35
On Sep 29, 2026, at 15:39, David Hildenbrand (Arm) [off-list ref] wrote:
On 9/27/26 04:54, Muchun Song wrote:
quoted
The vmemmap optimization helpers currently live in mm/sparse.h,
which is an internal MM header. That works for MM code, but
prevents powerpc from using the same interfaces without including a
private header.
Move the declarations and inline helpers to vmemmap-optimization.h.
This is a preparatory change for powerpc, which has its own vmemmap
optimization implementation and needs to use the common vmemmap
optimization interfaces from architecture code.
Which raises the question why powerpc was special and will remain special. Wha's
the big problem here that powerpc must do special things?
Good question. I also don't think PowerPC needs special handling,
but when HVO logic was introduced for PowerPC, it handled HVO on
its own. From my preliminary analysis, the reason it didn't reuse
the generic logic initially may be related to the fact that
PowerPC's section size is 16M. With a 64k base page, a single page
can cover the vmemmap range of multiple sections, and the current
generic logic doesn't cover this case.
However, completely removing PowerPC's special handling is already
in my follow-up plan. We need to wait for the current series to enter
the mainline, and then we can proceed gradually.
Change itself looks good.
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
The powerpc radix compound vmemmap population path still finds a reusable
tail page by walking the vmemmap page tables.
Switch it to the common vmemmap_shared_tail_page() helper instead, so it
can use the shared vmemmap page directly to simplify the code.
This removes the powerpc-specific tail-page lookup and its fallback path
and aligns the device DAX vmemmap optimization path with HugeTLB.
Signed-off-by: Muchun Song <redacted>
---
I did not grasp all the complexity of the old approach, but what i read here all
makes sense to me.
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
--
Cheers,
David
From: Muchun Song <muchun.song@linux.dev> Date: 2026-09-29 10:03:49
On Sep 29, 2026, at 16:44, Muchun Song [off-list ref] wrote:
quoted
On Sep 29, 2026, at 15:39, David Hildenbrand (Arm) [off-list ref] wrote:
On 9/27/26 04:54, Muchun Song wrote:
quoted
The vmemmap optimization helpers currently live in mm/sparse.h,
which is an internal MM header. That works for MM code, but
prevents powerpc from using the same interfaces without including a
private header.
Move the declarations and inline helpers to vmemmap-optimization.h.
This is a preparatory change for powerpc, which has its own vmemmap
optimization implementation and needs to use the common vmemmap
optimization interfaces from architecture code.
Which raises the question why powerpc was special and will remain special. Wha's
the big problem here that powerpc must do special things?
Good question. I also don't think PowerPC needs special handling,
but when HVO logic was introduced for PowerPC, it handled HVO on
its own. From my preliminary analysis, the reason it didn't reuse
the generic logic initially may be related to the fact that
PowerPC's section size is 16M. With a 64k base page, a single page
can cover the vmemmap range of multiple sections, and the current
generic logic doesn't cover this case.
I looked at the code in my local branch for removing the PowerPC
vmemmap optimization handling, and I found another issue that needs
to be addressed.
Since PowerPC vmemmap optimization is restricted to Radix, this only
needs to cover the Radix page-table implementation.
The generic vmemmap path currently allocates intermediate page-table
pages with vmemmap_alloc_block_zero(). This bypasses the normal
page-table constructors.
PowerPC Radix uses early_alloc_pgtable() before slab is available. For
runtime population, it uses pud_alloc(), pmd_alloc(), and
pte_alloc_kernel(). These helpers initialize the page-table metadata
and fragment reference counts expected by pud_free(), pmd_free(), and
pte_free_kernel() during hot-remove.
To address this, I plan to update the generic path so that it uses
the normal page-table helpers once slab is available, while retaining
memblock-backed allocations during early boot. Once allocation and
teardown are correctly paired, PowerPC Radix should be able to call
vmemmap_populate_hugepages() directly and remove its duplicate HVO
page-table walk.
Thanks,
Muchun
However, completely removing PowerPC's special handling is already
in my follow-up plan. We need to wait for the current series to enter
the mainline, and then we can proceed gradually.
quoted
Change itself looks good.
Acked-by: David Hildenbrand (Arm) <david@kernel.org>