Thread (39 messages) 39 messages, 4 authors, 9d ago

Re: [PATCH v5 02/12] mm/sparse-vmemmap: allocate shared tail page array dynamically

flat view

From: Muchun Song <muchun.song@linux.dev>
Date: 2026-09-29 08:01:08
Also in: linux-doc, linux-mm, lkml

On Sep 29, 2026, at 15:21, David Hildenbrand (Arm) [off-list ref] wrote:

On 9/27/26 04:54, Muchun Song wrote:
quoted
Commit 622026e87c40 ("mm/hugetlb: remove fake head pages") added the
per-zone vmemmap_tails array. Its size depends on MAX_FOLIO_ORDER, which
had been moved to mmzone.h in preparation for the array.

PUD_ORDER is defined by linux/pgtable.h, which cannot be included from
mmzone.h without creating an include cycle. It was therefore open-coded
as PUD_SHIFT - PAGE_SHIFT.

This removed the dependency on PUD_ORDER, but not the underlying
dependency on architecture page-table definitions. PUD_SHIFT is
generally provided by architecture page-table headers, which are not
guaranteed to have been included when mmzone.h is parsed.

The dependency remained hidden because vmemmap_tails was originally
guarded by CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP. Under that condition,
MAX_FOLIO_ORDER resolves to either MAX_PAGE_ORDER or the fixed HugeTLB
limit, rather than the PUD_SHIFT-based definition.

Device DAX, however, does not require CONFIG_HUGETLB_PAGE. When it is
converted to use section-based vmemmap optimization, MAX_FOLIO_ORDER can
resolve to PUD_SHIFT - PAGE_SHIFT while it is being used to size
vmemmap_tails. This would make struct zone depend on architecture
page-table definitions being available when mmzone.h is parsed.

Replace the embedded array with a pointer and allocate it on first use.
This moves the order-count evaluation into sparse-vmemmap.c, after the
architecture page-table definitions are available, and removes the
dependency from mmzone.h.

Removing the compile-time array also removes the original reason for
keeping MAX_FOLIO_ORDER and the vmemmap optimization sizing definitions
in mmzone.h. Follow-up cleanups can place each definition in the header
owned by its respective subsystem.

Signed-off-by: Muchun Song <redacted>
---
v5:
- Add this patch to fix the RISC-V build failure under the configuration
 reported by the kernel test robot
---
include/linux/mmzone.h |  7 +------
mm/sparse-vmemmap.c    | 35 +++++++++++++++++++++++++++++++----
2 files changed, 32 insertions(+), 10 deletions(-)
diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index acd94cecc0d3..68807ff7f946 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -113,11 +113,6 @@
(VMEMMAP_OPTIMIZATION_PAGES * PAGE_SIZE / sizeof(struct page))
#define VMEMMAP_OPTIMIZATION_MIN_ORDER (ilog2(VMEMMAP_OPTIMIZATION_NR_STRUCT_PAGES) + 1)

-#define __VMEMMAP_OPTIMIZATION_NR_ORDERS \
- 	(MAX_FOLIO_ORDER - VMEMMAP_OPTIMIZATION_MIN_ORDER + 1)
-#define VMEMMAP_OPTIMIZATION_NR_ORDERS \
- 	(__VMEMMAP_OPTIMIZATION_NR_ORDERS > 0 ? __VMEMMAP_OPTIMIZATION_NR_ORDERS : 0)
-
enum migratetype {
MIGRATE_UNMOVABLE,
MIGRATE_MOVABLE,
@@ -1156,7 +1151,7 @@ struct zone {
atomic_long_t vm_stat[NR_VM_ZONE_STAT_ITEMS];
atomic_long_t vm_numa_event[NR_VM_NUMA_EVENT_ITEMS];
#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP
- 	struct page *vmemmap_tails[VMEMMAP_OPTIMIZATION_NR_ORDERS];
+ 	struct page **vmemmap_tails;
#endif
} ____cacheline_internodealigned_in_smp;

[...]
quoted
struct page __ref *vmemmap_shared_tail_page(unsigned int order, struct zone *zone)
{
	void *addr;
- 	struct page *page;
+ 	struct page *page, **pages;
	const unsigned int idx = order - VMEMMAP_OPTIMIZATION_MIN_ORDER;

	 (WARN_ON_ONCE(idx >= VMEMMAP_OPTIMIZATION_NR_ORDERS))
		return NULL;

- 	page = READ_ONCE(zone->vmemmap_tails[idx]);
+ 	pages = READ_ONCE(zone->vmemmap_tails) ? : vmemmap_tails_alloc(zone);
This reads much nicer if you handle the  READ_ONCE(zone->vmemmap_tails) inside
the function.

pages = vmemmap_tails(zone);
Sounds good.
So just place the entire logic of obtaining the array in there.
No problem.

Thanks,
Muchun

Apart from that LGTM.

-- 
Cheers,

David

Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help