From: Baoquan He <hidden> Date: 2024-03-25 14:57:20
In function free_area_init_core(), the code calculating
zone->managed_pages and the subtracting dma_reserve from DMA zone looks
very confusing.
From git history, the code calculating zone->managed_pages was for
zone->present_pages originally. The early rough assignment is for
optimize zone's pcp and water mark setting. Later, managed_pages was
introduced into zone to represent the number of managed pages by buddy.
Now, zone->managed_pages is zeroed out and reset in mem_init() when
calling memblock_free_all(). zone's pcp and wmark setting relying on
actual zone->managed_pages are done later than mem_init() invocation.
So we don't need rush to early calculate and set zone->managed_pages,
just set it as zone->present_pages, will adjust it in mem_init().
And also add a new function calc_nr_kernel_pages() to count up free but
not reserved pages in memblock, then assign it to nr_all_pages and
nr_kernel_pages after memmap pages are allocated.
Changelog:
----------
v1->v2:
=======
These are all suggested by Mike, thanks to him.
- Swap the order of patch 1 and 2 in v1 to describe code change better,
Mike suggested this.
- Change to initializ zone->managed_pages as 0 in free_area_init_core()
as there isn't any page added into buddy system. And also improve the
ambiguous description in log. These are all in patch 4.
Baoquan He (6):
x86: remove unneeded memblock_find_dma_reserve()
mm/mm_init.c: remove the useless dma_reserve
mm/mm_init.c: add new function calc_nr_all_pages()
mm/mm_init.c: remove meaningless calculation of zone->managed_pages in
free_area_init_core()
mm/mm_init.c: remove unneeded calc_memmap_size()
mm/mm_init.c: remove arch_reserved_kernel_pages()
arch/powerpc/include/asm/mmu.h | 4 --
arch/powerpc/kernel/fadump.c | 5 --
arch/x86/include/asm/pgtable.h | 1 -
arch/x86/kernel/setup.c | 2 -
arch/x86/mm/init.c | 47 -------------
include/linux/mm.h | 4 --
mm/mm_init.c | 125 ++++++++-------------------------
7 files changed, 29 insertions(+), 159 deletions(-)
--
2.41.0
From: Baoquan He <hidden> Date: 2024-03-25 14:57:24
Variable dma_reserve and its usage was introduced in commit 0e0b864e069c
("[PATCH] Account for memmap and optionally the kernel image as holes").
Its original purpose was to accounting for the reserved pages in DMA
zone to make DMA zone's watermarks calculation more accurate on x86.
However, currently there's zone->managed_pages to account for all
available pages for buddy, zone->present_pages to account for all
present physical pages in zone. What is more important, on x86,
calculating and setting the zone->managed_pages is a temporary move,
all zone's managed_pages will be zeroed out and reset to the actual
value according to how many pages are added to buddy allocator in
mem_init(). Before mem_init(), no buddy alloction is requested. And
zone's pcp and watermark setting are all done after mem_init(). So,
no need to worry about the DMA zone's setting accuracy during
free_area_init().
Hence, remove memblock_find_dma_reserve() to stop calculating and
setting dma_reserve.
Signed-off-by: Baoquan He <redacted>
---
arch/x86/include/asm/pgtable.h | 1 -
arch/x86/kernel/setup.c | 2 --
arch/x86/mm/init.c | 47 ----------------------------------
3 files changed, 50 deletions(-)
From: Baoquan He <hidden> Date: 2024-03-25 14:57:32
Now nobody calls set_dma_reserve() to set value for dma_reserve, remove
set_dma_reserve(), global variable dma_reserve and the codes using it.
Signed-off-by: Baoquan He <redacted>
---
include/linux/mm.h | 1 -
mm/mm_init.c | 23 -----------------------
2 files changed, 24 deletions(-)
@@ -3210,7 +3210,6 @@ static inline int early_pfn_to_nid(unsigned long pfn)externint__meminitearly_pfn_to_nid(unsignedlongpfn);#endif-externvoidset_dma_reserve(unsignedlongnew_dma_reserve);externvoidmem_init(void);externvoid__initmmap_init(void);
From: Baoquan He <hidden> Date: 2024-03-25 14:57:35
This is a preparation to calculate nr_kernel_pages and nr_all_pages,
both of which will be used later in alloc_large_system_hash().
nr_all_pages counts up all free but not reserved memory in memblock
allocator, including HIGHMEM memory. While nr_kernel_pages counts up
all free but not reserved low memory in memblock allocator, excluding
HIGHMEM memory.
Signed-off-by: Baoquan He <redacted>
---
mm/mm_init.c | 24 ++++++++++++++++++++++++
1 file changed, 24 insertions(+)
From: Baoquan He <hidden> Date: 2024-03-25 14:57:40
Currently, in free_area_init_core(), when initialize zone's field, a
rough value is set to zone->managed_pages. That value is calculated by
(zone->present_pages - memmap_pages).
In the meantime, add the value to nr_all_pages and nr_kernel_pages which
represent all free pages of system (only low memory or including HIGHMEM
memory separately). Both of them are gonna be used in
alloc_large_system_hash().
However, the rough calculation and setting of zone->managed_pages is
meaningless because
a) memmap pages are allocated on units of node in sparse_init() or
alloc_node_mem_map(pgdat); The simple (zone->present_pages -
memmap_pages) is too rough to make sense for zone;
b) the set zone->managed_pages will be zeroed out and reset with
acutal value in mem_init() via memblock_free_all(). Before the
resetting, no buddy allocation request is issued.
Here, remove the meaningless and complicated calculation of
(zone->present_pages - memmap_pages), initialize zone->managed_pages as 0
which reflect its actual value because no any page is added into buddy
system right now. It will be reset in mem_init().
And also remove the assignment of nr_all_pages and nr_kernel_pages in
free_area_init_core(). Instead, call the newly added calc_nr_kernel_pages()
to count up all free but not reserved memory in memblock and assign to
nr_all_pages and nr_kernel_pages. The counting excludes memmap_pages,
and other kernel used data, which is more accurate than old way and
simpler, and can also cover the ppc required arch_reserved_kernel_pages()
case.
And also clean up the outdated code comment above free_area_init_core().
And free_area_init_core() is easy to understand now, no need to add
words to explain.
Signed-off-by: Baoquan He <redacted>
---
mm/mm_init.c | 46 +++++-----------------------------------------
1 file changed, 5 insertions(+), 41 deletions(-)
@@ -1584,41 +1575,13 @@ static void __init free_area_init_core(struct pglist_data *pgdat)for(j=0;j<MAX_NR_ZONES;j++){structzone*zone=pgdat->node_zones+j;-unsignedlongsize,freesize,memmap_pages;--size=zone->spanned_pages;-freesize=zone->present_pages;--/*-*Adjustfreesizesothatitaccountsforhowmuchmemory-*isusedbythiszoneformemmap.Thisaffectsthewatermark-*andper-cpuinitialisations-*/-memmap_pages=calc_memmap_size(size,freesize);-if(!is_highmem_idx(j)){-if(freesize>=memmap_pages){-freesize-=memmap_pages;-if(memmap_pages)-pr_debug(" %s zone: %lu pages used for memmap\n",-zone_names[j],memmap_pages);-}else-pr_warn(" %s zone: %lu memmap pages exceeds freesize %lu\n",-zone_names[j],memmap_pages,freesize);-}--if(!is_highmem_idx(j))-nr_kernel_pages+=freesize;-/* Charge for highmem memmap if there are enough kernel pages */-elseif(nr_kernel_pages>memmap_pages*2)-nr_kernel_pages-=memmap_pages;-nr_all_pages+=freesize;+unsignedlongsize=zone->spanned_pages;/*-*Setanapproximatevalueforlowmemhere,itwillbeadjusted-*whenthebootmemallocatorfreespagesintothebuddysystem.-*Andallhighmempageswillbemanagedbythebuddysystem.+*Initializezone->managed_pagesas0,itwillbereset+*whenmemblockallocatorfreespagesintobuddysystem.*/-zone_init_internals(zone,j,nid,freesize);+zone_init_internals(zone,j,nid,0);if(!size)continue;
@@ -1915,6 +1878,7 @@ void __init free_area_init(unsigned long *max_zone_pfn)check_for_memory(pgdat);}+calc_nr_kernel_pages();memmap_init();/* disable hash distribution for systems with a single node */
From: Baoquan He <hidden> Date: 2024-03-25 14:57:46
Since the current calculation of calc_nr_kernel_pages() has taken into
consideration of kernel reserved memory, no need to have
arch_reserved_kernel_pages() any more.
Signed-off-by: Baoquan He <redacted>
---
arch/powerpc/include/asm/mmu.h | 4 ----
arch/powerpc/kernel/fadump.c | 5 -----
include/linux/mm.h | 3 ---
mm/mm_init.c | 12 ------------
4 files changed, 24 deletions(-)
From: Mike Rapoport <rppt@kernel.org> Date: 2024-03-26 06:44:53
On Mon, Mar 25, 2024 at 10:56:41PM +0800, Baoquan He wrote:
Variable dma_reserve and its usage was introduced in commit 0e0b864e069c
("[PATCH] Account for memmap and optionally the kernel image as holes").
Its original purpose was to accounting for the reserved pages in DMA
zone to make DMA zone's watermarks calculation more accurate on x86.
However, currently there's zone->managed_pages to account for all
available pages for buddy, zone->present_pages to account for all
present physical pages in zone. What is more important, on x86,
calculating and setting the zone->managed_pages is a temporary move,
all zone's managed_pages will be zeroed out and reset to the actual
value according to how many pages are added to buddy allocator in
mem_init(). Before mem_init(), no buddy alloction is requested. And
zone's pcp and watermark setting are all done after mem_init(). So,
no need to worry about the DMA zone's setting accuracy during
free_area_init().
Hence, remove memblock_find_dma_reserve() to stop calculating and
setting dma_reserve.
Signed-off-by: Baoquan He <redacted>
Reviewed-by: Mike Rapoport (IBM) <rppt@kernel.org>
From: Mike Rapoport <rppt@kernel.org> Date: 2024-03-26 06:45:33
On Mon, Mar 25, 2024 at 10:56:42PM +0800, Baoquan He wrote:
Now nobody calls set_dma_reserve() to set value for dma_reserve, remove
set_dma_reserve(), global variable dma_reserve and the codes using it.
Signed-off-by: Baoquan He <redacted>
Reviewed-by: Mike Rapoport (IBM) <rppt@kernel.org>
@@ -3210,7 +3210,6 @@ static inline int early_pfn_to_nid(unsigned long pfn)externint__meminitearly_pfn_to_nid(unsignedlongpfn);#endif-externvoidset_dma_reserve(unsignedlongnew_dma_reserve);externvoidmem_init(void);externvoid__initmmap_init(void);
From: Mike Rapoport <rppt@kernel.org> Date: 2024-03-26 06:58:10
Hi Baoquan,
On Mon, Mar 25, 2024 at 10:56:43PM +0800, Baoquan He wrote:
This is a preparation to calculate nr_kernel_pages and nr_all_pages,
both of which will be used later in alloc_large_system_hash().
nr_all_pages counts up all free but not reserved memory in memblock
allocator, including HIGHMEM memory. While nr_kernel_pages counts up
all free but not reserved low memory in memblock allocator, excluding
HIGHMEM memory.
Sorry I've missed this in the previous review, but I think this patch and
the patch "remove unneeded calc_memmap_size()" can be merged into "remove
meaningless calculation of zone->managed_pages in free_area_init_core()"
with an appropriate update of the commit message.
With the current patch splitting there will be compilation warning about unused
function for this and the next patch.
From: Mike Rapoport <rppt@kernel.org> Date: 2024-03-26 06:58:36
On Mon, Mar 25, 2024 at 10:56:46PM +0800, Baoquan He wrote:
Since the current calculation of calc_nr_kernel_pages() has taken into
consideration of kernel reserved memory, no need to have
arch_reserved_kernel_pages() any more.
Signed-off-by: Baoquan He <redacted>
Reviewed-by: Mike Rapoport (IBM) <rppt@kernel.org>
From: Baoquan He <hidden> Date: 2024-03-26 13:49:22
On 03/26/24 at 08:57am, Mike Rapoport wrote:
Hi Baoquan,
On Mon, Mar 25, 2024 at 10:56:43PM +0800, Baoquan He wrote:
quoted
This is a preparation to calculate nr_kernel_pages and nr_all_pages,
both of which will be used later in alloc_large_system_hash().
nr_all_pages counts up all free but not reserved memory in memblock
allocator, including HIGHMEM memory. While nr_kernel_pages counts up
all free but not reserved low memory in memblock allocator, excluding
HIGHMEM memory.
Sorry I've missed this in the previous review, but I think this patch and
the patch "remove unneeded calc_memmap_size()" can be merged into "remove
meaningless calculation of zone->managed_pages in free_area_init_core()"
with an appropriate update of the commit message.
With the current patch splitting there will be compilation warning about unused
function for this and the next patch.
Thanks for careful checking.
We need to make patch bisect-able to not break compiling so that people can
spot the cirminal commit, that's for sure. Do we need care about the
compiling warning from intermediate patch in one series? Not sure about
it. I always suggest people to seperate out this kind of newly added
function to a standalone patch for better reviewing and later checking,
and I saw a lot of commits like this by searching with
'git log --oneline | grep helper'
From: Mike Rapoport <rppt@kernel.org> Date: 2024-03-27 15:40:55
On Mon, Mar 25, 2024 at 10:56:43PM +0800, Baoquan He wrote:
This is a preparation to calculate nr_kernel_pages and nr_all_pages,
both of which will be used later in alloc_large_system_hash().
nr_all_pages counts up all free but not reserved memory in memblock
allocator, including HIGHMEM memory. While nr_kernel_pages counts up
all free but not reserved low memory in memblock allocator, excluding
HIGHMEM memory.
Signed-off-by: Baoquan He <redacted>
Reviewed-by: Mike Rapoport (IBM) <rppt@kernel.org>
From: Mike Rapoport <rppt@kernel.org> Date: 2024-03-27 15:41:21
On Mon, Mar 25, 2024 at 10:56:44PM +0800, Baoquan He wrote:
Currently, in free_area_init_core(), when initialize zone's field, a
rough value is set to zone->managed_pages. That value is calculated by
(zone->present_pages - memmap_pages).
In the meantime, add the value to nr_all_pages and nr_kernel_pages which
represent all free pages of system (only low memory or including HIGHMEM
memory separately). Both of them are gonna be used in
alloc_large_system_hash().
However, the rough calculation and setting of zone->managed_pages is
meaningless because
a) memmap pages are allocated on units of node in sparse_init() or
alloc_node_mem_map(pgdat); The simple (zone->present_pages -
memmap_pages) is too rough to make sense for zone;
b) the set zone->managed_pages will be zeroed out and reset with
acutal value in mem_init() via memblock_free_all(). Before the
resetting, no buddy allocation request is issued.
Here, remove the meaningless and complicated calculation of
(zone->present_pages - memmap_pages), initialize zone->managed_pages as 0
which reflect its actual value because no any page is added into buddy
system right now. It will be reset in mem_init().
And also remove the assignment of nr_all_pages and nr_kernel_pages in
free_area_init_core(). Instead, call the newly added calc_nr_kernel_pages()
to count up all free but not reserved memory in memblock and assign to
nr_all_pages and nr_kernel_pages. The counting excludes memmap_pages,
and other kernel used data, which is more accurate than old way and
simpler, and can also cover the ppc required arch_reserved_kernel_pages()
case.
And also clean up the outdated code comment above free_area_init_core().
And free_area_init_core() is easy to understand now, no need to add
words to explain.
Signed-off-by: Baoquan He <redacted>
Reviewed-by: Mike Rapoport (IBM) <rppt@kernel.org>
From: Mike Rapoport <rppt@kernel.org> Date: 2024-03-27 15:41:45
On Mon, Mar 25, 2024 at 10:56:46PM +0800, Baoquan He wrote:
Since the current calculation of calc_nr_kernel_pages() has taken into
consideration of kernel reserved memory, no need to have
arch_reserved_kernel_pages() any more.
Signed-off-by: Baoquan He <redacted>
Reviewed-by: Mike Rapoport (IBM) <rppt@kernel.org>
From: Baoquan He <hidden> Date: 2024-03-28 08:32:50
On 03/25/24 at 10:56pm, Baoquan He wrote:
quoted hunk
Currently, in free_area_init_core(), when initialize zone's field, a
rough value is set to zone->managed_pages. That value is calculated by
(zone->present_pages - memmap_pages).
In the meantime, add the value to nr_all_pages and nr_kernel_pages which
represent all free pages of system (only low memory or including HIGHMEM
memory separately). Both of them are gonna be used in
alloc_large_system_hash().
However, the rough calculation and setting of zone->managed_pages is
meaningless because
a) memmap pages are allocated on units of node in sparse_init() or
alloc_node_mem_map(pgdat); The simple (zone->present_pages -
memmap_pages) is too rough to make sense for zone;
b) the set zone->managed_pages will be zeroed out and reset with
acutal value in mem_init() via memblock_free_all(). Before the
resetting, no buddy allocation request is issued.
Here, remove the meaningless and complicated calculation of
(zone->present_pages - memmap_pages), initialize zone->managed_pages as 0
which reflect its actual value because no any page is added into buddy
system right now. It will be reset in mem_init().
And also remove the assignment of nr_all_pages and nr_kernel_pages in
free_area_init_core(). Instead, call the newly added calc_nr_kernel_pages()
to count up all free but not reserved memory in memblock and assign to
nr_all_pages and nr_kernel_pages. The counting excludes memmap_pages,
and other kernel used data, which is more accurate than old way and
simpler, and can also cover the ppc required arch_reserved_kernel_pages()
case.
And also clean up the outdated code comment above free_area_init_core().
And free_area_init_core() is easy to understand now, no need to add
words to explain.
Signed-off-by: Baoquan He <redacted>
---
mm/mm_init.c | 46 +++++-----------------------------------------
1 file changed, 5 insertions(+), 41 deletions(-)
@@ -1584,41 +1575,13 @@ static void __init free_area_init_core(struct pglist_data *pgdat)for(j=0;j<MAX_NR_ZONES;j++){structzone*zone=pgdat->node_zones+j;-unsignedlongsize,freesize,memmap_pages;--size=zone->spanned_pages;-freesize=zone->present_pages;--/*-*Adjustfreesizesothatitaccountsforhowmuchmemory-*isusedbythiszoneformemmap.Thisaffectsthewatermark-*andper-cpuinitialisations-*/-memmap_pages=calc_memmap_size(size,freesize);-if(!is_highmem_idx(j)){-if(freesize>=memmap_pages){-freesize-=memmap_pages;-if(memmap_pages)-pr_debug(" %s zone: %lu pages used for memmap\n",-zone_names[j],memmap_pages);-}else-pr_warn(" %s zone: %lu memmap pages exceeds freesize %lu\n",-zone_names[j],memmap_pages,freesize);-}--if(!is_highmem_idx(j))-nr_kernel_pages+=freesize;-/* Charge for highmem memmap if there are enough kernel pages */-elseif(nr_kernel_pages>memmap_pages*2)-nr_kernel_pages-=memmap_pages;-nr_all_pages+=freesize;+unsignedlongsize=zone->spanned_pages;/*-*Setanapproximatevalueforlowmemhere,itwillbeadjusted-*whenthebootmemallocatorfreespagesintothebuddysystem.-*Andallhighmempageswillbemanagedbythebuddysystem.+*Initializezone->managed_pagesas0,itwillbereset+*whenmemblockallocatorfreespagesintobuddysystem.*/-zone_init_internals(zone,j,nid,freesize);+zone_init_internals(zone,j,nid,0);
Here, we should initialize zone->managed_pages as zone->present_pages
because later page_group_by_mobility_disabled need be set according to
zone->managed_pages. Otherwise page_group_by_mobility_disabled will be
set to 1 always. I will sent out v3.
From a17b0921b4bd00596330f61ee9ea4b82386a9fed Mon Sep 17 00:00:00 2001
From: Baoquan He <redacted>
Date: Thu, 28 Mar 2024 16:20:15 +0800
Subject: [PATCH] mm/mm_init.c: set zone's ->managed_pages as ->present_pages
for now
Content-type: text/plain
Because page_group_by_mobility_disabled need be set according to zone's
managed_pages later.
Signed-off-by: Baoquan He <redacted>
---
mm/mm_init.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Baoquan He <hidden> Date: 2024-03-28 09:12:28
Currently, in free_area_init_core(), when initialize zone's field, a
rough value is set to zone->managed_pages. That value is calculated by
(zone->present_pages - memmap_pages).
In the meantime, add the value to nr_all_pages and nr_kernel_pages which
represent all free pages of system (only low memory or including HIGHMEM
memory separately). Both of them are gonna be used in
alloc_large_system_hash().
However, the rough calculation and setting of zone->managed_pages is
meaningless because
a) memmap pages are allocated on units of node in sparse_init() or
alloc_node_mem_map(pgdat); The simple (zone->present_pages -
memmap_pages) is too rough to make sense for zone;
b) the set zone->managed_pages will be zeroed out and reset with
acutal value in mem_init() via memblock_free_all(). Before the
resetting, no buddy allocation request is issued.
Here, remove the meaningless and complicated calculation of
(zone->present_pages - memmap_pages), directly set zone->managed_pages
as zone->present_pages for now. It will be adjusted in mem_init().
And also remove the assignment of nr_all_pages and nr_kernel_pages in
free_area_init_core(). Instead, call the newly added calc_nr_kernel_pages()
to count up all free but not reserved memory in memblock and assign to
nr_all_pages and nr_kernel_pages. The counting excludes memmap_pages,
and other kernel used data, which is more accurate than old way and
simpler, and can also cover the ppc required arch_reserved_kernel_pages()
case.
And also clean up the outdated code comment above free_area_init_core().
And free_area_init_core() is easy to understand now, no need to add
words to explain.
Signed-off-by: Baoquan He <redacted>
---
v2->v3:
- Change to initialize zone->managed_pages as zone->present_pages for now
because later page_group_by_mobility_disabled need be set according to
zone->managed_pages. Otherwise it will cause setting
page_group_by_mobility_disabled to 1 always.
mm/mm_init.c | 46 +++++-----------------------------------------
1 file changed, 5 insertions(+), 41 deletions(-)
@@ -1584,41 +1575,13 @@ static void __init free_area_init_core(struct pglist_data *pgdat)for(j=0;j<MAX_NR_ZONES;j++){structzone*zone=pgdat->node_zones+j;-unsignedlongsize,freesize,memmap_pages;--size=zone->spanned_pages;-freesize=zone->present_pages;--/*-*Adjustfreesizesothatitaccountsforhowmuchmemory-*isusedbythiszoneformemmap.Thisaffectsthewatermark-*andper-cpuinitialisations-*/-memmap_pages=calc_memmap_size(size,freesize);-if(!is_highmem_idx(j)){-if(freesize>=memmap_pages){-freesize-=memmap_pages;-if(memmap_pages)-pr_debug(" %s zone: %lu pages used for memmap\n",-zone_names[j],memmap_pages);-}else-pr_warn(" %s zone: %lu memmap pages exceeds freesize %lu\n",-zone_names[j],memmap_pages,freesize);-}--if(!is_highmem_idx(j))-nr_kernel_pages+=freesize;-/* Charge for highmem memmap if there are enough kernel pages */-elseif(nr_kernel_pages>memmap_pages*2)-nr_kernel_pages-=memmap_pages;-nr_all_pages+=freesize;+unsignedlongsize=zone->spanned_pages;/*-*Setanapproximatevalueforlowmemhere,itwillbeadjusted-*whenthebootmemallocatorfreespagesintothebuddysystem.-*Andallhighmempageswillbemanagedbythebuddysystem.+*Initializezone->managed_pagesas0,itwillbereset+*whenmemblockallocatorfreespagesintobuddysystem.*/-zone_init_internals(zone,j,nid,freesize);+zone_init_internals(zone,j,nid,zone->present_pages);if(!size)continue;
@@ -1915,6 +1878,7 @@ void __init free_area_init(unsigned long *max_zone_pfn)check_for_memory(pgdat);}+calc_nr_kernel_pages();memmap_init();/* disable hash distribution for systems with a single node */
From: Mike Rapoport <rppt@kernel.org> Date: 2024-03-28 09:53:47
On Thu, Mar 28, 2024 at 04:32:38PM +0800, Baoquan He wrote:
On 03/25/24 at 10:56pm, Baoquan He wrote:
quoted
/*
- * Set an approximate value for lowmem here, it will be adjusted
- * when the bootmem allocator frees pages into the buddy system.
- * And all highmem pages will be managed by the buddy system.
+ * Initialize zone->managed_pages as 0 , it will be reset
+ * when memblock allocator frees pages into buddy system.
*/
- zone_init_internals(zone, j, nid, freesize);
+ zone_init_internals(zone, j, nid, 0);
Here, we should initialize zone->managed_pages as zone->present_pages
because later page_group_by_mobility_disabled need be set according to
zone->managed_pages. Otherwise page_group_by_mobility_disabled will be
set to 1 always. I will sent out v3.
With zone->managed_pages set to zone->present_pages we won't account for
the reserved memory for initialization of page_group_by_mobility_disabled.
As watermarks are still not initialized at the time build_all_zonelists()
is called, we may use nr_all_pages - nr_kernel_pages instead of
nr_free_zone_pages(), IMO.
quoted hunk
From a17b0921b4bd00596330f61ee9ea4b82386a9fed Mon Sep 17 00:00:00 2001
From: Baoquan He <redacted>
Date: Thu, 28 Mar 2024 16:20:15 +0800
Subject: [PATCH] mm/mm_init.c: set zone's ->managed_pages as ->present_pages
for now
Content-type: text/plain
Because page_group_by_mobility_disabled need be set according to zone's
managed_pages later.
Signed-off-by: Baoquan He <redacted>
---
mm/mm_init.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Baoquan He <hidden> Date: 2024-03-28 14:46:54
On 03/28/24 at 11:53am, Mike Rapoport wrote:
On Thu, Mar 28, 2024 at 04:32:38PM +0800, Baoquan He wrote:
quoted
On 03/25/24 at 10:56pm, Baoquan He wrote:
quoted
/*
- * Set an approximate value for lowmem here, it will be adjusted
- * when the bootmem allocator frees pages into the buddy system.
- * And all highmem pages will be managed by the buddy system.
+ * Initialize zone->managed_pages as 0 , it will be reset
+ * when memblock allocator frees pages into buddy system.
*/
- zone_init_internals(zone, j, nid, freesize);
+ zone_init_internals(zone, j, nid, 0);
Here, we should initialize zone->managed_pages as zone->present_pages
because later page_group_by_mobility_disabled need be set according to
zone->managed_pages. Otherwise page_group_by_mobility_disabled will be
set to 1 always. I will sent out v3.
With zone->managed_pages set to zone->present_pages we won't account for
the reserved memory for initialization of page_group_by_mobility_disabled.
The old zone->managed_pages didn't account for the reserved pages
either. It's calculated by (zone->present_pages - memmap_pages). memmap
pages only is only a very small portion, e.g on x86_64, 4K page size,
assuming size of struct page is 64, then it's 1/64 of system memory.
On arm64, 64K page size, it's 1/1024 of system memory.
And about the setting of page_group_by_mobility_disabled, the compared
value pageblock_nr_pages * MIGRATE_TYPES which is very small. On x86_64,
it's 4M*6=24M; on arm64 with 64K size and 128M*6=768M which should be
the biggest among ARCH-es.
if (vm_total_pages < (pageblock_nr_pages * MIGRATE_TYPES))
page_group_by_mobility_disabled = 1;
else
page_group_by_mobility_disabled = 0;
So page_group_by_mobility_disabled could be set to 1 only on system with
very little memory which is very rarely seen. And setting
zone->managed_pages as zone->present_pages is very close to its old
value: (zone->present_pages - memmap_pages). Here we don't need be very
accurate, just a rough value.
As watermarks are still not initialized at the time build_all_zonelists()
is called, we may use nr_all_pages - nr_kernel_pages instead of
nr_free_zone_pages(), IMO.
nr_all_pages should be fine if we take this way. nr_kernel_pages is a
misleading name, it's all low memory pages excluding kernel reserved
apges. nr_all_pages is all memory pages including highmema and exluding
kernel reserved pages.
Both is fine to me. The first one is easier, simply setting
zone->managed_pages as zone->present_pages. The 2nd way is a little more
accurate.
quoted
From a17b0921b4bd00596330f61ee9ea4b82386a9fed Mon Sep 17 00:00:00 2001
From: Baoquan He <redacted>
Date: Thu, 28 Mar 2024 16:20:15 +0800
Subject: [PATCH] mm/mm_init.c: set zone's ->managed_pages as ->present_pages
for now
Content-type: text/plain
Because page_group_by_mobility_disabled need be set according to zone's
managed_pages later.
Signed-off-by: Baoquan He <redacted>
---
mm/mm_init.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)