(1) Background
For the arm64, the hugetlb page size can be 32M (PMD + Contiguous bit).
In the 4K page environment, the max page order is 10 (max_order - 1),
so 32M page is the gigantic page.
The arm64 MMU supports a Contiguous bit which is a hint that the TTE
is one of a set of contiguous entries which can be cached in a single
TLB entry. Please refer to the arm64v8 mannul :
DDI0487A_f_armv8_arm.pdf (in page D4-1811)
(2) The bug
After I tested the libhugetlbfs, I found the test case "counter.sh"
will fail with the gigantic page (32M page in arm64 board).
The counter.sh is just a wrapper for counter.c.
You can find them in:
https://github.com/libhugetlbfs/libhugetlbfs/blob/master/tests/counters.chttps://github.com/libhugetlbfs/libhugetlbfs/blob/master/tests/counters.sh
The error log shows below:
----------------------------------------------------------
...........................................
LD_PRELOAD=libhugetlbfs.so shmoverride_unlinked (32M: 64): PASS
LD_PRELOAD=libhugetlbfs.so HUGETLB_SHM=yes shmoverride_unlinked (32M: 64): PASS
quota.sh (32M: 64): PASS
counters.sh (32M: 64): FAIL mmap failed: Invalid argument
********** TEST SUMMARY
* 32M
* 32-bit 64-bit
* Total testcases: 0 87
* Skipped: 0 0
* PASS: 0 86
* FAIL: 0 1
* Killed by signal: 0 0
* Bad configuration: 0 0
* Expected FAIL: 0 0
* Unexpected PASS: 0 0
* Strange test result: 0 0
**********
----------------------------------------------------------
The failure is caused by:
1) kernel fails to allocate a gigantic page for the surplus case.
And the gather_surplus_pages() will return NULL in the end.
2) The condition checks for some functions are wrong:
return_unused_surplus_pages()
nr_overcommit_hugepages_store()
hugetlb_overcommit_handler()
This patch set adds support for gigantic surplus hugetlb pages,
allowing the counter.sh unit test to pass.
Test this patch set with Juno-r1 board.
v2 -- > v3:
1.) In patch 2, change argument "no_init" to "do_prep"
2.) In patch 3, also change alloc_fresh_huge_page().
In the v2, this patch only changes the alloc_fresh_gigantic_page().
3.) Merge old patch #4,#5 into the last one.
4.) Follow Babka's suggestion, do the NULL check for @mask.
5.) others.
v1 -- > v2:
1.) fix the compiler error in X86.
2.) add new patches for NUMA.
The patch #2 ~ #5 are new patches.
Huang Shijie (4):
mm: hugetlb: rename some allocation functions
mm: hugetlb: add a new parameter for some functions
mm: hugetlb: change the return type for some functions
mm: hugetlb: support gigantic surplus pages
include/linux/mempolicy.h | 8 +++
mm/hugetlb.c | 146 +++++++++++++++++++++++++++++++++++-----------
mm/mempolicy.c | 44 ++++++++++++++
3 files changed, 163 insertions(+), 35 deletions(-)
--
2.5.5
After a future patch, the __alloc_buddy_huge_page() will not necessarily
use the buddy allocator.
So this patch removes the "buddy" from these functions:
__alloc_buddy_huge_page -> __alloc_huge_page
__alloc_buddy_huge_page_no_mpol -> __alloc_huge_page_no_mpol
__alloc_buddy_huge_page_with_mpol -> __alloc_huge_page_with_mpol
This patch also adds the description for alloc_gigantic_page().
This patch makes preparation for the later patch.
Signed-off-by: Huang Shijie <redacted>
---
mm/hugetlb.c | 30 ++++++++++++++++++++----------
1 file changed, 20 insertions(+), 10 deletions(-)
This patch adds a new parameter, the "do_prep", for these functions:
alloc_fresh_gigantic_page_node()
alloc_fresh_gigantic_page()
The prep_new_huge_page() does some initialization for the new page.
But sometime, we do not need it to do so, such as in the surplus case
in later patch.
With this parameter, the prep_new_huge_page() can be called by needed:
If the "do_prep" is true, calls the prep_new_huge_page() in
the alloc_fresh_gigantic_page_node();
This patch makes preparation for the later patches.
Signed-off-by: Huang Shijie <redacted>
---
mm/hugetlb.c | 14 ++++++++------
1 file changed, 8 insertions(+), 6 deletions(-)
This patch changes the return type to "struct page*" for
alloc_fresh_gigantic_page()/alloc_fresh_huge_page().
This patch makes preparation for later patch.
Signed-off-by: Huang Shijie <redacted>
---
mm/hugetlb.c | 29 ++++++++++++++---------------
1 file changed, 14 insertions(+), 15 deletions(-)
When testing the gigantic page whose order is too large for the buddy
allocator, the libhugetlbfs test case "counter.sh" will fail.
The counter.sh is just a wrapper for counter.c, you can find them in:
https://github.com/libhugetlbfs/libhugetlbfs/blob/master/tests/counters.chttps://github.com/libhugetlbfs/libhugetlbfs/blob/master/tests/counters.sh
Please see the error log below:
............................................
........
quota.sh (32M: 64): PASS
counters.sh (32M: 64): FAIL mmap failed: Invalid argument
********** TEST SUMMARY
* 32M
* 32-bit 64-bit
* Total testcases: 0 87
* Skipped: 0 0
* PASS: 0 86
* FAIL: 0 1
* Killed by signal: 0 0
* Bad configuration: 0 0
* Expected FAIL: 0 0
* Unexpected PASS: 0 0
* Strange test result: 0 0
**********
............................................
The failure is caused by:
1) kernel fails to allocate a gigantic page for the surplus case.
And the gather_surplus_pages() will return NULL in the end.
2) The condition checks for "over-commit" is wrong.
This patch does following things:
1) This patch changes the condition checks for:
return_unused_surplus_pages()
nr_overcommit_hugepages_store()
hugetlb_overcommit_handler()
2) This patch introduces two helper functions:
huge_nodemask() and __hugetlb_alloc_gigantic_page().
Please see the descritions in the two functions.
3) This patch uses __hugetlb_alloc_gigantic_page() to allocate the
gigantic page in the __alloc_huge_page(). After this patch,
gather_surplus_pages() can return a gigantic page for the surplus case.
After this patch, the counter.sh can pass for the gigantic page.
Signed-off-by: Huang Shijie <redacted>
---
include/linux/mempolicy.h | 8 +++++
mm/hugetlb.c | 77 +++++++++++++++++++++++++++++++++++++++++++----
mm/mempolicy.c | 44 +++++++++++++++++++++++++++
3 files changed, 123 insertions(+), 6 deletions(-)
@@ -1506,6 +1506,69 @@ int dissolve_free_huge_pages(unsigned long start_pfn, unsigned long end_pfn)/**Thereare3waysthiscangetcalled:+*+*1.WhentheNUMAisnotenabled,usealloc_gigantic_page()toget+*thegiganticpage.+*+*2.TheNUMAisenabled,butthevmaisNULL.+*Initializethe@mask,andusealloc_fresh_gigantic_page()toget+*thegiganticpage.+*+*3.TheNUMAisenabled,andthevmaisvalid.+*Usethe@vma'smemorypolicy.+*Get@maskbyhuge_nodemask(),andusealloc_fresh_gigantic_page()+*togetthegiganticpage.+*/+staticstructpage*__hugetlb_alloc_gigantic_page(structhstate*h,+structvm_area_struct*vma,unsignedlongaddr,intnid)+{+NODEMASK_ALLOC(nodemask_t,mask,GFP_KERNEL|__GFP_NORETRY);+structpage*page=NULL;++/* Not NUMA */+if(!IS_ENABLED(CONFIG_NUMA)){+if(nid==NUMA_NO_NODE)+nid=numa_mem_id();++page=alloc_gigantic_page(nid,huge_page_order(h));+if(page)+prep_compound_gigantic_page(page,huge_page_order(h));+gotogot_page;+}++/* NUMA && !vma */+if(!vma){+/* First, check the mask */+if(!mask){+mask=&node_states[N_MEMORY];+}else{+if(nid==NUMA_NO_NODE){+if(!init_nodemask_of_mempolicy(mask)){+NODEMASK_FREE(mask);+mask=&node_states[N_MEMORY];+}+}else{+init_nodemask_of_node(mask,nid);+}+}++page=alloc_fresh_gigantic_page(h,mask,false);+gotogot_page;+}++/* NUMA && vma */+if(mask&&huge_nodemask(vma,addr,mask))+page=alloc_fresh_gigantic_page(h,mask,false);++got_page:+if(mask!=&node_states[N_MEMORY])+NODEMASK_FREE(mask);++returnpage;+}++/*+*Thereare3waysthiscangetcalled:*1.Withvma+addr:weusetheVMA'smemorypolicy*2.With!vma,butnid=NUMA_NO_NODE:Wetrytoallocateahuge*pagefromanynode,andletthebuddyallocatoritselffigure
@@ -2966,7 +3031,7 @@ int hugetlb_overcommit_handler(struct ctl_table *table, int write,tmp=h->nr_overcommit_huge_pages;-if(write&&hstate_is_gigantic(h))+if(write&&hstate_is_gigantic(h)&&!gigantic_page_supported())return-EINVAL;table->data=&tmp;
From: Michal Hocko <mhocko@suse.com> Date: 2016-12-05 09:31:10
On Mon 05-12-16 17:17:07, Huang Shijie wrote:
[...]
The failure is caused by:
1) kernel fails to allocate a gigantic page for the surplus case.
And the gather_surplus_pages() will return NULL in the end.
2) The condition checks for some functions are wrong:
return_unused_surplus_pages()
nr_overcommit_hugepages_store()
hugetlb_overcommit_handler()
OK, so how is this any different from gigantic (1G) hugetlb pages on
x86_64? Do we need the same functionality or is it just 32MB not being
handled in the same way as 1G?
Thanks!
--
Michal Hocko
SUSE Labs
On Mon, Dec 05, 2016 at 05:31:01PM +0800, Michal Hocko wrote:
On Mon 05-12-16 17:17:07, Huang Shijie wrote:
[...]
quoted
The failure is caused by:
1) kernel fails to allocate a gigantic page for the surplus case.
And the gather_surplus_pages() will return NULL in the end.
2) The condition checks for some functions are wrong:
return_unused_surplus_pages()
nr_overcommit_hugepages_store()
hugetlb_overcommit_handler()
OK, so how is this any different from gigantic (1G) hugetlb pages on
I think there is no different from gigantic (1G) hugetlb pages on
x86_64. Do anyone ever tested the 1G hugetlb pages in x86_64 with the "counter.sh"
before?
x86_64? Do we need the same functionality or is it just 32MB not being
handled in the same way as 1G?
Yes, we need this functionality for gigantic pages, no matter it is
X86_64 or S390 or arm64, no matter it is 32MB or 1G. :)
But anyway, I will try to find some machine and try the 1G gigantic page
on ARM64.
Thanks
Huang Shijie
On Mon, Dec 05, 2016 at 05:31:01PM +0800, Michal Hocko wrote:
On Mon 05-12-16 17:17:07, Huang Shijie wrote:
[...]
quoted
The failure is caused by:
1) kernel fails to allocate a gigantic page for the surplus case.
And the gather_surplus_pages() will return NULL in the end.
2) The condition checks for some functions are wrong:
return_unused_surplus_pages()
nr_overcommit_hugepages_store()
hugetlb_overcommit_handler()
OK, so how is this any different from gigantic (1G) hugetlb pages on
x86_64? Do we need the same functionality or is it just 32MB not being
handled in the same way as 1G?
I tested this patch set on the Softiron board(ARM64) which has 16G memory.
I appended "hugepagesz=1G hugepages=6" in the kernel cmdline, the arm64
will use the PUD_SIZE for the hugetlb page.
The 1G page size can run well, I post the log here:
--------------------------------------------------------
counters.sh (1024M: 64): PASS
********** TEST SUMMARY
* 1024M
* 32-bit 64-bit
* Total testcases: 0 1
* Skipped: 0 0
* PASS: 0 1
* FAIL: 0 0
* Killed by signal: 0 0
* Bad configuration: 0 0
* Expected FAIL: 0 0
* Unexpected PASS: 0 0
* Strange test result: 0 0
**********
--------------------------------------------------------
My desktop is x86_64, but its memory is just 8G.
I will expand its memory capacity, and continue to
the test for x86_64.
Thanks
Huang Shijie
From: Michal Hocko <mhocko@suse.com> Date: 2016-12-07 15:02:43
On Tue 06-12-16 18:03:59, Huang Shijie wrote:
On Mon, Dec 05, 2016 at 05:31:01PM +0800, Michal Hocko wrote:
quoted
On Mon 05-12-16 17:17:07, Huang Shijie wrote:
[...]
quoted
The failure is caused by:
1) kernel fails to allocate a gigantic page for the surplus case.
And the gather_surplus_pages() will return NULL in the end.
2) The condition checks for some functions are wrong:
return_unused_surplus_pages()
nr_overcommit_hugepages_store()
hugetlb_overcommit_handler()
OK, so how is this any different from gigantic (1G) hugetlb pages on
I think there is no different from gigantic (1G) hugetlb pages on
x86_64. Do anyone ever tested the 1G hugetlb pages in x86_64 with the "counter.sh"
before?
I suspect nobody has because the gigantic page support is still somehow
coarse and from a quick look into the code we only support pre-allocated
giga pages. In other words surplus pages and their accounting is not
supported at all.
I haven't yet checked your patchset but I can tell you one thing.
Surplus and subpool pages code is tricky as hell. And it is not just a
matter of teaching the huge page allocation code to do the right thing.
There are subtle details all over the place. E.g. we currently
do not free giga pages AFAICS. In fact I believe that the giga pages are
kind of implanted to the existing code without any higher level
consistency. This should change long term. But I am worried it is much
more work.
Now I might be wrong because I might misremember things which might have
been changed recently but please make sure you describe the current
state and changes of giga pages when touching this area much better if
you want to pursue this route...
Thanks!
--
Michal Hocko
SUSE Labs
On Wed, Dec 07, 2016 at 11:02:38PM +0800, Michal Hocko wrote:
On Tue 06-12-16 18:03:59, Huang Shijie wrote:
quoted
On Mon, Dec 05, 2016 at 05:31:01PM +0800, Michal Hocko wrote:
quoted
On Mon 05-12-16 17:17:07, Huang Shijie wrote:
[...]
quoted
The failure is caused by:
1) kernel fails to allocate a gigantic page for the surplus case.
And the gather_surplus_pages() will return NULL in the end.
2) The condition checks for some functions are wrong:
return_unused_surplus_pages()
nr_overcommit_hugepages_store()
hugetlb_overcommit_handler()
add the > >
quoted
quoted
OK, so how is this any different from gigantic (1G) hugetlb pages on
I think there is no different from gigantic (1G) hugetlb pages on
x86_64. Do anyone ever tested the 1G hugetlb pages in x86_64 with the "counter.sh"
before?
I suspect nobody has because the gigantic page support is still somehow
coarse and from a quick look into the code we only support pre-allocated
Yes, the x86_64 even does not support the gigantic page.
The default x86_64_defconfig does not enable the CONFIG_CMA.
I enabled the CONFIG_CMA, and did the test for gigantic page in x86_64.
(I appended "hugepagesz=1G hugepages=4" in the kernel cmdline.)
The result is got with my 16G x86_64 desktop:
-------------------------------------------------
counters.sh (1024M: 32): FAIL mmap failed: Cannot allocate memory
counters.sh (1024M: 64): PASS
********** TEST SUMMARY
* 1024M
* 32-bit 64-bit
* Total testcases: 1 1
* Skipped: 0 0
* PASS: 0 1
* FAIL: 1 0
* Killed by signal: 0 0
* Bad configuration: 0 0
* Expected FAIL: 0 0
* Unexpected PASS: 0 0
* Test not present: 0 0
* Strange test result: 0 0
**********
-------------------------------------------------
The test passes for 64bit, but fails for 32bit (but I think it's okay,
since 1G hugetlb page is too large for the 32bit).
giga pages. In other words surplus pages and their accounting is not
supported at all.
Yes.
I haven't yet checked your patchset but I can tell you one thing.
Could you please review the patch set when you have time? Thanks a lot.
Surplus and subpool pages code is tricky as hell. And it is not just a
Agree.
Do we really need so many accountings? such as reserve/ovorcommit/surplus.
matter of teaching the huge page allocation code to do the right thing.
There are subtle details all over the place. E.g. we currently
do not free giga pages AFAICS. In fact I believe that the giga pages are
Please correct me if I am wrong. :)
I think the free-giga-pages can work well.
Please see the code in update_and_free_page().
Could you please list all the subtle details you think the code is wrong?
I can check them one by one.
kind of implanted to the existing code without any higher level
consistency. This should change long term. But I am worried it is much
What's type of the "higher level consistency" we should care about?
Thanks
Huang Shijie
more work.
Now I might be wrong because I might misremember things which might have
been changed recently but please make sure you describe the current
state and changes of giga pages when touching this area much better if
you want to pursue this route...
From: Michal Hocko <mhocko@suse.com> Date: 2016-12-08 09:52:59
On Thu 08-12-16 17:36:24, Huang Shijie wrote:
On Wed, Dec 07, 2016 at 11:02:38PM +0800, Michal Hocko wrote:
[...]
quoted
I haven't yet checked your patchset but I can tell you one thing.
Could you please review the patch set when you have time? Thanks a lot.
From a quick glance you do not handle the reservation code at all. You
just make sure that the allocation doesn't fail unconditionally. I might
be wrong here and Naoya resp. Mike will know much better but this seems
far from enough to me.
quoted
Surplus and subpool pages code is tricky as hell. And it is not just a
Agree.
Do we really need so many accountings? such as reserve/ovorcommit/surplus.
If we want to make giga page the first class citizen then the whole
reservation/surplus code has to independent on the page size.
quoted
matter of teaching the huge page allocation code to do the right thing.
There are subtle details all over the place. E.g. we currently
do not free giga pages AFAICS. In fact I believe that the giga pages are
Please correct me if I am wrong. :)
I think the free-giga-pages can work well.
Please see the code in update_and_free_page().
Hmm, I have missed that part. I guess you are right but I would have to
look much closer. Hugetlb code tends to be full of surprises.
Could you please list all the subtle details you think the code is wrong?
I can check them one by one.
Well, this would take me quite some time and basically restudy the whole
hugetlb code again. What you are trying to achieve is not a simple "fix
a test case" thing. You are trying to implement full featured giga pages
suport. And as I've said this requires a deeper understanding of the
current code and clean it up considerably wrt. giga pages. This is
definitely desirable plan longterm and I would like to encourage you for
that but it is not a simple project at the same time.
--
Michal Hocko
SUSE Labs
On Thu, Dec 08, 2016 at 10:52:54AM +0100, Michal Hocko wrote:
On Thu 08-12-16 17:36:24, Huang Shijie wrote:
quoted
On Wed, Dec 07, 2016 at 11:02:38PM +0800, Michal Hocko wrote:
[...]
quoted
quoted
I haven't yet checked your patchset but I can tell you one thing.
Could you please review the patch set when you have time? Thanks a lot.
From a quick glance you do not handle the reservation code at all. You
Thanks, I will study the code again, and try to find What we need to do
with the reservation code.
just make sure that the allocation doesn't fail unconditionally. I might
be wrong here and Naoya resp. Mike will know much better but this seems
far from enough to me.
Well, this would take me quite some time and basically restudy the whole
hugetlb code again. What you are trying to achieve is not a simple "fix
a test case" thing. You are trying to implement full featured giga pages
suport. And as I've said this requires a deeper understanding of the
current code and clean it up considerably wrt. giga pages. This is
definitely desirable plan longterm and I would like to encourage you for
that but it is not a simple project at the same time.
Okay, I will try to implement the full featured giga pages support. :)
But I feel confused at the "full featured". If the patch set can pass
all the giga pages tests in the libhugetlbfs, can we say it is "full
featured"? Or some one reviews this patch set, and say it is full
featured support for the giga pages.
Thanks
Huang Shijie