When the HPT size is explicitly passed on from the userspace, currently
the KVM_PPC_ALLOCATE_HTAB will try to allocate the requested size of HPT
from reserved CMA area and if that is not possible, the allocation just
fails. With the commit 572abd563befd56 ("KVM: PPC: Book3S HV: Don't fall
back to smaller HPT size in allocation ioctl"), it does not even try to
allocate the same order pages from the page allocator before failing for
good. Same order allocation should be attempted from the page allocator
as a fallback option when the CMA allocation attempt fails.
Signed-off-by: Anshuman Khandual <redacted>
---
- This change saves guests from failing to start after migration
arch/powerpc/kvm/book3s_64_mmu_hv.c | 8 ++++++++
1 file changed, 8 insertions(+)
On Mon, Sep 12, 2016 at 9:13 PM, Anshuman Khandual
[off-list ref] wrote:
quoted hunk
When the HPT size is explicitly passed on from the userspace, currently
the KVM_PPC_ALLOCATE_HTAB will try to allocate the requested size of HPT
from reserved CMA area and if that is not possible, the allocation just
fails. With the commit 572abd563befd56 ("KVM: PPC: Book3S HV: Don't fall
back to smaller HPT size in allocation ioctl"), it does not even try to
allocate the same order pages from the page allocator before failing for
good. Same order allocation should be attempted from the page allocator
as a fallback option when the CMA allocation attempt fails.
Signed-off-by: Anshuman Khandual <redacted>
---
- This change saves guests from failing to start after migration
arch/powerpc/kvm/book3s_64_mmu_hv.c | 8 ++++++++
1 file changed, 8 insertions(+)
When the HPT size is explicitly passed on from the userspace, currently
the KVM_PPC_ALLOCATE_HTAB will try to allocate the requested size of HPT
from reserved CMA area and if that is not possible, the allocation just
fails. With the commit 572abd563befd56 ("KVM: PPC: Book3S HV: Don't fall
back to smaller HPT size in allocation ioctl"), it does not even try to
allocate the same order pages from the page allocator before failing for
good. Same order allocation should be attempted from the page allocator
as a fallback option when the CMA allocation attempt fails.
IMO we should fix the reason for these CMA allocation failure. We are just
doing work around here.
quoted hunk
Signed-off-by: Anshuman Khandual <redacted>
---
- This change saves guests from failing to start after migration
arch/powerpc/kvm/book3s_64_mmu_hv.c | 8 ++++++++
1 file changed, 8 insertions(+)
@@ -68,16 +68,18 @@ long kvmppc_alloc_hpt(struct kvm *kvm, u32 *htab_orderp)memset((void*)hpt,0,(1ul<<order));kvm->arch.hpt_cma_alloc=1;}--/* Lastly try successively smaller sizes from the page allocator */-/* Only do this if userspace didn't specify a size via ioctl */-while(!hpt&&order>PPC_MIN_HPT_ORDER&&!htab_orderp){+/*+*Trysuccessivelysmallersizesfromthepageallocator.+*Ifasizewasspecifiedviaanioctl,wejusttrythat+*specificsize+*/+-while(!hpt&&order>PPC_MIN_HPT_ORDER){hpt=__get_free_pages(GFP_KERNEL|__GFP_ZERO|__GFP_REPEAT|__GFP_NOWARN,order-PAGE_SHIFT);-if(!hpt)---order;+if(htab_orderp)+break;+--order;}-if(!hpt)return-ENOMEM;
On Mon, Sep 12, 2016 at 9:13 PM, Anshuman Khandual
[off-list ref] wrote:
quoted
quoted
When the HPT size is explicitly passed on from the userspace, currently
the KVM_PPC_ALLOCATE_HTAB will try to allocate the requested size of HPT
from reserved CMA area and if that is not possible, the allocation just
fails. With the commit 572abd563befd56 ("KVM: PPC: Book3S HV: Don't fall
back to smaller HPT size in allocation ioctl"), it does not even try to
allocate the same order pages from the page allocator before failing for
good. Same order allocation should be attempted from the page allocator
as a fallback option when the CMA allocation attempt fails.
Signed-off-by: Anshuman Khandual <redacted>
---
- This change saves guests from failing to start after migration
arch/powerpc/kvm/book3s_64_mmu_hv.c | 8 ++++++++
1 file changed, 8 insertions(+)
@@ -78,6 +78,14 @@ long kvmppc_alloc_hpt(struct kvm *kvm, u32 *htab_orderp)--order;}+/*+*Fallbackincasetheuserspacehasprovidedasizeviaioctl.+*Tryallocatingthesameorderpagesfromthepageallocator.+*/+if(!hpt&&order>PPC_MIN_HPT_ORDER&&htab_orderp)+hpt=__get_free_pages(GFP_KERNEL|__GFP_ZERO|__GFP_REPEAT|+__GFP_NOWARN,order-PAGE_SHIFT);+
How often does this succeed? Please provide data. I presume this for
During continuous guest VM migration test from source host to destination host
this patch was able to prevent guest creation failure after migration on the
destination host which was failing after 2-3 days. We have not seen the failure
till now even after 3-4 days.
the case where guest pages are pinned?
Hmm, need to check that in the test setup. There was nothing running inside the
guests though. IIUC, HPT size of the guest is computed based on the max memory
the guest is ever going to have irrespective of the RAM usage before migration.
How does pinning effect the HPT size ?
On Mon, Sep 12, 2016 at 9:13 PM, Anshuman Khandual
[off-list ref] wrote:
quoted
quoted
When the HPT size is explicitly passed on from the userspace, currently
the KVM_PPC_ALLOCATE_HTAB will try to allocate the requested size of HPT
from reserved CMA area and if that is not possible, the allocation just
fails. With the commit 572abd563befd56 ("KVM: PPC: Book3S HV: Don't fall
back to smaller HPT size in allocation ioctl"), it does not even try to
allocate the same order pages from the page allocator before failing for
good. Same order allocation should be attempted from the page allocator
as a fallback option when the CMA allocation attempt fails.
Signed-off-by: Anshuman Khandual <redacted>
---
- This change saves guests from failing to start after migration
arch/powerpc/kvm/book3s_64_mmu_hv.c | 8 ++++++++
1 file changed, 8 insertions(+)
@@ -78,6 +78,14 @@ long kvmppc_alloc_hpt(struct kvm *kvm, u32 *htab_orderp)--order;}+/*+*Fallbackincasetheuserspacehasprovidedasizeviaioctl.+*Tryallocatingthesameorderpagesfromthepageallocator.+*/+if(!hpt&&order>PPC_MIN_HPT_ORDER&&htab_orderp)+hpt=__get_free_pages(GFP_KERNEL|__GFP_ZERO|__GFP_REPEAT|+__GFP_NOWARN,order-PAGE_SHIFT);+
How often does this succeed? Please provide data. I presume this for
During continuous guest VM migration test from source host to destination host
this patch was able to prevent guest creation failure after migration on the
destination host which was failing after 2-3 days. We have not seen the failure
till now even after 3-4 days.
OK.. the CMA failures need analysis. Are we just ignoring a CMA bug? IOW, why
would CMA allocation fail -- CMA size is too small to accommodate the required
number of allocations?
quoted
the case where guest pages are pinned?
Hmm, need to check that in the test setup. There was nothing running inside the
guests though. IIUC, HPT size of the guest is computed based on the max memory
the guest is ever going to have irrespective of the RAM usage before migration.
How does pinning effect the HPT size ?
If the pinned pages (from anywhere) belong to CMA, then CMA allocations would start failing
Balbir Singh
On Mon, Sep 12, 2016 at 9:13 PM, Anshuman Khandual
[off-list ref] wrote:
quoted
quoted
When the HPT size is explicitly passed on from the userspace, currently
the KVM_PPC_ALLOCATE_HTAB will try to allocate the requested size of HPT
from reserved CMA area and if that is not possible, the allocation just
fails. With the commit 572abd563befd56 ("KVM: PPC: Book3S HV: Don't fall
back to smaller HPT size in allocation ioctl"), it does not even try to
allocate the same order pages from the page allocator before failing for
good. Same order allocation should be attempted from the page allocator
as a fallback option when the CMA allocation attempt fails.
Signed-off-by: Anshuman Khandual <redacted>
---
- This change saves guests from failing to start after migration
arch/powerpc/kvm/book3s_64_mmu_hv.c | 8 ++++++++
1 file changed, 8 insertions(+)
@@ -78,6 +78,14 @@ long kvmppc_alloc_hpt(struct kvm *kvm, u32 *htab_orderp)--order;}+/*+*Fallbackincasetheuserspacehasprovidedasizeviaioctl.+*Tryallocatingthesameorderpagesfromthepageallocator.+*/+if(!hpt&&order>PPC_MIN_HPT_ORDER&&htab_orderp)+hpt=__get_free_pages(GFP_KERNEL|__GFP_ZERO|__GFP_REPEAT|+__GFP_NOWARN,order-PAGE_SHIFT);+
How often does this succeed? Please provide data. I presume this for
During continuous guest VM migration test from source host to destination host
this patch was able to prevent guest creation failure after migration on the
destination host which was failing after 2-3 days. We have not seen the failure
till now even after 3-4 days.
OK.. the CMA failures need analysis. Are we just ignoring a CMA bug? IOW, why
Sure, it does need analysis. But there will be situations where CMA
allocation request can fail, thats why we will need fallback option.
That the same reason why we have fall back options of attempting from
page allocator (in decreasing order every time) when the size is not
specified as part of the ioctl. Why the case should be any different
when the size is specified in the ioctl().
would CMA allocation fail -- CMA size is too small to accommodate the required
number of allocations?
The same size seems to be good enough for first couple of days and
then it fails. Probably some __GFP_MOVABLE allocation got pinned
later on.
quoted
quoted
the case where guest pages are pinned?
Hmm, need to check that in the test setup. There was nothing running inside the
guests though. IIUC, HPT size of the guest is computed based on the max memory
the guest is ever going to have irrespective of the RAM usage before migration.
How does pinning effect the HPT size ?
If the pinned pages (from anywhere) belong to CMA, then CMA allocations would start failing
Right and with the current design of CMA we can do nothing about it,
unless we make sure the pages allocated to satisfy guest real memory
do not come from CMA area at all.
When the HPT size is explicitly passed on from the userspace, currently
the KVM_PPC_ALLOCATE_HTAB will try to allocate the requested size of HPT
from reserved CMA area and if that is not possible, the allocation just
fails. With the commit 572abd563befd56 ("KVM: PPC: Book3S HV: Don't fall
back to smaller HPT size in allocation ioctl"), it does not even try to
allocate the same order pages from the page allocator before failing for
good. Same order allocation should be attempted from the page allocator
as a fallback option when the CMA allocation attempt fails.
IMO we should fix the reason for these CMA allocation failure. We are just
IMHO, irrespective of the ability of CMA to satisfy allocation requests,
a fall back option from page allocator should always be there when the
guest is failing to start due to unavailability of memory. In my previous
response in this thread, also pointed out how there need to be a parity
between what we do for cases when ioctl calls specify size or not from
having a fallback option.
quoted hunk
doing work around here.
quoted
Signed-off-by: Anshuman Khandual <redacted>
---
- This change saves guests from failing to start after migration
arch/powerpc/kvm/book3s_64_mmu_hv.c | 8 ++++++++
1 file changed, 8 insertions(+)
@@ -68,16 +68,18 @@ long kvmppc_alloc_hpt(struct kvm *kvm, u32 *htab_orderp)memset((void*)hpt,0,(1ul<<order));kvm->arch.hpt_cma_alloc=1;}--/* Lastly try successively smaller sizes from the page allocator */-/* Only do this if userspace didn't specify a size via ioctl */-while(!hpt&&order>PPC_MIN_HPT_ORDER&&!htab_orderp){+/*+*Trysuccessivelysmallersizesfromthepageallocator.+*Ifasizewasspecifiedviaanioctl,wejusttrythat+*specificsize+*/+-while(!hpt&&order>PPC_MIN_HPT_ORDER){hpt=__get_free_pages(GFP_KERNEL|__GFP_ZERO|__GFP_REPEAT|__GFP_NOWARN,order-PAGE_SHIFT);-if(!hpt)---order;+if(htab_orderp)+break;+--order;}-if(!hpt)return-ENOMEM;
Initially thought about this way but then decided not to change this
existing code block instead just one more. But anything is fine, I
can just change this next time around.
On Tue, Sep 13, 2016 at 3:49 PM, Anshuman Khandual
[off-list ref] wrote:
On 09/13/2016 10:04 AM, Balbir Singh wrote:
quoted
On 13/09/16 14:07, Anshuman Khandual wrote:
quoted
On 09/12/2016 05:03 PM, Balbir Singh wrote:
quoted
On Mon, Sep 12, 2016 at 9:13 PM, Anshuman Khandual
[off-list ref] wrote:
quoted
quoted
When the HPT size is explicitly passed on from the userspace, currently
the KVM_PPC_ALLOCATE_HTAB will try to allocate the requested size of HPT
from reserved CMA area and if that is not possible, the allocation just
fails. With the commit 572abd563befd56 ("KVM: PPC: Book3S HV: Don't fall
back to smaller HPT size in allocation ioctl"), it does not even try to
allocate the same order pages from the page allocator before failing for
good. Same order allocation should be attempted from the page allocator
as a fallback option when the CMA allocation attempt fails.
Signed-off-by: Anshuman Khandual <redacted>
---
- This change saves guests from failing to start after migration
arch/powerpc/kvm/book3s_64_mmu_hv.c | 8 ++++++++
1 file changed, 8 insertions(+)
@@ -78,6 +78,14 @@ long kvmppc_alloc_hpt(struct kvm *kvm, u32 *htab_orderp)--order;}+/*+*Fallbackincasetheuserspacehasprovidedasizeviaioctl.+*Tryallocatingthesameorderpagesfromthepageallocator.+*/+if(!hpt&&order>PPC_MIN_HPT_ORDER&&htab_orderp)+hpt=__get_free_pages(GFP_KERNEL|__GFP_ZERO|__GFP_REPEAT|+__GFP_NOWARN,order-PAGE_SHIFT);+
How often does this succeed? Please provide data. I presume this for
During continuous guest VM migration test from source host to destination host
this patch was able to prevent guest creation failure after migration on the
destination host which was failing after 2-3 days. We have not seen the failure
till now even after 3-4 days.
OK.. the CMA failures need analysis. Are we just ignoring a CMA bug? IOW, why
Sure, it does need analysis. But there will be situations where CMA
allocation request can fail, thats why we will need fallback option.
Please elaborate those situations. This patch needs more explanation
as to why we should fallback -- what are those short comings of CMA
allocation. Can anyone using CMA face them and have to design a fallback?
That the same reason why we have fall back options of attempting from
page allocator (in decreasing order every time) when the size is not
specified as part of the ioctl. Why the case should be any different
when the size is specified in the ioctl().
quoted
would CMA allocation fail -- CMA size is too small to accommodate the required
number of allocations?
The same size seems to be good enough for first couple of days and
then it fails. Probably some __GFP_MOVABLE allocation got pinned
later on.
Please analyze and let us know
quoted
quoted
quoted
the case where guest pages are pinned?
Hmm, need to check that in the test setup. There was nothing running inside the
guests though. IIUC, HPT size of the guest is computed based on the max memory
the guest is ever going to have irrespective of the RAM usage before migration.
How does pinning effect the HPT size ?
If the pinned pages (from anywhere) belong to CMA, then CMA allocations would start failing
Right and with the current design of CMA we can do nothing about it,
unless we make sure the pages allocated to satisfy guest real memory
do not come from CMA area at all.
I have patches to move non-THP pages out of CMA
Balbir
From: Michael Ellerman <mpe@ellerman.id.au> Date: 2016-09-13 23:57:50
Anshuman Khandual [off-list ref] writes:
When the HPT size is explicitly passed on from the userspace, currently
the KVM_PPC_ALLOCATE_HTAB will try to allocate the requested size of HPT
from reserved CMA area and if that is not possible, the allocation just
fails. With the commit 572abd563befd56 ("KVM: PPC: Book3S HV: Don't fall
back to smaller HPT size in allocation ioctl"), it does not even try to
allocate the same order pages from the page allocator before failing for
good. Same order allocation should be attempted from the page allocator
as a fallback option when the CMA allocation attempt fails.
It looks like if CMA is not configured we will just fail instantly.
So this does look like something we should fix.
But I think it is just a bug in commit 572abd563bef ("KVM: PPC: Book3S
HV: Don't fall back to smaller HPT size in allocation ioctl"), which did:
@@ -70,7 +70,8 @@ long kvmppc_alloc_hpt(struct kvm *kvm, u32 *htab_orderp)}/* Lastly try successively smaller sizes from the page allocator */-while(!hpt&&order>PPC_MIN_HPT_ORDER){+/* Only do this if userspace didn't specify a size via ioctl */+while(!hpt&&order>PPC_MIN_HPT_ORDER&&!htab_orderp){hpt=__get_free_pages(GFP_KERNEL|__GFP_ZERO|__GFP_REPEAT|__GFP_NOWARN,order-PAGE_SHIFT);if(!hpt)
Instead of guarding the loop entry with !htab_orderp, it should have
allowed the loop to enter, but prevented it from iterating if the
allocation fails and htab_orderp != 0.
cheers
From: Paul Mackerras <hidden> Date: 2016-09-14 00:35:47
On Wed, Sep 14, 2016 at 09:57:48AM +1000, Michael Ellerman wrote:
quoted hunk
Anshuman Khandual [off-list ref] writes:
quoted
When the HPT size is explicitly passed on from the userspace, currently
the KVM_PPC_ALLOCATE_HTAB will try to allocate the requested size of HPT
from reserved CMA area and if that is not possible, the allocation just
fails. With the commit 572abd563befd56 ("KVM: PPC: Book3S HV: Don't fall
back to smaller HPT size in allocation ioctl"), it does not even try to
allocate the same order pages from the page allocator before failing for
good. Same order allocation should be attempted from the page allocator
as a fallback option when the CMA allocation attempt fails.
It looks like if CMA is not configured we will just fail instantly.
So this does look like something we should fix.
But I think it is just a bug in commit 572abd563bef ("KVM: PPC: Book3S
HV: Don't fall back to smaller HPT size in allocation ioctl"), which did:
@@ -70,7 +70,8 @@ long kvmppc_alloc_hpt(struct kvm *kvm, u32 *htab_orderp)}/* Lastly try successively smaller sizes from the page allocator */-while(!hpt&&order>PPC_MIN_HPT_ORDER){+/* Only do this if userspace didn't specify a size via ioctl */+while(!hpt&&order>PPC_MIN_HPT_ORDER&&!htab_orderp){hpt=__get_free_pages(GFP_KERNEL|__GFP_ZERO|__GFP_REPEAT|__GFP_NOWARN,order-PAGE_SHIFT);if(!hpt)
Instead of guarding the loop entry with !htab_orderp, it should have
allowed the loop to enter, but prevented it from iterating if the
allocation fails and htab_orderp != 0.
When the HPT size is explicitly passed on from the userspace, currently
the KVM_PPC_ALLOCATE_HTAB will try to allocate the requested size of HPT
from reserved CMA area and if that is not possible, the allocation just
fails. With the commit 572abd563befd56 ("KVM: PPC: Book3S HV: Don't fall
back to smaller HPT size in allocation ioctl"), it does not even try to
allocate the same order pages from the page allocator before failing for
good. Same order allocation should be attempted from the page allocator
as a fallback option when the CMA allocation attempt fails.
It looks like if CMA is not configured we will just fail instantly.
Right and also we have this fallback registered any way. I wonder
why we are still debating about the need of a fallback mechanism
when we already have got one.
So this does look like something we should fix.
But I think it is just a bug in commit 572abd563bef ("KVM: PPC: Book3S
HV: Don't fall back to smaller HPT size in allocation ioctl"), which did:
Hmm, I think its something the commit missed to accommodate for.
But maybe yes, its a bug in the commit.
@@ -70,7 +70,8 @@ long kvmppc_alloc_hpt(struct kvm *kvm, u32 *htab_orderp)}/* Lastly try successively smaller sizes from the page allocator */-while(!hpt&&order>PPC_MIN_HPT_ORDER){+/* Only do this if userspace didn't specify a size via ioctl */+while(!hpt&&order>PPC_MIN_HPT_ORDER&&!htab_orderp){hpt=__get_free_pages(GFP_KERNEL|__GFP_ZERO|__GFP_REPEAT|__GFP_NOWARN,order-PAGE_SHIFT);if(!hpt)
Instead of guarding the loop entry with !htab_orderp, it should have
allowed the loop to enter, but prevented it from iterating if the
allocation fails and htab_orderp != 0.
Right and thats what Aneesh's proposed patch (in the other thread) does.
On Wed, Sep 14, 2016 at 09:57:48AM +1000, Michael Ellerman wrote:
quoted
quoted
Anshuman Khandual [off-list ref] writes:
quoted
quoted
When the HPT size is explicitly passed on from the userspace, currently
the KVM_PPC_ALLOCATE_HTAB will try to allocate the requested size of HPT
from reserved CMA area and if that is not possible, the allocation just
fails. With the commit 572abd563befd56 ("KVM: PPC: Book3S HV: Don't fall
back to smaller HPT size in allocation ioctl"), it does not even try to
allocate the same order pages from the page allocator before failing for
good. Same order allocation should be attempted from the page allocator
as a fallback option when the CMA allocation attempt fails.
It looks like if CMA is not configured we will just fail instantly.
So this does look like something we should fix.
But I think it is just a bug in commit 572abd563bef ("KVM: PPC: Book3S
HV: Don't fall back to smaller HPT size in allocation ioctl"), which did:
@@ -70,7 +70,8 @@ long kvmppc_alloc_hpt(struct kvm *kvm, u32 *htab_orderp)}/* Lastly try successively smaller sizes from the page allocator */-while(!hpt&&order>PPC_MIN_HPT_ORDER){+/* Only do this if userspace didn't specify a size via ioctl */+while(!hpt&&order>PPC_MIN_HPT_ORDER&&!htab_orderp){hpt=__get_free_pages(GFP_KERNEL|__GFP_ZERO|__GFP_REPEAT|__GFP_NOWARN,order-PAGE_SHIFT);if(!hpt)
Instead of guarding the loop entry with !htab_orderp, it should have
allowed the loop to enter, but prevented it from iterating if the
allocation fails and htab_orderp != 0.
You're right. I'll fix it.
Thanks Paul, so I will not be sending follow up patch on this.
Not sure whether this mail ever went. Sending it again.
quoted
quoted
When the HPT size is explicitly passed on from the userspace, currently
the KVM_PPC_ALLOCATE_HTAB will try to allocate the requested size of HPT
from reserved CMA area and if that is not possible, the allocation just
fails. With the commit 572abd563befd56 ("KVM: PPC: Book3S HV: Don't fall
back to smaller HPT size in allocation ioctl"), it does not even try to
allocate the same order pages from the page allocator before failing for
good. Same order allocation should be attempted from the page allocator
as a fallback option when the CMA allocation attempt fails.
It looks like if CMA is not configured we will just fail instantly.
Right and also we have this fallback registered any way. I wonder
why we are still debating about the need of a fallback mechanism
when we already have got one.
So this does look like something we should fix.
But I think it is just a bug in commit 572abd563bef ("KVM: PPC: Book3S
HV: Don't fall back to smaller HPT size in allocation ioctl"), which did:
Hmm, I think its something the commit missed to accommodate for.
But maybe yes, its a bug in the commit.
@@ -70,7 +70,8 @@ long kvmppc_alloc_hpt(struct kvm *kvm, u32 *htab_orderp)}/* Lastly try successively smaller sizes from the page allocator */-while(!hpt&&order>PPC_MIN_HPT_ORDER){+/* Only do this if userspace didn't specify a size via ioctl */+while(!hpt&&order>PPC_MIN_HPT_ORDER&&!htab_orderp){hpt=__get_free_pages(GFP_KERNEL|__GFP_ZERO|__GFP_REPEAT|__GFP_NOWARN,order-PAGE_SHIFT);if(!hpt)
Instead of guarding the loop entry with !htab_orderp, it should have
allowed the loop to enter, but prevented it from iterating if the
allocation fails and htab_orderp != 0.
Right and thats what Aneesh's proposed patch (in the other thread) does.