From: David Hildenbrand <hidden> Date: 2021-01-28 16:49:05
Let's count the number of CMA pages per zone and print them in
/proc/zoneinfo.
Having access to the total number of CMA pages per zone is helpful for
debugging purposes to know where exactly the CMA pages ended up, and to
figure out how many pages of a zone might behave differently, even after
some of these pages might already have been allocated.
As one example, CMA pages part of a kernel zone cannot be used for
ordinary kernel allocations but instead behave more like ZONE_MOVABLE.
For now, we are only able to get the global nr+free cma pages from
/proc/meminfo and the free cma pages per zone from /proc/zoneinfo.
Example after this patch when booting a 6 GiB QEMU VM with
"hugetlb_cma=2G":
# cat /proc/zoneinfo | grep cma
cma 0
nr_free_cma 0
cma 0
nr_free_cma 0
cma 524288
nr_free_cma 493016
cma 0
cma 0
# cat /proc/meminfo | grep Cma
CmaTotal: 2097152 kB
CmaFree: 1972064 kB
Note: We track/print only with CONFIG_CMA; "nr_free_cma" in /proc/zoneinfo
is currently also printed without CONFIG_CMA.
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Thomas Gleixner <redacted>
Cc: "Peter Zijlstra (Intel)" <peterz@infradead.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Oscar Salvador <osalvador@suse.de>
Cc: Michal Hocko <mhocko@kernel.org>
Cc: Wei Yang <redacted>
Cc: linux-api@vger.kernel.org
Signed-off-by: David Hildenbrand <redacted>
---
v1 -> v2:
- Print/track only with CONFIG_CMA
- Extend patch description
---
include/linux/mmzone.h | 6 ++++++
mm/page_alloc.c | 1 +
mm/vmstat.c | 5 +++++
3 files changed, 12 insertions(+)
From: Oscar Salvador <osalvador@suse.de> Date: 2021-01-28 21:44:17
On Thu, Jan 28, 2021 at 05:45:33PM +0100, David Hildenbrand wrote:
Let's count the number of CMA pages per zone and print them in
/proc/zoneinfo.
Having access to the total number of CMA pages per zone is helpful for
debugging purposes to know where exactly the CMA pages ended up, and to
figure out how many pages of a zone might behave differently, even after
some of these pages might already have been allocated.
As one example, CMA pages part of a kernel zone cannot be used for
ordinary kernel allocations but instead behave more like ZONE_MOVABLE.
For now, we are only able to get the global nr+free cma pages from
/proc/meminfo and the free cma pages per zone from /proc/zoneinfo.
Example after this patch when booting a 6 GiB QEMU VM with
"hugetlb_cma=2G":
# cat /proc/zoneinfo | grep cma
cma 0
nr_free_cma 0
cma 0
nr_free_cma 0
cma 524288
nr_free_cma 493016
cma 0
cma 0
# cat /proc/meminfo | grep Cma
CmaTotal: 2097152 kB
CmaFree: 1972064 kB
Note: We track/print only with CONFIG_CMA; "nr_free_cma" in /proc/zoneinfo
is currently also printed without CONFIG_CMA.
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Thomas Gleixner <redacted>
Cc: "Peter Zijlstra (Intel)" <peterz@infradead.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Oscar Salvador <osalvador@suse.de>
Cc: Michal Hocko <mhocko@kernel.org>
Cc: Wei Yang <redacted>
Cc: linux-api@vger.kernel.org
Signed-off-by: David Hildenbrand <redacted>
IMHO looks better to me, thanks:
Reviewed-by: Oscar Salvador <osalvador@suse.de>
Hmm, not sure about this. If cma is only printed for CONFIG_CMA, we can't
distinguish between (1) a kernel without your patch without including some
version checking and (2) a kernel without CONFIG_CMA enabled. IOW,
"cma 0" carries value: we know immediately that we do not have any CMA
pages on this zone, period.
/proc/zoneinfo is also not known for its conciseness so I think printing
"cma 0" even for !CONFIG_CMA is helpful :)
I think this #ifdef should be removed and it should call into a
zone_cma_pages(struct zone *zone) which returns 0UL if disabled.
Hmm, not sure about this. If cma is only printed for CONFIG_CMA, we can't
distinguish between (1) a kernel without your patch without including some
version checking and (2) a kernel without CONFIG_CMA enabled. IOW,
"cma 0" carries value: we know immediately that we do not have any CMA
pages on this zone, period.
/proc/zoneinfo is also not known for its conciseness so I think printing
"cma 0" even for !CONFIG_CMA is helpful :)
I think this #ifdef should be removed and it should call into a
zone_cma_pages(struct zone *zone) which returns 0UL if disabled.
Yeah, that’s also what I proposed in a sub-thread here.
The last option would be going the full mile and not printing nr_free_cma. Code might get a bit uglier though, but we could also remove that stats counter ;)
I don‘t particularly care, while printing „0“ might be easier, removing nr_free_cma might be cleaner.
But then, maybe there are tools that expect that value to be around on any kernel?
Thoughts?
Thanks
Hmm, not sure about this. If cma is only printed for CONFIG_CMA, we can't
distinguish between (1) a kernel without your patch without including some
version checking and (2) a kernel without CONFIG_CMA enabled. IOW,
"cma 0" carries value: we know immediately that we do not have any CMA
pages on this zone, period.
/proc/zoneinfo is also not known for its conciseness so I think printing
"cma 0" even for !CONFIG_CMA is helpful :)
I think this #ifdef should be removed and it should call into a
zone_cma_pages(struct zone *zone) which returns 0UL if disabled.
Yeah, that’s also what I proposed in a sub-thread here.
Ah, I certainly think your original intuition was correct.
The last option would be going the full mile and not printing nr_free_cma. Code might get a bit uglier though, but we could also remove that stats counter ;)
I don‘t particularly care, while printing „0“ might be easier, removing nr_free_cma might be cleaner.
But then, maybe there are tools that expect that value to be around on any kernel?
Yeah, that's probably undue risk, the ship has sailed and there's no
significant upside.
I still think "cma 0" in /proc/zoneinfo carries value, though, especially
for NUMA and it looks like this is how it's done in linux-next. With a
single read of the file, userspace can make the determination what CMA
pages exist on this node.
In general, I think the rule-of-thumb is that the fewer ifdefs in
/proc/zoneinfo, the easier it is for userspace to parse it.
(I made that change to /proc/zoneinfo to even print non-existant zones for
each node because otherwise you cannot determine what the indices of
things like vm.lowmem_reserve_ratio represent.)
Hmm, not sure about this. If cma is only printed for CONFIG_CMA, we can't
distinguish between (1) a kernel without your patch without including some
version checking and (2) a kernel without CONFIG_CMA enabled. IOW,
"cma 0" carries value: we know immediately that we do not have any CMA
pages on this zone, period.
/proc/zoneinfo is also not known for its conciseness so I think printing
"cma 0" even for !CONFIG_CMA is helpful :)
I think this #ifdef should be removed and it should call into a
zone_cma_pages(struct zone *zone) which returns 0UL if disabled.
Yeah, that’s also what I proposed in a sub-thread here.
Ah, I certainly think your original intuition was correct.
quoted
The last option would be going the full mile and not printing nr_free_cma. Code might get a bit uglier though, but we could also remove that stats counter ;)
I don‘t particularly care, while printing „0“ might be easier, removing nr_free_cma might be cleaner.
But then, maybe there are tools that expect that value to be around on any kernel?
Yeah, that's probably undue risk, the ship has sailed and there's no
significant upside.
I still think "cma 0" in /proc/zoneinfo carries value, though, especially
for NUMA and it looks like this is how it's done in linux-next. With a
single read of the file, userspace can make the determination what CMA
pages exist on this node.
In general, I think the rule-of-thumb is that the fewer ifdefs in
/proc/zoneinfo, the easier it is for userspace to parse it.
Makes sense, I‘ll send an updated version tomorrow - thanks!
(I made that change to /proc/zoneinfo to even print non-existant zones for
each node because otherwise you cannot determine what the indices of
things like vm.lowmem_reserve_ratio represent.)