Re: [v3 0/9] parallelized "struct page" zeroing

[v3 0/9] parallelized "struct page" zeroing · Pavel Tatashin <hidden> · 2017-05-05
[v3 6/9] sparc64: teach sparc not to zero struct pages memory · Pavel Tatashin <hidden> · 2017-05-05
[v3 7/9] x86: teach x86 not to zero struct pages memory · Pavel Tatashin <hidden> · 2017-05-05
[v3 1/9] sparc64: simplify vmemmap_populate · Pavel Tatashin <hidden> · 2017-05-05
[v3 4/9] mm: do not zero vmemmap_buf · Pavel Tatashin <hidden> · 2017-05-05
[v3 5/9] mm: zero struct pages during initialization · Pavel Tatashin <hidden> · 2017-05-05
[v3 2/9] mm: defining memblock_virt_alloc_try_nid_raw · Pavel Tatashin <hidden> · 2017-05-05
[v3 8/9] powerpc: teach platforms not to zero struct pages memory · Pavel Tatashin <hidden> · 2017-05-05
[v3 9/9] s390: teach platforms not to zero struct pages memory · Pavel Tatashin <hidden> · 2017-05-05
Re: [v3 9/9] s390: teach platforms not to zero struct pages memory · Heiko Carstens <hidden> · 2017-05-08
Re: [v3 9/9] s390: teach platforms not to zero struct pages memory · Pasha Tatashin <hidden> · 2017-05-15
Re: [v3 9/9] s390: teach platforms not to zero struct pages memory · Heiko Carstens <hidden> · 2017-05-15
Re: [v3 9/9] s390: teach platforms not to zero struct pages memory · Pasha Tatashin <hidden> · 2017-05-16
[v3 3/9] mm: add "zero" argument to vmemmap allocators · Pavel Tatashin <hidden> · 2017-05-05
Re: [v3 3/9] mm: add "zero" argument to vmemmap allocators · kbuild test robot <hidden> · 2017-05-13
Re: [v3 0/9] parallelized "struct page" zeroing · Michal Hocko <mhocko@kernel.org> · 2017-05-09
Re: [v3 0/9] parallelized "struct page" zeroing · Pasha Tatashin <hidden> · 2017-05-09
Re: [v3 0/9] parallelized "struct page" zeroing · Michal Hocko <mhocko@kernel.org> · 2017-05-10
Re: [v3 0/9] parallelized "struct page" zeroing · Pasha Tatashin <hidden> · 2017-05-10
Re: [v3 0/9] parallelized "struct page" zeroing · Michal Hocko <mhocko@kernel.org> · 2017-05-10
Re: [v3 0/9] parallelized "struct page" zeroing · Pasha Tatashin <hidden> · 2017-05-10
Re: [v3 0/9] parallelized "struct page" zeroing · David Miller <davem@davemloft.net> · 2017-05-10
Re: [v3 0/9] parallelized "struct page" zeroing · Pasha Tatashin <hidden> · 2017-05-11
Re: [v3 0/9] parallelized "struct page" zeroing · Pasha Tatashin <hidden> · 2017-05-11
Re: [v3 0/9] parallelized "struct page" zeroing · David Miller <davem@davemloft.net> · 2017-05-12
Re: [v3 0/9] parallelized "struct page" zeroing · Pasha Tatashin <hidden> · 2017-05-12
Re: [v3 0/9] parallelized "struct page" zeroing · David Miller <davem@davemloft.net> · 2017-05-12
Re: [v3 0/9] parallelized "struct page" zeroing · Benjamin Herrenschmidt <hidden> · 2017-05-16
Re: [v3 0/9] parallelized "struct page" zeroing · David Miller <davem@davemloft.net> · 2017-05-12
Re: [v3 0/9] parallelized "struct page" zeroing · David Miller <davem@davemloft.net> · 2017-05-10
Re: [v3 0/9] parallelized "struct page" zeroing · Matthew Wilcox <willy@infradead.org> · 2017-05-10
Re: [v3 0/9] parallelized "struct page" zeroing · David Miller <davem@davemloft.net> · 2017-05-10
Re: [v3 0/9] parallelized "struct page" zeroing · Matthew Wilcox <willy@infradead.org> · 2017-05-10
Re: [v3 0/9] parallelized "struct page" zeroing · Michal Hocko <mhocko@kernel.org> · 2017-05-11
Re: [v3 0/9] parallelized "struct page" zeroing · David Miller <davem@davemloft.net> · 2017-05-11
Re: [v3 0/9] parallelized "struct page" zeroing · Pasha Tatashin <hidden> · 2017-05-15
Re: [v3 0/9] parallelized "struct page" zeroing · Michal Hocko <mhocko@kernel.org> · 2017-05-15
Re: [v3 0/9] parallelized "struct page" zeroing · Pasha Tatashin <hidden> · 2017-05-15
Re: [v3 0/9] parallelized "struct page" zeroing · Michal Hocko <mhocko@kernel.org> · 2017-05-16
Re: [v3 0/9] parallelized "struct page" zeroing · Pasha Tatashin <hidden> · 2017-05-26
Re: [v3 0/9] parallelized "struct page" zeroing · Michal Hocko <mhocko@kernel.org> · 2017-05-29
Re: [v3 0/9] parallelized "struct page" zeroing · Pasha Tatashin <hidden> · 2017-05-30
Re: [v3 0/9] parallelized "struct page" zeroing · Michal Hocko <mhocko@kernel.org> · 2017-05-31
Re: [v3 0/9] parallelized "struct page" zeroing · David Miller <davem@davemloft.net> · 2017-05-31
Re: [v3 0/9] parallelized "struct page" zeroing · Pasha Tatashin <hidden> · 2017-06-01
Re: [v3 0/9] parallelized "struct page" zeroing · Michal Hocko <mhocko@kernel.org> · 2017-06-01

From: Pasha Tatashin <hidden>
Date: 2017-05-10 15:01:59
Also in: linux-mm, linux-s390, lkml, sparclinux


On 05/10/2017 10:57 AM, Michal Hocko wrote:

On Wed 10-05-17 09:42:22, Pasha Tatashin wrote:

quoted

Well, I didn't object to this particular part. I was mostly concerned
about
http://lkml.kernel.org/r/1494003796-748672-4-git-send-email-pasha.tatashin@oracle.com
and the "zero" argument for other functions. I guess we can do without
that. I _think_ that we should simply _always_ initialize the page at the
__init_single_page time rather than during the allocation. That would
require dropping __GFP_ZERO for non-memblock allocations. Or do you
think we could regress for single threaded initialization?

Hi Michal,

Thats exactly right, I am worried that we will regress when there is no
parallelized initialization of "struct pages" if we force unconditionally do
memset() in __init_single_page(). The overhead of calling memset() on a
smaller chunks (64-bytes) may cause the regression, this is why I opted only
for parallelized case to zero this metadata. This way, we are guaranteed to
see great improvements from this change without having regressions on
platforms and builds that do not support parallelized initialization of
"struct pages".

Have you measured that? I do not think it would be super hard to
measure. I would be quite surprised if this added much if anything at
all as the whole struct page should be in the cache line already. We do
set reference count and other struct members. Almost nobody should be
looking at our page at this time and stealing the cache line. On the
other hand a large memcpy will basically wipe everything away from the
cpu cache. Or am I missing something?

Perhaps you are right, and I will measure on x86. But, I suspect hit can 
become unacceptable on some platfoms: there is an overhead of calling a 
function, even if it is leaf-optimized, and there is an overhead in 
memset() to check for alignments of size and address, types of setting 
(zeroing vs. non-zeroing), etc., that adds up quickly.

Pasha

`h`	back out one level
`j`	next message in thread
`k`	previous message in thread
`l`	drill in
`Esc`	close help / fold thread tree
`?`	toggle this help