These changes enable the dynamic creation of movable nodes on power.
On x86, the ACPI SRAT memory affinity structure can mark memory
hotpluggable, allowing the kernel to possibly create movable nodes at
boot.
While power has no analog of this SRAT information, we can still create
a movable memory node, post boot, by hotplugging all of the node's
memory into ZONE_MOVABLE.
We provide a way to describe the extents and numa associativity of such
a node in the device tree, while deferring the memory addition to take
place through hotplug.
In v1, this patchset introduced a new dt compatible id to explicitly
create a memoryless node at boot. Here, things have been simplified to
be applicable regardless of the status of node hotplug on power. We
still intend to enable hotadding a pgdat, but that's now untangled as a
separate topic.
v4:
* Rename of_fdt_is_available() to of_fdt_device_is_available().
Rename of_flat_dt_is_available() to of_flat_dt_device_is_available().
* Instead of restoring top-down allocation, ensure it never goes
bottom-up in the first place, by making movable_node arch-specific.
* Use MEMORY_HOTPLUG instead of PPC64 in the mm/Kconfig patch.
v3:
* http://lkml.kernel.org/r/1474828616-16608-1-git-send-email-arbab@linux.vnet.ibm.com
* Use Rob Herring's suggestions to improve the node availability check.
* More verbose commit log in the patch enabling CONFIG_MOVABLE_NODE.
* Add a patch to restore top-down allocation the way x86 does.
v2:
* http://lkml.kernel.org/r/1473883618-14998-1-git-send-email-arbab@linux.vnet.ibm.com
* Use the "status" property of standard dt memory nodes instead of
introducing a new "ibm,hotplug-aperture" compatible id.
* Remove the patch which explicitly creates a memoryless node. This set
no longer has any bearing on whether the pgdat is created at boot or
at the time of memory addition.
v1:
* http://lkml.kernel.org/r/1470680843-28702-1-git-send-email-arbab@linux.vnet.ibm.com
Reza Arbab (5):
drivers/of: introduce of_fdt_device_is_available()
drivers/of: do not add memory for unavailable nodes
powerpc/mm: allow memory hotplug into a memoryless node
mm: make processing of movable_node arch-specific
mm: enable CONFIG_MOVABLE_NODE on non-x86 arches
arch/powerpc/mm/numa.c | 13 +------------
arch/x86/mm/numa.c | 35 ++++++++++++++++++++++++++++++++++-
drivers/of/fdt.c | 29 ++++++++++++++++++++++++++---
include/linux/of_fdt.h | 2 ++
mm/Kconfig | 2 +-
mm/memory_hotplug.c | 31 -------------------------------
6 files changed, 64 insertions(+), 48 deletions(-)
--
1.8.3.1
In __fdt_scan_reserved_mem(), the availability of a node is determined
by testing its "status" property.
Move this check into its own function, borrowing logic from the
unflattened version, of_device_is_available().
Another caller will be added in a subsequent patch.
Signed-off-by: Reza Arbab <redacted>
Acked-by: Rob Herring <robh@kernel.org>
---
drivers/of/fdt.c | 26 +++++++++++++++++++++++---
include/linux/of_fdt.h | 2 ++
2 files changed, 25 insertions(+), 3 deletions(-)
Currently, CONFIG_MOVABLE_NODE depends on X86_64. In preparation to
enable it for other arches, we need to factor a detail which is unique
to x86 out of the generic mm code.
Specifically, as documented in kernel-parameters.txt, the use of
"movable_node" should remain restricted to x86:
movable_node [KNL,X86] Boot-time switch to enable the effects
of CONFIG_MOVABLE_NODE=y. See mm/Kconfig for details.
This option tells x86 to find movable nodes identified by the ACPI SRAT.
On other arches, it would have no benefit, only the undesired side
effect of setting bottom-up memblock allocation.
Since #ifdef CONFIG_MOVABLE_NODE will no longer be enough to restrict
this option to x86, move it to an arch-specific compilation unit
instead.
Signed-off-by: Reza Arbab <redacted>
---
arch/x86/mm/numa.c | 35 ++++++++++++++++++++++++++++++++++-
mm/memory_hotplug.c | 31 -------------------------------
2 files changed, 34 insertions(+), 32 deletions(-)
@@ -1738,37 +1738,6 @@ static bool can_offline_normal(struct zone *zone, unsigned long nr_pages)}#endif /* CONFIG_MOVABLE_NODE */-staticint__initcmdline_parse_movable_node(char*p)-{-#ifdef CONFIG_MOVABLE_NODE-/*-*Memoryusedbythekernelcannotbehot-removedbecauseLinux-*cannotmigratethekernelpages.Whenmemoryhotplugis-*enabled,weshouldpreventmemblockfromallocatingmemory-*forthekernel.-*-*ACPISRATrecordsallhotpluggablememoryranges.Butbefore-*SRATisparsed,wedon'tknowaboutit.-*-*Thekernelimageisloadedintomemoryatveryearlytime.We-*cannotpreventthisanyway.SoonNUMAsystem,wesetany-*nodethekernelresidesinasun-hotpluggable.-*-*Sinceonmodernservers,onenodecouldhavedouble-digit-*gigabytesmemory,wecanassumethememoryaroundthekernel-*imageisalsoun-hotpluggable.SobeforeSRATisparsed,just-*allocatememorynearthekernelimagetotrythebesttokeep-*thekernelawayfromhotpluggablememory.-*/-memblock_set_bottom_up(true);-movable_node_enabled=true;-#else-pr_warn("movable_node option not supported\n");-#endif-return0;-}-early_param("movable_node",cmdline_parse_movable_node);-/* check which state of node_states will be changed when offline memory */staticvoidnode_states_check_changes_offline(unsignedlongnr_pages,structzone*zone,structmemory_notify*arg)
Remove the check which prevents us from hotplugging into an empty node.
This limitation has been questioned before [1], and judging by the
response, there doesn't seem to be a reason we can't remove it. No issues
have been found in light testing.
[1] http://lkml.kernel.org/r/CAGZKiBrmkSa1yyhbf5hwGxubcjsE5SmkSMY4tpANERMe2UG4bg@mail.gmail.comhttp://lkml.kernel.org/r/20160511215051.GF22115@arbab-laptop.austin.ibm.com
Signed-off-by: Reza Arbab <redacted>
Reviewed-by: Aneesh Kumar K.V <redacted>
Acked-by: Balbir Singh <bsingharora@gmail.com>
Cc: Nathan Fontenot <redacted>
Cc: Bharata B Rao <redacted>
---
arch/powerpc/mm/numa.c | 13 +------------
1 file changed, 1 insertion(+), 12 deletions(-)
@@ -1121,7 +1121,7 @@ static int hot_add_node_scn_to_nid(unsigned long scn_addr)inthot_add_scn_to_nid(unsignedlongscn_addr){structdevice_node*memory=NULL;-intnid,found=0;+intnid;if(!numa_enabled||(min_common_depth<0))returnfirst_online_node;
@@ -1137,17 +1137,6 @@ int hot_add_scn_to_nid(unsigned long scn_addr)if(nid<0||!node_online(nid))nid=first_online_node;-if(NODE_DATA(nid)->node_spanned_pages)-returnnid;--for_each_online_node(nid){-if(NODE_DATA(nid)->node_spanned_pages){-found=1;-break;-}-}--BUG_ON(!found);returnnid;}
To support movable memory nodes (CONFIG_MOVABLE_NODE), at least one of
the following must be true:
1. We're on x86. This arch has the capability to identify movable nodes
at boot by parsing the ACPI SRAT, if the movable_node option is used.
2. Our config supports memory hotplug, which means that a movable node
can be created by hotplugging all of its memory into ZONE_MOVABLE.
Fix the Kconfig definition of CONFIG_MOVABLE_NODE, which currently
recognizes (1), but not (2).
Signed-off-by: Reza Arbab <redacted>
---
mm/Kconfig | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
@@ -153,7 +153,7 @@ config MOVABLE_NODEbool"Enable to assign a node which has only movable memory"depends onHAVE_MEMBLOCKdepends onNO_BOOTMEM-depends onX86_64+depends onX86_64||MEMORY_HOTPLUGdepends onNUMAdefaultnhelp
Respect the standard dt "status" property when scanning memory nodes in
early_init_dt_scan_memory(), so that if the node is unavailable, no
memory will be added.
The use case at hand is accelerator or device memory, which may be
unusable until post-boot initialization of the memory link. Such a node
can be described in the dt as any other, given its status is "disabled".
Per the device tree specification,
"disabled"
Indicates that the device is not presently operational, but it
might become operational in the future (for example, something
is not plugged in, or switched off).
Once such memory is made operational, it can then be hotplugged.
Signed-off-by: Reza Arbab <redacted>
---
drivers/of/fdt.c | 3 +++
1 file changed, 3 insertions(+)
Currently, CONFIG_MOVABLE_NODE depends on X86_64. In preparation to
enable it for other arches, we need to factor a detail which is unique
to x86 out of the generic mm code.
Specifically, as documented in kernel-parameters.txt, the use of
"movable_node" should remain restricted to x86:
movable_node [KNL,X86] Boot-time switch to enable the effects
of CONFIG_MOVABLE_NODE=y. See mm/Kconfig for details.
This option tells x86 to find movable nodes identified by the ACPI SRAT.
On other arches, it would have no benefit, only the undesired side
effect of setting bottom-up memblock allocation.
Since #ifdef CONFIG_MOVABLE_NODE will no longer be enough to restrict
this option to x86, move it to an arch-specific compilation unit
instead.
@@ -1738,37 +1738,6 @@ static bool can_offline_normal(struct zone *zone, unsigned long nr_pages)}#endif /* CONFIG_MOVABLE_NODE */-staticint__initcmdline_parse_movable_node(char*p)-{-#ifdef CONFIG_MOVABLE_NODE-/*-*Memoryusedbythekernelcannotbehot-removedbecauseLinux-*cannotmigratethekernelpages.Whenmemoryhotplugis-*enabled,weshouldpreventmemblockfromallocatingmemory-*forthekernel.-*-*ACPISRATrecordsallhotpluggablememoryranges.Butbefore-*SRATisparsed,wedon'tknowaboutit.-*-*Thekernelimageisloadedintomemoryatveryearlytime.We-*cannotpreventthisanyway.SoonNUMAsystem,wesetany-*nodethekernelresidesinasun-hotpluggable.-*-*Sinceonmodernservers,onenodecouldhavedouble-digit-*gigabytesmemory,wecanassumethememoryaroundthekernel-*imageisalsoun-hotpluggable.SobeforeSRATisparsed,just-*allocatememorynearthekernelimagetotrythebesttokeep-*thekernelawayfromhotpluggablememory.-*/-memblock_set_bottom_up(true);-movable_node_enabled=true;-#else-pr_warn("movable_node option not supported\n");-#endif-return0;-}-early_param("movable_node",cmdline_parse_movable_node);-/* check which state of node_states will be changed when offline memory */staticvoidnode_states_check_changes_offline(unsignedlongnr_pages,structzone*zone,structmemory_notify*arg)
To support movable memory nodes (CONFIG_MOVABLE_NODE), at least one of
the following must be true:
1. We're on x86. This arch has the capability to identify movable nodes
at boot by parsing the ACPI SRAT, if the movable_node option is used.
2. Our config supports memory hotplug, which means that a movable node
can be created by hotplugging all of its memory into ZONE_MOVABLE.
Fix the Kconfig definition of CONFIG_MOVABLE_NODE, which currently
recognizes (1), but not (2).
We now enable a lot of new code on different arch, such as the new node list
N_MEMORY.
Reviewed-by: Aneesh Kumar K.V <redacted>
@@ -153,7 +153,7 @@ config MOVABLE_NODEbool"Enable to assign a node which has only movable memory"depends onHAVE_MEMBLOCKdepends onNO_BOOTMEM-depends onX86_64+depends onX86_64||MEMORY_HOTPLUGdepends onNUMAdefaultnhelp
Currently, CONFIG_MOVABLE_NODE depends on X86_64. In preparation to
enable it for other arches, we need to factor a detail which is unique
to x86 out of the generic mm code.
Specifically, as documented in kernel-parameters.txt, the use of
"movable_node" should remain restricted to x86:
movable_node [KNL,X86] Boot-time switch to enable the effects
of CONFIG_MOVABLE_NODE=y. See mm/Kconfig for details.
This option tells x86 to find movable nodes identified by the ACPI SRAT.
On other arches, it would have no benefit, only the undesired side
effect of setting bottom-up memblock allocation.
Since #ifdef CONFIG_MOVABLE_NODE will no longer be enough to restrict
this option to x86, move it to an arch-specific compilation unit
instead.
Signed-off-by: Reza Arbab <redacted>
To support movable memory nodes (CONFIG_MOVABLE_NODE), at least one of
the following must be true:
1. We're on x86. This arch has the capability to identify movable nodes
at boot by parsing the ACPI SRAT, if the movable_node option is used.
2. Our config supports memory hotplug, which means that a movable node
can be created by hotplugging all of its memory into ZONE_MOVABLE.
Fix the Kconfig definition of CONFIG_MOVABLE_NODE, which currently
recognizes (1), but not (2).
Signed-off-by: Reza Arbab <redacted>
---
mm/Kconfig | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
@@ -153,7 +153,7 @@ config MOVABLE_NODEbool"Enable to assign a node which has only movable memory"depends onHAVE_MEMBLOCKdepends onNO_BOOTMEM-depends onX86_64+depends onX86_64||MEMORY_HOTPLUGdepends onNUMAdefaultnhelp
From: Rob Herring <robh+dt@kernel.org> Date: 2016-10-11 13:59:08
On Thu, Oct 6, 2016 at 1:36 PM, Reza Arbab [off-list ref] wrote:
Respect the standard dt "status" property when scanning memory nodes in
early_init_dt_scan_memory(), so that if the node is unavailable, no
memory will be added.
The use case at hand is accelerator or device memory, which may be
unusable until post-boot initialization of the memory link. Such a node
can be described in the dt as any other, given its status is "disabled".
Per the device tree specification,
"disabled"
Indicates that the device is not presently operational, but it
might become operational in the future (for example, something
is not plugged in, or switched off).
Once such memory is made operational, it can then be hotplugged.
Signed-off-by: Reza Arbab <redacted>
---
drivers/of/fdt.c | 3 +++
1 file changed, 3 insertions(+)
Remove the check which prevents us from hotplugging into an empty node.
This limitation has been questioned before [1], and judging by the
response, there doesn't seem to be a reason we can't remove it. No issues
have been found in light testing.
[1] http://lkml.kernel.org/r/CAGZKiBrmkSa1yyhbf5hwGxubcjsE5SmkSMY4tpANERMe2UG4bg@mail.gmail.comhttp://lkml.kernel.org/r/20160511215051.GF22115@arbab-laptop.austin.ibm.com
Signed-off-by: Reza Arbab <redacted>
Reviewed-by: Aneesh Kumar K.V <redacted>
Acked-by: Balbir Singh <bsingharora@gmail.com>
Cc: Nathan Fontenot <redacted>
Cc: Bharata B Rao <redacted>
---
arch/powerpc/mm/numa.c | 13 +------------
1 file changed, 1 insertion(+), 12 deletions(-)
@@ -1121,7 +1121,7 @@ static int hot_add_node_scn_to_nid(unsigned long scn_addr)inthot_add_scn_to_nid(unsignedlongscn_addr){structdevice_node*memory=NULL;-intnid,found=0;+intnid;if(!numa_enabled||(min_common_depth<0))returnfirst_online_node;
@@ -1137,17 +1137,6 @@ int hot_add_scn_to_nid(unsigned long scn_addr)if(nid<0||!node_online(nid))nid=first_online_node;-if(NODE_DATA(nid)->node_spanned_pages)-returnnid;--for_each_online_node(nid){-if(NODE_DATA(nid)->node_spanned_pages){-found=1;-break;-}-}--BUG_ON(!found);returnnid;
FYI, these checks were temporary to begin with
I found this in git history
b226e462124522f2f23153daff31c311729dfa2f (powerpc: don't add memory to empty node/zone)
Balbir Singh.
On Thu, Oct 20, 2016 at 02:30:42PM +1100, Balbir Singh wrote:
FYI, these checks were temporary to begin with
I found this in git history
b226e462124522f2f23153daff31c311729dfa2f (powerpc: don't add memory to empty node/zone)
Nice find! I spent some time digging, but this had eluded me.
--
Reza Arbab
Hi Reza,
On Thu, 6 Oct 2016 01:36:32 PM Reza Arbab wrote:
Respect the standard dt "status" property when scanning memory nodes in
early_init_dt_scan_memory(), so that if the node is unavailable, no
memory will be added.
What happens if a kernel without this patch is booted on a system with some
status="disabled" device-nodes? Do older kernels just ignore this memory or do
they try to use it?
From what I can tell it seems that kernels without this patch will try and use
this memory even if it is marked in the device-tree as status="disabled" which
could lead to problems for older kernels when we start exporting this property
from firmware.
Arguably this might not be such a problem in practice as we probably don't
have many (if any) existing kernels that will boot on hardware exporting these
properties. However given this patch seems fairly independent perhaps it is
worth sending as a separate fix if it is not going to make it into this
release?
Regards,
Alistair
quoted hunk
The use case at hand is accelerator or device memory, which may be
unusable until post-boot initialization of the memory link. Such a node
can be described in the dt as any other, given its status is "disabled".
Per the device tree specification,
"disabled"
Indicates that the device is not presently operational, but it
might become operational in the future (for example, something
is not plugged in, or switched off).
Once such memory is made operational, it can then be hotplugged.
Signed-off-by: Reza Arbab <redacted>
---
drivers/of/fdt.c | 3 +++
1 file changed, 3 insertions(+)
Hi Alistair,
On Fri, Oct 21, 2016 at 05:22:54PM +1100, Alistair Popple wrote:
From what I can tell it seems that kernels without this patch will try
and use this memory even if it is marked in the device-tree as
status="disabled" which could lead to problems for older kernels when
we start exporting this property from firmware.
Arguably this might not be such a problem in practice as we probably
don't have many (if any) existing kernels that will boot on hardware
exporting these properties.
Yes, I think you've got it right.
However given this patch seems fairly independent perhaps it is worth
sending as a separate fix if it is not going to make it into this
release?
Michael,
If this set as a whole is going to miss the release, would it be helpful
for me to resend 1/5 and 2/5 as a separate set? They are the minimum
needed to prevent the possible forward compatibility issue Alistair
describes.
--
Reza Arbab
From: Michael Ellerman <mpe@ellerman.id.au> Date: 2016-10-24 10:24:15
Alistair Popple [off-list ref] writes:
Hi Reza,
On Thu, 6 Oct 2016 01:36:32 PM Reza Arbab wrote:
quoted
Respect the standard dt "status" property when scanning memory nodes in
early_init_dt_scan_memory(), so that if the node is unavailable, no
memory will be added.
What happens if a kernel without this patch is booted on a system with some
status="disabled" device-nodes? Do older kernels just ignore this memory or do
they try to use it?
From what I can tell it seems that kernels without this patch will try and use
this memory even if it is marked in the device-tree as status="disabled" which
could lead to problems for older kernels when we start exporting this property
from firmware.
The code already looks for "linux,usable-memory" in preference to "reg".
Can you use that instead?
That would have the advantage that existing kernels already understand
it.
Another problem with using "status" is we could have device trees out
there that have status = disabled and we don't know about it, and by
changing the kernel to use that property we break people's systems.
Though for memory nodes my guess is that's not true, but you never know ...
cheers
On Mon, Oct 24, 2016 at 09:24:04PM +1100, Michael Ellerman wrote:
The code already looks for "linux,usable-memory" in preference to
"reg". Can you use that instead?
Yes, we could set the size of "linux,usable-memory" to zero instead of
setting status to "disabled".
I'll send a v5 of this set which drops 1/5 and 2/5. That would be the
only difference here.
That would have the advantage that existing kernels already understand
it.
Another problem with using "status" is we could have device trees out
there that have status = disabled and we don't know about it, and by
changing the kernel to use that property we break people's systems.
Though for memory nodes my guess is that's not true, but you never know ...
From: Michael Ellerman <mpe@ellerman.id.au> Date: 2016-10-25 09:40:08
Balbir Singh [off-list ref] writes:
FYI, these checks were temporary to begin with
I found this in git history
b226e462124522f2f23153daff31c311729dfa2f (powerpc: don't add memory to empty node/zone)
Nice thanks for digging it up.
commit b226e462124522f2f23153daff31c311729dfa2f
Author: Mike Kravetz [off-list ref]
AuthorDate: Fri Dec 16 14:30:35 2005 -0800
^^^^
That is why maintainers don't like to merge "temporary" patches :)
cheers
Currently, CONFIG_MOVABLE_NODE depends on X86_64. In preparation to
enable it for other arches, we need to factor a detail which is unique
to x86 out of the generic mm code.
Specifically, as documented in kernel-parameters.txt, the use of
"movable_node" should remain restricted to x86:
movable_node [KNL,X86] Boot-time switch to enable the effects
of CONFIG_MOVABLE_NODE=y. See mm/Kconfig for details.
This option tells x86 to find movable nodes identified by the ACPI SRAT.
On other arches, it would have no benefit, only the undesired side
effect of setting bottom-up memblock allocation.
Since #ifdef CONFIG_MOVABLE_NODE will no longer be enough to restrict
this option to x86, move it to an arch-specific compilation unit
instead.
Signed-off-by: Reza Arbab <redacted>
Acked-by: Balbir Singh <bsingharora@gmail.com>
After the ack, I realized there were some more checks needed, IOW
questions for you :)
1. Have you checked to see if our memblock allocations spill
over to probably hotpluggable nodes?
2. Shouldn't we be marking nodes discovered as movable via
memblock_mark_hotplug()?
Balbir Singh.
On Tue, Oct 25, 2016 at 11:15:40PM +1100, Balbir Singh wrote:
After the ack, I realized there were some more checks needed, IOW
questions for you :)
Hey! No takebacks!
The short answer is that neither of these is a concern.
Longer; if you use "movable_node", x86 can identify these nodes at boot.
They call memblock_mark_hotplug() while parsing the SRAT. Then, when the
zones are initialized, those markings are used to determine ZONE_MOVABLE.
We have no analog of this SRAT information, so our movable nodes can
only be created post boot, by hotplugging and explicitly onlining with
online_movable.
1. Have you checked to see if our memblock allocations spill
over to probably hotpluggable nodes?
Since our nodes don't exist at boot, we don't have that short window
before the zones are drawn where the node has normal memory, and a
kernel allocation might occur within.
2. Shouldn't we be marking nodes discovered as movable via
memblock_mark_hotplug()?
Again, this early boot marking mechanism only applies to movable_node.
--
Reza Arbab
On Tue, Oct 25, 2016 at 11:15:40PM +1100, Balbir Singh wrote:
quoted
After the ack, I realized there were some more checks needed, IOW
questions for you :)
Hey! No takebacks!
I still believe we need your changes, I was wondering if we've tested
it against normal memory nodes and checked if any memblock
allocations end up there. Michael showed me some memblock
allocations on node 1 of a two node machine with movable_node
I'll double check at my end. See my question below
The short answer is that neither of these is a concern.
Longer; if you use "movable_node", x86 can identify these nodes at boot. They call memblock_mark_hotplug() while parsing the SRAT. Then, when the zones are initialized, those markings are used to determine ZONE_MOVABLE.
We have no analog of this SRAT information, so our movable nodes can only be created post boot, by hotplugging and explicitly onlining with online_movable.
Is this true for all of system memory as well or only for nodes
hotplugged later?
Balbir Singh.
On Tue, Oct 25, 2016 at 11:15:40PM +1100, Balbir Singh wrote:
quoted
After the ack, I realized there were some more checks needed, IOW
questions for you :)
Hey! No takebacks!
I still believe we need your changes, I was wondering if we've tested
it against normal memory nodes and checked if any memblock
allocations end up there. Michael showed me some memblock
allocations on node 1 of a two node machine with movable_node
I'll double check at my end. See my question below
The short answer is that neither of these is a concern.
Longer; if you use "movable_node", x86 can identify these nodes at boot. They call memblock_mark_hotplug() while parsing the SRAT. Then, when the zones are initialized, those markings are used to determine ZONE_MOVABLE.
We have no analog of this SRAT information, so our movable nodes can only be created post boot, by hotplugging and explicitly onlining with online_movable.
Is this true for all of system memory as well or only for nodes
hotplugged later?
Balbir Singh.
On Wed, Oct 26, 2016 at 09:34:18AM +1100, Balbir Singh wrote:
I still believe we need your changes, I was wondering if we've tested
it against normal memory nodes and checked if any memblock
allocations end up there. Michael showed me some memblock
allocations on node 1 of a two node machine with movable_node
The movable_node option is x86-only. Both of those nodes contain normal
memory, so allocations on both are allowed.
quoted
Longer; if you use "movable_node", x86 can identify these nodes at
boot. They call memblock_mark_hotplug() while parsing the SRAT. Then,
when the zones are initialized, those markings are used to determine
ZONE_MOVABLE.
We have no analog of this SRAT information, so our movable nodes can
only be created post boot, by hotplugging and explicitly onlining
with online_movable.
Is this true for all of system memory as well or only for nodes
hotplugged later?
As far as I know, power has nothing like the SRAT that tells us, at
boot, which memory is hotpluggable. So there is nothing to wire the
movable_node option up to.
Of course, any memory you hotplug afterwards is, by definition,
hotpluggable. So we can still create movable nodes that way.
--
Reza Arbab
From: Michael Ellerman <mpe@ellerman.id.au> Date: 2016-10-26 10:53:24
Reza Arbab [off-list ref] writes:
On Wed, Oct 26, 2016 at 09:34:18AM +1100, Balbir Singh wrote:
quoted
I still believe we need your changes, I was wondering if we've tested
it against normal memory nodes and checked if any memblock
allocations end up there. Michael showed me some memblock
allocations on node 1 of a two node machine with movable_node
The movable_node option is x86-only. Both of those nodes contain normal
memory, so allocations on both are allowed.
quoted
quoted
Longer; if you use "movable_node", x86 can identify these nodes at
boot. They call memblock_mark_hotplug() while parsing the SRAT. Then,
when the zones are initialized, those markings are used to determine
ZONE_MOVABLE.
We have no analog of this SRAT information, so our movable nodes can
only be created post boot, by hotplugging and explicitly onlining
with online_movable.
Is this true for all of system memory as well or only for nodes
hotplugged later?
As far as I know, power has nothing like the SRAT that tells us, at
boot, which memory is hotpluggable.
On pseries we have the ibm,dynamic-memory device tree property, which
can contain ranges of memory that are not yet "assigned to the
partition" - ie. can be hotplugged later.
So in general that statement is not true.
But I think you're focused on bare-metal, in which case you might be
right. But that doesn't mean we couldn't have a similar property, if
skiboot/hostboot knew what the ranges of memory were going to be.
cheers
On Wed, Oct 26, 2016 at 09:52:53PM +1100, Michael Ellerman wrote:
quoted
As far as I know, power has nothing like the SRAT that tells us, at
boot, which memory is hotpluggable.
On pseries we have the ibm,dynamic-memory device tree property, which
can contain ranges of memory that are not yet "assigned to the
partition" - ie. can be hotplugged later.
So in general that statement is not true.
But I think you're focused on bare-metal, in which case you might be
right. But that doesn't mean we couldn't have a similar property, if
skiboot/hostboot knew what the ranges of memory were going to be.
Yes, sorry, I should have qualified that statement to say I wasn't
talking about pseries.
I can amend this set to actually implement movable_node on power too,
but we'd have to settle on a name for the dt property. Is
"linux,movable-node" too on the nose?
--
Reza Arbab