Abdul reported a warning on a shared lpar.
"WARNING: workqueue cpumask: online intersect > possible intersect".
This is because per node workqueue possible mask is set very early in the
boot process even before the system was querying the home node
associativity. However per node workqueue online cpumask gets updated
dynamically. Hence there is a chance when per node workqueue online cpumask
is a superset of per node workqueue possible mask.
Link for v1: https://patchwork.ozlabs.org/patch/1151658
Changelog: v1->v2
- Handled comments from Nathan Lynch.
Link for v2: http://lkml.kernel.org/r/20190829055023.6171-1-srikar@linux.vnet.ibm.com
Changelog: v2->v3
- Handled comments from Nathan Lynch.
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Nicholas Piggin <npiggin@gmail.com>
Cc: Abdul Haleem <redacted>
Cc: Nathan Lynch <redacted>
Cc: linuxppc-dev@lists.ozlabs.org
Cc: Satheesh Rajendran <redacted>
Srikar Dronamraju (5):
powerpc/vphn: Check for error from hcall_vphn
powerpc/numa: Handle extra hcall_vphn error cases
powerpc/numa: Use cpu node map of first sibling thread
powerpc/numa: Early request for home node associativity
powerpc/numa: Remove late request for home node associativity
arch/powerpc/include/asm/topology.h | 4 --
arch/powerpc/kernel/smp.c | 5 --
arch/powerpc/mm/numa.c | 96 ++++++++++++++++++++++++++---------
arch/powerpc/platforms/pseries/vphn.c | 3 +-
4 files changed, 74 insertions(+), 34 deletions(-)
--
1.8.3.1
There is no value in unpacking associativity, if
H_HOME_NODE_ASSOCIATIVITY hcall has returned an error.
Signed-off-by: Srikar Dronamraju <redacted>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Nicholas Piggin <npiggin@gmail.com>
Cc: Nathan Lynch <redacted>
Cc: linuxppc-dev@lists.ozlabs.org
Cc: Satheesh Rajendran <redacted>
Reported-by: Abdul Haleem <redacted>
---
Changelog (v2->v1):
- Split the patch into 2(Suggested by Nathan).
arch/powerpc/platforms/pseries/vphn.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
Currently code handles H_FUNCTION, H_SUCCESS, H_HARDWARE return codes.
However hcall_vphn can return other return codes. Now it also handles
H_PARAMETER return code. Also the rest return codes are handled under the
default case.
Signed-off-by: Srikar Dronamraju <redacted>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Nicholas Piggin <npiggin@gmail.com>
Cc: Nathan Lynch <redacted>
Cc: linuxppc-dev@lists.ozlabs.org
Cc: Satheesh Rajendran <redacted>
Reported-by: Abdul Haleem <redacted>
---
Changelog (v2->v1):
Handled comments from Nathan:
- Split patch from patch 1.
- Corrected a problem where I missed calling stop_topology_update().
- Using pr_err_ratelimited instead of printk.
arch/powerpc/mm/numa.c | 25 ++++++++++++++++---------
1 file changed, 16 insertions(+), 9 deletions(-)
Currently the kernel detects if its running on a shared lpar platform
and requests home node associativity before the scheduler sched_domains
are setup. However between the time NUMA setup is initialized and the
request for home node associativity, workqueue initializes its per node
cpumask. The per node workqueue possible cpumask may turn invalid
after home node associativity resulting in weird situations like
workqueue possible cpumask being a subset of workqueue online cpumask.
This can be fixed by requesting home node associativity earlier just
before NUMA setup. However at the NUMA setup time, kernel may not be in
a position to detect if its running on a shared lpar platform. So
request for home node associativity and if the request fails, fallback
on the device tree property.
Signed-off-by: Srikar Dronamraju <redacted>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Nicholas Piggin <npiggin@gmail.com>
Cc: Nathan Lynch <redacted>
Cc: linuxppc-dev@lists.ozlabs.org
Cc: Satheesh Rajendran <redacted>
Reported-by: Abdul Haleem <redacted>
---
Changelog (v1->v2):
- Handled comments from Nathan Lynch
* Dont depend on pacas to be setup for the hwid
Changelog (v2->v3):
- Handled comments from Nathan Lynch
* Use first thread of the core for cpu-to-node map.
* get hardware-id in numa_setup_cpu
arch/powerpc/mm/numa.c | 45 ++++++++++++++++++++++++++++++++++++++++-----
1 file changed, 40 insertions(+), 5 deletions(-)
@@ -461,13 +461,27 @@ static int of_drconf_to_nid_single(struct drmem_lmb *lmb)returnnid;}+staticintvphn_get_nid(longhwid)+{+__be32associativity[VPHN_ASSOC_BUFSIZE]={0};+longrc;++rc=hcall_vphn(hwid,VPHN_FLAG_VCPU,associativity);+if(rc==H_SUCCESS)+returnassociativity_to_nid(associativity);++returnNUMA_NO_NODE;+}+/**Figureouttowhichdomainacpubelongsandstickitthere.+*cpu_to_phys_idisonlyvalidbetweensmp_setup_cpu_maps()and+*smp_setup_pacas().Ifcalledoutsidethiswindow,setget_hwidtotrue.*Returntheidofthedomainused.*/-staticintnuma_setup_cpu(unsignedlonglcpu)+staticintnuma_setup_cpu(unsignedlonglcpu,boolget_hwid){-structdevice_node*cpu;+structdevice_node*cpu=NULL;intfcpu=cpu_first_thread_sibling(lcpu);intnid=NUMA_NO_NODE;
@@ -485,6 +499,27 @@ static int numa_setup_cpu(unsigned long lcpu)returnnid;}+/*+*Onasharedlpar,devicetreewillnothavenodeassociativity.+*Atthistimelppaca,orits__old_statusfieldmaynotbe+*updated.Hencekernelcannotdetectifitsonasharedlpar.So+*requestanexplicitassociativityirrespectiveofwhetherthe+*lparissharedordedicated.Usethedevicetreepropertyasa+*fallback.+*/+if(firmware_has_feature(FW_FEATURE_VPHN)){+longhwid;++if(get_hwid)+hwid=get_hard_smp_processor_id(lcpu);+else+hwid=cpu_to_phys_id[lcpu];+nid=vphn_get_nid(hwid);+}++if(nid!=NUMA_NO_NODE)+gotoout_present;+cpu=of_get_cpu_node(lcpu,NULL);if(!cpu){
@@ -496,6 +531,7 @@ static int numa_setup_cpu(unsigned long lcpu)}nid=of_node_to_nid_single(cpu);+of_node_put(cpu);out_present:if(nid<0||!node_possible(nid))
@@ -512,7 +548,6 @@ static int numa_setup_cpu(unsigned long lcpu)map_cpu_to_node(fcpu,nid);map_cpu_to_node(lcpu,nid);-of_node_put(cpu);out:returnnid;}
@@ -543,7 +578,7 @@ static int ppc_numa_cpu_prepare(unsigned int cpu){intnid;-nid=numa_setup_cpu(cpu);+nid=numa_setup_cpu(cpu,true);verify_cpu_node_mapping(cpu,nid);return0;}
All the sibling threads of a core have to be part of the same node.
To ensure that all the sibling threads map to the same node, always
lookup/update the cpu-to-node map of the first thread in the core.
Signed-off-by: Srikar Dronamraju <redacted>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Nicholas Piggin <npiggin@gmail.com>
Cc: Nathan Lynch <redacted>
Cc: linuxppc-dev@lists.ozlabs.org
Cc: Satheesh Rajendran <redacted>
Reported-by: Abdul Haleem <redacted>
---
arch/powerpc/mm/numa.c | 19 +++++++++++++++++--
1 file changed, 17 insertions(+), 2 deletions(-)
@@ -467,15 +467,20 @@ static int of_drconf_to_nid_single(struct drmem_lmb *lmb)*/staticintnuma_setup_cpu(unsignedlonglcpu){-intnid=NUMA_NO_NODE;structdevice_node*cpu;+intfcpu=cpu_first_thread_sibling(lcpu);+intnid=NUMA_NO_NODE;/**Ifavalidcpu-to-nodemappingisalreadyavailable,useit*directlyinsteadofqueryingthefirmware,sinceitrepresents*themostrecentmappingnotifiedtousbytheplatform(eg:VPHN).+*Sincecpu_to_nodebindingremainsthesameforallthreadsinthe+*core.Ifavalidcpu-to-nodemappingisalreadyavailable,for+*thefirstthreadinthecore,useit.*/-if((nid=numa_cpu_lookup_table[lcpu])>=0){+nid=numa_cpu_lookup_table[fcpu];+if(nid>=0){map_cpu_to_node(lcpu,nid);returnnid;}
@@ -496,6 +501,16 @@ static int numa_setup_cpu(unsigned long lcpu)if(nid<0||!node_possible(nid))nid=first_online_node;+/*+*Updateforthefirstthreadofthecore.Allthreadsofacore+*havetobepartofthesamenode.Thisnotonlyavoidsquerying+*foreveryotherthreadinthecore,butalwaysavoidsacase+*wherevirtualnodeassociativitychangecausessubsequentthreads+*ofacoretobeassociatedwithdifferentnid.+*/+if(fcpu!=lcpu)+map_cpu_to_node(fcpu,nid);+map_cpu_to_node(lcpu,nid);of_node_put(cpu);out:
With commit ("powerpc/numa: Early request for home node associativity"),
commit 2ea626306810 ("powerpc/topology: Get topology for shared
processors at boot") which was requesting home node associativity
becomes redundant.
Hence remove the late request for home node associativity.
Signed-off-by: Srikar Dronamraju <redacted>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Nicholas Piggin <npiggin@gmail.com>
Cc: Nathan Lynch <redacted>
Cc: linuxppc-dev@lists.ozlabs.org
Cc: Satheesh Rajendran <redacted>
Reported-by: Abdul Haleem <redacted>
---
arch/powerpc/include/asm/topology.h | 4 ----
arch/powerpc/kernel/smp.c | 5 -----
arch/powerpc/mm/numa.c | 9 ---------
3 files changed, 18 deletions(-)
Hi Srikar,
Srikar Dronamraju [off-list ref] writes:
quoted hunk
@@ -467,15 +467,20 @@ static int of_drconf_to_nid_single(struct drmem_lmb *lmb) */ static int numa_setup_cpu(unsigned long lcpu) {- int nid = NUMA_NO_NODE; struct device_node *cpu;+ int fcpu = cpu_first_thread_sibling(lcpu);+ int nid = NUMA_NO_NODE; /* * If a valid cpu-to-node mapping is already available, use it * directly instead of querying the firmware, since it represents * the most recent mapping notified to us by the platform (eg: VPHN).+ * Since cpu_to_node binding remains the same for all threads in the+ * core. If a valid cpu-to-node mapping is already available, for+ * the first thread in the core, use it. */- if ((nid = numa_cpu_lookup_table[lcpu]) >= 0) {+ nid = numa_cpu_lookup_table[fcpu];+ if (nid >= 0) { map_cpu_to_node(lcpu, nid); return nid; }
Yes, we need to something like this to prevent a VPHN change that occurs
concurrently with onlining a core's threads from messing us up.
Is it a good assumption that the first thread of a sibling group will
have its mapping initialized first? I think the answer is yes for boot,
but hotplug... not so sure.
quoted hunk
@@ -496,6 +501,16 @@ static int numa_setup_cpu(unsigned long lcpu) if (nid < 0 || !node_possible(nid)) nid = first_online_node;+ /*+ * Update for the first thread of the core. All threads of a core+ * have to be part of the same node. This not only avoids querying+ * for every other thread in the core, but always avoids a case+ * where virtual node associativity change causes subsequent threads+ * of a core to be associated with different nid.+ */+ if (fcpu != lcpu)+ map_cpu_to_node(fcpu, nid);+
OK, I see that this somewhat addresses my concern above. But changing
this mapping for a remote cpu is unsafe except under specific
circumstances. I think this should first assert:
* numa_cpu_lookup_table[fcpu] == NUMA_NO_NODE
* cpu_online(fcpu) == false
to document and enforce the conditions that must hold for this to be OK.
- if ((nid = numa_cpu_lookup_table[lcpu]) >= 0) {
+ nid = numa_cpu_lookup_table[fcpu];
+ if (nid >= 0) {
map_cpu_to_node(lcpu, nid);
return nid;
}
Yes, we need to something like this to prevent a VPHN change that occurs
concurrently with onlining a core's threads from messing us up.
Is it a good assumption that the first thread of a sibling group will
have its mapping initialized first? I think the answer is yes for boot,
but hotplug... not so sure.
quoted
@@ -496,6 +501,16 @@ static int numa_setup_cpu(unsigned long lcpu) if (nid < 0 || !node_possible(nid)) nid = first_online_node;+ /*+ * Update for the first thread of the core. All threads of a core+ * have to be part of the same node. This not only avoids querying+ * for every other thread in the core, but always avoids a case+ * where virtual node associativity change causes subsequent threads+ * of a core to be associated with different nid.+ */+ if (fcpu != lcpu)+ map_cpu_to_node(fcpu, nid);+
OK, I see that this somewhat addresses my concern above. But changing
this mapping for a remote cpu is unsafe except under specific
circumstances. I think this should first assert:
* numa_cpu_lookup_table[fcpu] == NUMA_NO_NODE
* cpu_online(fcpu) == false
to document and enforce the conditions that must hold for this to be OK.
I do understand that we shouldn't be modifying the nid for a different cpu.
We just checked above that the mapping for the first cpu doesnt exist.
If the first cpu (or remote cpu as you coin it) was online, then its
mapping should have existed and we return even before we come here.
nid = numa_cpu_lookup_table[fcpu];
if (nid >= 0) {
map_cpu_to_node(lcpu, nid);
return nid;
}
Currently numa_setup_cpus is only called at very early boot and in cpu
hotplug. At hotplug time, the oneline of cpus is serialized. Right? Do we
see a chance of remote cpu changing its state as we set its nid here?
Also lets say if we assert and for some unknown reason the assertion fails.
How do we handle the failure case? We cant get out without setting
the nid. We cant continue setting the nid. Should we panic the system given
that the check a few lines above is now turning out to be false? Probably
no, as I think we can live with it.
Any thoughts?
--
Thanks and Regards
Srikar Dronamraju
Hi Srikar,
Srikar Dronamraju [off-list ref] writes:
quoted
quoted
@@ -496,6 +501,16 @@ static int numa_setup_cpu(unsigned long lcpu) if (nid < 0 || !node_possible(nid)) nid = first_online_node;+ /*+ * Update for the first thread of the core. All threads of a core+ * have to be part of the same node. This not only avoids querying+ * for every other thread in the core, but always avoids a case+ * where virtual node associativity change causes subsequent threads+ * of a core to be associated with different nid.+ */+ if (fcpu != lcpu)+ map_cpu_to_node(fcpu, nid);+
OK, I see that this somewhat addresses my concern above. But changing
this mapping for a remote cpu is unsafe except under specific
circumstances. I think this should first assert:
* numa_cpu_lookup_table[fcpu] == NUMA_NO_NODE
* cpu_online(fcpu) == false
to document and enforce the conditions that must hold for this to be OK.
I do understand that we shouldn't be modifying the nid for a different cpu.
We just checked above that the mapping for the first cpu doesnt exist.
If the first cpu (or remote cpu as you coin it) was online, then its
mapping should have existed and we return even before we come here.
I agree that is how the code will work with your change, and I'm fine
with simply warning if fcpu is offline.
The point is to make this rule more explicit in the code for the benefit
of future readers and to catch violations of it by future changes. There
is a fair amount of code remaining in this file and elsewhere in
arch/powerpc that was written under the impression that changing the
cpu-node relationship at runtime is OK.
nid = numa_cpu_lookup_table[fcpu];
if (nid >= 0) {
map_cpu_to_node(lcpu, nid);
return nid;
}
Currently numa_setup_cpus is only called at very early boot and in cpu
hotplug. At hotplug time, the oneline of cpus is serialized. Right? Do we
see a chance of remote cpu changing its state as we set its nid here?
Also lets say if we assert and for some unknown reason the assertion fails.
How do we handle the failure case? We cant get out without setting
the nid. We cant continue setting the nid. Should we panic the system given
that the check a few lines above is now turning out to be false? Probably
no, as I think we can live with it.
Any thoughts?
I think just WARN_ON(cpu_online(fcpu)) would be satisfactory. In my
experience, the downstream effects of violating this condition are
varied and quite difficult to debug. Seems only appropriate to emit a
warning and stack trace before the OS inevitably becomes unstable.
I think just WARN_ON(cpu_online(fcpu)) would be satisfactory. In my
experience, the downstream effects of violating this condition are
varied and quite difficult to debug. Seems only appropriate to emit a
warning and stack trace before the OS inevitably becomes unstable.
I still have to try but wouldn't this be a problem for the boot-cpu?
I mean boot-cpu would be marked online while it tries to do numa_setup_cpu.
No?
--
Thanks and Regards
Srikar Dronamraju
I think just WARN_ON(cpu_online(fcpu)) would be satisfactory. In my
experience, the downstream effects of violating this condition are
varied and quite difficult to debug. Seems only appropriate to emit a
warning and stack trace before the OS inevitably becomes unstable.
I still have to try but wouldn't this be a problem for the boot-cpu?
I mean boot-cpu would be marked online while it tries to do numa_setup_cpu.
No?
This is what I mean:
+ if (fcpu != lcpu) {
+ WARN_ON(cpu_online(fcpu));
+ map_cpu_to_node(fcpu, nid);
+ }
I.e. if we're modifying the mapping for a remote cpu, warn if it's
online.
I don't think this would warn on the boot cpu -- I would expect fcpu and
lcpu to be the same and this branch would not be taken.
I think just WARN_ON(cpu_online(fcpu)) would be satisfactory. In my
experience, the downstream effects of violating this condition are
varied and quite difficult to debug. Seems only appropriate to emit a
warning and stack trace before the OS inevitably becomes unstable.
I still have to try but wouldn't this be a problem for the boot-cpu?
I mean boot-cpu would be marked online while it tries to do numa_setup_cpu.
No?
This is what I mean:
+ if (fcpu != lcpu) {
+ WARN_ON(cpu_online(fcpu));
+ map_cpu_to_node(fcpu, nid);
+ }
Yes this should work. Will send an updated patch with this change.
I.e. if we're modifying the mapping for a remote cpu, warn if it's
online.
I don't think this would warn on the boot cpu -- I would expect fcpu and
lcpu to be the same and this branch would not be taken.