Re: [RFC/PATCH] powerpc/smp: Add SD_SHARE_PKG_RESOURCES flag to MC sched-domain
From: Vincent Guittot <vincent.guittot@linaro.org>
Date: 2021-04-12 12:22:06
Also in:
lkml
On Mon, 12 Apr 2021 at 11:37, Mel Gorman [off-list ref] wrote:
On Mon, Apr 12, 2021 at 11:54:36AM +0530, Srikar Dronamraju wrote:quoted
* Gautham R. Shenoy [off-list ref] [2021-04-02 11:07:54]:quoted
To remedy this, this patch proposes that the LLC be moved to the MC level which is a group of cores in one half of the chip. SMT (SMT4) --> MC (Hemisphere)[LLC] --> DIEI think marking Hemisphere as a LLC in a P10 scenario is a good idea.quoted
While there is no cache being shared at this level, this is still the level where some amount of cache-snooping takes place and it is relatively faster to access the data from the caches of the cores within this domain. With this change, we no longer see regressions on P10 for applications which require single threaded performance.Peter, Valentin, Vincent, Mel, etal On architectures where we have multiple levels of cache access latencies within a DIE, (For example: one within the current LLC or SMT core and the other at MC or Hemisphere, and finally across hemispheres), do you have any suggestions on how we could handle the same in the core scheduler?
I would say that SD_SHARE_PKG_RESOURCES is there for that and doesn't only rely on cache
quoted
Minimally I think it would be worth detecting when there are multiple LLCs per node and detecting that in generic code as a static branch. In select_idle_cpu, consider taking two passes -- first on the LLC domain and if no idle CPU is found then taking a second pass if the search depth
We have done a lot of changes to reduce and optimize the fast path and I don't think re adding another layer in the fast path makes sense as you will end up unrolling the for_each_domain behind some static_banches. SD_SHARE_PKG_RESOURCES should be set to the last level where we can efficiently move task between CPUs at wakeup
allows within the node with the LLC CPUs masked out. While there would be a latency hit because cache is not shared, it would still be a CPU local to memory that is idle. That would potentially be beneficial on Zen* as well without having to introduce new domains in the topology hierarchy.
What is the current sched_domain topology description for zen ?
-- Mel Gorman SUSE Labs