From: Patrick Bellasi <hidden> Date: 2019-01-15 10:15:30
Hi all, this is a respin of:
https://lore.kernel.org/lkml/20181029183311.29175-1-patrick.bellasi@arm.com/
which addresses all the comments collected in the previous posting and during
the LPC presentation [1].
It's based on v5.0-rc2, the full tree is available here:
git://linux-arm.org/linux-pb.git lkml/utilclamp_v6
http://www.linux-arm.org/git?p=linux-pb.git;a=shortlog;h=refs/heads/lkml/utilclamp_v6
Changes in this version are:
- rebased on top of recently merged EAS code [3] and better integrated with it
- squashed bucketization patch into previous patches
- wholesale s/group/bucket/
- wholesale s/_{get,put}/_{inc,dec}/ to match refcount APIs
- updated cmpxchg loops to looks like "do { } while (cmpxchg(ptr, old, new) != old)"
- switched to usage of try_cmpxchg()
- use SCHED_WARN_ON() instead of CONFIG_SCHED_DEBUG guarded blocks
- moved UCLAMP_FLAG_IDLE management into dedicated functions, i.e.
uclamp_idle_value() and uclamp_idle_reset()
- switched from rq::uclamp::flags to rq::uclamp_flags,
since now rq::uclamp is a per-clamp_id array
- added size check in sched_copy_attr()
- ensure se_count will never underflow
- better comment invariant conditions
- consistently use unary (++/--) operators
- redefined UCLAMP_GROUPS_COUNT range to be [5..20]
- added and make use of the bit_for() macro
- replaced some ifdifery with IS_ENABLED() checks
- overall documentation review to match new subsystem/maintainer
handbook for tip/sched/core
Thanks to all the valuable comments, hopefully this should be a reasonably
stable version for all the core scheduler bits. Thus, I hope we should
be in a good position to unlock Tejun [2] to delve into the review of
the proposed cgroup integration, but let see what Peter and Ingo think
before.
Cheers Patrick
Series Organization
===================
The series is organized into these main sections:
- Patches [01-07]: Per task (primary) API
- Patches [08-09]: Schedutil integration for CFS and RT tasks
- Patches [10-11]: EAS's energy_compute() integration
- Patches [12-16]: Per task group (secondary) API
Newcomer's Short Abstract
=========================
The Linux scheduler tracks a "utilization" signal for each scheduling entity
(SE), e.g. tasks, to know how much CPU time they use. This signal allows the
scheduler to know how "big" a task is and, in principle, it can support
advanced task placement strategies by selecting the best CPU to run a task.
Some of these strategies are represented by the Energy Aware Scheduler [3].
When the schedutil cpufreq governor is in use, the utilization signal allows
the Linux scheduler to also drive frequency selection. The CPU utilization
signal, which represents the aggregated utilization of tasks scheduled on that
CPU, is used to select the frequency which best fits the workload generated by
the tasks.
The current translation of utilization values into a frequency selection is
simple: we go to max for RT tasks or to the minimum frequency which can
accommodate the utilization of DL+FAIR tasks.
However, utilisation values by themselves cannot convey the desired
power/performance behaviours of each task as intended by user-space.
As such they are not ideally suited for task placement decisions.
Task placement and frequency selection policies in the kernel can be improved
by taking into consideration hints coming from authorised user-space elements,
like for example the Android middleware or more generally any "System
Management Software" (SMS) framework.
Utilization clamping is a mechanism which allows to "clamp" (i.e. filter) the
utilization generated by RT and FAIR tasks within a range defined by user-space.
The clamped utilization value can then be used, for example, to enforce a
minimum and/or maximum frequency depending on which tasks are active on a CPU.
The main use-cases for utilization clamping are:
- boosting: better interactive response for small tasks which
are affecting the user experience.
Consider for example the case of a small control thread for an external
accelerator (e.g. GPU, DSP, other devices). Here, from the task utilization
the scheduler does not have a complete view of what the task's requirements
are and, if it's a small utilization task, it keeps selecting a more energy
efficient CPU, with smaller capacity and lower frequency, thus negatively
impacting the overall time required to complete task activations.
- capping: increase energy efficiency for background tasks not affecting the
user experience.
Since running on a lower capacity CPU at a lower frequency is more energy
efficient, when the completion time is not a main goal, then capping the
utilization considered for certain (maybe big) tasks can have positive
effects, both on energy consumption and thermal headroom.
This feature allows also to make RT tasks more energy friendly on mobile
systems where running them on high capacity CPUs and at the maximum
frequency is not required.
From these two use-cases, it's worth noticing that frequency selection
biasing, introduced by patches 9 and 10 of this series, is just one possible
usage of utilization clamping. Another compelling extension of utilization
clamping is in helping the scheduler in macking tasks placement decisions.
Utilization is (also) a task specific property the scheduler uses to know
how much CPU bandwidth a task requires, at least as long as there is idle time.
Thus, the utilization clamp values, defined either per-task or per-task_group,
can represent tasks to the scheduler as being bigger (or smaller) than what
they actually are.
Utilization clamping thus enables interesting additional optimizations, for
example on asymmetric capacity systems like Arm big.LITTLE and DynamIQ CPUs,
where:
- boosting: try to run small/foreground tasks on higher-capacity CPUs to
complete them faster despite being less energy efficient.
- capping: try to run big/background tasks on low-capacity CPUs to save power
and thermal headroom for more important tasks
This series does not present this additional usage of utilization clamping but
it's an integral part of the EAS feature set, where [1] is one of its main
components.
Android kernels use SchedTune, a solution similar to utilization clamping, to
bias both 'frequency selection' and 'task placement'. This series provides the
foundation to add similar features to mainline while focusing, for the
time being, just on schedutil integration.
References
==========
[1] "Expressing per-task/per-cgroup performance hints"
Linux Plumbers Conference 2018
https://linuxplumbersconf.org/event/2/contributions/128/
[2] Message-ID: [off-list ref]
https://lore.kernel.org/lkml/20180911162827.GJ1100574@devbig004.ftw2.facebook.com/
[3] https://lore.kernel.org/lkml/20181203095628.11858-1-quentin.perret@arm.com/
Patrick Bellasi (16):
sched/core: Allow sched_setattr() to use the current policy
sched/core: uclamp: Extend sched_setattr() to support utilization
clamping
sched/core: uclamp: Map TASK's clamp values into CPU's clamp buckets
sched/core: uclamp: Add CPU's clamp buckets refcounting
sched/core: uclamp: Update CPU's refcount on clamp changes
sched/core: uclamp: Enforce last task UCLAMP_MAX
sched/core: uclamp: Add system default clamps
sched/cpufreq: uclamp: Add utilization clamping for FAIR tasks
sched/cpufreq: uclamp: Add utilization clamping for RT tasks
sched/core: Add uclamp_util_with()
sched/fair: Add uclamp support to energy_compute()
sched/core: uclamp: Extend CPU's cgroup controller
sched/core: uclamp: Propagate parent clamps
sched/core: uclamp: Map TG's clamp values into CPU's clamp buckets
sched/core: uclamp: Use TG's clamps to restrict TASK's clamps
sched/core: uclamp: Update CPU's refcount on TG's clamp changes
Documentation/admin-guide/cgroup-v2.rst | 46 ++
include/linux/log2.h | 37 +
include/linux/sched.h | 87 +++
include/linux/sched/sysctl.h | 11 +
include/linux/sched/task.h | 6 +
include/linux/sched/topology.h | 6 -
include/uapi/linux/sched.h | 12 +-
include/uapi/linux/sched/types.h | 65 +-
init/Kconfig | 75 ++
init/init_task.c | 1 +
kernel/exit.c | 1 +
kernel/sched/core.c | 947 +++++++++++++++++++++++-
kernel/sched/cpufreq_schedutil.c | 46 +-
kernel/sched/fair.c | 41 +-
kernel/sched/rt.c | 4 +
kernel/sched/sched.h | 136 +++-
kernel/sysctl.c | 16 +
17 files changed, 1480 insertions(+), 57 deletions(-)
--
2.19.2
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:15:34
The sched_setattr() syscall mandates that a policy is always specified.
This requires to always know which policy a task will have when
attributes are configured and it makes it impossible to add more generic
task attributes valid across different scheduling policies.
Reading the policy before setting generic tasks attributes is racy since
we cannot be sure it is not changed concurrently.
Introduce the required support to change generic task attributes without
affecting the current task policy. This is done by adding an attribute flag
(SCHED_FLAG_KEEP_POLICY) to enforce the usage of the current policy.
This is done by extending to the sched_setattr() non-POSIX syscall with
the SETPARAM_POLICY policy already used by the sched_setparam() POSIX
syscall.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
---
Changes in v6:
Message-ID: [off-list ref]
- rename SCHED_FLAG_TUNE_POLICY in SCHED_FLAG_KEEP_POLICY
- moved at the beginning of the series
---
include/uapi/linux/sched.h | 6 +++++-
kernel/sched/core.c | 11 ++++++++++-
2 files changed, 15 insertions(+), 2 deletions(-)
@@ -40,6 +40,8 @@/* SCHED_ISO: reserved but not implemented yet */#define SCHED_IDLE 5#define SCHED_DEADLINE 6+/* Must be the last entry: used to sanity check attr.policy values */+#define SCHED_POLICY_MAX 7/* Can be ORed in to make sure the process is reverted back to SCHED_NORMAL on fork */#define SCHED_RESET_ON_FORK 0x40000000
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:15:38
The SCHED_DEADLINE scheduling class provides an advanced and formal
model to define tasks requirements that can translate into proper
decisions for both task placements and frequencies selections. Other
classes have a more simplified model based on the POSIX concept of
priorities.
Such a simple priority based model however does not allow to exploit
most advanced features of the Linux scheduler like, for example, driving
frequencies selection via the schedutil cpufreq governor. However, also
for non SCHED_DEADLINE tasks, it's still interesting to define tasks
properties to support scheduler decisions.
Utilization clamping exposes to user-space a new set of per-task
attributes the scheduler can use as hints about the expected/required
utilization for a task. This allows to implement a "proactive" per-task
frequency control policy, a more advanced policy than the current one
based just on "passive" measured task utilization. For example, it's
possible to boost interactive tasks (e.g. to get better performance) or
cap background tasks (e.g. to be more energy/thermal efficient).
Introduce a new API to set utilization clamping values for a specified
task by extending sched_setattr(), a syscall which already allows to
define task specific properties for different scheduling classes. A new
pair of attributes allows to specify a minimum and maximum utilization
the scheduler can consider for a task.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
---
Changes in v6:
Message-ID: [off-list ref]
- add size check in sched_copy_attr()
Others:
- typos and changelog cleanups
---
include/linux/sched.h | 16 ++++++++
include/uapi/linux/sched.h | 4 +-
include/uapi/linux/sched/types.h | 65 +++++++++++++++++++++++++++-----
init/Kconfig | 21 +++++++++++
init/init_task.c | 5 +++
kernel/sched/core.c | 43 +++++++++++++++++++++
6 files changed, 144 insertions(+), 10 deletions(-)
@@ -640,6 +640,27 @@ config HAVE_UNSTABLE_SCHED_CLOCKconfigGENERIC_SCHED_CLOCKbool+menu"Scheduler features"++configUCLAMP_TASK+bool"Enable utilization clamping for RT/FAIR tasks"+depends onCPU_FREQ_GOV_SCHEDUTIL+help+Thisfeatureenablestheschedulertotracktheclampedutilization+ofeachCPUbasedonRUNNABLEtasksscheduledonthatCPU.++Withthisoption,theusercanspecifytheminandmaxCPU+utilizationallowedforRUNNABLEtasks.Themaxutilizationdefines+themaximumfrequencyataskshouldusewhiletheminutilization+definestheminimumfrequencyitshoulduse.++Bothminandmaxutilizationclampvaluesarehintstothescheduler,+aimingatimprovingitsfrequencyselectionpolicy,buttheydonot+enforceorgrantanyspecificbandwidthfortasks.++Ifindoubt,sayN.++endmenu## For architectures that want to enable the support for NUMA-affine scheduler# balancing logic:
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:15:41
Utilization clamping requires each CPU to know which clamp values are
assigned to tasks RUNNABLE on that CPU. A per-CPU array of reference
counters can be used where each entry tracks how many RUNNABLE tasks
require the same clamp value on each CPU. However, the range of clamp
values is too wide to track all the possible values in a per-CPU array.
Trade-off clamping precision for run-time and space efficiency using a
"bucketization and mapping" mechanism to translate "clamp values" into
"clamp buckets", each one representing a range of possible clamp values.
While the bucketization allows to use only a minimal set of clamp
buckets at run-time, the mapping ensures that the clamp buckets in use
are always at the beginning of the per-CPU array.
The minimum set of clamp buckets used at run-time depends on their
granularity and how many clamp values the target system expects to
use. Since on most systems we expect only a few different clamp
values, the bucketization and mapping mechanism increases our chances
to have all the required data fitting in one cache line.
For example, if we have only 20% and 25% clamped tasks, by setting:
CONFIG_UCLAMP_BUCKETS_COUNT 20
we allocate 20 clamp buckets with 5% resolution each, however we will
use only 2 of them at run-time, since their 5% resolution is enough to
always distinguish the clamp values in use, and they will both fit
into a single cache line for each CPU.
Introduce the "bucketization and mapping" mechanisms which are required
for the implement of the per-CPU operations.
Add a new "uclamp_enabled" sched_class attribute to mark which class
will contribute to clamping the CPU utilization. Move few callbacks
around to ensure that the most used callbacks are all in the same cache
line along with the new attribute.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
---
Changes in v6:
Message-ID: [off-list ref]
- added bucketization support since the beginning to avoid
semi-functional code in this patch
Message-ID: [off-list ref]
- update cmpxchg loops to use "do { } while (cmpxchg(ptr, old, new) != old)"
- switch to usage of try_cmpxchg()
Message-ID: [off-list ref]
- use SCHED_WARN_ON() instead of CONFIG_SCHED_DEBUG guarded blocks
- ensure se_count never underflow
Message-ID: <20181112000910.GC3038@worktop>
- wholesale s/group/bucket/
Message-ID: <20181111164754.GA3038@worktop>
- consistently use unary (++/--) operators
Message-ID: <20181107142428.GG14309@e110439-lin>
- added some better comments for invariant conditions
Message-ID: <20181107145612.GJ14309@e110439-lin>
- ensure UCLAMP_BUCKETS_COUNT >= 1
Others:
- added and make use of the bit_for() macro
- wholesale s/_{get,put}/_{inc,dec}/ to match refcount APIs
- documentation review and cleanup
---
include/linux/log2.h | 37 ++++++
include/linux/sched.h | 44 ++++++-
include/linux/sched/task.h | 6 +
include/linux/sched/topology.h | 6 -
include/uapi/linux/sched.h | 6 +-
init/Kconfig | 32 +++++
init/init_task.c | 4 -
kernel/exit.c | 1 +
kernel/sched/core.c | 234 ++++++++++++++++++++++++++++++---
kernel/sched/fair.c | 4 +
kernel/sched/sched.h | 19 ++-
11 files changed, 362 insertions(+), 31 deletions(-)
@@ -660,7 +660,39 @@ config UCLAMP_TASKIfindoubt,sayN.+configUCLAMP_BUCKETS_COUNT+int"Number of supported utilization clamp buckets"+range520+default5+depends onUCLAMP_TASK+help+Definesthenumberofclampbucketstouse.Therangeofeachbucket+willbeSCHED_CAPACITY_SCALE/UCLAMP_BUCKETS_COUNT.Thehigherthe+numberofclampbucketsthefinertheirgranularityandthehigher+theprecisionofclampingaggregationandtrackingatrun-time.++Forexample,withthedefaultconfigurationwewillhave5clamp+bucketstracking20%utilizationeach.A25%boostedtaskswillbe+refcountedinthe[20..39]%bucketandwillsetthebucketclamp+effectivevalueto25%.+Ifasecond30%boostedtaskshouldbeco-scheduledonthesameCPU,+thattaskwillberefcountedinthesamebucketofthefirsttaskand+itwillboostthebucketclampeffectivevalueto30%.+Theclampeffectivevalueofabucketisresettoitsnominalvalue+(20%intheexampleabove)whenthereareanymoretasksrefcountedin+thatbucket.++Anadditionalboost/cappingmargincanbeaddedtosometasks.Inthe+exampleabovethe25%taskwillbeboostedto30%untilitexitsthe+CPU.Ifthatshouldbeconsiderednotacceptableoncertainsystems,+it'salwayspossibletoreducethemarginbyincreasingthenumberof+clampbucketstotradeoffusedmemoryforrun-timetracking+precision.++Ifindoubt,usethedefaultvalue.+endmenu+## For architectures that want to enable the support for NUMA-affine scheduler# balancing logic:
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:15:44
Utilization clamping allows to clamp the CPU's utilization within a
[util_min, util_max] range, depending on the set of RUNNABLE tasks on
that CPU. Each task references two "clamp buckets" defining its minimum
and maximum (util_{min,max}) utilization "clamp values". A CPU's clamp
bucket is active if there is at least one RUNNABLE tasks enqueued on
that CPU and refcounting that bucket.
When a task is {en,de}queued {on,from} a CPU, the set of active clamp
buckets on that CPU can change. Since each clamp bucket enforces a
different utilization clamp value, when the set of active clamp buckets
changes, a new "aggregated" clamp value is computed for that CPU.
Clamp values are always MAX aggregated for both util_min and util_max.
This ensures that no tasks can affect the performance of other
co-scheduled tasks which are more boosted (i.e. with higher util_min
clamp) or less capped (i.e. with higher util_max clamp).
Each task has a:
task_struct::uclamp[clamp_id]::bucket_id
to track the "bucket index" of the CPU's clamp bucket it refcounts while
enqueued, for each clamp index (clamp_id).
Each CPU's rq has a:
rq::uclamp[clamp_id]::bucket[bucket_id].tasks
to track how many tasks, currently RUNNABLE on that CPU, refcount each
clamp bucket (bucket_id) of a clamp index (clamp_id).
Each CPU's rq has also a:
rq::uclamp[clamp_id]::bucket[bucket_id].value
to track the clamp value of each clamp bucket (bucket_id) of a clamp
index (clamp_id).
The unordered array rq::uclamp::bucket[clamp_id][] is scanned every time
we need to find a new MAX aggregated clamp value for a clamp_id. This
operation is required only when we dequeue the last task of a clamp
bucket tracking the current MAX aggregated clamp value. In these cases,
the CPU is either entering IDLE or going to schedule a less boosted or
more clamped task.
The expected number of different clamp values, configured at build time,
is small enough to fit the full unordered array into a single cache
line. In most use-cases we expect less than 10 different clamp values
for each clamp_id.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
---
Changes in v6:
Message-ID: <20181113151127.GA7681@darkstar>
- use SCHED_WARN_ON() instead of CONFIG_SCHED_DEBUG guarded WARN()s
- add some better inline documentation to explain per-CPU initializations
- add some local variables to use library's max() for aggregation on
bitfields attirbutes
Message-ID: <20181112000910.GC3038@worktop>
- wholesale s/group/bucket/
Message-ID: <20181111164754.GA3038@worktop>
- consistently use unary (++/--) operators
Others:
- updated from rq::uclamp::group[clamp_id][group_id]
to rq::uclamp[clamp_id]::bucket[bucket_id]
which better matches the layout already used for tasks, i.e.
p::uclamp[clamp_id]::value
- use {WRITE,READ}_ONCE() for rq's clamp access
- update layout of rq::uclamp_cpu to better match that of tasks,
i.e now access CPU's clamp buckets as:
rq->uclamp[clamp_id]{.bucket[bucket_id].value}
which matches:
p->uclamp[clamp_id]
---
include/linux/sched.h | 6 ++
kernel/sched/core.c | 152 ++++++++++++++++++++++++++++++++++++++++++
kernel/sched/sched.h | 49 ++++++++++++++
3 files changed, 207 insertions(+)
@@ -766,6 +766,124 @@ static inline unsigned int uclamp_bucket_value(unsigned int clamp_value)returnUCLAMP_BUCKET_DELTA*(clamp_value/UCLAMP_BUCKET_DELTA);}+staticinlinevoiduclamp_cpu_update(structrq*rq,unsignedintclamp_id)+{+unsignedintmax_value=0;+unsignedintbucket_id;++for(bucket_id=0;bucket_id<UCLAMP_BUCKETS;++bucket_id){+unsignedintbucket_value;++if(!rq->uclamp[clamp_id].bucket[bucket_id].tasks)+continue;++/* Both min and max clamps are MAX aggregated */+bucket_value=rq->uclamp[clamp_id].bucket[bucket_id].value;+max_value=max(max_value,bucket_value);+if(max_value>=SCHED_CAPACITY_SCALE)+break;+}+WRITE_ONCE(rq->uclamp[clamp_id].value,max_value);+}++/*+*WhenataskisenqueuedonaCPU'srq,theclampbucketcurrentlydefinedby+*thetask'suclamp::bucket_idisreferencecountedonthatCPU.Thisalso+*immediatelyupdatestheCPU'sclampvalueifrequired.+*+*Sincetasksknowtheirspecificvaluerequestedfromuser-space,wetrack+*withineachbucketthemaximumvaluefortasksrefcountedinthatbucket.+*/+staticinlinevoiduclamp_cpu_inc_id(structtask_struct*p,structrq*rq,+unsignedintclamp_id)+{+unsignedintcpu_clamp,grp_clamp,tsk_clamp;+unsignedintbucket_id;++if(unlikely(!p->uclamp[clamp_id].mapped))+return;++bucket_id=p->uclamp[clamp_id].bucket_id;+p->uclamp[clamp_id].active=true;++rq->uclamp[clamp_id].bucket[bucket_id].tasks++;++/* CPU's clamp buckets track the max effective clamp value */+tsk_clamp=p->uclamp[clamp_id].value;+grp_clamp=rq->uclamp[clamp_id].bucket[bucket_id].value;+rq->uclamp[clamp_id].bucket[bucket_id].value=max(grp_clamp,tsk_clamp);++/* Update CPU clamp value if required */+cpu_clamp=READ_ONCE(rq->uclamp[clamp_id].value);+WRITE_ONCE(rq->uclamp[clamp_id].value,max(cpu_clamp,tsk_clamp));+}++/*+*WhenataskisdequeuedfromaCPU'srq,theCPU'sclampbucketreference+*countedbythetaskisreleased.Ifthisisthelasttaskreference+*countingtheCPU'smaxactiveclampvalue,thentheCPU'sclampvalueis+*updated.+*BoththetasksreferencecounterandtheCPU'scachedclampvaluesare+*expectedtobealwaysvalid,ifwedetecttheyarenotweskiptheupdates,+*enforceaconsistentstateandwarn.+*/+staticinlinevoiduclamp_cpu_dec_id(structtask_struct*p,structrq*rq,+unsignedintclamp_id)+{+unsignedintclamp_value;+unsignedintbucket_id;++if(unlikely(!p->uclamp[clamp_id].mapped))+return;++bucket_id=p->uclamp[clamp_id].bucket_id;+p->uclamp[clamp_id].active=false;++SCHED_WARN_ON(!rq->uclamp[clamp_id].bucket[bucket_id].tasks);+if(likely(rq->uclamp[clamp_id].bucket[bucket_id].tasks))+rq->uclamp[clamp_id].bucket[bucket_id].tasks--;++/* We accept to (possibly) overboost tasks still RUNNABLE */+if(likely(rq->uclamp[clamp_id].bucket[bucket_id].tasks))+return;+clamp_value=rq->uclamp[clamp_id].bucket[bucket_id].value;++/* The CPU's clamp value is expected to always track the max */+SCHED_WARN_ON(clamp_value>rq->uclamp[clamp_id].value);++if(clamp_value>=READ_ONCE(rq->uclamp[clamp_id].value)){+/*+*ResetCPU'sclampbucketvaluetoitsnominalvaluewhenever+*thereareanymoreRUNNABLEtasksrefcountingit.+*/+rq->uclamp[clamp_id].bucket[bucket_id].value=+uclamp_maps[clamp_id][bucket_id].value;+uclamp_cpu_update(rq,clamp_id);+}+}++staticinlinevoiduclamp_cpu_inc(structrq*rq,structtask_struct*p)+{+unsignedintclamp_id;++if(unlikely(!p->sched_class->uclamp_enabled))+return;++for(clamp_id=0;clamp_id<UCLAMP_CNT;++clamp_id)+uclamp_cpu_inc_id(p,rq,clamp_id);+}++staticinlinevoiduclamp_cpu_dec(structrq*rq,structtask_struct*p)+{+unsignedintclamp_id;++if(unlikely(!p->sched_class->uclamp_enabled))+return;++for(clamp_id=0;clamp_id<UCLAMP_CNT;++clamp_id)+uclamp_cpu_dec_id(p,rq,clamp_id);+}+staticvoiduclamp_bucket_dec(unsignedintclamp_id,unsignedintbucket_id){unionuclamp_map*uc_maps=&uclamp_maps[clamp_id][0];
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:15:49
Utilization clamp values enforced on a CPU by a task can be updated, for
example via a sched_setattr() syscall, while a task is RUNNABLE on that
CPU. A clamp value change always implies a clamp bucket refcount update
to ensure the new constraints are enforced.
Hook into uclamp_bucket_get() to trigger a CPU refcount syncup, via
uclamp_cpu_{inc,dec}_id(), whenever a task is RUNNABLE.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
---
Changes in v6:
Other:
- wholesale s/group/bucket/
- wholesale s/_{get,put}/_{inc,dec}/ to match refcount APIs
- small documentation updates
---
kernel/sched/core.c | 48 +++++++++++++++++++++++++++++++++++++++------
1 file changed, 42 insertions(+), 6 deletions(-)
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:15:53
When the task sleeps, it removes its max utilization clamp from its CPU.
However, the blocked utilization on that CPU can be higher than the max
clamp value enforced while the task was running. This allows undesired
CPU frequency increases while a CPU is idle, for example, when another
CPU on the same frequency domain triggers a frequency update, since
schedutil can now see the full not clamped blocked utilization of the
idle CPU.
Fix this by using
uclamp_cpu_dec_id(p, rq, UCLAMP_MAX)
uclamp_cpu_update(rq, UCLAMP_MAX, clamp_value)
to detect when a CPU has no more RUNNABLE clamped tasks and to flag this
condition.
Don't track any minimum utilization clamps since an idle CPU never
requires a minimum frequency. The decay of the blocked utilization is
good enough to reduce the CPU frequency.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
---
Changes in v6:
Others:
- moved UCLAMP_FLAG_IDLE management into dedicated functions:
uclamp_idle_value() and uclamp_idle_reset()
- switched from rq::uclamp::flags to rq::uclamp_flags, since now
rq::uclamp is a per-clamp_id array
---
kernel/sched/core.c | 51 +++++++++++++++++++++++++++++++++++++++++---
kernel/sched/sched.h | 2 ++
2 files changed, 50 insertions(+), 3 deletions(-)
@@ -766,9 +766,45 @@ static inline unsigned int uclamp_bucket_value(unsigned int clamp_value)returnUCLAMP_BUCKET_DELTA*(clamp_value/UCLAMP_BUCKET_DELTA);}-staticinlinevoiduclamp_cpu_update(structrq*rq,unsignedintclamp_id)+staticinlineunsignedint+uclamp_idle_value(structrq*rq,unsignedintclamp_id,unsignedintclamp_value)+{+/*+*Avoidblockedutilizationpushingupthefrequencywhenwego+*idle(whichdropsthemax-clamp)byretainingthelastknown+*max-clamp.+*/+if(clamp_id==UCLAMP_MAX){+rq->uclamp_flags|=UCLAMP_FLAG_IDLE;+returnclamp_value;+}++returnuclamp_none(UCLAMP_MIN);+}++staticinlinevoiduclamp_idle_reset(structrq*rq,unsignedintclamp_id,+unsignedintclamp_value)+{+/* Reset max-clamp retention only on idle exit */+if(!(rq->uclamp_flags&UCLAMP_FLAG_IDLE))+return;++WRITE_ONCE(rq->uclamp[clamp_id].value,clamp_value);++/*+*ThisfunctioniscalledforbothUCLAMP_MIN(before)andUCLAMP_MAX+*(after).Theidleflagisresetonlythesecondtime,whenweknow+*thatUCLAMP_MINhasbeenalreadyupdated.+*/+if(clamp_id==UCLAMP_MAX)+rq->uclamp_flags&=~UCLAMP_FLAG_IDLE;+}++staticinlinevoiduclamp_cpu_update(structrq*rq,unsignedintclamp_id,+unsignedintclamp_value){unsignedintmax_value=0;+boolbuckets_active=false;unsignedintbucket_id;for(bucket_id=0;bucket_id<UCLAMP_BUCKETS;++bucket_id){
@@ -776,6 +812,7 @@ static inline void uclamp_cpu_update(struct rq *rq, unsigned int clamp_id)if(!rq->uclamp[clamp_id].bucket[bucket_id].tasks)continue;+buckets_active=true;/* Both min and max clamps are MAX aggregated */bucket_value=rq->uclamp[clamp_id].bucket[bucket_id].value;
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:15:58
Each time a frequency update is required via schedutil, a frequency is
selected to (possibly) satisfy the utilization reported by each
scheduling class. However, when utilization clamping is in use, the
frequency selection should consider userspace utilization clamping
hints. This will allow, for example, to:
- boost tasks which are directly affecting the user experience
by running them at least at a minimum "requested" frequency
- cap low priority tasks not directly affecting the user experience
by running them only up to a maximum "allowed" frequency
These constraints are meant to support a per-task based tuning of the
frequency selection thus supporting a fine grained definition of
performance boosting vs energy saving strategies in kernel space.
Add support to clamp the utilization and IOWait boost of RUNNABLE FAIR
tasks within the boundaries defined by their aggregated utilization
clamp constraints.
Based on the max(min_util, max_util) of each task, max-aggregated the
CPU clamp value in a way to give the boosted tasks the performance they
need when they happen to be co-scheduled with other capped tasks.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Rafael J. Wysocki <redacted>
---
Changes in v6:
Message-ID: <20181107113849.GC14309@e110439-lin>
- sanity check util_max >= util_min
Others:
- wholesale s/group/bucket/
- wholesale s/_{get,put}/_{inc,dec}/ to match refcount APIs
---
kernel/sched/cpufreq_schedutil.c | 27 ++++++++++++++++++++++++---
kernel/sched/sched.h | 23 +++++++++++++++++++++++
2 files changed, 47 insertions(+), 3 deletions(-)
@@ -218,8 +218,15 @@ unsigned long schedutil_freq_util(int cpu, unsigned long util_cfs,*CFStasksandweusethesamemetrictotracktheeffective*utilization(PELTwindowsaresynchronized)wecandirectlyaddthem*toobtaintheCPU'sactualutilization.+*+*CFSutilizationcanbeboostedorcapped,dependingonutilization+*clampconstraintsrequestedbycurrentlyRUNNABLEtasks.+*WhentherearenoCFSRUNNABLEtasks,clampsarereleasedand+*frequencywillbegracefullyreducedwiththeutilizationdecay.*/-util=util_cfs;+util=(type==ENERGY_UTIL)+?util_cfs+:uclamp_util(rq,util_cfs);util+=cpu_util_rt(rq);dl_util=cpu_util_dl(rq);
@@ -327,6 +334,7 @@ static void sugov_iowait_boost(struct sugov_cpu *sg_cpu, u64 time,unsignedintflags){boolset_iowait_boost=flags&SCHED_CPUFREQ_IOWAIT;+unsignedintmax_boost;/* Reset boost if the CPU appears to have been idle enough */if(sg_cpu->iowait_boost&&
@@ -342,11 +350,24 @@ static void sugov_iowait_boost(struct sugov_cpu *sg_cpu, u64 time,return;sg_cpu->iowait_boost_pending=true;+/*+*BoostFAIRtasksonlyuptotheCPUclampedutilization.+*+*SinceDLtaskshaveamuchmoreadvancedbandwidthcontrol,it's+*safetoassumethatIOboostdoesnotapplytothosetasks.+*Instead,sinceRTtasksarenotutilizationclamped,wedon'twant+*toapplyclampingonIOboostwhilethereisblockedRT+*utilization.+*/+max_boost=sg_cpu->iowait_boost_max;+if(!cpu_util_rt(cpu_rq(sg_cpu->cpu)))+max_boost=uclamp_util(cpu_rq(sg_cpu->cpu),max_boost);+/* Double the boost at each request */if(sg_cpu->iowait_boost){sg_cpu->iowait_boost<<=1;-if(sg_cpu->iowait_boost>sg_cpu->iowait_boost_max)-sg_cpu->iowait_boost=sg_cpu->iowait_boost_max;+if(sg_cpu->iowait_boost>max_boost)+sg_cpu->iowait_boost=max_boost;return;}
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:16:02
Schedutil enforces a maximum frequency when RT tasks are RUNNABLE.
This mandatory policy can be made tunable from userspace to define a max
frequency which is still reasonable for the execution of a specific RT
workload while being also power/energy friendly.
Extend the usage of util_{min,max} to the RT scheduling class.
Add uclamp_default_perf, a special set of clamp values to be used
for tasks requiring maximum performance, i.e. by default all the non
clamped RT tasks.
Since utilization clamping applies now to both CFS and RT tasks,
schedutil clamps the combined utilization of these two classes.
The IOWait boost value is also subject to clamping for RT tasks.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Rafael J. Wysocki <redacted>
---
Changes in v6:
Others:
- wholesale s/group/bucket/
- wholesale s/_{get,put}/_{inc,dec}/ to match refcount APIs
---
kernel/sched/core.c | 20 ++++++++++++++++----
kernel/sched/cpufreq_schedutil.c | 27 +++++++++++++--------------
kernel/sched/rt.c | 4 ++++
3 files changed, 33 insertions(+), 18 deletions(-)
@@ -746,6 +746,7 @@ unsigned int sysctl_sched_uclamp_util_max = SCHED_CAPACITY_SCALE;*Tasksspecificclampvaluesarerequiredtobewithinthisrange*/staticstructuclamp_seuclamp_default[UCLAMP_CNT];+staticstructuclamp_seuclamp_default_perf[UCLAMP_CNT];/***Referencecountutilizationclampbuckets
@@ -858,16 +859,23 @@ static inline voiduclamp_effective_get(structtask_struct*p,unsignedintclamp_id,unsignedint*clamp_value,unsignedint*bucket_id){+structuclamp_se*default_clamp;+/* Task specific clamp value */*clamp_value=p->uclamp[clamp_id].value;*bucket_id=p->uclamp[clamp_id].bucket_id;+/* RT tasks have different default values */+default_clamp=task_has_rt_policy(p)+?uclamp_default_perf+:uclamp_default;+/* System default restriction */-if(unlikely(*clamp_value<uclamp_default[UCLAMP_MIN].value||-*clamp_value>uclamp_default[UCLAMP_MAX].value)){+if(unlikely(*clamp_value<default_clamp[UCLAMP_MIN].value||+*clamp_value>default_clamp[UCLAMP_MAX].value)){/* Keep it simple: unconditionally enforce system defaults */-*clamp_value=uclamp_default[clamp_id].value;-*bucket_id=uclamp_default[clamp_id].bucket_id;+*clamp_value=default_clamp[clamp_id].value;+*bucket_id=default_clamp[clamp_id].bucket_id;}}
@@ -1282,6 +1290,10 @@ static void __init init_uclamp(void)uc_se=&uclamp_default[clamp_id];uclamp_bucket_inc(NULL,uc_se,clamp_id,uclamp_none(clamp_id));++/* RT tasks by default will go to max frequency */+uc_se=&uclamp_default_perf[clamp_id];+uclamp_bucket_inc(NULL,uc_se,clamp_id,uclamp_none(UCLAMP_MAX));}}
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:16:07
Currently uclamp_util() allows to clamp a specified utilization
considering the clamp values requested by RUNNABLE tasks in a CPU.
Sometimes however, it could be interesting to verify how clamp values
will change when a task is going to be running on a given CPU.
For example, the Energy Aware Scheduler (EAS) is interested in
evaluating and comparing the energy impact of different scheduling
decisions.
Add uclamp_util_with() which allows to clamp a given utilization by
considering the possible impact on CPU clamp values of a specified task.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
---
kernel/sched/core.c | 4 ++--
kernel/sched/sched.h | 21 ++++++++++++++++++++-
2 files changed, 22 insertions(+), 3 deletions(-)
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:16:11
The Energy Aware Scheduler (AES) estimates the energy impact of waking
up a task on a given CPU. This estimation is based on:
a) an (active) power consumptions defined for each CPU frequency
b) an estimation of which frequency will be used on each CPU
c) an estimation of the busy time (utilization) of each CPU
Utilization clamping can affect both b) and c) estimations. A CPU is
expected to run:
- on an higher than required frequency, but for a shorter time, in case
its estimated utilization will be smaller then the minimum utilization
enforced by uclamp
- on a smaller than required frequency, but for a longer time, in case
its estimated utilization is bigger then the maximum utilization
enforced by uclamp
While effects on busy time for both boosted/capped tasks are already
considered by compute_energy(), clamping effects on frequency selection
are currently ignored by that function.
Fix it by considering how CPU clamp values will be affected by a
task waking up and being RUNNABLE on that CPU.
Do that by refactoring schedutil_freq_util() to take an additional
task_struct* which allows EAS to evaluate the impact on clamp values of
a task being eventually queued in a CPU. Clamp values are applied to the
RT+CFS utilization only when a FREQUENCY_UTIL is required by
compute_energy().
Since we are at that:
- rename schedutil_freq_util() into schedutil_cpu_util(),
since it's not only used for frequency selection.
- use "unsigned int" instead of "unsigned long" whenever the tracked
utilization value is not expected to overflow 32bit.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Rafael J. Wysocki <redacted>
---
kernel/sched/cpufreq_schedutil.c | 26 +++++++++++-----------
kernel/sched/fair.c | 37 ++++++++++++++++++++++++++------
kernel/sched/sched.h | 19 +++++-----------
3 files changed, 48 insertions(+), 34 deletions(-)
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:16:15
In order to properly support hierarchical resources control, the cgroup
delegation model requires that attribute writes from a child group never
fail but still are (potentially) constrained based on parent's assigned
resources. This requires to properly propagate and aggregate parent
attributes down to its descendants.
Let's implement this mechanism by adding a new "effective" clamp value
for each task group. The effective clamp value is defined as the smaller
value between the clamp value of a group and the effective clamp value
of its parent. This is the actual clamp value enforced on tasks in a
task group.
Since it can be interesting for userspace, e.g. system management
software, to know exactly what the currently propagated/enforced
configuration is, the effective clamp values are exposed to user-space
by means of a new pair of read-only attributes
cpu.util.{min,max}.effective.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Tejun Heo <tj@kernel.org>
---
Changes in v6:
Others:
- wholesale s/group/bucket/
- wholesale s/_{get,put}/_{inc,dec}/ to match refcount APIs
---
Documentation/admin-guide/cgroup-v2.rst | 25 ++++++-
include/linux/sched.h | 10 ++-
kernel/sched/core.c | 89 +++++++++++++++++++++++--
3 files changed, 117 insertions(+), 7 deletions(-)
@@ -984,22 +984,43 @@ All time durations are in microseconds. A read-write single value file which exists on non-root cgroups. The default is "0", i.e. no utilization boosting.- The minimum utilization in the range [0, 1024].+ The requested minimum utilization in the range [0, 1024]. This interface allows reading and setting minimum utilization clamp values similar to the sched_setattr(2). This minimum utilization value is used to clamp the task specific minimum utilization clamp.+ cpu.util.min.effective+ A read-only single value file which exists on non-root cgroups and+ reports minimum utilization clamp value currently enforced on a task+ group.++ The actual minimum utilization in the range [0, 1024].++ This value can be lower then cpu.util.min in case a parent cgroup+ allows only smaller minimum utilization values.+ cpu.util.max A read-write single value file which exists on non-root cgroups. The default is "1024". i.e. no utilization capping- The maximum utilization in the range [0, 1024].+ The requested maximum utilization in the range [0, 1024]. This interface allows reading and setting maximum utilization clamp values similar to the sched_setattr(2). This maximum utilization value is used to clamp the task specific maximum utilization clamp.+ cpu.util.max.effective+ A read-only single value file which exists on non-root cgroups and+ reports maximum utilization clamp value currently enforced on a task+ group.++ The actual maximum utilization in the range [0, 1024].++ This value can be lower then cpu.util.max in case a parent cgroup+ is enforcing a more restrictive clamping on max utilization.++ Memory ------
@@ -625,7 +625,15 @@ struct uclamp_se {unsignedintbucket_id:bits_per(UCLAMP_BUCKETS);unsignedintmapped:1;unsignedintactive:1;-/* Clamp bucket and value actually used by a RUNNABLE task */+/*+*Clampbucketandvalueactuallyusedbyaschedulingentity,+*i.e.a(RUNNABLE)taskorataskgroup.+*Fortaskgroups,thisisthevalue(possibly)enforcedbya+*parenttaskgroup.+*Foratask,thisisthevalue(possibly)enforcedbythe+*taskgroupthetaskiscurrentlypartoforbythesystem+*defaultclampvalues,whicheveristhemostrestrictive.+*/struct{unsignedintvalue:bits_per(SCHED_CAPACITY_SCALE);unsignedintbucket_id:bits_per(UCLAMP_BUCKETS);
@@ -6890,6 +6891,8 @@ static inline int alloc_uclamp_sched_group(struct task_group *tg,parent->uclamp[clamp_id].value;tg->uclamp[clamp_id].bucket_id=parent->uclamp[clamp_id].bucket_id;+tg->uclamp[clamp_id].effective.value=+parent->uclamp[clamp_id].effective.value;}#endif
@@ -7143,6 +7146,45 @@ static void cpu_cgroup_attach(struct cgroup_taskset *tset)}#ifdef CONFIG_UCLAMP_TASK_GROUP+staticvoidcpu_util_update_hier(structcgroup_subsys_state*css,+intclamp_id,unsignedintvalue)+{+structcgroup_subsys_state*top_css=css;+structuclamp_se*uc_se,*uc_parent;++css_for_each_descendant_pre(css,top_css){+/*+*Thefirstvisitedtaskgroupistop_css,whichclampvalue+*istheonepassedasparameter.Fordescendenttask+*groupsweconsidertheircurrentvalue.+*/+uc_se=&css_tg(css)->uclamp[clamp_id];+if(css!=top_css)+value=uc_se->value;++/*+*Skipthewholesubtreesifthecurrenteffectiveclampis+*alreadymatchingtheTG'sclampvalue.+*Inthiscase,allthesubtreesalreadyhavetop_value,ora+*morerestrictivevalue,aseffectiveclamp.+*/+uc_parent=&css_tg(css)->parent->uclamp[clamp_id];+if(uc_se->effective.value==value&&+uc_parent->effective.value>=value){+css=css_rightmost_descendant(css);+continue;+}++/* Propagate the most restrictive effective value */+if(uc_parent->effective.value<value)+value=uc_parent->effective.value;+if(uc_se->effective.value==value)+continue;++uc_se->effective.value=value;+}+}+staticintcpu_util_min_write_u64(structcgroup_subsys_state*css,structcftype*cftype,u64min_value){
@@ -7162,6 +7204,9 @@ static int cpu_util_min_write_u64(struct cgroup_subsys_state *css,gotoout;}+/* Update effective clamps to track the most restrictive value */+cpu_util_update_hier(css,UCLAMP_MIN,min_value);+out:rcu_read_unlock();
@@ -7187,6 +7232,9 @@ static int cpu_util_max_write_u64(struct cgroup_subsys_state *css,gotoout;}+/* Update effective clamps to track the most restrictive value */+cpu_util_update_hier(css,UCLAMP_MAX,max_value);+out:rcu_read_unlock();
@@ -7194,14 +7242,17 @@ static int cpu_util_max_write_u64(struct cgroup_subsys_state *css,}staticinlineu64cpu_uclamp_read(structcgroup_subsys_state*css,-enumuclamp_idclamp_id)+enumuclamp_idclamp_id,+booleffective){structtask_group*tg;u64util_clamp;rcu_read_lock();tg=css_tg(css);-util_clamp=tg->uclamp[clamp_id].value;+util_clamp=effective+?tg->uclamp[clamp_id].effective.value+:tg->uclamp[clamp_id].value;rcu_read_unlock();returnutil_clamp;
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:16:20
Utilization clamping requires to map each different clamp value into one
of the available clamp buckets used at {en,de}queue time (fast-path).
Each time a TG's clamp value sysfs attribute is updated via:
cpu_util_{min,max}_write_u64()
we need to update the task group reference to the new value's clamp
bucket and release the reference to the previous one.
Ensure that, whenever a task group is assigned a specific clamp_value,
this is properly translated into a unique clamp bucket to be used in the
fast-path. Do it by slightly refactoring uclamp_bucket_inc() to make the
(*task_struct) parameter optional and by reusing the code already
available for the per-task API.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Tejun Heo <tj@kernel.org>
---
Changes in v6:
Others:
- wholesale s/group/bucket/
- wholesale s/_{get,put}/_{inc,dec}/ to match refcount APIs
---
include/linux/sched.h | 4 ++--
kernel/sched/core.c | 53 +++++++++++++++++++++++++++++++++----------
2 files changed, 43 insertions(+), 14 deletions(-)
@@ -7176,12 +7190,15 @@ static void cpu_util_update_hier(struct cgroup_subsys_state *css,}/* Propagate the most restrictive effective value */-if(uc_parent->effective.value<value)+if(uc_parent->effective.value<value){value=uc_parent->effective.value;+bucket_id=uc_parent->effective.bucket_id;+}if(uc_se->effective.value==value)continue;uc_se->effective.value=value;+uc_se->effective.bucket_id=bucket_id;}}
@@ -7194,6 +7211,7 @@ static int cpu_util_min_write_u64(struct cgroup_subsys_state *css,if(min_value>SCHED_CAPACITY_SCALE)return-ERANGE;+mutex_lock(&uclamp_mutex);rcu_read_lock();tg=css_tg(css);
@@ -7204,11 +7222,16 @@ static int cpu_util_min_write_u64(struct cgroup_subsys_state *css,gotoout;}+/* Update TG's reference count */+uclamp_bucket_inc(NULL,&tg->uclamp[UCLAMP_MIN],UCLAMP_MIN,min_value);+/* Update effective clamps to track the most restrictive value */-cpu_util_update_hier(css,UCLAMP_MIN,min_value);+cpu_util_update_hier(css,UCLAMP_MIN,tg->uclamp[UCLAMP_MIN].bucket_id,+min_value);out:rcu_read_unlock();+mutex_unlock(&uclamp_mutex);returnret;}
@@ -7222,6 +7245,7 @@ static int cpu_util_max_write_u64(struct cgroup_subsys_state *css,if(max_value>SCHED_CAPACITY_SCALE)return-ERANGE;+mutex_lock(&uclamp_mutex);rcu_read_lock();tg=css_tg(css);
@@ -7232,11 +7256,16 @@ static int cpu_util_max_write_u64(struct cgroup_subsys_state *css,gotoout;}+/* Update TG's reference count */+uclamp_bucket_inc(NULL,&tg->uclamp[UCLAMP_MAX],UCLAMP_MAX,max_value);+/* Update effective clamps to track the most restrictive value */-cpu_util_update_hier(css,UCLAMP_MAX,max_value);+cpu_util_update_hier(css,UCLAMP_MAX,tg->uclamp[UCLAMP_MAX].bucket_id,+max_value);out:rcu_read_unlock();+mutex_unlock(&uclamp_mutex);returnret;}
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:16:24
On updates of task group (TG) clamp values, ensure that these new values
are enforced on all RUNNABLE tasks of the task group, i.e. all RUNNABLE
tasks are immediately boosted and/or clamped as requested.
Do that by slightly refactoring uclamp_bucket_inc(). An additional
parameter *cgroup_subsys_state (css) is used to walk the list of tasks
in the TGs and update the RUNNABLE ones. Do that by taking the rq
lock for each task, the same mechanism used for cpu affinity masks
updates.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Tejun Heo <tj@kernel.org>
---
Changes in v6:
Others:
- wholesale s/group/bucket/
- wholesale s/_{get,put}/_{inc,dec}/ to match refcount APIs
- small documentation updates
---
kernel/sched/core.c | 56 +++++++++++++++++++++++++++++++++------------
1 file changed, 42 insertions(+), 14 deletions(-)
@@ -1111,7 +1111,22 @@ static void uclamp_bucket_dec(unsigned int clamp_id, unsigned int bucket_id)&uc_map_old.data,uc_map_new.data));}-staticvoiduclamp_bucket_inc(structtask_struct*p,structuclamp_se*uc_se,+staticinlinevoiduclamp_bucket_inc_tg(structcgroup_subsys_state*css,+intclamp_id,unsignedintbucket_id)+{+structcss_task_iterit;+structtask_struct*p;++/* Update clamp buckets for RUNNABLE tasks in this TG */+css_task_iter_start(css,0,&it);+while((p=css_task_iter_next(&it)))+uclamp_task_update_active(p,clamp_id);+css_task_iter_end(&it);+}++staticvoiduclamp_bucket_inc(structtask_struct*p,+structcgroup_subsys_state*css,+structuclamp_se*uc_se,unsignedintclamp_id,unsignedintclamp_value){unionuclamp_map*uc_maps=&uclamp_maps[clamp_id][0];
@@ -1221,11 +1239,11 @@ int sched_uclamp_handler(struct ctl_table *table, int write,}if(old_min!=sysctl_sched_uclamp_util_min){-uclamp_bucket_inc(NULL,&uclamp_default[UCLAMP_MIN],+uclamp_bucket_inc(NULL,NULL,&uclamp_default[UCLAMP_MIN],UCLAMP_MIN,sysctl_sched_uclamp_util_min);}if(old_max!=sysctl_sched_uclamp_util_max){-uclamp_bucket_inc(NULL,&uclamp_default[UCLAMP_MAX],+uclamp_bucket_inc(NULL,NULL,&uclamp_default[UCLAMP_MAX],UCLAMP_MAX,sysctl_sched_uclamp_util_max);}gotodone;
@@ -1260,12 +1278,12 @@ static int __setscheduler_uclamp(struct task_struct *p,mutex_lock(&uclamp_mutex);if(attr->sched_flags&SCHED_FLAG_UTIL_CLAMP_MIN){p->uclamp[UCLAMP_MIN].user_defined=true;-uclamp_bucket_inc(p,&p->uclamp[UCLAMP_MIN],+uclamp_bucket_inc(p,NULL,&p->uclamp[UCLAMP_MIN],UCLAMP_MIN,lower_bound);}if(attr->sched_flags&SCHED_FLAG_UTIL_CLAMP_MAX){p->uclamp[UCLAMP_MAX].user_defined=true;-uclamp_bucket_inc(p,&p->uclamp[UCLAMP_MAX],+uclamp_bucket_inc(p,NULL,&p->uclamp[UCLAMP_MAX],UCLAMP_MAX,upper_bound);}mutex_unlock(&uclamp_mutex);
@@ -1326,19 +1344,23 @@ static void __init init_uclamp(void)memset(uclamp_maps,0,sizeof(uclamp_maps));for(clamp_id=0;clamp_id<UCLAMP_CNT;++clamp_id){uc_se=&init_task.uclamp[clamp_id];-uclamp_bucket_inc(NULL,uc_se,clamp_id,uclamp_none(clamp_id));+uclamp_bucket_inc(NULL,NULL,uc_se,clamp_id,+uclamp_none(clamp_id));uc_se=&uclamp_default[clamp_id];-uclamp_bucket_inc(NULL,uc_se,clamp_id,uclamp_none(clamp_id));+uclamp_bucket_inc(NULL,NULL,uc_se,clamp_id,+uclamp_none(clamp_id));/* RT tasks by default will go to max frequency */uc_se=&uclamp_default_perf[clamp_id];-uclamp_bucket_inc(NULL,uc_se,clamp_id,uclamp_none(UCLAMP_MAX));+uclamp_bucket_inc(NULL,NULL,uc_se,clamp_id,+uclamp_none(UCLAMP_MAX));#ifdef CONFIG_UCLAMP_TASK_GROUP/* Init root TG's clamp bucket */uc_se=&root_task_group.uclamp[clamp_id];-uclamp_bucket_inc(NULL,uc_se,clamp_id,uclamp_none(UCLAMP_MAX));+uclamp_bucket_inc(NULL,NULL,uc_se,clamp_id,+uclamp_none(UCLAMP_MAX));uc_se->effective.bucket_id=uc_se->bucket_id;uc_se->effective.value=uc_se->value;#endif
@@ -6937,8 +6959,8 @@ static inline int alloc_uclamp_sched_group(struct task_group *tg,intclamp_id;for(clamp_id=0;clamp_id<UCLAMP_CNT;++clamp_id){-uclamp_bucket_inc(NULL,&tg->uclamp[clamp_id],clamp_id,-parent->uclamp[clamp_id].value);+uclamp_bucket_inc(NULL,NULL,&tg->uclamp[clamp_id],+clamp_id,parent->uclamp[clamp_id].value);tg->uclamp[clamp_id].effective.value=parent->uclamp[clamp_id].effective.value;tg->uclamp[clamp_id].effective.bucket_id=
@@ -7263,7 +7289,8 @@ static int cpu_util_min_write_u64(struct cgroup_subsys_state *css,}/* Update TG's reference count */-uclamp_bucket_inc(NULL,&tg->uclamp[UCLAMP_MIN],UCLAMP_MIN,min_value);+uclamp_bucket_inc(NULL,css,&tg->uclamp[UCLAMP_MIN],+UCLAMP_MIN,min_value);/* Update effective clamps to track the most restrictive value */cpu_util_update_hier(css,UCLAMP_MIN,tg->uclamp[UCLAMP_MIN].bucket_id,
@@ -7297,7 +7324,8 @@ static int cpu_util_max_write_u64(struct cgroup_subsys_state *css,}/* Update TG's reference count */-uclamp_bucket_inc(NULL,&tg->uclamp[UCLAMP_MAX],UCLAMP_MAX,max_value);+uclamp_bucket_inc(NULL,css,&tg->uclamp[UCLAMP_MAX],+UCLAMP_MAX,max_value);/* Update effective clamps to track the most restrictive value */cpu_util_update_hier(css,UCLAMP_MAX,tg->uclamp[UCLAMP_MAX].bucket_id,
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:16:28
When a task specific clamp value is configured via sched_setattr(2),
this value is accounted in the corresponding clamp bucket every time the
task is {en,de}qeued. However, when cgroups are also in use, the task
specific clamp values could be restricted by the task_group (TG)
clamp values.
Update uclamp_cpu_inc() to aggregate task and TG clamp values. Every
time a task is enqueued, it's accounted in the clamp_bucket defining the
smaller clamp between the task specific value and its TG effective
value. This allows to:
1. ensure cgroup clamps are always used to restrict task specific
requests, i.e. boosted only up to the effective granted value or
clamped at least to a certain value
2. implement a "nice-like" policy, where tasks are still allowed to
request less then what enforced by their current TG
This mimics what already happens for a task's CPU affinity mask when the
task is also in a cpuset, i.e. cgroup attributes are always used to
restrict per-task attributes.
Do this by exploiting the concept of "effective" clamp, which is already
used by a TG to track parent enforced restrictions.
Apply task group clamp restrictions only to tasks belonging to a child
group. While, for tasks in the root group or in an autogroup, only
system defaults are enforced.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Tejun Heo <tj@kernel.org>
---
Changes in v6:
Others:
- wholesale s/group/bucket/
---
include/linux/sched.h | 10 ++++++++++
kernel/sched/core.c | 42 +++++++++++++++++++++++++++++++++++++++++-
2 files changed, 51 insertions(+), 1 deletion(-)
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:16:39
The cgroup CPU bandwidth controller allows to assign a specified
(maximum) bandwidth to the tasks of a group. However this bandwidth is
defined and enforced only on a temporal base, without considering the
actual frequency a CPU is running on. Thus, the amount of computation
completed by a task within an allocated bandwidth can be very different
depending on the actual frequency the CPU is running that task.
The amount of computation can be affected also by the specific CPU a
task is running on, especially when running on asymmetric capacity
systems like Arm's big.LITTLE.
With the availability of schedutil, the scheduler is now able
to drive frequency selections based on actual task utilization.
Moreover, the utilization clamping support provides a mechanism to
bias the frequency selection operated by schedutil depending on
constraints assigned to the tasks currently RUNNABLE on a CPU.
Giving the mechanisms described above, it is now possible to extend the
cpu controller to specify the minimum (or maximum) utilization which
should be considered for tasks RUNNABLE on a cpu.
This makes it possible to better defined the actual computational
power assigned to task groups, thus improving the cgroup CPU bandwidth
controller which is currently based just on time constraints.
Extend the CPU controller with a couple of new attributes util.{min,max}
which allows to enforce utilization boosting and capping for all the
tasks in a group. Specifically:
- util.min: defines the minimum utilization which should be considered
i.e. the RUNNABLE tasks of this group will run at least at a
minimum frequency which corresponds to the min_util
utilization
- util.max: defines the maximum utilization which should be considered
i.e. the RUNNABLE tasks of this group will run up to a
maximum frequency which corresponds to the max_util
utilization
These attributes:
a) are available only for non-root nodes, both on default and legacy
hierarchies, while system wide clamps are defined by a generic
interface which does not depends on cgroups
b) do not enforce any constraints and/or dependencies between the parent
and its child nodes, thus relying:
- on permission settings defined by the system management software,
to define if subgroups can configure their clamp values
- on the delegation model, to ensure that effective clamps are
updated to consider both subgroup requests and parent group
constraints
c) have higher priority than task-specific clamps, defined via
sched_setattr(), thus allowing to control and restrict task requests
This patch provides the basic support to expose the two new attributes
and to validate their run-time updates, while we do not (yet) actually
allocated clamp buckets.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Tejun Heo <tj@kernel.org>
---
NOTEs:
1) The delegation model described above is provided in one of the
following patches of this series.
2) Utilization clamping constraints are useful not only to bias frequency
selection, when a task is running, but also to better support certain
scheduler decisions regarding task placement. For example, on
asymmetric capacity systems, a utilization clamp value can be
conveniently used to enforce important interactive tasks on more capable
CPUs or to run low priority and background tasks on more energy
efficient CPUs.
The ultimate goal of utilization clamping is thus to enable:
- boosting: by selecting an higher capacity CPU and/or higher execution
frequency for small tasks which are affecting the user
interactive experience.
- capping: by selecting more energy efficiency CPUs or lower execution
frequency, for big tasks which are mainly related to
background activities, and thus without a direct impact on
the user experience.
Thus, a proper extension of the cpu controller with utilization clamping
support will make this controller even more suitable for integration
with advanced system management software (e.g. Android).
Indeed, an informed user-space can provide rich information hints to the
scheduler regarding the tasks it's going to schedule.
The bits related to task placement biasing are left for a further
extension once the basic support introduced by this series will be
merged. Anyway they will not affect the integration with cgroups.
Changes in v6:
Others:
- wholesale s/group/bucket/
- wholesale s/_{get,put}/_{inc,dec}/ to match refcount APIs
---
Documentation/admin-guide/cgroup-v2.rst | 25 +++++
init/Kconfig | 22 ++++
kernel/sched/core.c | 131 ++++++++++++++++++++++++
kernel/sched/sched.h | 5 +
4 files changed, 183 insertions(+)
@@ -909,6 +909,12 @@ controller implements weight and absolute bandwidth limit models for normal scheduling policy and absolute bandwidth allocation model for realtime scheduling policy.+Cycles distribution is based, by default, on a temporal base and it+does not account for the frequency at which tasks are executed.+The (optional) utilization clamping support allows to enforce a minimum+bandwidth, which should always be provided by a CPU, and a maximum bandwidth,+which should never be exceeded by a CPU.+ WARNING: cgroup2 doesn't yet support control of realtime processes and the cpu controller can only be enabled when all RT processes are in the root cgroup. Be aware that system management software may already
@@ -974,6 +980,25 @@ All time durations are in microseconds. Shows pressure stall information for CPU. See Documentation/accounting/psi.txt for details.+ cpu.util.min+ A read-write single value file which exists on non-root cgroups.+ The default is "0", i.e. no utilization boosting.++ The minimum utilization in the range [0, 1024].++ This interface allows reading and setting minimum utilization clamp+ values similar to the sched_setattr(2). This minimum utilization+ value is used to clamp the task specific minimum utilization clamp.++ cpu.util.max+ A read-write single value file which exists on non-root cgroups.+ The default is "1024". i.e. no utilization capping++ The maximum utilization in the range [0, 1024].++ This interface allows reading and setting maximum utilization clamp+ values similar to the sched_setattr(2). This maximum utilization+ value is used to clamp the task specific maximum utilization clamp. Memory ------
@@ -866,6 +866,28 @@ config RT_GROUP_SCHEDendif#CGROUP_SCHED+configUCLAMP_TASK_GROUP+bool"Utilization clamping per group of tasks"+depends onCGROUP_SCHED+depends onUCLAMP_TASK+defaultn+help+Thisfeatureenablestheschedulertotracktheclampedutilization+ofeachCPUbasedonRUNNABLEtaskscurrentlyscheduledonthatCPU.++Whenthisoptionisenabled,theusercanspecifyaminandmax+CPUbandwidthwhichisallowedforeachsingletaskinagroup.+Themaxbandwidthallowstoclampthemaximumfrequencyatask+canuse,whiletheminbandwidthallowstodefineaminimum+frequencyataskwillalwaysuse.++Whentaskgroupbasedutilizationclampingisenabled,aneventually+specifiedtask-specificclampvalueisconstrainedbythecgroup+specifiedclampvalue.Bothminimumandmaximumtaskclampingcannot+bebiggerthanthecorrespondingclampingdefinedattaskgrouplevel.++Ifindoubt,sayN.+configCGROUP_PIDSbool"PIDs controller"help
@@ -1294,6 +1294,13 @@ static void __init init_uclamp(void)/* RT tasks by default will go to max frequency */uc_se=&uclamp_default_perf[clamp_id];uclamp_bucket_inc(NULL,uc_se,clamp_id,uclamp_none(UCLAMP_MAX));++#ifdef CONFIG_UCLAMP_TASK_GROUP+/* Init root TG's clamp bucket */+uc_se=&root_task_group.uclamp[clamp_id];+uc_se->value=uclamp_none(clamp_id);+uc_se->bucket_id=0;+#endif}}
@@ -6872,6 +6879,23 @@ void ia64_set_curr_task(int cpu, struct task_struct *p)/* task_group_lock serializes the addition/removal of task groups */staticDEFINE_SPINLOCK(task_group_lock);+staticinlineintalloc_uclamp_sched_group(structtask_group*tg,+structtask_group*parent)+{+#ifdef CONFIG_UCLAMP_TASK_GROUP+intclamp_id;++for(clamp_id=0;clamp_id<UCLAMP_CNT;++clamp_id){+tg->uclamp[clamp_id].value=+parent->uclamp[clamp_id].value;+tg->uclamp[clamp_id].bucket_id=+parent->uclamp[clamp_id].bucket_id;+}+#endif++return1;+}+staticvoidsched_free_group(structtask_group*tg){free_fair_sched_group(tg);
From: Patrick Bellasi <hidden> Date: 2019-01-15 10:17:00
Tasks without a user-defined clamp value are considered not clamped
and by default their utilization can have any value in the
[0..SCHED_CAPACITY_SCALE] range.
Tasks with a user-defined clamp value are allowed to request any value
in that range, and we unconditionally enforce the required clamps.
However, a "System Management Software" could be interested in limiting
the range of clamp values allowed for all tasks.
Add a privileged interface to define a system default configuration via:
/proc/sys/kernel/sched_uclamp_util_{min,max}
which works as an unconditional clamp range restriction for all tasks.
If a task specific value is not compliant with the system default range,
it will be forced to the corresponding system default value.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
---
The current restriction could be too aggressive since, for example if a
task has a util_min which is higher then the system default max, it
will be forced to the system default min unconditionally.
Let say we have:
Task Clamp: min=30, max=40
System Clamps: min=10, max=20
In principle we should set the task's min=20, since the system allows
boosts up to 20%. In the current implementation, however, since the task
mins exceed the system max, we just go for task min=10.
We should probably better restrict util_min to the maximum system
default value, but that would make the code more complex since it
required to track a cross clamp_id dependency.
Let's keep this as a possible future extension whenever we should really
see the need for it.
Changes in v6:
Others:
- wholesale s/group/bucket/
- make use of the bit_for() macro
---
include/linux/sched.h | 5 ++
include/linux/sched/sysctl.h | 11 +++
kernel/sched/core.c | 137 ++++++++++++++++++++++++++++++++++-
kernel/sysctl.c | 16 ++++
4 files changed, 166 insertions(+), 3 deletions(-)
@@ -625,6 +625,11 @@ struct uclamp_se {unsignedintbucket_id:bits_per(UCLAMP_BUCKETS);unsignedintmapped:1;unsignedintactive:1;+/* Clamp bucket and value actually used by a RUNNABLE task */+struct{+unsignedintvalue:bits_per(SCHED_CAPACITY_SCALE);+unsignedintbucket_id:bits_per(UCLAMP_BUCKETS);+}effective;};#endif /* CONFIG_UCLAMP_TASK */
@@ -827,6 +844,72 @@ static inline void uclamp_cpu_update(struct rq *rq, unsigned int clamp_id,WRITE_ONCE(rq->uclamp[clamp_id].value,max_value);}+/*+*Theeffectiveclampbucketindexofataskdependson,byincreasing+*priority:+*-thetaskspecificclampvalue,explicitlyrequestedfromuserspace+*-thesystemdefaultclampvalue,definedbythesysadmin+*+*Asasideeffect,updatethetask'seffectivevalue:+*task_struct::uclamp::effective::value+*torepresenttheclampvalueofthetaskeffectivebucketindex.+*/+staticinlinevoid+uclamp_effective_get(structtask_struct*p,unsignedintclamp_id,+unsignedint*clamp_value,unsignedint*bucket_id)+{+/* Task specific clamp value */+*clamp_value=p->uclamp[clamp_id].value;+*bucket_id=p->uclamp[clamp_id].bucket_id;++/* System default restriction */+if(unlikely(*clamp_value<uclamp_default[UCLAMP_MIN].value||+*clamp_value>uclamp_default[UCLAMP_MAX].value)){+/* Keep it simple: unconditionally enforce system defaults */+*clamp_value=uclamp_default[clamp_id].value;+*bucket_id=uclamp_default[clamp_id].bucket_id;+}+}++staticinlinevoid+uclamp_effective_assign(structtask_struct*p,unsignedintclamp_id)+{+unsignedintclamp_value,bucket_id;++uclamp_effective_get(p,clamp_id,&clamp_value,&bucket_id);++p->uclamp[clamp_id].effective.value=clamp_value;+p->uclamp[clamp_id].effective.bucket_id=bucket_id;+}++staticinlineunsignedintuclamp_effective_bucket_id(structtask_struct*p,+unsignedintclamp_id)+{+unsignedintclamp_value,bucket_id;++/* Task currently refcounted: use back-annotate effective value */+if(p->uclamp[clamp_id].active)+returnp->uclamp[clamp_id].effective.bucket_id;++uclamp_effective_get(p,clamp_id,&clamp_value,&bucket_id);++returnbucket_id;+}++staticunsignedintuclamp_effective_value(structtask_struct*p,+unsignedintclamp_id)+{+unsignedintclamp_value,bucket_id;++/* Task currently refcounted: use back-annotate effective value */+if(p->uclamp[clamp_id].active)+returnp->uclamp[clamp_id].effective.value;++uclamp_effective_get(p,clamp_id,&clamp_value,&bucket_id);++returnclamp_value;+}+/**WhenataskisenqueuedonaCPU'srq,theclampbucketcurrentlydefinedby*thetask'suclamp::bucket_idisreferencecountedonthatCPU.Thisalso
@@ -843,14 +926,15 @@ static inline void uclamp_cpu_inc_id(struct task_struct *p, struct rq *rq,if(unlikely(!p->uclamp[clamp_id].mapped))return;+uclamp_effective_assign(p,clamp_id);-bucket_id=p->uclamp[clamp_id].bucket_id;+bucket_id=uclamp_effective_bucket_id(p,clamp_id);p->uclamp[clamp_id].active=true;rq->uclamp[clamp_id].bucket[bucket_id].tasks++;/* Reset clamp holds on idle exit */-tsk_clamp=p->uclamp[clamp_id].value;+tsk_clamp=uclamp_effective_value(p,clamp_id);uclamp_idle_reset(rq,clamp_id,tsk_clamp);/* CPU's clamp buckets track the max effective clamp value */
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-21 10:15:18
On Tue, Jan 15, 2019 at 10:15:00AM +0000, Patrick Bellasi wrote:
+/*
+ * Number of utilization clamp buckets.
+ *
+ * The first clamp bucket (bucket_id=0) is used to track non clamped tasks, i.e.
+ * util_{min,max} (0,SCHED_CAPACITY_SCALE). Thus we allocate one more bucket in
+ * addition to the compile time configured number.
+ */
+#define UCLAMP_BUCKETS (CONFIG_UCLAMP_BUCKETS_COUNT + 1)
+
+/*
+ * Utilization clamp bucket
+ * @value: clamp value tracked by a clamp bucket
+ * @bucket_id: the bucket index used by the fast-path
+ * @mapped: the bucket index is valid
+ *
+ * A utilization clamp bucket maps a:
+ * clamp value (value), i.e.
+ * util_{min,max} value requested from userspace
+ * to a:
+ * clamp bucket index (bucket_id), i.e.
+ * index of the per-cpu RUNNABLE tasks refcounting array
+ *
+ * The mapped bit is set whenever a task has been mapped on a clamp bucket for
+ * the first time. When this bit is set, any:
+ * uclamp_bucket_inc() - for a new clamp value
+ * is matched by a:
+ * uclamp_bucket_dec() - for the old clamp value
+ */
+struct uclamp_se {
+ unsigned int value : bits_per(SCHED_CAPACITY_SCALE);
+ unsigned int bucket_id : bits_per(UCLAMP_BUCKETS);
+ unsigned int mapped : 1;
+};
Do we want something like:
BUILD_BUG_ON(sizeof(struct uclamp_se) == sizeof(unsigned int));
And/or put a limit on CONFIG_UCLAMP_BUCKETS_COUNT that guarantees that ?
From: Patrick Bellasi <hidden> Date: 2019-01-21 12:27:26
On 21-Jan 11:15, Peter Zijlstra wrote:
On Tue, Jan 15, 2019 at 10:15:00AM +0000, Patrick Bellasi wrote:
quoted
+/*
+ * Number of utilization clamp buckets.
+ *
+ * The first clamp bucket (bucket_id=0) is used to track non clamped tasks, i.e.
+ * util_{min,max} (0,SCHED_CAPACITY_SCALE). Thus we allocate one more bucket in
+ * addition to the compile time configured number.
+ */
+#define UCLAMP_BUCKETS (CONFIG_UCLAMP_BUCKETS_COUNT + 1)
+
+/*
+ * Utilization clamp bucket
+ * @value: clamp value tracked by a clamp bucket
+ * @bucket_id: the bucket index used by the fast-path
+ * @mapped: the bucket index is valid
+ *
+ * A utilization clamp bucket maps a:
+ * clamp value (value), i.e.
+ * util_{min,max} value requested from userspace
+ * to a:
+ * clamp bucket index (bucket_id), i.e.
+ * index of the per-cpu RUNNABLE tasks refcounting array
+ *
+ * The mapped bit is set whenever a task has been mapped on a clamp bucket for
+ * the first time. When this bit is set, any:
+ * uclamp_bucket_inc() - for a new clamp value
+ * is matched by a:
+ * uclamp_bucket_dec() - for the old clamp value
+ */
+struct uclamp_se {
+ unsigned int value : bits_per(SCHED_CAPACITY_SCALE);
+ unsigned int bucket_id : bits_per(UCLAMP_BUCKETS);
+ unsigned int mapped : 1;
+};
Do we want something like:
BUILD_BUG_ON(sizeof(struct uclamp_se) == sizeof(unsigned int));
Mmm... isn't "!=" what you mean ?
We cannot use less then an unsigned int for the fields above... am I
missing something?
And/or put a limit on CONFIG_UCLAMP_BUCKETS_COUNT that guarantees that ?
The number of buckets is currently KConfig limited to a max of 20, which gives:
UCLAMP_BUCKETS: 21 => 5bits
Thus, even on 32 bit targets and assuming 21bits for an "extended"
SCHED_CAPACITY_SCALE range we should always fit into an unsigned int
and have at least 6 bits for flags.
Are you afraid of some compiler magic related to bitfields packing ?
--
#include <best/regards.h>
Patrick Bellasi
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-21 12:51:40
On Mon, Jan 21, 2019 at 12:27:10PM +0000, Patrick Bellasi wrote:
On 21-Jan 11:15, Peter Zijlstra wrote:
quoted
On Tue, Jan 15, 2019 at 10:15:00AM +0000, Patrick Bellasi wrote:
quoted
+/*
+ * Number of utilization clamp buckets.
+ *
+ * The first clamp bucket (bucket_id=0) is used to track non clamped tasks, i.e.
+ * util_{min,max} (0,SCHED_CAPACITY_SCALE). Thus we allocate one more bucket in
+ * addition to the compile time configured number.
+ */
+#define UCLAMP_BUCKETS (CONFIG_UCLAMP_BUCKETS_COUNT + 1)
+
+/*
+ * Utilization clamp bucket
+ * @value: clamp value tracked by a clamp bucket
+ * @bucket_id: the bucket index used by the fast-path
+ * @mapped: the bucket index is valid
+ *
+ * A utilization clamp bucket maps a:
+ * clamp value (value), i.e.
+ * util_{min,max} value requested from userspace
+ * to a:
+ * clamp bucket index (bucket_id), i.e.
+ * index of the per-cpu RUNNABLE tasks refcounting array
+ *
+ * The mapped bit is set whenever a task has been mapped on a clamp bucket for
+ * the first time. When this bit is set, any:
+ * uclamp_bucket_inc() - for a new clamp value
+ * is matched by a:
+ * uclamp_bucket_dec() - for the old clamp value
+ */
+struct uclamp_se {
+ unsigned int value : bits_per(SCHED_CAPACITY_SCALE);
+ unsigned int bucket_id : bits_per(UCLAMP_BUCKETS);
+ unsigned int mapped : 1;
+};
Do we want something like:
BUILD_BUG_ON(sizeof(struct uclamp_se) == sizeof(unsigned int));
Mmm... isn't "!=" what you mean ?
Quite.
We cannot use less then an unsigned int for the fields above... am I
missing something?
I wanted to ensure we don't accidentally use more.
quoted
And/or put a limit on CONFIG_UCLAMP_BUCKETS_COUNT that guarantees that ?
The number of buckets is currently KConfig limited to a max of 20, which gives:
UCLAMP_BUCKETS: 21 => 5bits
Thus, even on 32 bit targets and assuming 21bits for an "extended"
SCHED_CAPACITY_SCALE range we should always fit into an unsigned int
and have at least 6 bits for flags.
Are you afraid of some compiler magic related to bitfields packing ?
Nah, I missed the Kconfig limit and was afraid that some weird configs
would end up with massively huge structures.
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-21 14:59:42
On Tue, Jan 15, 2019 at 10:15:01AM +0000, Patrick Bellasi wrote:
quoted hunk
@@ -835,6 +954,28 @@ static void uclamp_bucket_inc(struct uclamp_se *uc_se, unsigned int clamp_id, } while (!atomic_long_try_cmpxchg(&uc_maps[bucket_id].adata, &uc_map_old.data, uc_map_new.data));+ /*+ * Ensure each CPU tracks the correct value for this clamp bucket.+ * This initialization of per-CPU variables is required only when a+ * clamp value is requested for the first time from a slow-path.+ */
+static void uclamp_bucket_inc(struct uclamp_se *uc_se, unsigned int clamp_id,
+ unsigned int clamp_value)
+{
+ union uclamp_map *uc_maps = &uclamp_maps[clamp_id][0];
+ unsigned int prev_bucket_id = uc_se->bucket_id;
+ union uclamp_map uc_map_old, uc_map_new;
+ unsigned int free_bucket_id;
+ unsigned int bucket_value;
+ unsigned int bucket_id;
+
+ bucket_value = uclamp_bucket_value(clamp_value);
Aahh!!
So why don't you do:
bucket_id = clamp_value / UCLAMP_BUCKET_DELTA;
bucket_value = bucket_id * UCLAMP_BUCKET_DELTA;
+ do {
+ /* Find the bucket_id of an already mapped clamp bucket... */
+ free_bucket_id = UCLAMP_BUCKETS;
+ for (bucket_id = 0; bucket_id < UCLAMP_BUCKETS; ++bucket_id) {
+ uc_map_old.data = atomic_long_read(&uc_maps[bucket_id].adata);
+ if (free_bucket_id == UCLAMP_BUCKETS && !uc_map_old.se_count)
+ free_bucket_id = bucket_id;
+ if (uc_map_old.value == bucket_value)
+ break;
+ }
+
+ /* ... or allocate a new clamp bucket */
+ if (bucket_id >= UCLAMP_BUCKETS) {
+ /*
+ * A valid clamp bucket must always be available.
+ * If we cannot find one: refcounting is broken and we
+ * warn once. The sched_entity will be tracked in the
+ * fast-path using its previous clamp bucket, or not
+ * tracked at all if not yet mapped (i.e. it's new).
+ */
+ if (unlikely(free_bucket_id == UCLAMP_BUCKETS)) {
+ SCHED_WARN_ON(free_bucket_id == UCLAMP_BUCKETS);
+ return;
+ }
+ bucket_id = free_bucket_id;
+ uc_map_old.data = atomic_long_read(&uc_maps[bucket_id].adata);
+ }
And then skip all this?
+
+ uc_map_new.se_count = uc_map_old.se_count + 1;
+ uc_map_new.value = bucket_value;
+
+ } while (!atomic_long_try_cmpxchg(&uc_maps[bucket_id].adata,
+ &uc_map_old.data, uc_map_new.data));
+
+ uc_se->value = clamp_value;
+ uc_se->bucket_id = bucket_id;
+
+ if (uc_se->mapped)
+ uclamp_bucket_dec(clamp_id, prev_bucket_id);
+
+ /*
+ * Task's sched_entity are refcounted in the fast-path only when they
+ * have got a valid clamp_bucket assigned.
+ */
+ uc_se->mapped = true;
+}
From: Patrick Bellasi <hidden> Date: 2019-01-21 15:23:20
On 21-Jan 15:59, Peter Zijlstra wrote:
On Tue, Jan 15, 2019 at 10:15:01AM +0000, Patrick Bellasi wrote:
quoted
@@ -835,6 +954,28 @@ static void uclamp_bucket_inc(struct uclamp_se *uc_se, unsigned int clamp_id, } while (!atomic_long_try_cmpxchg(&uc_maps[bucket_id].adata, &uc_map_old.data, uc_map_new.data));+ /*+ * Ensure each CPU tracks the correct value for this clamp bucket.+ * This initialization of per-CPU variables is required only when a+ * clamp value is requested for the first time from a slow-path.+ */
I'm confused; why is this needed?
That's a lazy initialization of the per-CPU uclamp data for a given
bucket, i.e. the clamp value assigned to a bucket, which happens only
when new clamp values are requested... usually only at system
boot/configuration time.
For example, let say we have these buckets mapped to given clamp
values:
bucket_#0: clamp value: 10% (mapped)
bucket_#1: clamp value: 20% (mapped)
bucket_#2: clamp value: 30% (mapped)
and then let's assume all the users of bucket_#1 are "destroyed", i.e.
there are no more tasks, system defaults or cgroups asking for a
20% clamp value. The corresponding bucket will become free:
bucket_#0: clamp value: 10% (mapped)
bucket_#1: clamp value: 20% (free)
bucket_#2: clamp value: 30% (mapped)
If, in the future, we ask for a new clamp value, let say a task ask
for a 40% clamp value, then we need to map that value into a bucket.
Since bucket_#1 is free we can use it to fill up the hold and keep all
the buckets in use at the beginning of a cache line.
However, since now bucket_#1 tracks a different clamp value (40
instead of 20) we need to walk all the CPUs and updated the cached
value:
bucket_#0: clamp value: 10% (mapped)
bucket_#1: clamp value: 40% (mapped)
bucket_#2: clamp value: 30% (mapped)
Is that more clear ?
In the following code:
quoted
+ if (unlikely(!uc_map_old.se_count)) {
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
This condition is matched by clamp buckets which needs the
initialization described above. These are buckets without a client so
fare and that have been selected to map/track a new clamp value.
That's why we have an unlikely... quite likely tasks/cgroups will keep
asking for the same (limited number of) clamp values and thus we find
a bucket already properly initialized for them.
quoted
+ for_each_possible_cpu(cpu) {
+ struct uclamp_cpu *uc_cpu =
+ &cpu_rq(cpu)->uclamp[clamp_id];
+
+ /* CPU's tasks count must be 0 for free buckets */
+ SCHED_WARN_ON(uc_cpu->bucket[bucket_id].tasks);
+ if (unlikely(uc_cpu->bucket[bucket_id].tasks))
+ uc_cpu->bucket[bucket_id].tasks = 0;
That's a safety check, we expect that (free) buckets do not refcount
any task. That's one of the conditions for a bucket to be considered
free. Here we do just a sanity check, that's because we use unlikely.
If the check matches there is a data corruption, which is reported by
the previous SCHED_WARN_ON and "fixed" by the if branch.
In my tests I have s/SCHED_WARN_ON/BUG_ON/ and never hit that bug...
thus the refcounting code should be ok and this check is there just to
be more on the safe side for future changes.
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-21 15:33:16
On Tue, Jan 15, 2019 at 10:15:02AM +0000, Patrick Bellasi wrote:
+static inline void
+uclamp_task_update_active(struct task_struct *p, unsigned int clamp_id)
+{
+ struct rq_flags rf;
+ struct rq *rq;
+
+ /*
+ * Lock the task and the CPU where the task is (or was) queued.
+ *
+ * We might lock the (previous) rq of a !RUNNABLE task, but that's the
+ * price to pay to safely serialize util_{min,max} updates with
+ * enqueues, dequeues and migration operations.
+ * This is the same locking schema used by __set_cpus_allowed_ptr().
+ */
+ rq = task_rq_lock(p, &rf);
+
+ /*
+ * Setting the clamp bucket is serialized by task_rq_lock().
+ * If the task is not yet RUNNABLE and its task_struct is not
+ * affecting a valid clamp bucket, the next time it's enqueued,
+ * it will already see the updated clamp bucket value.
+ */
+ if (!p->uclamp[clamp_id].active)
+ goto done;
+
+ uclamp_cpu_dec_id(p, rq, clamp_id);
+ uclamp_cpu_inc_id(p, rq, clamp_id);
+
+done:
+ task_rq_unlock(rq, p, &rf);
+}
+static void uclamp_bucket_inc(struct uclamp_se *uc_se, unsigned int clamp_id,
+ unsigned int clamp_value)
+{
+ union uclamp_map *uc_maps = &uclamp_maps[clamp_id][0];
+ unsigned int prev_bucket_id = uc_se->bucket_id;
+ union uclamp_map uc_map_old, uc_map_new;
+ unsigned int free_bucket_id;
+ unsigned int bucket_value;
+ unsigned int bucket_id;
+
+ bucket_value = uclamp_bucket_value(clamp_value);
Aahh!!
So why don't you do:
bucket_id = clamp_value / UCLAMP_BUCKET_DELTA;
bucket_value = bucket_id * UCLAMP_BUCKET_DELTA;
The mapping done here is meant to keep at the beginning of the cache
line all and only the buckets we use. Let say we have configured the
system to track 20 buckets, to have a 5% clamping resolution, but then
we use only two values at run-time, e.g. 13% and 87%.
With the mapping done here the per-CPU variables will have to consider
only 2 buckets:
bucket_#00: clamp value: 10% (mapped)
bucket_#01: clamp value: 85% (mapped)
bucket_#02: (free)
...
bucket_#20: (free)
While without the mapping we will have:
bucket_#00: (free)
bucket_#01: clamp value: 10 (mapped)
bucket_#02: (free)
... big hole crossing a cache line ....
bucket_#16: (free)
bucket_#17: clamp value: 85 (mapped)
bucket_#18: (free)
...
bucket_#20: (free)
Addressing is simple without mapping but we can have performance
issues in the hot-path, since sometimes we need to scan all the
buckets to figure out the new max.
The mapping done here is meant to keep all the used slots at the very
beginning of a cache line to speed up that max computation when
required.
quoted
+ do {
+ /* Find the bucket_id of an already mapped clamp bucket... */
+ free_bucket_id = UCLAMP_BUCKETS;
+ for (bucket_id = 0; bucket_id < UCLAMP_BUCKETS; ++bucket_id) {
+ uc_map_old.data = atomic_long_read(&uc_maps[bucket_id].adata);
+ if (free_bucket_id == UCLAMP_BUCKETS && !uc_map_old.se_count)
+ free_bucket_id = bucket_id;
+ if (uc_map_old.value == bucket_value)
+ break;
+ }
+
+ /* ... or allocate a new clamp bucket */
+ if (bucket_id >= UCLAMP_BUCKETS) {
+ /*
+ * A valid clamp bucket must always be available.
+ * If we cannot find one: refcounting is broken and we
+ * warn once. The sched_entity will be tracked in the
+ * fast-path using its previous clamp bucket, or not
+ * tracked at all if not yet mapped (i.e. it's new).
+ */
+ if (unlikely(free_bucket_id == UCLAMP_BUCKETS)) {
+ SCHED_WARN_ON(free_bucket_id == UCLAMP_BUCKETS);
+ return;
+ }
+ bucket_id = free_bucket_id;
+ uc_map_old.data = atomic_long_read(&uc_maps[bucket_id].adata);
+ }
And then skip all this?
quoted
+
+ uc_map_new.se_count = uc_map_old.se_count + 1;
+ uc_map_new.value = bucket_value;
+
+ } while (!atomic_long_try_cmpxchg(&uc_maps[bucket_id].adata,
+ &uc_map_old.data, uc_map_new.data));
+
+ uc_se->value = clamp_value;
+ uc_se->bucket_id = bucket_id;
+
+ if (uc_se->mapped)
+ uclamp_bucket_dec(clamp_id, prev_bucket_id);
+
+ /*
+ * Task's sched_entity are refcounted in the fast-path only when they
+ * have got a valid clamp_bucket assigned.
+ */
+ uc_se->mapped = true;
+}
From: Patrick Bellasi <hidden> Date: 2019-01-21 15:44:20
On 21-Jan 16:33, Peter Zijlstra wrote:
On Tue, Jan 15, 2019 at 10:15:02AM +0000, Patrick Bellasi wrote:
quoted
+static inline void
+uclamp_task_update_active(struct task_struct *p, unsigned int clamp_id)
+{
+ struct rq_flags rf;
+ struct rq *rq;
+
+ /*
+ * Lock the task and the CPU where the task is (or was) queued.
+ *
+ * We might lock the (previous) rq of a !RUNNABLE task, but that's the
+ * price to pay to safely serialize util_{min,max} updates with
+ * enqueues, dequeues and migration operations.
+ * This is the same locking schema used by __set_cpus_allowed_ptr().
+ */
+ rq = task_rq_lock(p, &rf);
+
+ /*
+ * Setting the clamp bucket is serialized by task_rq_lock().
+ * If the task is not yet RUNNABLE and its task_struct is not
+ * affecting a valid clamp bucket, the next time it's enqueued,
+ * it will already see the updated clamp bucket value.
+ */
+ if (!p->uclamp[clamp_id].active)
+ goto done;
+
+ uclamp_cpu_dec_id(p, rq, clamp_id);
+ uclamp_cpu_inc_id(p, rq, clamp_id);
+
+done:
+ task_rq_unlock(rq, p, &rf);
+}
But.... __sched_setscheduler() actually does the whole dequeue + enqueue
thing already ?!? See where it does __setscheduler().
This is slow-path accounting, not fast path.
There are two refcounting going on here:
1) mapped buckets:
clamp_value <--(M1)--> bucket_id
2) RUNNABLE tasks:
bucket_id <--(M2)--> RUNNABLE tasks in a bucket
What we fix here is the refcounting for the buckets mapping. If a task
does not have a task specific clamp value it does not refcount any
bucket. The moment we assign a task specific clamp value, we need to
refcount the task in the bucket corresponding to that clamp value.
This will keep the bucket in use at least as long as the task will
need that clamp value.
--
#include <best/regards.h>
Patrick Bellasi
With the default of 5, this UCLAMP_BUCKETS := 6, so struct uclamp_cpu
ends up being 7 'unsigned long's, or 56 bytes on 64bit (with a 4 byte
hole).
Yes, that's dimensioned and configured to fit into a single cache line
for all the possible 5 (by default) clamp values of a clamp index
(i.e. min or max util).
quoted
+#endif /* CONFIG_UCLAMP_TASK */
+
/*
* This is the main, per-CPU runqueue data structure.
*
@@ -835,6 +879,11 @@ struct rq { unsigned long nr_load_updates; u64 nr_switches;+#ifdef CONFIG_UCLAMP_TASK+ /* Utilization clamp values based on CPU's RUNNABLE tasks */+ struct uclamp_cpu uclamp[UCLAMP_CNT] ____cacheline_aligned;
Which makes this 112 bytes with 8 bytes in 2 holes, which is short of 2
64 byte cachelines.
Right, we have 2 cache lines where:
- the first $L tracks 5 different util_min values
- the second $L tracks 5 different util_max values
Is that the best layout?
It changed few times and that's what I found more reasonable for both
for fitting the default configuration and also for code readability.
Notice that we access RQ and SE clamp values with the same patter,
for example:
{rq|p}->uclamp[clamp_idx].value
Are you worried about the holes or something else specific ?
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-21 16:12:46
On Mon, Jan 21, 2019 at 03:23:11PM +0000, Patrick Bellasi wrote:
On 21-Jan 15:59, Peter Zijlstra wrote:
quoted
On Tue, Jan 15, 2019 at 10:15:01AM +0000, Patrick Bellasi wrote:
quoted
@@ -835,6 +954,28 @@ static void uclamp_bucket_inc(struct uclamp_se *uc_se, unsigned int clamp_id, } while (!atomic_long_try_cmpxchg(&uc_maps[bucket_id].adata, &uc_map_old.data, uc_map_new.data));+ /*+ * Ensure each CPU tracks the correct value for this clamp bucket.+ * This initialization of per-CPU variables is required only when a+ * clamp value is requested for the first time from a slow-path.+ */
I'm confused; why is this needed?
That's a lazy initialization of the per-CPU uclamp data for a given
bucket, i.e. the clamp value assigned to a bucket, which happens only
when new clamp values are requested... usually only at system
boot/configuration time.
For example, let say we have these buckets mapped to given clamp
values:
bucket_#0: clamp value: 10% (mapped)
bucket_#1: clamp value: 20% (mapped)
bucket_#2: clamp value: 30% (mapped)
and then let's assume all the users of bucket_#1 are "destroyed", i.e.
there are no more tasks, system defaults or cgroups asking for a
20% clamp value. The corresponding bucket will become free:
bucket_#0: clamp value: 10% (mapped)
bucket_#1: clamp value: 20% (free)
bucket_#2: clamp value: 30% (mapped)
If, in the future, we ask for a new clamp value, let say a task ask
for a 40% clamp value, then we need to map that value into a bucket.
Since bucket_#1 is free we can use it to fill up the hold and keep all
the buckets in use at the beginning of a cache line.
However, since now bucket_#1 tracks a different clamp value (40
instead of 20) we need to walk all the CPUs and updated the cached
value:
bucket_#0: clamp value: 10% (mapped)
bucket_#1: clamp value: 40% (mapped)
bucket_#2: clamp value: 30% (mapped)
Is that more clear ?
Yes, and I realized this a little while after sending this; but I'm not
sure I have an answer to why though.
That is; why isn't the whole thing hard coded to have:
bucket_n: clamp value: n*UCLAMP_BUCKET_DELTA
We already do that division anyway (clamp_value / UCLAMP_BUCKET_DELTA),
and from that we instantly have the right bucket index. And that allows
us to initialize all this beforehand.
and keep all
the buckets in use at the beginning of a cache line.
That; is that the rationale for all this? Note that per the defaults
everything is in a single line already.
From: Patrick Bellasi <hidden> Date: 2019-01-21 16:33:46
On 21-Jan 17:12, Peter Zijlstra wrote:
On Mon, Jan 21, 2019 at 03:23:11PM +0000, Patrick Bellasi wrote:
quoted
On 21-Jan 15:59, Peter Zijlstra wrote:
quoted
On Tue, Jan 15, 2019 at 10:15:01AM +0000, Patrick Bellasi wrote:
quoted
@@ -835,6 +954,28 @@ static void uclamp_bucket_inc(struct uclamp_se *uc_se, unsigned int clamp_id, } while (!atomic_long_try_cmpxchg(&uc_maps[bucket_id].adata, &uc_map_old.data, uc_map_new.data));+ /*+ * Ensure each CPU tracks the correct value for this clamp bucket.+ * This initialization of per-CPU variables is required only when a+ * clamp value is requested for the first time from a slow-path.+ */
I'm confused; why is this needed?
That's a lazy initialization of the per-CPU uclamp data for a given
bucket, i.e. the clamp value assigned to a bucket, which happens only
when new clamp values are requested... usually only at system
boot/configuration time.
For example, let say we have these buckets mapped to given clamp
values:
bucket_#0: clamp value: 10% (mapped)
bucket_#1: clamp value: 20% (mapped)
bucket_#2: clamp value: 30% (mapped)
and then let's assume all the users of bucket_#1 are "destroyed", i.e.
there are no more tasks, system defaults or cgroups asking for a
20% clamp value. The corresponding bucket will become free:
bucket_#0: clamp value: 10% (mapped)
bucket_#1: clamp value: 20% (free)
bucket_#2: clamp value: 30% (mapped)
If, in the future, we ask for a new clamp value, let say a task ask
for a 40% clamp value, then we need to map that value into a bucket.
Since bucket_#1 is free we can use it to fill up the hold and keep all
the buckets in use at the beginning of a cache line.
However, since now bucket_#1 tracks a different clamp value (40
instead of 20) we need to walk all the CPUs and updated the cached
value:
bucket_#0: clamp value: 10% (mapped)
bucket_#1: clamp value: 40% (mapped)
bucket_#2: clamp value: 30% (mapped)
Is that more clear ?
Yes, and I realized this a little while after sending this; but I'm not
sure I have an answer to why though.
That is; why isn't the whole thing hard coded to have:
bucket_n: clamp value: n*UCLAMP_BUCKET_DELTA
We already do that division anyway (clamp_value / UCLAMP_BUCKET_DELTA),
and from that we instantly have the right bucket index. And that allows
us to initialize all this beforehand.
quoted
and keep all
the buckets in use at the beginning of a cache line.
That; is that the rationale for all this? Note that per the defaults
everything is in a single line already.
Yes, that's because of the loop in:
dequeue_task()
uclamp_cpu_dec()
uclamp_cpu_dec_id()
uclamp_cpu_update()
where buckets needs sometimes to be scanned to find a new max.
Consider also that, with mapping, we can more easily increase the
buckets count to 20 in order to have a finer clamping granularity if
needed without warring too much about performance impact especially
when we use anyway few different clamp values.
So, I agree that mapping adds (code) complexity but it can also save
few cycles in the fast path... do you think it's not worth the added
complexity?
TBH I never did a proper profiling w/-w/o mapping... I'm just worried
in principle for a loop on 20 entries spanning 4 cache lines. :/
NOTE: the loop is currently going through all the entries anyway,
but we can add later a guard to bail out once we covered the
number of active entries.
--
#include <best/regards.h>
Patrick Bellasi
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-22 09:37:14
On Mon, Jan 21, 2019 at 03:44:12PM +0000, Patrick Bellasi wrote:
On 21-Jan 16:33, Peter Zijlstra wrote:
quoted
On Tue, Jan 15, 2019 at 10:15:02AM +0000, Patrick Bellasi wrote:
quoted
+static inline void
+uclamp_task_update_active(struct task_struct *p, unsigned int clamp_id)
+{
+ struct rq_flags rf;
+ struct rq *rq;
+
+ /*
+ * Lock the task and the CPU where the task is (or was) queued.
+ *
+ * We might lock the (previous) rq of a !RUNNABLE task, but that's the
+ * price to pay to safely serialize util_{min,max} updates with
+ * enqueues, dequeues and migration operations.
+ * This is the same locking schema used by __set_cpus_allowed_ptr().
+ */
+ rq = task_rq_lock(p, &rf);
+
+ /*
+ * Setting the clamp bucket is serialized by task_rq_lock().
+ * If the task is not yet RUNNABLE and its task_struct is not
+ * affecting a valid clamp bucket, the next time it's enqueued,
+ * it will already see the updated clamp bucket value.
+ */
+ if (!p->uclamp[clamp_id].active)
+ goto done;
+
+ uclamp_cpu_dec_id(p, rq, clamp_id);
+ uclamp_cpu_inc_id(p, rq, clamp_id);
+
+done:
+ task_rq_unlock(rq, p, &rf);
+}
But.... __sched_setscheduler() actually does the whole dequeue + enqueue
thing already ?!? See where it does __setscheduler().
This is slow-path accounting, not fast path.
Sure; but that's still no reason for duplicate or unneeded code.
There are two refcounting going on here:
1) mapped buckets:
clamp_value <--(M1)--> bucket_id
2) RUNNABLE tasks:
bucket_id <--(M2)--> RUNNABLE tasks in a bucket
What we fix here is the refcounting for the buckets mapping. If a task
does not have a task specific clamp value it does not refcount any
bucket. The moment we assign a task specific clamp value, we need to
refcount the task in the bucket corresponding to that clamp value.
This will keep the bucket in use at least as long as the task will
need that clamp value.
Sure, I get that. What I don't get is why you're adding that (2) here.
Like said, __sched_setscheduler() already does a dequeue/enqueue under
rq->lock, which should already take care of that.
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-22 09:45:16
On Mon, Jan 21, 2019 at 04:33:38PM +0000, Patrick Bellasi wrote:
On 21-Jan 17:12, Peter Zijlstra wrote:
quoted
On Mon, Jan 21, 2019 at 03:23:11PM +0000, Patrick Bellasi wrote:
quoted
quoted
and keep all
the buckets in use at the beginning of a cache line.
That; is that the rationale for all this? Note that per the defaults
everything is in a single line already.
Yes, that's because of the loop in:
dequeue_task()
uclamp_cpu_dec()
uclamp_cpu_dec_id()
uclamp_cpu_update()
where buckets needs sometimes to be scanned to find a new max.
Consider also that, with mapping, we can more easily increase the
buckets count to 20 in order to have a finer clamping granularity if
needed without warring too much about performance impact especially
when we use anyway few different clamp values.
So, I agree that mapping adds (code) complexity but it can also save
few cycles in the fast path... do you think it's not worth the added
complexity?
Then maybe split this out in a separate patch? Do the trivial linear
bucket thing first and then do this smarty pants thing on top.
One problem with the scheme is that it doesn't defrag; so if you get a
peak usage, you can still end up with only two active buckets in
different lines.
Also; if it is it's own patch, you get a much better view of the
additional complexity and a chance to justify it ;-)
Also; would it make sense to do s/cpu/rq/ on much of this? All this
uclamp_cpu_*() stuff really is per rq and takes rq arguments, so why
does it have cpu in the name... no strong feelings, just noticed it and
thought is a tad inconsistent.
With the default of 5, this UCLAMP_BUCKETS := 6, so struct uclamp_cpu
ends up being 7 'unsigned long's, or 56 bytes on 64bit (with a 4 byte
hole).
Yes, that's dimensioned and configured to fit into a single cache line
for all the possible 5 (by default) clamp values of a clamp index
(i.e. min or max util).
And I suppose you picked 5 because 20% is a 'nice' number? whereas
16./666/% is a bit odd?
quoted
quoted
+#endif /* CONFIG_UCLAMP_TASK */
+
/*
* This is the main, per-CPU runqueue data structure.
*
@@ -835,6 +879,11 @@ struct rq { unsigned long nr_load_updates; u64 nr_switches;+#ifdef CONFIG_UCLAMP_TASK+ /* Utilization clamp values based on CPU's RUNNABLE tasks */+ struct uclamp_cpu uclamp[UCLAMP_CNT] ____cacheline_aligned;
Which makes this 112 bytes with 8 bytes in 2 holes, which is short of 2
64 byte cachelines.
Right, we have 2 cache lines where:
- the first $L tracks 5 different util_min values
- the second $L tracks 5 different util_max values
Well, not quite so, if you want that you should put
____cacheline_aligned on struct uclamp_cpu. Such that the individual
array entries are each aligned, the above only alignes the whole array,
so the second uclamp_cpu is spread over both lines.
But I think this is actually better, since you have to scan both
min/max anyway, and allowing one the straddle a line you have to touch
anyway, allows for using less lines in total.
Consider for example the case where UCLAMP_BUCKETS=8, then each
uclamp_cpu would be 9 words or 72 bytes. If you force align the member,
then you end up with 4 lines, whereas now it would be 3.
quoted
Is that the best layout?
It changed few times and that's what I found more reasonable for both
for fitting the default configuration and also for code readability.
Notice that we access RQ and SE clamp values with the same patter,
for example:
{rq|p}->uclamp[clamp_idx].value
Are you worried about the holes or something else specific ?
Not sure; just mostly asking if this was by design or by accident.
One thing I did wonder though; since bucket[0] is counting the tasks
that are unconstrained and it's bucket value is basically fixed (0 /
1024), can't we abuse that value field to store uclamp_cpu::value ?
OTOH, doing that might make the code really ugly with all them:
if (!bucket_id)
exceptions all over the place.
From: Patrick Bellasi <hidden> Date: 2019-01-22 10:31:15
On 22-Jan 10:45, Peter Zijlstra wrote:
On Mon, Jan 21, 2019 at 04:33:38PM +0000, Patrick Bellasi wrote:
quoted
On 21-Jan 17:12, Peter Zijlstra wrote:
quoted
On Mon, Jan 21, 2019 at 03:23:11PM +0000, Patrick Bellasi wrote:
quoted
quoted
quoted
and keep all
the buckets in use at the beginning of a cache line.
That; is that the rationale for all this? Note that per the defaults
everything is in a single line already.
Yes, that's because of the loop in:
dequeue_task()
uclamp_cpu_dec()
uclamp_cpu_dec_id()
uclamp_cpu_update()
where buckets needs sometimes to be scanned to find a new max.
Consider also that, with mapping, we can more easily increase the
buckets count to 20 in order to have a finer clamping granularity if
needed without warring too much about performance impact especially
when we use anyway few different clamp values.
So, I agree that mapping adds (code) complexity but it can also save
few cycles in the fast path... do you think it's not worth the added
complexity?
Then maybe split this out in a separate patch? Do the trivial linear
bucket thing first and then do this smarty pants thing on top.
One problem with the scheme is that it doesn't defrag; so if you get a
peak usage, you can still end up with only two active buckets in
different lines.
You right, that was saved for a later optimization. :/
Mainly in consideration of the fact that, at least for the main usage
we have in mind on Android, we will likely configure all the required
clamps once for all at boot time.
Also; if it is it's own patch, you get a much better view of the
additional complexity and a chance to justify it ;-)
What about ditching the mapping for the time being and see if we
get a real overhead hit in the future ?
At that point we will revamp the mapping patch with also a proper
defrag support.
Also; would it make sense to do s/cpu/rq/ on much of this? All this
uclamp_cpu_*() stuff really is per rq and takes rq arguments, so why
does it have cpu in the name... no strong feelings, just noticed it and
thought is a tad inconsistent.
The idea behind using "cpu" instead of "rq" was that we use those only at
root rq level and the clamps are aggregated per-CPU.
I remember one of the first versions used "cpu" instead of "rq" as a
parameter name and you proposed to change it as an optimization since
we call it from dequeue_task() where we already have a *rq.
... but, since we have those uclamp data within struct rq, I think you
are right: it makes more sense to rename the functions.
Will do it in v7, thanks.
--
#include <best/regards.h>
Patrick Bellasi
From: Rafael J. Wysocki <hidden> Date: 2019-01-22 10:38:44
On Tuesday, January 15, 2019 11:15:05 AM CET Patrick Bellasi wrote:
quoted hunk
Each time a frequency update is required via schedutil, a frequency is
selected to (possibly) satisfy the utilization reported by each
scheduling class. However, when utilization clamping is in use, the
frequency selection should consider userspace utilization clamping
hints. This will allow, for example, to:
- boost tasks which are directly affecting the user experience
by running them at least at a minimum "requested" frequency
- cap low priority tasks not directly affecting the user experience
by running them only up to a maximum "allowed" frequency
These constraints are meant to support a per-task based tuning of the
frequency selection thus supporting a fine grained definition of
performance boosting vs energy saving strategies in kernel space.
Add support to clamp the utilization and IOWait boost of RUNNABLE FAIR
tasks within the boundaries defined by their aggregated utilization
clamp constraints.
Based on the max(min_util, max_util) of each task, max-aggregated the
CPU clamp value in a way to give the boosted tasks the performance they
need when they happen to be co-scheduled with other capped tasks.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Rafael J. Wysocki <redacted>
---
Changes in v6:
Message-ID: <20181107113849.GC14309@e110439-lin>
- sanity check util_max >= util_min
Others:
- wholesale s/group/bucket/
- wholesale s/_{get,put}/_{inc,dec}/ to match refcount APIs
---
kernel/sched/cpufreq_schedutil.c | 27 ++++++++++++++++++++++++---
kernel/sched/sched.h | 23 +++++++++++++++++++++++
2 files changed, 47 insertions(+), 3 deletions(-)
@@ -218,8 +218,15 @@ unsigned long schedutil_freq_util(int cpu, unsigned long util_cfs,*CFStasksandweusethesamemetrictotracktheeffective*utilization(PELTwindowsaresynchronized)wecandirectlyaddthem*toobtaintheCPU'sactualutilization.+*+*CFSutilizationcanbeboostedorcapped,dependingonutilization+*clampconstraintsrequestedbycurrentlyRUNNABLEtasks.+*WhentherearenoCFSRUNNABLEtasks,clampsarereleasedand+*frequencywillbegracefullyreducedwiththeutilizationdecay.*/-util=util_cfs;+util=(type==ENERGY_UTIL)+?util_cfs+:uclamp_util(rq,util_cfs);util+=cpu_util_rt(rq);dl_util=cpu_util_dl(rq);
@@ -327,6 +334,7 @@ static void sugov_iowait_boost(struct sugov_cpu *sg_cpu, u64 time,unsignedintflags){boolset_iowait_boost=flags&SCHED_CPUFREQ_IOWAIT;+unsignedintmax_boost;/* Reset boost if the CPU appears to have been idle enough */if(sg_cpu->iowait_boost&&
@@ -342,11 +350,24 @@ static void sugov_iowait_boost(struct sugov_cpu *sg_cpu, u64 time,return;sg_cpu->iowait_boost_pending=true;+/*+*BoostFAIRtasksonlyuptotheCPUclampedutilization.+*+*SinceDLtaskshaveamuchmoreadvancedbandwidthcontrol,it's+*safetoassumethatIOboostdoesnotapplytothosetasks.+*Instead,sinceRTtasksarenotutilizationclamped,wedon'twant+*toapplyclampingonIOboostwhilethereisblockedRT+*utilization.+*/+max_boost=sg_cpu->iowait_boost_max;+if(!cpu_util_rt(cpu_rq(sg_cpu->cpu)))+max_boost=uclamp_util(cpu_rq(sg_cpu->cpu),max_boost);+/* Double the boost at each request */if(sg_cpu->iowait_boost){sg_cpu->iowait_boost<<=1;-if(sg_cpu->iowait_boost>sg_cpu->iowait_boost_max)-sg_cpu->iowait_boost=sg_cpu->iowait_boost_max;+if(sg_cpu->iowait_boost>max_boost)+sg_cpu->iowait_boost=max_boost;return;}
IMO it would be better to combine this patch with the next one.
At least some things in it I was about to ask about would go away
then. :-)
Besides, I don't really see a reason for the split here.
From: Patrick Bellasi <hidden> Date: 2019-01-22 10:43:14
On 22-Jan 10:37, Peter Zijlstra wrote:
On Mon, Jan 21, 2019 at 03:44:12PM +0000, Patrick Bellasi wrote:
quoted
On 21-Jan 16:33, Peter Zijlstra wrote:
quoted
On Tue, Jan 15, 2019 at 10:15:02AM +0000, Patrick Bellasi wrote:
quoted
+static inline void
+uclamp_task_update_active(struct task_struct *p, unsigned int clamp_id)
+{
+ struct rq_flags rf;
+ struct rq *rq;
+
+ /*
+ * Lock the task and the CPU where the task is (or was) queued.
+ *
+ * We might lock the (previous) rq of a !RUNNABLE task, but that's the
+ * price to pay to safely serialize util_{min,max} updates with
+ * enqueues, dequeues and migration operations.
+ * This is the same locking schema used by __set_cpus_allowed_ptr().
+ */
+ rq = task_rq_lock(p, &rf);
+
+ /*
+ * Setting the clamp bucket is serialized by task_rq_lock().
+ * If the task is not yet RUNNABLE and its task_struct is not
+ * affecting a valid clamp bucket, the next time it's enqueued,
+ * it will already see the updated clamp bucket value.
+ */
+ if (!p->uclamp[clamp_id].active)
+ goto done;
+
+ uclamp_cpu_dec_id(p, rq, clamp_id);
+ uclamp_cpu_inc_id(p, rq, clamp_id);
+
+done:
+ task_rq_unlock(rq, p, &rf);
+}
But.... __sched_setscheduler() actually does the whole dequeue + enqueue
thing already ?!? See where it does __setscheduler().
This is slow-path accounting, not fast path.
Sure; but that's still no reason for duplicate or unneeded code.
quoted
There are two refcounting going on here:
1) mapped buckets:
clamp_value <--(M1)--> bucket_id
2) RUNNABLE tasks:
bucket_id <--(M2)--> RUNNABLE tasks in a bucket
What we fix here is the refcounting for the buckets mapping. If a task
does not have a task specific clamp value it does not refcount any
bucket. The moment we assign a task specific clamp value, we need to
refcount the task in the bucket corresponding to that clamp value.
This will keep the bucket in use at least as long as the task will
need that clamp value.
Sure, I get that. What I don't get is why you're adding that (2) here.
Like said, __sched_setscheduler() already does a dequeue/enqueue under
rq->lock, which should already take care of that.
Oh, ok... got it what you mean now.
With:
[PATCH v6 01/16] sched/core: Allow sched_setattr() to use the current policy
[off-list ref]
we can call __sched_setscheduler() with:
attr->sched_flags & SCHED_FLAG_KEEP_POLICY
whenever we want just to change the clamp values of a task without
changing its class. Thus, we can end up returning from
__sched_setscheduler() without doing an actual dequeue/enqueue.
This is likely the most common use-case.
I'll better check if I can propagate this info and avoid M2 if we
actually did a dequeue/enqueue.
Cheers Patrick
--
#include <best/regards.h>
Patrick Bellasi
With the default of 5, this UCLAMP_BUCKETS := 6, so struct uclamp_cpu
ends up being 7 'unsigned long's, or 56 bytes on 64bit (with a 4 byte
hole).
Yes, that's dimensioned and configured to fit into a single cache line
for all the possible 5 (by default) clamp values of a clamp index
(i.e. min or max util).
And I suppose you picked 5 because 20% is a 'nice' number? whereas
16./666/% is a bit odd?
Yes, UCLAMP_BUCKETS:=6 gives me 5 20% buckets:
0-19%, 20-39%, 40-59%, 60-79%, 80-99%
plus a 100% bucket to track the max boosted tasks.
Does that makes sense ?
quoted
quoted
quoted
+#endif /* CONFIG_UCLAMP_TASK */
+
/*
* This is the main, per-CPU runqueue data structure.
*
@@ -835,6 +879,11 @@ struct rq { unsigned long nr_load_updates; u64 nr_switches;+#ifdef CONFIG_UCLAMP_TASK+ /* Utilization clamp values based on CPU's RUNNABLE tasks */+ struct uclamp_cpu uclamp[UCLAMP_CNT] ____cacheline_aligned;
Which makes this 112 bytes with 8 bytes in 2 holes, which is short of 2
64 byte cachelines.
Right, we have 2 cache lines where:
- the first $L tracks 5 different util_min values
- the second $L tracks 5 different util_max values
Well, not quite so, if you want that you should put
____cacheline_aligned on struct uclamp_cpu. Such that the individual
array entries are each aligned, the above only alignes the whole array,
so the second uclamp_cpu is spread over both lines.
That's true... I was considering more important to save space if we
have a buckets number which can fit in let say 3 cache lines.
... but if you prefer the other way around I'll move it.
But I think this is actually better, since you have to scan both
min/max anyway, and allowing one the straddle a line you have to touch
anyway, allows for using less lines in total.
Right.
Consider for example the case where UCLAMP_BUCKETS=8, then each
uclamp_cpu would be 9 words or 72 bytes. If you force align the member,
then you end up with 4 lines, whereas now it would be 3.
Exactly :)
quoted
quoted
Is that the best layout?
It changed few times and that's what I found more reasonable for both
for fitting the default configuration and also for code readability.
Notice that we access RQ and SE clamp values with the same patter,
for example:
{rq|p}->uclamp[clamp_idx].value
Are you worried about the holes or something else specific ?
Not sure; just mostly asking if this was by design or by accident.
One thing I did wonder though; since bucket[0] is counting the tasks
that are unconstrained and it's bucket value is basically fixed (0 /
1024), can't we abuse that value field to store uclamp_cpu::value ?
Mmm... should be possible, just worried about adding special cases
which can make the code even more complex of what it's not already.
.... moreover, if we ditch the mapping, the 1024 will be indexed at
the top of the array... so...
OTOH, doing that might make the code really ugly with all them:
if (!bucket_id)
exceptions all over the place.
Exactly... I should read all your comments before replying :)
--
#include <best/regards.h>
Patrick Bellasi
From: Patrick Bellasi <hidden> Date: 2019-01-22 11:02:14
On 22-Jan 11:37, Rafael J. Wysocki wrote:
On Tuesday, January 15, 2019 11:15:05 AM CET Patrick Bellasi wrote:
quoted
Each time a frequency update is required via schedutil, a frequency is
selected to (possibly) satisfy the utilization reported by each
scheduling class. However, when utilization clamping is in use, the
frequency selection should consider userspace utilization clamping
hints. This will allow, for example, to:
- boost tasks which are directly affecting the user experience
by running them at least at a minimum "requested" frequency
- cap low priority tasks not directly affecting the user experience
by running them only up to a maximum "allowed" frequency
These constraints are meant to support a per-task based tuning of the
frequency selection thus supporting a fine grained definition of
performance boosting vs energy saving strategies in kernel space.
Add support to clamp the utilization and IOWait boost of RUNNABLE FAIR
tasks within the boundaries defined by their aggregated utilization
clamp constraints.
Based on the max(min_util, max_util) of each task, max-aggregated the
CPU clamp value in a way to give the boosted tasks the performance they
need when they happen to be co-scheduled with other capped tasks.
Signed-off-by: Patrick Bellasi <redacted>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Rafael J. Wysocki <redacted>
---
Changes in v6:
Message-ID: <20181107113849.GC14309@e110439-lin>
- sanity check util_max >= util_min
Others:
- wholesale s/group/bucket/
- wholesale s/_{get,put}/_{inc,dec}/ to match refcount APIs
---
kernel/sched/cpufreq_schedutil.c | 27 ++++++++++++++++++++++++---
kernel/sched/sched.h | 23 +++++++++++++++++++++++
2 files changed, 47 insertions(+), 3 deletions(-)
@@ -218,8 +218,15 @@ unsigned long schedutil_freq_util(int cpu, unsigned long util_cfs,*CFStasksandweusethesamemetrictotracktheeffective*utilization(PELTwindowsaresynchronized)wecandirectlyaddthem*toobtaintheCPU'sactualutilization.+*+*CFSutilizationcanbeboostedorcapped,dependingonutilization+*clampconstraintsrequestedbycurrentlyRUNNABLEtasks.+*WhentherearenoCFSRUNNABLEtasks,clampsarereleasedand+*frequencywillbegracefullyreducedwiththeutilizationdecay.*/-util=util_cfs;+util=(type==ENERGY_UTIL)+?util_cfs+:uclamp_util(rq,util_cfs);util+=cpu_util_rt(rq);dl_util=cpu_util_dl(rq);
@@ -327,6 +334,7 @@ static void sugov_iowait_boost(struct sugov_cpu *sg_cpu, u64 time,unsignedintflags){boolset_iowait_boost=flags&SCHED_CPUFREQ_IOWAIT;+unsignedintmax_boost;/* Reset boost if the CPU appears to have been idle enough */if(sg_cpu->iowait_boost&&
@@ -342,11 +350,24 @@ static void sugov_iowait_boost(struct sugov_cpu *sg_cpu, u64 time,return;sg_cpu->iowait_boost_pending=true;+/*+*BoostFAIRtasksonlyuptotheCPUclampedutilization.+*+*SinceDLtaskshaveamuchmoreadvancedbandwidthcontrol,it's+*safetoassumethatIOboostdoesnotapplytothosetasks.+*Instead,sinceRTtasksarenotutilizationclamped,wedon'twant+*toapplyclampingonIOboostwhilethereisblockedRT+*utilization.+*/+max_boost=sg_cpu->iowait_boost_max;+if(!cpu_util_rt(cpu_rq(sg_cpu->cpu)))+max_boost=uclamp_util(cpu_rq(sg_cpu->cpu),max_boost);+/* Double the boost at each request */if(sg_cpu->iowait_boost){sg_cpu->iowait_boost<<=1;-if(sg_cpu->iowait_boost>sg_cpu->iowait_boost_max)-sg_cpu->iowait_boost=sg_cpu->iowait_boost_max;+if(sg_cpu->iowait_boost>max_boost)+sg_cpu->iowait_boost=max_boost;return;}
IMO it would be better to combine this patch with the next one.
Main reason was to better document in the changelog what we do for the
two different classes...
At least some things in it I was about to ask about would go away
then. :-)
... but if it creates confusion I can certainly merge them.
Or maybe clarify better in this patch what's not clear: may I ask what
were your questions ?
Besides, I don't really see a reason for the split here.
Was mainly to make the changes required for RT more self-contained.
For that class only, not for FAIR, we have additional code in the
following patch which add uclamp_default_perf which are system
defaults used to track/account tasks requesting the maximum frequency.
Again, I can either better clarify the above patch or just merge the
two together: what do you prefer ?
--
#include <best/regards.h>
Patrick Bellasi
From: "Rafael J. Wysocki" <rafael@kernel.org> Date: 2019-01-22 11:05:10
On Tue, Jan 22, 2019 at 12:02 PM Patrick Bellasi
[off-list ref] wrote:
On 22-Jan 11:37, Rafael J. Wysocki wrote:
quoted
On Tuesday, January 15, 2019 11:15:05 AM CET Patrick Bellasi wrote:
[cut]
quoted
IMO it would be better to combine this patch with the next one.
Main reason was to better document in the changelog what we do for the
two different classes...
quoted
At least some things in it I was about to ask about would go away
then. :-)
... but if it creates confusion I can certainly merge them.
Or maybe clarify better in this patch what's not clear: may I ask what
were your questions ?
quoted
Besides, I don't really see a reason for the split here.
Was mainly to make the changes required for RT more self-contained.
For that class only, not for FAIR, we have additional code in the
following patch which add uclamp_default_perf which are system
defaults used to track/account tasks requesting the maximum frequency.
Again, I can either better clarify the above patch or just merge the
two together: what do you prefer ?
On Tuesday 15 Jan 2019 at 10:15:08 (+0000), Patrick Bellasi wrote:
The Energy Aware Scheduler (AES) estimates the energy impact of waking
s/AES/EAS :-)
[...]
+ for_each_cpu_and(cpu, pd_mask, cpu_online_mask) {
+ cfs_util = cpu_util_next(cpu, p, dst_cpu);
+
+ /*
+ * Busy time computation: utilization clamping is not
+ * required since the ratio (sum_util / cpu_capacity)
+ * is already enough to scale the EM reported power
+ * consumption at the (eventually clamped) cpu_capacity.
+ */
Right.
+ sum_util += schedutil_cpu_util(cpu, cfs_util, cpu_cap,
+ ENERGY_UTIL, NULL);
+
+ /*
+ * Performance domain frequency: utilization clamping
+ * must be considered since it affects the selection
+ * of the performance domain frequency.
+ */
So that actually affects the way we deal with RT I think. I assume the
idea is to say if you don't want to reflect the RT-go-to-max-freq thing
in EAS (which is what we do now) you should set the min cap for RT to 0.
Is that correct ?
I'm fine with this conceptually but maybe the specific case of RT should
be mentioned somewhere in the commit message or so ? I think it's
important to say that clearly since this patch changes the default
behaviour.
+ cpu_util = schedutil_cpu_util(cpu, cfs_util, cpu_cap,
+ FREQUENCY_UTIL,
+ cpu == dst_cpu ? p : NULL);
+ max_util = max(max_util, cpu_util);
}
energy += em_pd_energy(pd->em_pd, max_util, sum_util);
From: Patrick Bellasi <hidden> Date: 2019-01-22 12:45:53
On 22-Jan 12:13, Quentin Perret wrote:
On Tuesday 15 Jan 2019 at 10:15:08 (+0000), Patrick Bellasi wrote:
quoted
The Energy Aware Scheduler (AES) estimates the energy impact of waking
[...]
quoted
+ for_each_cpu_and(cpu, pd_mask, cpu_online_mask) {
+ cfs_util = cpu_util_next(cpu, p, dst_cpu);
+
+ /*
+ * Busy time computation: utilization clamping is not
+ * required since the ratio (sum_util / cpu_capacity)
+ * is already enough to scale the EM reported power
+ * consumption at the (eventually clamped) cpu_capacity.
+ */
Right.
quoted
+ sum_util += schedutil_cpu_util(cpu, cfs_util, cpu_cap,
+ ENERGY_UTIL, NULL);
+
+ /*
+ * Performance domain frequency: utilization clamping
+ * must be considered since it affects the selection
+ * of the performance domain frequency.
+ */
So that actually affects the way we deal with RT I think. I assume the
idea is to say if you don't want to reflect the RT-go-to-max-freq thing
in EAS (which is what we do now) you should set the min cap for RT to 0.
Is that correct ?
By default configuration, RT tasks still go to max when uclamp is
enabled, since they get a util_min=1024.
If we want to save power on RT tasks, we can set a smaller util_min...
but not necessarily 0. A util_min=0 for RT tasks means to use just
cpu_util_rt() for that class.
I'm fine with this conceptually but maybe the specific case of RT should
be mentioned somewhere in the commit message or so ? I think it's
important to say that clearly since this patch changes the default
behaviour.
Default behavior for RT should not be affected. While a capping is
possible for those tasks... where do you see issues ?
Here we are just figuring out what's the capacity the task will run
at, if we will have clamped RT tasks will not be the max but: is that
a problem ?
quoted
+ cpu_util = schedutil_cpu_util(cpu, cfs_util, cpu_cap,
+ FREQUENCY_UTIL,
+ cpu == dst_cpu ? p : NULL);
+ max_util = max(max_util, cpu_util);
}
energy += em_pd_energy(pd->em_pd, max_util, sum_util);
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-22 13:28:41
On Tue, Jan 22, 2019 at 10:43:05AM +0000, Patrick Bellasi wrote:
On 22-Jan 10:37, Peter Zijlstra wrote:
quoted
Sure, I get that. What I don't get is why you're adding that (2) here.
Like said, __sched_setscheduler() already does a dequeue/enqueue under
rq->lock, which should already take care of that.
Oh, ok... got it what you mean now.
With:
[PATCH v6 01/16] sched/core: Allow sched_setattr() to use the current policy
[off-list ref]
we can call __sched_setscheduler() with:
attr->sched_flags & SCHED_FLAG_KEEP_POLICY
whenever we want just to change the clamp values of a task without
changing its class. Thus, we can end up returning from
__sched_setscheduler() without doing an actual dequeue/enqueue.
I don't see that happening.. when KEEP_POLICY we set attr.sched_policy =
SETPARAM_POLICY. That is then checked again in __setscheduler_param(),
which is in the middle of that dequeue/enqueue.
Also, and this might be 'broken', SETPARAM_POLICY _does_ reset all the
other attributes, it only preserves policy, but it will (re)set nice
level for example (see that same function).
So maybe we want to introduce another (few?) FLAG_KEEP flag(s) that
preserve the other bits; I'm thinking at least KEEP_PARAM and KEEP_UTIL
or something.
On Tuesday 22 Jan 2019 at 12:45:46 (+0000), Patrick Bellasi wrote:
On 22-Jan 12:13, Quentin Perret wrote:
quoted
On Tuesday 15 Jan 2019 at 10:15:08 (+0000), Patrick Bellasi wrote:
quoted
The Energy Aware Scheduler (AES) estimates the energy impact of waking
[...]
quoted
quoted
+ for_each_cpu_and(cpu, pd_mask, cpu_online_mask) {
+ cfs_util = cpu_util_next(cpu, p, dst_cpu);
+
+ /*
+ * Busy time computation: utilization clamping is not
+ * required since the ratio (sum_util / cpu_capacity)
+ * is already enough to scale the EM reported power
+ * consumption at the (eventually clamped) cpu_capacity.
+ */
Right.
quoted
+ sum_util += schedutil_cpu_util(cpu, cfs_util, cpu_cap,
+ ENERGY_UTIL, NULL);
+
+ /*
+ * Performance domain frequency: utilization clamping
+ * must be considered since it affects the selection
+ * of the performance domain frequency.
+ */
So that actually affects the way we deal with RT I think. I assume the
idea is to say if you don't want to reflect the RT-go-to-max-freq thing
in EAS (which is what we do now) you should set the min cap for RT to 0.
Is that correct ?
By default configuration, RT tasks still go to max when uclamp is
enabled, since they get a util_min=1024.
If we want to save power on RT tasks, we can set a smaller util_min...
but not necessarily 0. A util_min=0 for RT tasks means to use just
cpu_util_rt() for that class.
Ah, sorry, I guess my message was misleading. I'm saying this is
changing the way _EAS_ deals with RT tasks. Right now we don't actually
consider the RT-go-to-max thing at all in the EAS prediction. Your
patch is changing that AFAICT. It actually changes the way EAS sees RT
tasks even without uclamp ...
But I'm not hostile to the idea since it's possible to enable uclamp and
set the min cap to 0 for RT if you want. So there is a story there.
However, I think this needs be documented somewhere, at the very least.
Thanks,
Quentin
@@ -625,6 +625,11 @@ struct uclamp_se {unsignedintbucket_id:bits_per(UCLAMP_BUCKETS);unsignedintmapped:1;unsignedintactive:1;+/* Clamp bucket and value actually used by a RUNNABLE task */+struct{+unsignedintvalue:bits_per(SCHED_CAPACITY_SCALE);+unsignedintbucket_id:bits_per(UCLAMP_BUCKETS);+}effective;
I am confuzled by this thing.. so uclamp_se already has a value,bucket,
which per the prior code is the effective one.
Now; I think I see why you want another value; you need the second to
store the original value for when the system limits change and we must
re-evaluate.
So why are you not adding something like:
unsigned int orig_value : bits_per(SCHED_CAPACITY_SCALE);
+unsigned int sysctl_sched_uclamp_util_min;
+unsigned int sysctl_sched_uclamp_util_max = SCHED_CAPACITY_SCALE;
+static inline void
+uclamp_effective_get(struct task_struct *p, unsigned int clamp_id,
+ unsigned int *clamp_value, unsigned int *bucket_id)
+{
+ /* Task specific clamp value */
+ *clamp_value = p->uclamp[clamp_id].value;
+ *bucket_id = p->uclamp[clamp_id].bucket_id;
+
+ /* System default restriction */
+ if (unlikely(*clamp_value < uclamp_default[UCLAMP_MIN].value ||
+ *clamp_value > uclamp_default[UCLAMP_MAX].value)) {
+ /* Keep it simple: unconditionally enforce system defaults */
+ *clamp_value = uclamp_default[clamp_id].value;
+ *bucket_id = uclamp_default[clamp_id].bucket_id;
+ }
+}
That would then turn into something like:
unsigned int high = READ_ONCE(sysctl_sched_uclamp_util_max);
unsigned int low = READ_ONCE(sysctl_sched_uclamp_util_min);
uclamp_se->orig_value = value;
uclamp_se->value = clamp(value, low, high);
And then determine bucket_id based on value.
+int sched_uclamp_handler(struct ctl_table *table, int write,
+ void __user *buffer, size_t *lenp,
+ loff_t *ppos)
+{
+ int old_min, old_max;
+ int result = 0;
+
+ mutex_lock(&uclamp_mutex);
+
+ old_min = sysctl_sched_uclamp_util_min;
+ old_max = sysctl_sched_uclamp_util_max;
+
+ result = proc_dointvec(table, write, buffer, lenp, ppos);
+ if (result)
+ goto undo;
+ if (!write)
+ goto done;
+
+ if (sysctl_sched_uclamp_util_min > sysctl_sched_uclamp_util_max ||
+ sysctl_sched_uclamp_util_max > SCHED_CAPACITY_SCALE) {
+ result = -EINVAL;
+ goto undo;
+ }
+
+ if (old_min != sysctl_sched_uclamp_util_min) {
+ uclamp_bucket_inc(NULL, &uclamp_default[UCLAMP_MIN],
+ UCLAMP_MIN, sysctl_sched_uclamp_util_min);
+ }
+ if (old_max != sysctl_sched_uclamp_util_max) {
+ uclamp_bucket_inc(NULL, &uclamp_default[UCLAMP_MAX],
+ UCLAMP_MAX, sysctl_sched_uclamp_util_max);
+ }
From: Patrick Bellasi <hidden> Date: 2019-01-22 14:01:23
On 22-Jan 14:28, Peter Zijlstra wrote:
On Tue, Jan 22, 2019 at 10:43:05AM +0000, Patrick Bellasi wrote:
quoted
On 22-Jan 10:37, Peter Zijlstra wrote:
quoted
quoted
Sure, I get that. What I don't get is why you're adding that (2) here.
Like said, __sched_setscheduler() already does a dequeue/enqueue under
rq->lock, which should already take care of that.
Oh, ok... got it what you mean now.
With:
[PATCH v6 01/16] sched/core: Allow sched_setattr() to use the current policy
[off-list ref]
we can call __sched_setscheduler() with:
attr->sched_flags & SCHED_FLAG_KEEP_POLICY
whenever we want just to change the clamp values of a task without
changing its class. Thus, we can end up returning from
__sched_setscheduler() without doing an actual dequeue/enqueue.
I don't see that happening.. when KEEP_POLICY we set attr.sched_policy =
SETPARAM_POLICY. That is then checked again in __setscheduler_param(),
which is in the middle of that dequeue/enqueue.
Yes, I think I've forgot a check before we actually dequeue the task.
The current code does:
---8<---
SYSCALL_DEFINE3(sched_setattr)
// A) request to keep the same policy
if (attr.sched_flags & SCHED_FLAG_KEEP_POLICY)
attr.sched_policy = SETPARAM_POLICY;
sched_setattr()
// B) actually enforce the same policy
if (policy < 0)
policy = oldpolicy = p->policy;
// C) tune the clamp values
if (attr->sched_flags & SCHED_FLAG_UTIL_CLAMP)
retval = __setscheduler_uclamp(p, attr);
// D) tune attributes if policy is the same
if (unlikely(policy == p->policy))
if (fair_policy(policy) && attr->sched_nice != task_nice(p))
goto change;
if (rt_policy(policy) && attr->sched_priority != p->rt_priority)
goto change;
if (dl_policy(policy) && dl_param_changed(p, attr))
goto change;
return 0;
change:
// E) dequeue/enqueue task
---8<---
So, probably in D) I've missed a check on SCHED_FLAG_KEEP_POLICY to
enforce a return in that case...
Also, and this might be 'broken', SETPARAM_POLICY _does_ reset all the
other attributes, it only preserves policy, but it will (re)set nice
level for example (see that same function).
Mmm... right... my bad! :/
So maybe we want to introduce another (few?) FLAG_KEEP flag(s) that
preserve the other bits; I'm thinking at least KEEP_PARAM and KEEP_UTIL
or something.
Yes, I would say we have two options:
1) SCHED_FLAG_KEEP_POLICY enforces all the scheduling class specific
attributes, but cross class attributes (e.g. uclamp)
2) add SCHED_KEEP_NICE, SCHED_KEEP_PRIO, and SCED_KEEP_PARAMS
and use them in the if conditions in D)
In both cases the goal should be to return from code block D).
What do you prefer?
--
#include <best/regards.h>
Patrick Bellasi
From: Patrick Bellasi <hidden> Date: 2019-01-22 14:26:14
On 22-Jan 13:29, Quentin Perret wrote:
On Tuesday 22 Jan 2019 at 12:45:46 (+0000), Patrick Bellasi wrote:
quoted
On 22-Jan 12:13, Quentin Perret wrote:
quoted
On Tuesday 15 Jan 2019 at 10:15:08 (+0000), Patrick Bellasi wrote:
quoted
The Energy Aware Scheduler (AES) estimates the energy impact of waking
[...]
quoted
quoted
+ for_each_cpu_and(cpu, pd_mask, cpu_online_mask) {
+ cfs_util = cpu_util_next(cpu, p, dst_cpu);
+
+ /*
+ * Busy time computation: utilization clamping is not
+ * required since the ratio (sum_util / cpu_capacity)
+ * is already enough to scale the EM reported power
+ * consumption at the (eventually clamped) cpu_capacity.
+ */
Right.
quoted
+ sum_util += schedutil_cpu_util(cpu, cfs_util, cpu_cap,
+ ENERGY_UTIL, NULL);
+
+ /*
+ * Performance domain frequency: utilization clamping
+ * must be considered since it affects the selection
+ * of the performance domain frequency.
+ */
So that actually affects the way we deal with RT I think. I assume the
idea is to say if you don't want to reflect the RT-go-to-max-freq thing
in EAS (which is what we do now) you should set the min cap for RT to 0.
Is that correct ?
By default configuration, RT tasks still go to max when uclamp is
enabled, since they get a util_min=1024.
If we want to save power on RT tasks, we can set a smaller util_min...
but not necessarily 0. A util_min=0 for RT tasks means to use just
cpu_util_rt() for that class.
Ah, sorry, I guess my message was misleading. I'm saying this is
changing the way _EAS_ deals with RT tasks. Right now we don't actually
consider the RT-go-to-max thing at all in the EAS prediction. Your
patch is changing that AFAICT. It actually changes the way EAS sees RT
tasks even without uclamp ...
Lemme see if I get it right.
Currently, whenever we look at CPU utilization for ENERGY_UTIL, we
always use cpu_util_rt() for RT tasks:
---8<---
util = util_cfs;
util += cpu_util_rt(rq);
util += dl_util;
---8<---
Thus, even when RT tasks are RUNNABLE, we don't always assume the CPU
running at the max capacity but just whatever is the aggregated
utilization across all the classes.
With uclamp, we have:
---8<---
util = cpu_util_rt(rq) + util_cfs;
if (type == FREQUENCY_UTIL)
util = uclamp_util_with(rq, util, p);
dl_util = cpu_util_dl(rq);
if (type == ENERGY_UTIL)
util += dl_util;
---8<---
So, I would say that, in terms of ENERGY_UTIL we do the same both
w/ and w/o uclamp. Isn't it?
But I'm not hostile to the idea since it's possible to enable uclamp and
set the min cap to 0 for RT if you want. So there is a story there.
However, I think this needs be documented somewhere, at the very least.
The only difference I see is that the actual frequency could be
different (lower then max) when a clamped RT task is RUNNABLE.
Are you worried that running RT on a lower freq could have side
effects on the estimated busy time the CPU ?
I also still don't completely get why you say it could be useful to
"set the min cap to 0 for RT if you want"
IMO this will be an even bigger difference wrt mainline, since the RT
tasks will never have a granted minimum freq but just whatever
utilization we measure for them.
--
#include <best/regards.h>
Patrick Bellasi
On Tuesday 22 Jan 2019 at 14:26:06 (+0000), Patrick Bellasi wrote:
On 22-Jan 13:29, Quentin Perret wrote:
quoted
On Tuesday 22 Jan 2019 at 12:45:46 (+0000), Patrick Bellasi wrote:
quoted
On 22-Jan 12:13, Quentin Perret wrote:
quoted
On Tuesday 15 Jan 2019 at 10:15:08 (+0000), Patrick Bellasi wrote:
quoted
The Energy Aware Scheduler (AES) estimates the energy impact of waking
[...]
quoted
quoted
+ for_each_cpu_and(cpu, pd_mask, cpu_online_mask) {
+ cfs_util = cpu_util_next(cpu, p, dst_cpu);
+
+ /*
+ * Busy time computation: utilization clamping is not
+ * required since the ratio (sum_util / cpu_capacity)
+ * is already enough to scale the EM reported power
+ * consumption at the (eventually clamped) cpu_capacity.
+ */
Right.
quoted
+ sum_util += schedutil_cpu_util(cpu, cfs_util, cpu_cap,
+ ENERGY_UTIL, NULL);
+
+ /*
+ * Performance domain frequency: utilization clamping
+ * must be considered since it affects the selection
+ * of the performance domain frequency.
+ */
So that actually affects the way we deal with RT I think. I assume the
idea is to say if you don't want to reflect the RT-go-to-max-freq thing
in EAS (which is what we do now) you should set the min cap for RT to 0.
Is that correct ?
By default configuration, RT tasks still go to max when uclamp is
enabled, since they get a util_min=1024.
If we want to save power on RT tasks, we can set a smaller util_min...
but not necessarily 0. A util_min=0 for RT tasks means to use just
cpu_util_rt() for that class.
Ah, sorry, I guess my message was misleading. I'm saying this is
changing the way _EAS_ deals with RT tasks. Right now we don't actually
consider the RT-go-to-max thing at all in the EAS prediction. Your
patch is changing that AFAICT. It actually changes the way EAS sees RT
tasks even without uclamp ...
Lemme see if I get it right.
Currently, whenever we look at CPU utilization for ENERGY_UTIL, we
always use cpu_util_rt() for RT tasks:
---8<---
util = util_cfs;
util += cpu_util_rt(rq);
util += dl_util;
---8<---
Thus, even when RT tasks are RUNNABLE, we don't always assume the CPU
running at the max capacity but just whatever is the aggregated
utilization across all the classes.
With uclamp, we have:
---8<---
util = cpu_util_rt(rq) + util_cfs;
if (type == FREQUENCY_UTIL)
util = uclamp_util_with(rq, util, p);
dl_util = cpu_util_dl(rq);
if (type == ENERGY_UTIL)
util += dl_util;
---8<---
So, I would say that, in terms of ENERGY_UTIL we do the same both
w/ and w/o uclamp. Isn't it?
Yes but now you use FREQUENCY_UTIL for computing 'max_util' in the EAS
prediction.
Let's take an example. You have a perf domain with two CPUs. One CPU is
busy running a RT task, the other CPU runs a CFS task. Right now in
compute_energy() we only use ENERGY_UTIL, so 'max_util' ends up being
the max between the utilization of the two tasks. So we don't predict
we're going to max freq.
With your patch, we use FREQUENCY_UTIL to compute 'max_util', so we
_will_ predict that we're going to max freq. And we will do that even if
CONFIG_UCLAMP_TASK=n.
The default EAS calculation will be different with your patch when there
are runnable RT tasks in the system. This needs to be documented, I
think.
quoted
But I'm not hostile to the idea since it's possible to enable uclamp and
set the min cap to 0 for RT if you want. So there is a story there.
However, I think this needs be documented somewhere, at the very least.
The only difference I see is that the actual frequency could be
different (lower then max) when a clamped RT task is RUNNABLE.
Are you worried that running RT on a lower freq could have side
effects on the estimated busy time the CPU ?
I also still don't completely get why you say it could be useful to
"set the min cap to 0 for RT if you want"
I'm not saying it's useful, I'm saying userspace can decide to do that
if it thinks it is a good idea. The default should be min_cap = 1024 for
RT, no questions. But you _can_ change it at runtime if you want to.
That's my point. And doing that basically provides the same behaviour as
what we have right now in terms of EAS calculation (but it changes the
freq selection obviously) which is why I'm not fundamentally opposed to
your patch.
So in short, I'm fine with the behavioural change, but please at least
mention it somewhere :-)
Thanks,
Quentin
@@ -625,6 +625,11 @@ struct uclamp_se {unsignedintbucket_id:bits_per(UCLAMP_BUCKETS);unsignedintmapped:1;unsignedintactive:1;+/* Clamp bucket and value actually used by a RUNNABLE task */+struct{+unsignedintvalue:bits_per(SCHED_CAPACITY_SCALE);+unsignedintbucket_id:bits_per(UCLAMP_BUCKETS);+}effective;
I am confuzled by this thing.. so uclamp_se already has a value,bucket,
which per the prior code is the effective one.
Now; I think I see why you want another value; you need the second to
store the original value for when the system limits change and we must
re-evaluate.
Yes, that's one reason, the other one being to properly support
CGroup when we add them in the following patches.
Effective will always track the value/bucket in which the task has
been refcounted at enqueue time and it depends on the aggregated
value.
So why are you not adding something like:
unsigned int orig_value : bits_per(SCHED_CAPACITY_SCALE);
Would say that can be enough if we decide to ditch the mapping and use
a linear mapping. In that case the value will always be enough to find
in which bucket a task has been accounted.
quoted
+unsigned int sysctl_sched_uclamp_util_min;
quoted
+unsigned int sysctl_sched_uclamp_util_max = SCHED_CAPACITY_SCALE;
quoted
+static inline void
+uclamp_effective_get(struct task_struct *p, unsigned int clamp_id,
+ unsigned int *clamp_value, unsigned int *bucket_id)
+{
+ /* Task specific clamp value */
+ *clamp_value = p->uclamp[clamp_id].value;
+ *bucket_id = p->uclamp[clamp_id].bucket_id;
+
+ /* System default restriction */
+ if (unlikely(*clamp_value < uclamp_default[UCLAMP_MIN].value ||
+ *clamp_value > uclamp_default[UCLAMP_MAX].value)) {
+ /* Keep it simple: unconditionally enforce system defaults */
+ *clamp_value = uclamp_default[clamp_id].value;
+ *bucket_id = uclamp_default[clamp_id].bucket_id;
+ }
+}
That would then turn into something like:
unsigned int high = READ_ONCE(sysctl_sched_uclamp_util_max);
unsigned int low = READ_ONCE(sysctl_sched_uclamp_util_min);
uclamp_se->orig_value = value;
uclamp_se->value = clamp(value, low, high);
And then determine bucket_id based on value.
Right... if I ditch the mapping that should work.
quoted
+int sched_uclamp_handler(struct ctl_table *table, int write,
+ void __user *buffer, size_t *lenp,
+ loff_t *ppos)
+{
+ int old_min, old_max;
+ int result = 0;
+
+ mutex_lock(&uclamp_mutex);
+
+ old_min = sysctl_sched_uclamp_util_min;
+ old_max = sysctl_sched_uclamp_util_max;
+
+ result = proc_dointvec(table, write, buffer, lenp, ppos);
+ if (result)
+ goto undo;
+ if (!write)
+ goto done;
+
+ if (sysctl_sched_uclamp_util_min > sysctl_sched_uclamp_util_max ||
+ sysctl_sched_uclamp_util_max > SCHED_CAPACITY_SCALE) {
+ result = -EINVAL;
+ goto undo;
+ }
+
+ if (old_min != sysctl_sched_uclamp_util_min) {
+ uclamp_bucket_inc(NULL, &uclamp_default[UCLAMP_MIN],
+ UCLAMP_MIN, sysctl_sched_uclamp_util_min);
+ }
+ if (old_max != sysctl_sched_uclamp_util_max) {
+ uclamp_bucket_inc(NULL, &uclamp_default[UCLAMP_MAX],
+ UCLAMP_MAX, sysctl_sched_uclamp_util_max);
+ }
Should you not update all tasks?
That's true, but that's also an expensive operation, that's why now
I'm doing only lazy updates at next enqueue time.
Do you think that could be acceptable?
Perhaps I can sanity check all the CPU to ensure that they all have a
current clamp value within the new enforced range. This kind-of
anticipate the idea to have an in-kernel API which has higher priority
and allows to set clamp values across all the CPUs...
--
#include <best/regards.h>
Patrick Bellasi
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-22 14:57:53
On Tue, Jan 22, 2019 at 02:01:15PM +0000, Patrick Bellasi wrote:
On 22-Jan 14:28, Peter Zijlstra wrote:
quoted
On Tue, Jan 22, 2019 at 10:43:05AM +0000, Patrick Bellasi wrote:
quoted
On 22-Jan 10:37, Peter Zijlstra wrote:
quoted
quoted
Sure, I get that. What I don't get is why you're adding that (2) here.
Like said, __sched_setscheduler() already does a dequeue/enqueue under
rq->lock, which should already take care of that.
Oh, ok... got it what you mean now.
With:
[PATCH v6 01/16] sched/core: Allow sched_setattr() to use the current policy
[off-list ref]
we can call __sched_setscheduler() with:
attr->sched_flags & SCHED_FLAG_KEEP_POLICY
whenever we want just to change the clamp values of a task without
changing its class. Thus, we can end up returning from
__sched_setscheduler() without doing an actual dequeue/enqueue.
I don't see that happening.. when KEEP_POLICY we set attr.sched_policy =
SETPARAM_POLICY. That is then checked again in __setscheduler_param(),
which is in the middle of that dequeue/enqueue.
Yes, I think I've forgot a check before we actually dequeue the task.
The current code does:
---8<---
SYSCALL_DEFINE3(sched_setattr)
// A) request to keep the same policy
if (attr.sched_flags & SCHED_FLAG_KEEP_POLICY)
attr.sched_policy = SETPARAM_POLICY;
sched_setattr()
// B) actually enforce the same policy
if (policy < 0)
policy = oldpolicy = p->policy;
// C) tune the clamp values
if (attr->sched_flags & SCHED_FLAG_UTIL_CLAMP)
retval = __setscheduler_uclamp(p, attr);
// D) tune attributes if policy is the same
if (unlikely(policy == p->policy))
if (fair_policy(policy) && attr->sched_nice != task_nice(p))
goto change;
if (rt_policy(policy) && attr->sched_priority != p->rt_priority)
goto change;
if (dl_policy(policy) && dl_param_changed(p, attr))
goto change;
if (util_changed)
goto change;
?
return 0;
change:
// E) dequeue/enqueue task
---8<---
So, probably in D) I've missed a check on SCHED_FLAG_KEEP_POLICY to
enforce a return in that case...
quoted
Also, and this might be 'broken', SETPARAM_POLICY _does_ reset all the
other attributes, it only preserves policy, but it will (re)set nice
level for example (see that same function).
Mmm... right... my bad! :/
quoted
So maybe we want to introduce another (few?) FLAG_KEEP flag(s) that
preserve the other bits; I'm thinking at least KEEP_PARAM and KEEP_UTIL
or something.
Yes, I would say we have two options:
1) SCHED_FLAG_KEEP_POLICY enforces all the scheduling class specific
attributes, but cross class attributes (e.g. uclamp)
2) add SCHED_KEEP_NICE, SCHED_KEEP_PRIO, and SCED_KEEP_PARAMS
and use them in the if conditions in D)
So the current KEEP_POLICY basically provides sched_setparam(), and
given we have that as a syscall, that is supposedly a useful
functionality.
Also, NICE/PRIO/DL* is all the same thing and depends on the policy,
KEEP_PARAM should cover the lot
And I suppose the UTIL_CLAMP is !KEEP_UTIL; we could go either way
around with that flag.
In both cases the goal should be to return from code block D).
I don't think so; we really do want to 'goto change' for util changes
too I think. Why duplicate part of that logic?
From: Patrick Bellasi <hidden> Date: 2019-01-22 15:01:45
On 22-Jan 14:39, Quentin Perret wrote:
On Tuesday 22 Jan 2019 at 14:26:06 (+0000), Patrick Bellasi wrote:
quoted
On 22-Jan 13:29, Quentin Perret wrote:
quoted
On Tuesday 22 Jan 2019 at 12:45:46 (+0000), Patrick Bellasi wrote:
quoted
On 22-Jan 12:13, Quentin Perret wrote:
quoted
On Tuesday 15 Jan 2019 at 10:15:08 (+0000), Patrick Bellasi wrote:
quoted
The Energy Aware Scheduler (AES) estimates the energy impact of waking
[...]
quoted
quoted
Ah, sorry, I guess my message was misleading. I'm saying this is
changing the way _EAS_ deals with RT tasks. Right now we don't actually
consider the RT-go-to-max thing at all in the EAS prediction. Your
patch is changing that AFAICT. It actually changes the way EAS sees RT
tasks even without uclamp ...
Lemme see if I get it right.
Currently, whenever we look at CPU utilization for ENERGY_UTIL, we
always use cpu_util_rt() for RT tasks:
---8<---
util = util_cfs;
util += cpu_util_rt(rq);
util += dl_util;
---8<---
Thus, even when RT tasks are RUNNABLE, we don't always assume the CPU
running at the max capacity but just whatever is the aggregated
utilization across all the classes.
With uclamp, we have:
---8<---
util = cpu_util_rt(rq) + util_cfs;
if (type == FREQUENCY_UTIL)
util = uclamp_util_with(rq, util, p);
dl_util = cpu_util_dl(rq);
if (type == ENERGY_UTIL)
util += dl_util;
---8<---
So, I would say that, in terms of ENERGY_UTIL we do the same both
w/ and w/o uclamp. Isn't it?
Yes but now you use FREQUENCY_UTIL for computing 'max_util' in the EAS
prediction.
Right, I overlook that "little" detail... :/
Let's take an example. You have a perf domain with two CPUs. One CPU is
busy running a RT task, the other CPU runs a CFS task. Right now in
compute_energy() we only use ENERGY_UTIL, so 'max_util' ends up being
the max between the utilization of the two tasks. So we don't predict
we're going to max freq.
+1
With your patch, we use FREQUENCY_UTIL to compute 'max_util', so we
_will_ predict that we're going to max freq.
Right, with the default conf yes.
And we will do that even if CONFIG_UCLAMP_TASK=n.
While this should not happen, as I wrote in the RT integration patch,
that's happening because I'm missing some compilation guard or
similar. In this configurations we should always go to max... will
look into that.
The default EAS calculation will be different with your patch when there
are runnable RT tasks in the system. This needs to be documented, I
think.
Sure...
quoted
quoted
But I'm not hostile to the idea since it's possible to enable uclamp and
set the min cap to 0 for RT if you want. So there is a story there.
However, I think this needs be documented somewhere, at the very least.
The only difference I see is that the actual frequency could be
different (lower then max) when a clamped RT task is RUNNABLE.
Are you worried that running RT on a lower freq could have side
effects on the estimated busy time the CPU ?
I also still don't completely get why you say it could be useful to
"set the min cap to 0 for RT if you want"
I'm not saying it's useful, I'm saying userspace can decide to do that
if it thinks it is a good idea. The default should be min_cap = 1024 for
RT, no questions. But you _can_ change it at runtime if you want to.
That's my point. And doing that basically provides the same behaviour as
what we have right now in terms of EAS calculation (but it changes the
freq selection obviously) which is why I'm not fundamentally opposed to
your patch.
Well, I think it's tricky to say whether the current or new approach
is better... it probably depends on the use-case.
So in short, I'm fine with the behavioural change, but please at least
mention it somewhere :-)
Anyway... agree, it's just that to add some documentation I need to
get what you are pointing out ;)
Will come up with some additional text to be added to the changelog
Maybe we can add a more detailed explanation of the different
behaviors you can get in the EAS documentation which is coming to
mainline ?
@@ -625,6 +625,11 @@ struct uclamp_se {unsignedintbucket_id:bits_per(UCLAMP_BUCKETS);unsignedintmapped:1;unsignedintactive:1;+/* Clamp bucket and value actually used by a RUNNABLE task */+struct{+unsignedintvalue:bits_per(SCHED_CAPACITY_SCALE);+unsignedintbucket_id:bits_per(UCLAMP_BUCKETS);+}effective;
I am confuzled by this thing.. so uclamp_se already has a value,bucket,
which per the prior code is the effective one.
Now; I think I see why you want another value; you need the second to
store the original value for when the system limits change and we must
re-evaluate.
Yes, that's one reason, the other one being to properly support
CGroup when we add them in the following patches.
Effective will always track the value/bucket in which the task has
been refcounted at enqueue time and it depends on the aggregated
value.
quoted
Should you not update all tasks?
That's true, but that's also an expensive operation, that's why now
I'm doing only lazy updates at next enqueue time.
Aaah, so you refcount on the original value, which allows you to skip
fixing up all tasks. I missed that bit.
Do you think that could be acceptable?
Think so, it's a sysctl poke, 'nobody' ever does that.
On Tuesday 22 Jan 2019 at 15:01:37 (+0000), Patrick Bellasi wrote:
quoted
I'm not saying it's useful, I'm saying userspace can decide to do that
if it thinks it is a good idea. The default should be min_cap = 1024 for
RT, no questions. But you _can_ change it at runtime if you want to.
That's my point. And doing that basically provides the same behaviour as
what we have right now in terms of EAS calculation (but it changes the
freq selection obviously) which is why I'm not fundamentally opposed to
your patch.
Well, I think it's tricky to say whether the current or new approach
is better... it probably depends on the use-case.
Agreed.
quoted
So in short, I'm fine with the behavioural change, but please at least
mention it somewhere :-)
Anyway... agree, it's just that to add some documentation I need to
get what you are pointing out ;)
Will come up with some additional text to be added to the changelog
Sounds good.
Maybe we can add a more detailed explanation of the different
behaviors you can get in the EAS documentation which is coming to
mainline ?
Yeah, if you feel like it, I guess that won't hurt :-)
Thanks,
Quentin
@@ -218,8 +218,15 @@ unsigned long schedutil_freq_util(int cpu, unsigned long util_cfs,*CFStasksandweusethesamemetrictotracktheeffective*utilization(PELTwindowsaresynchronized)wecandirectlyaddthem*toobtaintheCPU'sactualutilization.+*+*CFSutilizationcanbeboostedorcapped,dependingonutilization+*clampconstraintsrequestedbycurrentlyRUNNABLEtasks.+*WhentherearenoCFSRUNNABLEtasks,clampsarereleasedand+*frequencywillbegracefullyreducedwiththeutilizationdecay.*/-util=util_cfs;+util=(type==ENERGY_UTIL)+?util_cfs+:uclamp_util(rq,util_cfs);
That's pretty horrible; what's wrong with:
util = util_cfs;
if (type == FREQUENCY_UTIL)
util = uclamp_util(rq, util);
That should generate the same code, but is (IMO) far easier to read.
From: Patrick Bellasi <hidden> Date: 2019-01-22 15:33:23
On 22-Jan 15:57, Peter Zijlstra wrote:
On Tue, Jan 22, 2019 at 02:01:15PM +0000, Patrick Bellasi wrote:
quoted
On 22-Jan 14:28, Peter Zijlstra wrote:
quoted
On Tue, Jan 22, 2019 at 10:43:05AM +0000, Patrick Bellasi wrote:
quoted
On 22-Jan 10:37, Peter Zijlstra wrote:
quoted
quoted
Sure, I get that. What I don't get is why you're adding that (2) here.
Like said, __sched_setscheduler() already does a dequeue/enqueue under
rq->lock, which should already take care of that.
Oh, ok... got it what you mean now.
With:
[PATCH v6 01/16] sched/core: Allow sched_setattr() to use the current policy
[off-list ref]
we can call __sched_setscheduler() with:
attr->sched_flags & SCHED_FLAG_KEEP_POLICY
whenever we want just to change the clamp values of a task without
changing its class. Thus, we can end up returning from
__sched_setscheduler() without doing an actual dequeue/enqueue.
I don't see that happening.. when KEEP_POLICY we set attr.sched_policy =
SETPARAM_POLICY. That is then checked again in __setscheduler_param(),
which is in the middle of that dequeue/enqueue.
Yes, I think I've forgot a check before we actually dequeue the task.
The current code does:
---8<---
SYSCALL_DEFINE3(sched_setattr)
// A) request to keep the same policy
if (attr.sched_flags & SCHED_FLAG_KEEP_POLICY)
attr.sched_policy = SETPARAM_POLICY;
sched_setattr()
// B) actually enforce the same policy
if (policy < 0)
policy = oldpolicy = p->policy;
// C) tune the clamp values
if (attr->sched_flags & SCHED_FLAG_UTIL_CLAMP)
retval = __setscheduler_uclamp(p, attr);
// D) tune attributes if policy is the same
if (unlikely(policy == p->policy))
if (fair_policy(policy) && attr->sched_nice != task_nice(p))
goto change;
if (rt_policy(policy) && attr->sched_priority != p->rt_priority)
goto change;
if (dl_policy(policy) && dl_param_changed(p, attr))
goto change;
if (util_changed)
goto change;
?
quoted
return 0;
change:
// E) dequeue/enqueue task
---8<---
So, probably in D) I've missed a check on SCHED_FLAG_KEEP_POLICY to
enforce a return in that case...
quoted
Also, and this might be 'broken', SETPARAM_POLICY _does_ reset all the
other attributes, it only preserves policy, but it will (re)set nice
level for example (see that same function).
Mmm... right... my bad! :/
quoted
So maybe we want to introduce another (few?) FLAG_KEEP flag(s) that
preserve the other bits; I'm thinking at least KEEP_PARAM and KEEP_UTIL
or something.
Yes, I would say we have two options:
1) SCHED_FLAG_KEEP_POLICY enforces all the scheduling class specific
attributes, but cross class attributes (e.g. uclamp)
2) add SCHED_KEEP_NICE, SCHED_KEEP_PRIO, and SCED_KEEP_PARAMS
and use them in the if conditions in D)
So the current KEEP_POLICY basically provides sched_setparam(), and
But it's not exposed user-space.
given we have that as a syscall, that is supposedly a useful
functionality.
For uclamp is definitively useful to change clamps without the need to
read beforehand the current policy params and use them in a following
set syscall... which is racy pattern.
Also, NICE/PRIO/DL* is all the same thing and depends on the policy,
KEEP_PARAM should cover the lot
Right, that makes sense.
And I suppose the UTIL_CLAMP is !KEEP_UTIL; we could go either way
around with that flag.
What about getting rid of the racy case above by exposing userspace
only the new UTIL_CLAMP and, on:
sched_setscheduler(flags: UTIL_CLAMP)
we enforce the other two flags from the syscall:
---8<---
SYSCALL_DEFINE3(sched_setattr)
if (attr.sched_flags & SCHED_FLAG_KEEP_POLICY) {
attr.sched_policy = SETPARAM_POLICY;
attr.sched_flags |= (KEEP_POLICY|KEEP_PARAMS);
}
---8<---
This will not make possible to change class and set flags in one go,
but honestly that's likely a very limited use-case, isn't it ?
quoted
In both cases the goal should be to return from code block D).
I don't think so; we really do want to 'goto change' for util changes
too I think. Why duplicate part of that logic?
But that will force a dequeue/enqueue... isn't too much overhead just
to change a clamp value? Perhaps we can also end up with some wired
side-effects like the task being preempted ?
Consider also that the uclamp_task_update_active() added by this patch
not only has lower overhead but it will be use also by cgroups where
we want to force update all the tasks on a cgroup's clamp change.
--
#include <best/regards.h>
Patrick Bellasi
@@ -625,6 +625,11 @@ struct uclamp_se {unsignedintbucket_id:bits_per(UCLAMP_BUCKETS);unsignedintmapped:1;unsignedintactive:1;+/* Clamp bucket and value actually used by a RUNNABLE task */+struct{+unsignedintvalue:bits_per(SCHED_CAPACITY_SCALE);+unsignedintbucket_id:bits_per(UCLAMP_BUCKETS);+}effective;
I am confuzled by this thing.. so uclamp_se already has a value,bucket,
which per the prior code is the effective one.
Now; I think I see why you want another value; you need the second to
store the original value for when the system limits change and we must
re-evaluate.
Yes, that's one reason, the other one being to properly support
CGroup when we add them in the following patches.
Effective will always track the value/bucket in which the task has
been refcounted at enqueue time and it depends on the aggregated
value.
quoted
quoted
Should you not update all tasks?
That's true, but that's also an expensive operation, that's why now
I'm doing only lazy updates at next enqueue time.
Aaah, so you refcount on the original value, which allows you to skip
fixing up all tasks. I missed that bit.
Right, effective is always tracking the bucket we refcounted at
enqueue time.
We can still argue that, the moment we change a clamp, a task should
be updated without waiting for a dequeue/enqueue cycle.
IMO, that could be a limitation only for tasks which never sleep, but
that's a very special case.
Instead, as you'll see, in the cgroup integration we force update all
RUNNABLE tasks. Although that's expensive, since we are in the domain
of the "delegation model" and "containers resources control", there
it's probably more worth to pay than here.
quoted
Do you think that could be acceptable?
Think so, it's a sysctl poke, 'nobody' ever does that.
Cool, so... I'll keep lazy update for system default.
--
#include <best/regards.h>
Patrick Bellasi
@@ -218,8 +218,15 @@ unsigned long schedutil_freq_util(int cpu, unsigned long util_cfs,*CFStasksandweusethesamemetrictotracktheeffective*utilization(PELTwindowsaresynchronized)wecandirectlyaddthem*toobtaintheCPU'sactualutilization.+*+*CFSutilizationcanbeboostedorcapped,dependingonutilization+*clampconstraintsrequestedbycurrentlyRUNNABLEtasks.+*WhentherearenoCFSRUNNABLEtasks,clampsarereleasedand+*frequencywillbegracefullyreducedwiththeutilizationdecay.*/-util=util_cfs;+util=(type==ENERGY_UTIL)+?util_cfs+:uclamp_util(rq,util_cfs);
That's pretty horrible; what's wrong with:
util = util_cfs;
if (type == FREQUENCY_UTIL)
util = uclamp_util(rq, util);
That should generate the same code, but is (IMO) far easier to read.
Yes, right... and that's also the pattern we end up with the
following patch on RT integration.
However, as suggested by Rafael, I'll squash these two patches
together and we will get rid of the above for free ;)
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-22 17:13:29
On Tue, Jan 15, 2019 at 10:15:05AM +0000, Patrick Bellasi wrote:
quoted hunk
@@ -342,11 +350,24 @@ static void sugov_iowait_boost(struct sugov_cpu *sg_cpu, u64 time, return; sg_cpu->iowait_boost_pending = true;+ /*+ * Boost FAIR tasks only up to the CPU clamped utilization.+ *+ * Since DL tasks have a much more advanced bandwidth control, it's+ * safe to assume that IO boost does not apply to those tasks.
I'm not buying that argument. IO-boost isn't related to b/w management.
IO-boot is more about compensating for hidden dependencies, and those
don't get less hidden for using a different scheduling class.
Now, arguably DL should not be doing IO in the first place, but that's a
whole different discussion.
+ * Instead, since RT tasks are not utilization clamped, we don't want
+ * to apply clamping on IO boost while there is blocked RT
+ * utilization.
+ */
+ max_boost = sg_cpu->iowait_boost_max;
+ if (!cpu_util_rt(cpu_rq(sg_cpu->cpu)))
+ max_boost = uclamp_util(cpu_rq(sg_cpu->cpu), max_boost);
+
/* Double the boost at each request */
if (sg_cpu->iowait_boost) {
sg_cpu->iowait_boost <<= 1;
- if (sg_cpu->iowait_boost > sg_cpu->iowait_boost_max)
- sg_cpu->iowait_boost = sg_cpu->iowait_boost_max;
+ if (sg_cpu->iowait_boost > max_boost)
+ sg_cpu->iowait_boost = max_boost;
return;
}
From: Patrick Bellasi <hidden> Date: 2019-01-22 18:18:39
On 22-Jan 18:13, Peter Zijlstra wrote:
On Tue, Jan 15, 2019 at 10:15:05AM +0000, Patrick Bellasi wrote:
quoted
@@ -342,11 +350,24 @@ static void sugov_iowait_boost(struct sugov_cpu *sg_cpu, u64 time, return; sg_cpu->iowait_boost_pending = true;+ /*+ * Boost FAIR tasks only up to the CPU clamped utilization.+ *+ * Since DL tasks have a much more advanced bandwidth control, it's+ * safe to assume that IO boost does not apply to those tasks.
I'm not buying that argument. IO-boost isn't related to b/w management.
IO-boot is more about compensating for hidden dependencies, and those
don't get less hidden for using a different scheduling class.
Now, arguably DL should not be doing IO in the first place, but that's a
whole different discussion.
My understanding is that IOBoost is there to help tasks doing many
and _frequent_ IO operations, which are relatively _not so much_
computational intensive on the CPU.
Those tasks generate a small utilization and, without IOBoost, will be
executed at a lower frequency and will add undesired latency on
triggering the next IO operation.
Isn't mainly that the reason for it?
IO operations have also to be _frequent_ since we don't got to max OPP
at the very first wakeup from IO. We double frequency and get to max
only if we have a stable stream of IO operations.
IMHO, it makes perfectly sense to use DL for these kind of operations
but I would expect that, since you care about latency we should come
up with a proper description of the required bandwidth... eventually
accounting for an additional headroom to compensate for "hidden
dependencies"... without relaying on a quite dummy policy like
IOBoost to get our DL tasks working.
At the end, DL is now quite good in driving the freq as high has it
needs... and by closing userspace feedback loops it can also
compensate for all sort of fluctuations and noise... as demonstrated
by Alessio during last OSPM:
http://retis.sssup.it/luca/ospm-summit/2018/Downloads/OSPM_deadline_audio.pdf
quoted
+ * Instead, since RT tasks are not utilization clamped, we don't want
+ * to apply clamping on IO boost while there is blocked RT
+ * utilization.
+ */
+ max_boost = sg_cpu->iowait_boost_max;
+ if (!cpu_util_rt(cpu_rq(sg_cpu->cpu)))
+ max_boost = uclamp_util(cpu_rq(sg_cpu->cpu), max_boost);
+
/* Double the boost at each request */
if (sg_cpu->iowait_boost) {
sg_cpu->iowait_boost <<= 1;
- if (sg_cpu->iowait_boost > sg_cpu->iowait_boost_max)
- sg_cpu->iowait_boost = sg_cpu->iowait_boost_max;
+ if (sg_cpu->iowait_boost > max_boost)
+ sg_cpu->iowait_boost = max_boost;
return;
}
Hurmph... so I'm not sold on this bit.
If a task is not clamped we execute it at its required utilization or
even max frequency in case of wakeup from IO.
When a task is util_max clamped instead, we are saying that we don't
care to run it above the specified clamp value and, if possible, we
should run it below that capacity level.
If that's the case, why this clamping hints should not be enforced on
IO wakeups too?
At the end it's still a user-space decision, we basically allow
userspace to defined what's the max IO boost they like to get.
--
#include <best/regards.h>
Patrick Bellasi
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-23 09:16:43
On Tue, Jan 22, 2019 at 03:33:15PM +0000, Patrick Bellasi wrote:
On 22-Jan 15:57, Peter Zijlstra wrote:
quoted
On Tue, Jan 22, 2019 at 02:01:15PM +0000, Patrick Bellasi wrote:
quoted
quoted
Yes, I would say we have two options:
1) SCHED_FLAG_KEEP_POLICY enforces all the scheduling class specific
attributes, but cross class attributes (e.g. uclamp)
2) add SCHED_KEEP_NICE, SCHED_KEEP_PRIO, and SCED_KEEP_PARAMS
and use them in the if conditions in D)
So the current KEEP_POLICY basically provides sched_setparam(), and
But it's not exposed user-space.
Correct; not until your first patch indeed.
quoted
given we have that as a syscall, that is supposedly a useful
functionality.
For uclamp is definitively useful to change clamps without the need to
read beforehand the current policy params and use them in a following
set syscall... which is racy pattern.
Right; but my argument was mostly that if sched_setparam() is a useful
interface, a 'pure' KEEP_POLICY would be too and your (1) looses that.
quoted
And I suppose the UTIL_CLAMP is !KEEP_UTIL; we could go either way
around with that flag.
What about getting rid of the racy case above by exposing userspace
only the new UTIL_CLAMP and, on:
sched_setscheduler(flags: UTIL_CLAMP)
we enforce the other two flags from the syscall:
---8<---
SYSCALL_DEFINE3(sched_setattr)
if (attr.sched_flags & SCHED_FLAG_KEEP_POLICY) {
attr.sched_policy = SETPARAM_POLICY;
attr.sched_flags |= (KEEP_POLICY|KEEP_PARAMS);
}
---8<---
This will not make possible to change class and set flags in one go,
but honestly that's likely a very limited use-case, isn't it ?
So I must admit to not seeing much use for sched_setparam() (and its
equivalents) myself, but given it is an existing interface, I also think
it would be nice to cover that functionality in the sched_setattr()
call.
That is; I know of userspace priority-ceiling implementations using
sched_setparam(), but I don't see any reason why that wouldn't also work
with sched_setscheduler() (IOW always also set the policy).
quoted
quoted
In both cases the goal should be to return from code block D).
I don't think so; we really do want to 'goto change' for util changes
too I think. Why duplicate part of that logic?
But that will force a dequeue/enqueue... isn't too much overhead just
to change a clamp value?
These syscalls aren't what I consider fast paths anyway. However, there
are people that rely on the scheduler syscalls not to schedule
themselves, or rather be non-blocking (see for example that prio-ceiling
implementation).
And in that respect the newly introduced uclamp_mutex does appear to be
a problem.
Also; do you expect these clamp values to be changed often?
Perhaps we can also end up with some wired
s/wired/weird/, right?
side-effects like the task being preempted ?
Nothing worse than any other random reschedule would cause.
Consider also that the uclamp_task_update_active() added by this patch
not only has lower overhead but it will be use also by cgroups where
we want to force update all the tasks on a cgroup's clamp change.
I haven't gotten that far; but I would prefer not to have two different
'change' paths in __sched_setscheduler().
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-23 09:22:21
On Tue, Jan 22, 2019 at 03:41:29PM +0000, Patrick Bellasi wrote:
On 22-Jan 16:13, Peter Zijlstra wrote:
quoted
On Tue, Jan 22, 2019 at 02:43:29PM +0000, Patrick Bellasi wrote:
quoted
quoted
Do you think that could be acceptable?
Think so, it's a sysctl poke, 'nobody' ever does that.
Cool, so... I'll keep lazy update for system default.
Ah, I think I misunderstood. I meant to say that since nobody ever pokes
at sysctl's it doesn't matter if its a little more expensive and iterate
everything.
Also; if you always keep everything up-to-date, you can avoid doing that
duplicate accounting.
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-23 09:52:35
On Tue, Jan 22, 2019 at 06:18:31PM +0000, Patrick Bellasi wrote:
On 22-Jan 18:13, Peter Zijlstra wrote:
quoted
On Tue, Jan 15, 2019 at 10:15:05AM +0000, Patrick Bellasi wrote:
quoted
@@ -342,11 +350,24 @@ static void sugov_iowait_boost(struct sugov_cpu *sg_cpu, u64 time, return; sg_cpu->iowait_boost_pending = true;+ /*+ * Boost FAIR tasks only up to the CPU clamped utilization.+ *+ * Since DL tasks have a much more advanced bandwidth control, it's+ * safe to assume that IO boost does not apply to those tasks.
I'm not buying that argument. IO-boost isn't related to b/w management.
IO-boot is more about compensating for hidden dependencies, and those
don't get less hidden for using a different scheduling class.
Now, arguably DL should not be doing IO in the first place, but that's a
whole different discussion.
My understanding is that IOBoost is there to help tasks doing many
and _frequent_ IO operations, which are relatively _not so much_
computational intensive on the CPU.
Those tasks generate a small utilization and, without IOBoost, will be
executed at a lower frequency and will add undesired latency on
triggering the next IO operation.
Isn't mainly that the reason for it?
http://lkml.kernel.org/r/20170522082154.f57cqovterd2qajv@hirez.programming.kicks-ass.net
Using a lower frequency will allow the IO device to go idle while we try
and get the next request going.
The connection between IO device and task/freq selection is hidden/lost.
We could potentially do better here, but fundamentally a completion
doesn't have an 'owner', there can be multiple waiters etc.
We loose (through our software architecture, and this we could possibly
improve, although it would be fairly invasive) the device busy state,
and it would be the device that raises the CPU frequency (to the point
where request submission is no longer the bottle neck to staying busy).
Currently all we do is mark a task as sleeping on IO and loose any
and all device relations/metrics.
So I don't think the task clamping should affect the IO boosting, as
that is meant to represent the device state, not the task utilization.
IMHO, it makes perfectly sense to use DL for these kind of operations
but I would expect that, since you care about latency we should come
up with a proper description of the required bandwidth... eventually
accounting for an additional headroom to compensate for "hidden
dependencies"... without relaying on a quite dummy policy like
IOBoost to get our DL tasks working.
Deadline is about determinsm, (file/disk) IO is typically the
anti-thesis of that.
Audio is a special in that it is indeed a deterministic device, also, I
don't think ALSA touches the IO-wait code, that is typically all
filesystem stuff.
quoted
quoted
+ * Instead, since RT tasks are not utilization clamped, we don't want
+ * to apply clamping on IO boost while there is blocked RT
+ * utilization.
+ */
+ max_boost = sg_cpu->iowait_boost_max;
+ if (!cpu_util_rt(cpu_rq(sg_cpu->cpu)))
+ max_boost = uclamp_util(cpu_rq(sg_cpu->cpu), max_boost);
+
/* Double the boost at each request */
if (sg_cpu->iowait_boost) {
sg_cpu->iowait_boost <<= 1;
- if (sg_cpu->iowait_boost > sg_cpu->iowait_boost_max)
- sg_cpu->iowait_boost = sg_cpu->iowait_boost_max;
+ if (sg_cpu->iowait_boost > max_boost)
+ sg_cpu->iowait_boost = max_boost;
return;
}
Hurmph... so I'm not sold on this bit.
If a task is not clamped we execute it at its required utilization or
even max frequency in case of wakeup from IO.
When a task is util_max clamped instead, we are saying that we don't
care to run it above the specified clamp value and, if possible, we
should run it below that capacity level.
If that's the case, why this clamping hints should not be enforced on
IO wakeups too?
At the end it's still a user-space decision, we basically allow
userspace to defined what's the max IO boost they like to get.
Because it is the wrong knob for it.
Ideally we'd extend the IO-wait state to include the device-busy state
at the time of sleep. At the very least double state io_schedule() state
space from 1 to 2 bits, where we not only indicate: yes this is an
IO-sleep, but also can indicate device saturation. When the device is
saturated, we don't need to boost further.
(this binary state will ofcourse cause oscilations where we drop the
freq, drop device saturation, then ramp the freq, regain device
saturation etc..)
However, doing this is going to require fairly massive surgery on our
whole IO stack.
Also; how big of a problem is 'supriouos' boosting really? Joel tried to
introduce a boost_max tunable, but the grandual boosting thing was good
enough at the time.
Or the combined thing:
- util = util_cfs;
- util += cpu_util_rt(rq);
+ util = cpu_util_rt(rq);
+ if (type == FREQUENCY_UTIL) {
+ util += cpu_util_cfs(rq);
+ util = uclamp_util(rq, util);
+ } else {
+ util += util_cfs;
+ }
Leaves me confused.
When type == FREQ, util_cfs should already be cpu_util_cfs(), per
sugov_get_util().
So should that not end up like:
util = util_cfs;
util += cpu_util_rt(rq);
+ if (type == FREQUENCY_UTIL)
+ util = uclamp_util(rq, util);
instead?
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-23 13:33:44
On Tue, Jan 15, 2019 at 10:15:07AM +0000, Patrick Bellasi wrote:
+static __always_inline
+unsigned int uclamp_util_with(struct rq *rq, unsigned int util,
+ struct task_struct *p)
{
unsigned int min_util = READ_ONCE(rq->uclamp[UCLAMP_MIN].value);
unsigned int max_util = READ_ONCE(rq->uclamp[UCLAMP_MAX].value);
+ if (p) {
+ min_util = max(min_util, uclamp_effective_value(p, UCLAMP_MIN));
+ max_util = max(max_util, uclamp_effective_value(p, UCLAMP_MAX));
+ }
+
Like I think you mentioned earlier; this doesn't look right at all.
Should that not be something like:
lo = READ_ONCE(rq->uclamp[UCLAMP_MIN].value);
hi = READ_ONCE(rq->uclamp[UCLAMP_MAX].value);
min_util = clamp(uclamp_effective(p, UCLAMP_MIN), lo, hi);
max_util = clamp(uclamp_effective(p, UCLAMP_MAX), lo, hi);
From: Patrick Bellasi <hidden> Date: 2019-01-23 14:14:35
On 23-Jan 10:16, Peter Zijlstra wrote:
On Tue, Jan 22, 2019 at 03:33:15PM +0000, Patrick Bellasi wrote:
quoted
On 22-Jan 15:57, Peter Zijlstra wrote:
quoted
On Tue, Jan 22, 2019 at 02:01:15PM +0000, Patrick Bellasi wrote:
quoted
quoted
quoted
Yes, I would say we have two options:
1) SCHED_FLAG_KEEP_POLICY enforces all the scheduling class specific
attributes, but cross class attributes (e.g. uclamp)
2) add SCHED_KEEP_NICE, SCHED_KEEP_PRIO, and SCED_KEEP_PARAMS
and use them in the if conditions in D)
So the current KEEP_POLICY basically provides sched_setparam(), and
But it's not exposed user-space.
Correct; not until your first patch indeed.
quoted
quoted
given we have that as a syscall, that is supposedly a useful
functionality.
For uclamp is definitively useful to change clamps without the need to
read beforehand the current policy params and use them in a following
set syscall... which is racy pattern.
Right; but my argument was mostly that if sched_setparam() is a useful
interface, a 'pure' KEEP_POLICY would be too and your (1) looses that.
Ok, that's an argument in favour of option (2).
quoted
quoted
And I suppose the UTIL_CLAMP is !KEEP_UTIL; we could go either way
around with that flag.
What about getting rid of the racy case above by exposing userspace
only the new UTIL_CLAMP and, on:
sched_setscheduler(flags: UTIL_CLAMP)
we enforce the other two flags from the syscall:
---8<---
SYSCALL_DEFINE3(sched_setattr)
if (attr.sched_flags & SCHED_FLAG_KEEP_POLICY) {
attr.sched_policy = SETPARAM_POLICY;
attr.sched_flags |= (KEEP_POLICY|KEEP_PARAMS);
}
---8<---
This will not make possible to change class and set flags in one go,
but honestly that's likely a very limited use-case, isn't it ?
So I must admit to not seeing much use for sched_setparam() (and its
equivalents) myself, but given it is an existing interface, I also think
it would be nice to cover that functionality in the sched_setattr()
call.
Which will make them sort-of equivalent... meaning: both the POSIX
sched_setparam() and the !POSIX sched_setattr() will allow to change
params/attributes without changing the policy.
That is; I know of userspace priority-ceiling implementations using
sched_setparam(), but I don't see any reason why that wouldn't also work
with sched_setscheduler() (IOW always also set the policy).
The sched_setscheduler() requires a policy to be explicitely defined,
it's a mandatory parameter and has to be specified.
Unless a RT task could be blocked by a FAIR one and you need
sched_setscheduler() to boost both prio and class (which looks like a
poor RT design to begin with) why would you use sched_setscheduler()
instead of sched_setparam()?
They are both POSIX calls and, AFAIU, sched_setparam() seems to be
designed exactly for those kind on use cases.
quoted
quoted
quoted
In both cases the goal should be to return from code block D).
I don't think so; we really do want to 'goto change' for util changes
too I think. Why duplicate part of that logic?
But that will force a dequeue/enqueue... isn't too much overhead just
to change a clamp value?
These syscalls aren't what I consider fast paths anyway. However, there
are people that rely on the scheduler syscalls not to schedule
themselves, or rather be non-blocking (see for example that prio-ceiling
implementation).
And in that respect the newly introduced uclamp_mutex does appear to be
a problem.
Mmm... could be... I'll look better into it. Could be that that mutex
is not really required. We don't need to serialize task specific
clamp changes and anyway the protected code never sleeps and uses
atomic instruction.
Also; do you expect these clamp values to be changed often?
Not really, the most common use cases are:
a) a resource manager (e.g. the Android run-time) set clamps for a
bunch of tasks whenever you switch for example from one app to
antoher... but that will be done via cgroups (i.e. different path)
b) a task can relax his constraints to save energy (something
conceptually similar to use a deferrable timers)
In both cases I expect a limited call frequency.
quoted
Perhaps we can also end up with some wired
s/wired/weird/, right?
Right :)
quoted
side-effects like the task being preempted ?
Nothing worse than any other random reschedule would cause.
quoted
Consider also that the uclamp_task_update_active() added by this patch
not only has lower overhead but it will be use also by cgroups where
we want to force update all the tasks on a cgroup's clamp change.
I haven't gotten that far; but I would prefer not to have two different
'change' paths in __sched_setscheduler().
Yes, I agree that two paths in __sched_setscheduler() could be
confusing. Still we have to consider that here we are adding
"not class specific" attributes.
What if we keep "not class specific" code completely outside of
__sched_setscheduler() and do something like:
---8<---
int sched_setattr(struct task_struct *p, const struct sched_attr *attr)
{
retval = __sched_setattr(p, attr);
if (retval)
return retval;
return __sched_setscheduler(p, attr, true, true);
}
EXPORT_SYMBOL_GPL(sched_setattr);
---8<---
where __sched_setattr() will collect all the tunings which do not
require an enqueue/dequeue, so far only the new uclamp settings, while
the rest remains under __sched_setscheduler().
Thoughts ?
--
#include <best/regards.h>
Patrick Bellasi
From: Patrick Bellasi <hidden> Date: 2019-01-23 14:19:33
On 23-Jan 10:22, Peter Zijlstra wrote:
On Tue, Jan 22, 2019 at 03:41:29PM +0000, Patrick Bellasi wrote:
quoted
On 22-Jan 16:13, Peter Zijlstra wrote:
quoted
On Tue, Jan 22, 2019 at 02:43:29PM +0000, Patrick Bellasi wrote:
quoted
quoted
quoted
Do you think that could be acceptable?
Think so, it's a sysctl poke, 'nobody' ever does that.
Cool, so... I'll keep lazy update for system default.
Ah, I think I misunderstood. I meant to say that since nobody ever pokes
at sysctl's it doesn't matter if its a little more expensive and iterate
everything.
Here I was more worried about the code complexity/overhead... for
something actually not very used/useful.
Also; if you always keep everything up-to-date, you can avoid doing that
duplicate accounting.
To update everything we will have to walk all the CPUs and update all
the RUNNABLE tasks currently enqueued, which are either RT or CFS.
That's way more expensive both in code and time then what we do for
cgroups, where at least we have a limited scope since the cgroup
already provides a (usually limited) list of tasks to consider.
Do you think it's really worth to have ?
Perhaps we can add it in a second step, once we have the core bits in
and we really see a need for a specific use-case.
--
#include <best/regards.h>
Patrick Bellasi
From: Patrick Bellasi <hidden> Date: 2019-01-23 14:25:04
On 23-Jan 10:52, Peter Zijlstra wrote:
On Tue, Jan 22, 2019 at 06:18:31PM +0000, Patrick Bellasi wrote:
quoted
On 22-Jan 18:13, Peter Zijlstra wrote:
quoted
On Tue, Jan 15, 2019 at 10:15:05AM +0000, Patrick Bellasi wrote:
[...]
quoted
If a task is not clamped we execute it at its required utilization or
even max frequency in case of wakeup from IO.
When a task is util_max clamped instead, we are saying that we don't
care to run it above the specified clamp value and, if possible, we
should run it below that capacity level.
If that's the case, why this clamping hints should not be enforced on
IO wakeups too?
At the end it's still a user-space decision, we basically allow
userspace to defined what's the max IO boost they like to get.
Because it is the wrong knob for it.
Ideally we'd extend the IO-wait state to include the device-busy state
at the time of sleep. At the very least double state io_schedule() state
space from 1 to 2 bits, where we not only indicate: yes this is an
IO-sleep, but also can indicate device saturation. When the device is
saturated, we don't need to boost further.
(this binary state will ofcourse cause oscilations where we drop the
freq, drop device saturation, then ramp the freq, regain device
saturation etc..)
However, doing this is going to require fairly massive surgery on our
whole IO stack.
Also; how big of a problem is 'supriouos' boosting really? Joel tried to
introduce a boost_max tunable, but the grandual boosting thing was good
enough at the time.
Ok then, I'll drop the clamp on IOBoost... you right and moreover we
can always investigate for a better solution in the future with a
real use-case on hand.
Cheers.
--
#include <best/regards.h>
Patrick Bellasi
Or the combined thing:
- util = util_cfs;
- util += cpu_util_rt(rq);
+ util = cpu_util_rt(rq);
+ if (type == FREQUENCY_UTIL) {
+ util += cpu_util_cfs(rq);
+ util = uclamp_util(rq, util);
+ } else {
+ util += util_cfs;
+ }
Leaves me confused.
When type == FREQ, util_cfs should already be cpu_util_cfs(), per
sugov_get_util().
So should that not end up like:
util = util_cfs;
util += cpu_util_rt(rq);
+ if (type == FREQUENCY_UTIL)
+ util = uclamp_util(rq, util);
instead?
You right, I get to that core after the patches which integrate
compute_energy(). The chuck above was the version before the EM got
merged but I missed to backport the change once I rebased on
tip/sched/core with the EM in.
Sorry for the confusion, will fix in v7.
Cheers
--
#include <best/regards.h>
Patrick Bellasi
From: Patrick Bellasi <hidden> Date: 2019-01-23 14:40:19
On 23-Jan 11:49, Peter Zijlstra wrote:
On Tue, Jan 15, 2019 at 10:15:06AM +0000, Patrick Bellasi wrote:
quoted
@@ -858,16 +859,23 @@ static inline void uclamp_effective_get(struct task_struct *p, unsigned int clamp_id, unsigned int *clamp_value, unsigned int *bucket_id) {+ struct uclamp_se *default_clamp;+ /* Task specific clamp value */ *clamp_value = p->uclamp[clamp_id].value; *bucket_id = p->uclamp[clamp_id].bucket_id;+ /* RT tasks have different default values */+ default_clamp = task_has_rt_policy(p)+ ? uclamp_default_perf+ : uclamp_default;+ /* System default restriction */- if (unlikely(*clamp_value < uclamp_default[UCLAMP_MIN].value ||- *clamp_value > uclamp_default[UCLAMP_MAX].value)) {+ if (unlikely(*clamp_value < default_clamp[UCLAMP_MIN].value ||+ *clamp_value > default_clamp[UCLAMP_MAX].value)) { /* Keep it simple: unconditionally enforce system defaults */- *clamp_value = uclamp_default[clamp_id].value;- *bucket_id = uclamp_default[clamp_id].bucket_id;+ *clamp_value = default_clamp[clamp_id].value;+ *bucket_id = default_clamp[clamp_id].bucket_id; } }
So I still don't much like the whole effective thing;
:/
I find back-annotation useful in many cases since we have different
sources for possible clamp values:
1. task specific
2. cgroup defined
3. system defaults
4. system power default
I don't think we can avoid to somehow back annotate on which bucket a
task has been refcounted... it makes dequeue so much easier, it helps
in ensuring that the refcouning is consistent and enable lazy updates.
but I think you should use rt_task() instead of
task_has_rt_policy().
Right... will do.
--
#include <best/regards.h>
Patrick Bellasi
From: Patrick Bellasi <hidden> Date: 2019-01-23 14:51:14
On 23-Jan 14:33, Peter Zijlstra wrote:
On Tue, Jan 15, 2019 at 10:15:07AM +0000, Patrick Bellasi wrote:
quoted
+static __always_inline
+unsigned int uclamp_util_with(struct rq *rq, unsigned int util,
+ struct task_struct *p)
{
unsigned int min_util = READ_ONCE(rq->uclamp[UCLAMP_MIN].value);
unsigned int max_util = READ_ONCE(rq->uclamp[UCLAMP_MAX].value);
+ if (p) {
+ min_util = max(min_util, uclamp_effective_value(p, UCLAMP_MIN));
+ max_util = max(max_util, uclamp_effective_value(p, UCLAMP_MAX));
+ }
+
Like I think you mentioned earlier; this doesn't look right at all.
What we wanna do here is to compute what _will_ be the clamp values of
a CPU if we enqueue *p on it.
The code above starts from the current CPU clamp value and mimics what
uclamp will do in case we move the task there... which is always a max
aggregation.
Should that not be something like:
lo = READ_ONCE(rq->uclamp[UCLAMP_MIN].value);
hi = READ_ONCE(rq->uclamp[UCLAMP_MAX].value);
min_util = clamp(uclamp_effective(p, UCLAMP_MIN), lo, hi);
max_util = clamp(uclamp_effective(p, UCLAMP_MAX), lo, hi);
Here you end up with a restriction of the task clamp (effective)
clamps values considering the CPU clamps... which is different.
Why do you think we should do that?... perhaps I'm missing something.
--
#include <best/regards.h>
Patrick Bellasi
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-23 18:59:50
On Wed, Jan 23, 2019 at 02:14:26PM +0000, Patrick Bellasi wrote:
quoted
quoted
Consider also that the uclamp_task_update_active() added by this patch
not only has lower overhead but it will be use also by cgroups where
we want to force update all the tasks on a cgroup's clamp change.
I haven't gotten that far; but I would prefer not to have two different
'change' paths in __sched_setscheduler().
Yes, I agree that two paths in __sched_setscheduler() could be
confusing. Still we have to consider that here we are adding
"not class specific" attributes.
But that change thing is not class specific; the whole:
rq = task_rq_lock(p, &rf);
queued = task_on_rq_queued(p);
running = task_current(rq, p);
if (queued)
dequeue_task(rq, p, queue_flags);
if (running)
put_prev_task(rq, p);
/* @p is in it's invariant state; frob it's state */
if (queued)
enqueue_task(rq, p, queue_flags);
if (running)
set_curr_task(rq, p);
task_rq_unlock(rq, p, &rf);
pattern is all over the place; it is just because C sucks that that
isn't more explicitly shared (do_set_cpus_allowed(), rt_mutex_setprio(),
set_user_nice(), __sched_setscheduler(), sched_setnuma(),
sched_move_task()).
This is _the_ pattern for changing state and is not class specific at
all.
From: Peter Zijlstra <peterz@infradead.org> Date: 2019-01-23 19:10:16
On Wed, Jan 23, 2019 at 02:19:24PM +0000, Patrick Bellasi wrote:
On 23-Jan 10:22, Peter Zijlstra wrote:
quoted
On Tue, Jan 22, 2019 at 03:41:29PM +0000, Patrick Bellasi wrote:
quoted
On 22-Jan 16:13, Peter Zijlstra wrote:
quoted
On Tue, Jan 22, 2019 at 02:43:29PM +0000, Patrick Bellasi wrote:
quoted
quoted
quoted
Do you think that could be acceptable?
Think so, it's a sysctl poke, 'nobody' ever does that.
Cool, so... I'll keep lazy update for system default.
Ah, I think I misunderstood. I meant to say that since nobody ever pokes
at sysctl's it doesn't matter if its a little more expensive and iterate
everything.
Here I was more worried about the code complexity/overhead... for
something actually not very used/useful.
quoted
Also; if you always keep everything up-to-date, you can avoid doing that
duplicate accounting.
To update everything we will have to walk all the CPUs and update all
the RUNNABLE tasks currently enqueued, which are either RT or CFS.
That's way more expensive both in code and time then what we do for
cgroups, where at least we have a limited scope since the cgroup
already provides a (usually limited) list of tasks to consider.
Do you think it's really worth to have ?
Dunno; the whole double bucket thing seems a bit weird to me; but maybe
it will all look better without the mapping stuff.