On Wed, Aug 10 2016 at 12:09 -0600, Sudeep Holla wrote:
On 10/08/16 17:40, Lina Iyer wrote:
quoted
Hi Sudeep,
On Wed, Aug 10 2016 at 09:15 -0600, Sudeep Holla wrote:
quoted
Hi Lina,
I have few concerns mainly due to the lack of description and not the
binding per say.
[...]
quoted
It is pretty clear that CPUs cannot not define the domain idle states.
Domains define their own idle states. Just as you mention above. CPU is
just a single component in its domain. There may be other devices like
PMUs, Coresights etc that also may have a say in the idle state the
domain may be put in, when the devices are idle. As such, adding domain
idle states to the CPU's idle state property is not appropriate.
No I am not saying we need to add domain idle states to the CPU's idle
state property. I am saying we need to remove cpu-idle-states or ignore
it when PD is present. And get all the idle state information for PD.
I am objecting the split we are creating across CPU and higher level
power domains. And this binding document is incomplete as it skips all
those details. We just need PD handle in CPU and no idle state
information there. Create PD hierarchy and have all idle state
information at one place.
Let me think about this a bit and see what I can come up with.
quoted
Our kernel has runtime PM for devices and then there is CPUidle, both
are diverging without one knowing about the other. We have to start
unifying them inorder to have better holistic power management in the
SoC. To that regard, we have to start imagining CPUs as just another
device, albeit a special device. But for our purposes in determining
domain idle state, it will just be a device attached to the domain.
Absolutely agree on that. No arguments. I am asking to go a step ahead
to include even cpu/core level power domains not just cluster/higher
level domains.
quoted
quoted
We need to have all the idle state information at one place and in this
case PD seems more appropriate instead of splitting them across.
That approach isn't correct. Where will we put the idle states of other
devices that are also part of the domain? We are thinking about a model,
where every device defines its own idle states and we define
relationships between those idle states and their parents' idle states.
Yes I understand. You confused me here. Won't that be one-to-one
relationship ? If not, how is that dealt in the current bindings ?
quoted
Ofcourse, devices don't have idle states today, but that is something we
have been pondering over.
Yes we these binding should be easily extensible, I don't see any issue.
quoted
quoted
We can also keep the code clean and not break compatibility. Whenever
both PD and CPU contains idle-states, PD must take precedence.
Why?
The CPU and PD states are orthogonal. While the PD state is dependent on
the CPU state, the latter is not true. Devices determine their own
states. Based on the individual device states, we then determine the
state of the parent and bubble up on the hierarchy.
I may be missing something. Now with your example in the binding, if
another device shares the cluster PD, can it have different idle states?
If so how does it map ?
In general whatever binding we come up must not just address OS
coordinated mode. Also I was thinking to have better coverage in the
description by having a bit more complex system like:
cluster0
CLUSTER_RET(Retention)
CLUSTER_PG(Power Gate)
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
cluster1
CLUSTER_RET
CLUSTER_PG
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
Platform Co-ordinate supports the following states and we should be
able to determine that from the binding:
CORE_RET
CORE_PG
CORE_RET + CLUSTER_RET
The problem that we have to sove here is knowing that
CORE_RET + CLUSTER_PG (hypothetically) an invalid combination. Kevin and
I debated it in the earlier RFC and we dont have a good way to solve
this generically for all devices.
On Wed, Aug 10 2016 at 12:09 -0600, Sudeep Holla wrote:
quoted
[...]
quoted
cluster0
CLUSTER_RET(Retention)
CLUSTER_PG(Power Gate)
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
cluster1
CLUSTER_RET
CLUSTER_PG
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
Platform Co-ordinate supports the following states and we should be
able to determine that from the binding:
CORE_RET
CORE_PG
CORE_RET + CLUSTER_RET
The problem that we have to sove here is knowing that CORE_RET +
CLUSTER_PG (hypothetically) an invalid combination. Kevin and
I debated it in the earlier RFC and we dont have a good way to solve
this generically for all devices.
Yes, I agree it's complex. But that needs to be solved IMO.
I can think of 2 possible solutions:
1. Index the states(which people have not liked, but as along as we
don't use it in the code as it for any other purpose, it should be
fine) and then have each state mentioning what parent state can be
entered at this child state(i.e. starting index and all states below
it)
2. Something similar to (1) but without index instead phandles.
Again these are just thoughts, others may think of some better
solution(s). Sorry I haven't followed all the previous threads in detail.
--
Regards,
Sudeep
Hi Lina,
Apologies, I sent this reply before and automatically included an "IMPORTANT
NOTICE" footer, please disregard that email, here's the same thing without the
footer.
On Thu, Aug 11, 2016 at 03:10:23PM -0600, Lina Iyer wrote:
On Wed, Aug 10 2016 at 12:09 -0600, Sudeep Holla wrote:
quoted
On 10/08/16 17:40, Lina Iyer wrote:
quoted
Hi Sudeep,
On Wed, Aug 10 2016 at 09:15 -0600, Sudeep Holla wrote:
quoted
Hi Lina,
I have few concerns mainly due to the lack of description and not the
binding per say.
[...]
quoted
It is pretty clear that CPUs cannot not define the domain idle states.
Domains define their own idle states. Just as you mention above. CPU is
just a single component in its domain. There may be other devices like
PMUs, Coresights etc that also may have a say in the idle state the
domain may be put in, when the devices are idle. As such, adding domain
idle states to the CPU's idle state property is not appropriate.
No I am not saying we need to add domain idle states to the CPU's idle
state property. I am saying we need to remove cpu-idle-states or ignore
it when PD is present. And get all the idle state information for PD.
I am objecting the split we are creating across CPU and higher level
power domains. And this binding document is incomplete as it skips all
those details. We just need PD handle in CPU and no idle state
information there. Create PD hierarchy and have all idle state
information at one place.
Let me think about this a bit and see what I can come up with.
quoted
quoted
Our kernel has runtime PM for devices and then there is CPUidle, both
are diverging without one knowing about the other. We have to start
unifying them inorder to have better holistic power management in the
SoC. To that regard, we have to start imagining CPUs as just another
device, albeit a special device. But for our purposes in determining
domain idle state, it will just be a device attached to the domain.
Absolutely agree on that. No arguments. I am asking to go a step ahead
to include even cpu/core level power domains not just cluster/higher
level domains.
quoted
quoted
We need to have all the idle state information at one place and in this
case PD seems more appropriate instead of splitting them across.
That approach isn't correct. Where will we put the idle states of other
devices that are also part of the domain? We are thinking about a model,
where every device defines its own idle states and we define
relationships between those idle states and their parents' idle states.
Yes I understand. You confused me here. Won't that be one-to-one
relationship ? If not, how is that dealt in the current bindings ?
quoted
Ofcourse, devices don't have idle states today, but that is something we
have been pondering over.
Yes we these binding should be easily extensible, I don't see any issue.
quoted
quoted
We can also keep the code clean and not break compatibility. Whenever
both PD and CPU contains idle-states, PD must take precedence.
Why?
The CPU and PD states are orthogonal. While the PD state is dependent on
the CPU state, the latter is not true. Devices determine their own
states. Based on the individual device states, we then determine the
state of the parent and bubble up on the hierarchy.
I may be missing something. Now with your example in the binding, if
another device shares the cluster PD, can it have different idle states?
If so how does it map ?
In general whatever binding we come up must not just address OS
coordinated mode. Also I was thinking to have better coverage in
the description by having a bit more complex system like:
cluster0
CLUSTER_RET(Retention)
CLUSTER_PG(Power Gate)
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
cluster1
CLUSTER_RET
CLUSTER_PG
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
Platform Co-ordinate supports the following states and we should
be able to determine that from the binding:
CORE_RET
CORE_PG
CORE_RET + CLUSTER_RET
The problem that we have to sove here is knowing that CORE_RET +
CLUSTER_PG (hypothetically) an invalid combination. Kevin and
I debated it in the earlier RFC and we dont have a good way to solve
this generically for all devices.
This is interesting. I had been working on the assumption that a parent
power domain cannot enter any idle state until its children were all in
their deepest idle state. I now realise that it's easy to imagine
platforms where this isn't the case.
However, I don't understand how your current bindings solve this issue
and why using domain-power-states for all states (i.e. ignoring
cpu-idle-states and putting CPU idle states in the domain-idle-states of
a per-CPU power domain - I believe this is what Sudeep is suggesting)
makes it any more difficult.
Could you link to this previous discussion you mentioned? I'm having
trouble finding it (R.I.P Gmane).
On Fri, Aug 12 2016 at 06:35 -0600, Brendan Jackman wrote:
quoted
quoted
In general whatever binding we come up must not just address OS
coordinated mode. Also I was thinking to have better coverage in
the description by having a bit more complex system like:
cluster0
CLUSTER_RET(Retention)
CLUSTER_PG(Power Gate)
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
cluster1
CLUSTER_RET
CLUSTER_PG
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
Platform Co-ordinate supports the following states and we should
be able to determine that from the binding:
CORE_RET
CORE_PG
CORE_RET + CLUSTER_RET
The problem that we have to sove here is knowing that CORE_RET +
CLUSTER_PG (hypothetically) an invalid combination. Kevin and
I debated it in the earlier RFC and we dont have a good way to solve
this generically for all devices.
This is interesting. I had been working on the assumption that a parent
power domain cannot enter any idle state until its children were all in
their deepest idle state. I now realise that it's easy to imagine
platforms where this isn't the case.
However, I don't understand how your current bindings solve this issue
and why using domain-power-states for all states (i.e. ignoring
cpu-idle-states and putting CPU idle states in the domain-idle-states of
a per-CPU power domain - I believe this is what Sudeep is suggesting)
makes it any more difficult.
You are right, my current bindings don't solve it. I imagined one would
solve it by writing their own CPU PM Domain governor. In the context of
platform coordinated, we dont have a choice in Linux. May be the
firmware can assert that intelligence in not choosing those states. So,
we may have states added to cpuidle that are invalid and never get
chosen by the firmware. I am not sure, but may be that is acceptable.
Could you link to this previous discussion you mentioned? I'm having
trouble finding it (R.I.P Gmane).
Sigh. So hard to search. Let me see where it is, if it in mail or IRC
communication.
Thanks,
Lina
On Fri, Aug 12 2016 at 04:08 -0600, Sudeep Holla wrote:
On 11/08/16 22:10, Lina Iyer wrote:
quoted
On Wed, Aug 10 2016 at 12:09 -0600, Sudeep Holla wrote:
quoted
[...]
quoted
quoted
cluster0
CLUSTER_RET(Retention)
CLUSTER_PG(Power Gate)
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
cluster1
CLUSTER_RET
CLUSTER_PG
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
Platform Co-ordinate supports the following states and we should be
able to determine that from the binding:
CORE_RET
CORE_PG
CORE_RET + CLUSTER_RET
The problem that we have to sove here is knowing that CORE_RET +
CLUSTER_PG (hypothetically) an invalid combination. Kevin and
I debated it in the earlier RFC and we dont have a good way to solve
this generically for all devices.
Yes, I agree it's complex. But that needs to be solved IMO.
I can think of 2 possible solutions:
1. Index the states(which people have not liked, but as along as we
don't use it in the code as it for any other purpose, it should be
fine) and then have each state mentioning what parent state can be
entered at this child state(i.e. starting index and all states below
it)
This is how QCOM solved it downstream.
2. Something similar to (1) but without index instead phandles.
The problem is when you have non-CPU devices in the device tree and
since they do not have a way to represent states like CPU, we did not
have a clear path to that. Hence we punted that to later. Whatever we
do, we should solve it for a generic PM domain, not just CPU domains.
Thanks,
Lina
Again these are just thoughts, others may think of some better
solution(s). Sorry I haven't followed all the previous threads in detail.
--
Regards,
Sudeep
On Fri, Aug 12 2016 at 04:08 -0600, Sudeep Holla wrote:
quoted
On 11/08/16 22:10, Lina Iyer wrote:
quoted
On Wed, Aug 10 2016 at 12:09 -0600, Sudeep Holla wrote:
quoted
[...]
quoted
quoted
cluster0
CLUSTER_RET(Retention)
CLUSTER_PG(Power Gate)
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
cluster1
CLUSTER_RET
CLUSTER_PG
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
Platform Co-ordinate supports the following states and we should be
able to determine that from the binding:
CORE_RET
CORE_PG
CORE_RET + CLUSTER_RET
The problem that we have to sove here is knowing that CORE_RET +
CLUSTER_PG (hypothetically) an invalid combination. Kevin and
I debated it in the earlier RFC and we dont have a good way to solve
this generically for all devices.
Yes, I agree it's complex. But that needs to be solved IMO.
I can think of 2 possible solutions:
1. Index the states(which people have not liked, but as along as we
don't use it in the code as it for any other purpose, it should be
fine) and then have each state mentioning what parent state can be
entered at this child state(i.e. starting index and all states below
it)
This is how QCOM solved it downstream.
Yes even ACPI has indices to solve this.
quoted
2. Something similar to (1) but without index instead phandles.
The problem is when you have non-CPU devices in the device tree and
since they do not have a way to represent states like CPU, we did not
have a clear path to that. Hence we punted that to later. Whatever we
do, we should solve it for a generic PM domain, not just CPU domains.
Yes bindings defined here should be applicable for devices to, but only
CPU's will have this hierarchy while the devices need not bother about
hierarchy. However the parent power domain can ever the state which is
least common denominator of all it's children power domain. That's my
understanding. No ?
--
Regards,
Sudeep
On Mon, Aug 15 2016 at 10:14 -0600, Sudeep Holla wrote:
On 15/08/16 17:08, Lina Iyer wrote:
quoted
On Fri, Aug 12 2016 at 04:08 -0600, Sudeep Holla wrote:
quoted
On 11/08/16 22:10, Lina Iyer wrote:
quoted
On Wed, Aug 10 2016 at 12:09 -0600, Sudeep Holla wrote:
quoted
[...]
quoted
quoted
cluster0
CLUSTER_RET(Retention)
CLUSTER_PG(Power Gate)
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
cluster1
CLUSTER_RET
CLUSTER_PG
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
Platform Co-ordinate supports the following states and we should be
able to determine that from the binding:
CORE_RET
CORE_PG
CORE_RET + CLUSTER_RET
The problem that we have to sove here is knowing that CORE_RET +
CLUSTER_PG (hypothetically) an invalid combination. Kevin and
I debated it in the earlier RFC and we dont have a good way to solve
this generically for all devices.
Yes, I agree it's complex. But that needs to be solved IMO.
I can think of 2 possible solutions:
1. Index the states(which people have not liked, but as along as we
don't use it in the code as it for any other purpose, it should be
fine) and then have each state mentioning what parent state can be
entered at this child state(i.e. starting index and all states below
it)
This is how QCOM solved it downstream.
Yes even ACPI has indices to solve this.
quoted
quoted
2. Something similar to (1) but without index instead phandles.
The problem is when you have non-CPU devices in the device tree and
since they do not have a way to represent states like CPU, we did not
have a clear path to that. Hence we punted that to later. Whatever we
do, we should solve it for a generic PM domain, not just CPU domains.
Yes bindings defined here should be applicable for devices to, but only
CPU's will have this hierarchy while the devices need not bother about
hierarchy. However the parent power domain can ever the state which is
least common denominator of all it's children power domain. That's my
understanding. No ?
That is correct. But say if all the CPUs choose CORE_RET + CLUSTER_PG,
which is invalid and the firmware has to ignore it and does CORE_RET +
CLUSTER_RET instead, then Linux may have an inconsistent view of the
state selection.
Thanks,
Lina
Hi Lina,
Agh, sorry, sent with the "IMPORTANT NOTICE" again, still getting used to
mailing lists.. here's the message again without it.
On Mon, Aug 15, 2016 at 04:40:14PM -0600, Lina Iyer wrote:
On Mon, Aug 15 2016 at 10:14 -0600, Sudeep Holla wrote:
quoted
On 15/08/16 17:08, Lina Iyer wrote:
quoted
On Fri, Aug 12 2016 at 04:08 -0600, Sudeep Holla wrote:
quoted
On 11/08/16 22:10, Lina Iyer wrote:
quoted
On Wed, Aug 10 2016 at 12:09 -0600, Sudeep Holla wrote:
quoted
[...]
quoted
quoted
cluster0
CLUSTER_RET(Retention)
CLUSTER_PG(Power Gate)
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
cluster1
CLUSTER_RET
CLUSTER_PG
core0
CORE_RET
CORE_PG
core1
CORE_RET
CORE_PG
Platform Co-ordinate supports the following states and we should be
able to determine that from the binding:
CORE_RET
CORE_PG
CORE_RET + CLUSTER_RET
The problem that we have to sove here is knowing that CORE_RET +
CLUSTER_PG (hypothetically) an invalid combination. Kevin and
I debated it in the earlier RFC and we dont have a good way to solve
this generically for all devices.
Yes, I agree it's complex. But that needs to be solved IMO.
I can think of 2 possible solutions:
1. Index the states(which people have not liked, but as along as we
don't use it in the code as it for any other purpose, it should be
fine) and then have each state mentioning what parent state can be
entered at this child state(i.e. starting index and all states below
it)
This is how QCOM solved it downstream.
Yes even ACPI has indices to solve this.
quoted
quoted
2. Something similar to (1) but without index instead phandles.
The problem is when you have non-CPU devices in the device tree and
since they do not have a way to represent states like CPU, we did not
have a clear path to that. Hence we punted that to later. Whatever we
do, we should solve it for a generic PM domain, not just CPU domains.
Yes bindings defined here should be applicable for devices to, but only
CPU's will have this hierarchy while the devices need not bother about
hierarchy. However the parent power domain can ever the state which is
least common denominator of all it's children power domain. That's my
understanding. No?
Are you saying that the parent can enter the shallowest idle state that all its
children are in (I.e if all its children are in "retention" then it can enter
"retention")? I don't know what the reality is on existing platforms but it
doesn't sound like 100% safe assumption to make. Also I don't think you can
necessarily correlate idle states at different domain levels - i.e. here we've
matched up the idea of "retention" at core level with that of "retention" at
cluster level. I may have misunderstood you there..
That is correct. But say if all the CPUs choose CORE_RET + CLUSTER_PG,
which is invalid and the firmware has to ignore it and does CORE_RET +
CLUSTER_RET instead, then Linux may have an inconsistent view of the
state selection.
Perhaps a better starting point would be to go with the assumption that a parent
PD can only enter any idle state once its children are in their deepest idle
states.
So in the example above we'd end up with
CORE_RET
CORE_PG
CORE_PG + CLUSTER_RET
CORE_PG + CLUSTER_PG
(Missing out on CORE_RET + CLUSTER_RET, even though that's a valid combination
from the hardware's perspective)
Then a later addition to the bindings as discussed above could enable the
possibility of those combinations to be expressed.
Hi Lina,
On Mon, Aug 15, 2016 at 04:40:14PM -0600, Lina Iyer wrote:
quoted
On Mon, Aug 15 2016 at 10:14 -0600, Sudeep Holla wrote:
[,,,]
quoted
quoted
Yes even ACPI has indices to solve this.
quoted
quoted
2. Something similar to (1) but without index instead phandles.
The problem is when you have non-CPU devices in the device tree and
since they do not have a way to represent states like CPU, we did not
have a clear path to that. Hence we punted that to later. Whatever we
do, we should solve it for a generic PM domain, not just CPU domains.
Yes bindings defined here should be applicable for devices to, but only
CPU's will have this hierarchy while the devices need not bother about
hierarchy. However the parent power domain can ever the state which is
least common denominator of all it's children power domain. That's my
understanding. No?
Are you saying that the parent can enter the shallowest idle state that all its
children are in (I.e if all its children are in "retention" then it can enter
"retention")? I don't know what the reality is on existing platforms but it
doesn't sound like 100% safe assumption to make.
I was referring to non-CPU/device power states above. For CPU we do need
a mechanism in place to indicate the dependency.
Also I don't think you can
necessarily correlate idle states at different domain levels - i.e. here we've
matched up the idea of "retention" at core level with that of "retention" at
cluster level. I may have misunderstood you there..
Correct for CPUs. For normal devices and their power domains, it could
straight forward. E.g if many devices are at-least at state D1(few may
be at state D2 or above), the parent can enter D1.(D0-runnning and D1-D3
are low power states in the above example)
quoted
That is correct. But say if all the CPUs choose CORE_RET + CLUSTER_PG,
which is invalid and the firmware has to ignore it and does CORE_RET +
CLUSTER_RET instead, then Linux may have an inconsistent view of the
state selection.
1. First, CORE_RET + CLUSTER_PG should not be registered as valid idle
state.
2. We do have inconsistent view already for platform co-ordinated idle
In-fact it could happen even with OSC mode I believe. Platform can
always demote the state, so OS can never get the exact view unless it
queries the firmware for that explicitly(e.g. PSCI_STATS)
Perhaps a better starting point would be to go with the assumption that a parent
PD can only enter any idle state once its children are in their deepest idle
states.
So in the example above we'd end up with
CORE_RET
CORE_PG
CORE_PG + CLUSTER_RET
CORE_PG + CLUSTER_PG
Yes this assumption seems good enough to me. At-least no invalid
combination is ensured.
(Missing out on CORE_RET + CLUSTER_RET, even though that's a valid combination
from the hardware's perspective)
Yes, if it's a real issue then we need proper bindings to deal with
that. Otherwise we can manage without the extra information.
Then a later addition to the bindings as discussed above could enable the
possibility of those combinations to be expressed.
Seems feasible solution to me, but better to make this explicit in the
binding and check with few others. It looks fair enough assumption IMO.
--
--
Regards,
Sudeep