From: Eran Ben Elisha <hidden> Date: 2018-09-13 13:27:20
The health spec is targeted for Real Time Alerting, in order to know when
something bad had happened to a PCI device
- Provide alert debug information
- Self healing
- If problem needs vendor support, provide a way to gather all needed debugging
information.
The health contains sensors which sense for malfunction. Once sensor triggered,
actions such as logs and correction can be taken.
Sensors are sensing the health state and can trigger correction action.
The sensors are divided into the following groups
- Hardware sensor - a sensor which is triggered by the device due to
malfunction.
- Software sensor - a sensor which is triggered by the software due to
malfunction.
Both group of sensors can be triggered due to error event or due to a periodic check.
Actions are the way to handle sensor events. Action can be in one of the
following groups:
- Dump - SW trace, SW dump, HW trace, HW dump
- Reset - Surgical correction (e.g. modify Q, flush Q, reset of device, etc)
Actions can be performed by SW or HW.
User is allowed to enable or disable sensors and sensor2action mapping.
This RFC man page patch describes the suggested API of devlink-health in order
to control sensors and actions.
Eran Ben Elisha (1):
man: Add devlink health man page
man/man8/devlink-health.8 | 171 ++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 171 insertions(+)
create mode 100644 man/man8/devlink-health.8
--
1.8.3.1
From: Eran Ben Elisha <hidden> Date: 2018-09-13 13:27:30
Add devlink-health man page. Devlink-health tool will control device
health attributes, sensors, actions and logging.
Signed-off-by: Eran Ben Elisha <redacted>
-------------------------------------------------------
Copy paste man output to here for easier review process of the RFC.
DEVLINK-HEALTH(8) Linux DEVLINK-HEALTH(8)
NAME
devlink-health - devlink health configuration
SYNOPSIS
devlink [ OPTIONS ] health { COMMAND | help }
OPTIONS := { -V[ersion] | -n[no-nice-names] }
devlink health show [ DEV ] [ sensor NAME ]
devlink health sensor set DEV name NAME [ action NAME { active | inactive } ]"
devlink health action set DEV name NAME period PERIOD count COUNT fail { ignore | down }
devlink health action reinit DEV name NAME
devlink health help
DESCRIPTION
devlink-health tool allows user to configure the way driver treats unexpected status. The tool allows configuration of the sensors that can trigger health activity. Set for each sensor the follow up operations, such as,
reset and dump of info. In addition, set the health activity termination action.
devlink health show - Display devlink health sensors and actions attributes
DEV - Specifies the devlink device to show. If this argument is omitted, all devices are listed.
Format is:
BUS_NAME/BUS_ADDRESS
sensor NAME - Specifies the devlink sensor to show.
devlink health sensor set - sets devlink health sensor attributes
DEV Specifies the devlink device to show.
name NAME
Name of the sensor to set.
action NAME { active | inactive }
Specify which actions to activate and which to deactivate once a sensor was triggered. actions can be dump, reset, etc.
devlink health action set - sets devlink action attributes
DEV Specifies the devlink device to set.
name NAME
Specifies the devlink action to set.
period PERIOD
The period on which we limit the amount of performed actions, measured in seconds.
count COUNT
The maximum amount of actions performed in a limit time frame.
fail { ignore | down }
Specify the behavior once count limit was reached.
ignore - Ignore errors without execution of any action.
down - Driver will remain in nonoperational state.
devlink health action reinit - reset devlink action attributes (period, count, fail, etc)
DEV Specifies the devlink device to set.
name NAME
Specifies the devlink action to set.
EXAMPLES
devlink health show
Shows the health state of all devlink devices on the system.
devlink health show pci/0000:01:00.0
Shows the health state of specified devlink device.
devlink health sensor set pci/0000:01:00.0 name TX_COMP_ERROR action reset off action dump on
Sets TX_COMP_ERROR sensor parameters for a specific device.
devlink health action set pci/0000:01:00.0 name reset period 3600 count 5 fail ignore
Sets health attributes for reset action.
SEE ALSO
devlink(8), devlink-port(8), devlink-sb(8), devlink-monitor(8), devlink-dev(8),
AUTHOR
Eran ben Elisha [off-list ref]
iproute2 15 Aug 2018 DEVLINK-HEALTH(8)
---
man/man8/devlink-health.8 | 171 ++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 171 insertions(+)
create mode 100644 man/man8/devlink-health.8
@@ -0,0 +1,171 @@+.THDEVLINK\-HEALTH8"15 Aug 2018""iproute2""Linux"+.SHNAME+devlink-health \- devlink health configuration+.SHSYNOPSIS+.sp+.adl+.in+8+.ti-8+.Bdevlink+.RI"[ "OPTIONS" ]"+.BRhealth+.RI" { "COMMAND" | "+.BRhelp" }"+.sp++.ti-8+.IROPTIONS" := { "+\fB\-V\fR[\fIersion\fR] |+\fB\-n\fR[\fIno-nice-names\fR] }++.ti-8+.Bdevlinkhealthshow+.RI"[ "DEV" ]"+.RI"[ "+.Bsensor+.IRNAME+.RI"]"++.ti-8+.Bdevlinkhealthsensorset+.IRDEV+.Bname+.IRNAME+.RI"[ "+.BRaction+.IRNAME+.R"{"active"|"inactive"}"]"++.ti-8+.Bdevlinkhealthactionset+.IRDEV+.Bname+.IRNAME+.BRperiod+.IRPERIOD+.BRcount+.IRCOUNT+.BRfail" { "+.IRignore+.BR"| "+.IRdown+.R"} "++.ti-8+.Bdevlinkhealthactionreinit+.IRDEV+.Bname+.IRNAME++.ti-8+.Bdevlinkhealthhelp++.SH"DESCRIPTION"+.Bdevlink-health+tool allows user to configure the way driver treats unexpected status. The tool allows configuration of the sensors that can trigger health activity. Set for each sensor the follow up operations, such as, reset and dump of info. In addition, set the health activity termination action.++.SSdevlinkhealthshow-Displaydevlinkhealthsensorsandactionsattributes+.PP+.B"DEV"+- Specifies the devlink device to show.+If this argument is omitted, all devices are listed.++.in+4+Format is:+.in+2+BUS_NAME/BUS_ADDRESS++.PP+.BRsensor+.IR"NAME"+- Specifies the devlink sensor to show.++.SSdevlinkhealthsensorset-setsdevlinkhealthsensorattributes++.TP+.B"DEV"+Specifies the devlink device to show.++.TP+.BIname" NAME"+Name of the sensor to set.++.TP+.BRaction+.IRNAME+.R"{"active"|"inactive"} "+.in+4+Specify which actions to activate and which to deactivate once a sensor was triggered. actions can be dump, reset, etc.++.SSdevlinkhealthactionset-setsdevlinkactionattributes++.TP+.B"DEV"+Specifies the devlink device to set.++.TP+.BIname" NAME"+Specifies the devlink action to set.++.TP+.BIperiod" PERIOD"+The period on which we limit the amount of performed actions, measured in seconds.++.TP+.BIcount" COUNT"+The maximum amount of actions performed in a limit time frame.++.TP+.BRfail+.R"{"ignore"|"down"}"+.in+4+Specify the behavior once count limit was reached.++.Iignore+- Ignore errors without execution of any action.++.Idown+- Driver will remain in nonoperational state.++.SSdevlinkhealthactionreinit-resetdevlinkactionattributes(period,count,fail,etc)++.TP+.B"DEV"+Specifies the devlink device to set.++.TP+.BIname" NAME"+Specifies the devlink action to set.++.SH"EXAMPLES"+.PP+devlink health show+.RS4+Shows the health state of all devlink devices on the system.+.RE+.PP+devlink health show pci/0000:01:00.0+.RS4+Shows the health state of specified devlink device.+.RE+.PP+devlink health sensor set pci/0000:01:00.0 name TX_COMP_ERROR action reset off action dump on+.RS4+Sets TX_COMP_ERROR sensor parameters for a specific device.+.RE+.PP+devlink health action set pci/0000:01:00.0 name reset period 3600 count 5 fail ignore+.RS4+Sets health attributes for reset action.+.RE++.SHSEEALSO+.BRdevlink(8),+.BRdevlink-port(8),+.BRdevlink-sb(8),+.BRdevlink-monitor(8),+.BRdevlink-dev(8),+.br++.SHAUTHOR+Eran ben Elisha <eranbe@mellanox.com>
From: Tobin C. Harding <hidden> Date: 2018-09-13 15:36:29
On Thu, Sep 13, 2018 at 11:18:16AM +0300, Eran Ben Elisha wrote:
Add devlink-health man page. Devlink-health tool will control device
health attributes, sensors, actions and logging.
Signed-off-by: Eran Ben Elisha <redacted>
-------------------------------------------------------
Copy paste man output to here for easier review process of the RFC.
DEVLINK-HEALTH(8) Linux DEVLINK-HEALTH(8)
NAME
devlink-health - devlink health configuration
SYNOPSIS
devlink [ OPTIONS ] health { COMMAND | help }
OPTIONS := { -V[ersion] | -n[no-nice-names] }
devlink health show [ DEV ] [ sensor NAME ]
devlink health sensor set DEV name NAME [ action NAME { active | inactive } ]"
devlink health action set DEV name NAME period PERIOD count COUNT fail { ignore | down }
devlink health action reinit DEV name NAME
devlink health help
DESCRIPTION
devlink-health tool allows user to configure the way driver treats unexpected status. The tool allows configuration of the sensors that can trigger health activity. Set for each sensor the follow up operations, such as,
reset and dump of info. In addition, set the health activity termination action.
devlink health show - Display devlink health sensors and actions attributes
DEV - Specifies the devlink device to show. If this argument is omitted, all devices are listed.
Format is:
BUS_NAME/BUS_ADDRESS
sensor NAME - Specifies the devlink sensor to show.
Perhaps the commands should include the optional arguments so when
reading the description one doesn't have to scroll to the top of the
page all the time
e.g
devlink health show [ DEV ] [ sensor NAME ] - Display devlink health sensors and actions attributes
devlink health sensor set - sets devlink health sensor attributes
DEV Specifies the devlink device to show.
set
name NAME
Name of the sensor to set.
action NAME { active | inactive }
Specify which actions to activate and which to deactivate once a sensor was triggered. actions can be dump, reset, etc.
devlink health action set - sets devlink action attributes
DEV Specifies the devlink device to set.
name NAME
Specifies the devlink action to set.
This is a little unclear to me?
period PERIOD
The period on which we limit the amount of performed actions, measured in seconds.
count COUNT
The maximum amount of actions performed in a limit time frame.
Perhaps
The maximum number of actions performed in a limited time frame.
fail { ignore | down }
Specify the behavior once count limit was reached.
ignore - Ignore errors without execution of any action.
down - Driver will remain in nonoperational state.
devlink health action reinit - reset devlink action attributes (period, count, fail, etc)
DEV Specifies the devlink device to set.
name NAME
Specifies the devlink action to set.
Perhaps s/set/reinitialise/g for the above two descriptions.
Hope this helps,
Tobin.
From: Eran Ben Elisha <hidden> Date: 2018-09-13 17:08:18
On 9/13/2018 1:27 PM, Tobin C. Harding wrote:
On Thu, Sep 13, 2018 at 11:18:16AM +0300, Eran Ben Elisha wrote:
quoted
Add devlink-health man page. Devlink-health tool will control device
health attributes, sensors, actions and logging.
Signed-off-by: Eran Ben Elisha <redacted>
-------------------------------------------------------
Copy paste man output to here for easier review process of the RFC.
DEVLINK-HEALTH(8) Linux DEVLINK-HEALTH(8)
NAME
devlink-health - devlink health configuration
SYNOPSIS
devlink [ OPTIONS ] health { COMMAND | help }
OPTIONS := { -V[ersion] | -n[no-nice-names] }
devlink health show [ DEV ] [ sensor NAME ]
devlink health sensor set DEV name NAME [ action NAME { active | inactive } ]"
devlink health action set DEV name NAME period PERIOD count COUNT fail { ignore | down }
devlink health action reinit DEV name NAME
devlink health help
DESCRIPTION
devlink-health tool allows user to configure the way driver treats unexpected status. The tool allows configuration of the sensors that can trigger health activity. Set for each sensor the follow up operations, such as,
reset and dump of info. In addition, set the health activity termination action.
devlink health show - Display devlink health sensors and actions attributes
DEV - Specifies the devlink device to show. If this argument is omitted, all devices are listed.
Format is:
BUS_NAME/BUS_ADDRESS
sensor NAME - Specifies the devlink sensor to show.
Perhaps the commands should include the optional arguments so when
reading the description one doesn't have to scroll to the top of the
page all the time
e.g
devlink health show [ DEV ] [ sensor NAME ] - Display devlink health sensors and actions attributes
I followed the scheme presented in all other devlink man pages.
see devlink-region, devlink-port, etc.
From my perspective, I am fine with adding it to devlink-health, need
ack from the devlink maintainer to see if he likes it...
quoted
devlink health sensor set - sets devlink health sensor attributes
DEV Specifies the devlink device to show.
set
quoted
name NAME
Name of the sensor to set.
action NAME { active | inactive }
Specify which actions to activate and which to deactivate once a sensor was triggered. actions can be dump, reset, etc.
devlink health action set - sets devlink action attributes
DEV Specifies the devlink device to set.
name NAME
Specifies the devlink action to set.
This is a little unclear to me?
what is not clear? the term 'action' or the naming? can you elaborate?
quoted
period PERIOD
The period on which we limit the amount of performed actions, measured in seconds.
count COUNT
The maximum amount of actions performed in a limit time frame.
Perhaps
The maximum number of actions performed in a limited time frame.
ack
quoted
fail { ignore | down }
Specify the behavior once count limit was reached.
ignore - Ignore errors without execution of any action.
down - Driver will remain in nonoperational state.
devlink health action reinit - reset devlink action attributes (period, count, fail, etc)
DEV Specifies the devlink device to set.
name NAME
Specifies the devlink action to set.
Perhaps s/set/reinitialise/g for the above two descriptions.
From: Andrew Lunn <andrew@lunn.ch> Date: 2018-09-13 17:17:28
devlink health sensor set pci/0000:01:00.0 name TX_COMP_ERROR action reset off action dump on
Sets TX_COMP_ERROR sensor parameters for a specific device.
I hope the real sensors have more understandable names. If i remember
correctly, the same sort of comment was given for resource
management. It was pretty unclear what the resource names actually
mean. Is an average user going to have any idea how to actually use
these sensors and actions?
Can you give more examples of sensors. We should understand if there
are any overlaps with hwmon.
Andrew
From: Eran Ben Elisha <hidden> Date: 2018-09-13 17:59:07
On 9/13/2018 3:08 PM, Andrew Lunn wrote:
quoted
devlink health sensor set pci/0000:01:00.0 name TX_COMP_ERROR action reset off action dump on
Sets TX_COMP_ERROR sensor parameters for a specific device.
I hope the real sensors have more understandable names. If i remember
correctly, the same sort of comment was given for resource
management. It was pretty unclear what the resource names actually
mean. Is an average user going to have any idea how to actually use
these sensors and actions?
well, hopefully. the whole point is to have it fully controlled by the
user. However, names for the command should be short. I guess we shall
have it documented (challenge is to fit to multi vendors).
Can you give more examples of sensors. We should understand if there
are any overlaps with hwmon.
I restate here that we shall have SW sensors as well, and not only HW
sensors.
This is what I had in mind:
1. command interface error
2. command interface timeout
3. stuck TX queue (like tx_timeout)
4. stuck TX completion queue (driver did not process packets in a
reasonable time period)
5. stuck RX queue
6. RX completion error
7. TX completion error
8. HW / FW catastrophic error report
9. completion queue overrun
Eran
From: Andrew Lunn <andrew@lunn.ch> Date: 2018-09-13 18:34:23
On Thu, Sep 13, 2018 at 03:49:37PM +0300, Eran Ben Elisha wrote:
On 9/13/2018 3:08 PM, Andrew Lunn wrote:
quoted
quoted
devlink health sensor set pci/0000:01:00.0 name TX_COMP_ERROR action reset off action dump on
Sets TX_COMP_ERROR sensor parameters for a specific device.
I hope the real sensors have more understandable names. If i remember
correctly, the same sort of comment was given for resource
management. It was pretty unclear what the resource names actually
mean. Is an average user going to have any idea how to actually use
these sensors and actions?
well, hopefully. the whole point is to have it fully controlled by the user.
However, names for the command should be short. I guess we shall have it
documented (challenge is to fit to multi vendors).
quoted
Can you give more examples of sensors. We should understand if there
are any overlaps with hwmon.
I restate here that we shall have SW sensors as well, and not only HW
sensors.
This is what I had in mind:
1. command interface error
2. command interface timeout
3. stuck TX queue (like tx_timeout)
4. stuck TX completion queue (driver did not process packets in a reasonable
time period)
5. stuck RX queue
6. RX completion error
7. TX completion error
8. HW / FW catastrophic error report
9. completion queue overrun
Hi Eran
I'm having trouble differentiating between these SW sensors and bugs
which need fixing. What causes a command interface error? Sending it a
command it does not understand? A wrongly formatted command? A command
the version of the firmware does not support? These all sound just
like plain old bugs which need fixing, not something which needs a
framework to detect them and try to recover from them by resetting
something.
I would of expected that all the issues are about physical
properties. Something similar to SMART for hard disks. The power
supplies are starting to droop, suggesting it might die soon. The
tacho on the fan suggests the FAN is not rotating as fast as it
should, so the motor is going to die soon. An SFP is giving i2c
errors, suggesting it is not seated correctly. The card as a whole is
overheating, despite the fan working, suggesting the ambient
temperature is just too high.
Andrew
From: Eran Ben Elisha <hidden> Date: 2018-09-13 19:40:27
On 9/13/2018 4:24 PM, Andrew Lunn wrote:
On Thu, Sep 13, 2018 at 03:49:37PM +0300, Eran Ben Elisha wrote:
quoted
On 9/13/2018 3:08 PM, Andrew Lunn wrote:
quoted
quoted
devlink health sensor set pci/0000:01:00.0 name TX_COMP_ERROR action reset off action dump on
Sets TX_COMP_ERROR sensor parameters for a specific device.
I hope the real sensors have more understandable names. If i remember
correctly, the same sort of comment was given for resource
management. It was pretty unclear what the resource names actually
mean. Is an average user going to have any idea how to actually use
these sensors and actions?
well, hopefully. the whole point is to have it fully controlled by the user.
However, names for the command should be short. I guess we shall have it
documented (challenge is to fit to multi vendors).
quoted
Can you give more examples of sensors. We should understand if there
are any overlaps with hwmon.
I restate here that we shall have SW sensors as well, and not only HW
sensors.
This is what I had in mind:
1. command interface error
2. command interface timeout
3. stuck TX queue (like tx_timeout)
4. stuck TX completion queue (driver did not process packets in a reasonable
time period)
5. stuck RX queue
6. RX completion error
7. TX completion error
8. HW / FW catastrophic error report
9. completion queue overrun
Hi Eran
I'm having trouble differentiating between these SW sensors and bugs
which need fixing. What causes a command interface error? Sending it a
command it does not understand? A wrongly formatted command? A command
the version of the firmware does not support? These all sound just
like plain old bugs which need fixing, not something which needs a
framework to detect them and try to recover from them by resetting
something.
Such issues do exist in production environment, and need to be handled
even if root cause is a bug which will be fixed in latest release. My
feature should help developers / administrator to control and recover
their live systems, by auto correction and logging support.
Goal is:
- Provide alert debug information
- Self healing
- If problem needs vendor support, provide a way to gather all needed
debugging information.
I would of expected that all the issues are about physical
properties. Something similar to SMART for hard disks. The power
supplies are starting to droop, suggesting it might die soon. The
tacho on the fan suggests the FAN is not rotating as fast as it
should, so the motor is going to die soon. An SFP is giving i2c
errors, suggesting it is not seated correctly. The card as a whole is
overheating, despite the fan working, suggesting the ambient
temperature is just too high.
AFAIU, the kind of sensors you suggest here requires manual fix /
physically approaching to the setup, replace HW, install new Fan, etc.
Monitor such events is easy, driver can just log events from HW to the
dmesg and end its handle there.
None of these is a real networking issue I would like to handle with
devlink-health.
Eran
From: Andrew Lunn <andrew@lunn.ch> Date: 2018-09-13 20:22:50
quoted
quoted
quoted
quoted
devlink health sensor set pci/0000:01:00.0 name TX_COMP_ERROR action reset off action dump on
Sets TX_COMP_ERROR sensor parameters for a specific device.
quoted
quoted
This is what I had in mind:
1. command interface error
2. command interface timeout
3. stuck TX queue (like tx_timeout)
4. stuck TX completion queue (driver did not process packets in a reasonable
time period)
5. stuck RX queue
6. RX completion error
7. TX completion error
8. HW / FW catastrophic error report
9. completion queue overrun
Such issues do exist in production environment, and need to be handled even
if root cause is a bug which will be fixed in latest release. My feature
should help developers / administrator to control and recover their live
systems, by auto correction and logging support.
Goal is:
- Provide alert debug information
- Self healing
- If problem needs vendor support, provide a way to gather all needed
debugging information.
So maybe you have the wrong name for this. Health is nice in terms of
Marketing, but we are actually talking about bug recovery.
devlink bug sensor set pci/0000:01:00.0 name command_interface_error action reset off action dump on
devlink bug sensor set pci/0000:01:00.0 name command_interface_timeout action reset off action dump on
devlink bug sensor set pci/0000:01:00.0 name transmit_completion_error action reset off action dump on
devlink bug sensor set pci/0000:01:00.0 name completion_queue_overrun action reset off action dump on
seems a lot more understandable than:
devlink health set pci/0000:01:00.0 name TX_COMP_ERROR action reset off action dump on
Andrew
From: Jakub Kicinski <hidden> Date: 2018-09-13 22:46:49
On Thu, 13 Sep 2018 11:18:15 +0300, Eran Ben Elisha wrote:
The health spec is targeted for Real Time Alerting, in order to know when
something bad had happened to a PCI device
By spec you mean some standards body spec you implement or this
proposal is a spec?
- Provide alert debug information
- Self healing
- If problem needs vendor support, provide a way to gather all needed debugging
information.
The health contains sensors which sense for malfunction. Once sensor triggered,
actions such as logs and correction can be taken.
Sensors are sensing the health state and can trigger correction action.
The sensors are divided into the following groups
- Hardware sensor - a sensor which is triggered by the device due to
malfunction.
- Software sensor - a sensor which is triggered by the software due to
malfunction.
Both group of sensors can be triggered due to error event or due to a periodic check.
Actions are the way to handle sensor events. Action can be in one of the
following groups:
- Dump - SW trace, SW dump, HW trace, HW dump
- Reset - Surgical correction (e.g. modify Q, flush Q, reset of device, etc)
Actions can be performed by SW or HW.
User is allowed to enable or disable sensors and sensor2action mapping.
This RFC man page patch describes the suggested API of devlink-health in order
to control sensors and actions.
I like the idea of configuring response to events like this, although
I'm not sure the name sensor is appropriate here - perhaps exception or
error would be better? Are there going to be values reported?
I'm not so sure about HW sensors in relation to existing HWMON
infrastructure... I assume you're targeting things like say some HW
engine/block reporting it encountered an error? Sounds good, too.
Are the actions all envisioned to be performed by the driver?
Firmware? Hardware? I guess that distinction can be added later.
For FW/HW actions we would go back to the problem of persistence of
the setting since it was only implemented for params :S
Is the dump option going to tie back into region snapshots?
From: Tobin C. Harding <hidden> Date: 2018-09-14 03:18:05
On Thu, Sep 13, 2018 at 02:58:52PM +0300, Eran Ben Elisha wrote:
On 9/13/2018 1:27 PM, Tobin C. Harding wrote:
quoted
On Thu, Sep 13, 2018 at 11:18:16AM +0300, Eran Ben Elisha wrote:
quoted
Add devlink-health man page. Devlink-health tool will control device
health attributes, sensors, actions and logging.
Signed-off-by: Eran Ben Elisha <redacted>
-------------------------------------------------------
Copy paste man output to here for easier review process of the RFC.
DEVLINK-HEALTH(8) Linux DEVLINK-HEALTH(8)
NAME
devlink-health - devlink health configuration
SYNOPSIS
devlink [ OPTIONS ] health { COMMAND | help }
OPTIONS := { -V[ersion] | -n[no-nice-names] }
devlink health show [ DEV ] [ sensor NAME ]
devlink health sensor set DEV name NAME [ action NAME { active | inactive } ]"
devlink health action set DEV name NAME period PERIOD count COUNT fail { ignore | down }
devlink health action reinit DEV name NAME
devlink health help
DESCRIPTION
devlink-health tool allows user to configure the way driver treats unexpected status. The tool allows configuration of the sensors that can trigger health activity. Set for each sensor the follow up operations, such as,
reset and dump of info. In addition, set the health activity termination action.
devlink health show - Display devlink health sensors and actions attributes
DEV - Specifies the devlink device to show. If this argument is omitted, all devices are listed.
Format is:
BUS_NAME/BUS_ADDRESS
sensor NAME - Specifies the devlink sensor to show.
Perhaps the commands should include the optional arguments so when
reading the description one doesn't have to scroll to the top of the
page all the time
e.g
devlink health show [ DEV ] [ sensor NAME ] - Display devlink health sensors and actions attributes
I followed the scheme presented in all other devlink man pages.
see devlink-region, devlink-port, etc.
Oh ok, my mistake. I'd stick with what you have then. Thanks for
pointing this out.
From my perspective, I am fine with adding it to devlink-health, need ack
from the devlink maintainer to see if he likes it...
quoted
quoted
devlink health sensor set - sets devlink health sensor attributes
DEV Specifies the devlink device to show.
set
quoted
name NAME
Name of the sensor to set.
action NAME { active | inactive }
Specify which actions to activate and which to deactivate once a sensor was triggered. actions can be dump, reset, etc.
devlink health action set - sets devlink action attributes
DEV Specifies the devlink device to set.
name NAME
Specifies the devlink action to set.
This is a little unclear to me?
what is not clear? the term 'action' or the naming? can you elaborate?
It wasn't immediately clear what 'name' referred to. But following on
from discussion above this may be because I have not read any of the
other devlink man pages.
thanks,
Tobin.
From: Eran Ben Elisha <hidden> Date: 2018-09-16 14:36:48
On 9/13/2018 6:12 PM, Andrew Lunn wrote:
quoted
quoted
quoted
quoted
quoted
devlink health sensor set pci/0000:01:00.0 name TX_COMP_ERROR action reset off action dump on
Sets TX_COMP_ERROR sensor parameters for a specific device.
quoted
quoted
quoted
This is what I had in mind:
1. command interface error
2. command interface timeout
3. stuck TX queue (like tx_timeout)
4. stuck TX completion queue (driver did not process packets in a reasonable
time period)
5. stuck RX queue
6. RX completion error
7. TX completion error
8. HW / FW catastrophic error report
9. completion queue overrun
quoted
Such issues do exist in production environment, and need to be handled even
if root cause is a bug which will be fixed in latest release. My feature
should help developers / administrator to control and recover their live
systems, by auto correction and logging support.
Goal is:
- Provide alert debug information
- Self healing
- If problem needs vendor support, provide a way to gather all needed
debugging information.
So maybe you have the wrong name for this. Health is nice in terms of
Marketing, but we are actually talking about bug recovery.
The way I see it, this feature is responsible for the health of the
system from the pci/xxxx perspective.
I though about devlink-recover for example, but I really wouldn't like
to limit the feature to be called after one of its actions. The same for
devlink-bug, which highlights only part of the range of capabilities
(sensor).
My work is currently focused on error reporting and recovery, but I
wouldn't like to see the API limited for "bugs" only.
Eran
devlink bug sensor set pci/0000:01:00.0 name command_interface_error action reset off action dump on
devlink bug sensor set pci/0000:01:00.0 name command_interface_timeout action reset off action dump on
devlink bug sensor set pci/0000:01:00.0 name transmit_completion_error action reset off action dump on
devlink bug sensor set pci/0000:01:00.0 name completion_queue_overrun action reset off action dump on
seems a lot more understandable than:
devlink health set pci/0000:01:00.0 name TX_COMP_ERROR action reset off action dump on
Andrew
From: Eran Ben Elisha <hidden> Date: 2018-09-16 15:59:46
On 9/13/2018 8:36 PM, Jakub Kicinski wrote:
On Thu, 13 Sep 2018 11:18:15 +0300, Eran Ben Elisha wrote:
quoted
The health spec is targeted for Real Time Alerting, in order to know when
something bad had happened to a PCI device
By spec you mean some standards body spec you implement or this
proposal is a spec?
This proposal is a spec
quoted
- Provide alert debug information
- Self healing
- If problem needs vendor support, provide a way to gather all needed debugging
information.
The health contains sensors which sense for malfunction. Once sensor triggered,
actions such as logs and correction can be taken.
Sensors are sensing the health state and can trigger correction action.
The sensors are divided into the following groups
- Hardware sensor - a sensor which is triggered by the device due to
malfunction.
- Software sensor - a sensor which is triggered by the software due to
malfunction.
Both group of sensors can be triggered due to error event or due to a periodic check.
Actions are the way to handle sensor events. Action can be in one of the
following groups:
- Dump - SW trace, SW dump, HW trace, HW dump
- Reset - Surgical correction (e.g. modify Q, flush Q, reset of device, etc)
Actions can be performed by SW or HW.
User is allowed to enable or disable sensors and sensor2action mapping.
This RFC man page patch describes the suggested API of devlink-health in order
to control sensors and actions.
I like the idea of configuring response to events like this, although
I'm not sure the name sensor is appropriate here - perhaps exception or
error would be better?
I was trying to avoid the negativity description. Have it called sensor
to avoid restricting the API for errors / exceptions only. I got the
same type of comment from Andrew as well devlink-health->devlink-bug.
But if other vendors driver developers don't see it can be expanded to
sensor which are not errors, then I guess we can refactor the names.
Are there going to be values reported?
It depends on the sensor. If it has data that would help in the debug,
then I assume yes, via the dumps.
I'm not so sure about HW sensors in relation to existing HWMON
infrastructure... I assume you're targeting things like say some HW
engine/block reporting it encountered an error? Sounds good, too.
yes, exactly.
Are the actions all envisioned to be performed by the driver?
Firmware? Hardware? I guess that distinction can be added later.
For FW/HW actions we would go back to the problem of persistence of
the setting since it was only implemented for params :S
The problem is not with FW action, the problem is when you try to set
sensor2action mapping for the FW/HW. this will need persistence
configuration mode. Sensor2action in SW shall be run-time mode (at least
as a start).
But it sound as this need some more tuning, to make it clear.
Is the dump option going to tie back into region snapshots?
no necessarily, dumping SW objects as well can be helpful
From: Stephen Hemminger <stephen@networkplumber.org> Date: 2018-09-17 00:53:45
On Thu, 13 Sep 2018 10:36:04 -0700
Jakub Kicinski [off-list ref] wrote:
On Thu, 13 Sep 2018 11:18:15 +0300, Eran Ben Elisha wrote:
quoted
The health spec is targeted for Real Time Alerting, in order to know when
something bad had happened to a PCI device
By spec you mean some standards body spec you implement or this
proposal is a spec?
quoted
- Provide alert debug information
- Self healing
- If problem needs vendor support, provide a way to gather all needed debugging
information.
The health contains sensors which sense for malfunction. Once sensor triggered,
actions such as logs and correction can be taken.
Sensors are sensing the health state and can trigger correction action.
The sensors are divided into the following groups
- Hardware sensor - a sensor which is triggered by the device due to
malfunction.
- Software sensor - a sensor which is triggered by the software due to
malfunction.
Both group of sensors can be triggered due to error event or due to a periodic check.
Actions are the way to handle sensor events. Action can be in one of the
following groups:
- Dump - SW trace, SW dump, HW trace, HW dump
- Reset - Surgical correction (e.g. modify Q, flush Q, reset of device, etc)
Actions can be performed by SW or HW.
User is allowed to enable or disable sensors and sensor2action mapping.
This RFC man page patch describes the suggested API of devlink-health in order
to control sensors and actions.
I like the idea of configuring response to events like this, although
I'm not sure the name sensor is appropriate here - perhaps exception or
error would be better? Are there going to be values reported?
I'm not so sure about HW sensors in relation to existing HWMON
infrastructure... I assume you're targeting things like say some HW
engine/block reporting it encountered an error? Sounds good, too.
Are the actions all envisioned to be performed by the driver?
Firmware? Hardware? I guess that distinction can be added later.
For FW/HW actions we would go back to the problem of persistence of
the setting since it was only implemented for params :S
Is the dump option going to tie back into region snapshots?
Why is this going under iproute rather than using one of the existing sensor API's.
For example Intel NIC's have thermal sensors etc.
From: Andrew Lunn <andrew@lunn.ch> Date: 2018-09-17 01:21:33
Why is this going under iproute rather than using one of the existing sensor API's.
For example Intel NIC's have thermal sensors etc.
Hi Stephen
These are not that sort of sensors. This is part of the naming problem
here. It is not really to do with health, it is about exceptions and
bugs. And the sensors are more like timeouts and watchdogs.
It is clear that the current names lead to a lot of confusion. Maybe:
health -> exception
sensor -> condition
Andrew
From: Eran Ben Elisha <hidden> Date: 2018-09-25 18:07:57
On 9/16/2018 1:37 PM, Eran Ben Elisha wrote:
On 9/13/2018 8:36 PM, Jakub Kicinski wrote:
quoted
On Thu, 13 Sep 2018 11:18:15 +0300, Eran Ben Elisha wrote:
quoted
The health spec is targeted for Real Time Alerting, in order to know
when
something bad had happened to a PCI device
By spec you mean some standards body spec you implement or this
proposal is a spec?
This proposal is a spec
quoted
quoted
- Provide alert debug information
- Self healing
- If problem needs vendor support, provide a way to gather all needed
debugging
information.
The health contains sensors which sense for malfunction. Once sensor
triggered,
actions such as logs and correction can be taken.
Sensors are sensing the health state and can trigger correction action.
The sensors are divided into the following groups
- Hardware sensor - a sensor which is triggered by the device due to
malfunction.
- Software sensor - a sensor which is triggered by the software due to
malfunction.
Both group of sensors can be triggered due to error event or due to a
periodic check.
Actions are the way to handle sensor events. Action can be in one of the
following groups:
- Dump - SW trace, SW dump, HW trace, HW dump
- Reset - Surgical correction (e.g. modify Q, flush Q, reset of
device, etc)
Actions can be performed by SW or HW.
User is allowed to enable or disable sensors and sensor2action mapping.
This RFC man page patch describes the suggested API of devlink-health
in order
to control sensors and actions.
I like the idea of configuring response to events like this, although
I'm not sure the name sensor is appropriate here - perhaps exception or
error would be better?
I was trying to avoid the negativity description. Have it called sensor
to avoid restricting the API for errors / exceptions only. I got the
same type of comment from Andrew as well devlink-health->devlink-bug.
But if other vendors driver developers don't see it can be expanded to
sensor which are not errors, then I guess we can refactor the names.
Are there going to be values reported?
It depends on the sensor. If it has data that would help in the debug,
then I assume yes, via the dumps.
quoted
I'm not so sure about HW sensors in relation to existing HWMON
infrastructure... I assume you're targeting things like say some HW
engine/block reporting it encountered an error? Sounds good, too.
yes, exactly.
quoted
Are the actions all envisioned to be performed by the driver?
Firmware? Hardware? I guess that distinction can be added later.
For FW/HW actions we would go back to the problem of persistence of
the setting since it was only implemented for params :S
The problem is not with FW action, the problem is when you try to set
sensor2action mapping for the FW/HW. this will need persistence
configuration mode. Sensor2action in SW shall be run-time mode (at least
as a start).
But it sound as this need some more tuning, to make it clear.
Revisiting this (before sending V2). My guideline is that persistency
inside the device is needed only when a persistence information is
needed before the driver loads. For any other configuration (i.e post HW
boot), one can use standard Linux scripts in order to control its
persistence information.
If any new sensor will be added that requires pre HW boot information,
the API can be extended later.
quoted
Is the dump option going to tie back into region snapshots?
no necessarily, dumping SW objects as well can be helpful
From: Eran Ben Elisha <hidden> Date: 2018-09-25 18:25:06
On 9/16/2018 10:57 PM, Andrew Lunn wrote:
quoted
Why is this going under iproute rather than using one of the existing sensor API's.
For example Intel NIC's have thermal sensors etc.
Hi Stephen
These are not that sort of sensors. This is part of the naming problem
here. It is not really to do with health, it is about exceptions and
bugs. And the sensors are more like timeouts and watchdogs.
It is clear that the current names lead to a lot of confusion. Maybe:
health -> exception
sensor -> condition
Andrew
I think those names renaming can work well.
(Sorry for that response, Local holiday season...)
Eran