The drivers currently rely on irq_set_affinity_hint() to either set the
affinity_hint that is consumed by the userspace and/or to enforce a custom
affinity.
irq_set_affinity_hint() as the name suggests is originally introduced to
only set the affinity_hint to help the userspace in guiding the interrupts
and not the affinity itself. However, since the commit
e2e64a932556 "genirq: Set initial affinity in irq_set_affinity_hint()"
irq_set_affinity_hint() also started applying the provided cpumask (if not
NULL) as the affinity for the interrupts. The issue that this commit was
trying to solve is to allow the drivers to enforce their affinity mask to
distribute the interrupts across the CPUs such that they don't always end
up on CPU0. This issue has been resolved within the irq subsystem since the
commit
a0c9259dc4e1 "irq/matrix: Spread interrupts on allocation"
Hence, there is no need for the drivers to overwrite the affinity to spread
as it is dynamically performed at the time of allocation.
Also, irq_set_affinity_hint() setting affinity unconditionally introduces
issues for the drivers that only want to set their affinity_hint and not the
affinity itself as for these driver interrupts the default_smp_affinity_mask
is completely ignored (for detailed investigation please refer to [1]).
Unfortunately reverting the commit e2e64a932556 is not an option at this
point for two reasons [2]:
- Several drivers for a valid reason (performance) rely on this API to
enforce their affinity mask
- Until very recently this was the only exported interface that was
available
To clear this out Thomas has come up with the following interfaces:
- irq_set_affinity(): only sets affinity of an IRQ [3]
- irq_update_affinity_hint(): Only sets the hint [4]
- irq_set_affinity_and_hint(): Sets both affinity and the hint mask [4]
The first API is already merged in the linux-next tree and the patch
that introduces the other two interfaces are included with this patch-set.
To move to the stage where we can safely get rid of the
irq_set_affinity_hint(), which has been marked deprecated, we have to
move all its consumers to these new interfaces. In this patch-set, I have
done that for a few drivers and will hopefully try to move the remaining of
them in the coming days.
Testing
-------
In terms of testing, I have performed some basic testing on x86 to verify
things such as the interrupts are evenly spread on all CPUs, hint mask is
correctly set etc. for the drivers - i40e, iavf, mlx5, mlx4, ixgbe, i40iw
and enic on top of:
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git
So more testing is probably required for these and the drivers that I didn't
test and any help will be much appreciated.
Notes
-----
- I was told that i40iw driver is going to be replaced by irdma, however,
the new driver didn't land in Linus's tree yet. Once it does I will send
a follow up patch for that as well.
- For the mpt3sas driver I decided to go with the usage of
irq_set_affinity_and_hint over irq_set_affinity based on my little
analysis of it and the megaraid driver. However, if we are sure that it
is not required then I can replace it with just irq_set_affinity as one
of its comment suggests.
Change from v1 [5]
------------------
- Fixed compilation error by adding the new interface definitions for cases
where CONFIG_SMP is not defined
- Fixed function usage in megaraid_sas and removed unnecessary variable
(Robin Murphy)
- Removed unwanted #if/endif from mlx4 (Leon Romanovsky)
- Other indentation related fixes
[1] https://lore.kernel.org/lkml/1a044a14-0884-eedb-5d30-28b4bec24b23@redhat.com/
[2] https://lore.kernel.org/linux-pci/d1d5e797-49ee-4968-88c6-c07119343492@arm.com/
[3] https://lore.kernel.org/linux-arm-kernel/20210518091725.046774792@linutronix.de/
[4] https://lore.kernel.org/patchwork/patch/1434326/
[5] https://lore.kernel.org/linux-scsi/20210617182242.8637-1-nitesh@redhat.com/
Nitesh Narayan Lal (13):
iavf: Use irq_update_affinity_hint
i40e: Use irq_update_affinity_hint
scsi: megaraid_sas: Use irq_set_affinity_and_hint
scsi: mpt3sas: Use irq_set_affinity_and_hint
RDMA/i40iw: Use irq_update_affinity_hint
enic: Use irq_update_affinity_hint
be2net: Use irq_update_affinity_hint
ixgbe: Use irq_update_affinity_hint
mailbox: Use irq_update_affinity_hint
scsi: lpfc: Use irq_set_affinity
hinic: Use irq_set_affinity_and_hint
net/mlx5: Use irq_update_affinity_hint
net/mlx4: Use irq_update_affinity_hint
Thomas Gleixner (1):
genirq: Provide new interfaces for affinity hints
drivers/infiniband/hw/i40iw/i40iw_main.c | 4 +-
drivers/mailbox/bcm-flexrm-mailbox.c | 4 +-
drivers/net/ethernet/cisco/enic/enic_main.c | 8 +--
drivers/net/ethernet/emulex/benet/be_main.c | 4 +-
drivers/net/ethernet/huawei/hinic/hinic_rx.c | 4 +-
drivers/net/ethernet/intel/i40e/i40e_main.c | 8 +--
drivers/net/ethernet/intel/iavf/iavf_main.c | 8 +--
drivers/net/ethernet/intel/ixgbe/ixgbe_main.c | 10 ++--
drivers/net/ethernet/mellanox/mlx4/eq.c | 8 ++-
.../net/ethernet/mellanox/mlx5/core/pci_irq.c | 6 +--
drivers/scsi/lpfc/lpfc_init.c | 4 +-
drivers/scsi/megaraid/megaraid_sas_base.c | 27 +++++-----
drivers/scsi/mpt3sas/mpt3sas_base.c | 21 ++++----
include/linux/interrupt.h | 53 ++++++++++++++++++-
kernel/irq/manage.c | 8 +--
15 files changed, 113 insertions(+), 64 deletions(-)
--
From: Thomas Gleixner <redacted>
The discussion about removing the side effect of irq_set_affinity_hint() of
actually applying the cpumask (if not NULL) as affinity to the interrupt,
unearthed a few unpleasantries:
1) The modular perf drivers rely on the current behaviour for the very
wrong reasons.
2) While none of the other drivers prevents user space from changing
the affinity, a cursorily inspection shows that there are at least
expectations in some drivers.
#1 needs to be cleaned up anyway, so that's not a problem
#2 might result in subtle regressions especially when irqbalanced (which
nowadays ignores the affinity hint) is disabled.
Provide new interfaces:
irq_update_affinity_hint() - Only sets the affinity hint pointer
irq_set_affinity_and_hint() - Set the pointer and apply the affinity to
the interrupt
Make irq_set_affinity_hint() a wrapper around irq_apply_affinity_hint() and
document it to be phased out.
Signed-off-by: Thomas Gleixner <redacted>
Signed-off-by: Nitesh Narayan Lal <redacted>
Link: https://lore.kernel.org/r/20210501021832.743094-1-jesse.brandeburg@intel.com
---
include/linux/interrupt.h | 53 ++++++++++++++++++++++++++++++++++++++-
kernel/irq/manage.c | 8 +++---
2 files changed, 56 insertions(+), 5 deletions(-)
@@ -328,7 +328,46 @@ extern int irq_force_affinity(unsigned int irq, const struct cpumask *cpumask);externintirq_can_set_affinity(unsignedintirq);externintirq_select_affinity(unsignedintirq);-externintirq_set_affinity_hint(unsignedintirq,conststructcpumask*m);+externint__irq_apply_affinity_hint(unsignedintirq,conststructcpumask*m,+boolsetaffinity);++/**+*irq_update_affinity_hint-Updatetheaffinityhint+*@irq:Interrupttoupdate+*@cpumask:cpumaskpointer(NULLtoclearthehint)+*+*Updatestheaffinityhint,butdoesnotchangetheaffinityoftheinterrupt.+*/+staticinlineint+irq_update_affinity_hint(unsignedintirq,conststructcpumask*m)+{+return__irq_apply_affinity_hint(irq,m,false);+}++/**+*irq_set_affinity_and_hint-Updatetheaffinityhintandapplytheprovided+*cpumasktotheinterrupt+*@irq:Interrupttoupdate+*@cpumask:cpumaskpointer(NULLtoclearthehint)+*+*Updatestheaffinityhintandif@cpumaskisnotNULLitappliesitas+*theaffinityofthatinterrupt.+*/+staticinlineint+irq_set_affinity_and_hint(unsignedintirq,conststructcpumask*m)+{+return__irq_apply_affinity_hint(irq,m,true);+}++/*+*Deprecated.Useirq_update_affinity_hint()orirq_set_affinity_and_hint()+*instead.+*/+staticinlineintirq_set_affinity_hint(unsignedintirq,conststructcpumask*m)+{+returnirq_set_affinity_and_hint(irq,m);+}+externintirq_update_affinity_desc(unsignedintirq,structirq_affinity_desc*affinity);
@@ -360,6 +399,18 @@ static inline int irq_can_set_affinity(unsigned int irq)staticinlineintirq_select_affinity(unsignedintirq){return0;}+staticinlineintirq_update_affinity_hint(unsignedintirq,+conststructcpumask*m)+{+return-EINVAL;+}++staticinlineintirq_set_affinity_and_hint(unsignedintirq,+conststructcpumask*m)+{+return-EINVAL;+}+staticinlineintirq_set_affinity_hint(unsignedintirq,conststructcpumask*m){
@@ -487,7 +487,8 @@ int irq_force_affinity(unsigned int irq, const struct cpumask *cpumask)}EXPORT_SYMBOL_GPL(irq_force_affinity);-intirq_set_affinity_hint(unsignedintirq,conststructcpumask*m)+int__irq_apply_affinity_hint(unsignedintirq,conststructcpumask*m,+boolsetaffinity){unsignedlongflags;structirq_desc*desc=irq_get_desc_lock(irq,&flags,IRQ_GET_DESC_CHECK_GLOBAL);
@@ -496,12 +497,11 @@ int irq_set_affinity_hint(unsigned int irq, const struct cpumask *m)return-EINVAL;desc->affinity_hint=m;irq_put_desc_unlock(desc,flags);-/* set the initial affinity to prevent every interrupt being on CPU0 */-if(m)+if(m&&setaffinity)__irq_set_affinity(irq,m,false);return0;}-EXPORT_SYMBOL_GPL(irq_set_affinity_hint);+EXPORT_SYMBOL_GPL(__irq_apply_affinity_hint);staticvoidirq_affinity_notify(structwork_struct*work){
The driver uses irq_set_affinity_hint() for two purposes:
- To set the affinity_hint which is consumed by the userspace for
distributing the interrupts
- To apply an affinity that it provides for the iavf interrupts
The latter is done to ensure that all the interrupts are evenly spread
across all available CPUs. However, since commit a0c9259dc4e1 ("irq/matrix:
Spread interrupts on allocation") the spreading of interrupts is
dynamically performed at the time of allocation. Hence, there is no need
for the drivers to enforce their own affinity for the spreading of
interrupts.
Also, irq_set_affinity_hint() applying the provided cpumask as an affinity
for the interrupt is an undocumented side effect. To remove this side
effect irq_set_affinity_hint() has been marked as deprecated and new
interfaces have been introduced. Hence, replace the irq_set_affinity_hint()
with the new interface irq_update_affinity_hint() that only sets the
pointer for the affinity_hint.
Signed-off-by: Nitesh Narayan Lal <redacted>
---
drivers/net/ethernet/intel/iavf/iavf_main.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
The driver uses irq_set_affinity_hint() for two purposes:
- To set the affinity_hint which is consumed by the userspace for
distributing the interrupts
- To apply an affinity that it provides for the i40e interrupts
The latter is done to ensure that all the interrupts are evenly spread
across all available CPUs. However, since commit a0c9259dc4e1 ("irq/matrix:
Spread interrupts on allocation") the spreading of interrupts is
dynamically performed at the time of allocation. Hence, there is no need
for the drivers to enforce their own affinity for the spreading of
interrupts.
Also, irq_set_affinity_hint() applying the provided cpumask as an affinity
for the interrupt is an undocumented side effect. To remove this side
effect irq_set_affinity_hint() has been marked as deprecated and new
interfaces have been introduced. Hence, replace the irq_set_affinity_hint()
with the new interface irq_update_affinity_hint() that only sets the
pointer for the affinity_hint.
Signed-off-by: Nitesh Narayan Lal <redacted>
---
drivers/net/ethernet/intel/i40e/i40e_main.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
The driver uses irq_set_affinity_hint() specifically for the high IOPS
queue interrupts for two purposes:
- To set the affinity_hint which is consumed by the userspace for
distributing the interrupts
- To apply an affinity that it provides
The driver enforces its own affinity to bind the high IOPS queue interrupts
to the local NUMA node. However, irq_set_affinity_hint() applying the
provided cpumask as an affinity for the interrupt is an undocumented side
effect.
To remove this side effect irq_set_affinity_hint() has been marked
as deprecated and new interfaces have been introduced. Hence, replace the
irq_set_affinity_hint() with the new interface irq_set_affinity_and_hint()
that clearly indicates the purpose of the usage and is meant to apply the
affinity and set the affinity_hint pointer. Also, replace
irq_set_affinity_hint() with irq_update_affinity_hint() when only
affinity_hint needs to be updated.
Change the megasas_set_high_iops_queue_affinity_hint function name to
megasas_set_high_iops_queue_affinity_and_hint to clearly indicate that the
function is setting both affinity and affinity_hint.
Signed-off-by: Nitesh Narayan Lal <redacted>
---
drivers/scsi/megaraid/megaraid_sas_base.c | 27 +++++++++++++----------
1 file changed, 15 insertions(+), 12 deletions(-)
The driver uses irq_set_affinity_hint() specifically for the high IOPS
queue interrupts for two purposes:
- To set the affinity_hint which is consumed by the userspace for
distributing the interrupts
- To apply an affinity that it provides
The driver enforces its own affinity to bind the high IOPS queue interrupts
to the local NUMA node. However, irq_set_affinity_hint() applying the
provided cpumask as an affinity (if not NULL) for the interrupt is an
undocumented side effect.
To remove this side effect irq_set_affinity_hint() has been marked
as deprecated and new interfaces have been introduced. Hence, replace the
irq_set_affinity_hint() with the new interface irq_set_affinity_and_hint()
that clearly indicates the purpose of the usage and is meant to apply the
affinity and set the affinity_hint pointer. Also, replace
irq_set_affinity_hint() with irq_update_affinity_hint() when only
affinity_hint needs to be updated.
Signed-off-by: Nitesh Narayan Lal <redacted>
---
drivers/scsi/mpt3sas/mpt3sas_base.c | 21 ++++++++++-----------
1 file changed, 10 insertions(+), 11 deletions(-)
The driver uses irq_set_affinity_hint() to update the affinity_hint mask
that is consumed by the userspace to distribute the interrupts. However,
under the hood irq_set_affinity_hint() also applies the provided cpumask
(if not NULL) as the affinity for the given interrupt which is an
undocumented side effect.
To remove this side effect irq_set_affinity_hint() has been marked
as deprecated and new interfaces have been introduced. Hence, replace the
irq_set_affinity_hint() with the new interface irq_update_affinity_hint()
that only updates the affinity_hint pointer.
Signed-off-by: Nitesh Narayan Lal <redacted>
---
drivers/infiniband/hw/i40iw/i40iw_main.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
The driver uses irq_set_affinity_hint() to update the affinity_hint mask
that is consumed by the userspace to distribute the interrupts. However,
under the hood irq_set_affinity_hint() also applies the provided cpumask
(if not NULL) as the affinity for the given interrupt which is an
undocumented side effect.
To remove this side effect irq_set_affinity_hint() has been marked
as deprecated and new interfaces have been introduced. Hence, replace the
irq_set_affinity_hint() with the new interface irq_update_affinity_hint()
that only updates the affinity_hint pointer.
Signed-off-by: Nitesh Narayan Lal <redacted>
---
drivers/net/ethernet/cisco/enic/enic_main.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
The driver uses irq_set_affinity_hint() to update the affinity_hint mask
that is consumed by the userspace to distribute the interrupts. However,
under the hood irq_set_affinity_hint() also applies the provided cpumask
(if not NULL) as the affinity for the given interrupt which is an
undocumented side effect.
To remove this side effect irq_set_affinity_hint() has been marked
as deprecated and new interfaces have been introduced. Hence, replace the
irq_set_affinity_hint() with the new interface irq_update_affinity_hint()
that only updates the affinity_hint pointer.
Signed-off-by: Nitesh Narayan Lal <redacted>
---
drivers/net/ethernet/emulex/benet/be_main.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
The driver uses irq_set_affinity_hint() to update the affinity_hint mask
that is consumed by the userspace to distribute the interrupts. However,
under the hood irq_set_affinity_hint() also applies the provided cpumask
(if not NULL) as the affinity for the given interrupt which is an
undocumented side effect.
To remove this side effect irq_set_affinity_hint() has been marked
as deprecated and new interfaces have been introduced. Hence, replace the
irq_set_affinity_hint() with the new interface irq_update_affinity_hint()
that only updates the affinity_hint pointer.
Signed-off-by: Nitesh Narayan Lal <redacted>
---
drivers/net/ethernet/intel/ixgbe/ixgbe_main.c | 10 +++++-----
1 file changed, 5 insertions(+), 5 deletions(-)
@@ -3243,8 +3243,8 @@ static int ixgbe_request_msix_irqs(struct ixgbe_adapter *adapter)/* If Flow Director is enabled, set interrupt affinity */if(adapter->flags&IXGBE_FLAG_FDIR_HASH_CAPABLE){/* assign the mask for this irq */-irq_set_affinity_hint(entry->vector,-&q_vector->affinity_mask);+irq_update_affinity_hint(entry->vector,+&q_vector->affinity_mask);}}
@@ -3260,8 +3260,8 @@ static int ixgbe_request_msix_irqs(struct ixgbe_adapter *adapter)free_queue_irqs:while(vector){vector--;-irq_set_affinity_hint(adapter->msix_entries[vector].vector,-NULL);+irq_update_affinity_hint(adapter->msix_entries[vector].vector,+NULL);free_irq(adapter->msix_entries[vector].vector,adapter->q_vector[vector]);}
@@ -3394,7 +3394,7 @@ static void ixgbe_free_irq(struct ixgbe_adapter *adapter)continue;/* clear the affinity_mask in the IRQ descriptor */-irq_set_affinity_hint(entry->vector,NULL);+irq_update_affinity_hint(entry->vector,NULL);free_irq(entry->vector,q_vector);}
The driver uses irq_set_affinity_hint() to:
- Set the affinity_hint which is consumed by the userspace for
distributing the interrupts
- Enforce affinity
As per commit 6ac17fe8c14a ("mailbox: bcm-flexrm-mailbox: Set IRQ affinity
hint for FlexRM ring IRQs") the latter is done to ensure that the FlexRM
ring interrupts are evenly spread across all available CPUs. However, since
commit a0c9259dc4e1 ("irq/matrix: Spread interrupts on allocation") the
spreading of interrupts is dynamically performed at the time of allocation.
Hence, there is no need for the drivers to enforce their own affinity for
the spreading of interrupts.
Also, irq_set_affinity_hint() applying the provided cpumask as an affinity
for the interrupt is an undocumented side effect. To remove this side
effect irq_set_affinity_hint() has been marked as deprecated and new
interfaces have been introduced. Hence, replace the irq_set_affinity_hint()
with the new interface irq_update_affinity_hint() that only sets the
affinity_hint pointer.
Signed-off-by: Nitesh Narayan Lal <redacted>
---
drivers/mailbox/bcm-flexrm-mailbox.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
@@ -1298,7 +1298,7 @@ static int flexrm_startup(struct mbox_chan *chan)val=(num_online_cpus()<val)?val/num_online_cpus():1;cpumask_set_cpu((ring->num/val)%num_online_cpus(),&ring->irq_aff_hint);-ret=irq_set_affinity_hint(ring->irq,&ring->irq_aff_hint);+ret=irq_update_affinity_hint(ring->irq,&ring->irq_aff_hint);if(ret){dev_err(ring->mbox->dev,"failed to set IRQ affinity hint for ring%d\n",
The driver uses irq_set_affinity_hint to set the affinity for the lpfc
interrupts to a mask corresponding to the local NUMA node to avoid
performance overhead on AMD architectures.
However, irq_set_affinity_hint() setting the affinity is an undocumented
side effect that this function also sets the affinity under the hood.
To remove this side effect irq_set_affinity_hint() has been marked as
deprecated and new interfaces have been introduced.
Also, as per the commit dcaa21367938 ("scsi: lpfc: Change default IRQ model
on AMD architectures"):
"On AMD architecture, revert the irq allocation to the normal style
(non-managed) and then use irq_set_affinity_hint() to set the cpu affinity
and disable user-space rebalancing."
we don't really need to set the affinity_hint as user-space rebalancing for
the lpfc interrupts is not desired.
Hence, replace the irq_set_affinity_hint() with irq_set_affinity() which
only applies the affinity for the interrupts.
Signed-off-by: Nitesh Narayan Lal <redacted>
---
drivers/scsi/lpfc/lpfc_init.c | 4 +---
1 file changed, 1 insertion(+), 3 deletions(-)
The driver uses irq_set_affinity_hint() to:
- Set the affinity_hint which is consumed by the userspace for
distributing the interrupts
- Enforce affinity
As per commit 352f58b0d9f2 ("net-next/hinic: Set Rxq irq to specific cpu
for NUMA"), the hinic driver enforces its own affinity to bind IRQs to the
local NUMA node. However, irq_set_affinity_hint() applying the provided
cpumask as an affinity for the interrupt is an undocumented side effect.
To remove this side effect irq_set_affinity_hint() has been marked as
deprecated and new interfaces have been introduced. Hence, replace the
irq_set_affinity_hint() with the new interface
irq_set_affinity_and_hint() that applies the affinity and updates the
affinity_hint pointer. Also, use irq_update_affinity() when only
affinity_hint needs to be updated.
Signed-off-by: Nitesh Narayan Lal <redacted>
---
drivers/net/ethernet/huawei/hinic/hinic_rx.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
The driver uses irq_set_affinity_hint() to update the affinity_hint mask
that is consumed by the userspace to distribute the interrupts. However,
under the hood irq_set_affinity_hint() also applies the provided cpumask
(if not NULL) as the affinity for the given interrupt which is an
undocumented side effect.
To remove this side effect irq_set_affinity_hint() has been marked
as deprecated and new interfaces have been introduced. Hence, replace the
irq_set_affinity_hint() with the new interface irq_update_affinity_hint()
that only updates the affinity_hint pointer.
Signed-off-by: Nitesh Narayan Lal <redacted>
Reviewed-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/pci_irq.c | 6 +++---
1 file changed, 3 insertions(+), 3 deletions(-)
The driver uses irq_set_affinity_hint() to update the affinity_hint mask
that is consumed by the userspace to distribute the interrupts. However,
under the hood irq_set_affinity_hint() also applies the provided cpumask
(if not NULL) as the affinity for the given interrupt which is an
undocumented side effect.
To remove this side effect irq_set_affinity_hint() has been marked
as deprecated and new interfaces have been introduced. Hence, replace the
irq_set_affinity_hint() with the new interface irq_update_affinity_hint()
that only updates the affinity_hint pointer.
Signed-off-by: Nitesh Narayan Lal <redacted>
Reviewed-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx4/eq.c | 8 +++-----
1 file changed, 3 insertions(+), 5 deletions(-)
On Tue, Jun 29, 2021 at 11:28 AM Nitesh Narayan Lal [off-list ref] wrote:
The drivers currently rely on irq_set_affinity_hint() to either set the
affinity_hint that is consumed by the userspace and/or to enforce a custom
affinity.
irq_set_affinity_hint() as the name suggests is originally introduced to
only set the affinity_hint to help the userspace in guiding the interrupts
and not the affinity itself. However, since the commit
e2e64a932556 "genirq: Set initial affinity in irq_set_affinity_hint()"
irq_set_affinity_hint() also started applying the provided cpumask (if not
NULL) as the affinity for the interrupts. The issue that this commit was
trying to solve is to allow the drivers to enforce their affinity mask to
distribute the interrupts across the CPUs such that they don't always end
up on CPU0. This issue has been resolved within the irq subsystem since the
commit
a0c9259dc4e1 "irq/matrix: Spread interrupts on allocation"
Hence, there is no need for the drivers to overwrite the affinity to spread
as it is dynamically performed at the time of allocation.
Also, irq_set_affinity_hint() setting affinity unconditionally introduces
issues for the drivers that only want to set their affinity_hint and not the
affinity itself as for these driver interrupts the default_smp_affinity_mask
is completely ignored (for detailed investigation please refer to [1]).
Unfortunately reverting the commit e2e64a932556 is not an option at this
point for two reasons [2]:
- Several drivers for a valid reason (performance) rely on this API to
enforce their affinity mask
- Until very recently this was the only exported interface that was
available
To clear this out Thomas has come up with the following interfaces:
- irq_set_affinity(): only sets affinity of an IRQ [3]
- irq_update_affinity_hint(): Only sets the hint [4]
- irq_set_affinity_and_hint(): Sets both affinity and the hint mask [4]
The first API is already merged in the linux-next tree and the patch
that introduces the other two interfaces are included with this patch-set.
To move to the stage where we can safely get rid of the
irq_set_affinity_hint(), which has been marked deprecated, we have to
move all its consumers to these new interfaces. In this patch-set, I have
done that for a few drivers and will hopefully try to move the remaining of
them in the coming days.
Testing
-------
In terms of testing, I have performed some basic testing on x86 to verify
things such as the interrupts are evenly spread on all CPUs, hint mask is
correctly set etc. for the drivers - i40e, iavf, mlx5, mlx4, ixgbe, i40iw
and enic on top of:
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git
So more testing is probably required for these and the drivers that I didn't
test and any help will be much appreciated.
Notes
-----
- I was told that i40iw driver is going to be replaced by irdma, however,
the new driver didn't land in Linus's tree yet. Once it does I will send
a follow up patch for that as well.
- For the mpt3sas driver I decided to go with the usage of
irq_set_affinity_and_hint over irq_set_affinity based on my little
analysis of it and the megaraid driver. However, if we are sure that it
is not required then I can replace it with just irq_set_affinity as one
of its comment suggests.
Change from v1 [5]
------------------
- Fixed compilation error by adding the new interface definitions for cases
where CONFIG_SMP is not defined
- Fixed function usage in megaraid_sas and removed unnecessary variable
(Robin Murphy)
- Removed unwanted #if/endif from mlx4 (Leon Romanovsky)
- Other indentation related fixes
[1] https://lore.kernel.org/lkml/1a044a14-0884-eedb-5d30-28b4bec24b23@redhat.com/
[2] https://lore.kernel.org/linux-pci/d1d5e797-49ee-4968-88c6-c07119343492@arm.com/
[3] https://lore.kernel.org/linux-arm-kernel/20210518091725.046774792@linutronix.de/
[4] https://lore.kernel.org/patchwork/patch/1434326/
[5] https://lore.kernel.org/linux-scsi/20210617182242.8637-1-nitesh@redhat.com/
Nitesh Narayan Lal (13):
iavf: Use irq_update_affinity_hint
i40e: Use irq_update_affinity_hint
scsi: megaraid_sas: Use irq_set_affinity_and_hint
scsi: mpt3sas: Use irq_set_affinity_and_hint
RDMA/i40iw: Use irq_update_affinity_hint
enic: Use irq_update_affinity_hint
be2net: Use irq_update_affinity_hint
ixgbe: Use irq_update_affinity_hint
mailbox: Use irq_update_affinity_hint
scsi: lpfc: Use irq_set_affinity
hinic: Use irq_set_affinity_and_hint
net/mlx5: Use irq_update_affinity_hint
net/mlx4: Use irq_update_affinity_hint
Thomas Gleixner (1):
genirq: Provide new interfaces for affinity hints
drivers/infiniband/hw/i40iw/i40iw_main.c | 4 +-
drivers/mailbox/bcm-flexrm-mailbox.c | 4 +-
drivers/net/ethernet/cisco/enic/enic_main.c | 8 +--
drivers/net/ethernet/emulex/benet/be_main.c | 4 +-
drivers/net/ethernet/huawei/hinic/hinic_rx.c | 4 +-
drivers/net/ethernet/intel/i40e/i40e_main.c | 8 +--
drivers/net/ethernet/intel/iavf/iavf_main.c | 8 +--
drivers/net/ethernet/intel/ixgbe/ixgbe_main.c | 10 ++--
drivers/net/ethernet/mellanox/mlx4/eq.c | 8 ++-
.../net/ethernet/mellanox/mlx5/core/pci_irq.c | 6 +--
drivers/scsi/lpfc/lpfc_init.c | 4 +-
drivers/scsi/megaraid/megaraid_sas_base.c | 27 +++++-----
drivers/scsi/mpt3sas/mpt3sas_base.c | 21 ++++----
include/linux/interrupt.h | 53 ++++++++++++++++++-
kernel/irq/manage.c | 8 +--
15 files changed, 113 insertions(+), 64 deletions(-)
--
Gentle ping.
Any comments or suggestions on any of the patches included in this series?
--
Thanks
Nitesh
Gentle ping.
Any comments or suggestions on any of the patches included in this series?
Please wait for -rc1, rebase and resend.
At least i40iw was deleted during merge window.
In -rc1 some non-trivial mlx5 changes also went in. I was going through
these changes and it seems after your patch
e4e3f24b822f: ("net/mlx5: Provide cpumask at EQ creation phase")
we do want to control the affinity for the mlx5 interrupts from the driver.
Is that correct? This would mean that we should use
irq_set_affinity_and_hint() instead
of irq_update_affinity_hint().
--
Thanks
Nitesh
Gentle ping.
Any comments or suggestions on any of the patches included in this series?
Please wait for -rc1, rebase and resend.
At least i40iw was deleted during merge window.
In -rc1 some non-trivial mlx5 changes also went in. I was going through
these changes and it seems after your patch
e4e3f24b822f: ("net/mlx5: Provide cpumask at EQ creation phase")
we do want to control the affinity for the mlx5 interrupts from the driver.
Is that correct?
We would like to create devices with correct affinity from the
beginning. For this, we will introduce extension to devlink to control
affinity that will be used prior initialization sequence.
Currently, netdev users who don't want irqbalance are digging into
their procfs, reconfigure affinity on already existing devices and
hope for the best.
This is even more cumbersome for the SIOV use case, where every physical
NIC PCI device will/can create thousands of lightweights netdevs that will
be forwarded to the containers later. These containers are limited to known
CPU cores, so no reason do not limit netdev device too.
The same goes for other sub-functions of that PCI device, like RDMA,
vdpa e.t.c.
This would mean that we should use irq_set_affinity_and_hint() instead
of irq_update_affinity_hint().
On Tue, Jul 13, 2021 at 1:01 AM Leon Romanovsky [off-list ref] wrote:
On Mon, Jul 12, 2021 at 05:27:05PM -0400, Nitesh Lal wrote:
quoted
Hi Leon,
<snip>
quoted
quoted
quoted
Gentle ping.
Any comments or suggestions on any of the patches included in this series?
Please wait for -rc1, rebase and resend.
At least i40iw was deleted during merge window.
In -rc1 some non-trivial mlx5 changes also went in. I was going through
these changes and it seems after your patch
e4e3f24b822f: ("net/mlx5: Provide cpumask at EQ creation phase")
we do want to control the affinity for the mlx5 interrupts from the driver.
Is that correct?
We would like to create devices with correct affinity from the
beginning. For this, we will introduce extension to devlink to control
affinity that will be used prior initialization sequence.
Currently, netdev users who don't want irqbalance are digging into
their procfs, reconfigure affinity on already existing devices and
hope for the best.
This is even more cumbersome for the SIOV use case, where every physical
NIC PCI device will/can create thousands of lightweights netdevs that will
be forwarded to the containers later. These containers are limited to known
CPU cores, so no reason do not limit netdev device too.
The same goes for other sub-functions of that PCI device, like RDMA,
vdpa e.t.c.
quoted
This would mean that we should use irq_set_affinity_and_hint() instead
of irq_update_affinity_hint().
I think so.
Thanks, will make that change in the patch and re-send.
I will also drop your reviewed-by for the mlx5 patch so that you can
have a look at it again, please let me know if you have any
objections.
--
Thanks
Nitesh