The CPU to TXQ binding behavior of ixgbe vs. igb NIC driver are
somehow different. Normally I setup NIC IRQ-to-CPU bindings 1-to-1,
with script set_irq_affinity [1].
For forcing use of a specific HW TXQ, I normally force the CPU binding
of the process, either with "taskset" or with "netperf -T lcpu,rcpu".
This works fine with driver ixgbe, but not with driver igb. That is
with igb, the program forced to specific CPU, can still use another
TXQ. What am I missing?
I'm monitoring this with both:
1) watch -d sudo tc -s -d q ls dev ethXX
2) https://github.com/ffainelli/bqlmon
[1] https://github.com/netoptimizer/network-testing/blob/master/bin/set_irq_affinity
--
Best regards,
Jesper Dangaard Brouer
MSc.CS, Sr. Network Kernel Developer at Red Hat
Author of http://www.iptv-analyzer.org
LinkedIn: http://www.linkedin.com/in/brouer
From: Eric Dumazet <hidden> Date: 2014-09-17 14:32:42
On Wed, 2014-09-17 at 15:26 +0200, Jesper Dangaard Brouer wrote:
The CPU to TXQ binding behavior of ixgbe vs. igb NIC driver are
somehow different. Normally I setup NIC IRQ-to-CPU bindings 1-to-1,
with script set_irq_affinity [1].
For forcing use of a specific HW TXQ, I normally force the CPU binding
of the process, either with "taskset" or with "netperf -T lcpu,rcpu".
This works fine with driver ixgbe, but not with driver igb. That is
with igb, the program forced to specific CPU, can still use another
TXQ. What am I missing?
I'm monitoring this with both:
1) watch -d sudo tc -s -d q ls dev ethXX
2) https://github.com/ffainelli/bqlmon
[1] https://github.com/netoptimizer/network-testing/blob/master/bin/set_irq_affinity
Have you setup XPS ?
echo 0001 >/sys/class/net/ethX/queues/tx-0/xps_cpus
echo 0002 >/sys/class/net/ethX/queues/tx-1/xps_cpus
echo 0004 >/sys/class/net/ethX/queues/tx-2/xps_cpus
echo 0008 >/sys/class/net/ethX/queues/tx-3/xps_cpus
echo 0010 >/sys/class/net/ethX/queues/tx-4/xps_cpus
echo 0020 >/sys/class/net/ethX/queues/tx-5/xps_cpus
echo 0040 >/sys/class/net/ethX/queues/tx-6/xps_cpus
echo 0080 >/sys/class/net/ethX/queues/tx-7/xps_cpus
Or something like that, depending on number of cpus and TX queues.
On Wed, 17 Sep 2014 07:32:39 -0700
Eric Dumazet [off-list ref] wrote:
On Wed, 2014-09-17 at 15:26 +0200, Jesper Dangaard Brouer wrote:
quoted
The CPU to TXQ binding behavior of ixgbe vs. igb NIC driver are
somehow different. Normally I setup NIC IRQ-to-CPU bindings 1-to-1,
with script set_irq_affinity [1].
For forcing use of a specific HW TXQ, I normally force the CPU binding
of the process, either with "taskset" or with "netperf -T lcpu,rcpu".
This works fine with driver ixgbe, but not with driver igb. That is
with igb, the program forced to specific CPU, can still use another
TXQ. What am I missing?
I'm monitoring this with both:
1) watch -d sudo tc -s -d q ls dev ethXX
2) https://github.com/ffainelli/bqlmon
[1] https://github.com/netoptimizer/network-testing/blob/master/bin/set_irq_affinity
Have you setup XPS ?
echo 0001 >/sys/class/net/ethX/queues/tx-0/xps_cpus
echo 0002 >/sys/class/net/ethX/queues/tx-1/xps_cpus
echo 0004 >/sys/class/net/ethX/queues/tx-2/xps_cpus
echo 0008 >/sys/class/net/ethX/queues/tx-3/xps_cpus
echo 0010 >/sys/class/net/ethX/queues/tx-4/xps_cpus
echo 0020 >/sys/class/net/ethX/queues/tx-5/xps_cpus
echo 0040 >/sys/class/net/ethX/queues/tx-6/xps_cpus
echo 0080 >/sys/class/net/ethX/queues/tx-7/xps_cpus
Or something like that, depending on number of cpus and TX queues.
Thanks, that worked! They were all default set to "000" for igb, but
set correctly/like-above for ixgbe (strange).
Did:
$ export DEV=eth1 ; export NR_CPUS=11 ; \
for txq in `seq 0 $NR_CPUS` ; do \
file=/sys/class/net/${DEV}/queues/tx-${txq}/xps_cpus \
mask=`printf %X $((1<<$txq))`
test -e $file && sudo sh -c "echo $mask > $file" && \
grep . -H $file ;\
done
/sys/class/net/eth1/queues/tx-0/xps_cpus:001
/sys/class/net/eth1/queues/tx-1/xps_cpus:002
/sys/class/net/eth1/queues/tx-2/xps_cpus:004
/sys/class/net/eth1/queues/tx-3/xps_cpus:008
/sys/class/net/eth1/queues/tx-4/xps_cpus:010
/sys/class/net/eth1/queues/tx-5/xps_cpus:020
/sys/class/net/eth1/queues/tx-6/xps_cpus:040
/sys/class/net/eth1/queues/tx-7/xps_cpus:080
--
Best regards,
Jesper Dangaard Brouer
MSc.CS, Sr. Network Kernel Developer at Red Hat
Author of http://www.iptv-analyzer.org
LinkedIn: http://www.linkedin.com/in/brouer
From: Alexander Duyck <hidden> Date: 2014-09-17 14:59:52
On 09/17/2014 07:32 AM, Eric Dumazet wrote:
On Wed, 2014-09-17 at 15:26 +0200, Jesper Dangaard Brouer wrote:
quoted
The CPU to TXQ binding behavior of ixgbe vs. igb NIC driver are
somehow different. Normally I setup NIC IRQ-to-CPU bindings 1-to-1,
with script set_irq_affinity [1].
For forcing use of a specific HW TXQ, I normally force the CPU binding
of the process, either with "taskset" or with "netperf -T lcpu,rcpu".
This works fine with driver ixgbe, but not with driver igb. That is
with igb, the program forced to specific CPU, can still use another
TXQ. What am I missing?
I'm monitoring this with both:
1) watch -d sudo tc -s -d q ls dev ethXX
2) https://github.com/ffainelli/bqlmon
[1] https://github.com/netoptimizer/network-testing/blob/master/bin/set_irq_affinity
Have you setup XPS ?
echo 0001 >/sys/class/net/ethX/queues/tx-0/xps_cpus
echo 0002 >/sys/class/net/ethX/queues/tx-1/xps_cpus
echo 0004 >/sys/class/net/ethX/queues/tx-2/xps_cpus
echo 0008 >/sys/class/net/ethX/queues/tx-3/xps_cpus
echo 0010 >/sys/class/net/ethX/queues/tx-4/xps_cpus
echo 0020 >/sys/class/net/ethX/queues/tx-5/xps_cpus
echo 0040 >/sys/class/net/ethX/queues/tx-6/xps_cpus
echo 0080 >/sys/class/net/ethX/queues/tx-7/xps_cpus
Or something like that, depending on number of cpus and TX queues.
That was what I was thinking as well.
ixgbe has ATR which makes use of XPS to setup the transmit queues for a
1:1 mapping. The receive side of the flow is routed back to the same Rx
queue through flow director mappings.
In the case of igb it only has RSS and doesn't set a default XPS
configuration. So you should probably setup XPS and you might also want
to try and make use of RPS to try and steer receive packets since the Rx
queues won't match the Tx queues.
Thanks,
Alex
On Wed, 17 Sep 2014 07:59:51 -0700
Alexander Duyck [off-list ref] wrote:
On 09/17/2014 07:32 AM, Eric Dumazet wrote:
quoted
On Wed, 2014-09-17 at 15:26 +0200, Jesper Dangaard Brouer wrote:
quoted
The CPU to TXQ binding behavior of ixgbe vs. igb NIC driver are
somehow different. Normally I setup NIC IRQ-to-CPU bindings 1-to-1,
with script set_irq_affinity [1].
For forcing use of a specific HW TXQ, I normally force the CPU binding
of the process, either with "taskset" or with "netperf -T lcpu,rcpu".
This works fine with driver ixgbe, but not with driver igb. That is
with igb, the program forced to specific CPU, can still use another
TXQ. What am I missing?
I'm monitoring this with both:
1) watch -d sudo tc -s -d q ls dev ethXX
2) https://github.com/ffainelli/bqlmon
[1] https://github.com/netoptimizer/network-testing/blob/master/bin/set_irq_affinity
Have you setup XPS ?
echo 0001 >/sys/class/net/ethX/queues/tx-0/xps_cpus
echo 0002 >/sys/class/net/ethX/queues/tx-1/xps_cpus
echo 0004 >/sys/class/net/ethX/queues/tx-2/xps_cpus
echo 0008 >/sys/class/net/ethX/queues/tx-3/xps_cpus
echo 0010 >/sys/class/net/ethX/queues/tx-4/xps_cpus
echo 0020 >/sys/class/net/ethX/queues/tx-5/xps_cpus
echo 0040 >/sys/class/net/ethX/queues/tx-6/xps_cpus
echo 0080 >/sys/class/net/ethX/queues/tx-7/xps_cpus
Or something like that, depending on number of cpus and TX queues.
That was what I was thinking as well.
ixgbe has ATR which makes use of XPS to setup the transmit queues for a
1:1 mapping. The receive side of the flow is routed back to the same Rx
queue through flow director mappings.
In the case of igb it only has RSS and doesn't set a default XPS
configuration. So you should probably setup XPS and you might also want
to try and make use of RPS to try and steer receive packets since the Rx
queues won't match the Tx queues.
After setting up XPS to CPU 1:1 binding, it works most of the time.
Meaning, most of the traffic will go through the TXQ I've bound the
process to, BUT some packets can still choose another TXQ (observed
monitoring tc output and blqmon).
Could this be related to the missing RPS setup?
Can I get some hints setting up RPS?
--
Best regards,
Jesper Dangaard Brouer
MSc.CS, Sr. Network Kernel Developer at Red Hat
Author of http://www.iptv-analyzer.org
LinkedIn: http://www.linkedin.com/in/brouer
On Wed, 17 Sep 2014 07:59:51 -0700
Alexander Duyck [off-list ref] wrote:
quoted
On 09/17/2014 07:32 AM, Eric Dumazet wrote:
quoted
On Wed, 2014-09-17 at 15:26 +0200, Jesper Dangaard Brouer wrote:
quoted
The CPU to TXQ binding behavior of ixgbe vs. igb NIC driver are
somehow different. Normally I setup NIC IRQ-to-CPU bindings 1-to-1,
with script set_irq_affinity [1].
For forcing use of a specific HW TXQ, I normally force the CPU binding
of the process, either with "taskset" or with "netperf -T lcpu,rcpu".
This works fine with driver ixgbe, but not with driver igb. That is
with igb, the program forced to specific CPU, can still use another
TXQ. What am I missing?
I'm monitoring this with both:
1) watch -d sudo tc -s -d q ls dev ethXX
2) https://github.com/ffainelli/bqlmon
[1] https://github.com/netoptimizer/network-testing/blob/master/bin/set_irq_affinity
Have you setup XPS ?
echo 0001 >/sys/class/net/ethX/queues/tx-0/xps_cpus
echo 0002 >/sys/class/net/ethX/queues/tx-1/xps_cpus
echo 0004 >/sys/class/net/ethX/queues/tx-2/xps_cpus
echo 0008 >/sys/class/net/ethX/queues/tx-3/xps_cpus
echo 0010 >/sys/class/net/ethX/queues/tx-4/xps_cpus
echo 0020 >/sys/class/net/ethX/queues/tx-5/xps_cpus
echo 0040 >/sys/class/net/ethX/queues/tx-6/xps_cpus
echo 0080 >/sys/class/net/ethX/queues/tx-7/xps_cpus
Or something like that, depending on number of cpus and TX queues.
That was what I was thinking as well.
ixgbe has ATR which makes use of XPS to setup the transmit queues for a
1:1 mapping. The receive side of the flow is routed back to the same Rx
queue through flow director mappings.
In the case of igb it only has RSS and doesn't set a default XPS
configuration. So you should probably setup XPS and you might also want
to try and make use of RPS to try and steer receive packets since the Rx
queues won't match the Tx queues.
After setting up XPS to CPU 1:1 binding, it works most of the time.
Meaning, most of the traffic will go through the TXQ I've bound the
process to, BUT some packets can still choose another TXQ (observed
monitoring tc output and blqmon).
Could this be related to the missing RPS setup?
It helped setting up RPS, but not 100%. I see small periods of packets
going out on other TXQs, and sometimes as before some heavy flow will
find its way to another TXQ.
Can I get some hints setting up RPS?
My setup command now maps both XPS and RPS 1:1 to CPUs.
# Setup both RPS and XPS with a 1:1 binding to CPUs
export DEV=eth1 ; export NR_CPUS=11 ; \
for txq in `seq 0 $NR_CPUS` ; do \
file_xps=/sys/class/net/${DEV}/queues/tx-${txq}/xps_cpus \
file_rps=/sys/class/net/${DEV}/queues/rx-${txq}/rps_cpus \
mask=`printf %X $((1<<$txq))`
test -e $file_xps && sudo sh -c "echo $mask > $file_xps" && grep . -H $file_xps ;\
test -e $file_rps && sudo sh -c "echo $mask > $file_rps" && grep . -H $file_rps ;\
done
Output:
/sys/class/net/eth1/queues/tx-0/xps_cpus:001
/sys/class/net/eth1/queues/rx-0/rps_cpus:001
/sys/class/net/eth1/queues/tx-1/xps_cpus:002
/sys/class/net/eth1/queues/rx-1/rps_cpus:002
/sys/class/net/eth1/queues/tx-2/xps_cpus:004
/sys/class/net/eth1/queues/rx-2/rps_cpus:004
/sys/class/net/eth1/queues/tx-3/xps_cpus:008
/sys/class/net/eth1/queues/rx-3/rps_cpus:008
/sys/class/net/eth1/queues/tx-4/xps_cpus:010
/sys/class/net/eth1/queues/rx-4/rps_cpus:010
/sys/class/net/eth1/queues/tx-5/xps_cpus:020
/sys/class/net/eth1/queues/rx-5/rps_cpus:020
/sys/class/net/eth1/queues/tx-6/xps_cpus:040
/sys/class/net/eth1/queues/rx-6/rps_cpus:040
/sys/class/net/eth1/queues/tx-7/xps_cpus:080
/sys/class/net/eth1/queues/rx-7/rps_cpus:080
--
Best regards,
Jesper Dangaard Brouer
MSc.CS, Sr. Network Kernel Developer at Red Hat
Author of http://www.iptv-analyzer.org
LinkedIn: http://www.linkedin.com/in/brouer
From: Eric Dumazet <hidden> Date: 2014-09-18 13:33:31
On Thu, 2014-09-18 at 08:56 +0200, Jesper Dangaard Brouer wrote:
After setting up XPS to CPU 1:1 binding, it works most of the time.
Meaning, most of the traffic will go through the TXQ I've bound the
process to, BUT some packets can still choose another TXQ (observed
monitoring tc output and blqmon).
Note that for TCP, there are packets sent by the process doing the
send(), or packets sent by cpu doing TX completion (because of TSQ),
but also packets sent by ACK processing done in the reverse way.
As Alexander explained, if the ACK packets are delivered into another
CPU, then you might select another TX queue.
This is mostly prevented because of ooo_okay logic, meaning that a busy
bulk flow should stick into a single TX queue, no matter of XPS says.
A TCP_RR workload is free to chose whatever queue, because every packet
starting a RR block has the ooo_okay set (As prior data was delivered
and acknowledged by the opposite peer)
Could this be related to the missing RPS setup?
No, for this to really work, you need hardware support, so that ACK
packets take the same RX queue than the sent packets.
Can I get some hints setting up RPS?
Documentation/networking/scaling.txt is full of hints...
From: Eric Dumazet <hidden> Date: 2014-09-18 13:41:31
On Thu, 2014-09-18 at 06:33 -0700, Eric Dumazet wrote:
Note that for TCP, there are packets sent by the process doing the
send(), or packets sent by cpu doing TX completion (because of TSQ),
but also packets sent by ACK processing done in the reverse way.
As Alexander explained, if the ACK packets are delivered into another
CPU, then you might select another TX queue.
This is mostly prevented because of ooo_okay logic, meaning that a busy
bulk flow should stick into a single TX queue, no matter of XPS says.
A TCP_RR workload is free to chose whatever queue, because every packet
starting a RR block has the ooo_okay set (As prior data was delivered
and acknowledged by the opposite peer)
Last but not least, there is the fact that networking stacks use
mod_timer() to arm timers, and that by default, timer migration is on
( cf /proc/sys/kernel/timer_migration )
We probably should use mod_timer_pinned(), but I could not really see
any difference.
From: Eric Dumazet <hidden> Date: 2014-09-18 15:42:33
On Thu, 2014-09-18 at 06:41 -0700, Eric Dumazet wrote:
Last but not least, there is the fact that networking stacks use
mod_timer() to arm timers, and that by default, timer migration is on
( cf /proc/sys/kernel/timer_migration )
We probably should use mod_timer_pinned(), but I could not really see
any difference.
On Thu, 18 Sep 2014 08:42:31 -0700
Eric Dumazet [off-list ref] wrote:
On Thu, 2014-09-18 at 06:41 -0700, Eric Dumazet wrote:
quoted
Last but not least, there is the fact that networking stacks use
mod_timer() to arm timers, and that by default, timer migration is on
( cf /proc/sys/kernel/timer_migration )
I don't have this proc file on my system, as I didn't select CONFIG_SCHED_DEBUG.
quoted
We probably should use mod_timer_pinned(), but I could not really see
any difference.
Hmm... actually its quite noticeable :
Interesting impact.
I'm looking for some 1G hardware without multiqueue, so I can get
around this measurement constraint. And possibly turning it down to
100Mbit/s, so I can more easily measure the HoL blocking effect.
From: Eric Dumazet <hidden> Date: 2014-09-18 16:07:07
On Thu, 2014-09-18 at 08:42 -0700, Eric Dumazet wrote:
quoted hunk
On Thu, 2014-09-18 at 06:41 -0700, Eric Dumazet wrote:
quoted
Last but not least, there is the fact that networking stacks use
mod_timer() to arm timers, and that by default, timer migration is on
( cf /proc/sys/kernel/timer_migration )
We probably should use mod_timer_pinned(), but I could not really see
any difference.
And/or changing all occurences of HRTIMER_MODE_ABS in net/sched
into HRTIMER_MODE_ABS_PINNED
Because we _want_ qdisc being restarted on the right cpu for sure.
From: Eric Dumazet <hidden> Date: 2014-09-18 16:34:25
On Thu, 2014-09-18 at 17:59 +0200, Jesper Dangaard Brouer wrote:
On Thu, 18 Sep 2014 08:42:31 -0700
Eric Dumazet [off-list ref] wrote:
quoted
On Thu, 2014-09-18 at 06:41 -0700, Eric Dumazet wrote:
quoted
Last but not least, there is the fact that networking stacks use
mod_timer() to arm timers, and that by default, timer migration is on
( cf /proc/sys/kernel/timer_migration )
I don't have this proc file on my system, as I didn't select CONFIG_SCHED_DEBUG.
Interesting... this timer_migration stuff seems a bit scary to me.
quoted
quoted
We probably should use mod_timer_pinned(), but I could not really see
any difference.
Hmm... actually its quite noticeable :
Interesting impact.
I'm looking for some 1G hardware without multiqueue, so I can get
around this measurement constraint. And possibly turning it down to
100Mbit/s, so I can more easily measure the HoL blocking effect.
ethtool -L eth0 rx 1 tx 1
(Or similar if combined is used)
On Thu, 18 Sep 2014 09:34:24 -0700 Eric Dumazet [off-list ref] wrote:
On Thu, 2014-09-18 at 17:59 +0200, Jesper Dangaard Brouer wrote:
quoted
On Thu, 18 Sep 2014 08:42:31 -0700
Eric Dumazet [off-list ref] wrote:
[...]
quoted
I'm looking for some 1G hardware without multiqueue, so I can get
around this measurement constraint. And possibly turning it down to
100Mbit/s, so I can more easily measure the HoL blocking effect.
ethtool -L eth0 rx 1 tx 1
(Or similar if combined is used)
Thanks! - that solves my qdisc measurement problem :-)
And yes, I had to use:
ethtool -L eth1 combined 1
--
Best regards,
Jesper Dangaard Brouer
MSc.CS, Sr. Network Kernel Developer at Red Hat
Author of http://www.iptv-analyzer.org
LinkedIn: http://www.linkedin.com/in/brouer