From: Eric Dumazet <hidden> Date: 2017-02-01 17:05:37
Hi all
Is anybody working on adding QUEUED spinlocks to powerpc 64bit ?
I've seen past attempts with ticket spinlocks
( https://patchwork.ozlabs.org/patch/449381/ and other related links )
But it looks ticket spinlocks are a thing of the past.
Thanks.
From: Benjamin Herrenschmidt <benh@kernel.crashing.org> Date: 2017-02-01 20:37:44
On Wed, 2017-02-01 at 09:05 -0800, Eric Dumazet wrote:
Hi all
Is anybody working on adding QUEUED spinlocks to powerpc 64bit ?
I've seen past attempts with ticket spinlocks
( https://patchwork.ozlabs.org/patch/449381/ and other related links
)
But it looks ticket spinlocks are a thing of the past.
From: Michael Ellerman <mpe@ellerman.id.au> Date: 2017-02-02 04:05:00
Benjamin Herrenschmidt [off-list ref] writes:
On Wed, 2017-02-01 at 09:05 -0800, Eric Dumazet wrote:
quoted
Hi all
Is anybody working on adding QUEUED spinlocks to powerpc 64bit ?
I've seen past attempts with ticket spinlocks
( https://patchwork.ozlabs.org/patch/449381/ and other related links
)
But it looks ticket spinlocks are a thing of the past.
From: Eric Dumazet <edumazet@google.com> Date: 2017-02-02 04:40:54
On Wed, Feb 1, 2017 at 8:04 PM, Michael Ellerman [off-list ref] wrote:
Benjamin Herrenschmidt [off-list ref] writes:
quoted
On Wed, 2017-02-01 at 09:05 -0800, Eric Dumazet wrote:
quoted
Hi all
Is anybody working on adding QUEUED spinlocks to powerpc 64bit ?
I've seen past attempts with ticket spinlocks
( https://patchwork.ozlabs.org/patch/449381/ and other related links
)
But it looks ticket spinlocks are a thing of the past.
Needs a good review, and the benchmark results were not all that
compelling - though perhaps they were just the wrong benchmarks.
A typical benchmark would be to use 200 concurrent netperf -t TCP_RR,
through a single qdisc (protected by a spinlock)
Non ticket/queued spinlocks behave quite bad in this scenario.
I can try this next week if you want.
From: Benjamin Herrenschmidt <benh@kernel.crashing.org> Date: 2017-02-02 04:42:30
On Wed, 2017-02-01 at 20:40 -0800, Eric Dumazet wrote:
A typical benchmark would be to use 200 concurrent netperf -t TCP_RR,
through a single qdisc (protected by a spinlock)
Non ticket/queued spinlocks behave quite bad in this scenario.
I can try this next week if you want.
On Wed, Feb 1, 2017 at 8:04 PM, Michael Ellerman [off-list ref] wrote:
quoted
Benjamin Herrenschmidt [off-list ref] writes:
quoted
On Wed, 2017-02-01 at 09:05 -0800, Eric Dumazet wrote:
quoted
Hi all
Is anybody working on adding QUEUED spinlocks to powerpc 64bit ?
I've seen past attempts with ticket spinlocks
( https://patchwork.ozlabs.org/patch/449381/ and other related links
)
But it looks ticket spinlocks are a thing of the past.
Needs a good review, and the benchmark results were not all that
compelling - though perhaps they were just the wrong benchmarks.
A typical benchmark would be to use 200 concurrent netperf -t TCP_RR,
through a single qdisc (protected by a spinlock)
hi all
I do some netperf tests and get some benchmark results.
I also attach my test script and netperf-result(Excel)
There are two machine. one runs netserver and the other runs netperf
benchmark. 1000Mbps network is connected with them.
#ip link infomation
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast
state UNKNOWN mode DEFAULT group default qlen 1000
link/ether ba:68:9c:14:32:02 brd ff:ff:ff:ff:ff:ff
According to the results, there is not much performance gap with each
other. And as we are only testing the throughput, the pvqspinlock shows
the overhead of its pv stuff. but qspinlock shows a little improvement
than spinlock. My simple summary in this testcase is
qspinlock > spinlock > pvqspinlock.
when run 200 concurrent netperf, I paste the total throughput here.
concurrent runners| total throughput | variance
-------------------------------------------
spinlock | 199 | 66882.8 | 89.93
-------------------------------------------
qspinlock | 199 | 66350.4 | 72.0239
-------------------------------------------
pvqspinlock | 199 | 64740.5 | 85.7837
You could see more data in nerperf.xlsx
thanks
xinhui
Non ticket/queued spinlocks behave quite bad in this scenario.
I can try this next week if you want.
From: Eric Dumazet <edumazet@google.com> Date: 2017-02-07 06:46:13
On Mon, Feb 6, 2017 at 10:21 PM, panxinhui [off-list ref] wrote:
hi all
I do some netperf tests and get some benchmark results.
I also attach my test script and netperf-result(Excel)
There are two machine. one runs netserver and the other runs netperf
benchmark. 1000Mbps network is connected with them.
#ip link infomation
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state
UNKNOWN mode DEFAULT group default qlen 1000
link/ether ba:68:9c:14:32:02 brd ff:ff:ff:ff:ff:ff
According to the results, there is not much performance gap with each other.
And as we are only testing the throughput, the pvqspinlock shows the
overhead of its pv stuff. but qspinlock shows a little improvement than
spinlock. My simple summary in this testcase is
qspinlock > spinlock > pvqspinlock.
when run 200 concurrent netperf, I paste the total throughput here.
concurrent runners| total throughput | variance
-------------------------------------------
spinlock | 199 | 66882.8 | 89.93
-------------------------------------------
qspinlock | 199 | 66350.4 | 72.0239
-------------------------------------------
pvqspinlock | 199 | 64740.5 | 85.7837
You could see more data in nerperf.xlsx
thanks
xinhui
Hi xinhui
1Gbit NIC is too slow for this use case. I would try a 10Gbit NIC at least...
Alternatively, you could use loopback interface. (netperf -H 127.0.0.1)
tc qd add dev lo root pfifo limit 10000
On Mon, Feb 6, 2017 at 10:21 PM, panxinhui [off-list ref] wrote:
quoted
hi all
I do some netperf tests and get some benchmark results.
I also attach my test script and netperf-result(Excel)
There are two machine. one runs netserver and the other runs netperf
benchmark. 1000Mbps network is connected with them.
#ip link infomation
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state
UNKNOWN mode DEFAULT group default qlen 1000
link/ether ba:68:9c:14:32:02 brd ff:ff:ff:ff:ff:ff
According to the results, there is not much performance gap with each other.
And as we are only testing the throughput, the pvqspinlock shows the
overhead of its pv stuff. but qspinlock shows a little improvement than
spinlock. My simple summary in this testcase is
qspinlock > spinlock > pvqspinlock.
when run 200 concurrent netperf, I paste the total throughput here.
concurrent runners| total throughput | variance
-------------------------------------------
spinlock | 199 | 66882.8 | 89.93
-------------------------------------------
qspinlock | 199 | 66350.4 | 72.0239
-------------------------------------------
pvqspinlock | 199 | 64740.5 | 85.7837
You could see more data in nerperf.xlsx
thanks
xinhui
Hi xinhui
1Gbit NIC is too slow for this use case. I would try a 10Gbit NIC at least...
Alternatively, you could use loopback interface. (netperf -H 127.0.0.1)
tc qd add dev lo root pfifo limit 10000
On Mon, Feb 6, 2017 at 10:21 PM, panxinhui [off-list ref] wrote:
quoted
hi all
I do some netperf tests and get some benchmark results.
I also attach my test script and netperf-result(Excel)
HI, all
I use loopback interface to run netperf tests,
#tc qd add dev lo root pfifo limit 10000
#ip link
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc pfifo state UNKNOWN mode DEFAULT group default qlen 1000
link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00
and put the result in netperf.xlsx(excel)
It is a 32 vcpus P8 machine, with 32Gib memory.
This time spinlock is the best one, qspinlock > pvqspinlock. So sad.
thanks
xinhui
quoted
There are two machine. one runs netserver and the other runs netperf
benchmark. 1000Mbps network is connected with them.
#ip link infomation
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state
UNKNOWN mode DEFAULT group default qlen 1000
link/ether ba:68:9c:14:32:02 brd ff:ff:ff:ff:ff:ff
According to the results, there is not much performance gap with each other.
And as we are only testing the throughput, the pvqspinlock shows the
overhead of its pv stuff. but qspinlock shows a little improvement than
spinlock. My simple summary in this testcase is
qspinlock > spinlock > pvqspinlock.
when run 200 concurrent netperf, I paste the total throughput here.
concurrent runners| total throughput | variance
-------------------------------------------
spinlock | 199 | 66882.8 | 89.93
-------------------------------------------
qspinlock | 199 | 66350.4 | 72.0239
-------------------------------------------
pvqspinlock | 199 | 64740.5 | 85.7837
You could see more data in nerperf.xlsx
thanks
xinhui
Hi xinhui
1Gbit NIC is too slow for this use case. I would try a 10Gbit NIC at least...
Alternatively, you could use loopback interface. (netperf -H 127.0.0.1)
tc qd add dev lo root pfifo limit 10000
On Mon, Feb 6, 2017 at 10:21 PM, panxinhui [off-list ref] wrote:
quoted
hi all
I do some netperf tests and get some benchmark results.
I also attach my test script and netperf-result(Excel)
HI, all
I use loopback interface to run netperf tests,
#tc qd add dev lo root pfifo limit 10000
#ip link
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc pfifo state UNKNOWN mode DEFAULT group default qlen 1000
link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00
and put the result in netperf.xlsx(excel)
It is a 32 vcpus P8 machine, with 32Gib memory.
This time spinlock is the best one, qspinlock > pvqspinlock. So sad.
This time, I have appiled some optimising patches on pvqspinlock.
When there is a high contention, the performance has a good improvement ans is very similar to spinlock.
Result is attached in netperf.xlsx
thanks
xinhui
thanks
xinhui
quoted
quoted
There are two machine. one runs netserver and the other runs netperf
benchmark. 1000Mbps network is connected with them.
#ip link infomation
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state
UNKNOWN mode DEFAULT group default qlen 1000
link/ether ba:68:9c:14:32:02 brd ff:ff:ff:ff:ff:ff
According to the results, there is not much performance gap with each other.
And as we are only testing the throughput, the pvqspinlock shows the
overhead of its pv stuff. but qspinlock shows a little improvement than
spinlock. My simple summary in this testcase is
qspinlock > spinlock > pvqspinlock.
when run 200 concurrent netperf, I paste the total throughput here.
concurrent runners| total throughput | variance
-------------------------------------------
spinlock | 199 | 66882.8 | 89.93
-------------------------------------------
qspinlock | 199 | 66350.4 | 72.0239
-------------------------------------------
pvqspinlock | 199 | 64740.5 | 85.7837
You could see more data in nerperf.xlsx
thanks
xinhui
Hi xinhui
1Gbit NIC is too slow for this use case. I would try a 10Gbit NIC at least...
Alternatively, you could use loopback interface. (netperf -H 127.0.0.1)
tc qd add dev lo root pfifo limit 10000