[RFC] implement QUEUED spinlocks on powerpc

10 messages, 5 authors, 2017-02-15 · open the first message on its own page

[RFC] implement QUEUED spinlocks on powerpc

From: Eric Dumazet <hidden>
Date: 2017-02-01 17:05:37

Hi all

Is anybody working on adding QUEUED spinlocks to powerpc 64bit ?

I've seen past attempts with ticket spinlocks
( https://patchwork.ozlabs.org/patch/449381/ and other related links )

But it looks ticket spinlocks are a thing of the past.

Thanks.

Re: [RFC] implement QUEUED spinlocks on powerpc

From: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Date: 2017-02-01 20:37:44

On Wed, 2017-02-01 at 09:05 -0800, Eric Dumazet wrote:
Hi all

Is anybody working on adding QUEUED spinlocks to powerpc 64bit ?

I've seen past attempts with ticket spinlocks
( https://patchwork.ozlabs.org/patch/449381/ and other related links
)

But it looks ticket spinlocks are a thing of the past.
Yes, we have a tentative implementation of qspinlock and pv variants:

https://patchwork.ozlabs.org/patch/703139/
https://patchwork.ozlabs.org/patch/703140/
https://patchwork.ozlabs.org/patch/703141/
https://patchwork.ozlabs.org/patch/703142/
https://patchwork.ozlabs.org/patch/703143/
https://patchwork.ozlabs.org/patch/703144/

Michael, what's the status with getting that merged ?

Cheers,
Ben.

Re: [RFC] implement QUEUED spinlocks on powerpc

From: Michael Ellerman <mpe@ellerman.id.au>
Date: 2017-02-02 04:05:00

Benjamin Herrenschmidt [off-list ref] writes:
On Wed, 2017-02-01 at 09:05 -0800, Eric Dumazet wrote:
quoted
Hi all

Is anybody working on adding QUEUED spinlocks to powerpc 64bit ?

I've seen past attempts with ticket spinlocks
( https://patchwork.ozlabs.org/patch/449381/ and other related links
)

But it looks ticket spinlocks are a thing of the past.
Yes, we have a tentative implementation of qspinlock and pv variants:

https://patchwork.ozlabs.org/patch/703139/
https://patchwork.ozlabs.org/patch/703140/
https://patchwork.ozlabs.org/patch/703141/
https://patchwork.ozlabs.org/patch/703142/
https://patchwork.ozlabs.org/patch/703143/
https://patchwork.ozlabs.org/patch/703144/

Michael, what's the status with getting that merged ?
Needs a good review, and the benchmark results were not all that
compelling - though perhaps they were just the wrong benchmarks.

cheers

Re: [RFC] implement QUEUED spinlocks on powerpc

From: Eric Dumazet <edumazet@google.com>
Date: 2017-02-02 04:40:54

On Wed, Feb 1, 2017 at 8:04 PM, Michael Ellerman [off-list ref] wrote:
Benjamin Herrenschmidt [off-list ref] writes:
quoted
On Wed, 2017-02-01 at 09:05 -0800, Eric Dumazet wrote:
quoted
Hi all

Is anybody working on adding QUEUED spinlocks to powerpc 64bit ?

I've seen past attempts with ticket spinlocks
( https://patchwork.ozlabs.org/patch/449381/ and other related links
)

But it looks ticket spinlocks are a thing of the past.
Yes, we have a tentative implementation of qspinlock and pv variants:

https://patchwork.ozlabs.org/patch/703139/
https://patchwork.ozlabs.org/patch/703140/
https://patchwork.ozlabs.org/patch/703141/
https://patchwork.ozlabs.org/patch/703142/
https://patchwork.ozlabs.org/patch/703143/
https://patchwork.ozlabs.org/patch/703144/

Michael, what's the status with getting that merged ?
Needs a good review, and the benchmark results were not all that
compelling - though perhaps they were just the wrong benchmarks.
A typical benchmark would be to use 200 concurrent netperf -t TCP_RR,
through a single qdisc (protected by a spinlock)

Non ticket/queued spinlocks behave quite bad in this scenario.

I can try this next week if you want.

Re: [RFC] implement QUEUED spinlocks on powerpc

From: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Date: 2017-02-02 04:42:30

On Wed, 2017-02-01 at 20:40 -0800, Eric Dumazet wrote:
A typical benchmark would be to use 200 concurrent netperf -t TCP_RR,
through a single qdisc (protected by a spinlock)

Non ticket/queued spinlocks behave quite bad in this scenario.

I can try this next week if you want.
That would be great !

Cheers,
Ben.

Re: [RFC] implement QUEUED spinlocks on powerpc

From: panxinhui <hidden>
Date: 2017-02-07 06:21:43

在 2017/2/2 下午12:40, Eric Dumazet 写道:
On Wed, Feb 1, 2017 at 8:04 PM, Michael Ellerman [off-list ref] wrote:
quoted
Benjamin Herrenschmidt [off-list ref] writes:
quoted
On Wed, 2017-02-01 at 09:05 -0800, Eric Dumazet wrote:
quoted
Hi all

Is anybody working on adding QUEUED spinlocks to powerpc 64bit ?

I've seen past attempts with ticket spinlocks
( https://patchwork.ozlabs.org/patch/449381/ and other related links
)

But it looks ticket spinlocks are a thing of the past.
Yes, we have a tentative implementation of qspinlock and pv variants:

https://patchwork.ozlabs.org/patch/703139/
https://patchwork.ozlabs.org/patch/703140/
https://patchwork.ozlabs.org/patch/703141/
https://patchwork.ozlabs.org/patch/703142/
https://patchwork.ozlabs.org/patch/703143/
https://patchwork.ozlabs.org/patch/703144/

Michael, what's the status with getting that merged ?
Needs a good review, and the benchmark results were not all that
compelling - though perhaps they were just the wrong benchmarks.
A typical benchmark would be to use 200 concurrent netperf -t TCP_RR,
through a single qdisc (protected by a spinlock)
hi all
	I do some netperf tests and get some benchmark results.
I also attach my test script and netperf-result(Excel)

There are two machine. one runs netserver and the other runs netperf 
benchmark. 1000Mbps network is connected with them.

#ip link infomation
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast 
state UNKNOWN mode DEFAULT group default qlen 1000
      link/ether ba:68:9c:14:32:02 brd ff:ff:ff:ff:ff:ff

According to the results, there is not much performance gap with each 
other. And as we are only testing the throughput, the pvqspinlock shows 
the overhead of its pv stuff. but qspinlock shows a little improvement 
than spinlock. My simple summary in this testcase is
qspinlock > spinlock > pvqspinlock.

when run 200 concurrent netperf, I paste the total throughput here.

	concurrent runners| total throughput | variance
-------------------------------------------
spinlock	| 199 | 66882.8 | 89.93
-------------------------------------------
qspinlock 	| 199 | 66350.4 | 72.0239
-------------------------------------------
pvqspinlock	| 199 | 64740.5 | 85.7837

You could see more data in nerperf.xlsx

thanks
xinhui
Non ticket/queued spinlocks behave quite bad in this scenario.

I can try this next week if you want.

Re: [RFC] implement QUEUED spinlocks on powerpc

From: Eric Dumazet <edumazet@google.com>
Date: 2017-02-07 06:46:13

On Mon, Feb 6, 2017 at 10:21 PM, panxinhui [off-list ref] wrote:
hi all
        I do some netperf tests and get some benchmark results.
I also attach my test script and netperf-result(Excel)

There are two machine. one runs netserver and the other runs netperf
benchmark. 1000Mbps network is connected with them.

#ip link infomation
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state
UNKNOWN mode DEFAULT group default qlen 1000
     link/ether ba:68:9c:14:32:02 brd ff:ff:ff:ff:ff:ff

According to the results, there is not much performance gap with each other.
And as we are only testing the throughput, the pvqspinlock shows the
overhead of its pv stuff. but qspinlock shows a little improvement than
spinlock. My simple summary in this testcase is
qspinlock > spinlock > pvqspinlock.

when run 200 concurrent netperf, I paste the total throughput here.

        concurrent runners| total throughput | variance
-------------------------------------------
spinlock        | 199 | 66882.8 | 89.93
-------------------------------------------
qspinlock       | 199 | 66350.4 | 72.0239
-------------------------------------------
pvqspinlock     | 199 | 64740.5 | 85.7837

You could see more data in nerperf.xlsx

thanks
xinhui

Hi xinhui

1Gbit NIC is too slow for this use case. I would try a 10Gbit NIC at least...

Alternatively, you could use loopback interface.  (netperf -H 127.0.0.1)

tc qd add dev lo root pfifo limit 10000

Re: [RFC] implement QUEUED spinlocks on powerpc

From: panxinhui <hidden>
Date: 2017-02-07 07:22:31


在 2017/2/7 下午2:46, Eric Dumazet 写道:
On Mon, Feb 6, 2017 at 10:21 PM, panxinhui [off-list ref] wrote:
quoted
hi all
        I do some netperf tests and get some benchmark results.
I also attach my test script and netperf-result(Excel)

There are two machine. one runs netserver and the other runs netperf
benchmark. 1000Mbps network is connected with them.

#ip link infomation
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state
UNKNOWN mode DEFAULT group default qlen 1000
     link/ether ba:68:9c:14:32:02 brd ff:ff:ff:ff:ff:ff

According to the results, there is not much performance gap with each other.
And as we are only testing the throughput, the pvqspinlock shows the
overhead of its pv stuff. but qspinlock shows a little improvement than
spinlock. My simple summary in this testcase is
qspinlock > spinlock > pvqspinlock.

when run 200 concurrent netperf, I paste the total throughput here.

        concurrent runners| total throughput | variance
-------------------------------------------
spinlock        | 199 | 66882.8 | 89.93
-------------------------------------------
qspinlock       | 199 | 66350.4 | 72.0239
-------------------------------------------
pvqspinlock     | 199 | 64740.5 | 85.7837

You could see more data in nerperf.xlsx

thanks
xinhui

Hi xinhui

1Gbit NIC is too slow for this use case. I would try a 10Gbit NIC at least...

Alternatively, you could use loopback interface.  (netperf -H 127.0.0.1)

tc qd add dev lo root pfifo limit 10000
great, thanks
xinhui

Re: [RFC] implement QUEUED spinlocks on powerpc

From: panxinhui <hidden>
Date: 2017-02-13 09:09:10


在 2017/2/7 下午2:46, Eric Dumazet 写道:
On Mon, Feb 6, 2017 at 10:21 PM, panxinhui [off-list ref] wrote:
quoted
hi all
        I do some netperf tests and get some benchmark results.
I also attach my test script and netperf-result(Excel)
HI, all
I use loopback interface to run netperf tests,
#tc qd add dev lo root pfifo limit 10000
#ip link
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc pfifo state UNKNOWN mode DEFAULT group default qlen 1000
    link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00

and put the result in netperf.xlsx(excel)

It is a 32 vcpus P8 machine, with 32Gib memory.

This time spinlock is the best one, qspinlock > pvqspinlock. So sad.

thanks
xinhui
quoted
There are two machine. one runs netserver and the other runs netperf
benchmark. 1000Mbps network is connected with them.

#ip link infomation
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state
UNKNOWN mode DEFAULT group default qlen 1000
     link/ether ba:68:9c:14:32:02 brd ff:ff:ff:ff:ff:ff

According to the results, there is not much performance gap with each other.
And as we are only testing the throughput, the pvqspinlock shows the
overhead of its pv stuff. but qspinlock shows a little improvement than
spinlock. My simple summary in this testcase is
qspinlock > spinlock > pvqspinlock.

when run 200 concurrent netperf, I paste the total throughput here.

        concurrent runners| total throughput | variance
-------------------------------------------
spinlock        | 199 | 66882.8 | 89.93
-------------------------------------------
qspinlock       | 199 | 66350.4 | 72.0239
-------------------------------------------
pvqspinlock     | 199 | 64740.5 | 85.7837

You could see more data in nerperf.xlsx

thanks
xinhui

Hi xinhui

1Gbit NIC is too slow for this use case. I would try a 10Gbit NIC at least...

Alternatively, you could use loopback interface.  (netperf -H 127.0.0.1)

tc qd add dev lo root pfifo limit 10000

Re: [RFC] implement QUEUED spinlocks on powerpc

From: panxinhui <hidden>
Date: 2017-02-15 10:17:39


在 2017/2/13 下午5:08, panxinhui 写道:

在 2017/2/7 下午2:46, Eric Dumazet 写道:
quoted
On Mon, Feb 6, 2017 at 10:21 PM, panxinhui [off-list ref] wrote:
quoted
hi all
        I do some netperf tests and get some benchmark results.
I also attach my test script and netperf-result(Excel)
HI, all
I use loopback interface to run netperf tests,
#tc qd add dev lo root pfifo limit 10000
#ip link
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc pfifo state UNKNOWN mode DEFAULT group default qlen 1000
    link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00

and put the result in netperf.xlsx(excel)

It is a 32 vcpus P8 machine, with 32Gib memory.

This time spinlock is the best one, qspinlock > pvqspinlock. So sad.
This time, I have appiled some optimising patches on pvqspinlock.
When there is a high contention, the performance has a good improvement ans is very similar to spinlock.

Result is attached in netperf.xlsx

thanks
xinhui
thanks
xinhui
quoted
quoted
There are two machine. one runs netserver and the other runs netperf
benchmark. 1000Mbps network is connected with them.

#ip link infomation
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state
UNKNOWN mode DEFAULT group default qlen 1000
     link/ether ba:68:9c:14:32:02 brd ff:ff:ff:ff:ff:ff

According to the results, there is not much performance gap with each other.
And as we are only testing the throughput, the pvqspinlock shows the
overhead of its pv stuff. but qspinlock shows a little improvement than
spinlock. My simple summary in this testcase is
qspinlock > spinlock > pvqspinlock.

when run 200 concurrent netperf, I paste the total throughput here.

        concurrent runners| total throughput | variance
-------------------------------------------
spinlock        | 199 | 66882.8 | 89.93
-------------------------------------------
qspinlock       | 199 | 66350.4 | 72.0239
-------------------------------------------
pvqspinlock     | 199 | 64740.5 | 85.7837

You could see more data in nerperf.xlsx

thanks
xinhui

Hi xinhui

1Gbit NIC is too slow for this use case. I would try a 10Gbit NIC at least...

Alternatively, you could use loopback interface.  (netperf -H 127.0.0.1)

tc qd add dev lo root pfifo limit 10000
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help