From: Mihai Maruseac <hidden> Date: 2011-10-07 15:21:20
Instead of using the dev->next chain and trying to resync at each call to
dev_seq_start, use this hash and store bucket number and bucket offset in
seq->private field.
Tests revealed the following results for ifconfig > /dev/null
* 1000 interfaces:
* 0.114s without patch
* 0.020s with patch
* 3000 interfaces:
* 0.489s without patch
* 0.048s with patch
* 5000 interfaces:
* 1.363s without patch
* 0.131s with patch
As one can notice the improvement is of 1 order of magnitude.
Signed-off-by: Mihai Maruseac <redacted>
---
net/core/dev.c | 73 +++++++++++++++++++++++++++++++++++++++++++-------------
1 files changed, 56 insertions(+), 17 deletions(-)
From: Stephen Hemminger <hidden> Date: 2011-10-07 16:24:52
On Fri, 7 Oct 2011 18:20:49 +0300
Mihai Maruseac [off-list ref] wrote:
Instead of using the dev->next chain and trying to resync at each call to
dev_seq_start, use this hash and store bucket number and bucket offset in
seq->private field.
Tests revealed the following results for ifconfig > /dev/null
* 1000 interfaces:
* 0.114s without patch
* 0.020s with patch
* 3000 interfaces:
* 0.489s without patch
* 0.048s with patch
* 5000 interfaces:
* 1.363s without patch
* 0.131s with patch
As one can notice the improvement is of 1 order of magnitude.
Good idea,
This will change the ordering of entries in /proc which may upset
some program, not a critical flaw but worth noting.
Rather than recording the bucket and offset of last entry, another
alternative would be to just record the ifindex.
Also ifconfig is considered deprecated and replaced by ip commands
for general use.
From: Mihai Maruseac <hidden> Date: 2011-10-08 07:22:49
On Fri, Oct 7, 2011 at 7:24 PM, Stephen Hemminger [off-list ref] wrote:
On Fri, 7 Oct 2011 18:20:49 +0300
Mihai Maruseac [off-list ref] wrote:
quoted
Instead of using the dev->next chain and trying to resync at each call to
dev_seq_start, use this hash and store bucket number and bucket offset in
seq->private field.
Tests revealed the following results for ifconfig > /dev/null
* 1000 interfaces:
* 0.114s without patch
* 0.020s with patch
* 3000 interfaces:
* 0.489s without patch
* 0.048s with patch
* 5000 interfaces:
* 1.363s without patch
* 0.131s with patch
As one can notice the improvement is of 1 order of magnitude.
Good idea,
This will change the ordering of entries in /proc which may upset
some program, not a critical flaw but worth noting.
Rather than recording the bucket and offset of last entry, another
alternative would be to just record the ifindex.
Also ifconfig is considered deprecated and replaced by ip commands
for general use.
Thanks,
This is a patch from a series of improvements to both the ifconfig and
ip commands. The ip part will come later, after being properly
implemented and tested.
--
Mihai
From: Mihai Maruseac <hidden> Date: 2011-10-10 08:43:41
On Fri, Oct 7, 2011 at 7:24 PM, Stephen Hemminger [off-list ref] wrote:
On Fri, 7 Oct 2011 18:20:49 +0300
Mihai Maruseac [off-list ref] wrote:
quoted
Instead of using the dev->next chain and trying to resync at each call to
dev_seq_start, use this hash and store bucket number and bucket offset in
seq->private field.
As one can notice the improvement is of 1 order of magnitude.
Good idea,
This will change the ordering of entries in /proc which may upset
some program, not a critical flaw but worth noting.
Rather than recording the bucket and offset of last entry, another
alternative would be to just record the ifindex.
I tried to record the ifindex but I think that using it and
dev_get_by_index can result in an infinite loop or a NULL
dereferrence. If a device is removed and ifindex points to it we'll
get a NULL from dev_get_by_index. Checking for NULL and calling again
dev_get_by_index will end in an infinite loop at the end of the hlist.
Augmenting the structure to also contain the number of indexes when
the seq_file is opened returns to the current situation with two ints.
Also, it is more prone to bugs caused by device removal while
printing.
--
Mihai
From: Stephen Hemminger <hidden> Date: 2011-10-11 05:55:41
On Mon, 10 Oct 2011 11:43:20 +0300
Mihai Maruseac [off-list ref] wrote:
On Fri, Oct 7, 2011 at 7:24 PM, Stephen Hemminger [off-list ref] wrote:
quoted
On Fri, 7 Oct 2011 18:20:49 +0300
Mihai Maruseac [off-list ref] wrote:
quoted
Instead of using the dev->next chain and trying to resync at each call to
dev_seq_start, use this hash and store bucket number and bucket offset in
seq->private field.
As one can notice the improvement is of 1 order of magnitude.
Good idea,
This will change the ordering of entries in /proc which may upset
some program, not a critical flaw but worth noting.
Rather than recording the bucket and offset of last entry, another
alternative would be to just record the ifindex.
I tried to record the ifindex but I think that using it and
dev_get_by_index can result in an infinite loop or a NULL
dereferrence. If a device is removed and ifindex points to it we'll
get a NULL from dev_get_by_index. Checking for NULL and calling again
dev_get_by_index will end in an infinite loop at the end of the hlist.
Augmenting the structure to also contain the number of indexes when
the seq_file is opened returns to the current situation with two ints.
Also, it is more prone to bugs caused by device removal while
printing.
If ifindex has been deleted, the code should fall back to delivering
the next offset. That means for the rare case it would have the old
behavior of linear searching. There is similar code already to
deal with /proc/net/route; it continues from the last address.
From: Mihai Maruseac <hidden> Date: 2011-10-12 09:49:30
Instead of using the dev->next chain and trying to resync at each call to
dev_seq_start, use the ifindex, keeping the last index in seq->private field.
Tests revealed the following results for ifconfig > /dev/null
* 1000 interfaces:
* 0.114s without patch
* 0.089s with patch
* 3000 interfaces:
* 0.489s without patch
* 0.110s with patch
* 5000 interfaces:
* 1.363s without patch
* 0.250s with patch
* 128000 interfaces (other setup):
* ~100s without patch
* ~30s with patch
---
net/core/dev.c | 53 ++++++++++++++++++++++++++++++++---------------------
1 files changed, 32 insertions(+), 21 deletions(-)
From: Mihai Maruseac <hidden> Date: 2011-10-12 09:57:41
Instead of using the dev->next chain and trying to resync at each call to
dev_seq_start, use the ifindex, keeping the last index in seq->private field.
Tests revealed the following results for ifconfig > /dev/null
* 1000 interfaces:
* 0.114s without patch
* 0.089s with patch
* 3000 interfaces:
* 0.489s without patch
* 0.110s with patch
* 5000 interfaces:
* 1.363s without patch
* 0.250s with patch
* 128000 interfaces (other setup):
* ~100s without patch
* ~30s with patch
Signed-off-by: Mihai Maruseac <redacted>
---
net/core/dev.c | 53 ++++++++++++++++++++++++++++++++---------------------
1 files changed, 32 insertions(+), 21 deletions(-)
From: Mihai Maruseac <hidden> Date: 2011-10-12 10:04:37
On Wed, 2011-10-12 at 02:57 -0700, Mihai Maruseac wrote:
Instead of using the dev->next chain and trying to resync at each call to
dev_seq_start, use the ifindex, keeping the last index in seq->private field.
Please consider only my last patch email, the first one was missing
Signed-off-by line.
Thanks,
Mihai
From: Mihai Maruseac <hidden> Date: 2011-10-12 10:08:44
On Mon, 2011-10-10 at 22:55 -0700, Stephen Hemminger wrote:
If ifindex has been deleted, the code should fall back to delivering
the next offset. That means for the rare case it would have the old
behavior of linear searching. There is similar code already to
deal with /proc/net/route; it continues from the last address.
We rewrote the patch to use the ifindex. Since we are at the beginning,
the patch was sent in another thread. We'll learn how to properly send
the mail.
--
Thanks,
Mihai
From: Stephen Hemminger <hidden> Date: 2011-10-12 15:50:29
On Wed, 12 Oct 2011 12:57:26 +0300
Mihai Maruseac [off-list ref] wrote:
It looks like you lost the ability to seek back and read the header
(start token). How is the header handled, is possible to rewind
the file and read it over again?
Looks good a couple of minor nits.
1. The function should not be inline, since it is in no way performance
critical. The compiler will probably inline it anyway.
2. dev does not have to be initialized since it is assigned a few
lines later. Most programmers are trained now to always initialize
variables, but often it is unnecessary.
3. The name next_dev() is a little generic; maybe a better name.
From: Mihai Maruseac <hidden> Date: 2011-10-14 09:53:52
Instead of using the dev->next chain and trying to resync at each call to
dev_seq_start, use the ifindex, keeping the last index in seq->private field.
Tests revealed the following results for ifconfig > /dev/null
* 1000 interfaces:
* 0.114s without patch
* 0.089s with patch
* 3000 interfaces:
* 0.489s without patch
* 0.110s with patch
* 5000 interfaces:
* 1.363s without patch
* 0.250s with patch
* 128000 interfaces (other setup):
* ~100s without patch
* ~30s with patch
Signed-off-by: Mihai Maruseac <redacted>
---
net/core/dev.c | 55 ++++++++++++++++++++++++++++++++++---------------------
1 files changed, 34 insertions(+), 21 deletions(-)
From: Mihai Maruseac <hidden> Date: 2011-10-14 12:21:03
On Wed, Oct 12, 2011 at 6:50 PM, Stephen Hemminger
[off-list ref] wrote:
On Wed, 12 Oct 2011 12:57:26 +0300
Mihai Maruseac [off-list ref] wrote:
It looks like you lost the ability to seek back and read the header
(start token). How is the header handled, is possible to rewind
the file and read it over again?
We tested with a simple program reading a part of the /proc file and
then doing a seek to the start and rereading. It worked with latest
submitted patch.
Looks good a couple of minor nits.
1. The function should not be inline, since it is in no way performance
critical. The compiler will probably inline it anyway.
2. dev does not have to be initialized since it is assigned a few
lines later. Most programmers are trained now to always initialize
variables, but often it is unnecessary.
3. The name next_dev() is a little generic; maybe a better name.
From: Eric Dumazet <hidden> Date: 2011-10-14 12:52:58
Le vendredi 14 octobre 2011 à 12:53 +0300, Mihai Maruseac a écrit :
quoted hunk
Instead of using the dev->next chain and trying to resync at each call to
dev_seq_start, use the ifindex, keeping the last index in seq->private field.
Tests revealed the following results for ifconfig > /dev/null
* 1000 interfaces:
* 0.114s without patch
* 0.089s with patch
* 3000 interfaces:
* 0.489s without patch
* 0.110s with patch
* 5000 interfaces:
* 1.363s without patch
* 0.250s with patch
* 128000 interfaces (other setup):
* ~100s without patch
* ~30s with patch
Signed-off-by: Mihai Maruseac <redacted>
---
net/core/dev.c | 55 ++++++++++++++++++++++++++++++++++---------------------
1 files changed, 34 insertions(+), 21 deletions(-)
@@ -4041,6 +4041,37 @@ static int dev_ifconf(struct net *net, char __user *arg)}#ifdef CONFIG_PROC_FS++structdev_iter_state{+structseq_net_privatep;+intifindex;+};++staticstructnet_device*__dev_seq_next(structseq_file*seq,loff_t*pos)+{+structdev_iter_state*state=seq->private;+structnet*net=seq_file_net(seq);+structnet_device*dev;+loff_toff;++dev=dev_get_by_index_rcu(net,state->ifindex);+if(likely(dev))+gotofound;++off=0;+for_each_netdev_rcu(net,dev)+if(off++==*pos){+state->ifindex=dev->ifindex;+gotofound;+}++returnNULL;+found:+state->ifindex++;
This assumes device ifindexes are contained in a small range
[N .. N + X]
I understand this can help some benchmarks, but in real world this wont
help that much once ifindexes are 'fragmented' (If really this multi
thousand devices stuff is for real)
Listen, we currently have 256 slots in the hash table.
Can we try to make 'offset' something like (slot_number<<24) +
(position in hash chain [slot_number]), instead of (position in devices
global list)
From: Daniel Baluta <hidden> Date: 2011-10-17 08:03:57
This assumes device ifindexes are contained in a small range
[N .. N + X]
I understand this can help some benchmarks, but in real world this wont
help that much once ifindexes are 'fragmented' (If really this multi
thousand devices stuff is for real)
Listen, we currently have 256 slots in the hash table.
Can we try to make 'offset' something like (slot_number<<24) +
(position in hash chain [slot_number]), instead of (position in devices
global list)
Eric, we can refine the idea of our first patch [1], where we recorded
the (bucket, offset) pair. Stephen, do you agree with this?
thanks,
Daniel.
[1] http://patchwork.ozlabs.org/patch/118331/
From: Stephen Hemminger <hidden> Date: 2011-10-17 15:12:12
On Mon, 17 Oct 2011 11:03:54 +0300
Daniel Baluta [off-list ref] wrote:
quoted
This assumes device ifindexes are contained in a small range
[N .. N + X]
I understand this can help some benchmarks, but in real world this wont
help that much once ifindexes are 'fragmented' (If really this multi
thousand devices stuff is for real)
Listen, we currently have 256 slots in the hash table.
Can we try to make 'offset' something like (slot_number<<24) +
(position in hash chain [slot_number]), instead of (position in devices
global list)
Eric, we can refine the idea of our first patch [1], where we recorded
the (bucket, offset) pair. Stephen, do you agree with this?
thanks,
Daniel.
[1] http://patchwork.ozlabs.org/patch/118331/
Using buckets is fine, my idea about ifindex was just to try and
preserve the order, but it doesn't matter.