[RFC] Introduce to batch variants of accept() and epoll_ctl() syscall

7 messages, 4 authors, 2012-07-09 · open the first message on its own page

[RFC] Introduce to batch variants of accept() and epoll_ctl() syscall

From: Li Yu <hidden>
Date: 2012-06-15 04:13:20

Hi,

  We encounter a performance problem in a large scale computer
cluster, which needs to handle a lot of incoming concurrent TCP
connection requests.

  The top shows the kernel is most cpu hog, the testing is simple,
just a accept() -> epoll_ctl(ADD) loop, the ratio of cpu util sys% to
si% is about 2:5.

  I also asked some experienced webserver/proxy developers in my team
for suggestions, it seem that behavior of many userland programs already
called accept() multiple times after it is waked up by
epoll_wait(). And the common action is adding the fd that accept()
return into epoll interface by epoll_ctl() syscall then.

  Therefore, I think that we'd better to introduce to batch variants of
accept() and epoll_ctl() syscall, just like sendmmsg() or recvmmsg().

  For accept(), we may need a new syscall, it may like this,

  struct accept_result {
      int fd;
      struct sockaddr addr;
      socklen_t addr_len;
  };

  int maccept4(int fd, int flags, int nr_accept_result, struct
accept_result *results);

  For epoll_ctl(), there are two means to extend it, I prefer to extend
current interface instead of introduce to new syscall. We may introduce
to a new flag EPOLL_CTL_BATCH. If userland call epoll_ctl() with this
flag set, the meaning of last two arguments of epoll_ctl() change, .e.g:

  struct batch_epoll_event batch_event[] = {
         {
              .fd = a_newsock_fd;
              .epoll_event = { ... };
         },
         ...
  };

  ret = epoll_ctl(fd, EPOLL_CTL_ADD|EPOLL_CTL_BATCH, nr_batch_events,
batch_events);

  Thanks.

Yu

Re: [RFC] Introduce to batch variants of accept() and epoll_ctl() syscall

From: Changli Gao <hidden>
Date: 2012-06-15 04:30:09

On Fri, Jun 15, 2012 at 12:13 PM, Li Yu [off-list ref] wrote:
Hi,

 We encounter a performance problem in a large scale computer
cluster, which needs to handle a lot of incoming concurrent TCP
connection requests.

 The top shows the kernel is most cpu hog, the testing is simple,
just a accept() -> epoll_ctl(ADD) loop, the ratio of cpu util sys% to
si% is about 2:5.

 I also asked some experienced webserver/proxy developers in my team
for suggestions, it seem that behavior of many userland programs already
called accept() multiple times after it is waked up by
epoll_wait(). And the common action is adding the fd that accept()
return into epoll interface by epoll_ctl() syscall then.

 Therefore, I think that we'd better to introduce to batch variants of
accept() and epoll_ctl() syscall, just like sendmmsg() or recvmmsg().

 For accept(), we may need a new syscall, it may like this,

 struct accept_result {
     int fd;
     struct sockaddr addr;
     socklen_t addr_len;
 };

 int maccept4(int fd, int flags, int nr_accept_result, struct
accept_result *results);

 For epoll_ctl(), there are two means to extend it, I prefer to extend
current interface instead of introduce to new syscall. We may introduce
to a new flag EPOLL_CTL_BATCH. If userland call epoll_ctl() with this
flag set, the meaning of last two arguments of epoll_ctl() change, .e.g:

 struct batch_epoll_event batch_event[] = {
        {
             .fd = a_newsock_fd;
             .epoll_event = { ... };
        },
        ...
 };

 ret = epoll_ctl(fd, EPOLL_CTL_ADD|EPOLL_CTL_BATCH, nr_batch_events,
batch_events);
I think it is good idea. Would you please implement a prototype and
give some numbers? This kind of data may help selling this idea.
Thanks.

-- 
Regards,
Changli Gao(xiaosuo@gmail.com)

Re: [RFC] Introduce to batch variants of accept() and epoll_ctl() syscall

From: Li Yu <hidden>
Date: 2012-06-15 05:37:49

于 2012年06月15日 12:29, Changli Gao 写道:
On Fri, Jun 15, 2012 at 12:13 PM, Li Yu[off-list ref]  wrote:
quoted
Hi,

  We encounter a performance problem in a large scale computer
cluster, which needs to handle a lot of incoming concurrent TCP
connection requests.

  The top shows the kernel is most cpu hog, the testing is simple,
just a accept() ->  epoll_ctl(ADD) loop, the ratio of cpu util sys% to
si% is about 2:5.

  I also asked some experienced webserver/proxy developers in my team
for suggestions, it seem that behavior of many userland programs already
called accept() multiple times after it is waked up by
epoll_wait(). And the common action is adding the fd that accept()
return into epoll interface by epoll_ctl() syscall then.

  Therefore, I think that we'd better to introduce to batch variants of
accept() and epoll_ctl() syscall, just like sendmmsg() or recvmmsg().

  For accept(), we may need a new syscall, it may like this,

  struct accept_result {
      int fd;
      struct sockaddr addr;
      socklen_t addr_len;
  };

  int maccept4(int fd, int flags, int nr_accept_result, struct
accept_result *results);

  For epoll_ctl(), there are two means to extend it, I prefer to extend
current interface instead of introduce to new syscall. We may introduce
to a new flag EPOLL_CTL_BATCH. If userland call epoll_ctl() with this
flag set, the meaning of last two arguments of epoll_ctl() change, .e.g:

  struct batch_epoll_event batch_event[] = {
         {
              .fd = a_newsock_fd;
              .epoll_event = { ... };
         },
         ...
  };

  ret = epoll_ctl(fd, EPOLL_CTL_ADD|EPOLL_CTL_BATCH, nr_batch_events,
batch_events);
I think it is good idea. Would you please implement a prototype and
give some numbers? This kind of data may help selling this idea.
Thanks.
Of course, I think that implementing them should not be a hard work :)

Em. I really do not know whether it is necessary to introduce to a new 
syscall here. An alternative solution to add new socket option to handle 
such batch requirement, so applications also can detect if kernel has 
this extended ability with a easy getsockopt() call.

Any way, I am going to try to write a prototype first.

Thanks

Yu

RE: [RFC] Introduce to batch variants of accept() and epoll_ctl() syscall

From: David Laight <hidden>
Date: 2012-06-15 08:36:39

 
  We encounter a performance problem in a large scale computer
cluster, which needs to handle a lot of incoming concurrent TCP
connection requests.

  The top shows the kernel is most cpu hog, the testing is simple,
just a accept() -> epoll_ctl(ADD) loop, the ratio of cpu util sys% to
si% is about 2:5.

  I also asked some experienced webserver/proxy developers in my team
for suggestions, it seem that behavior of many userland 
programs already
called accept() multiple times after it is waked up by
epoll_wait(). And the common action is adding the fd that accept()
return into epoll interface by epoll_ctl() syscall then.

  Therefore, I think that we'd better to introduce to batch 
variants of
accept() and epoll_ctl() syscall, just like sendmmsg() or recvmmsg().
...

Having seen the support added to NetBSD for sendmmsg() and
recvmmsg() (and I'm told the linux code is much the same),
I'm surprised that just cutting out a system call entry/exit
and fd lookup is significant above the rest of the costs
involved in sending a message (which I presume is UDP here).
I'd be even more surprised if it is significant for an
incoming connection.

	David

Re: [RFC] Introduce to batch variants of accept() and epoll_ctl() syscall

From: Eric Dumazet <hidden>
Date: 2012-06-15 08:52:12

On Fri, 2012-06-15 at 13:37 +0800, Li Yu wrote:
Of course, I think that implementing them should not be a hard work :)

Em. I really do not know whether it is necessary to introduce to a new 
syscall here. An alternative solution to add new socket option to handle 
such batch requirement, so applications also can detect if kernel has 
this extended ability with a easy getsockopt() call.

Any way, I am going to try to write a prototype first.
Before that, could you post the result of "perf top", or "perf
record ...;perf report"
 The top shows the kernel is most cpu hog, the testing is simple,
just a accept() -> epoll_ctl(ADD) loop, the ratio of cpu util sys% to
si% is about 2:5.
This ratio is not meaningful, if we dont know where time is spent.


I doubt epoll_ctl(ADD) is a problem here...

If it is, batching the fds wont speed the thing anyway...

I believe accept() is the problem here, because it contends with the
softirq processing the tcp session handshake.

Re: [RFC] Introduce to batch variants of accept() and epoll_ctl() syscall

From: Li Yu <hidden>
Date: 2012-07-06 09:38:46

于 2012年06月15日 16:51, Eric Dumazet 写道:
On Fri, 2012-06-15 at 13:37 +0800, Li Yu wrote:
quoted
Of course, I think that implementing them should not be a hard work :)

Em. I really do not know whether it is necessary to introduce to a new
syscall here. An alternative solution to add new socket option to handle
such batch requirement, so applications also can detect if kernel has
this extended ability with a easy getsockopt() call.

Any way, I am going to try to write a prototype first.
Before that, could you post the result of "perf top", or "perf
record ...;perf report"
Sorry for I just have time to write a benchmark to reproduce this
problem on my test bed, below are results of "perf record -g -C 0".
kernel is 3.4.0:

Events: 7K cycles
+  54.87%  swapper  [kernel.kallsyms]  [k] poll_idle
-   3.10%   :22984  [kernel.kallsyms]  [k] _raw_spin_lock
    - _raw_spin_lock
       - 64.62% sch_direct_xmit
            dev_queue_xmit
            ip_finish_output
            ip_output
          - ip_local_out
             + 49.48% ip_queue_xmit
             + 37.48% ip_build_and_send_pkt
             + 13.04% ip_send_skb

I can not reproduce complete same high CPU usage on my testing 
environment, but top show that it has similar ratio of sys% and
si% on one CPU:

Tasks: 125 total,   2 running, 123 sleeping,   0 stopped,   0 zombie
Cpu0  :  1.0%us, 30.7%sy,  0.0%ni, 18.8%id,  0.0%wa,  0.0%hi, 49.5%si, 
0.0%st

Well, it seem that I must acknowledge I was wrong here. however,
I recall that I indeed ever encountered this in another benchmarking a
small packets performance.

I guess, this is since TX softirq and syscall context contend same lock
in sch_direct_xmit(), is this right?

thanks

Yu
quoted
  The top shows the kernel is most cpu hog, the testing is simple,
just a accept() -> epoll_ctl(ADD) loop, the ratio of cpu util sys% to
si% is about 2:5.
This ratio is not meaningful, if we dont know where time is spent.


I doubt epoll_ctl(ADD) is a problem here...

If it is, batching the fds wont speed the thing anyway...

I believe accept() is the problem here, because it contends with the
softirq processing the tcp session handshake.


Re: [RFC] Introduce to batch variants of accept() and epoll_ctl() syscall

From: Li Yu <hidden>
Date: 2012-07-09 03:36:55

于 2012年07月06日 17:38, Li Yu 写道:
于 2012年06月15日 16:51, Eric Dumazet 写道:
quoted
On Fri, 2012-06-15 at 13:37 +0800, Li Yu wrote:
quoted
Of course, I think that implementing them should not be a hard work :)

Em. I really do not know whether it is necessary to introduce to a new
syscall here. An alternative solution to add new socket option to handle
such batch requirement, so applications also can detect if kernel has
this extended ability with a easy getsockopt() call.

Any way, I am going to try to write a prototype first.
Before that, could you post the result of "perf top", or "perf
record ...;perf report"
Sorry for I just have time to write a benchmark to reproduce this
problem on my test bed, below are results of "perf record -g -C 0".
kernel is 3.4.0:

Events: 7K cycles
+  54.87%  swapper  [kernel.kallsyms]  [k] poll_idle
-   3.10%   :22984  [kernel.kallsyms]  [k] _raw_spin_lock
    - _raw_spin_lock
       - 64.62% sch_direct_xmit
            dev_queue_xmit
            ip_finish_output
            ip_output
          - ip_local_out
             + 49.48% ip_queue_xmit
             + 37.48% ip_build_and_send_pkt
             + 13.04% ip_send_skb

I can not reproduce complete same high CPU usage on my testing
environment, but top show that it has similar ratio of sys% and
si% on one CPU:

Tasks: 125 total,   2 running, 123 sleeping,   0 stopped,   0 zombie
Cpu0  :  1.0%us, 30.7%sy,  0.0%ni, 18.8%id,  0.0%wa,  0.0%hi, 49.5%si,
0.0%st

Well, it seem that I must acknowledge I was wrong here. however,
I recall that I indeed ever encountered this in another benchmarking a
small packets performance.

I guess, this is since TX softirq and syscall context contend same lock
in sch_direct_xmit(), is this right?
Em, do we have some means to decrease the lock contention here?
thanks

Yu
quoted
quoted
  The top shows the kernel is most cpu hog, the testing is simple,
just a accept() -> epoll_ctl(ADD) loop, the ratio of cpu util sys% to
si% is about 2:5.
This ratio is not meaningful, if we dont know where time is spent.


I doubt epoll_ctl(ADD) is a problem here...

If it is, batching the fds wont speed the thing anyway...

I believe accept() is the problem here, because it contends with the
softirq processing the tcp session handshake.


Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help