RE: [RFC][PATCH] add tracepoint to __sk_mem_schedule
From: Satoru Moriya <hidden>
Date: 2011-06-15 20:17:14
On 06/15/2011 04:04 PM, Neil Horman wrote:
On Wed, Jun 15, 2011 at 03:18:09PM -0400, Satoru Moriya wrote:quoted
Right. We can catch packet drop events via dropwatch/kfree_skb tracepoint. OTOH, what I'd like to know is why kernel dropped packets. Basically, we're able to know the reason because kfree_skb tracepoint records its return address. But in this case it is difficult to know a detailed reason because there are some possibilities. I'll try to explain UDP case.Ok, but you're patch only gets us incrmentally closer to that goal. You can tell if theres too much memory pressue to accept another packet on a socket, but it doesn't tell you if for instance if sk_filter indicated an error, or something else failed. In short it sounds to me like you're trying to debug a specific problem, rather than add something generally useful.
You're right.
quoted
As you said, we're able to catch the packet drop in __udp_queue_rcv_skb and it means ip_queue_rcv_skb/sock_queue_rcv_skb returned negative value. In sock_queue_rcv_skb there are 3 possibilities where it returns nagative val but we can't separate them from the kfree_skb tracepoint. MoreoverThis is true, but you're patch doesn't do that either (not that either feature holds that as a specific goal).
Right.
quoted
sock_queue_rcv_skb calls __sk_mem_schedule and there are several if-statement to decide whether kernel should drop a packet in it. I'd like to know the condition when it returned error because some of them are related to sysctl knob e.g. /proc/sys/net/ipv4/udp_mem, udp_rmem_min, udp_wmem_min for UDP case and we can tune kernel behavior easily. Also they are needed to show root cause of packet drop to our customers. Does it make sense?Yes, it makes sense, but I think the patch could be made generally more useful.
OK. I'll update my patch to make it more useful. Thanks, Satoru