From: Martin Zaharinov <hidden> Date: 2021-08-05 20:54:00
Hi Net dev team
Please check this error :
Last time I write for this problem : https://www.spinics.net/lists/netdev/msg707513.html
But not find any solution.
Config of server is : Bonding port channel (LACP) > Accel PPP server > Huawei switch.
Server is work fine users is down/up 500+ users .
But in one moment server make spike and affect other vlans in same server .
And in accel I see many row with this error.
Is there options to find and fix this bug.
With accel team I discus this problem and they claim it is kernel bug and need to find solution with Kernel dev team.
[2021-08-05 13:52:05.294] vlan912: 24b205903d09718e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.298] vlan912: 24b205903d097162: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.626] vlan641: 24b205903d09711b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:11.000] vlan912: 24b205903d097105: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:17.852] vlan912: 24b205903d0971ae: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:21.113] vlan641: 24b205903d09715b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:27.963] vlan912: 24b205903d09718d: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.249] vlan496: 24b205903d097184: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.992] vlan420: 24b205903d09718a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:33.937] vlan640: 24b205903d0971cd: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.032] vlan912: 24b205903d097182: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.420] vlan912: 24b205903d0971d5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan912: 24b205903d09713a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan614: 24b205903d0971e5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.102] vlan912: 24b205903d097190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097153: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097141: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.852] vlan912: 24b205903d097198: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.977] vlan637: 24b205903d097148: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:44.528] vlan637: 24b205903d0971c3: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
Martin
On Thu, Aug 05, 2021 at 11:53:50PM +0300, Martin Zaharinov wrote:
Hi Net dev team
Please check this error :
Last time I write for this problem : https://www.spinics.net/lists/netdev/msg707513.html
But not find any solution.
Config of server is : Bonding port channel (LACP) > Accel PPP server > Huawei switch.
Server is work fine users is down/up 500+ users .
But in one moment server make spike and affect other vlans in same server .
And in accel I see many row with this error.
Is there options to find and fix this bug.
With accel team I discus this problem and they claim it is kernel bug and need to find solution with Kernel dev team.
[2021-08-05 13:52:05.294] vlan912: 24b205903d09718e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.298] vlan912: 24b205903d097162: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.626] vlan641: 24b205903d09711b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:11.000] vlan912: 24b205903d097105: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:17.852] vlan912: 24b205903d0971ae: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:21.113] vlan641: 24b205903d09715b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:27.963] vlan912: 24b205903d09718d: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.249] vlan496: 24b205903d097184: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.992] vlan420: 24b205903d09718a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:33.937] vlan640: 24b205903d0971cd: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.032] vlan912: 24b205903d097182: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.420] vlan912: 24b205903d0971d5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan912: 24b205903d09713a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan614: 24b205903d0971e5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.102] vlan912: 24b205903d097190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097153: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097141: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.852] vlan912: 24b205903d097198: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.977] vlan637: 24b205903d097148: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:44.528] vlan637: 24b205903d0971c3: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
These are userspace error messages, not kernel messages.
What kernel version are you using?
thanks,
greg k-h
From: Martin Zaharinov <hidden> Date: 2021-08-06 05:41:01
Hi Greg
Latest kernel 5.13.8.
I try old version from 5.10 to 5.13 and its same error.
Martin
On 6 Aug 2021, at 7:40, Greg KH [off-list ref] wrote:
On Thu, Aug 05, 2021 at 11:53:50PM +0300, Martin Zaharinov wrote:
quoted
Hi Net dev team
Please check this error :
Last time I write for this problem : https://www.spinics.net/lists/netdev/msg707513.html
But not find any solution.
Config of server is : Bonding port channel (LACP) > Accel PPP server > Huawei switch.
Server is work fine users is down/up 500+ users .
But in one moment server make spike and affect other vlans in same server .
And in accel I see many row with this error.
Is there options to find and fix this bug.
With accel team I discus this problem and they claim it is kernel bug and need to find solution with Kernel dev team.
[2021-08-05 13:52:05.294] vlan912: 24b205903d09718e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.298] vlan912: 24b205903d097162: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.626] vlan641: 24b205903d09711b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:11.000] vlan912: 24b205903d097105: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:17.852] vlan912: 24b205903d0971ae: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:21.113] vlan641: 24b205903d09715b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:27.963] vlan912: 24b205903d09718d: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.249] vlan496: 24b205903d097184: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.992] vlan420: 24b205903d09718a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:33.937] vlan640: 24b205903d0971cd: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.032] vlan912: 24b205903d097182: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.420] vlan912: 24b205903d0971d5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan912: 24b205903d09713a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan614: 24b205903d0971e5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.102] vlan912: 24b205903d097190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097153: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097141: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.852] vlan912: 24b205903d097198: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.977] vlan637: 24b205903d097148: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:44.528] vlan637: 24b205903d0971c3: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
These are userspace error messages, not kernel messages.
What kernel version are you using?
thanks,
greg k-h
From: Martin Zaharinov <hidden> Date: 2021-08-08 15:14:15
Add Pali Rohár,
If have any idea .
Martin
On 6 Aug 2021, at 7:40, Greg KH [off-list ref] wrote:
On Thu, Aug 05, 2021 at 11:53:50PM +0300, Martin Zaharinov wrote:
quoted
Hi Net dev team
Please check this error :
Last time I write for this problem : https://www.spinics.net/lists/netdev/msg707513.html
But not find any solution.
Config of server is : Bonding port channel (LACP) > Accel PPP server > Huawei switch.
Server is work fine users is down/up 500+ users .
But in one moment server make spike and affect other vlans in same server .
And in accel I see many row with this error.
Is there options to find and fix this bug.
With accel team I discus this problem and they claim it is kernel bug and need to find solution with Kernel dev team.
[2021-08-05 13:52:05.294] vlan912: 24b205903d09718e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.298] vlan912: 24b205903d097162: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.626] vlan641: 24b205903d09711b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:11.000] vlan912: 24b205903d097105: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:17.852] vlan912: 24b205903d0971ae: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:21.113] vlan641: 24b205903d09715b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:27.963] vlan912: 24b205903d09718d: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.249] vlan496: 24b205903d097184: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.992] vlan420: 24b205903d09718a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:33.937] vlan640: 24b205903d0971cd: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.032] vlan912: 24b205903d097182: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.420] vlan912: 24b205903d0971d5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan912: 24b205903d09713a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan614: 24b205903d0971e5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.102] vlan912: 24b205903d097190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097153: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097141: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.852] vlan912: 24b205903d097198: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.977] vlan637: 24b205903d097148: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:44.528] vlan637: 24b205903d0971c3: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
These are userspace error messages, not kernel messages.
What kernel version are you using?
thanks,
greg k-h
Hello!
On Sunday 08 August 2021 18:14:09 Martin Zaharinov wrote:
Add Pali Rohár,
If have any idea .
Martin
quoted
On 6 Aug 2021, at 7:40, Greg KH [off-list ref] wrote:
On Thu, Aug 05, 2021 at 11:53:50PM +0300, Martin Zaharinov wrote:
quoted
Hi Net dev team
Please check this error :
Last time I write for this problem : https://www.spinics.net/lists/netdev/msg707513.html
But not find any solution.
Config of server is : Bonding port channel (LACP) > Accel PPP server > Huawei switch.
Server is work fine users is down/up 500+ users .
But in one moment server make spike and affect other vlans in same server .
When this error started to happen? After kernel upgrade? After pppd
upgrade? Or after system upgrade? Or when more users started to
connecting?
quoted
quoted
And in accel I see many row with this error.
Is there options to find and fix this bug.
With accel team I discus this problem and they claim it is kernel bug and need to find solution with Kernel dev team.
[2021-08-05 13:52:05.294] vlan912: 24b205903d09718e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.298] vlan912: 24b205903d097162: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.626] vlan641: 24b205903d09711b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:11.000] vlan912: 24b205903d097105: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:17.852] vlan912: 24b205903d0971ae: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:21.113] vlan641: 24b205903d09715b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:27.963] vlan912: 24b205903d09718d: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.249] vlan496: 24b205903d097184: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.992] vlan420: 24b205903d09718a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:33.937] vlan640: 24b205903d0971cd: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.032] vlan912: 24b205903d097182: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.420] vlan912: 24b205903d0971d5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan912: 24b205903d09713a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan614: 24b205903d0971e5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.102] vlan912: 24b205903d097190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097153: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097141: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.852] vlan912: 24b205903d097198: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.977] vlan637: 24b205903d097148: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:44.528] vlan637: 24b205903d0971c3: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
These are userspace error messages, not kernel messages.
What kernel version are you using?
Yes, we need to know, what kernel version are you using.
quoted
thanks,
greg k-h
And also another question, what version of pppd daemon are you using?
Also, are you able to dump state of ppp channels and ppp units? It is
needed to know to which tty device, file descriptor (or socket
extension) is (or should be) particular ppp channel bounded.
From: Martin Zaharinov <hidden> Date: 2021-08-08 15:29:36
Hi Pali
Kernel 5.13.8
The problem is from kernel 5.8 > I try all major update 5.9, 5.10, 5.11 ,5.12
I use accel-pppd daemon (not pppd) .
And yes after users started to connecting .
When system boot and connect first time all user connect without any problem .
In time of work user disconnect and connect (power cut , fiber cut or other problem in network) , but in time of spike (may be make lock or other problem ) disconnect ~ 400-500 users and affect other users. Process go to load over 100% and In statistic I see many finishing connection and many start connection.
And in this time in log get many lines with ioctl(PPPIOCCONNECT): Transport endpoint is not connected. After finish (unlock or other) stop to see this error and system is back to normal. And connect all disconnected users.
Martin
On 8 Aug 2021, at 18:23, Pali Rohár [off-list ref] wrote:
Hello!
On Sunday 08 August 2021 18:14:09 Martin Zaharinov wrote:
quoted
Add Pali Rohár,
If have any idea .
Martin
quoted
On 6 Aug 2021, at 7:40, Greg KH [off-list ref] wrote:
On Thu, Aug 05, 2021 at 11:53:50PM +0300, Martin Zaharinov wrote:
quoted
Hi Net dev team
Please check this error :
Last time I write for this problem : https://www.spinics.net/lists/netdev/msg707513.html
But not find any solution.
Config of server is : Bonding port channel (LACP) > Accel PPP server > Huawei switch.
Server is work fine users is down/up 500+ users .
But in one moment server make spike and affect other vlans in same server .
When this error started to happen? After kernel upgrade? After pppd
upgrade? Or after system upgrade? Or when more users started to
connecting?
quoted
quoted
quoted
And in accel I see many row with this error.
Is there options to find and fix this bug.
With accel team I discus this problem and they claim it is kernel bug and need to find solution with Kernel dev team.
[2021-08-05 13:52:05.294] vlan912: 24b205903d09718e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.298] vlan912: 24b205903d097162: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.626] vlan641: 24b205903d09711b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:11.000] vlan912: 24b205903d097105: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:17.852] vlan912: 24b205903d0971ae: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:21.113] vlan641: 24b205903d09715b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:27.963] vlan912: 24b205903d09718d: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.249] vlan496: 24b205903d097184: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.992] vlan420: 24b205903d09718a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:33.937] vlan640: 24b205903d0971cd: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.032] vlan912: 24b205903d097182: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.420] vlan912: 24b205903d0971d5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan912: 24b205903d09713a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan614: 24b205903d0971e5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.102] vlan912: 24b205903d097190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097153: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097141: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.852] vlan912: 24b205903d097198: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.977] vlan637: 24b205903d097148: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:44.528] vlan637: 24b205903d0971c3: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
These are userspace error messages, not kernel messages.
What kernel version are you using?
Yes, we need to know, what kernel version are you using.
quoted
quoted
thanks,
greg k-h
And also another question, what version of pppd daemon are you using?
Also, are you able to dump state of ppp channels and ppp units? It is
needed to know to which tty device, file descriptor (or socket
extension) is (or should be) particular ppp channel bounded.
On Sunday 08 August 2021 18:29:30 Martin Zaharinov wrote:
Hi Pali
Kernel 5.13.8
The problem is from kernel 5.8 > I try all major update 5.9, 5.10, 5.11 ,5.12
I use accel-pppd daemon (not pppd) .
I'm not using accel-pppd, so cannot help here.
I would suggest to try "git bisect" kernel version which started to be
problematic for accel-pppd.
Providing state of ppp channels and ppp units could help to debug this
issue, but I'm not sure if accel-pppd has this debug feature. IIRC only
process which has ppp file descriptors can retrieve and dump this
information.
And yes after users started to connecting .
When system boot and connect first time all user connect without any problem .
In time of work user disconnect and connect (power cut , fiber cut or other problem in network) , but in time of spike (may be make lock or other problem ) disconnect ~ 400-500 users and affect other users. Process go to load over 100% and In statistic I see many finishing connection and many start connection.
And in this time in log get many lines with ioctl(PPPIOCCONNECT): Transport endpoint is not connected. After finish (unlock or other) stop to see this error and system is back to normal. And connect all disconnected users.
Martin
quoted
On 8 Aug 2021, at 18:23, Pali Rohár [off-list ref] wrote:
Hello!
On Sunday 08 August 2021 18:14:09 Martin Zaharinov wrote:
quoted
Add Pali Rohár,
If have any idea .
Martin
quoted
On 6 Aug 2021, at 7:40, Greg KH [off-list ref] wrote:
On Thu, Aug 05, 2021 at 11:53:50PM +0300, Martin Zaharinov wrote:
quoted
Hi Net dev team
Please check this error :
Last time I write for this problem : https://www.spinics.net/lists/netdev/msg707513.html
But not find any solution.
Config of server is : Bonding port channel (LACP) > Accel PPP server > Huawei switch.
Server is work fine users is down/up 500+ users .
But in one moment server make spike and affect other vlans in same server .
When this error started to happen? After kernel upgrade? After pppd
upgrade? Or after system upgrade? Or when more users started to
connecting?
quoted
quoted
quoted
And in accel I see many row with this error.
Is there options to find and fix this bug.
With accel team I discus this problem and they claim it is kernel bug and need to find solution with Kernel dev team.
[2021-08-05 13:52:05.294] vlan912: 24b205903d09718e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.298] vlan912: 24b205903d097162: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.626] vlan641: 24b205903d09711b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:11.000] vlan912: 24b205903d097105: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:17.852] vlan912: 24b205903d0971ae: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:21.113] vlan641: 24b205903d09715b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:27.963] vlan912: 24b205903d09718d: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.249] vlan496: 24b205903d097184: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.992] vlan420: 24b205903d09718a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:33.937] vlan640: 24b205903d0971cd: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.032] vlan912: 24b205903d097182: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.420] vlan912: 24b205903d0971d5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan912: 24b205903d09713a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan614: 24b205903d0971e5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.102] vlan912: 24b205903d097190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097153: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097141: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.852] vlan912: 24b205903d097198: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.977] vlan637: 24b205903d097148: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:44.528] vlan637: 24b205903d0971c3: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
These are userspace error messages, not kernel messages.
What kernel version are you using?
Yes, we need to know, what kernel version are you using.
quoted
quoted
thanks,
greg k-h
And also another question, what version of pppd daemon are you using?
Also, are you able to dump state of ppp channels and ppp units? It is
needed to know to which tty device, file descriptor (or socket
extension) is (or should be) particular ppp channel bounded.
From: Martin Zaharinov <hidden> Date: 2021-08-10 18:27:20
Add Guillaume Nault
On 9 Aug 2021, at 18:15, Pali Rohár [off-list ref] wrote:
On Sunday 08 August 2021 18:29:30 Martin Zaharinov wrote:
quoted
Hi Pali
Kernel 5.13.8
The problem is from kernel 5.8 > I try all major update 5.9, 5.10, 5.11 ,5.12
I use accel-pppd daemon (not pppd) .
I'm not using accel-pppd, so cannot help here.
I would suggest to try "git bisect" kernel version which started to be
problematic for accel-pppd.
Providing state of ppp channels and ppp units could help to debug this
issue, but I'm not sure if accel-pppd has this debug feature. IIRC only
process which has ppp file descriptors can retrieve and dump this
information.
quoted
And yes after users started to connecting .
When system boot and connect first time all user connect without any problem .
In time of work user disconnect and connect (power cut , fiber cut or other problem in network) , but in time of spike (may be make lock or other problem ) disconnect ~ 400-500 users and affect other users. Process go to load over 100% and In statistic I see many finishing connection and many start connection.
And in this time in log get many lines with ioctl(PPPIOCCONNECT): Transport endpoint is not connected. After finish (unlock or other) stop to see this error and system is back to normal. And connect all disconnected users.
Martin
quoted
On 8 Aug 2021, at 18:23, Pali Rohár [off-list ref] wrote:
Hello!
On Sunday 08 August 2021 18:14:09 Martin Zaharinov wrote:
quoted
Add Pali Rohár,
If have any idea .
Martin
quoted
On 6 Aug 2021, at 7:40, Greg KH [off-list ref] wrote:
On Thu, Aug 05, 2021 at 11:53:50PM +0300, Martin Zaharinov wrote:
quoted
Hi Net dev team
Please check this error :
Last time I write for this problem : https://www.spinics.net/lists/netdev/msg707513.html
But not find any solution.
Config of server is : Bonding port channel (LACP) > Accel PPP server > Huawei switch.
Server is work fine users is down/up 500+ users .
But in one moment server make spike and affect other vlans in same server .
When this error started to happen? After kernel upgrade? After pppd
upgrade? Or after system upgrade? Or when more users started to
connecting?
quoted
quoted
quoted
And in accel I see many row with this error.
Is there options to find and fix this bug.
With accel team I discus this problem and they claim it is kernel bug and need to find solution with Kernel dev team.
[2021-08-05 13:52:05.294] vlan912: 24b205903d09718e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.298] vlan912: 24b205903d097162: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.626] vlan641: 24b205903d09711b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:11.000] vlan912: 24b205903d097105: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:17.852] vlan912: 24b205903d0971ae: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:21.113] vlan641: 24b205903d09715b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:27.963] vlan912: 24b205903d09718d: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.249] vlan496: 24b205903d097184: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.992] vlan420: 24b205903d09718a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:33.937] vlan640: 24b205903d0971cd: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.032] vlan912: 24b205903d097182: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.420] vlan912: 24b205903d0971d5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan912: 24b205903d09713a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan614: 24b205903d0971e5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.102] vlan912: 24b205903d097190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097153: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097141: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.852] vlan912: 24b205903d097198: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.977] vlan637: 24b205903d097148: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:44.528] vlan637: 24b205903d0971c3: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
These are userspace error messages, not kernel messages.
What kernel version are you using?
Yes, we need to know, what kernel version are you using.
quoted
quoted
thanks,
greg k-h
And also another question, what version of pppd daemon are you using?
Also, are you able to dump state of ppp channels and ppp units? It is
needed to know to which tty device, file descriptor (or socket
extension) is (or should be) particular ppp channel bounded.
From: Martin Zaharinov <hidden> Date: 2021-08-11 11:10:41
And one more that see.
Problem is come when accel start finishing sessions,
Now in server have 2k users and restart on one of vlans 3 Olt with 400 users and affect other vlans ,
And problem is start when start destroying dead sessions from vlan with 3 Olt and this affect all other vlans.
May be kernel destroy old session slow and entrained other users by locking other sessions.
is there a way to speed up the closing of stopped/dead sessions.
Martin
On 9 Aug 2021, at 18:15, Pali Rohár [off-list ref] wrote:
On Sunday 08 August 2021 18:29:30 Martin Zaharinov wrote:
quoted
Hi Pali
Kernel 5.13.8
The problem is from kernel 5.8 > I try all major update 5.9, 5.10, 5.11 ,5.12
I use accel-pppd daemon (not pppd) .
I'm not using accel-pppd, so cannot help here.
I would suggest to try "git bisect" kernel version which started to be
problematic for accel-pppd.
Providing state of ppp channels and ppp units could help to debug this
issue, but I'm not sure if accel-pppd has this debug feature. IIRC only
process which has ppp file descriptors can retrieve and dump this
information.
quoted
And yes after users started to connecting .
When system boot and connect first time all user connect without any problem .
In time of work user disconnect and connect (power cut , fiber cut or other problem in network) , but in time of spike (may be make lock or other problem ) disconnect ~ 400-500 users and affect other users. Process go to load over 100% and In statistic I see many finishing connection and many start connection.
And in this time in log get many lines with ioctl(PPPIOCCONNECT): Transport endpoint is not connected. After finish (unlock or other) stop to see this error and system is back to normal. And connect all disconnected users.
Martin
quoted
On 8 Aug 2021, at 18:23, Pali Rohár [off-list ref] wrote:
Hello!
On Sunday 08 August 2021 18:14:09 Martin Zaharinov wrote:
quoted
Add Pali Rohár,
If have any idea .
Martin
quoted
On 6 Aug 2021, at 7:40, Greg KH [off-list ref] wrote:
On Thu, Aug 05, 2021 at 11:53:50PM +0300, Martin Zaharinov wrote:
quoted
Hi Net dev team
Please check this error :
Last time I write for this problem : https://www.spinics.net/lists/netdev/msg707513.html
But not find any solution.
Config of server is : Bonding port channel (LACP) > Accel PPP server > Huawei switch.
Server is work fine users is down/up 500+ users .
But in one moment server make spike and affect other vlans in same server .
When this error started to happen? After kernel upgrade? After pppd
upgrade? Or after system upgrade? Or when more users started to
connecting?
quoted
quoted
quoted
And in accel I see many row with this error.
Is there options to find and fix this bug.
With accel team I discus this problem and they claim it is kernel bug and need to find solution with Kernel dev team.
[2021-08-05 13:52:05.294] vlan912: 24b205903d09718e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.298] vlan912: 24b205903d097162: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.626] vlan641: 24b205903d09711b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:11.000] vlan912: 24b205903d097105: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:17.852] vlan912: 24b205903d0971ae: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:21.113] vlan641: 24b205903d09715b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:27.963] vlan912: 24b205903d09718d: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.249] vlan496: 24b205903d097184: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.992] vlan420: 24b205903d09718a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:33.937] vlan640: 24b205903d0971cd: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.032] vlan912: 24b205903d097182: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.420] vlan912: 24b205903d0971d5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan912: 24b205903d09713a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan614: 24b205903d0971e5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.102] vlan912: 24b205903d097190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097153: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097141: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.852] vlan912: 24b205903d097198: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.977] vlan637: 24b205903d097148: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:44.528] vlan637: 24b205903d0971c3: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
These are userspace error messages, not kernel messages.
What kernel version are you using?
Yes, we need to know, what kernel version are you using.
quoted
quoted
thanks,
greg k-h
And also another question, what version of pppd daemon are you using?
Also, are you able to dump state of ppp channels and ppp units? It is
needed to know to which tty device, file descriptor (or socket
extension) is (or should be) particular ppp channel bounded.
On Tue, Aug 10, 2021 at 09:27:14PM +0300, Martin Zaharinov wrote:
Add Guillaume Nault
quoted
On 9 Aug 2021, at 18:15, Pali Rohár [off-list ref] wrote:
On Sunday 08 August 2021 18:29:30 Martin Zaharinov wrote:
quoted
quoted
quoted
quoted
quoted
[2021-08-05 13:52:05.294] vlan912: 24b205903d09718e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.298] vlan912: 24b205903d097162: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:05.626] vlan641: 24b205903d09711b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:11.000] vlan912: 24b205903d097105: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:17.852] vlan912: 24b205903d0971ae: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:21.113] vlan641: 24b205903d09715b: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:27.963] vlan912: 24b205903d09718d: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.249] vlan496: 24b205903d097184: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:30.992] vlan420: 24b205903d09718a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:33.937] vlan640: 24b205903d0971cd: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.032] vlan912: 24b205903d097182: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:40.420] vlan912: 24b205903d0971d5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan912: 24b205903d09713a: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:42.799] vlan614: 24b205903d0971e5: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.102] vlan912: 24b205903d097190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097153: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.850] vlan479: 24b205903d097141: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.852] vlan912: 24b205903d097198: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:43.977] vlan637: 24b205903d097148: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-08-05 13:52:44.528] vlan637: 24b205903d0971c3: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
The PPPIOCCONNECT ioctl returns -ENOTCONN if the ppp channel has been
unregistered.
From a user space point of view, this means that accel-ppp establishes
PPPoE sessions, starts negociating PPP connection parameters on top of
them (LCP and authentication) and finally the PPPoE sessions get
disconnected before accel-ppp connects them to ppp units (units are
roughly the "pppX" network devices).
Unregistration of PPPoE channels can happen for the following reasons:
* Changing some parameters of the network interface used by the
PPPoE connection: MAC address, MTU, bringing the device down.
* Reception of a PADT (PPPoE disconnection message sent from the peer).
* Closing the PPPoE socket.
* Re-connecting a PPPoE socket with a different session ID (this
unregisters the previous channel and creates a new one, so that
shouldn't be the problem you're facing here).
Given that this seems to affect all PPPoE connections, I guess
something happened to the underlying network interface (1st bullet
point).
On Wed, Aug 11, 2021 at 02:10:32PM +0300, Martin Zaharinov wrote:
And one more that see.
Problem is come when accel start finishing sessions,
Now in server have 2k users and restart on one of vlans 3 Olt with 400 users and affect other vlans ,
And problem is start when start destroying dead sessions from vlan with 3 Olt and this affect all other vlans.
May be kernel destroy old session slow and entrained other users by locking other sessions.
is there a way to speed up the closing of stopped/dead sessions.
What are the CPU stats when that happen? Is it users space or kernel
space that keeps it busy?
One easy way to check is to run "mpstat 1" for a few seconds when the
problem occurs.
From: Martin Zaharinov <hidden> Date: 2021-09-07 06:16:24
Hi
Sorry for delay but not easy to catch moment .
See this is mpstatl 1 :
Linux 5.14.1 (demobng) 09/07/21 _x86_64_ (12 CPU)
11:12:16 CPU %usr %nice %sys %iowait %irq %soft %steal %guest %gnice %idle
11:12:17 all 0.17 0.00 6.66 0.00 0.00 4.13 0.00 0.00 0.00 89.05
11:12:18 all 0.25 0.00 8.36 0.00 0.00 4.88 0.00 0.00 0.00 86.51
11:12:19 all 0.26 0.00 9.62 0.00 0.00 3.91 0.00 0.00 0.00 86.21
11:12:20 all 0.85 0.00 6.00 0.00 0.00 4.31 0.00 0.00 0.00 88.84
11:12:21 all 0.08 0.00 4.45 0.00 0.00 4.79 0.00 0.00 0.00 90.67
11:12:22 all 0.17 0.00 9.50 0.00 0.00 4.58 0.00 0.00 0.00 85.75
11:12:23 all 0.00 0.00 6.92 0.00 0.00 2.48 0.00 0.00 0.00 90.61
11:12:24 all 0.17 0.00 5.45 0.00 0.00 4.27 0.00 0.00 0.00 90.11
11:12:25 all 0.25 0.00 5.38 0.00 0.00 4.79 0.00 0.00 0.00 89.58
11:12:26 all 0.60 0.00 1.45 0.00 0.00 2.65 0.00 0.00 0.00 95.30
11:12:27 all 0.42 0.00 6.91 0.00 0.00 4.47 0.00 0.00 0.00 88.20
11:12:28 all 0.00 0.00 6.75 0.00 0.00 4.18 0.00 0.00 0.00 89.07
11:12:29 all 0.17 0.00 3.52 0.00 0.00 5.11 0.00 0.00 0.00 91.20
11:12:30 all 1.45 0.00 10.14 0.00 0.00 3.49 0.00 0.00 0.00 84.92
11:12:31 all 0.09 0.00 5.11 0.00 0.00 4.77 0.00 0.00 0.00 90.03
11:12:32 all 0.25 0.00 3.11 0.00 0.00 4.46 0.00 0.00 0.00 92.17
Average: all 0.32 0.00 6.21 0.00 0.00 4.21 0.00 0.00 0.00 89.26
I attache and one screenshot from perf top (Screenshot is send on preview mail)
And I see in lsmod
pppoe 20480 8198
pppox 16384 1 pppoe
ppp_generic 45056 16364 pppox,pppoe
slhc 16384 1 ppp_generic
To slow remove pppoe session .
And from log :
[2021-09-07 11:01:11.129] vlan3020: ebdd1c5d8b5900f6: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:01:53.621] vlan643: ebdd1c5d8b59014e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:00.359] vlan1616: ebdd1c5d8b590195: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:05.859] vlan3020: ebdd1c5d8b5900d8: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:08.258] vlan3005: ebdd1c5d8b590190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:13.820] vlan643: ebdd1c5d8b590152: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:15.839] vlan727: ebdd1c5d8b590144: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:20.139] vlan1693: ebdd1c5d8b59019f: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
On 11 Aug 2021, at 19:48, Guillaume Nault [off-list ref] wrote:
On Wed, Aug 11, 2021 at 02:10:32PM +0300, Martin Zaharinov wrote:
quoted
And one more that see.
Problem is come when accel start finishing sessions,
Now in server have 2k users and restart on one of vlans 3 Olt with 400 users and affect other vlans ,
And problem is start when start destroying dead sessions from vlan with 3 Olt and this affect all other vlans.
May be kernel destroy old session slow and entrained other users by locking other sessions.
is there a way to speed up the closing of stopped/dead sessions.
What are the CPU stats when that happen? Is it users space or kernel
space that keeps it busy?
One easy way to check is to run "mpstat 1" for a few seconds when the
problem occurs.
On 7 Sep 2021, at 9:16, Martin Zaharinov [off-list ref] wrote:
Hi
Sorry for delay but not easy to catch moment .
See this is mpstatl 1 :
Linux 5.14.1 (demobng) 09/07/21 _x86_64_ (12 CPU)
11:12:16 CPU %usr %nice %sys %iowait %irq %soft %steal %guest %gnice %idle
11:12:17 all 0.17 0.00 6.66 0.00 0.00 4.13 0.00 0.00 0.00 89.05
11:12:18 all 0.25 0.00 8.36 0.00 0.00 4.88 0.00 0.00 0.00 86.51
11:12:19 all 0.26 0.00 9.62 0.00 0.00 3.91 0.00 0.00 0.00 86.21
11:12:20 all 0.85 0.00 6.00 0.00 0.00 4.31 0.00 0.00 0.00 88.84
11:12:21 all 0.08 0.00 4.45 0.00 0.00 4.79 0.00 0.00 0.00 90.67
11:12:22 all 0.17 0.00 9.50 0.00 0.00 4.58 0.00 0.00 0.00 85.75
11:12:23 all 0.00 0.00 6.92 0.00 0.00 2.48 0.00 0.00 0.00 90.61
11:12:24 all 0.17 0.00 5.45 0.00 0.00 4.27 0.00 0.00 0.00 90.11
11:12:25 all 0.25 0.00 5.38 0.00 0.00 4.79 0.00 0.00 0.00 89.58
11:12:26 all 0.60 0.00 1.45 0.00 0.00 2.65 0.00 0.00 0.00 95.30
11:12:27 all 0.42 0.00 6.91 0.00 0.00 4.47 0.00 0.00 0.00 88.20
11:12:28 all 0.00 0.00 6.75 0.00 0.00 4.18 0.00 0.00 0.00 89.07
11:12:29 all 0.17 0.00 3.52 0.00 0.00 5.11 0.00 0.00 0.00 91.20
11:12:30 all 1.45 0.00 10.14 0.00 0.00 3.49 0.00 0.00 0.00 84.92
11:12:31 all 0.09 0.00 5.11 0.00 0.00 4.77 0.00 0.00 0.00 90.03
11:12:32 all 0.25 0.00 3.11 0.00 0.00 4.46 0.00 0.00 0.00 92.17
Average: all 0.32 0.00 6.21 0.00 0.00 4.21 0.00 0.00 0.00 89.26
I attache and one screenshot from perf top (Screenshot is send on preview mail)
And I see in lsmod
pppoe 20480 8198
pppox 16384 1 pppoe
ppp_generic 45056 16364 pppox,pppoe
slhc 16384 1 ppp_generic
To slow remove pppoe session .
And from log :
[2021-09-07 11:01:11.129] vlan3020: ebdd1c5d8b5900f6: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:01:53.621] vlan643: ebdd1c5d8b59014e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:00.359] vlan1616: ebdd1c5d8b590195: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:05.859] vlan3020: ebdd1c5d8b5900d8: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:08.258] vlan3005: ebdd1c5d8b590190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:13.820] vlan643: ebdd1c5d8b590152: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:15.839] vlan727: ebdd1c5d8b590144: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:20.139] vlan1693: ebdd1c5d8b59019f: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
quoted
On 11 Aug 2021, at 19:48, Guillaume Nault [off-list ref] wrote:
On Wed, Aug 11, 2021 at 02:10:32PM +0300, Martin Zaharinov wrote:
quoted
And one more that see.
Problem is come when accel start finishing sessions,
Now in server have 2k users and restart on one of vlans 3 Olt with 400 users and affect other vlans ,
And problem is start when start destroying dead sessions from vlan with 3 Olt and this affect all other vlans.
May be kernel destroy old session slow and entrained other users by locking other sessions.
is there a way to speed up the closing of stopped/dead sessions.
What are the CPU stats when that happen? Is it users space or kernel
space that keeps it busy?
One easy way to check is to run "mpstat 1" for a few seconds when the
problem occurs.
From: Martin Zaharinov <hidden> Date: 2021-09-11 06:26:52
Hi Guillaume
Main problem is overload of service because have many finishing ppp (customer) last two day down from 40-50 to 100-200 users and make problem when is happen if try to type : ip a wait 10-20 sec to start list interface .
But how to find where is a problem any locking or other.
And is there options to make fast remove ppp interface from kernel to reduce this load.
Martin
On 7 Sep 2021, at 9:16, Martin Zaharinov [off-list ref] wrote:
Hi
Sorry for delay but not easy to catch moment .
See this is mpstatl 1 :
Linux 5.14.1 (demobng) 09/07/21 _x86_64_ (12 CPU)
11:12:16 CPU %usr %nice %sys %iowait %irq %soft %steal %guest %gnice %idle
11:12:17 all 0.17 0.00 6.66 0.00 0.00 4.13 0.00 0.00 0.00 89.05
11:12:18 all 0.25 0.00 8.36 0.00 0.00 4.88 0.00 0.00 0.00 86.51
11:12:19 all 0.26 0.00 9.62 0.00 0.00 3.91 0.00 0.00 0.00 86.21
11:12:20 all 0.85 0.00 6.00 0.00 0.00 4.31 0.00 0.00 0.00 88.84
11:12:21 all 0.08 0.00 4.45 0.00 0.00 4.79 0.00 0.00 0.00 90.67
11:12:22 all 0.17 0.00 9.50 0.00 0.00 4.58 0.00 0.00 0.00 85.75
11:12:23 all 0.00 0.00 6.92 0.00 0.00 2.48 0.00 0.00 0.00 90.61
11:12:24 all 0.17 0.00 5.45 0.00 0.00 4.27 0.00 0.00 0.00 90.11
11:12:25 all 0.25 0.00 5.38 0.00 0.00 4.79 0.00 0.00 0.00 89.58
11:12:26 all 0.60 0.00 1.45 0.00 0.00 2.65 0.00 0.00 0.00 95.30
11:12:27 all 0.42 0.00 6.91 0.00 0.00 4.47 0.00 0.00 0.00 88.20
11:12:28 all 0.00 0.00 6.75 0.00 0.00 4.18 0.00 0.00 0.00 89.07
11:12:29 all 0.17 0.00 3.52 0.00 0.00 5.11 0.00 0.00 0.00 91.20
11:12:30 all 1.45 0.00 10.14 0.00 0.00 3.49 0.00 0.00 0.00 84.92
11:12:31 all 0.09 0.00 5.11 0.00 0.00 4.77 0.00 0.00 0.00 90.03
11:12:32 all 0.25 0.00 3.11 0.00 0.00 4.46 0.00 0.00 0.00 92.17
Average: all 0.32 0.00 6.21 0.00 0.00 4.21 0.00 0.00 0.00 89.26
I attache and one screenshot from perf top (Screenshot is send on preview mail)
And I see in lsmod
pppoe 20480 8198
pppox 16384 1 pppoe
ppp_generic 45056 16364 pppox,pppoe
slhc 16384 1 ppp_generic
To slow remove pppoe session .
And from log :
[2021-09-07 11:01:11.129] vlan3020: ebdd1c5d8b5900f6: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:01:53.621] vlan643: ebdd1c5d8b59014e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:00.359] vlan1616: ebdd1c5d8b590195: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:05.859] vlan3020: ebdd1c5d8b5900d8: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:08.258] vlan3005: ebdd1c5d8b590190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:13.820] vlan643: ebdd1c5d8b590152: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:15.839] vlan727: ebdd1c5d8b590144: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:20.139] vlan1693: ebdd1c5d8b59019f: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
quoted
On 11 Aug 2021, at 19:48, Guillaume Nault [off-list ref] wrote:
On Wed, Aug 11, 2021 at 02:10:32PM +0300, Martin Zaharinov wrote:
quoted
And one more that see.
Problem is come when accel start finishing sessions,
Now in server have 2k users and restart on one of vlans 3 Olt with 400 users and affect other vlans ,
And problem is start when start destroying dead sessions from vlan with 3 Olt and this affect all other vlans.
May be kernel destroy old session slow and entrained other users by locking other sessions.
is there a way to speed up the closing of stopped/dead sessions.
What are the CPU stats when that happen? Is it users space or kernel
space that keeps it busy?
One easy way to check is to run "mpstat 1" for a few seconds when the
problem occurs.
On 11 Sep 2021, at 9:26, Martin Zaharinov [off-list ref] wrote:
Hi Guillaume
Main problem is overload of service because have many finishing ppp (customer) last two day down from 40-50 to 100-200 users and make problem when is happen if try to type : ip a wait 10-20 sec to start list interface .
But how to find where is a problem any locking or other.
And is there options to make fast remove ppp interface from kernel to reduce this load.
Martin
On 7 Sep 2021, at 9:16, Martin Zaharinov [off-list ref] wrote:
Hi
Sorry for delay but not easy to catch moment .
See this is mpstatl 1 :
Linux 5.14.1 (demobng) 09/07/21 _x86_64_ (12 CPU)
11:12:16 CPU %usr %nice %sys %iowait %irq %soft %steal %guest %gnice %idle
11:12:17 all 0.17 0.00 6.66 0.00 0.00 4.13 0.00 0.00 0.00 89.05
11:12:18 all 0.25 0.00 8.36 0.00 0.00 4.88 0.00 0.00 0.00 86.51
11:12:19 all 0.26 0.00 9.62 0.00 0.00 3.91 0.00 0.00 0.00 86.21
11:12:20 all 0.85 0.00 6.00 0.00 0.00 4.31 0.00 0.00 0.00 88.84
11:12:21 all 0.08 0.00 4.45 0.00 0.00 4.79 0.00 0.00 0.00 90.67
11:12:22 all 0.17 0.00 9.50 0.00 0.00 4.58 0.00 0.00 0.00 85.75
11:12:23 all 0.00 0.00 6.92 0.00 0.00 2.48 0.00 0.00 0.00 90.61
11:12:24 all 0.17 0.00 5.45 0.00 0.00 4.27 0.00 0.00 0.00 90.11
11:12:25 all 0.25 0.00 5.38 0.00 0.00 4.79 0.00 0.00 0.00 89.58
11:12:26 all 0.60 0.00 1.45 0.00 0.00 2.65 0.00 0.00 0.00 95.30
11:12:27 all 0.42 0.00 6.91 0.00 0.00 4.47 0.00 0.00 0.00 88.20
11:12:28 all 0.00 0.00 6.75 0.00 0.00 4.18 0.00 0.00 0.00 89.07
11:12:29 all 0.17 0.00 3.52 0.00 0.00 5.11 0.00 0.00 0.00 91.20
11:12:30 all 1.45 0.00 10.14 0.00 0.00 3.49 0.00 0.00 0.00 84.92
11:12:31 all 0.09 0.00 5.11 0.00 0.00 4.77 0.00 0.00 0.00 90.03
11:12:32 all 0.25 0.00 3.11 0.00 0.00 4.46 0.00 0.00 0.00 92.17
Average: all 0.32 0.00 6.21 0.00 0.00 4.21 0.00 0.00 0.00 89.26
I attache and one screenshot from perf top (Screenshot is send on preview mail)
And I see in lsmod
pppoe 20480 8198
pppox 16384 1 pppoe
ppp_generic 45056 16364 pppox,pppoe
slhc 16384 1 ppp_generic
To slow remove pppoe session .
And from log :
[2021-09-07 11:01:11.129] vlan3020: ebdd1c5d8b5900f6: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:01:53.621] vlan643: ebdd1c5d8b59014e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:00.359] vlan1616: ebdd1c5d8b590195: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:05.859] vlan3020: ebdd1c5d8b5900d8: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:08.258] vlan3005: ebdd1c5d8b590190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:13.820] vlan643: ebdd1c5d8b590152: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:15.839] vlan727: ebdd1c5d8b590144: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:20.139] vlan1693: ebdd1c5d8b59019f: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
quoted
On 11 Aug 2021, at 19:48, Guillaume Nault [off-list ref] wrote:
On Wed, Aug 11, 2021 at 02:10:32PM +0300, Martin Zaharinov wrote:
quoted
And one more that see.
Problem is come when accel start finishing sessions,
Now in server have 2k users and restart on one of vlans 3 Olt with 400 users and affect other vlans ,
And problem is start when start destroying dead sessions from vlan with 3 Olt and this affect all other vlans.
May be kernel destroy old session slow and entrained other users by locking other sessions.
is there a way to speed up the closing of stopped/dead sessions.
What are the CPU stats when that happen? Is it users space or kernel
space that keeps it busy?
One easy way to check is to run "mpstat 1" for a few seconds when the
problem occurs.
In case need to know why system is overloaded when deconfig ppp interface.
Does it help if you disable conntrack?
Best regards,
Martin
quoted
On 11 Sep 2021, at 9:26, Martin Zaharinov [off-list ref] wrote:
Hi Guillaume
Main problem is overload of service because have many finishing ppp (customer) last two day down from 40-50 to 100-200 users and make problem when is happen if try to type : ip a wait 10-20 sec to start list interface .
But how to find where is a problem any locking or other.
And is there options to make fast remove ppp interface from kernel to reduce this load.
Martin
On 7 Sep 2021, at 9:16, Martin Zaharinov [off-list ref] wrote:
Hi
Sorry for delay but not easy to catch moment .
See this is mpstatl 1 :
Linux 5.14.1 (demobng) 09/07/21 _x86_64_ (12 CPU)
11:12:16 CPU %usr %nice %sys %iowait %irq %soft %steal %guest %gnice %idle
11:12:17 all 0.17 0.00 6.66 0.00 0.00 4.13 0.00 0.00 0.00 89.05
11:12:18 all 0.25 0.00 8.36 0.00 0.00 4.88 0.00 0.00 0.00 86.51
11:12:19 all 0.26 0.00 9.62 0.00 0.00 3.91 0.00 0.00 0.00 86.21
11:12:20 all 0.85 0.00 6.00 0.00 0.00 4.31 0.00 0.00 0.00 88.84
11:12:21 all 0.08 0.00 4.45 0.00 0.00 4.79 0.00 0.00 0.00 90.67
11:12:22 all 0.17 0.00 9.50 0.00 0.00 4.58 0.00 0.00 0.00 85.75
11:12:23 all 0.00 0.00 6.92 0.00 0.00 2.48 0.00 0.00 0.00 90.61
11:12:24 all 0.17 0.00 5.45 0.00 0.00 4.27 0.00 0.00 0.00 90.11
11:12:25 all 0.25 0.00 5.38 0.00 0.00 4.79 0.00 0.00 0.00 89.58
11:12:26 all 0.60 0.00 1.45 0.00 0.00 2.65 0.00 0.00 0.00 95.30
11:12:27 all 0.42 0.00 6.91 0.00 0.00 4.47 0.00 0.00 0.00 88.20
11:12:28 all 0.00 0.00 6.75 0.00 0.00 4.18 0.00 0.00 0.00 89.07
11:12:29 all 0.17 0.00 3.52 0.00 0.00 5.11 0.00 0.00 0.00 91.20
11:12:30 all 1.45 0.00 10.14 0.00 0.00 3.49 0.00 0.00 0.00 84.92
11:12:31 all 0.09 0.00 5.11 0.00 0.00 4.77 0.00 0.00 0.00 90.03
11:12:32 all 0.25 0.00 3.11 0.00 0.00 4.46 0.00 0.00 0.00 92.17
Average: all 0.32 0.00 6.21 0.00 0.00 4.21 0.00 0.00 0.00 89.26
I attache and one screenshot from perf top (Screenshot is send on preview mail)
And I see in lsmod
pppoe 20480 8198
pppox 16384 1 pppoe
ppp_generic 45056 16364 pppox,pppoe
slhc 16384 1 ppp_generic
To slow remove pppoe session .
And from log :
[2021-09-07 11:01:11.129] vlan3020: ebdd1c5d8b5900f6: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:01:53.621] vlan643: ebdd1c5d8b59014e: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:00.359] vlan1616: ebdd1c5d8b590195: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:05.859] vlan3020: ebdd1c5d8b5900d8: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:08.258] vlan3005: ebdd1c5d8b590190: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:13.820] vlan643: ebdd1c5d8b590152: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:15.839] vlan727: ebdd1c5d8b590144: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
[2021-09-07 11:02:20.139] vlan1693: ebdd1c5d8b59019f: ioctl(PPPIOCCONNECT): Transport endpoint is not connected
quoted
On 11 Aug 2021, at 19:48, Guillaume Nault [off-list ref] wrote:
On Wed, Aug 11, 2021 at 02:10:32PM +0300, Martin Zaharinov wrote:
quoted
And one more that see.
Problem is come when accel start finishing sessions,
Now in server have 2k users and restart on one of vlans 3 Olt with 400 users and affect other vlans ,
And problem is start when start destroying dead sessions from vlan with 3 Olt and this affect all other vlans.
May be kernel destroy old session slow and entrained other users by locking other sessions.
is there a way to speed up the closing of stopped/dead sessions.
What are the CPU stats when that happen? Is it users space or kernel
space that keeps it busy?
One easy way to check is to run "mpstat 1" for a few seconds when the
problem occurs.
And on time of problem when try to write : ip a
to list interface wait 15-20 sec i finaly have options to simulate but users is angry when down internet.
From: Martin Zaharinov <hidden> Date: 2021-09-14 10:01:38
Hi Nault and Florian
Nault :
No not test need conntrack to log user traffic.
Florian:
If you make patch send to test please.
Martin
On 14 Sep 2021, at 12:50, Florian Westphal [off-list ref] wrote:
Guillaume Nault [off-list ref] wrote:
quoted
quoted
And on time of problem when try to write : ip a
to list interface wait 15-20 sec i finaly have options to simulate but users is angry when down internet.
From: Martin Zaharinov <hidden> Date: 2021-09-14 10:53:26
Florian Hi
One more
please see i try to remove nf_nat and xt_MASQUERADE
and on time of problem need 50-80 sec to remove and overload system .
see perf from this moment:
PerfTop: 1738 irqs/sec kernel:85.0% exact: 100.0% lost: 0/0 drop: 0/0 [4000Hz cycles], (all, 12 CPUs)
---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
40.63% [nf_conntrack] [k] nf_ct_iterate_cleanup
21.23% [kernel] [k] __local_bh_enable_ip
10.93% [kernel] [k] __cond_resched
9.20% [kernel] [k] _raw_spin_lock
8.91% [kernel] [k] rcu_all_qs
5.83% [nf_conntrack] [k] nf_conntrack_lock
0.10% [kernel] [k] mutex_spin_on_owner
0.08% telegraf [.] 0x0000000000021bf0
0.06% [kernel] [k] osq_lock
0.06% [kernel] [k] kallsyms_expand_symbol.constprop.0
0.05% [kernel] [k] format_decode
0.04% [kernel] [k] rtnl_fill_ifinfo.constprop.0.isra.0
0.04% perf [.] 0x00000000000bc7b3
0.04% [kernel] [k] memcpy_erms
0.03% [kernel] [k] string
0.03% [kernel] [k] menu_select
0.03% [kernel] [k] nla_put
0.03% [kernel] [k] vsnprintf
Martin
On 14 Sep 2021, at 12:50, Florian Westphal [off-list ref] wrote:
Guillaume Nault [off-list ref] wrote:
quoted
quoted
And on time of problem when try to write : ip a
to list interface wait 15-20 sec i finaly have options to simulate but users is angry when down internet.