From: Thadeu Lima de Souza Cascardo <hidden> Date: 2021-03-17 18:55:22
During forced garbage collection, neighbours with more than a reference are
not removed. It's possible to DoS the neighbour table by using ARP spoofing
in such a way that there is always a timer pending for all neighbours,
preventing any of them from being removed. That will cause any new
neighbour creation to fail.
Use the same code as used by neigh_flush_dev, which deletes the timer and
cleans the queue when there are still references left.
With the same ARP spoofing technique, it was still possible to reach a valid
destination when this fix was applied, with no more table overflows.
Signed-off-by: Thadeu Lima de Souza Cascardo <redacted>
---
net/core/neighbour.c | 117 +++++++++++++++++++------------------------
1 file changed, 51 insertions(+), 66 deletions(-)
From: Thadeu Lima de Souza Cascardo <hidden> Date: 2021-03-17 18:55:22
IFF_POINTOPOINT interfaces use NUD_NOARP entries for IPv6. It's possible to
fill up the neighbour table with enough entries that it will overflow for
valid connections after that.
This behaviour is more prevalent after commit 58956317c8de ("neighbor:
Improve garbage collection") is applied, as it prevents removal from
entries that are not NUD_FAILED, unless they are more than 5s old.
Fixes: 58956317c8de (neighbor: Improve garbage collection)
Reported-by: Kasper Dupont <redacted>
Signed-off-by: Thadeu Lima de Souza Cascardo <redacted>
---
net/core/neighbour.c | 1 +
1 file changed, 1 insertion(+)
From: David Ahern <hidden> Date: 2021-03-17 23:42:57
On 3/17/21 12:53 PM, Thadeu Lima de Souza Cascardo wrote:
During forced garbage collection, neighbours with more than a reference are
not removed. It's possible to DoS the neighbour table by using ARP spoofing
in such a way that there is always a timer pending for all neighbours,
preventing any of them from being removed. That will cause any new
neighbour creation to fail.
Use the same code as used by neigh_flush_dev, which deletes the timer and
cleans the queue when there are still references left.
With the same ARP spoofing technique, it was still possible to reach a valid
destination when this fix was applied, with no more table overflows.
And how fast are neighbor entries removed with this patch? The current
code gives a neighbor entry a minimum lifetime to allow it to exist long
enough to be confirmed. Removing the minimum lifetime means neighbor
entries are constantly churning which is just as bad as the arp spoofing
problem.
quoted hunk
Signed-off-by: Thadeu Lima de Souza Cascardo <redacted>
---
net/core/neighbour.c | 117 +++++++++++++++++++------------------------
1 file changed, 51 insertions(+), 66 deletions(-)
From: Thadeu Lima de Souza Cascardo <hidden> Date: 2021-03-22 21:35:03
On Wed, Mar 17, 2021 at 05:42:00PM -0600, David Ahern wrote:
On 3/17/21 12:53 PM, Thadeu Lima de Souza Cascardo wrote:
quoted
During forced garbage collection, neighbours with more than a reference are
not removed. It's possible to DoS the neighbour table by using ARP spoofing
in such a way that there is always a timer pending for all neighbours,
preventing any of them from being removed. That will cause any new
neighbour creation to fail.
Use the same code as used by neigh_flush_dev, which deletes the timer and
cleans the queue when there are still references left.
With the same ARP spoofing technique, it was still possible to reach a valid
destination when this fix was applied, with no more table overflows.
And how fast are neighbor entries removed with this patch? The current
code gives a neighbor entry a minimum lifetime to allow it to exist long
enough to be confirmed. Removing the minimum lifetime means neighbor
entries are constantly churning which is just as bad as the arp spoofing
problem.
The patch should not change the rules of removing entries, so they are removed
only after 5 seconds. When trying to reach the other endpoint of a veth device,
it usually takes between 0 and 6 failures (with 1s interval) before succeeding,
and then, succeeding in succession.
I will be honest and say that I wasn't able yet to find out the exact order of
events in respect to neighbor state and lifetime updates, but the code still
only removes entries if they have been updated more than 5s before "now".
The change here is that entries are removed even if there is a reference to it
because of the timer. The timer is then removed and the entry can be removed.
I still see lots of table overflows with the tests I was able to run lately, so
not sure what changed in respect to when I first tested this patch. However,
then this patch is not applied, I cannot reach the other veth endpoint at all.
And, eventually, I loose remote access to the machine. With the patch applied,
the system is at least accessible, though it may "lag" sometimes, indicating
that access has been lost temporarily, but restored eventually.
Cascardo.
quoted
Signed-off-by: Thadeu Lima de Souza Cascardo <redacted>
---
net/core/neighbour.c | 117 +++++++++++++++++++------------------------
1 file changed, 51 insertions(+), 66 deletions(-)
On 17/03/21 15.53, Thadeu Lima de Souza Cascardo wrote:
quoted hunk
IFF_POINTOPOINT interfaces use NUD_NOARP entries for IPv6. It's possible to
fill up the neighbour table with enough entries that it will overflow for
valid connections after that.
This behaviour is more prevalent after commit 58956317c8de ("neighbor:
Improve garbage collection") is applied, as it prevents removal from
entries that are not NUD_FAILED, unless they are more than 5s old.
Fixes: 58956317c8de (neighbor: Improve garbage collection)
Reported-by: Kasper Dupont <redacted>
Signed-off-by: Thadeu Lima de Souza Cascardo <redacted>
---
net/core/neighbour.c | 1 +
1 file changed, 1 insertion(+)
@@ -256,6 +256,7 @@ static int neigh_forced_gc(struct neigh_table *tbl)write_lock(&n->lock);if((n->nud_state==NUD_FAILED)||+(n->nud_state==NUD_NOARP)||(tbl->is_multicast&&tbl->is_multicast(n->primary_key))||time_after(tref,n->updated))
--
2.27.0
Is there any update regarding this change?
I noticed this regression when it was used in a DoS attack on one of
my servers which I had upgraded from Ubuntu 18.04 to 20.04.
I have verified that Ubuntu 18.04 is not subject to this attack and
Ubuntu 20.04 is vulnerable. I have also verified that the one-line
change which Cascardo has provided fixes the vulnerability on Ubuntu
20.04.
Kind regards
Kasper
From: David Ahern <hidden> Date: 2021-04-19 17:10:19
On 4/19/21 9:44 AM, Kasper Dupont wrote:
On 17/03/21 15.53, Thadeu Lima de Souza Cascardo wrote:
quoted
IFF_POINTOPOINT interfaces use NUD_NOARP entries for IPv6. It's possible to
fill up the neighbour table with enough entries that it will overflow for
valid connections after that.
This behaviour is more prevalent after commit 58956317c8de ("neighbor:
Improve garbage collection") is applied, as it prevents removal from
entries that are not NUD_FAILED, unless they are more than 5s old.
Fixes: 58956317c8de (neighbor: Improve garbage collection)
Reported-by: Kasper Dupont <redacted>
Signed-off-by: Thadeu Lima de Souza Cascardo <redacted>
---
net/core/neighbour.c | 1 +
1 file changed, 1 insertion(+)
@@ -256,6 +256,7 @@ static int neigh_forced_gc(struct neigh_table *tbl)write_lock(&n->lock);if((n->nud_state==NUD_FAILED)||+(n->nud_state==NUD_NOARP)||(tbl->is_multicast&&tbl->is_multicast(n->primary_key))||time_after(tref,n->updated))
--
2.27.0
Is there any update regarding this change?
I noticed this regression when it was used in a DoS attack on one of
my servers which I had upgraded from Ubuntu 18.04 to 20.04.
I have verified that Ubuntu 18.04 is not subject to this attack and
Ubuntu 20.04 is vulnerable. I have also verified that the one-line
change which Cascardo has provided fixes the vulnerability on Ubuntu
20.04.
your testing included both patches or just this one?
On 17/03/21 15.53, Thadeu Lima de Souza Cascardo wrote:
quoted
IFF_POINTOPOINT interfaces use NUD_NOARP entries for IPv6. It's possible to
fill up the neighbour table with enough entries that it will overflow for
valid connections after that.
This behaviour is more prevalent after commit 58956317c8de ("neighbor:
Improve garbage collection") is applied, as it prevents removal from
entries that are not NUD_FAILED, unless they are more than 5s old.
Fixes: 58956317c8de (neighbor: Improve garbage collection)
Reported-by: Kasper Dupont <redacted>
Signed-off-by: Thadeu Lima de Souza Cascardo <redacted>
---
net/core/neighbour.c | 1 +
1 file changed, 1 insertion(+)
@@ -256,6 +256,7 @@ static int neigh_forced_gc(struct neigh_table *tbl)write_lock(&n->lock);if((n->nud_state==NUD_FAILED)||+(n->nud_state==NUD_NOARP)||(tbl->is_multicast&&tbl->is_multicast(n->primary_key))||time_after(tref,n->updated))
--
2.27.0
Is there any update regarding this change?
I noticed this regression when it was used in a DoS attack on one of
my servers which I had upgraded from Ubuntu 18.04 to 20.04.
I have verified that Ubuntu 18.04 is not subject to this attack and
Ubuntu 20.04 is vulnerable. I have also verified that the one-line
change which Cascardo has provided fixes the vulnerability on Ubuntu
20.04.
your testing included both patches or just this one?
I applied only this one line change on top of the kernel in Ubuntu
20.04. The behavior I observed was that without the patch the kernel
was vulnerable and with that patch I was unable to reproduce the
problem.
The other longer patch is for a different issue which Cascardo
discovered while working on the one I had reported. I don't have an
environment set up where I can reproduce the issue addressed by that
larger patch.
From: David Ahern <hidden> Date: 2021-04-20 04:27:20
On 4/19/21 10:52 AM, Kasper Dupont wrote:
On 19/04/21 10.10, David Ahern wrote:
quoted
On 4/19/21 9:44 AM, Kasper Dupont wrote:
quoted
Is there any update regarding this change?
I noticed this regression when it was used in a DoS attack on one of
my servers which I had upgraded from Ubuntu 18.04 to 20.04.
I have verified that Ubuntu 18.04 is not subject to this attack and
Ubuntu 20.04 is vulnerable. I have also verified that the one-line
change which Cascardo has provided fixes the vulnerability on Ubuntu
20.04.
your testing included both patches or just this one?
I applied only this one line change on top of the kernel in Ubuntu
20.04. The behavior I observed was that without the patch the kernel
was vulnerable and with that patch I was unable to reproduce the
problem.
This patch should be re-submitted standalone for -net
The other longer patch is for a different issue which Cascardo
discovered while working on the one I had reported. I don't have an
environment set up where I can reproduce the issue addressed by that
larger patch.