From: Jiri Pirko <redacted>
Ido says:
In kernel 4.9 the switchdev-specific FIB offload mechanism was replaced
by a new FIB notification chain to which modules could register in order
to be notified about the addition and deletion of FIB entries. The
motivation for this change was that switchdev drivers need to be able to
reflect the entire FIB table and not only FIBs configured on top of the
port netdevs themselves. This is useful in case of in-band management.
The fundamental problem with this approach is that upon registration
listeners lose all the information previously sent in the chain and
thus have an incomplete view of the FIB tables, which can result in
packet loss. This patchset fixes that by introducing a new API to dump
the FIB tables.
The entire dump process is done under RCU and thus the FIB notification
chain is converted to be atomic. The listeners are modified accordingly.
This is done in the first seven patches.
The eighth and ninth patches add a change sequence counter to ensure the
integrity of the FIB dump and a sysctl to set the number of retries,
respectively. The tenth patch finally introduces the FIB dump itself.
The last two patches modify current listeners of the FIB notification
chain to invoke the dump during their init.
v2->v3:
- Add sysctl to set the number of FIB dump retries
(Hannes Frederic Sowa).
- Read the sequence counter under RTNL to ensure synchronization
between the dump process and other processes changing the routing
tables (Hannes Frederic Sowa).
- Pass a callback to the dump function to be executed prior to a retry.
- Limit the dump to a single net namespace.
v1->v2:
- Add a sequence counter to ensure the integrity of the FIB dump
(David S. Miller, Hannes Frederic Sowa).
- Protect notifications from re-ordering in listeners by using an
ordered workqueue (Hannes Frederic Sowa).
- Introduce fib_info_hold() (Jiri Pirko).
- Relieve rocker from the need to invoke the FIB dump by registering
to the FIB notification chain prior to ports creation.
Ido Schimmel (12):
ipv4: fib: Export free_fib_info()
ipv4: fib: Add fib_info_hold() helper
mlxsw: core: Create an ordered workqueue for FIB offload
mlxsw: spectrum_router: Implement FIB offload in deferred work
rocker: Create an ordered workqueue for FIB offload
rocker: Implement FIB offload in deferred work
ipv4: fib: Convert FIB notification chain to be atomic
ipv4: fib: Allow for consistent FIB dumping
ipv4: fib: Add sysctl to limit number of FIB dump retries
ipv4: fib: Add an API to request a FIB dump
mlxsw: spectrum_router: Request a dump of FIB tables during init
rocker: Register FIB notifier before creating ports
Documentation/networking/ip-sysctl.txt | 8 ++
drivers/net/ethernet/mellanox/mlxsw/core.c | 22 +++
drivers/net/ethernet/mellanox/mlxsw/core.h | 2 +
.../net/ethernet/mellanox/mlxsw/spectrum_router.c | 95 +++++++++++--
drivers/net/ethernet/rocker/rocker.h | 1 +
drivers/net/ethernet/rocker/rocker_main.c | 78 +++++++++--
drivers/net/ethernet/rocker/rocker_ofdpa.c | 1 +
include/net/ip_fib.h | 9 ++
include/net/netns/ipv4.h | 4 +
net/ipv4/fib_frontend.c | 3 +
net/ipv4/fib_semantics.c | 1 +
net/ipv4/fib_trie.c | 147 ++++++++++++++++++++-
net/ipv4/sysctl_net_ipv4.c | 7 +
13 files changed, 352 insertions(+), 26 deletions(-)
--
2.7.4
From: Ido Schimmel <redacted>
The FIB notification chain is going to be converted to an atomic chain,
which means switchdev drivers will have to offload FIB entries in
deferred work, as hardware operations entail sleeping.
However, while the work is queued fib info might be freed, so a
reference must be taken. To release the reference (and potentially free
the fib info) fib_info_put() will be called, which in turn calls
free_fib_info().
Export free_fib_info() so that modules will be able to invoke
fib_info_put().
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
net/ipv4/fib_semantics.c | 1 +
1 file changed, 1 insertion(+)
From: Ido Schimmel <redacted>
As explained in the previous commit, modules are going to need to take a
reference on fib info and then drop it using fib_info_put().
Add the fib_info_hold() helper to make the code more readable and also
symmetric with fib_info_put().
Signed-off-by: Ido Schimmel <redacted>
Suggested-by: Jiri Pirko <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
include/net/ip_fib.h | 5 +++++
1 file changed, 5 insertions(+)
From: Ido Schimmel <redacted>
We're going to start processing FIB entries addition / deletion events
in deferred work. These work items must be processed in the order they
were submitted or otherwise we can have differences between the kernel's
FIB table and the device's.
Solve this by creating an ordered workqueue to which these work items
will be submitted to. Note that we can't simply convert the current
workqueue to be ordered, as EMADs re-transmissions are also processed in
deferred work.
Later on, we can migrate other work items to this workqueue, such as FDB
notification processing and nexthop resolution, since they all take the
same lock anyway.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
drivers/net/ethernet/mellanox/mlxsw/core.c | 22 ++++++++++++++++++++++
drivers/net/ethernet/mellanox/mlxsw/core.h | 2 ++
2 files changed, 24 insertions(+)
From: Ido Schimmel <redacted>
FIB offload is currently done in process context with RTNL held, but
we're about to dump the FIB tables in RCU critical section, so we can no
longer sleep.
Instead, defer the operation to process context using deferred work. Make
sure fib info isn't freed while the work is queued by taking a reference
on it and releasing it after the operation is done.
Deferring the operation is valid because the upper layers always assume
the operation was successful. If it's not, then the driver-specific
abort mechanism is called and all routed traffic is directed to slow
path.
The work items are submitted to an ordered workqueue to prevent a
mismatch between the kernel's FIB table and the device's.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
.../net/ethernet/mellanox/mlxsw/spectrum_router.c | 72 +++++++++++++++++++---
1 file changed, 62 insertions(+), 10 deletions(-)
@@ -593,6 +593,14 @@ static void mlxsw_sp_router_fib_flush(struct mlxsw_sp *mlxsw_sp);staticvoidmlxsw_sp_vrs_fini(structmlxsw_sp*mlxsw_sp){+/* At this stage we're guaranteed not to have new incoming+*FIBnotificationsandtheworkqueueisfreefromFIBs+*sittingontopofmlxswnetdevs.However,wecanstill+*haveotherFIBsqueued.Flushthequeuebeforeflushing+*thedevice'stables.Noneedforlocks,aswe'retheonly+*writer.+*/+mlxsw_core_flush_owq();mlxsw_sp_router_fib_flush(mlxsw_sp);kfree(mlxsw_sp->router.vrs);}
@@ -1948,30 +1956,74 @@ static void __mlxsw_sp_router_fini(struct mlxsw_sp *mlxsw_sp)kfree(mlxsw_sp->rifs);}-staticintmlxsw_sp_router_fib_event(structnotifier_block*nb,-unsignedlongevent,void*ptr)+structmlxsw_sp_fib_event_work{+structdelayed_workdw;+structfib_entry_notifier_infofen_info;+structmlxsw_sp*mlxsw_sp;+unsignedlongevent;+};++staticvoidmlxsw_sp_router_fib_event_work(structwork_struct*work){-structmlxsw_sp*mlxsw_sp=container_of(nb,structmlxsw_sp,fib_nb);-structfib_entry_notifier_info*fen_info=ptr;+structmlxsw_sp_fib_event_work*fib_work=+container_of(work,structmlxsw_sp_fib_event_work,dw.work);+structmlxsw_sp*mlxsw_sp=fib_work->mlxsw_sp;interr;-if(!net_eq(fen_info->info.net,&init_net))-returnNOTIFY_DONE;--switch(event){+/* Protect internal structures from changes */+rtnl_lock();+switch(fib_work->event){caseFIB_EVENT_ENTRY_ADD:-err=mlxsw_sp_router_fib4_add(mlxsw_sp,fen_info);+err=mlxsw_sp_router_fib4_add(mlxsw_sp,&fib_work->fen_info);if(err)mlxsw_sp_router_fib4_abort(mlxsw_sp);+fib_info_put(fib_work->fen_info.fi);break;caseFIB_EVENT_ENTRY_DEL:-mlxsw_sp_router_fib4_del(mlxsw_sp,fen_info);+mlxsw_sp_router_fib4_del(mlxsw_sp,&fib_work->fen_info);+fib_info_put(fib_work->fen_info.fi);break;caseFIB_EVENT_RULE_ADD:/* fall through */caseFIB_EVENT_RULE_DEL:mlxsw_sp_router_fib4_abort(mlxsw_sp);break;}+rtnl_unlock();+kfree(fib_work);+}++/* Called with rcu_read_lock() */+staticintmlxsw_sp_router_fib_event(structnotifier_block*nb,+unsignedlongevent,void*ptr)+{+structmlxsw_sp*mlxsw_sp=container_of(nb,structmlxsw_sp,fib_nb);+structmlxsw_sp_fib_event_work*fib_work;+structfib_notifier_info*info=ptr;++if(!net_eq(info->net,&init_net))+returnNOTIFY_DONE;++fib_work=kzalloc(sizeof(*fib_work),GFP_ATOMIC);+if(WARN_ON(!fib_work))+returnNOTIFY_BAD;++INIT_DELAYED_WORK(&fib_work->dw,mlxsw_sp_router_fib_event_work);+fib_work->mlxsw_sp=mlxsw_sp;+fib_work->event=event;++switch(event){+caseFIB_EVENT_ENTRY_ADD:/* fall through */+caseFIB_EVENT_ENTRY_DEL:+memcpy(&fib_work->fen_info,ptr,sizeof(fib_work->fen_info));+/* Take referece on fib_info to prevent it from being+*freedwhileworkisqueued.Releaseitafterwards.+*/+fib_info_hold(fib_work->fen_info.fi);+break;+}++mlxsw_core_schedule_odw(&fib_work->dw,0);+returnNOTIFY_DONE;}
From: Ido Schimmel <redacted>
As explained in the previous commits, we need to process FIB entries
addition / deletion events in FIFO order or otherwise we can have a
mismatch between the kernel's FIB table and the device's.
Create an ordered workqueue for rocker to which these work items will be
submitted to.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
drivers/net/ethernet/rocker/rocker.h | 1 +
drivers/net/ethernet/rocker/rocker_main.c | 11 +++++++++++
2 files changed, 12 insertions(+)
From: Ido Schimmel <redacted>
Convert rocker to offload FIBs in deferred work in a similar fashion to
mlxsw, which was converted in the previous commits.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
drivers/net/ethernet/rocker/rocker_main.c | 58 +++++++++++++++++++++++++-----
drivers/net/ethernet/rocker/rocker_ofdpa.c | 1 +
2 files changed, 51 insertions(+), 8 deletions(-)
@@ -2166,28 +2166,70 @@ static const struct switchdev_ops rocker_port_switchdev_ops = {.switchdev_port_obj_dump=rocker_port_obj_dump,};-staticintrocker_router_fib_event(structnotifier_block*nb,-unsignedlongevent,void*ptr)+structrocker_fib_event_work{+structwork_structwork;+structfib_entry_notifier_infofen_info;+structrocker*rocker;+unsignedlongevent;+};++staticvoidrocker_router_fib_event_work(structwork_struct*work){-structrocker*rocker=container_of(nb,structrocker,fib_nb);-structfib_entry_notifier_info*fen_info=ptr;+structrocker_fib_event_work*fib_work=+container_of(work,structrocker_fib_event_work,work);+structrocker*rocker=fib_work->rocker;interr;-switch(event){+/* Protect internal structures from changes */+rtnl_lock();+switch(fib_work->event){caseFIB_EVENT_ENTRY_ADD:-err=rocker_world_fib4_add(rocker,fen_info);+err=rocker_world_fib4_add(rocker,&fib_work->fen_info);if(err)rocker_world_fib4_abort(rocker);-else+fib_info_put(fib_work->fen_info.fi);break;caseFIB_EVENT_ENTRY_DEL:-rocker_world_fib4_del(rocker,fen_info);+rocker_world_fib4_del(rocker,&fib_work->fen_info);+fib_info_put(fib_work->fen_info.fi);break;caseFIB_EVENT_RULE_ADD:/* fall through */caseFIB_EVENT_RULE_DEL:rocker_world_fib4_abort(rocker);break;}+rtnl_unlock();+kfree(fib_work);+}++/* Called with rcu_read_lock() */+staticintrocker_router_fib_event(structnotifier_block*nb,+unsignedlongevent,void*ptr)+{+structrocker*rocker=container_of(nb,structrocker,fib_nb);+structrocker_fib_event_work*fib_work;++fib_work=kzalloc(sizeof(*fib_work),GFP_ATOMIC);+if(WARN_ON(!fib_work))+returnNOTIFY_BAD;++INIT_WORK(&fib_work->work,rocker_router_fib_event_work);+fib_work->rocker=rocker;+fib_work->event=event;++switch(event){+caseFIB_EVENT_ENTRY_ADD:/* fall through */+caseFIB_EVENT_ENTRY_DEL:+memcpy(&fib_work->fen_info,ptr,sizeof(fib_work->fen_info));+/* Take referece on fib_info to prevent it from being+*freedwhileworkisqueued.Releaseitafterwards.+*/+fib_info_hold(fib_work->fen_info.fi);+break;+}++queue_work(rocker->rocker_owq,&fib_work->work);+returnNOTIFY_DONE;}
From: Ido Schimmel <redacted>
In order not to hold RTNL for long periods of time we're going to dump
the FIB tables using RCU.
Convert the FIB notification chain to be atomic, as we can't block in
RCU critical sections.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
net/ipv4/fib_trie.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
From: Ido Schimmel <redacted>
The next patch will enable listeners of the FIB notification chain to
request a dump of the FIB tables. However, since RTNL isn't taken during
the dump, it's possible for the FIB tables to change mid-dump, which
will result in inconsistency between the listener's table and the
kernel's.
Allow listeners to know about changes that occurred mid-dump, by adding
a change sequence counter to each net namespace. The counter is
incremented just before a notification is sent in the FIB chain.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
include/net/netns/ipv4.h | 3 +++
net/ipv4/fib_frontend.c | 2 ++
net/ipv4/fib_trie.c | 1 +
3 files changed, 6 insertions(+)
@@ -1219,6 +1219,8 @@ static int __net_init ip_fib_net_init(struct net *net)interr;size_tsize=sizeof(structhlist_head)*FIB_TABLE_HASHSZ;+net->ipv4.fib_seq=0;+/* Avoid false sharing : Use at least a full cache line */size=max_t(size_t,size,L1_CACHE_BYTES);
From: Ido Schimmel <redacted>
When dumping the FIB tables in the next commit, the dump will be
considered invalid if notifications were sent in the FIB notification
chain mid-dump. In systems where routing changes are frequent, the dump
might need to be restarted multiple times.
Add sysctl to limit the number of FIB dump retries, thereby preventing
callers from looping for long periods of time.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
Documentation/networking/ip-sysctl.txt | 8 ++++++++
include/net/netns/ipv4.h | 1 +
net/ipv4/fib_frontend.c | 1 +
net/ipv4/sysctl_net_ipv4.c | 7 +++++++
4 files changed, 17 insertions(+)
@@ -73,6 +73,14 @@ fib_multipath_use_neigh - BOOLEAN 0 - disabled 1 - enabled+fib_dump_max_retries - INTEGER+ Maximum number of retries until the FIB dump is aborted. For a+ given net namespace, a FIB dump is considered invalid if+ notifications were sent in the FIB notification chain mid-dump.+ The dump will be retried until it is successful or maximum+ number of retries has been reached.+ Default: 5+ route/max_size - INTEGER Maximum number of routes allowed in the kernel. Increase this when using large numbers of interfaces and/or routes.
@@ -1219,6 +1219,7 @@ static int __net_init ip_fib_net_init(struct net *net)interr;size_tsize=sizeof(structhlist_head)*FIB_TABLE_HASHSZ;+net->ipv4.sysctl_fib_dump_max_retries=5;net->ipv4.fib_seq=0;/* Avoid false sharing : Use at least a full cache line */
From: Ido Schimmel <redacted>
Commit b90eb7549499 ("fib: introduce FIB notification infrastructure")
introduced a new notification chain to notify listeners (f.e., switchdev
drivers) about addition and deletion of routes.
However, upon registration to the chain the FIB tables can already be
populated, which means potential listeners will have an incomplete view
of the tables.
Solve that by adding an API to request a FIB dump. The dump itself it
done using RCU in order not to starve consumers that need RTNL to make
progress.
The integrity of the dump is ensured by reading the FIB change sequence
counter before and after the dump. This allows us to avoid the
problematic situation in which the dumping process sends a ENTRY_ADD
notification following ENTRY_DEL generated by another process holding
RTNL.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
include/net/ip_fib.h | 4 ++
net/ipv4/fib_trie.c | 138 +++++++++++++++++++++++++++++++++++++++++++++++++++
2 files changed, 142 insertions(+)
@@ -1902,6 +1984,62 @@ int fib_table_flush(struct net *net, struct fib_table *tb)returnfound;}+staticvoidfib_leaf_notify(structnet*net,structkey_vector*l,+structfib_table*tb,structnotifier_block*nb,+enumfib_event_typeevent_type)+{+structfib_alias*fa;++hlist_for_each_entry_rcu(fa,&l->leaf,fa_list){+structfib_info*fi=fa->fa_info;++if(!fi)+continue;++/* local and main table can share the same trie,+*sodon'tnotifytwiceforthesameentry.+*/+if(tb->tb_id!=fa->tb_id)+continue;++call_fib_entry_notifier(nb,net,event_type,l->key,+KEYLENGTH-fa->fa_slen,fi,fa->fa_tos,+fa->fa_type,fa->tb_id,0);+}+}++staticvoidfib_table_notify(structnet*net,structfib_table*tb,+structnotifier_block*nb,+enumfib_event_typeevent_type)+{+structtrie*t=(structtrie*)tb->tb_data;+structkey_vector*l,*tp=t->kv;+t_keykey=0;++while((l=leaf_walk_rcu(&tp,key))!=NULL){+fib_leaf_notify(net,l,tb,nb,event_type);++key=l->key+1;+/* stop in case of wrap around */+if(key<l->key)+break;+}+}++staticvoidfib_notify(structnet*net,structnotifier_block*nb,+enumfib_event_typeevent_type)+{+unsignedinth;++for(h=0;h<FIB_TABLE_HASHSZ;h++){+structhlist_head*head=&net->ipv4.fib_table_hash[h];+structfib_table*tb;++hlist_for_each_entry_rcu(tb,head,tb_hlist)+fib_table_notify(net,tb,nb,event_type);+}+}+staticvoid__trie_free_rcu(structrcu_head*head){structfib_table*tb=container_of(head,structfib_table,rcu);
From: Ido Schimmel <redacted>
Make sure the device has a complete view of the FIB tables by invoking
their dump during module init.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
.../net/ethernet/mellanox/mlxsw/spectrum_router.c | 23 ++++++++++++++++++++++
1 file changed, 23 insertions(+)
@@ -2027,8 +2027,23 @@ static int mlxsw_sp_router_fib_event(struct notifier_block *nb,returnNOTIFY_DONE;}+staticvoidmlxsw_sp_router_fib_dump_flush(structnotifier_block*nb)+{+structmlxsw_sp*mlxsw_sp=container_of(nb,structmlxsw_sp,fib_nb);++/* Flush pending FIB notifications and then flush the device's+*tablebeforerequestinganotherdump.DothatwithRTNLheld,+*asFIBnotificationblockisalreadyregistered.+*/+mlxsw_core_flush_owq();+rtnl_lock();+mlxsw_sp_router_fib_flush(mlxsw_sp);+rtnl_unlock();+}+intmlxsw_sp_router_init(structmlxsw_sp*mlxsw_sp){+fib_dump_cb_t*cb=mlxsw_sp_router_fib_dump_flush;interr;INIT_LIST_HEAD(&mlxsw_sp->router.nexthop_neighs_list);
@@ -2048,8 +2063,16 @@ int mlxsw_sp_router_init(struct mlxsw_sp *mlxsw_sp)mlxsw_sp->fib_nb.notifier_call=mlxsw_sp_router_fib_event;register_fib_notifier(&mlxsw_sp->fib_nb);+if(!fib_notifier_dump(&mlxsw_sp->fib_nb,&init_net,cb)){+err=-EBUSY;+gotoerr_fib_notifier_dump;+}+return0;+err_fib_notifier_dump:+unregister_fib_notifier(&mlxsw_sp->fib_nb);+mlxsw_sp_neigh_fini(mlxsw_sp);err_neigh_init:mlxsw_sp_vrs_fini(mlxsw_sp);err_vrs_init:
From: Ido Schimmel <redacted>
Unlike mlxsw, rocker only supports the reflection of routes pointing to
its own netdevs. Therefore, instead of requesting a FIB dump during
init, simply register the FIB notifier before creating the ports.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
drivers/net/ethernet/rocker/rocker_main.c | 9 +++++----
1 file changed, 5 insertions(+), 4 deletions(-)
From: Hannes Frederic Sowa <hidden> Date: 2016-11-30 15:37:57
On 30.11.2016 11:09, Jiri Pirko wrote:
quoted hunk
From: Ido Schimmel <redacted>
Make sure the device has a complete view of the FIB tables by invoking
their dump during module init.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
.../net/ethernet/mellanox/mlxsw/spectrum_router.c | 23 ++++++++++++++++++++++
1 file changed, 23 insertions(+)
@@ -2027,8 +2027,23 @@ static int mlxsw_sp_router_fib_event(struct notifier_block *nb,returnNOTIFY_DONE;}+staticvoidmlxsw_sp_router_fib_dump_flush(structnotifier_block*nb)+{+structmlxsw_sp*mlxsw_sp=container_of(nb,structmlxsw_sp,fib_nb);++/* Flush pending FIB notifications and then flush the device's+*tablebeforerequestinganotherdump.DothatwithRTNLheld,+*asFIBnotificationblockisalreadyregistered.+*/+mlxsw_core_flush_owq();+rtnl_lock();+mlxsw_sp_router_fib_flush(mlxsw_sp);+rtnl_unlock();+}+intmlxsw_sp_router_init(structmlxsw_sp*mlxsw_sp){+fib_dump_cb_t*cb=mlxsw_sp_router_fib_dump_flush;interr;INIT_LIST_HEAD(&mlxsw_sp->router.nexthop_neighs_list);
@@ -2048,8 +2063,16 @@ int mlxsw_sp_router_init(struct mlxsw_sp *mlxsw_sp)mlxsw_sp->fib_nb.notifier_call=mlxsw_sp_router_fib_event;register_fib_notifier(&mlxsw_sp->fib_nb);
Sorry to pick in here again:
There is a race here. You need to protect the registration of the fib
notifier as well by the sequence counter. Updates here are not ordered
in relation to this code below.
I think just move the register notification into the fib_notifier_dump
function, rename it to fib_notifier_init and use it here:
Hi Hannes,
On Wed, Nov 30, 2016 at 04:37:48PM +0100, Hannes Frederic Sowa wrote:
On 30.11.2016 11:09, Jiri Pirko wrote:
quoted
From: Ido Schimmel <redacted>
Make sure the device has a complete view of the FIB tables by invoking
their dump during module init.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
.../net/ethernet/mellanox/mlxsw/spectrum_router.c | 23 ++++++++++++++++++++++
1 file changed, 23 insertions(+)
@@ -2027,8 +2027,23 @@ static int mlxsw_sp_router_fib_event(struct notifier_block *nb,returnNOTIFY_DONE;}+staticvoidmlxsw_sp_router_fib_dump_flush(structnotifier_block*nb)+{+structmlxsw_sp*mlxsw_sp=container_of(nb,structmlxsw_sp,fib_nb);++/* Flush pending FIB notifications and then flush the device's+*tablebeforerequestinganotherdump.DothatwithRTNLheld,+*asFIBnotificationblockisalreadyregistered.+*/+mlxsw_core_flush_owq();+rtnl_lock();+mlxsw_sp_router_fib_flush(mlxsw_sp);+rtnl_unlock();+}+intmlxsw_sp_router_init(structmlxsw_sp*mlxsw_sp){+fib_dump_cb_t*cb=mlxsw_sp_router_fib_dump_flush;interr;INIT_LIST_HEAD(&mlxsw_sp->router.nexthop_neighs_list);
@@ -2048,8 +2063,16 @@ int mlxsw_sp_router_init(struct mlxsw_sp *mlxsw_sp)mlxsw_sp->fib_nb.notifier_call=mlxsw_sp_router_fib_event;register_fib_notifier(&mlxsw_sp->fib_nb);
Sorry to pick in here again:
There is a race here. You need to protect the registration of the fib
notifier as well by the sequence counter. Updates here are not ordered
in relation to this code below.
You mean updates that can be received after you registered the notifier
and until the dump started? I'm aware of that and that's OK. This
listener should be able to handle duplicates.
I've a follow up patchset that introduces a new event in switchdev
notification chain called SWITCHDEV_SYNC, which is sent when port
netdevs are enslaved / released from a master device (points in time
where kernel<->device can get out of sync). It will invoke
re-propagation of configuration from different parts of the stack
(e.g. bridge driver, 8021q driver, fib/neigh code), which can result
in duplicates.
I think just move the register notification into the fib_notifier_dump
function, rename it to fib_notifier_init and use it here:
I separated the two on purpose. For example, rocker only needs to
register notifier, but doesn't need the dump.
From: Hannes Frederic Sowa <hidden> Date: 2016-11-30 16:50:42
On 30.11.2016 17:32, Ido Schimmel wrote:
Hi Hannes,
On Wed, Nov 30, 2016 at 04:37:48PM +0100, Hannes Frederic Sowa wrote:
quoted
On 30.11.2016 11:09, Jiri Pirko wrote:
quoted
From: Ido Schimmel <redacted>
Make sure the device has a complete view of the FIB tables by invoking
their dump during module init.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
.../net/ethernet/mellanox/mlxsw/spectrum_router.c | 23 ++++++++++++++++++++++
1 file changed, 23 insertions(+)
@@ -2027,8 +2027,23 @@ static int mlxsw_sp_router_fib_event(struct notifier_block *nb,returnNOTIFY_DONE;}+staticvoidmlxsw_sp_router_fib_dump_flush(structnotifier_block*nb)+{+structmlxsw_sp*mlxsw_sp=container_of(nb,structmlxsw_sp,fib_nb);++/* Flush pending FIB notifications and then flush the device's+*tablebeforerequestinganotherdump.DothatwithRTNLheld,+*asFIBnotificationblockisalreadyregistered.+*/+mlxsw_core_flush_owq();+rtnl_lock();+mlxsw_sp_router_fib_flush(mlxsw_sp);+rtnl_unlock();+}+intmlxsw_sp_router_init(structmlxsw_sp*mlxsw_sp){+fib_dump_cb_t*cb=mlxsw_sp_router_fib_dump_flush;interr;INIT_LIST_HEAD(&mlxsw_sp->router.nexthop_neighs_list);
@@ -2048,8 +2063,16 @@ int mlxsw_sp_router_init(struct mlxsw_sp *mlxsw_sp)mlxsw_sp->fib_nb.notifier_call=mlxsw_sp_router_fib_event;register_fib_notifier(&mlxsw_sp->fib_nb);
Sorry to pick in here again:
There is a race here. You need to protect the registration of the fib
notifier as well by the sequence counter. Updates here are not ordered
in relation to this code below.
You mean updates that can be received after you registered the notifier
and until the dump started? I'm aware of that and that's OK. This
listener should be able to handle duplicates.
I am not concerned about duplicates, but about ordering deletes and
getting an add from the RCU code you will add the node to hw while it is
deleted in the software path. You probably will ignore the delete
because nothing is installed in hw and later add the node which was
actually deleted but just reordered which happend on another CPU, no?
I've a follow up patchset that introduces a new event in switchdev
notification chain called SWITCHDEV_SYNC, which is sent when port
netdevs are enslaved / released from a master device (points in time
where kernel<->device can get out of sync). It will invoke
re-propagation of configuration from different parts of the stack
(e.g. bridge driver, 8021q driver, fib/neigh code), which can result
in duplicates.
Okay, understood. I wonder how we can protect against accidentally abort
calls actually. E.g. if I start to inject routes into my routing domain
how can I make sure the box doesn't die after I try to insert enough
routes. Do we need to touch quagga etc?
Thanks,
Hannes
On Wed, Nov 30, 2016 at 05:49:56PM +0100, Hannes Frederic Sowa wrote:
On 30.11.2016 17:32, Ido Schimmel wrote:
quoted
On Wed, Nov 30, 2016 at 04:37:48PM +0100, Hannes Frederic Sowa wrote:
quoted
On 30.11.2016 11:09, Jiri Pirko wrote:
quoted
From: Ido Schimmel <redacted>
Make sure the device has a complete view of the FIB tables by invoking
their dump during module init.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
.../net/ethernet/mellanox/mlxsw/spectrum_router.c | 23 ++++++++++++++++++++++
1 file changed, 23 insertions(+)
@@ -2027,8 +2027,23 @@ static int mlxsw_sp_router_fib_event(struct notifier_block *nb,returnNOTIFY_DONE;}+staticvoidmlxsw_sp_router_fib_dump_flush(structnotifier_block*nb)+{+structmlxsw_sp*mlxsw_sp=container_of(nb,structmlxsw_sp,fib_nb);++/* Flush pending FIB notifications and then flush the device's+*tablebeforerequestinganotherdump.DothatwithRTNLheld,+*asFIBnotificationblockisalreadyregistered.+*/+mlxsw_core_flush_owq();+rtnl_lock();+mlxsw_sp_router_fib_flush(mlxsw_sp);+rtnl_unlock();+}+intmlxsw_sp_router_init(structmlxsw_sp*mlxsw_sp){+fib_dump_cb_t*cb=mlxsw_sp_router_fib_dump_flush;interr;INIT_LIST_HEAD(&mlxsw_sp->router.nexthop_neighs_list);
@@ -2048,8 +2063,16 @@ int mlxsw_sp_router_init(struct mlxsw_sp *mlxsw_sp)mlxsw_sp->fib_nb.notifier_call=mlxsw_sp_router_fib_event;register_fib_notifier(&mlxsw_sp->fib_nb);
Sorry to pick in here again:
There is a race here. You need to protect the registration of the fib
notifier as well by the sequence counter. Updates here are not ordered
in relation to this code below.
You mean updates that can be received after you registered the notifier
and until the dump started? I'm aware of that and that's OK. This
listener should be able to handle duplicates.
I am not concerned about duplicates, but about ordering deletes and
getting an add from the RCU code you will add the node to hw while it is
deleted in the software path. You probably will ignore the delete
because nothing is installed in hw and later add the node which was
actually deleted but just reordered which happend on another CPU, no?
Are you referring to reordering in the workqueue? We already covered
this using an ordered workqueue, which has one context of execution
system-wide.
quoted
I've a follow up patchset that introduces a new event in switchdev
notification chain called SWITCHDEV_SYNC, which is sent when port
netdevs are enslaved / released from a master device (points in time
where kernel<->device can get out of sync). It will invoke
re-propagation of configuration from different parts of the stack
(e.g. bridge driver, 8021q driver, fib/neigh code), which can result
in duplicates.
Okay, understood. I wonder how we can protect against accidentally abort
calls actually. E.g. if I start to inject routes into my routing domain
how can I make sure the box doesn't die after I try to insert enough
routes. Do we need to touch quagga etc?
The whole point of moving abort mechanism to the driver is that the
system won't die, but instead routing will be done in the kernel. If you
respect hardware limitations, then there's no reason for abort mechanism
to kick in.
From: David Miller <davem@davemloft.net> Date: 2016-12-01 20:04:48
Hannes and Ido,
It looks like we are very close to having this in mergable shape, can
you guys work out this final issue and figure out if it really is
a merge stopped or not?
Thanks.
From: Hannes Frederic Sowa <hidden> Date: 2016-12-01 20:40:53
On 01.12.2016 21:04, David Miller wrote:
Hannes and Ido,
It looks like we are very close to having this in mergable shape, can
you guys work out this final issue and figure out if it really is
a merge stopped or not?
Sure, if the fib notification register could be done under protection of
the sequence counter I don't see any more problems.
The sync handler is nice to have and can be done in a later patch series.
On Thu, Dec 01, 2016 at 09:40:48PM +0100, Hannes Frederic Sowa wrote:
On 01.12.2016 21:04, David Miller wrote:
quoted
Hannes and Ido,
It looks like we are very close to having this in mergable shape, can
you guys work out this final issue and figure out if it really is
a merge stopped or not?
Sure, if the fib notification register could be done under protection of
the sequence counter I don't see any more problems.
Did you maybe miss my reply yesterday? Because I was trying to
understand what "ordering" you're referring to, but didn't receive a
reply from you.
The sync handler is nice to have and can be done in a later patch series.
From: Hannes Frederic Sowa <hidden> Date: 2016-12-01 21:10:32
On 30.11.2016 17:32, Ido Schimmel wrote:
Hi Hannes,
On Wed, Nov 30, 2016 at 04:37:48PM +0100, Hannes Frederic Sowa wrote:
quoted
On 30.11.2016 11:09, Jiri Pirko wrote:
quoted
From: Ido Schimmel <redacted>
Make sure the device has a complete view of the FIB tables by invoking
their dump during module init.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
.../net/ethernet/mellanox/mlxsw/spectrum_router.c | 23 ++++++++++++++++++++++
1 file changed, 23 insertions(+)
@@ -2027,8 +2027,23 @@ static int mlxsw_sp_router_fib_event(struct notifier_block *nb,returnNOTIFY_DONE;}+staticvoidmlxsw_sp_router_fib_dump_flush(structnotifier_block*nb)+{+structmlxsw_sp*mlxsw_sp=container_of(nb,structmlxsw_sp,fib_nb);++/* Flush pending FIB notifications and then flush the device's+*tablebeforerequestinganotherdump.DothatwithRTNLheld,+*asFIBnotificationblockisalreadyregistered.+*/+mlxsw_core_flush_owq();+rtnl_lock();+mlxsw_sp_router_fib_flush(mlxsw_sp);+rtnl_unlock();+}+intmlxsw_sp_router_init(structmlxsw_sp*mlxsw_sp){+fib_dump_cb_t*cb=mlxsw_sp_router_fib_dump_flush;interr;INIT_LIST_HEAD(&mlxsw_sp->router.nexthop_neighs_list);
@@ -2048,8 +2063,16 @@ int mlxsw_sp_router_init(struct mlxsw_sp *mlxsw_sp)mlxsw_sp->fib_nb.notifier_call=mlxsw_sp_router_fib_event;register_fib_notifier(&mlxsw_sp->fib_nb);
Sorry to pick in here again:
There is a race here. You need to protect the registration of the fib
notifier as well by the sequence counter. Updates here are not ordered
in relation to this code below.
You mean updates that can be received after you registered the notifier
and until the dump started? I'm aware of that and that's OK. This
listener should be able to handle duplicates.
I am not concerned about duplicates, but about ordering deletes and
getting an add from the RCU code you will add the node to hw while it is
deleted in the software path. You probably will ignore the delete
because nothing is installed in hw and later add the node which was
actually deleted but just reordered which happend on another CPU, no?
I've a follow up patchset that introduces a new event in switchdev
notification chain called SWITCHDEV_SYNC, which is sent when port
netdevs are enslaved / released from a master device (points in time
where kernel<->device can get out of sync). It will invoke
re-propagation of configuration from different parts of the stack
(e.g. bridge driver, 8021q driver, fib/neigh code), which can result
in duplicates.
Okay, understood. I wonder how we can protect against accidentally abort
calls actually. E.g. if I start to inject routes into my routing domain
how can I make sure the box doesn't die after I try to insert enough
routes. Do we need to touch quagga etc?
Thanks,
Hannes
From: Hannes Frederic Sowa <hidden> Date: 2016-12-01 21:18:37
On 01.12.2016 21:54, Ido Schimmel wrote:
On Thu, Dec 01, 2016 at 09:40:48PM +0100, Hannes Frederic Sowa wrote:
quoted
On 01.12.2016 21:04, David Miller wrote:
quoted
Hannes and Ido,
It looks like we are very close to having this in mergable shape, can
you guys work out this final issue and figure out if it really is
a merge stopped or not?
Sure, if the fib notification register could be done under protection of
the sequence counter I don't see any more problems.
Did you maybe miss my reply yesterday? Because I was trying to
understand what "ordering" you're referring to, but didn't receive a
reply from you.
Oh, strange, I am pretty sure I replied to that. Let me resend it.
quoted
The sync handler is nice to have and can be done in a later patch series.
On Thu, Dec 01, 2016 at 10:09:19PM +0100, Hannes Frederic Sowa wrote:
On 01.12.2016 21:54, Ido Schimmel wrote:
quoted
On Thu, Dec 01, 2016 at 09:40:48PM +0100, Hannes Frederic Sowa wrote:
quoted
On 01.12.2016 21:04, David Miller wrote:
quoted
Hannes and Ido,
It looks like we are very close to having this in mergable shape, can
you guys work out this final issue and figure out if it really is
a merge stopped or not?
Sure, if the fib notification register could be done under protection of
the sequence counter I don't see any more problems.
Did you maybe miss my reply yesterday? Because I was trying to
understand what "ordering" you're referring to, but didn't receive a
reply from you.
Oh, strange, I am pretty sure I replied to that. Let me resend it.
From: Hannes Frederic Sowa <hidden> Date: 2016-12-01 21:57:57
On 30.11.2016 19:22, Ido Schimmel wrote:
On Wed, Nov 30, 2016 at 05:49:56PM +0100, Hannes Frederic Sowa wrote:
quoted
On 30.11.2016 17:32, Ido Schimmel wrote:
quoted
On Wed, Nov 30, 2016 at 04:37:48PM +0100, Hannes Frederic Sowa wrote:
quoted
On 30.11.2016 11:09, Jiri Pirko wrote:
quoted
From: Ido Schimmel <redacted>
Make sure the device has a complete view of the FIB tables by invoking
their dump during module init.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
.../net/ethernet/mellanox/mlxsw/spectrum_router.c | 23 ++++++++++++++++++++++
1 file changed, 23 insertions(+)
@@ -2027,8 +2027,23 @@ static int mlxsw_sp_router_fib_event(struct notifier_block *nb,returnNOTIFY_DONE;}+staticvoidmlxsw_sp_router_fib_dump_flush(structnotifier_block*nb)+{+structmlxsw_sp*mlxsw_sp=container_of(nb,structmlxsw_sp,fib_nb);++/* Flush pending FIB notifications and then flush the device's+*tablebeforerequestinganotherdump.DothatwithRTNLheld,+*asFIBnotificationblockisalreadyregistered.+*/+mlxsw_core_flush_owq();+rtnl_lock();+mlxsw_sp_router_fib_flush(mlxsw_sp);+rtnl_unlock();+}+intmlxsw_sp_router_init(structmlxsw_sp*mlxsw_sp){+fib_dump_cb_t*cb=mlxsw_sp_router_fib_dump_flush;interr;INIT_LIST_HEAD(&mlxsw_sp->router.nexthop_neighs_list);
@@ -2048,8 +2063,16 @@ int mlxsw_sp_router_init(struct mlxsw_sp *mlxsw_sp)mlxsw_sp->fib_nb.notifier_call=mlxsw_sp_router_fib_event;register_fib_notifier(&mlxsw_sp->fib_nb);
Sorry to pick in here again:
There is a race here. You need to protect the registration of the fib
notifier as well by the sequence counter. Updates here are not ordered
in relation to this code below.
You mean updates that can be received after you registered the notifier
and until the dump started? I'm aware of that and that's OK. This
listener should be able to handle duplicates.
I am not concerned about duplicates, but about ordering deletes and
getting an add from the RCU code you will add the node to hw while it is
deleted in the software path. You probably will ignore the delete
because nothing is installed in hw and later add the node which was
actually deleted but just reordered which happend on another CPU, no?
Are you referring to reordering in the workqueue? We already covered
this using an ordered workqueue, which has one context of execution
system-wide.
Ups, sorry, I missed that mail. Probably read it on the mobile phone and
it became invisible for me later on. Busy day... ;)
The reordering in the workqueue seems fine to me and also still necessary.
Basically, if you delete a node right now the kernel might simply do a
RCU_INIT_POINTER(ptr_location, NULL), which has absolutely no barriers
or synchronization with the reader side. Thus you might get a callback
from the notifier for a delete event on the one CPU and you end up
queueing this fib entry after the delete queue, because the RCU walk
isn't protected by any means.
Looking closer at this series again, I overlooked the fact that you
fetch fib_seq using a rtnl_lock and rtnl_unlock pair, which first of all
orders fetching of fib_seq and thus the RCU dumping after any concurrent
executing fib table update, also the mutex_lock and unlock provide
proper acquire and release fences, so the CPU indeed sees the effect of
a RCU_INIT_POINTER update done on another CPU, because they pair with
the rtnl_unlock which might happen on the other CPU.
My question is if this is a bit of luck and if we should make this
explicit by putting the registration itself under the protection of the
sequence counter. I favor the additional protection, e.g. if we some day
actually we optimize the fib_seq code? Otherwise we might probably
document this fact. :)
quoted
quoted
I've a follow up patchset that introduces a new event in switchdev
notification chain called SWITCHDEV_SYNC, which is sent when port
netdevs are enslaved / released from a master device (points in time
where kernel<->device can get out of sync). It will invoke
re-propagation of configuration from different parts of the stack
(e.g. bridge driver, 8021q driver, fib/neigh code), which can result
in duplicates.
Okay, understood. I wonder how we can protect against accidentally abort
calls actually. E.g. if I start to inject routes into my routing domain
how can I make sure the box doesn't die after I try to insert enough
routes. Do we need to touch quagga etc?
The whole point of moving abort mechanism to the driver is that the
system won't die, but instead routing will be done in the kernel. If you
respect hardware limitations, then there's no reason for abort mechanism
to kick in.
Quick follow-up question: How can I quickly find out the hw limitations
via the kernel api?
Thanks,
Hannes
On Thu, Dec 01, 2016 at 10:57:52PM +0100, Hannes Frederic Sowa wrote:
On 30.11.2016 19:22, Ido Schimmel wrote:
quoted
On Wed, Nov 30, 2016 at 05:49:56PM +0100, Hannes Frederic Sowa wrote:
quoted
On 30.11.2016 17:32, Ido Schimmel wrote:
quoted
On Wed, Nov 30, 2016 at 04:37:48PM +0100, Hannes Frederic Sowa wrote:
quoted
On 30.11.2016 11:09, Jiri Pirko wrote:
quoted
From: Ido Schimmel <redacted>
Make sure the device has a complete view of the FIB tables by invoking
their dump during module init.
Signed-off-by: Ido Schimmel <redacted>
Signed-off-by: Jiri Pirko <redacted>
---
.../net/ethernet/mellanox/mlxsw/spectrum_router.c | 23 ++++++++++++++++++++++
1 file changed, 23 insertions(+)
@@ -2027,8 +2027,23 @@ static int mlxsw_sp_router_fib_event(struct notifier_block *nb,returnNOTIFY_DONE;}+staticvoidmlxsw_sp_router_fib_dump_flush(structnotifier_block*nb)+{+structmlxsw_sp*mlxsw_sp=container_of(nb,structmlxsw_sp,fib_nb);++/* Flush pending FIB notifications and then flush the device's+*tablebeforerequestinganotherdump.DothatwithRTNLheld,+*asFIBnotificationblockisalreadyregistered.+*/+mlxsw_core_flush_owq();+rtnl_lock();+mlxsw_sp_router_fib_flush(mlxsw_sp);+rtnl_unlock();+}+intmlxsw_sp_router_init(structmlxsw_sp*mlxsw_sp){+fib_dump_cb_t*cb=mlxsw_sp_router_fib_dump_flush;interr;INIT_LIST_HEAD(&mlxsw_sp->router.nexthop_neighs_list);
@@ -2048,8 +2063,16 @@ int mlxsw_sp_router_init(struct mlxsw_sp *mlxsw_sp)mlxsw_sp->fib_nb.notifier_call=mlxsw_sp_router_fib_event;register_fib_notifier(&mlxsw_sp->fib_nb);
Sorry to pick in here again:
There is a race here. You need to protect the registration of the fib
notifier as well by the sequence counter. Updates here are not ordered
in relation to this code below.
You mean updates that can be received after you registered the notifier
and until the dump started? I'm aware of that and that's OK. This
listener should be able to handle duplicates.
I am not concerned about duplicates, but about ordering deletes and
getting an add from the RCU code you will add the node to hw while it is
deleted in the software path. You probably will ignore the delete
because nothing is installed in hw and later add the node which was
actually deleted but just reordered which happend on another CPU, no?
Are you referring to reordering in the workqueue? We already covered
this using an ordered workqueue, which has one context of execution
system-wide.
Ups, sorry, I missed that mail. Probably read it on the mobile phone and
it became invisible for me later on. Busy day... ;)
Yet another reason not to read emails on your phone ;)
The reordering in the workqueue seems fine to me and also still necessary.
Correct.
Basically, if you delete a node right now the kernel might simply do a
RCU_INIT_POINTER(ptr_location, NULL), which has absolutely no barriers
or synchronization with the reader side. Thus you might get a callback
from the notifier for a delete event on the one CPU and you end up
queueing this fib entry after the delete queue, because the RCU walk
isn't protected by any means.
Looking closer at this series again, I overlooked the fact that you
fetch fib_seq using a rtnl_lock and rtnl_unlock pair, which first of all
orders fetching of fib_seq and thus the RCU dumping after any concurrent
executing fib table update, also the mutex_lock and unlock provide
proper acquire and release fences, so the CPU indeed sees the effect of
a RCU_INIT_POINTER update done on another CPU, because they pair with
the rtnl_unlock which might happen on the other CPU.
Yep, Exactly. I had a feeling this is the issue you were referring to,
but then you were the one to suggest the use of RTNL, so I was quite
confused.
My question is if this is a bit of luck and if we should make this
explicit by putting the registration itself under the protection of the
sequence counter. I favor the additional protection, e.g. if we some day
actually we optimize the fib_seq code? Otherwise we might probably
document this fact. :)
Well, some listeners don't require a dump, but only registration
(rocker) and in the future we might only need a dump (e.g., port being
moved to a different net namespace). So I'm not sure if bundling both
together is a good idea.
Maybe we can keep register_fib_notifier() as-is and add 'bool register'
to fib_notifier_dump() so that when set, 'nb' is also registered after
RCU walk, but before we check if the dump is consistent (unregistered if
inconsistent)?
quoted
quoted
quoted
I've a follow up patchset that introduces a new event in switchdev
notification chain called SWITCHDEV_SYNC, which is sent when port
netdevs are enslaved / released from a master device (points in time
where kernel<->device can get out of sync). It will invoke
re-propagation of configuration from different parts of the stack
(e.g. bridge driver, 8021q driver, fib/neigh code), which can result
in duplicates.
Okay, understood. I wonder how we can protect against accidentally abort
calls actually. E.g. if I start to inject routes into my routing domain
how can I make sure the box doesn't die after I try to insert enough
routes. Do we need to touch quagga etc?
The whole point of moving abort mechanism to the driver is that the
system won't die, but instead routing will be done in the kernel. If you
respect hardware limitations, then there's no reason for abort mechanism
to kick in.
Quick follow-up question: How can I quickly find out the hw limitations
via the kernel api?
That's a good question. Currently, you can't. However, we already have a
mechanism in place to read device's capabilities from the firmware and
we can (and should) expose some of them to the user. The best API for
that would be devlink, as it can represent the entire device as opposed
to only a port netdev like other tools.
We're also working on making the pipeline more visible to the user, so
that it would be easier for users to understand and debug their
networks. I believe a colleague of mine (Matty) presented this during
the last netdev conference.
From: Hannes Frederic Sowa <hidden> Date: 2016-12-01 23:27:30
On 02.12.2016 00:14, Ido Schimmel wrote:
[...]
quoted
Basically, if you delete a node right now the kernel might simply do a
RCU_INIT_POINTER(ptr_location, NULL), which has absolutely no barriers
or synchronization with the reader side. Thus you might get a callback
from the notifier for a delete event on the one CPU and you end up
queueing this fib entry after the delete queue, because the RCU walk
isn't protected by any means.
Looking closer at this series again, I overlooked the fact that you
fetch fib_seq using a rtnl_lock and rtnl_unlock pair, which first of all
orders fetching of fib_seq and thus the RCU dumping after any concurrent
executing fib table update, also the mutex_lock and unlock provide
proper acquire and release fences, so the CPU indeed sees the effect of
a RCU_INIT_POINTER update done on another CPU, because they pair with
the rtnl_unlock which might happen on the other CPU.
Yep, Exactly. I had a feeling this is the issue you were referring to,
but then you were the one to suggest the use of RTNL, so I was quite
confused.
At that time I actually had in mind that the fib_register would happen
under the sequence lock, so I didn't look closely to the memory barrier
pairings. I kinda still consider this to be a happy accident. ;)
quoted
My question is if this is a bit of luck and if we should make this
explicit by putting the registration itself under the protection of the
sequence counter. I favor the additional protection, e.g. if we some day
actually we optimize the fib_seq code? Otherwise we might probably
document this fact. :)
Well, some listeners don't require a dump, but only registration
(rocker) and in the future we might only need a dump (e.g., port being
moved to a different net namespace). So I'm not sure if bundling both
together is a good idea.
Maybe we can keep register_fib_notifier() as-is and add 'bool register'
to fib_notifier_dump() so that when set, 'nb' is also registered after
RCU walk, but before we check if the dump is consistent (unregistered if
inconsistent)?
I really like that. Would you mind adding this?
[...]
quoted
Quick follow-up question: How can I quickly find out the hw limitations
via the kernel api?
That's a good question. Currently, you can't. However, we already have a
mechanism in place to read device's capabilities from the firmware and
we can (and should) expose some of them to the user. The best API for
that would be devlink, as it can represent the entire device as opposed
to only a port netdev like other tools.
We're also working on making the pipeline more visible to the user, so
that it would be easier for users to understand and debug their
networks. I believe a colleague of mine (Matty) presented this during
the last netdev conference.
On Fri, Dec 02, 2016 at 12:27:25AM +0100, Hannes Frederic Sowa wrote:
I really like that. Would you mind adding this?
Yes. I'll send another version to Jiri today after testing and hopefully
we can submit today / tomorrow. I think Linus is still undecided about
-rc8 and I would like to get this in 4.10.
quoted
quoted
Quick follow-up question: How can I quickly find out the hw limitations
via the kernel api?
That's a good question. Currently, you can't. However, we already have a
mechanism in place to read device's capabilities from the firmware and
we can (and should) expose some of them to the user. The best API for
that would be devlink, as it can represent the entire device as opposed
to only a port netdev like other tools.
We're also working on making the pipeline more visible to the user, so
that it would be easier for users to understand and debug their
networks. I believe a colleague of mine (Matty) presented this during
the last netdev conference.