From: Daniel Borkmann
...
But I think I found a different problem with this idea. It could
happen with net devices as well, but probably less likely as there
might be a better distribution of hold/puts among CPUs. However,
for TX_RING, if we pin the process to a particular CPU, and since
the destructor is invoked through ksoftirqd, we could end up with
a misbalance and if the process runs long enough eventually
overflow for one particular CPU. We could work around that, but I
think it's not worth the effort.
The sum will be 'correct' when summed across all the cpu even if
one of the values has wrapped - provided all the arithmetic is
unsigned and the variables all the same type.
David