"David Laight" [off-list ref] writes:
quoted
It was to match the comment we have few lines above :
/*
* The pid hash table is scaled according to the amount of memory in
the
quoted
* machine. From a minimum of 16 slots up to 4096 slots at one
gigabyte or
quoted
* more.
*/
The comment was actually correct until someone converted the code to use
alloc_large_system_hash.
These large hash tables are, IMHO, an indication that the
algorithm used is, perhaps, suboptimal.
Not least of the problems is actually finding a suitable
(and fast) hash function that will work with the actual
real-life data.
The pid table is a good example of something where a hash
table is unnecessary.
Linux should steal the code I put into NetBSD :-)
On this unrelated topic. What algorithm did you use on NetBSD for
dealing with pids?
A hash chain length of 1 for common sizes is a pretty attractive
algorithm, as it minimizes the number of cache line misses. Ages ago
when I modeled the linux situation that is what I was seeing.
The normal case for the linux pid hash table is that it has 4096
entries taking 16k or 32k of memory and the typical process load
has about 200 or so processes. Making it very easy in the common
case to have a single entry hash chain.
All of the other algorithms I know have a tree structure and thus
more cache misses (as you traverse the tree) and ultimately worse
real world performance.
Eric
quoted
The pid table is a good example of something where a hash
table is unnecessary.
Linux should steal the code I put into NetBSD :-)
On this unrelated topic. What algorithm did you use on NetBSD for
dealing with pids?
Basically I forced the hash chain length to one by allocating
a pid that hit an empty entry in the table.
So you start off with (say) 64 entries and use the low 6
bits to index the table. The higher bits are incremented
each time a 'slot' is reused.
Free entries are kept in a FIFO list.
So each entry either contains a pointer to the process,
or the high bits and the index of the next free slot.
(and the PGID pointer).
When there are only (say) 2 free entries, then the table
size is doubled, the pointers moved to the correct places,
the free ist fixed up, and the minimum number of free entries
doubled.
The overall effect:
- lookup is only ever a mask and index + compare.
- Allocate is always fast and fixed cost (except when
the table size has to be doubled).
- A pid value will never be reused within (about) 2000
allocates (for 16bit pids, much larger for 32bit ones).
- Allocated pid numbers tend to be random, certainly
very difficult to predict.
- Small memory footprint for small systems.
For pids we normally avoid issuing large values, but
will do so to avoid immediate re-use on systems that
have 1000s of active processes.
See lines 580-820 of
http://cvsweb.netbsd.org/bsdweb.cgi/src/sys/kern/kern_proc.c?annotate=1.
182&only_with_tag=MAIN
David
Le jeudi 01 mars 2012 à 08:55 +0000, David Laight a écrit :
quoted
quoted
The pid table is a good example of something where a hash
table is unnecessary.
Linux should steal the code I put into NetBSD :-)
On this unrelated topic. What algorithm did you use on NetBSD for
dealing with pids?
Basically I forced the hash chain length to one by allocating
a pid that hit an empty entry in the table.
So you start off with (say) 64 entries and use the low 6
bits to index the table. The higher bits are incremented
each time a 'slot' is reused.
Free entries are kept in a FIFO list.
So each entry either contains a pointer to the process,
or the high bits and the index of the next free slot.
(and the PGID pointer).
When there are only (say) 2 free entries, then the table
size is doubled, the pointers moved to the correct places,
the free ist fixed up, and the minimum number of free entries
doubled.
The overall effect:
- lookup is only ever a mask and index + compare.
- Allocate is always fast and fixed cost (except when
the table size has to be doubled).
- A pid value will never be reused within (about) 2000
allocates (for 16bit pids, much larger for 32bit ones).
- Allocated pid numbers tend to be random, certainly
very difficult to predict.
- Small memory footprint for small systems.
For pids we normally avoid issuing large values, but
will do so to avoid immediate re-use on systems that
have 1000s of active processes.
You describe a hash table mechanism still, and you made chain lengthes
be 0 or 1.
This GEN_ID/SLOT schem is the one used in IBM AIX for pid allocations
(with a 31 (or was it 32) bit range). They did not use FIFO, because
they tried to use lower slots of the proc table.
Hashes values in network land are unpredictable, so we cannot make sure
a slot contains at most one entry.
Note: If you manage 4 million tcp sockets in your server, hash table
must be at _least_ 4 million slots. I would not call that suboptimal as
you said in your previous mail.
We could argue that default sizes of these hash tables (ip route,
tcp, ...) are now more suited for high end uses instead of
desktop/embedded uses, because ram sizes increased so much last 10
years.
[ We added in commits 0ccfe61803ad & c9503e0fe05 a 512K limits for
tcp/ip route hash tables to stop insanity, but thats it ]
RCU lookups make dynamic resizes of these hash table a bit complex.
IIRC David Miller have submitted a patch but this work was not
completed. Not sure its an issue these days, as we prefer to work on ip
route cache (and its associated hash table) removal anyway.
For UDP/UDPLite it certainly is not an issue, since max size is 65536
slots. On a 32bit kernel, with at more 1GB of LOWMEM, max is 512 slots.
On Thu, Mar 01, 2012 at 04:33:38AM -0800, Eric Dumazet wrote:
Le jeudi 01 mars 2012 à 08:55 +0000, David Laight a écrit :
quoted
quoted
quoted
The pid table is a good example of something where a hash
table is unnecessary.
Linux should steal the code I put into NetBSD :-)
On this unrelated topic. What algorithm did you use on NetBSD for
dealing with pids?
Basically I forced the hash chain length to one by allocating
a pid that hit an empty entry in the table.
So you start off with (say) 64 entries and use the low 6
bits to index the table. The higher bits are incremented
each time a 'slot' is reused.
Free entries are kept in a FIFO list.
So each entry either contains a pointer to the process,
or the high bits and the index of the next free slot.
(and the PGID pointer).
When there are only (say) 2 free entries, then the table
size is doubled, the pointers moved to the correct places,
the free ist fixed up, and the minimum number of free entries
doubled.
The overall effect:
- lookup is only ever a mask and index + compare.
- Allocate is always fast and fixed cost (except when
the table size has to be doubled).
- A pid value will never be reused within (about) 2000
allocates (for 16bit pids, much larger for 32bit ones).
- Allocated pid numbers tend to be random, certainly
very difficult to predict.
- Small memory footprint for small systems.
For pids we normally avoid issuing large values, but
will do so to avoid immediate re-use on systems that
have 1000s of active processes.
You describe a hash table mechanism still, and you made chain lengthes
be 0 or 1.
This GEN_ID/SLOT schem is the one used in IBM AIX for pid allocations
(with a 31 (or was it 32) bit range). They did not use FIFO, because
they tried to use lower slots of the proc table.
Hashes values in network land are unpredictable, so we cannot make sure
a slot contains at most one entry.
Note: If you manage 4 million tcp sockets in your server, hash table
must be at _least_ 4 million slots. I would not call that suboptimal as
you said in your previous mail.
We could argue that default sizes of these hash tables (ip route,
tcp, ...) are now more suited for high end uses instead of
desktop/embedded uses, because ram sizes increased so much last 10
years.
[ We added in commits 0ccfe61803ad & c9503e0fe05 a 512K limits for
tcp/ip route hash tables to stop insanity, but thats it ]
RCU lookups make dynamic resizes of these hash table a bit complex.
IIRC David Miller have submitted a patch but this work was not
completed. Not sure its an issue these days, as we prefer to work on ip
route cache (and its associated hash table) removal anyway.
One approach is to put a pair of list_head structures in each object, then
proceed as Herbert Xu did: http://lists.openwall.net/netdev/2010/02/28/20.
There is another RCU-protected resizable hash table with the userspace
RCU library: http://lttng.org/urcu. And another approach is described
in a USENIX paper:
http://www.usenix.org/event/atc11/tech/final_files/Triplett.pdf
Thanx, Paul
For UDP/UDPLite it certainly is not an issue, since max size is 65536
slots. On a 32bit kernel, with at more 1GB of LOWMEM, max is 512 slots.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/