Re: [RFC, PATCH 00/19] Numa aware LRU lists and shrinkers

[RFC, PATCH 00/19] Numa aware LRU lists and shrinkers · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 01/19] dcache: convert dentry_stat.nr_unused to per-cpu counters · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 02/19] dentry: move to per-sb LRU locks · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 03/19] dcache: remove dentries from LRU before putting on dispose list · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 04/19] mm: new shrinker API · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 05/19] shrinker: convert superblock shrinkers to new API · Dave Chinner <david@fromorbit.com> · 2012-11-27
Re: [PATCH 05/19] shrinker: convert superblock shrinkers to new API · Glauber Costa <hidden> · 2012-12-20
Re: [PATCH 05/19] shrinker: convert superblock shrinkers to new API · Dave Chinner <david@fromorbit.com> · 2012-12-21
Re: [PATCH 05/19] shrinker: convert superblock shrinkers to new API · Glauber Costa <hidden> · 2012-12-21
[PATCH 06/19] list: add a new LRU list type · Dave Chinner <david@fromorbit.com> · 2012-11-27
Re: [PATCH 06/19] list: add a new LRU list type · Christoph Hellwig <hch@infradead.org> · 2012-11-28
[PATCH 07/19] inode: convert inode lru list to generic lru list code. · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 08/19] dcache: convert to use new lru list infrastructure · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 09/19] list_lru: per-node list infrastructure · Dave Chinner <david@fromorbit.com> · 2012-11-27
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Glauber Costa <hidden> · 2012-12-20
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Dave Chinner <david@fromorbit.com> · 2012-12-21
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Glauber Costa <hidden> · 2013-01-16
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Dave Chinner <david@fromorbit.com> · 2013-01-16
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Glauber Costa <hidden> · 2013-01-17
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Dave Chinner <david@fromorbit.com> · 2013-01-17
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Glauber Costa <hidden> · 2013-01-17
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Dave Chinner <david@fromorbit.com> · 2013-01-18
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Glauber Costa <hidden> · 2013-01-18
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Dave Chinner <david@fromorbit.com> · 2013-01-18
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Glauber Costa <hidden> · 2013-01-18
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Dave Chinner <david@fromorbit.com> · 2013-01-19
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Glauber Costa <hidden> · 2013-01-19
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Glauber Costa <hidden> · 2013-01-18
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Dave Chinner <david@fromorbit.com> · 2013-01-18
Re: [PATCH 09/19] list_lru: per-node list infrastructure · Glauber Costa <hidden> · 2013-01-18
[PATCH 11/19] fs: convert inode and dentry shrinking to be node aware · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 12/19] xfs: convert buftarg LRU to generic code · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 13/19] xfs: Node aware direct inode reclaim · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 14/19] xfs: use generic AG walk for background inode reclaim · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 15/19] xfs: convert dquot cache lru to list_lru · Dave Chinner <david@fromorbit.com> · 2012-11-27
Re: [PATCH 15/19] xfs: convert dquot cache lru to list_lru · Christoph Hellwig <hch@infradead.org> · 2012-11-28
[PATCH 16/19] fs: convert fs shrinkers to new scan/count API · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 17/19] drivers: convert shrinkers to new count/scan API · Dave Chinner <david@fromorbit.com> · 2012-11-27
Re: [PATCH 17/19] drivers: convert shrinkers to new count/scan API · Chris Wilson <hidden> · 2012-11-28
Re: [PATCH 17/19] drivers: convert shrinkers to new count/scan API · Dave Chinner <david@fromorbit.com> · 2012-11-28
Re: [PATCH 17/19] drivers: convert shrinkers to new count/scan API · Glauber Costa <hidden> · 2012-11-28
Re: [PATCH 17/19] drivers: convert shrinkers to new count/scan API · Dave Chinner <david@fromorbit.com> · 2012-11-28
Re: [PATCH 17/19] drivers: convert shrinkers to new count/scan API · Glauber Costa <hidden> · 2012-11-29
Re: [PATCH 17/19] drivers: convert shrinkers to new count/scan API · Dave Chinner <david@fromorbit.com> · 2012-11-29
Re: [PATCH 17/19] drivers: convert shrinkers to new count/scan API · Konrad Rzeszutek Wilk <konrad@kernel.org> · 2013-06-07
[PATCH 18/19] shrinker: convert remaining shrinkers to count/scan API · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 19/19] shrinker: Kill old ->shrink API. · Dave Chinner <david@fromorbit.com> · 2012-11-27
[PATCH 10/19] shrinker: add node awareness · Dave Chinner <david@fromorbit.com> · 2012-11-27
Re: [RFC, PATCH 00/19] Numa aware LRU lists and shrinkers · Glauber Costa <hidden> · 2012-12-20
Re: [RFC, PATCH 00/19] Numa aware LRU lists and shrinkers · Dave Chinner <david@fromorbit.com> · 2012-12-21
Re: [RFC, PATCH 00/19] Numa aware LRU lists and shrinkers · Glauber Costa <hidden> · 2012-12-21
Re: [RFC, PATCH 00/19] Numa aware LRU lists and shrinkers · Glauber Costa <hidden> · 2013-01-21
Re: [RFC, PATCH 00/19] Numa aware LRU lists and shrinkers · Dave Chinner <david@fromorbit.com> · 2013-01-21
Re: [RFC, PATCH 00/19] Numa aware LRU lists and shrinkers · Glauber Costa <hidden> · 2013-01-23
Re: [RFC, PATCH 00/19] Numa aware LRU lists and shrinkers · Dave Chinner <david@fromorbit.com> · 2013-01-23

From: Dave Chinner <david@fromorbit.com>
Date: 2013-01-21 23:21:21
Also in: linux-fsdevel, lkml

On Mon, Jan 21, 2013 at 08:08:53PM +0400, Glauber Costa wrote:

On 11/28/2012 03:14 AM, Dave Chinner wrote:

quoted

[PATCH 09/19] list_lru: per-node list infrastructure

This makes the generic LRU list much more scalable by changing it to
a {list,lock,count} tuple per node. There are no external API
changes to this changeover, so is transparent to current users.

[PATCH 10/19] shrinker: add node awareness
[PATCH 11/19] fs: convert inode and dentry shrinking to be node

Adds a nodemask to the struct shrink_control for callers of
shrink_slab to set appropriately for their reclaim context. This
nodemask is then passed by the inode and dentry cache reclaim code
to the generic LRU list code to implement node aware shrinking.

I have a follow up question that popped up from a discussion between me
and my very American friend Johnny Wheeler, also known as Johannes
Weiner (CC'd). I actually remember we discussing this, but don't fully
remember the outcome. And since I can't find it anywhere, it must have
been in a media other than e-mail. So I thought it would do no harm in
at least documenting it...

Why are we doing this per-node, instead of per-zone?

It seems to me that the goal is to collapse all zones of a node into a
single list, but since the number of zones is not terribly larger than
the number of nodes, and zones is where the pressure comes from, what do
we really gain from this?

The number is quite a bit higher - there are platforms with 5 zones
to a node. The reality is, though, for most platforms slab
allocations come from a single zone - they never come from ZONE_DMA,
ZONE_HIGHMEM or ZONE_MOVEABLE, so there is there is no good reason
for having cache LRUs for these zones. So, two zones at most.

And then there's the complexity issue - it's simple/trivial to user
per node lists, node masks, etc. It's an obvious abstraction that
everyone understands, is simle to understand, acheives exactly the
purpose that is needed and is not tied to the /current/
implementation of the current VM memory management code.

I don't see any good reason for tying LRUs to MM zones. the
original implementation of the per-node shrinkers by Nick Piggin did
this: the LRUs for the dentry and inode caches were embedded in the
struct zone, and it wasn't generically extensible because of that.
i.e. node-aware shrinkers were directly influenced by the zone
infrastructure and so the internal implementation of the mm
subsystem started leaking out and determining how completely
unrelated subsystems need to implement their own cache
management.....

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org.  For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>

`h`	back out one level
`j`	next message in thread
`k`	previous message in thread
`l`	drill in
`Esc`	close help / fold thread tree
`?`	toggle this help