From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:21
Slab defragmentation is mainly an issue if Linux is used as a fileserver
and large amounts of dentries, inodes and buffer heads accumulate. In some
load situations the slabs become very sparsely populated so that a lot of
memory is wasted by slabs that only contain one or a few objects. In
extreme cases the performance of a machine will become sluggish since
we are continually running reclaim. Slab defragmentation adds the
capability to recover wasted memory.
For lumpy reclaim slab defragmentation can be used to enhance the
ability to recover larger contiguous areas of memory. Lumpy reclaim currently
cannot do anything if a slab page is encountered. With slab defragmentation
that slab page can be removed and a large contiguous page freed. It may
be possible to have slab pages also part of ZONE_MOVABLE (Mel's defrag
scheme in 2.6.23) or the MOVABLE areas (antifrag patches in mm).
The trouble with this patchset is that it is difficult to validate.
Activities are only performed when special load situations are encountered.
Are there any tests that could give meaningful information about
the effectiveness of these measures? I have run various tests here
creating and deleting files and building kernels under low memory situations
to trigger these reclaim mechanisms but how does one measure their
effectiveness?
The patchset is also available via git
git pull git://git.kernel.org/pub/scm/linux/kernel/git/christoph/slab.git defrag
We currently support the following types of reclaim:
1. dentry cache
2. inode cache (with a generic interface to allow easy setup of more
filesystems than the currently supported ext2/3/4 reiserfs, XFS
and proc)
3. buffer_head
One typical mechanism that triggers slab defragmentation on my systems
is the daily run of
updatedb
Updatedb scans all files on the system which causes a high inode and dentry
use. After updatedb is complete we need to go back to the regular use
patterns (typical on my machine: kernel compiles). Those need the memory now
for different purposes. The inodes and dentries used for updatedb will
gradually be aged by the dentry/inode reclaim algorithm which will free
up the dentries and inode entries randomly through the slabs that were
allocated. As a result the slabs will become sparsely populated. If they
become empty then they can be freed but a lot of them will remain sparsely
populated. That is where slab defrag comes in: It removes the slabs with
just a few entries reclaiming more memory for other uses.
V4->V5:
- Support lumpy reclaim for slabs
- Support reclaim via slab_shrink()
- Add constructors to insure a consistent object state at all times.
V3->V4:
- Optimize scan for slabs that need defragmentation
- Add /sys/slab/*/defrag_ratio to allow setting defrag limits
per slab.
- Add support for buffer heads.
- Describe how the cleanup after the daily updatedb can be
improved by slab defragmentation.
V2->V3
- Support directory reclaim
- Add infrastructure to trigger defragmentation after slab shrinking if we
have slabs with a high degree of fragmentation.
V1->V2
- Clean up control flow using a state variable. Simplify API. Back to 2
functions that now take arrays of objects.
- Inode defrag support for a set of filesystems
- Fix up dentry defrag support to work on negative dentries by adding
a new dentry flag that indicates that a dentry is not in the process
of being freed or allocated.
--
From: Christoph Lameter <hidden> Date: 2007-09-01 01:41:09
Move the counting function for objects in partial slabs so that it is placed
before kmem_cache_shrink. We will need to use it to establish the
fragmentation ratio of per node slab lists.
Signed-off-by: Christoph Lameter <redacted>
---
mm/slub.c | 26 +++++++++++++-------------
1 files changed, 13 insertions(+), 13 deletions(-)
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:21
Create an ops field in /sys/slab/*/ops to contain all the operations defined
on a slab. This will be used to display the additional operations that we
will define soon.
Signed-off-by: Christoph Lameter <redacted>
---
mm/slub.c | 16 +++++++++-------
1 files changed, 9 insertions(+), 7 deletions(-)
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:21
Add the two methods needed for defragmentation and add the display of the
methods via the proc interface.
Add documentation explaining the use of these methods.
Signed-off-by: Christoph Lameter <redacted>
---
include/linux/slab.h | 3 +++
include/linux/slub_def.h | 32 ++++++++++++++++++++++++++++++++
mm/slub.c | 32 ++++++++++++++++++++++++++++++--
3 files changed, 65 insertions(+), 2 deletions(-)
@@ -50,6 +50,38 @@ struct kmem_cache {intobjects;/* Number of objects in slab */intrefcount;/* Refcount for slab cache destroy */void(*ctor)(void*,structkmem_cache*,unsignedlong);++/*+*Calledwithslablockheldandinterruptsdisabled.+*Noslaboperationmaybeperformedinget().+*+*Parameterspassedarethenumberofobjectstoprocess+*andanarrayofpointerstoobjectsforwhichwe+*needreferences.+*+*Returnsapointerthatispassedtothekickfunction.+*Ifallobjectscannotbemovedthenthepointermay+*indicatethatthiswontworkandthenkickcansimply+*removethereferencesthatwerealreadyobtained.+*+*Thearraypassedtoget()isalsopassedtokick().The+*functionmayremoveobjectsbysettingarrayelementstoNULL.+*/+void*(*get)(structkmem_cache*,intnr,void**);++/*+*Calledwithnolocksheldandinterruptsenabled.+*Anyoperationmaybeperformedinkick().+*+*Parameterspassedarethenumberofobjectsinthearray,+*thearrayofpointerstotheobjectsandthepointer+*returnedbyget().+*+*Successischeckedbyexaminingthenumberofremaining+*objectsintheslab.+*/+void(*kick)(structkmem_cache*,intnr,void**,void*private);+intinuse;/* Offset to metadata */intalign;/* Alignment */intdefrag_ratio;/*
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:21
-D lists caches that support defragmentation
-C lists caches that use a ctor.
Change field names for defrag_ratio and remote_node_defrag_ratio.
Add determination of the allocation ratio for slab. The allocation ratio
is the percentage of available slots for objects in use.
Signed-off-by: Christoph Lameter <redacted>
---
Documentation/vm/slabinfo.c | 52 ++++++++++++++++++++++++++++++++++++------
1 files changed, 44 insertions(+), 8 deletions(-)
@@ -56,6 +58,8 @@ int show_slab = 0;intskip_zero=1;intshow_numa=0;intshow_track=0;+intshow_defrag=0;+intshow_ctor=0;intshow_first_alias=0;intvalidate=0;intshrink=0;
@@ -90,18 +94,20 @@ void fatal(const char *x, ...)voidusage(void){printf("slabinfo 5/7/2007. (c) 2007 sgi. clameter@sgi.com\n\n"-"slabinfo [-ahnpvtsz] [-d debugopts] [slab-regexp]\n"+"slabinfo [-aCDefhilnosSrtTvz1] [-d debugopts] [slab-regexp]\n""-a|--aliases Show aliases\n"+"-C|--ctor Show slabs with ctors\n""-d<options>|--debug=<options> Set/Clear Debug options\n"-"-e|--empty Show empty slabs\n"+"-D|--defrag Show defragmentable caches\n"+"-e|--empty Show empty slabs\n""-f|--first-alias Show first alias\n""-h|--help Show usage information\n""-i|--inverted Inverted list\n""-l|--slabs Show slabs\n""-n|--numa Show NUMA information\n"-"-o|--ops Show kmem_cache_ops\n"+"-o|--ops Show kmem_cache_ops\n""-s|--shrink Shrink slabs\n"-"-r|--report Detailed report on single slabs\n"+"-r|--report Detailed report on single slabs\n""-S|--Size Sort by size\n""-t|--tracking Show alloc/free information\n""-T|--Totals Show summary information\n"
@@ -281,7 +287,7 @@ int line = 0;voidfirst_line(void){printf("Name Objects Objsize Space "-"Slabs/Part/Cpu O/S O %%Fr %%Ef Flg\n");+"Slabs/Part/Cpu O/S O %%Ra %%Ef Flg\n");}/*
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:21
Add a parameter to add_partial instead of having separate functions.
That allows the detailed control from multiple places when putting
slabs back to the partial list. If we put slabs back to the front
then they are likely used immediately for allocations. If they are
put at the end then we can maximize the time that the partial slabs
spent without allocations.
When deactivating slab we can put the slabs that had remote objects freed
to them at the end of the list so that the cachelines can cool down.
Slabs that had objects from the cpu freed to them are put in the front
of the list to be reused ASAP.
Signed-off-by: Christoph Lameter <redacted>
---
mm/slub.c | 31 +++++++++++++++----------------
1 file changed, 15 insertions(+), 16 deletions(-)
Index: linux-2.6/mm/slub.c
===================================================================
@@ -1360,6 +1357,8 @@ static void deactivate_slab(struct kmem_while(unlikely(c->freelist)){void**object;+tail=0;/* Hot objects. Put the slab first */+/* Retrieve object from cpu_freelist */object=c->freelist;c->freelist=c->freelist[c->offset];
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:21
The defrag_ratio is used to set the threshold when a slabcache should be
defragmented.
The allocation ratio is measured in a percentage of the available slots.
The percentage will be lower for slabs that are more fragmented.
Add a defrag ratio field and set it to 30% by default. A limit of 30%
that less than 3 out of 10 available slots for objects are in use.
Signed-off-by: Christoph Lameter <redacted>
---
include/linux/slub_def.h | 7 +++++++
mm/slub.c | 18 ++++++++++++++++++
2 files changed, 25 insertions(+), 0 deletions(-)
@@ -52,6 +52,13 @@ struct kmem_cache {void(*ctor)(void*,structkmem_cache*,unsignedlong);intinuse;/* Offset to metadata */intalign;/* Alignment */+intdefrag_ratio;/*+*objects/possible-objectslimit.Ifwehave+*lessthatthespecifiedpercentageof+*objectsallocatedthendefragpasses+*willstarttooccurduringreclaim.+*/+constchar*name;/* Name (only for display!) */structlist_headlist;/* List of slab caches */#ifdef CONFIG_SLUB_DEBUG
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:21
We need the defrag ratio for the non NUMA situation now. The NUMA defrag works
by allocating objects from partial slabs on remote nodes. Rename it to
remote_node_defrag_ratio
to be clear about this.
Signed-off-by: Christoph Lameter <redacted>
---
include/linux/slub_def.h | 5 ++++-
mm/slub.c | 17 +++++++++--------
2 files changed, 13 insertions(+), 9 deletions(-)
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:22
This patch triggers slab defragmentation from memory reclaim.
The logical point for this is after slab shrinking was performed in
vmscan.c. At that point the fragmentation ratio of a slab was increased
by objects being freed. So we call kmem_cache_defrag from there.
slab_shrink() from vmscan.c is called in some contexts to do
global shrinking of slabs and in others to do shrinking for
a particular zone. Pass the zone to slab_shrink, so that slab_shrink
can call kmem_cache_defrag() and restrict the defragmentation to
the node that is under memory pressure.
Signed-off-by: Christoph Lameter <redacted>
---
fs/drop_caches.c | 2 +-
include/linux/mm.h | 2 +-
include/linux/slab.h | 1 +
mm/vmscan.c | 27 ++++++++++++++++++++-------
4 files changed, 23 insertions(+), 9 deletions(-)
@@ -1202,7 +1202,7 @@ int in_gate_area_no_task(unsigned long addr);intdrop_caches_sysctl_handler(structctl_table*,int,structfile*,void__user*,size_t*,loff_t*);unsignedlongshrink_slab(unsignedlongscanned,gfp_tgfp_mask,-unsignedlonglru_pages);+unsignedlonglru_pages,structzone*zone);voiddrop_pagecache(void);voiddrop_slab(void);
@@ -210,6 +218,8 @@ unsigned long shrink_slab(unsigned long scanned, gfp_t gfp_mask,shrinker->nr+=total_scan;}up_read(&shrinker_rwsem);+if(gfp_mask&__GFP_FS)+kmem_cache_defrag(zone?zone_to_nid(zone):-1);returnret;}
@@ -1151,7 +1161,8 @@ unsigned long try_to_free_pages(struct zone **zones, int order, gfp_t gfp_mask)if(!priority)disable_swap_token();nr_reclaimed+=shrink_zones(priority,zones,&sc);-shrink_slab(sc.nr_scanned,gfp_mask,lru_pages);+shrink_slab(sc.nr_scanned,gfp_mask,lru_pages,+NULL);if(reclaim_state){nr_reclaimed+=reclaim_state->reclaimed_slab;reclaim_state->reclaimed_slab=0;
@@ -1559,7 +1570,7 @@ unsigned long shrink_all_memory(unsigned long nr_pages)/* If slab caches are huge, it's better to hit them first */while(nr_slab>=lru_pages){reclaim_state.reclaimed_slab=0;-shrink_slab(nr_pages,sc.gfp_mask,lru_pages);+shrink_slab(nr_pages,sc.gfp_mask,lru_pages,NULL);if(!reclaim_state.reclaimed_slab)break;
@@ -1597,7 +1608,7 @@ unsigned long shrink_all_memory(unsigned long nr_pages)reclaim_state.reclaimed_slab=0;shrink_slab(sc.nr_scanned,sc.gfp_mask,-count_lru_pages());+count_lru_pages(),NULL);ret+=reclaim_state.reclaimed_slab;if(ret>=nr_pages)gotoout;
@@ -1614,7 +1625,8 @@ unsigned long shrink_all_memory(unsigned long nr_pages)if(!ret){do{reclaim_state.reclaimed_slab=0;-shrink_slab(nr_pages,sc.gfp_mask,count_lru_pages());+shrink_slab(nr_pages,sc.gfp_mask,+count_lru_pages(),NULL);ret+=reclaim_state.reclaimed_slab;}while(ret<nr_pages&&reclaim_state.reclaimed_slab>0);}
@@ -1774,7 +1786,8 @@ static int __zone_reclaim(struct zone *zone, gfp_t gfp_mask, unsigned int order)*Notethatshrink_slabwillfreememoryonallzonesandmay*takealongtime.*/-while(shrink_slab(sc.nr_scanned,gfp_mask,order)&&+while(shrink_slab(sc.nr_scanned,gfp_mask,order,+zone)&&zone_page_state(zone,NR_SLAB_RECLAIMABLE)>slab_reclaimable-nr_pages);
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:22
Creates a special function kmem_cache_isolate_slab() and kmem_cache_reclaim()
to support lumpy reclaim.
In order to isolate pages we will have to handle slab page allocations in
such a way that we can determine if a slab is valid whenever we access it
regardless of its time in life.
A valid slab that can be freed has PageSlab(page) and page->inuse > 0 set.
So we need to make sure in allocate_slab that page->inuse is zero before
PageSlab is set otherwise kmem_cache_vacate may operate on a slab that
has not been properly setup yet.
kmem_cache_isolate_page() is called from lumpy reclaim to isolate pages
neighboring a page cache page that is being reclaimed. Lumpy reclaim will
gather the slabs and call kmem_cache_reclaim() on the list.
This means that we can remove a slab that is in the way of coalescing
together a higher order page.
Signed-off-by: Christoph Lameter <redacted>
---
include/linux/slab.h | 2 +
mm/slab.c | 13 +++++++
mm/slub.c | 88 +++++++++++++++++++++++++++++++++++++++++++++++----
mm/vmscan.c | 15 ++++++--
4 files changed, 109 insertions(+), 9 deletions(-)
Index: linux-2.6/include/linux/slab.h
===================================================================
@@ -657,6 +657,7 @@ static int __isolate_lru_page(struct pag*/staticunsignedlongisolate_lru_pages(unsignedlongnr_to_scan,structlist_head*src,structlist_head*dst,+structlist_head*slab_pages,unsignedlong*scanned,intorder,intmode){unsignedlongnr_taken=0;
@@ -730,7 +731,13 @@ static unsigned long isolate_lru_pages(ucase-EBUSY:/* else it is being freed elsewhere */list_move(&cursor_page->lru,src);+break;+default:+if(slab_pages&&+kmem_cache_isolate_slab(cursor_page)==0)+list_add(&cursor_page->lru,+slab_pages);break;}}
@@ -766,6 +773,7 @@ static unsigned long shrink_inactive_lisstructzone*zone,structscan_control*sc){LIST_HEAD(page_list);+LIST_HEAD(slab_list);structpagevecpvec;unsignedlongnr_scanned=0;unsignedlongnr_reclaimed=0;
@@ -783,7 +791,7 @@ static unsigned long shrink_inactive_lisnr_taken=isolate_lru_pages(sc->swap_cluster_max,&zone->inactive_list,-&page_list,&nr_scan,sc->order,+&page_list,&slab_list,&nr_scan,sc->order,(sc->order>PAGE_ALLOC_COSTLY_ORDER)?ISOLATE_BOTH:ISOLATE_INACTIVE);nr_active=clear_active_flags(&page_list);
@@ -793,6 +801,7 @@ static unsigned long shrink_inactive_lis-(nr_taken-nr_active));zone->pages_scanned+=nr_scan;spin_unlock_irq(&zone->lru_lock);+kmem_cache_reclaim(&slab_list);nr_scanned+=nr_scan;nr_freed=shrink_page_list(&page_list,sc);
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:22
SLUB uses compound pages for larger slabs. We need to increment
the page count of these pages in order to make sure that they are not
freed under us for reclaim from within lumpy reclaim.
(The patch is also part of the large blocksize patchset)
Signed-off-by: Christoph Lameter <redacted>
---
include/linux/mm.h | 2 +-
1 files changed, 1 insertions(+), 1 deletions(-)
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:22
When we defragmenting slabs then it is advantageous to have all
defragmentable slabs together at the beginning of the list so that we do not
have to scan the complete list. When adding a slab cache put defragmentale
caches first and others last.
Determine the maximum number of objects in defragmentable slabs. This allows
to size the allocation of arrays holding refs to these objects later.
Signed-off-by: Christoph Lameter <redacted>
---
mm/slub.c | 19 +++++++++++++++++--
1 files changed, 17 insertions(+), 2 deletions(-)
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:22
Slab defragmentation (aside from Lumpy Reclaim) may occur:
1. Unconditionally when kmem_cache_shrink is called on a slab cache by the
kernel calling kmem_cache_shrink.
2. Use of the slabinfo command line to trigger slab shrinking.
3. Per node defrag conditionally when kmem_cache_defrag(<node>) is called.
Defragmentation is only performed if the fragmentation of the slab
is lower than the specified percentage. Fragmentation ratios are measured
by calculating the percentage of objects in use compared to the total
number of objects that the slab cache could hold.
kmem_cache_defrag takes a node parameter. This can either be -1 if
defragmentation should be performed on all nodes, or a node number.
If a node number was specified then defragmentation is only performed
on a specific node.
Slab defragmentation is a memory intensive operation that can be
sped up in a NUMA system if mostly node local memory is accessed. That
is the case if we just have reclaimed reclaim on a node.
In order for a slabcache to support defragmentation a couple of functions
must be setup via a call to kmem_cache_setup_defrag(). These are
void *get(struct kmem_cache *s, int nr, void **objects)
Must obtain a reference to the listed objects. SLUB guarantees that
the objects are still allocated. However, other threads may be blocked
in slab_free attempting to free objects in the slab. These may succeed
as soon as get() returns to the slab allocator. The function must
be able to detect such situations and void the attempts to free such
objects (by for example voiding the corresponding entry in the objects
array).
No slab operations may be performed in get(). Interrupts
are disabled. What can be done is very limited. The slab lock
for the page with the object is taken. Any attempt to perform a slab
operation may lead to a deadlock.
get() returns a private pointer that is passed to kick. Should we
be unable to obtain all references then that pointer may indicate
to the kick() function that it should not attempt any object removal
or move but simply remove the reference counts.
void kick(struct kmem_cache *, int nr, void **objects, void *get_result)
After SLUB has established references to the objects in a
slab it will then drop all locks and use kick() to move objects out
of the slab. The existence of the object is guaranteed by virtue of
the earlier obtained references via get(). The callback may perform
any slab operation since no locks are held at the time of call.
The callback should remove the object from the slab in some way. This
may be accomplished by reclaiming the object and then running
kmem_cache_free() or reallocating it and then running
kmem_cache_free(). Reallocation is advantageous because the partial
slabs were just sorted to have the partial slabs with the most objects
first. Reallocation is likely to result in filling up a slab in
addition to freeing up one slab so that it also can be removed from
the partial list.
Kick() does not return a result. SLUB will check the number of
remaining objects in the slab. If all objects were removed then
we know that the operation was successful.
Signed-off-by: Christoph Lameter <redacted>
---
mm/slab.c | 5 +
mm/slub.c | 265 ++++++++++++++++++++++++++++++++++++++++++++++++++------------
2 files changed, 222 insertions(+), 48 deletions(-)
Index: linux-2.6/mm/slab.c
===================================================================
@@ -2639,75 +2639,244 @@ static unsigned long count_partial(struc}/*-*kmem_cache_shrinkremovesemptyslabsfromthepartiallistsandsorts-*theremainingslabsbythenumberofitemsinuse.Theslabswiththe-*mostitemsinusecomefirst.Newallocationswillthenfillthoseup-*andthustheycanberemovedfromthepartiallists.+*Vacateallobjectsinthegivenslab.*-*Theslabswiththeleastitemsareplacedlast.Thisresultsinthem-*beingallocatedfromlastincreasingthechancethatthelastobjects-*arefreedinthem.+*Thescratchareadpassedtolistfunctionissufficienttohold+*structlistheadtimesobjectsperslab.Weuseittoholdvoid**times+*objectsperslabplusabitmapforeachobject.*/-intkmem_cache_shrink(structkmem_cache*s)+staticintkmem_cache_vacate(structpage*page,void*scratch){-intnode;-inti;-structkmem_cache_node*n;+void**vector=scratch;+void*p;+void*addr=page_address(page);+structkmem_cache*s;+unsignedlong*map;+intleftover;+intobjects;+void*private;+unsignedlongflags;+inttail=1;++BUG_ON(!PageSlab(page)||!SlabFrozen(page));+local_irq_save(flags);+slab_lock(page);++s=page->slab;+map=scratch+s->objects*sizeof(void**);+if(!page->inuse||!s->kick)+gotoout;++/* Determine used objects */+bitmap_fill(map,s->objects);+for_each_free_object(p,s,page->freelist)+__clear_bit(slab_index(p,s,addr),map);++objects=0;+memset(vector,0,s->objects*sizeof(void**));+for_each_object(p,s,addr)+if(test_bit(slab_index(p,s,addr),map))+vector[objects++]=p;++private=s->get(s,objects,vector);++/*+*Gotreferences.Nowwecandroptheslablock.Theslab+*isfrozensoitcannotvanishfromunderusnorwill+*allocationsbeperformedontheslab.However,unlockingthe+*slabwillallowconcurrentslab_freestoproceed.+*/+slab_unlock(page);+local_irq_restore(flags);++/*+*PerformtheKICKcallbackstoremovetheobjects.+*/+s->kick(s,objects,vector,private);++local_irq_save(flags);+slab_lock(page);+tail=0;+out:+/*+*Checktheresultandunfreezetheslab+*/+leftover=page->inuse;+unfreeze_slab(s,page,tail);+local_irq_restore(flags);+returnleftover;+}++/*+*Reclaimobjectsfromalistofslabpagesthathavebeengathered.+*Mustbecalledwithslabsthathavebeenisolatedbefore.+*/+intkmem_cache_reclaim(structlist_head*zaplist)+{+intfreed=0;+void**scratch;structpage*page;-structpage*t;-structlist_head*slabs_by_inuse=-kmalloc(sizeof(structlist_head)*s->objects,GFP_KERNEL);+structpage*page2;++if(list_empty(zaplist))+return0;++scratch=alloc_scratch();+if(!scratch)+return0;++list_for_each_entry_safe(page,page2,zaplist,lru){+list_del(&page->lru);+if(kmem_cache_vacate(page,scratch)==0)+freed++;+}+kfree(scratch);+returnfreed;+}++/*+*Shrinktheslabcacheonaparticularnodeofthecache+*byreleasingslabswithzeroobjectsandtryingtoreclaim+*slabswithlessthanaquarterofobjectsallocated.+*/+staticunsignedlong__kmem_cache_shrink(structkmem_cache*s,+structkmem_cache_node*n)+{unsignedlongflags;+structpage*page,*page2;+LIST_HEAD(zaplist);+intfreed=0;+intinuse;-if(!slabs_by_inuse)-return-ENOMEM;+spin_lock_irqsave(&n->list_lock,flags);+list_for_each_entry_safe(page,page2,&n->partial,lru){+inuse=page->inuse;-flush_all(s);-for_each_online_node(node){-n=get_node(s,node);+if(inuse>s->objects/4)+continue;-if(!n->nr_partial)+if(!slab_trylock(page))continue;-for(i=0;i<s->objects;i++)-INIT_LIST_HEAD(slabs_by_inuse+i);+if(inuse){-spin_lock_irqsave(&n->list_lock,flags);+list_move(&page->lru,&zaplist);-/*-*Buildlistsindexedbytheitemsinuseineachslab.-*-*Notethatconcurrentfreesmayoccurwhileweholdthe-*list_lock.page->inusehereistheupperlimit.-*/-list_for_each_entry_safe(page,t,&n->partial,lru){-if(!page->inuse&&slab_trylock(page)){-/*-*Mustholdslablockherebecauseslab_free-*mayhavefreedthelastobjectandbe-*waitingtoreleasetheslab.-*/-list_del(&page->lru);+if(s->kick){n->nr_partial--;-slab_unlock(page);-discard_slab(s,page);-}else{-list_move(&page->lru,-slabs_by_inuse+page->inuse);+SetSlabFrozen(page);}+slab_unlock(page);++}else{+list_del(&page->lru);+slab_unlock(page);+discard_slab(s,page);+freed++;}+}++if(!s->kick)+/* Simply put the zaplist at the end */+list_splice(&zaplist,n->partial.prev);+spin_unlock_irqrestore(&n->list_lock,flags);++if(s->kick)/*-*Rebuildthepartiallistwiththeslabsfilledupmost-*firstandtheleastusedslabsattheend.+*Nowwecanfreeobjectsintheslabsonthezaplist+*(orwesimplyreorderthelist*/-for(i=s->objects-1;i>=0;i--)-list_splice(slabs_by_inuse+i,n->partial.prev);+freed+=kmem_cache_reclaim(&zaplist);-spin_unlock_irqrestore(&n->list_lock,flags);+returnfreed;+}+++staticunsignedlong__kmem_cache_defrag(structkmem_cache*s,intnode)+{+unsignedlongcapacity;+unsignedlongobjects_in_full_slabs;+unsignedlongratio;+structkmem_cache_node*n=get_node(s,node);++/*+*Aninsignificantnumberofpartialslabsmakes+*theslabnotinteresting.+*/+if(n->nr_partial<=MAX_PARTIAL)+return0;++capacity=atomic_long_read(&n->nr_slabs)*s->objects;+objects_in_full_slabs=+(atomic_long_read(&n->nr_slabs)-n->nr_partial)+*s->objects;+/*+*Worstcasecalculation:Ifwewouldbeovertheratio+*evenifallpartialslabswouldonlyhaveoneobject+*thenwecanskipthefurthertestthatwouldrequireascan+*throughallthepartialpagestructstosumuptheactual+*numberofobjectsinthepartialslabs.+*/+ratio=(objects_in_full_slabs+1*n->nr_partial)*100/capacity;+if(ratio>s->defrag_ratio)+return0;++/*+*Nowfortherealcalculation.Ifusageratioismorethanrequired+*thennodefragmentation+*/+ratio=(objects_in_full_slabs+count_partial(n))*100/capacity;+if(ratio>s->defrag_ratio)+return0;++return__kmem_cache_shrink(s,n)<<s->order;+}++/*+*Defragslabsconditionalonthefragmentationratiooneachnode.+*/+intkmem_cache_defrag(intnode)+{+structkmem_cache*s;+unsignedlongpages=0;++/*+*kmem_cache_defragmaybecalledfromthereclaimpathwhichmaybe+*calledforanypageallocatoralloc.Sothereisthedangerthatwe+*getcalledinasituationwhereslubalreadyacquiredtheslub_lock+*forotherpurposes.+*/+if(!down_read_trylock(&slub_lock))+return0;++list_for_each_entry(s,&slab_caches,list){+if(node==-1){+intnid;++for_each_online_node(nid)+pages+=__kmem_cache_defrag(s,nid);+}else+pages+=__kmem_cache_defrag(s,node);}+up_read(&slub_lock);+returnpages;+}+EXPORT_SYMBOL(kmem_cache_defrag);++/*+*kmem_cache_shrinkremovesemptyslabsfromthepartiallists.+*Iftheslabcachesupportdefragmentationthenobjectsare+*reclaimed.+*/+intkmem_cache_shrink(structkmem_cache*s)+{+intnode;++flush_all(s);+for_each_online_node(node)+__kmem_cache_shrink(s,get_node(s,node));-kfree(slabs_by_inuse);return0;}EXPORT_SYMBOL(kmem_cache_shrink);
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:23
Defragmentation support for buffer heads. We convert the references to
buffers to struct page references and try to remove the buffers from
those pages. If the pages are dirty then trigger writeout so that the
buffer heads can be removed later.
Signed-off-by: Christoph Lameter <redacted>
---
fs/buffer.c | 101 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 101 insertions(+)
Index: linux-2.6/fs/buffer.c
===================================================================
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:23
Add a flag SlabReclaimable() that is set on slabs with a method
that allows defrag/reclaim. Clear the flag if a reclaim action is not
successful in reducing the number of objects in a slab. The reclaim
flag is set again if all objects have been allocated from it.
Signed-off-by: Christoph Lameter <redacted>
---
mm/slub.c | 42 ++++++++++++++++++++++++++++++++++++------
1 file changed, 36 insertions(+), 6 deletions(-)
Index: linux-2.6/mm/slub.c
===================================================================
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:23
Slabs that are reclaimable fit the definition of the objects in
ZONE_MOVABLE. So set __GFP_MOVABLE on them (this only works
on platforms where there is no HIGHMEM. Hopefully that restriction
will vanish at some point).
Also add the SLAB_TEMPORARY flag for slab caches that allocate objects with
a short lifetime. Slabs with SLAB_TEMPORARY also are allocated with
__GFP_MOVABLE. Reclaim on them works by isolating the slab for awhile and
waiting for the objects to expire.
The skbuff_head_cache is a prime example of such a slab. Add the
SLAB_TEMPORARY flag to it.
Signed-off-by: Christoph Lameter <redacted>
---
include/linux/slab.h | 1 +
mm/slub.c | 8 +++++++-
net/core/skbuff.c | 2 +-
3 files changed, 9 insertions(+), 2 deletions(-)
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:24
This implements the ability to remove inodes in a particular slab
from inode cache. In order to remove an inode we may have to write out
the pages of an inode, the inode itself and remove the dentries referring
to the node.
Provide generic functionality that can be used by filesystems that have
their own inode caches to also tie into the defragmentation functions
that are made available here.
Signed-off-by: Christoph Lameter <redacted>
---
fs/inode.c | 95 +++++++++++++++++++++++++++++++++++++++++++++++++++++
include/linux/fs.h | 5 ++
2 files changed, 100 insertions(+)
Index: linux-2.6/fs/inode.c
===================================================================
@@ -1351,6 +1351,100 @@ static int __init set_ihash_entries(char}__setup("ihash_entries=",set_ihash_entries);+staticvoid*get_inodes(structkmem_cache*s,intnr,void**v)+{+inti;++spin_lock(&inode_lock);+for(i=0;i<nr;i++){+structinode*inode=v[i];++if(inode->i_state&(I_FREEING|I_CLEAR|I_WILL_FREE))+v[i]=NULL;+else+__iget(inode);+}+spin_unlock(&inode_lock);+returnNULL;+}++/*+*Functionforfilesystemsthatembeddstructinodeintotheirown+*structures.Theoffsetistheoffsetofthestructinodeinthefsinode.+*/+void*fs_get_inodes(structkmem_cache*s,intnr,void**v,+unsignedlongoffset)+{+inti;++for(i=0;i<nr;i++)+v[i]+=offset;++returnget_inodes(s,nr,v);+}+EXPORT_SYMBOL(fs_get_inodes);++voidkick_inodes(structkmem_cache*s,intnr,void**v,void*private)+{+structinode*inode;+inti;+intabort=0;+LIST_HEAD(freeable);+structsuper_block*sb;++for(i=0;i<nr;i++){+inode=v[i];+if(!inode)+continue;++if(inode_has_buffers(inode)||inode->i_data.nrpages){+if(remove_inode_buffers(inode))+invalidate_mapping_pages(&inode->i_data,+0,-1);+}++/* Invalidate children and dentry */+if(S_ISDIR(inode->i_mode)){+structdentry*d=d_find_alias(inode);++if(d){+d_invalidate(d);+dput(d);+}+}++if(inode->i_state&I_DIRTY)+write_inode_now(inode,1);++d_prune_aliases(inode);+}++mutex_lock(&iprune_mutex);+for(i=0;i<nr;i++){+inode=v[i];+if(!inode)+continue;++sb=inode->i_sb;+iput(inode);+if(abort||!(sb->s_flags&MS_ACTIVE))+continue;++spin_lock(&inode_lock);+abort=!can_unuse(inode);++if(!abort){+list_move(&inode->i_list,&freeable);+inode->i_state|=I_FREEING;+inodes_stat.nr_unused--;+}+spin_unlock(&inode_lock);+}+dispose_list(&freeable);+mutex_unlock(&iprune_mutex);+}+EXPORT_SYMBOL(kick_inodes);+/**Initializethewaitqueuesandinodehashtable.*/
@@ -1390,6 +1484,7 @@ void __init inode_init(unsigned long memSLAB_MEM_SPREAD),init_once);register_shrinker(&icache_shrinker);+kmem_cache_setup_defrag(inode_cachep,get_inodes,kick_inodes);/* Hash may have been set up in inode_init_early */if(!hashdist)
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:24
The constructor for buffer_head slabs was removed recently. We need
the constructor in order to insure that slab objects always have a definite
state even before we allocated them.
Signed-off-by: Christoph Lameter <redacted>
---
fs/buffer.c | 19 +++++++++++++++----
1 files changed, 15 insertions(+), 4 deletions(-)
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:25
In order to support defragmentation on the dentry cache we need to have
an determined object state at all times. Without a destructor the object
would have a random state after allocation.
So provide a constructor.
Signed-off-by: Christoph Lameter <redacted>
---
fs/dcache.c | 26 ++++++++++++++------------
1 files changed, 14 insertions(+), 12 deletions(-)
@@ -2098,14 +2104,10 @@ static void __init dcache_init(unsigned long mempages){intloop;-/* -*Aconstructorcouldbeaddedforstablestatelikethelists,-*butitisprobablynotworthitbecauseofthecachenature-*ofthedcache.-*/-dentry_cache=KMEM_CACHE(dentry,-SLAB_RECLAIM_ACCOUNT|SLAB_PANIC|SLAB_MEM_SPREAD);-+dentry_cache=kmem_cache_create("dentry_cache",sizeof(structdentry),+0,SLAB_RECLAIM_ACCOUNT|SLAB_PANIC|SLAB_MEM_SPREAD,+dcache_ctor);+register_shrinker(&dcache_shrinker);/* Hash may have been set up in dcache_init_early */
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:25
Extract the common code to remove a dentry from the lru into a new function
dentry_lru_remove().
Two call sites used list_del() instead of list_del_init(). AFAIK the
performance of both is the same. dentry_lru_remove() does a list_del_init().
As a result dentry->d_lru is now always empty when a dentry is freed.
A consistent state is useful to establish dentry state from slab defrag.
Signed-off-by: Christoph Lameter <redacted>
---
fs/dcache.c | 42 ++++++++++++++----------------------------
1 files changed, 14 insertions(+), 28 deletions(-)
@@ -211,13 +219,7 @@ repeat:unhash_it:__d_drop(dentry);kill_it:-/* If dentry was on d_lru list-*deleteitfromthere-*/-if(!list_empty(&dentry->d_lru)){-list_del(&dentry->d_lru);-dentry_stat.nr_unused--;-}+dentry_lru_remove(dentry);dentry=d_kill(dentry);if(dentry)gotorepeat;
@@ -285,10 +287,7 @@ int d_invalidate(struct dentry * dentry)staticinlinestructdentry*__dget_locked(structdentry*dentry){atomic_inc(&dentry->d_count);-if(!list_empty(&dentry->d_lru)){-dentry_stat.nr_unused--;-list_del_init(&dentry->d_lru);-}+dentry_lru_remove(dentry);returndentry;}
@@ -600,10 +596,7 @@ static void shrink_dcache_for_umount_subtree(struct dentry *dentry)/* detach this root from the system */spin_lock(&dcache_lock);-if(!list_empty(&dentry->d_lru)){-dentry_stat.nr_unused--;-list_del_init(&dentry->d_lru);-}+dentry_lru_remove(dentry);__d_drop(dentry);spin_unlock(&dcache_lock);
From: Christoph Lameter <hidden> Date: 2007-09-01 01:42:25
kick() is called after get() has been used and after the slab has dropped
all of its own locks. The dentry pruning for unused entries works in a
straightforward way.
Signed-off-by: Christoph Lameter <redacted>
---
fs/dcache.c | 100 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 99 insertions(+), 1 deletion(-)
Index: linux-2.6/fs/dcache.c
===================================================================
@@ -143,7 +143,10 @@ static struct dentry *d_kill(struct dentlist_del(&dentry->d_u.d_child);dentry_stat.nr_dentry--;/* For d_free, below */-/*drops the locks, at that point nobody can reach this dentry */+/*+*dropsthelocks,atthatpointnobody(asidefromdefrag)+*canreachthisdentry+*/dentry_iput(dentry);parent=dentry->d_parent;d_free(dentry);
@@ -2100,6 +2103,100 @@ static void __init dcache_init_early(voiINIT_HLIST_HEAD(&dentry_hashtable[loop]);}+/*+*Theslaballocatorisholdingofffrees.Wecansafelyexamine+*theobjectwithoutthedangerofitvanishingfromunderus.+*/+staticvoid*get_dentries(structkmem_cache*s,intnr,void**v)+{+structdentry*dentry;+inti;++spin_lock(&dcache_lock);+for(i=0;i<nr;i++){+dentry=v[i];++/*+*Threesortsofdentriescannotbereclaimed:+*+*1.dentriesthatareintheprocessofbeingallocated+*orbeingfreed.Inthatcasethedentryisneither+*ontheLRUnorhashed.+*+*2.Fakehashedentriesasusedforanonymousdentries+*andpipeI/O.Thefakehashedentrieshaved_flags+*settoindicateahashedentry.However,the+*d_hashfieldindicatesthattheentryisnothashed.+*+*3.dentriesthathaveabackingstorethatisnot+*writable.Thisistruefortmpsfsandotherin+*memoryfilesystems.Removingdentriesfromthem+*wouldloosedentriesforgood.+*/+if((d_unhashed(dentry)&&list_empty(&dentry->d_lru))||+(!d_unhashed(dentry)&&hlist_unhashed(&dentry->d_hash))||+(dentry->d_inode&&+!mapping_cap_writeback_dirty(dentry->d_inode->i_mapping)))+/* Ignore this dentry */+v[i]=NULL;+else+/* dget_locked will remove the dentry from the LRU */+dget_locked(dentry);+}+spin_unlock(&dcache_lock);+returnNULL;+}++/*+*Slabhasdroppedallthelocks.Getridoftherefcountobtained+*earlierandalsofreetheobject.+*/+staticvoidkick_dentries(structkmem_cache*s,+intnr,void**v,void*private)+{+structdentry*dentry;+inti;++/*+*Firstinvalidatethedentrieswithoutholdingthedcachelock+*/+for(i=0;i<nr;i++){+dentry=v[i];++if(dentry)+d_invalidate(dentry);+}++/*+*Ifwearethelastoneholdingareferencethenthedentriescan+*befreed.Weneedthedcache_lock.+*/+spin_lock(&dcache_lock);+for(i=0;i<nr;i++){+dentry=v[i];+if(!dentry)+continue;++spin_lock(&dentry->d_lock);+if(atomic_read(&dentry->d_count)>1){+spin_unlock(&dentry->d_lock);+spin_unlock(&dcache_lock);+dput(dentry);+spin_lock(&dcache_lock);+continue;+}++prune_one_dentry(dentry,1);+}+spin_unlock(&dcache_lock);++/*+*dentriesarefreedusingRCUsoweneedtowaituntilRCU+*operationsarecomplete+*/+synchronize_rcu();+}+staticvoid__initdcache_init(unsignedlongmempages){intloop;
@@ -2109,6 +2206,7 @@ static void __init dcache_init(unsigned dcache_ctor);register_shrinker(&dcache_shrinker);+kmem_cache_setup_defrag(dentry_cache,get_dentries,kick_dentries);/* Hash may have been set up in dcache_init_early */if(!hashdist)
From: Jörn Engel <hidden> Date: 2007-09-06 20:38:07
On Fri, 31 August 2007 18:41:07 -0700, Christoph Lameter wrote:
The trouble with this patchset is that it is difficult to validate.
Activities are only performed when special load situations are encountered.
Are there any tests that could give meaningful information about
the effectiveness of these measures? I have run various tests here
creating and deleting files and building kernels under low memory situations
to trigger these reclaim mechanisms but how does one measure their
effectiveness?
One could play with updatedb followed by a memhog. How much time passes
and how many slab objects have to be freed before the memhog has
allocated N% of physical memory? Both numbers are relevant. The first
indicates how quickly pages are reclaimed from slab caches, while the
second show how many objects remain cached for future lookups. Updatedb
aside, caching objects is done for solid performance reasons.
Creating a qemu image with little memory and a huge directory hierarchy
filled with 0-byte files may be a nice test system. Unless you beat me
to it I'll try to set it up once logfs is in merge-worthy shape.
Jörn
--
A quarrel is quickly settled when deserted by one party; there is
no battle unless there be two.
-- Seneca
-
To unsubscribe from this list: send the line "unsubscribe linux-fsdevel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
From: Rik van Riel <hidden> Date: 2007-09-19 15:08:49
Christoph Lameter wrote:
quoted hunk
Add a flag SlabReclaimable() that is set on slabs with a method
that allows defrag/reclaim. Clear the flag if a reclaim action is not
successful in reducing the number of objects in a slab. The reclaim
flag is set again if all objects have been allocated from it.
Signed-off-by: Christoph Lameter <redacted>
---
mm/slub.c | 42 ++++++++++++++++++++++++++++++++++++------
1 file changed, 36 insertions(+), 6 deletions(-)
Index: linux-2.6/mm/slub.c
===================================================================
Why is it safe to not use the normal page flag bit operators
for these page flags operations?
--
Politics is the struggle between those who want to make their country
the best in the world, and those who believe it already is. Each group
calls the other unpatriotic.