Changelog since V7
o Rebase to 3.3-rc2
o Take greater care propagating page->pfmemalloc to skb
o Propagate pfmemalloc from netdev_alloc_page to skb where possible
o Release RCU lock properly on preempt kernel
Changelog since V6
o Rebase to 3.1-rc8
o Use wake_up instead of wake_up_interruptible()
o Do not throttle kernel threads
o Avoid a potential race between kswapd going to sleep and processes being
throttled
Changelog since V5
o Rebase to 3.1-rc5
Changelog since V4
o Update comment clarifying what protocols can be used (Michal)
o Rebase to 3.0-rc3
Changelog since V3
o Propogate pfmemalloc from packet fragment pages to skb (Neil)
o Rebase to 3.0-rc2
Changelog since V2
o Document that __GFP_NOMEMALLOC overrides __GFP_MEMALLOC (Neil)
o Use wait_event_interruptible (Neil)
o Use !! when casting to bool to avoid any possibilitity of type
truncation (Neil)
o Nicer logic when using skb_pfmemalloc_protocol (Neil)
Changelog since V1
o Rebase on top of mmotm
o Use atomic_t for memalloc_socks (David Miller)
o Remove use of sk_memalloc_socks in vmscan (Neil Brown)
o Check throttle within prepare_to_wait (Neil Brown)
o Add statistics on throttling instead of printk
When a user or administrator requires swap for their application, they
create a swap partition and file, format it with mkswap and activate it
with swapon. Swap over the network is considered as an option in diskless
systems. The two likely scenarios are when blade servers are used as part
of a cluster where the form factor or maintenance costs do not allow the
use of disks and thin clients.
The Linux Terminal Server Project recommends the use of the
Network Block Device (NBD) for swap according to the manual at
https://sourceforge.net/projects/ltsp/files/Docs-Admin-Guide/LTSPManual.pdf/download
There is also documentation and tutorials on how to setup swap over NBD
at places like https://help.ubuntu.com/community/UbuntuLTSP/EnableNBDSWAP
The nbd-client also documents the use of NBD as swap. Despite this, the
fact is that a machine using NBD for swap can deadlock within minutes if
swap is used intensively. This patch series addresses the problem.
The core issue is that network block devices do not use mempools like normal
block devices do. As the host cannot control where they receive packets from,
they cannot reliably work out in advance how much memory they might need.
Some years ago, Peter Ziljstra developed a series of patches that supported
swap over an NFS that some distributions are carrying in their kernels. This
patch series borrows very heavily from Peter's work to support swapping
over NBD as a pre-requisite to supporting swap-over-NFS. The bulk of the
complexity is concerned with preserving memory that is allocated from the
PFMEMALLOC reserves for use by the network layer which is needed for both
NBD and NFS.
Patch 1 serialises access to min_free_kbytes. It's not strictly needed
by this series but as the series cares about watermarks in
general, it's a harmless fix. It could be merged independently.
Patch 2 adds knowledge of the PFMEMALLOC reserves to SLAB and SLUB to
preserve access to pages allocated under low memory situations
to callers that are freeing memory.
Patch 3 introduces __GFP_MEMALLOC to allow access to the PFMEMALLOC
reserves without setting PFMEMALLOC.
Patch 4 opens the possibility for softirqs to use PFMEMALLOC reserves
for later use by network packet processing.
Patch 5 ignores memory policies when ALLOC_NO_WATERMARKS is set.
Patches 6-11 allows network processing to use PFMEMALLOC reserves when
the socket has been marked as being used by the VM to clean
pages. If packets are received and stored in pages that were
allocated under low-memory situations and are unrelated to
the VM, the packets are dropped.
Patch 12 is a micro-optimisation to avoid a function call in the
common case.
Patch 13 tags NBD sockets as being SOCK_MEMALLOC so they can use
PFMEMALLOC if necessary.
Patch 14 notes that it is still possible for the PFMEMALLOC reserve
to be depleted. To prevent this, direct reclaimers get
throttled on a waitqueue if 50% of the PFMEMALLOC reserves are
depleted. It is expected that kswapd and the direct reclaimers
already running will clean enough pages for the low watermark
to be reached and the throttled processes are woken up.
Patch 15 adds a statistic to track how often processes get throttled
Some basic performance testing was run using kernel builds, netperf
on loopback for UDP and TCP, hackbench (pipes and sockets), iozone
and sysbench. Each of them were expected to use the sl*b allocators
reasonably heavily but there did not appear to be significant
performance variances.
For testing swap-over-NBD, a machine was booted with 2G of RAM with a
swapfile backed by NBD. 8*NUM_CPU processes were started that create
anonymous memory mappings and read them linearly in a loop. The total
size of the mappings were 4*PHYSICAL_MEMORY to use swap heavily under
memory pressure. Without the patches, the machine locks up within
minutes and runs to completion with them applied.
drivers/block/nbd.c | 6 +-
drivers/net/ethernet/chelsio/cxgb4/sge.c | 2 +-
drivers/net/ethernet/chelsio/cxgb4vf/sge.c | 2 +-
drivers/net/ethernet/intel/igb/igb_main.c | 2 +-
drivers/net/ethernet/intel/ixgbe/ixgbe_main.c | 2 +-
drivers/net/ethernet/intel/ixgbevf/ixgbevf_main.c | 3 +-
drivers/net/usb/cdc-phonet.c | 2 +-
drivers/usb/gadget/f_phonet.c | 2 +-
include/linux/gfp.h | 13 +-
include/linux/mm_types.h | 9 +
include/linux/mmzone.h | 1 +
include/linux/sched.h | 7 +
include/linux/skbuff.h | 66 ++++++-
include/linux/slub_def.h | 1 +
include/linux/vm_event_item.h | 1 +
include/net/sock.h | 19 ++
include/trace/events/gfpflags.h | 1 +
kernel/softirq.c | 3 +
mm/page_alloc.c | 57 ++++-
mm/slab.c | 235 ++++++++++++++++++---
mm/slub.c | 36 +++-
mm/vmscan.c | 72 +++++++
mm/vmstat.c | 1 +
net/core/dev.c | 52 ++++-
net/core/filter.c | 8 +
net/core/skbuff.c | 93 +++++++--
net/core/sock.c | 42 ++++
net/ipv4/tcp.c | 3 +-
net/ipv4/tcp_output.c | 16 +-
net/ipv6/tcp_ipv6.c | 12 +-
30 files changed, 675 insertions(+), 94 deletions(-)
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
There is a race between the min_free_kbytes sysctl, memory hotplug
and transparent hugepage support enablement. Memory hotplug uses a
zonelists_mutex to avoid a race when building zonelists. Reuse it to
serialise watermark updates.
[a.p.zijlstra@chello.nl: Older patch fixed the race with spinlock]
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
mm/page_alloc.c | 23 +++++++++++++++--------
1 files changed, 15 insertions(+), 8 deletions(-)
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
The reserve is proportionally distributed over all !highmem zones
in the system. So we need to allow an emergency allocation access to
all zones. In order to do that we need to break out of any mempolicy
boundaries we might have.
In my opinion that does not break mempolicies as those are user
oriented and not system oriented. That is, system allocations are
not guaranteed to be within mempolicy boundaries. For instance IRQs
do not even have a mempolicy.
So breaking out of mempolicy boundaries for 'rare' emergency
allocations, which are always system allocations (as opposed to user)
is ok.
Signed-off-by: Peter Zijlstra <redacted>
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
mm/page_alloc.c | 7 +++++++
1 files changed, 7 insertions(+), 0 deletions(-)
@@ -2256,6 +2256,13 @@ rebalance:/* Allocate without watermarks if the context allows */if(alloc_flags&ALLOC_NO_WATERMARKS){+/*+*IgnoremempoliciesifALLOC_NO_WATERMARKSonthegrounds+*theallocationishighpriorityandthesetypeof+*allocationsaresystemratherthanuserorientated+*/+zonelist=node_zonelist(numa_node_id(),gfp_mask);+page=__alloc_pages_high_priority(gfp_mask,order,zonelist,high_zoneidx,nodemask,preferred_zone,migratetype);
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
Allow specific sockets to be tagged SOCK_MEMALLOC and use
__GFP_MEMALLOC for their allocations. These sockets will be able to go
below watermarks and allocate from the emergency reserve. Such sockets
are to be used to service the VM (iow. to swap over). They must be
handled kernel side, exposing such a socket to user-space is a bug.
There is a risk that the reserves be depleted so for now, the
administrator is responsible for increasing min_free_kbytes as
necessary to prevent deadlock for their workloads.
[a.p.zijlstra@chello.nl: Original patches]
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
include/net/sock.h | 5 ++++-
net/core/sock.c | 22 ++++++++++++++++++++++
2 files changed, 26 insertions(+), 1 deletions(-)
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
Introduce sk_allocation(), this function allows to inject sock specific
flags to each sock related allocation. It is only used on allocation
paths that may be required for writing pages back to network storage.
Signed-off-by: Peter Zijlstra <redacted>
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
include/net/sock.h | 5 +++++
net/ipv4/tcp.c | 3 ++-
net/ipv4/tcp_output.c | 16 +++++++++-------
net/ipv6/tcp_ipv6.c | 12 +++++++++---
4 files changed, 25 insertions(+), 11 deletions(-)
@@ -696,7 +696,8 @@ struct sk_buff *sk_stream_alloc_skb(struct sock *sk, int size, gfp_t gfp)/* The TCP header must be at least 32-bit aligned. */size=ALIGN(size,4);-skb=alloc_skb_fclone(size+sk->sk_prot->max_header,gfp);+skb=alloc_skb_fclone(size+sk->sk_prot->max_header,+sk_allocation(sk,gfp));if(skb){if(sk_wmem_schedule(sk,skb->truesize)){/*
@@ -2753,7 +2755,7 @@ void tcp_send_ack(struct sock *sk)/* Send it off, this clears delayed acks for us. */TCP_SKB_CB(buff)->when=tcp_time_stamp;-tcp_transmit_skb(sk,buff,0,GFP_ATOMIC);+tcp_transmit_skb(sk,buff,0,sk_allocation(sk,GFP_ATOMIC));}/* This routine sends a packet with an out of date sequence
@@ -2773,7 +2775,7 @@ static int tcp_xmit_probe_skb(struct sock *sk, int urgent)structsk_buff*skb;/* We don't queue it, tcp_transmit_skb() sets ownership. */-skb=alloc_skb(MAX_TCP_HEADER,GFP_ATOMIC);+skb=alloc_skb(MAX_TCP_HEADER,sk_allocation(sk,GFP_ATOMIC));if(skb==NULL)return-1;
@@ -584,7 +584,8 @@ static int tcp_v6_md5_do_add(struct sock *sk, const struct in6_addr *peer,}else{/* reallocate new list if current one is full. */if(!tp->md5sig_info){-tp->md5sig_info=kzalloc(sizeof(*tp->md5sig_info),GFP_ATOMIC);+tp->md5sig_info=kzalloc(sizeof(*tp->md5sig_info),+sk_allocation(sk,GFP_ATOMIC));if(!tp->md5sig_info){kfree(newkey);return-ENOMEM;
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
The skb->pfmemalloc flag gets set to true iff during the slab
allocation of data in __alloc_skb that the the PFMEMALLOC reserves
were used. If page splitting is used, it is possible that pages will
be allocated from the PFMEMALLOC reserve without propagating this
information to the skb. This patch propagates page->pfmemalloc from
pages allocated for fragments to the skb.
It works by reintroducing and expanding the netdev_alloc_page() API
to take an skb. If the page was allocated from pfmemalloc reserves,
it is automatically copied. If the driver allocates the page before
the skb, it should call propagate_pfmemalloc_skb() after the skb is
allocated to ensure the flag is copied properly.
Failure to do so is not critical. The resulting driver may perform
slower if it is used for swap-over-NBD or swap-over-NFS but it should
not result in failure.
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
drivers/net/ethernet/chelsio/cxgb4/sge.c | 2 +-
drivers/net/ethernet/chelsio/cxgb4vf/sge.c | 2 +-
drivers/net/ethernet/intel/igb/igb_main.c | 2 +-
drivers/net/ethernet/intel/ixgbe/ixgbe_main.c | 2 +-
drivers/net/ethernet/intel/ixgbevf/ixgbevf_main.c | 3 +-
drivers/net/usb/cdc-phonet.c | 2 +-
drivers/usb/gadget/f_phonet.c | 2 +-
include/linux/skbuff.h | 38 +++++++++++++++++++++
8 files changed, 46 insertions(+), 7 deletions(-)
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
Getting and putting objects in SLAB currently requires a function call
but the bulk of the work is related to PFMEMALLOC reserves which are
only consumed when network-backed storage is critical. Use an inline
function to determine if the function call is required.
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
mm/slab.c | 28 ++++++++++++++++++++++++++--
1 files changed, 26 insertions(+), 2 deletions(-)
Under significant pressure when writing back to network-backed storage,
direct reclaimers may get throttled. This is expected to be a
short-lived event and the processes get woken up again but processes do
get stalled. This patch counts how many times such stalling occurs. It's
up to the administrator whether to reduce these stalls by increasing
min_free_kbytes.
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
include/linux/vm_event_item.h | 1 +
mm/vmscan.c | 1 +
mm/vmstat.c | 1 +
3 files changed, 3 insertions(+), 0 deletions(-)
If swap is backed by network storage such as NBD, there is a risk
that a large number of reclaimers can hang the system by consuming
all PF_MEMALLOC reserves. To avoid these hangs, the administrator
must tune min_free_kbytes in advance. This patch will throttle direct
reclaimers if half the PF_MEMALLOC reserves are in use as the system
is at risk of hanging.
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
include/linux/mmzone.h | 1 +
mm/page_alloc.c | 1 +
mm/vmscan.c | 71 ++++++++++++++++++++++++++++++++++++++++++++++++
3 files changed, 73 insertions(+), 0 deletions(-)
@@ -2425,6 +2425,49 @@ out:return0;}+staticboolpfmemalloc_watermark_ok(pg_data_t*pgdat)+{+structzone*zone;+unsignedlongpfmemalloc_reserve=0;+unsignedlongfree_pages=0;+inti;++for(i=0;i<=ZONE_NORMAL;i++){+zone=&pgdat->node_zones[i];+pfmemalloc_reserve+=min_wmark_pages(zone);+free_pages+=zone_page_state(zone,NR_FREE_PAGES);+}++return(free_pages>pfmemalloc_reserve/2)?true:false;+}++/*+*Throttledirectreclaimersifbackingstorageisbackedbythenetwork+*andthePFMEMALLOCreserveforthepreferrednodeisgettingdangerously+*depleted.kswapdwillcontinuetomakeprogressandwaketheprocesses+*whenthelowwatermarkisreached+*/+staticvoidthrottle_direct_reclaim(gfp_tgfp_mask,structzonelist*zonelist,+nodemask_t*nodemask)+{+structzone*zone;+inthigh_zoneidx=gfp_zone(gfp_mask);+DEFINE_WAIT(wait);++/* Kernel threads such as kjournald should not be throttled */+if(current->flags&PF_KTHREAD)+return;++/* Check if the pfmemalloc reserves are ok */+first_zones_zonelist(zonelist,high_zoneidx,NULL,&zone);+if(pfmemalloc_watermark_ok(zone->zone_pgdat))+return;++/* Throttle */+wait_event_killable(zone->zone_pgdat->pfmemalloc_wait,+pfmemalloc_watermark_ok(zone->zone_pgdat));+}+unsignedlongtry_to_free_pages(structzonelist*zonelist,intorder,gfp_tgfp_mask,nodemask_t*nodemask){
@@ -2443,6 +2486,15 @@ unsigned long try_to_free_pages(struct zonelist *zonelist, int order,.gfp_mask=sc.gfp_mask,};+throttle_direct_reclaim(gfp_mask,zonelist,nodemask);++/*+*Donotenterreclaimiffatalsignalispending.1isreturnedso+*thatthepageallocatordoesnotconsidertriggeringOOM+*/+if(fatal_signal_pending(current))+return1;+trace_mm_vmscan_direct_reclaim_begin(order,sc.may_writepage,gfp_mask);
@@ -2840,6 +2892,12 @@ loop_again:}}++/* Wake throttled direct reclaimers if low watermark is met */+if(waitqueue_active(&pgdat->pfmemalloc_wait)&&+pfmemalloc_watermark_ok(pgdat))+wake_up(&pgdat->pfmemalloc_wait);+if(all_zones_ok||(order&&pgdat_balanced(pgdat,balanced,*classzone_idx)))break;/* kswapd: all done *//*
@@ -2961,6 +3019,19 @@ static void kswapd_try_to_sleep(pg_data_t *pgdat, int order, int classzone_idx)trace_mm_vmscan_kswapd_sleep(pgdat->node_id);/*+*Thereisapotentialracebetweenwhenkswapdchecksit+*watermarksandaprocessgetsthrottled.Thereisalso+*apotentialraceifprocessesgetthrottled,kswapdwakes,+*alargeprocessexitstherbybalancingthezonesthatcauses+*kswapdtomissawakeup.Ifkswapdisgoingtosleep,no+*processshouldbesleepingonpfmemalloc_waitsowakethem+*nowifnecessary.Ifnecessary,processeswillwakekswapd+*andgetthrottledagain+*/+if(waitqueue_active(&pgdat->pfmemalloc_wait))+wake_up(&pgdat->pfmemalloc_wait);++/**vmstatcountersarenotperfectlyaccurateandtheestimated*valueforcounterssuchasNR_FREE_PAGEScandeviatefromthe*truevaluebynr_online_cpus*threshold.Toavoidthezone
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
Set SOCK_MEMALLOC on the NBD socket to allow access to PFMEMALLOC
reserves so pages backed by NBD, particularly if swap related, can
be cleaned to prevent the machine being deadlocked. It is still
possible that the PFMEMALLOC reserves get depleted resulting in
deadlock but this can be resolved by the administrator by increasing
min_free_kbytes.
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
drivers/block/nbd.c | 6 +++++-
1 files changed, 5 insertions(+), 1 deletions(-)
@@ -155,6 +155,7 @@ static int sock_xmit(struct nbd_device *lo, int send, void *buf, int size,structmsghdrmsg;structkveciov;sigset_tblocked,oldset;+unsignedlongpflags=current->flags;if(unlikely(!sock)){dev_err(disk_to_dev(lo->disk),
@@ -168,8 +169,9 @@ static int sock_xmit(struct nbd_device *lo, int send, void *buf, int size,siginitsetinv(&blocked,sigmask(SIGKILL));sigprocmask(SIG_SETMASK,&blocked,&oldset);+current->flags|=PF_MEMALLOC;do{-sock->sk->sk_allocation=GFP_NOIO;+sock->sk->sk_allocation=GFP_NOIO|__GFP_MEMALLOC;iov.iov_base=buf;iov.iov_len=size;msg.msg_name=NULL;
@@ -215,6 +217,7 @@ static int sock_xmit(struct nbd_device *lo, int send, void *buf, int size,}while(size>0);sigprocmask(SIG_SETMASK,&oldset,NULL);+tsk_restore_flags(current,pflags,PF_MEMALLOC);returnresult;}
@@ -405,6 +408,7 @@ static int nbd_do_it(struct nbd_device *lo)BUG_ON(lo->magic!=LO_MAGIC);+sk_set_memalloc(lo->sock->sk);lo->pid=task_pid_nr(current);ret=device_create_file(disk_to_dev(lo->disk),&pid_attr);if(ret){
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
In order to make sure pfmemalloc packets receive all memory
needed to proceed, ensure processing of pfmemalloc SKBs happens
under PF_MEMALLOC. This is limited to a subset of protocols that
are expected to be used for writing to swap. Taps are not allowed to
use PF_MEMALLOC as these are expected to communicate with userspace
processes which could be paged out.
[a.p.zijlstra@chello.nl: Ideas taken from various patches]
[jslaby@suse.cz: Lock imbalance fix]
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
include/net/sock.h | 5 +++++
net/core/dev.c | 52 ++++++++++++++++++++++++++++++++++++++++++++++------
net/core/sock.c | 16 ++++++++++++++++
3 files changed, 67 insertions(+), 6 deletions(-)
@@ -3171,14 +3188,27 @@ static int __netif_receive_skb(struct sk_buff *skb)booldeliver_exact=false;intret=NET_RX_DROP;__be16type;+unsignedlongpflags=current->flags;net_timestamp_check(!netdev_tstamp_prequeue,skb);trace_netif_receive_skb(skb);+/*+*PFMEMALLOCskbsarespecial,theyshould+*-bedeliveredtoSOCK_MEMALLOCsocketsonly+*-stayawayfromuserspace+*-haveboundedmemoryusage+*+*UsePF_MEMALLOCasthissavesusfrompropagatingtheallocation+*contextdowntoallallocationsites.+*/+if(skb_pfmemalloc(skb))+current->flags|=PF_MEMALLOC;+/* if we've gotten here through NAPI, check netpoll */if(netpoll_receive_skb(skb))-returnNET_RX_DROP;+gotoout;if(!skb->skb_iif)skb->skb_iif=skb->dev->ifindex;
@@ -3273,6 +3310,7 @@ ncls:if(pt_prev){ret=pt_prev->func(skb,skb->dev,pt_prev,orig_dev);}else{+drop:atomic_long_inc(&skb->dev->rx_dropped);kfree_skb(skb);/* Jamal, now you will not able to escape explaining
@@ -293,6 +293,22 @@ void sk_clear_memalloc(struct sock *sk)}EXPORT_SYMBOL_GPL(sk_clear_memalloc);+int__sk_backlog_rcv(structsock*sk,structsk_buff*skb)+{+intret;+unsignedlongpflags=current->flags;++/* these should have been dropped before queueing */+BUG_ON(!sock_flag(sk,SOCK_MEMALLOC));++current->flags|=PF_MEMALLOC;+ret=sk->sk_backlog_rcv(sk,skb);+tsk_restore_flags(current,pflags,PF_MEMALLOC);++returnret;+}+EXPORT_SYMBOL(__sk_backlog_rcv);+#if defined(CONFIG_CGROUPS)#if !defined(CONFIG_NET_CLS_CGROUP)intnet_cls_subsys_id=-1;
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
The skb->pfmemalloc flag gets set to true iff during the slab
allocation of data in __alloc_skb that the the PFMEMALLOC reserves
were used. If the packet is fragmented, it is possible that pages
will be allocated from the PFMEMALLOC reserve without propagating
this information to the skb. This patch propagates page->pfmemalloc
from pages allocated for fragments to the skb.
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
include/linux/skbuff.h | 11 +++++++++++
1 files changed, 11 insertions(+), 0 deletions(-)
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
Change the skb allocation API to indicate RX usage and use this to fall
back to the PFMEMALLOC reserve when needed. SKBs allocated from the
reserve are tagged in skb->pfmemalloc. If an SKB is allocated from
the reserve and the socket is later found to be unrelated to page
reclaim, the packet is dropped so that the memory remains available
for page reclaim. Network protocols are expected to recover from this
packet loss.
[a.p.zijlstra@chello.nl: Ideas taken from various patches]
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
include/linux/gfp.h | 3 ++
include/linux/skbuff.h | 17 +++++++--
include/net/sock.h | 6 +++
mm/internal.h | 3 --
net/core/filter.c | 8 ++++
net/core/skbuff.c | 93 ++++++++++++++++++++++++++++++++++++++++-------
net/core/sock.c | 4 ++
7 files changed, 114 insertions(+), 20 deletions(-)
@@ -385,6 +385,9 @@ void drain_local_pages(void *dummy);*/externgfp_tgfp_allowed_mask;+/* Returns true if the gfp_mask allows use of ALLOC_NO_WATERMARK */+boolgfp_pfmemalloc_allowed(gfp_tgfp_mask);+externvoidpm_restrict_gfp_mask(void);externvoidpm_restore_gfp_mask(void);
@@ -147,6 +147,43 @@ static void skb_under_panic(struct sk_buff *skb, int sz, void *here)BUG();}++/*+*kmalloc_reserveisawrapperaroundkmalloc_node_track_callerthattells+*thecallerifemergencypfmemallocreservesarebeingused.Ifitisand+*thesocketislaterfoundtobeSOCK_MEMALLOCthenPFMEMALLOCreserves+*maybeused.Otherwise,thepacketdatamaybediscardeduntilenough+*memoryisfree+*/+#define kmalloc_reserve(size, gfp, node, pfmemalloc) \+__kmalloc_reserve(size,gfp,node,_RET_IP_,pfmemalloc)+void*__kmalloc_reserve(size_tsize,gfp_tflags,intnode,unsignedlongip,+bool*pfmemalloc)+{+void*obj;+boolret_pfmemalloc=false;++/*+*Tryaregularallocation,whenthatfailsandwe'renotentitled+*tothereserves,fail.+*/+obj=kmalloc_node_track_caller(size,+flags|__GFP_NOMEMALLOC|__GFP_NOWARN,+node);+if(obj||!(gfp_pfmemalloc_allowed(flags)))+gotoout;++/* Try again but now we are using pfmemalloc reserves */+ret_pfmemalloc=true;+obj=kmalloc_node_track_caller(size,flags,node);++out:+if(pfmemalloc)+*pfmemalloc=ret_pfmemalloc;++returnobj;+}+/* Allocate a new skbuff. We do this ourselves so we can fill in a few*'private'fieldsandalsodomemorystatisticstofindallthe*[BEEP]leaks.
@@ -169,14 +208,19 @@ static void skb_under_panic(struct sk_buff *skb, int sz, void *here)*%GFP_ATOMIC.*/structsk_buff*__alloc_skb(unsignedintsize,gfp_tgfp_mask,-intfclone,intnode)+intflags,intnode){structkmem_cache*cache;structskb_shared_info*shinfo;structsk_buff*skb;u8*data;+boolpfmemalloc;++cache=(flags&SKB_ALLOC_FCLONE)+?skbuff_fclone_cache:skbuff_head_cache;-cache=fclone?skbuff_fclone_cache:skbuff_head_cache;+if(sk_memalloc_socks()&&(flags&SKB_ALLOC_RX))+gfp_mask|=__GFP_MEMALLOC;/* Get the HEAD */skb=kmem_cache_alloc_node(cache,gfp_mask&~__GFP_DMA,node);
@@ -191,7 +235,7 @@ struct sk_buff *__alloc_skb(unsigned int size, gfp_t gfp_mask,*/size=SKB_DATA_ALIGN(size);size+=SKB_DATA_ALIGN(sizeof(structskb_shared_info));-data=kmalloc_node_track_caller(size,gfp_mask,node);+data=kmalloc_reserve(size,gfp_mask,node,&pfmemalloc);if(!data)gotonodata;/* kmalloc(size) might give us more room than requested.
@@ -209,6 +253,7 @@ struct sk_buff *__alloc_skb(unsigned int size, gfp_t gfp_mask,memset(skb,0,offsetof(structsk_buff,tail));/* Account for allocated memory : skb + skb->head */skb->truesize=SKB_TRUESIZE(size);+skb->pfmemalloc=pfmemalloc;atomic_set(&skb->users,1);skb->head=data;skb->data=data;
@@ -952,7 +1012,10 @@ int pskb_expand_head(struct sk_buff *skb, int nhead, int ntail,gotoadjust_others;}-data=kmalloc(size+sizeof(structskb_shared_info),gfp_mask);+if(skb_pfmemalloc(skb))+gfp_mask|=__GFP_MEMALLOC;+data=kmalloc_reserve(size+sizeof(structskb_shared_info),gfp_mask,+NUMA_NO_NODE,NULL);if(!data)gotonodata;
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
This is needed to allow network softirq packet processing to make
use of PF_MEMALLOC.
Currently softirq context cannot use PF_MEMALLOC due to it not being
associated with a task, and therefore not having task flags to fiddle
with - thus the gfp to alloc flag mapping ignores the task flags when
in interrupts (hard or soft) context.
Allowing softirqs to make use of PF_MEMALLOC therefore requires some
trickery. We basically borrow the task flags from whatever process
happens to be preempted by the softirq.
So we modify the gfp to alloc flags mapping to not exclude task flags
in softirq context, and modify the softirq code to save, clear and
restore the PF_MEMALLOC flag.
The save and clear, ensures the preempted task's PF_MEMALLOC flag
doesn't leak into the softirq. The restore ensures a softirq's
PF_MEMALLOC flag cannot leak back into the preempted process.
Signed-off-by: Peter Zijlstra <redacted>
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
include/linux/sched.h | 7 +++++++
kernel/softirq.c | 3 +++
mm/page_alloc.c | 5 ++++-
3 files changed, 14 insertions(+), 1 deletions(-)
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
__GFP_MEMALLOC will allow the allocation to disregard the watermarks,
much like PF_MEMALLOC. It allows one to pass along the memalloc state
in object related allocation flags as opposed to task related flags,
such as sk->sk_allocation. This removes the need for ALLOC_PFMEMALLOC
as callers using __GFP_MEMALLOC can get the ALLOC_NO_WATERMARK flag
which is now enough to identify allocations related to page reclaim.
Signed-off-by: Peter Zijlstra <redacted>
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
include/linux/gfp.h | 10 ++++++++--
include/linux/mm_types.h | 2 +-
include/trace/events/gfpflags.h | 1 +
mm/page_alloc.c | 14 ++++++--------
mm/slab.c | 2 +-
5 files changed, 17 insertions(+), 12 deletions(-)
@@ -54,7 +54,7 @@ struct page {pgoff_tindex;/* Our offset within mapping. */void*freelist;/* slub first free object */boolpfmemalloc;/* If set by the page allocator,-*ALLOC_PFMEMALLOCwasset+*ALLOC_NO_WATERMARKSwasset*andthelowwatermarkwasnot*metimplyingthatthesystem*isundersomepressure.The
@@ -3050,7 +3050,7 @@ static int cache_grow(struct kmem_cache *cachep,if(!slabp)gotoopps1;-/* Record if ALLOC_PFMEMALLOC was set when allocating the slab */+/* Record if ALLOC_NO_WATERMARKS was set when allocating the slab */if(pfmemalloc){structarray_cache*ac=cpu_cache_get(cachep);slabp->pfmemalloc=true;
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
Allocations of pages below the min watermark run a risk of the
machine hanging due to a lack of memory. To prevent this, only
callers who have PF_MEMALLOC or TIF_MEMDIE set and are not processing
an interrupt are allowed to allocate with ALLOC_NO_WATERMARKS. Once
they are allocated to a slab though, nothing prevents other callers
consuming free objects within those slabs. This patch limits access
to slab pages that were alloced from the PFMEMALLOC reserves.
Pages allocated from the reserve are returned with page->pfmemalloc
set and it is up to the caller to determine how the page should be
protected. SLAB restricts access to any page with page->pfmemalloc set
to callers which are known to able to access the PFMEMALLOC reserve. If
one is not available, an attempt is made to allocate a new page rather
than use a reserve. SLUB is a bit more relaxed in that it only records
if the current per-CPU page was allocated from PFMEMALLOC reserve and
uses another partial slab if the caller does not have the necessary
GFP or process flags. This was found to be sufficient in tests to
avoid hangs due to SLUB generally maintaining smaller lists than SLAB.
In low-memory conditions it does mean that !PFMEMALLOC allocators
can fail a slab allocation even though free objects are available
because they are being preserved for callers that are freeing pages.
[a.p.zijlstra@chello.nl: Original implementation]
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
include/linux/mm_types.h | 9 ++
include/linux/slub_def.h | 1 +
mm/internal.h | 3 +
mm/page_alloc.c | 27 +++++-
mm/slab.c | 211 +++++++++++++++++++++++++++++++++++++++-------
mm/slub.c | 36 +++++++--
6 files changed, 244 insertions(+), 43 deletions(-)
@@ -53,6 +53,15 @@ struct page {union{pgoff_tindex;/* Our offset within mapping. */void*freelist;/* slub first free object */+boolpfmemalloc;/* If set by the page allocator,+*ALLOC_PFMEMALLOCwasset+*andthelowwatermarkwasnot+*metimplyingthatthesystem+*isundersomepressure.The+*callershouldtryensure+*thispageisonlyusedto+*freeotherpages.+*/};union{
@@ -46,6 +46,7 @@ struct kmem_cache_cpu {structpage*page;/* The slab from which we are allocating */structpage*partial;/* Partially allocated frozen slabs */intnode;/* The node of the page (or -1 for debug) */+boolpfmemalloc;/* Slab page had pfmemalloc set */#ifdef CONFIG_SLUB_STATSunsignedstat[NR_SLUB_STAT_ITEMS];#endif
@@ -229,6 +231,7 @@ struct slab {unsignedintinuse;/* num of objs active in slab */kmem_bufctl_tfree;unsignedshortnodeid;+boolpfmemalloc;/* Slab had pfmemalloc set */};structslab_rcu__slab_cover_slab_rcu;};
@@ -945,12 +970,100 @@ static struct array_cache *alloc_arraycache(int node, int entries,nc->avail=0;nc->limit=entries;nc->batchcount=batchcount;-nc->touched=0;+nc->touched=false;spin_lock_init(&nc->lock);}returnnc;}+/* Clears ac->pfmemalloc if no slabs have pfmalloc set */+staticvoidcheck_ac_pfmemalloc(structkmem_cache*cachep,+structarray_cache*ac)+{+structkmem_list3*l3=cachep->nodelists[numa_mem_id()];+structslab*slabp;++if(!ac->pfmemalloc)+return;++list_for_each_entry(slabp,&l3->slabs_full,list)+if(slabp->pfmemalloc)+return;++list_for_each_entry(slabp,&l3->slabs_partial,list)+if(slabp->pfmemalloc)+return;++list_for_each_entry(slabp,&l3->slabs_free,list)+if(slabp->pfmemalloc)+return;++ac->pfmemalloc=false;+}++staticvoid*ac_get_obj(structkmem_cache*cachep,structarray_cache*ac,+gfp_tflags,boolforce_refill)+{+inti;+void*objp=ac->entry[--ac->avail];++/* Ensure the caller is allowed to use objects from PFMEMALLOC slab */+if(unlikely(is_obj_pfmemalloc(objp))){+structkmem_list3*l3;++if(gfp_pfmemalloc_allowed(flags)){+clear_obj_pfmemalloc(&objp);+returnobjp;+}++/* The caller cannot use PFMEMALLOC objects, find another one */+for(i=1;i<ac->avail;i++){+/* If a !PFMEMALLOC object is found, swap them */+if(!is_obj_pfmemalloc(ac->entry[i])){+objp=ac->entry[i];+ac->entry[i]=ac->entry[ac->avail];+ac->entry[ac->avail]=objp;+returnobjp;+}+}++/*+*Ifthereareemptyslabsontheslabs_freelistandweare+*beingforcedtorefillthecache,markthisone!pfmemalloc.+*/+l3=cachep->nodelists[numa_mem_id()];+if(!list_empty(&l3->slabs_free)&&force_refill){+structslab*slabp=virt_to_slab(objp);+slabp->pfmemalloc=false;+clear_obj_pfmemalloc(&objp);+check_ac_pfmemalloc(cachep,ac);+returnobjp;+}++/* No !PFMEMALLOC objects available */+ac->avail++;+objp=NULL;+}++returnobjp;+}++staticvoidac_put_obj(structkmem_cache*cachep,structarray_cache*ac,+void*objp)+{+structslab*slabp;++/* If there are pfmemalloc slabs, check if the object is part of one */+if(unlikely(ac->pfmemalloc)){+slabp=virt_to_slab(objp);++if(slabp->pfmemalloc)+set_obj_pfmemalloc(&objp);+}++ac->entry[ac->avail++]=objp;+}+/**Transferobjectsinonearraycachetoanother.*Lockingmustbehandledbythecaller.
@@ -2924,7 +3040,7 @@ static int cache_grow(struct kmem_cache *cachep,*'nodeid'.*/if(!objp)-objp=kmem_getpages(cachep,local_flags,nodeid);+objp=kmem_getpages(cachep,local_flags,nodeid,&pfmemalloc);if(!objp)gotofailed;
@@ -2934,6 +3050,13 @@ static int cache_grow(struct kmem_cache *cachep,if(!slabp)gotoopps1;+/* Record if ALLOC_PFMEMALLOC was set when allocating the slab */+if(pfmemalloc){+structarray_cache*ac=cpu_cache_get(cachep);+slabp->pfmemalloc=true;+ac->pfmemalloc=true;+}+slab_map_pages(cachep,slabp,objp);cache_init_objs(cachep,slabp);
@@ -3071,16 +3194,19 @@ bad:#define check_slabp(x,y) do { } while(0)#endif-staticvoid*cache_alloc_refill(structkmem_cache*cachep,gfp_tflags)+staticvoid*cache_alloc_refill(structkmem_cache*cachep,gfp_tflags,+boolforce_refill){intbatchcount;structkmem_list3*l3;structarray_cache*ac;intnode;-retry:check_irq_off();node=numa_mem_id();+if(unlikely(force_refill))+gotoforce_grow;+retry:ac=cpu_cache_get(cachep);batchcount=ac->batchcount;if(!ac->touched&&batchcount>BATCHREFILL_LIMIT){
@@ -3098,7 +3224,7 @@ retry:/* See if we can refill from the shared array */if(l3->shared&&transfer_objects(ac,l3->shared,batchcount)){-l3->shared->touched=1;+l3->shared->touched=true;gotoalloc_done;}
@@ -3150,18 +3276,22 @@ alloc_done:if(unlikely(!ac->avail)){intx;-x=cache_grow(cachep,flags|GFP_THISNODE,node,NULL);+force_grow:+x=cache_grow(cachep,flags|GFP_THISNODE,node,NULL,false);/* cache_grow can reenable interrupts, then ac could change. */ac=cpu_cache_get(cachep);-if(!x&&ac->avail==0)/* no objects in sight? abort */++/* no objects in sight? abort */+if(!x&&(ac->avail==0||force_refill))returnNULL;if(!ac->avail)/* objects refilled by interrupt? */gotoretry;}-ac->touched=1;-returnac->entry[--ac->avail];+ac->touched=true;++returnac_get_obj(cachep,ac,flags,force_refill);}staticinlinevoidcache_alloc_debugcheck_before(structkmem_cache*cachep,
@@ -2200,6 +2214,16 @@ redo:gotonew_slab;}+/*+*Byrights,weshouldbesearchingforaslabpagethatwas+*PFMEMALLOCbutrightnow,wearelosingthepfmemalloc+*informationwhenthepageleavestheper-cpuallocator+*/+if(unlikely(!pfmemalloc_match(c,gfpflags))){+deactivate_slab(s,c);+gotonew_slab;+}+/* must check again c->freelist in case of cpu migration or IRQ */object=c->freelist;if(object)
@@ -2236,7 +2260,6 @@ new_slab:/* Then do expensive stuff like retrieving pages from the partial lists */object=get_partial(s,gfpflags,node,c);-if(unlikely(!object)){object=new_slab_objects(s,gfpflags,node,&c);
@@ -2784,10 +2807,11 @@ static void early_kmem_cache_node_alloc(int node){structpage*page;structkmem_cache_node*n;+boolpfmemalloc;/* Ignore this early in boot */BUG_ON(kmem_cache_node->size<sizeof(structkmem_cache_node));-page=new_slab(kmem_cache_node,GFP_NOWAIT,node);+page=new_slab(kmem_cache_node,GFP_NOWAIT,node,&pfmemalloc);BUG_ON(!page);if(page_to_nid(page)!=node){
--
1.7.3.4
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
On Tue, Feb 7, 2012 at 6:56 AM, Mel Gorman [off-list ref] wrote:
The core issue is that network block devices do not use mempools like normal
block devices do. As the host cannot control where they receive packets from,
they cannot reliably work out in advance how much memory they might need.
Patch 1 serialises access to min_free_kbytes. It's not strictly needed
by this series but as the series cares about watermarks in
general, it's a harmless fix. It could be merged independently.
Any light shed on tuning min_free_kbytes for every day work?
Patch 2 adds knowledge of the PFMEMALLOC reserves to SLAB and SLUB to
preserve access to pages allocated under low memory situations
to callers that are freeing memory.
Patch 3 introduces __GFP_MEMALLOC to allow access to the PFMEMALLOC
reserves without setting PFMEMALLOC.
Patch 4 opens the possibility for softirqs to use PFMEMALLOC reserves
for later use by network packet processing.
Patch 5 ignores memory policies when ALLOC_NO_WATERMARKS is set.
Patches 6-11 allows network processing to use PFMEMALLOC reserves when
the socket has been marked as being used by the VM to clean
pages. If packets are received and stored in pages that were
allocated under low-memory situations and are unrelated to
the VM, the packets are dropped.
Patch 12 is a micro-optimisation to avoid a function call in the
common case.
Patch 13 tags NBD sockets as being SOCK_MEMALLOC so they can use
PFMEMALLOC if necessary.
If it is feasible to bypass hang by tuning min_mem_kbytes, things may
become simpler if NICs are also tagged. Sock buffers, pre-allocated if
necessary just after NICs are turned on, are not handed back to kmem
cache but queued on local lists which are maintained by NIC driver, based
the on the info of min_mem_kbytes or similar, for tagged NICs.
Upside is no changes in VM core. Downsides?
Patch 14 notes that it is still possible for the PFMEMALLOC reserve
to be depleted. To prevent this, direct reclaimers get
throttled on a waitqueue if 50% of the PFMEMALLOC reserves are
depleted. It is expected that kswapd and the direct reclaimers
already running will clean enough pages for the low watermark
to be reached and the throttled processes are woken up.
Patch 15 adds a statistic to track how often processes get throttled
For testing swap-over-NBD, a machine was booted with 2G of RAM with a
swapfile backed by NBD. 8*NUM_CPU processes were started that create
anonymous memory mappings and read them linearly in a loop. The total
size of the mappings were 4*PHYSICAL_MEMORY to use swap heavily under
memory pressure. Without the patches, the machine locks up within
minutes and runs to completion with them applied.
While testing, what happens if the network wire is plugged off over
three minutes?
Thanks
Hillf
On Tue, Feb 07, 2012 at 08:45:18PM +0800, Hillf Danton wrote:
On Tue, Feb 7, 2012 at 6:56 AM, Mel Gorman [off-list ref] wrote:
quoted
The core issue is that network block devices do not use mempools like normal
block devices do. As the host cannot control where they receive packets from,
they cannot reliably work out in advance how much memory they might need.
Patch 1 serialises access to min_free_kbytes. It's not strictly needed
by this series but as the series cares about watermarks in
general, it's a harmless fix. It could be merged independently.
Any light shed on tuning min_free_kbytes for every day work?
For every day work, leave min_free_kbytes as the default.
quoted
Patch 2 adds knowledge of the PFMEMALLOC reserves to SLAB and SLUB to
preserve access to pages allocated under low memory situations
to callers that are freeing memory.
Patch 3 introduces __GFP_MEMALLOC to allow access to the PFMEMALLOC
reserves without setting PFMEMALLOC.
Patch 4 opens the possibility for softirqs to use PFMEMALLOC reserves
for later use by network packet processing.
Patch 5 ignores memory policies when ALLOC_NO_WATERMARKS is set.
Patches 6-11 allows network processing to use PFMEMALLOC reserves when
the socket has been marked as being used by the VM to clean
pages. If packets are received and stored in pages that were
allocated under low-memory situations and are unrelated to
the VM, the packets are dropped.
Patch 12 is a micro-optimisation to avoid a function call in the
common case.
Patch 13 tags NBD sockets as being SOCK_MEMALLOC so they can use
PFMEMALLOC if necessary.
If it is feasible to bypass hang by tuning min_mem_kbytes,
No. Increasing or descreasing min_free_kbytes changes the timing but it
will still hang.
things may
become simpler if NICs are also tagged.
That would mean making changes to every driver and they do not necessarily
know what higher level protocol like TCP they are transmitting. How is
that simpler? What is the benefit?
Sock buffers, pre-allocated if
necessary just after NICs are turned on, are not handed back to kmem
cache but queued on local lists which are maintained by NIC driver, based
the on the info of min_mem_kbytes or similar, for tagged NICs.
I think you are referring to doing something like SKB recycling within
the driver.
Upside is no changes in VM core. Downsides?
That wouls indead requires driver-specific changes and new core
infrastructure to deal with SKB recycling spreading the complexity over a
wider range of code. If all the SKBs are in use for SOCK_MEMALLOC purposes
for whatever reason and more cannot be allocated, it will still hang. So
downsides are it would be equally if not more complex than this approach
that it may still hang.
quoted
Patch 14 notes that it is still possible for the PFMEMALLOC reserve
to be depleted. To prevent this, direct reclaimers get
throttled on a waitqueue if 50% of the PFMEMALLOC reserves are
depleted. It is expected that kswapd and the direct reclaimers
already running will clean enough pages for the low watermark
to be reached and the throttled processes are woken up.
Patch 15 adds a statistic to track how often processes get throttled
For testing swap-over-NBD, a machine was booted with 2G of RAM with a
swapfile backed by NBD. 8*NUM_CPU processes were started that create
anonymous memory mappings and read them linearly in a loop. The total
size of the mappings were 4*PHYSICAL_MEMORY to use swap heavily under
memory pressure. Without the patches, the machine locks up within
minutes and runs to completion with them applied.
While testing, what happens if the network wire is plugged off over
three minutes?
I didn't test the scenario and I don't have a test machine available
right now to try but it is up to the userspace NBD client to manage the
reconnection. It is also up to the admin to prevent the NBD client being
killed by something like the OOM killer and to have it mlocked to avoid
the NBD client itself being swapped. NFS is able to handle this in-kernel
but NBD may be more fragile.
--
Mel Gorman
SUSE Labs
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Christoph Lameter <hidden> Date: 2012-02-07 16:28:01
On Mon, 6 Feb 2012, Mel Gorman wrote:
Pages allocated from the reserve are returned with page->pfmemalloc
set and it is up to the caller to determine how the page should be
protected. SLAB restricts access to any page with page->pfmemalloc set
pfmemalloc sounds like a page flag. If you would use one then the
preservation of the flag by copying it elsewhere may not be necessary and
the patches would be less invasive. Also you would not need to extend
and modify many of the structures.
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
This takes care of the case where we are allocating the page, but what
about if we are reusing the page? For this driver it might work better
to hold of on doing the association between the page and skb either
somewhere after the skb and the page have both been allocated, or in the
igb_clean_rx_irq path where we will have both the page and the data
accessible.
I am pretty sure this is incorrect. I believe you want bi->page, not
bi->page_dma. This one is closer though to what I had in mind for igb
and ixgbe in terms of making it so there is only one location that
generates the association.
Also a similar changes would be needed for the igbvf , e1000, and e1000e
drivers in the Intel tree.
Is this function even really needed? It seems like you already have
this covered in your earlier patches, specifically 9/15, which takes
care of associating the skb and the page pfmemalloc flags when you use
skb_fill_page_desc. It would be useful to narrow things down so that we
are associating this either at the allocation time or at the
fill_page_desc call instead of doing it at both.
Thanks,
Alex
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
On Tue, Feb 7, 2012 at 9:27 PM, Mel Gorman [off-list ref] wrote:
On Tue, Feb 07, 2012 at 08:45:18PM +0800, Hillf Danton wrote:
quoted
If it is feasible to bypass hang by tuning min_mem_kbytes,
No. Increasing or descreasing min_free_kbytes changes the timing but it
will still hang.
quoted
things may
become simpler if NICs are also tagged.
That would mean making changes to every driver and they do not necessarily
know what higher level protocol like TCP they are transmitting. How is
that simpler? What is the benefit?
The benefit is to avoid allocating sock buffer in softirq by recycling,
then the changes in VM core maybe less.
Thanks
Hillf
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
On Tue, Feb 07, 2012 at 10:27:56AM -0600, Christoph Lameter wrote:
On Mon, 6 Feb 2012, Mel Gorman wrote:
quoted
Pages allocated from the reserve are returned with page->pfmemalloc
set and it is up to the caller to determine how the page should be
protected. SLAB restricts access to any page with page->pfmemalloc set
pfmemalloc sounds like a page flag. If you would use one then the
preservation of the flag by copying it elsewhere may not be necessary and
the patches would be less invasive.
Using a page flag would simplify parts of the patch. The catch of course
is that it requires a page flag which are in tight supply and I do not
want to tie this to being 32-bit unnecessarily.
Also you would not need to extend
and modify many of the structures.
Lets see;
o struct page size would be unaffected
o struct kmem_cache_cpu could be left alone even though it's a small saving
o struct slab also be left alone
o struct array_cache could be left alone although I would point out that
it would make no difference in size as touched is changed to a bool to
fit pfmemalloc in
o It would still be necessary to do the object pointer tricks in slab.c
to avoid doing an excessive number of page lookups which is where much
of the complexity is
o The virt_to_slab could be replaced by looking up the page flag instead
and avoiding a level of indirection that would be pleasing
to an int and placed with struct kmem_cache
I agree that parts of the patch would be simplier although the
complexity of storing pfmemalloc within the obj pointer would probably
remain. However, the downside of requiring a page flag is very high. In
the event we increase the number of page flags - great, I'll use one but
right now I do not think the use of page flag is justified.
--
Mel Gorman
SUSE Labs
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Christoph Lameter <hidden> Date: 2012-02-08 15:14:38
On Wed, 8 Feb 2012, Mel Gorman wrote:
o struct kmem_cache_cpu could be left alone even though it's a small saving
Its multiplied by the number of caches and by the number of
processors.
o struct slab also be left alone
o struct array_cache could be left alone although I would point out that
it would make no difference in size as touched is changed to a bool to
fit pfmemalloc in
Both of these are performance critical structures in slab.
o It would still be necessary to do the object pointer tricks in slab.c
These trick are not done for slub. It seems that they are not necessary?
remain. However, the downside of requiring a page flag is very high. In
the event we increase the number of page flags - great, I'll use one but
right now I do not think the use of page flag is justified.
On 64 bit I think there is not much of an issue with another page flag.
Also consider that the slab allocators do not make full use of the other
page flags. We could overload one of the existing flags. I removed
slubs use of them last year. PG_active could be overloaded I think.
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
This takes care of the case where we are allocating the page, but what
about if we are reusing the page?
Then nothing... You're right, I did not consider that case.
For this driver it might work better
to hold of on doing the association between the page and skb either
somewhere after the skb and the page have both been allocated, or in the
igb_clean_rx_irq path where we will have both the page and the data
accessible.
Again, from looking through the code you appear to be right. Thanks for
the suggestion!
I am pretty sure this is incorrect. I believe you want bi->page, not
bi->page_dma. This one is closer though to what I had in mind for igb
and ixgbe in terms of making it so there is only one location that
generates the association.
You are on a roll of being right.
Also a similar changes would be needed for the igbvf , e1000, and e1000e
drivers in the Intel tree.
It's not *really* needed. As noted in the changelog, getting this wrong
has minor consequences. At worst, swap becomes a little slower but it
should not result in hangs.
It seems like you already have
this covered in your earlier patches, specifically 9/15, which takes
care of associating the skb and the page pfmemalloc flags when you use
skb_fill_page_desc.
Yes, this patch was an attempt to being thorough but the actual impact
is moving a bunch of complexity into drivers where it is difficult to
test and of marginal benefit.
It would be useful to narrow things down so that we
are associating this either at the allocation time or at the
fill_page_desc call instead of doing it at both.
I think you're right. I'm going to drop this patch entirely as the
benfit is marginal and not necessary for swap over network to work.
Thanks very much for the review.
--
Mel Gorman
SUSE Labs
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
On Wed, Feb 08, 2012 at 08:51:11PM +0800, Hillf Danton wrote:
On Tue, Feb 7, 2012 at 9:27 PM, Mel Gorman [off-list ref] wrote:
quoted
On Tue, Feb 07, 2012 at 08:45:18PM +0800, Hillf Danton wrote:
quoted
If it is feasible to bypass hang by tuning min_mem_kbytes,
No. Increasing or descreasing min_free_kbytes changes the timing but it
will still hang.
quoted
things may
become simpler if NICs are also tagged.
That would mean making changes to every driver and they do not necessarily
know what higher level protocol like TCP they are transmitting. How is
that simpler? What is the benefit?
The benefit is to avoid allocating sock buffer in softirq by recycling,
then the changes in VM core maybe less.
The VM is responsible for swapping. It's reasonable that the core
VM has responsibility for it without trying to shove complexity into
drivers or elsewhere unnecessarily. I see some benefit in following on
by recycling some skbs and only allocating from softirq if no recycled
skbs are available. That potentially improves performance but I do not
recycling as a replacement.
--
Mel Gorman
SUSE Labs
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
On Wed, Feb 08, 2012 at 09:14:32AM -0600, Christoph Lameter wrote:
On Wed, 8 Feb 2012, Mel Gorman wrote:
quoted
o struct kmem_cache_cpu could be left alone even though it's a small saving
Its multiplied by the number of caches and by the number of
processors.
quoted
o struct slab also be left alone
o struct array_cache could be left alone although I would point out that
it would make no difference in size as touched is changed to a bool to
fit pfmemalloc in
Both of these are performance critical structures in slab.
Ok, I looked into what is necessary to replace these with checking a page
flag and the cost shifts quite a bit and ends up being more expensive.
Right now, I use array_cache to record if there are any pfmemalloc
objects in the free list at all. If there are not, no expensive checks
are made. For example, in __ac_put_obj(), I check ac->pfmemalloc to see
if an expensive check is required. Using a page flag, the same check
requires a lookup with virt_to_page(). This in turns uses a
pfn_to_page() which depending on the memory model can be very expensive.
No matter what, it's more expensive than a simple check and this is in
the slab free path.
It is more complicated in check_ac_pfmemalloc() too although the performance
impact is less because it is a slow path. If ac->pfmemalloc is false,
the check of each slabp can be avoided. Without it, all the slabps must
be checked unconditionally and each slabp that is checked must call
virt_to_page().
Overall, the memory savings of moving to a page flag are miniscule but
the performance cost is far higher because of the use of virt_to_page().
quoted
o It would still be necessary to do the object pointer tricks in slab.c
These trick are not done for slub. It seems that they are not necessary?
In slub, it's sufficient to check kmem_cache_cpu to know whether the
objects in the list are pfmemalloc or not.
quoted
remain. However, the downside of requiring a page flag is very high. In
the event we increase the number of page flags - great, I'll use one but
right now I do not think the use of page flag is justified.
On 64 bit I think there is not much of an issue with another page flag.
There isn't, but on 32 bit there is.
Also consider that the slab allocators do not make full use of the other
page flags. We could overload one of the existing flags. I removed
slubs use of them last year. PG_active could be overloaded I think.
Yeah, you're right on the button there. I did my checking assuming that
PG_active+PG_slab were safe to use. The following is an untested patch that
I probably got details wrong in but it illustrates where virt_to_page()
starts cropping up.
It was a good idea and thanks for thinking of it but unfortunately the
implementation would be more expensive than what I have currently.
@@ -46,7 +46,6 @@ struct kmem_cache_cpu {structpage*page;/* The slab from which we are allocating */structpage*partial;/* Partially allocated frozen slabs */intnode;/* The node of the page (or -1 for debug) */-boolpfmemalloc;/* Slab page had pfmemalloc set */#ifdef CONFIG_SLUB_STATSunsignedstat[NR_SLUB_STAT_ITEMS];#endif
@@ -233,7 +233,6 @@ struct slab {unsignedintinuse;/* num of objs active in slab */kmem_bufctl_tfree;unsignedshortnodeid;-boolpfmemalloc;/* Slab had pfmemalloc set */};structslab_rcu__slab_cover_slab_rcu;};
@@ -978,6 +976,13 @@ static struct array_cache *alloc_arraycache(int node, int entries,returnnc;}+staticinlineboolis_slab_pfmemalloc(structslab*slabp)+{+structpage*page=virt_to_page(slabp->s_mem);++returnPageSlabPfmemalloc(page);+}+/* Clears ac->pfmemalloc if no slabs have pfmalloc set */staticvoidcheck_ac_pfmemalloc(structkmem_cache*cachep,structarray_cache*ac)
@@ -1066,15 +1067,11 @@ static inline void *ac_get_obj(struct kmem_cache *cachep,staticvoid*__ac_put_obj(structkmem_cache*cachep,structarray_cache*ac,void*objp){-structslab*slabp;+structpage*page=virt_to_page(objp);/* If there are pfmemalloc slabs, check if the object is part of one */-if(unlikely(ac->pfmemalloc)){-slabp=virt_to_slab(objp);--if(slabp->pfmemalloc)-set_obj_pfmemalloc(&objp);-}+if(PageSlabPfmemalloc(page))+set_obj_pfmemalloc(&objp);returnobjp;}
@@ -3074,13 +3074,6 @@ static int cache_grow(struct kmem_cache *cachep,if(!slabp)gotoopps1;-/* Record if ALLOC_NO_WATERMARKS was set when allocating the slab */-if(pfmemalloc){-structarray_cache*ac=cpu_cache_get(cachep);-slabp->pfmemalloc=true;-ac->pfmemalloc=true;-}-slab_map_pages(cachep,slabp,objp);cache_init_objs(cachep,slabp);--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Rik van Riel <hidden> Date: 2012-02-08 18:49:00
On 02/06/2012 05:56 PM, Mel Gorman wrote:
There is a race between the min_free_kbytes sysctl, memory hotplug
and transparent hugepage support enablement. Memory hotplug uses a
zonelists_mutex to avoid a race when building zonelists. Reuse it to
serialise watermark updates.
[a.p.zijlstra@chello.nl: Older patch fixed the race with spinlock]
Signed-off-by: Mel Gorman<mgorman@suse.de>
Reviewed-by: Rik van Riel <redacted>
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Christoph Lameter <hidden> Date: 2012-02-08 19:49:12
On Wed, 8 Feb 2012, Mel Gorman wrote:
Ok, I looked into what is necessary to replace these with checking a page
flag and the cost shifts quite a bit and ends up being more expensive.
That is only true if you go the slab route. Slab suffers from not having
the page struct pointer readily available. The changes are likely already
impacting slab performance without the virt_to_page patch.
In slub, it's sufficient to check kmem_cache_cpu to know whether the
objects in the list are pfmemalloc or not.
We try to minimize the size of kmem_cache_cpu. The page pointer is readily
available. We just removed the node field from kmem_cache_cpu because it
was less expensive to get the node number from the struct page field.
The same is certainly true for a PFMEMALLOC flag.
Yeah, you're right on the button there. I did my checking assuming that
PG_active+PG_slab were safe to use. The following is an untested patch that
I probably got details wrong in but it illustrates where virt_to_page()
starts cropping up.
Yes you need to come up with a way to not use virt_to_page otherwise slab
performance is significantly impacted. On NUMA we are already doing a page
struct lookup on free in slab. If you would save the page struct pointer
there and reuse it then you would not have an issue at least on free.
You still would need to determine which "struct slab" pointer is in use
which will also require similar lookups in varous places.
Transfer of the pfmemalloc flags (guess you must have a pfmemalloc
field in struct slab then) in slab is best be done when allocating and
freeing a slab page from the page allocator.
I think its rather trivial to add the support you want in a non intrusive
way to slub. Slab would require some more thought and discussion.
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
On Wed, Feb 08, 2012 at 01:49:05PM -0600, Christoph Lameter wrote:
On Wed, 8 Feb 2012, Mel Gorman wrote:
quoted
Ok, I looked into what is necessary to replace these with checking a page
flag and the cost shifts quite a bit and ends up being more expensive.
That is only true if you go the slab route.
Well, yes but both slab and slub have to be supported. I see no reason
why I would choose to make this a slab-only or slub-only feature. Slob is
not supported because it's not expected that a platform using slob is also
going to use network-based swap.
Slab suffers from not having
the page struct pointer readily available. The changes are likely already
impacting slab performance without the virt_to_page patch.
The performance impact only comes into play when swap is on a network
device and pfmemalloc reserves are in use. The rest of the time the check
on ac avoids all the cost and there is a micro-optimisation later to avoid
calling a function (patch 12).
quoted
In slub, it's sufficient to check kmem_cache_cpu to know whether the
objects in the list are pfmemalloc or not.
We try to minimize the size of kmem_cache_cpu. The page pointer is readily
available. We just removed the node field from kmem_cache_cpu because it
was less expensive to get the node number from the struct page field.
The same is certainly true for a PFMEMALLOC flag.
Ok, are you asking that I use the page flag for slub and leave kmem_cache_cpu
alone in the slub case? I can certainly check it out if that's what you
are asking for.
quoted
Yeah, you're right on the button there. I did my checking assuming that
PG_active+PG_slab were safe to use. The following is an untested patch that
I probably got details wrong in but it illustrates where virt_to_page()
starts cropping up.
Yes you need to come up with a way to not use virt_to_page otherwise slab
performance is significantly impacted.
I did come up with a way: the necessary information is in ac and slabp
on slab :/ . There are not exactly many ways that the information can
be recorded.
On NUMA we are already doing a page struct lookup on free in slab.
If you would save the page struct pointer
there and reuse it then you would not have an issue at least on free.
That information is only available on NUMA and only when there is more than
one node. Having cache_free_alien return the page for passing to ac_put_obj()
would also be ugly. The biggest downfall by far is that single-node machines
incur the cost of virt_to_page() where they did not have to before. This
is not a solution and it is not better than the current simply check on
a struct field.
You still would need to determine which "struct slab" pointer is in use
which will also require similar lookups in varous places.
Transfer of the pfmemalloc flags (guess you must have a pfmemalloc
field in struct slab then) in slab is best be done when allocating and
freeing a slab page from the page allocator.
The page->pfmemalloc is already been transferred to the slab in
cache_grow.
I think its rather trivial to add the support you want in a non intrusive
way to slub. Slab would require some more thought and discussion.
I'm slightly confused by this sentence. Support for slub is already in the
patch and as you say, it's fairly straight-forward. Supporting a page flag
and leaving kmem_cache_cpu alone may also be easier as kmem_cache_cpu->page
can be used instead of a kmem_cache_cpu->pfmemalloc field.
--
Mel Gorman
SUSE Labs
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Christoph Lameter <hidden> Date: 2012-02-08 22:31:26
On Wed, 8 Feb 2012, Mel Gorman wrote:
On Wed, Feb 08, 2012 at 01:49:05PM -0600, Christoph Lameter wrote:
quoted
On Wed, 8 Feb 2012, Mel Gorman wrote:
quoted
Ok, I looked into what is necessary to replace these with checking a page
flag and the cost shifts quite a bit and ends up being more expensive.
That is only true if you go the slab route.
Well, yes but both slab and slub have to be supported. I see no reason
why I would choose to make this a slab-only or slub-only feature. Slob is
not supported because it's not expected that a platform using slob is also
going to use network-based swap.
I think so far the patches in particular to slab.c are pretty significant
in impact.
quoted
Slab suffers from not having
the page struct pointer readily available. The changes are likely already
impacting slab performance without the virt_to_page patch.
The performance impact only comes into play when swap is on a network
device and pfmemalloc reserves are in use. The rest of the time the check
on ac avoids all the cost and there is a micro-optimisation later to avoid
calling a function (patch 12).
We have been down this road too many times. Logic is added to critical
paths and memory structures grow. This is not free. And for NBD swap
support? Pretty exotic use case.
Ok, are you asking that I use the page flag for slub and leave kmem_cache_cpu
alone in the slub case? I can certainly check it out if that's what you
are asking for.
No I am not asking for something. Still thinking about the best way to
address the issues. I think we can easily come up with a minimally
invasive patch for slub. Not sure about slab at this point. I think we
could avoid most of the new fields but this requires some tinkering. I
have a day @ home tomorrow which hopefully gives me a chance to
put some focus on this issue.
I did come up with a way: the necessary information is in ac and slabp
on slab :/ . There are not exactly many ways that the information can
be recorded.
Wish we had something that would not involve increasing the number of
fields in these slab structures.
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
On Wed, Feb 08, 2012 at 04:13:15PM -0600, Christoph Lameter wrote:
On Wed, 8 Feb 2012, Mel Gorman wrote:
quoted
On Wed, Feb 08, 2012 at 01:49:05PM -0600, Christoph Lameter wrote:
quoted
On Wed, 8 Feb 2012, Mel Gorman wrote:
quoted
Ok, I looked into what is necessary to replace these with checking a page
flag and the cost shifts quite a bit and ends up being more expensive.
That is only true if you go the slab route.
Well, yes but both slab and slub have to be supported. I see no reason
why I would choose to make this a slab-only or slub-only feature. Slob is
not supported because it's not expected that a platform using slob is also
going to use network-based swap.
I think so far the patches in particular to slab.c are pretty significant
in impact.
Ok, I am working on a solution that does not affect any of the existing
slab structures. Between that and the fact we check if there are any
memalloc_socks after patch 12, the impact for normal systems is an additional
branch in ac_get_obj() and ac_put_obj()
quoted
quoted
Slab suffers from not having
the page struct pointer readily available. The changes are likely already
impacting slab performance without the virt_to_page patch.
The performance impact only comes into play when swap is on a network
device and pfmemalloc reserves are in use. The rest of the time the check
on ac avoids all the cost and there is a micro-optimisation later to avoid
calling a function (patch 12).
We have been down this road too many times. Logic is added to critical
paths and memory structures grow. This is not free. And for NBD swap
support? Pretty exotic use case.
NFS support is the real target. NBD is the logical starting point and
NFS needs the same support.
quoted
Ok, are you asking that I use the page flag for slub and leave kmem_cache_cpu
alone in the slub case? I can certainly check it out if that's what you
are asking for.
No I am not asking for something. Still thinking about the best way to
address the issues. I think we can easily come up with a minimally
invasive patch for slub. Not sure about slab at this point. I think we
could avoid most of the new fields but this requires some tinkering. I
have a day @ home tomorrow which hopefully gives me a chance to
put some focus on this issue.
I think we can avoid adding any additional fields but array_cache needs
a new read-mostly global. Also, when network storage is in use and under
memory pressure, it might be slower as we will have lost granularity on
what slabs are using pfmemalloc. That is an acceptable compromise as it
moves the cost to users of network-based swap instead of normal usage.
quoted
I did come up with a way: the necessary information is in ac and slabp
on slab :/ . There are not exactly many ways that the information can
be recorded.
Wish we had something that would not involve increasing the number of
fields in these slab structures.
This is what I currently have. It's untested but builds. It reverts the
structures back to the way they were and uses page flags instead.
@@ -46,7 +46,6 @@ struct kmem_cache_cpu {structpage*page;/* The slab from which we are allocating */structpage*partial;/* Partially allocated frozen slabs */intnode;/* The node of the page (or -1 for debug) */-boolpfmemalloc;/* Slab page had pfmemalloc set */#ifdef CONFIG_SLUB_STATSunsignedstat[NR_SLUB_STAT_ITEMS];#endif
@@ -155,6 +155,12 @@#define ARCH_KMALLOC_FLAGS SLAB_HWCACHE_ALIGN#endif+/*+*trueifapagewasallocatedfrompfmemallocreservesfornetwork-based+*swap+*/+staticboolpfmemalloc_active;+/* Legal flag mask for kmem_cache_create(). */#if DEBUG# define CREATE_MASK (SLAB_RED_ZONE | \
@@ -233,7 +239,6 @@ struct slab {unsignedintinuse;/* num of objs active in slab */kmem_bufctl_tfree;unsignedshortnodeid;-boolpfmemalloc;/* Slab had pfmemalloc set */};structslab_rcu__slab_cover_slab_rcu;};
@@ -978,6 +982,13 @@ static struct array_cache *alloc_arraycache(int node, int entries,returnnc;}+staticinlineboolis_slab_pfmemalloc(structslab*slabp)+{+structpage*page=virt_to_page(slabp->s_mem);++returnPageSlabPfmemalloc(page);+}+/* Clears ac->pfmemalloc if no slabs have pfmalloc set */staticvoidcheck_ac_pfmemalloc(structkmem_cache*cachep,structarray_cache*ac)
@@ -1066,13 +1077,10 @@ static inline void *ac_get_obj(struct kmem_cache *cachep,staticvoid*__ac_put_obj(structkmem_cache*cachep,structarray_cache*ac,void*objp){-structslab*slabp;--/* If there are pfmemalloc slabs, check if the object is part of one */-if(unlikely(ac->pfmemalloc)){-slabp=virt_to_slab(objp);--if(slabp->pfmemalloc)+if(unlikely(pfmemalloc_active)){+/* Some pfmemalloc slabs exist, check if this is one */+structpage*page=virt_to_page(objp);+if(PageSlabPfmemalloc(page))set_obj_pfmemalloc(&objp);}
@@ -3075,11 +3086,8 @@ static int cache_grow(struct kmem_cache *cachep,gotoopps1;/* Record if ALLOC_NO_WATERMARKS was set when allocating the slab */-if(pfmemalloc){-structarray_cache*ac=cpu_cache_get(cachep);-slabp->pfmemalloc=true;-ac->pfmemalloc=true;-}+if(unlikely(pfmemalloc))+pfmemalloc_active=pfmemalloc;slab_map_pages(cachep,slabp,objp);
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Christoph Lameter <hidden> Date: 2012-02-09 19:54:04
On Thu, 9 Feb 2012, Mel Gorman wrote:
Ok, I am working on a solution that does not affect any of the existing
slab structures. Between that and the fact we check if there are any
memalloc_socks after patch 12, the impact for normal systems is an additional
branch in ac_get_obj() and ac_put_obj()
That sounds good in particular since some other things came up again,
sigh. Have not had time to see if an alternate approach works.
quoted
We have been down this road too many times. Logic is added to critical
paths and memory structures grow. This is not free. And for NBD swap
support? Pretty exotic use case.
NFS support is the real target. NBD is the logical starting point and
NFS needs the same support.
But this is already a pretty strange use case on multiple levels. Swap is
really detrimental to performance. Its a kind of emergency outlet that
gets worse with every new step that increases the differential in
performance between disk and memory. On top of that you want to add
special code in various subsystems to also do that over the network.
Sigh. I think we agreed a while back that we want to limit the amount of
I/O triggered from reclaim paths? AFAICT many filesystems do not support
writeout from reclaim anymore because of all the issues that arise at that
level.
We have numerous other mechanisms that can compress swap etc and provide
ways to work around the problem without I/O which has always be
troublesome and these fixes are likely only to work in a very limited
way causing a lot of maintenance effort because (given the exotic
nature) it is highly likely that there are cornercases that only will be
triggered in rare cases.
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
On Thu, Feb 09, 2012 at 01:53:58PM -0600, Christoph Lameter wrote:
On Thu, 9 Feb 2012, Mel Gorman wrote:
quoted
Ok, I am working on a solution that does not affect any of the existing
slab structures. Between that and the fact we check if there are any
memalloc_socks after patch 12, the impact for normal systems is an additional
branch in ac_get_obj() and ac_put_obj()
That sounds good in particular since some other things came up again,
sigh. Have not had time to see if an alternate approach works.
I have an updated version of this 02/15 patch below. It passed testing
and is a lot less invasive than the previous release. As you suggested,
it uses page flags and the bulk of the complexity is only executed if
someone is using network-backed storage.
quoted
quoted
We have been down this road too many times. Logic is added to critical
paths and memory structures grow. This is not free. And for NBD swap
support? Pretty exotic use case.
NFS support is the real target. NBD is the logical starting point and
NFS needs the same support.
But this is already a pretty strange use case on multiple levels. Swap is
really detrimental to performance. Its a kind of emergency outlet that
gets worse with every new step that increases the differential in
performance between disk and memory.
Performance is generally not the concern of the users of swap-over-N[FS|BD].
In the cases I am aware of, they just want an emergency overflow. One user
for example had an application with a sparse mapping larger than physical
memory. During the workload execution it would occasionally push small
parts out to swap and needed network-based swap due to the lack of a local
disk. The performance impact was not a concern because swapping was rare.
On top of that you want to add
special code in various subsystems to also do that over the network.
Sigh. I think we agreed a while back that we want to limit the amount of
I/O triggered from reclaim paths?
Specifically we wanted to reduce or stop page reclaim calling ->writepage()
for file-backed pages because it generated awful IO patterns and deep
call stacks. We still write anonymous pages from page reclaim because we
do not have a dedicated thread for writing to swap. It is expected that
the call stack for writing to network storage would be less than a
filesystem.
AFAICT many filesystems do not support
writeout from reclaim anymore because of all the issues that arise at that
level.
NBD is a block device so filesystem restrictions like you mention do not
apply. In NFS, the direct_IO paths are used to write pages not
->writepage so again the restriction does not apply.
We have numerous other mechanisms that can compress swap etc and provide
ways to work around the problem without I/O which has always be
troublesome and these fixes are likely only to work in a very limited
way causing a lot of maintenance effort because (given the exotic
nature) it is highly likely that there are cornercases that only will be
triggered in rare cases.
Compressing swap only gets you so far. For some workloads, at some point
the anonymous pages have to be written to swap somewhere. If there is a
local disk, great, use it. If there is no disk, then either
hardware-based solutions are needed (HBA that exposes the network as a
block device, works but is expensive), virtualisation is used (the host
os exposes a network-based swapfile as a block device to the guest but
only usable in virtualisation) or you need something like these patches.
Here is the revised 02/15 patch
=== CUT HERE ===
mm: sl[au]b: Add knowledge of PFMEMALLOC reserve pages
Allocations of pages below the min watermark run a risk of the
machine hanging due to a lack of memory. To prevent this, only
callers who have PF_MEMALLOC or TIF_MEMDIE set and are not processing
an interrupt are allowed to allocate with ALLOC_NO_WATERMARKS. Once
they are allocated to a slab though, nothing prevents other callers
consuming free objects within those slabs. This patch limits access
to slab pages that were alloced from the PFMEMALLOC reserves.
Pages allocated from the reserve are returned with page->pfmemalloc
set and it is up to the caller to determine how the page should be
protected. SLAB restricts access to any page with page->pfmemalloc set
to callers which are known to able to access the PFMEMALLOC reserve. If
one is not available, an attempt is made to allocate a new page rather
than use a reserve. SLUB is a bit more relaxed in that it only records
if the current per-CPU page was allocated from PFMEMALLOC reserve and
uses another partial slab if the caller does not have the necessary
GFP or process flags. This was found to be sufficient in tests to
avoid hangs due to SLUB generally maintaining smaller lists than SLAB.
In low-memory conditions it does mean that !PFMEMALLOC allocators
can fail a slab allocation even though free objects are available
because they are being preserved for callers that are freeing pages.
[a.p.zijlstra@chello.nl: Original implementation]
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
include/linux/mm_types.h | 9 ++
include/linux/page-flags.h | 28 +++++++
mm/internal.h | 3 +
mm/page_alloc.c | 27 +++++-
mm/slab.c | 190 +++++++++++++++++++++++++++++++++++++++-----
mm/slub.c | 27 ++++++-
6 files changed, 258 insertions(+), 26 deletions(-)
@@ -53,6 +53,15 @@ struct page {union{pgoff_tindex;/* Our offset within mapping. */void*freelist;/* slub first free object */+boolpfmemalloc;/* If set by the page allocator,+*ALLOC_PFMEMALLOCwasset+*andthelowwatermarkwasnot+*metimplyingthatthesystem+*isundersomepressure.The+*callershouldtryensure+*thispageisonlyusedto+*freeotherpages.+*/};union{
@@ -951,6 +980,98 @@ static struct array_cache *alloc_arraycache(int node, int entries,returnnc;}+staticinlineboolis_slab_pfmemalloc(structslab*slabp)+{+structpage*page=virt_to_page(slabp->s_mem);++returnPageSlabPfmemalloc(page);+}++/* Clears ac->pfmemalloc if no slabs have pfmalloc set */+staticvoidcheck_ac_pfmemalloc(structkmem_cache*cachep,+structarray_cache*ac)+{+structkmem_list3*l3=cachep->nodelists[numa_mem_id()];+structslab*slabp;++if(!pfmemalloc_active)+return;++list_for_each_entry(slabp,&l3->slabs_full,list)+if(is_slab_pfmemalloc(slabp))+return;++list_for_each_entry(slabp,&l3->slabs_partial,list)+if(is_slab_pfmemalloc(slabp))+return;++list_for_each_entry(slabp,&l3->slabs_free,list)+if(is_slab_pfmemalloc(slabp))+return;++pfmemalloc_active=false;+}++staticvoid*ac_get_obj(structkmem_cache*cachep,structarray_cache*ac,+gfp_tflags,boolforce_refill)+{+inti;+void*objp=ac->entry[--ac->avail];++/* Ensure the caller is allowed to use objects from PFMEMALLOC slab */+if(unlikely(is_obj_pfmemalloc(objp))){+structkmem_list3*l3;++if(gfp_pfmemalloc_allowed(flags)){+clear_obj_pfmemalloc(&objp);+returnobjp;+}++/* The caller cannot use PFMEMALLOC objects, find another one */+for(i=1;i<ac->avail;i++){+/* If a !PFMEMALLOC object is found, swap them */+if(!is_obj_pfmemalloc(ac->entry[i])){+objp=ac->entry[i];+ac->entry[i]=ac->entry[ac->avail];+ac->entry[ac->avail]=objp;+returnobjp;+}+}++/*+*Ifthereareemptyslabsontheslabs_freelistandweare+*beingforcedtorefillthecache,markthisone!pfmemalloc.+*/+l3=cachep->nodelists[numa_mem_id()];+if(!list_empty(&l3->slabs_free)&&force_refill){+structslab*slabp=virt_to_slab(objp);+ClearPageSlabPfmemalloc(virt_to_page(slabp->s_mem));+clear_obj_pfmemalloc(&objp);+check_ac_pfmemalloc(cachep,ac);+returnobjp;+}++/* No !PFMEMALLOC objects available */+ac->avail++;+objp=NULL;+}++returnobjp;+}++staticvoidac_put_obj(structkmem_cache*cachep,structarray_cache*ac,+void*objp)+{+if(unlikely(pfmemalloc_active)){+/* Some pfmemalloc slabs exist, check if this is one */+structpage*page=virt_to_page(objp);+if(PageSlabPfmemalloc(page))+set_obj_pfmemalloc(&objp);+}++ac->entry[ac->avail++]=objp;+}+/**Transferobjectsinonearraycachetoanother.*Lockingmustbehandledbythecaller.
@@ -1760,6 +1881,10 @@ static void *kmem_getpages(struct kmem_cache *cachep, gfp_t flags, int nodeid)if(!page)returnNULL;+/* Record if ALLOC_PFMEMALLOC was set when allocating the slab */+if(unlikely(page->pfmemalloc))+pfmemalloc_active=true;+nr_pages=(1<<cachep->gfporder);if(cachep->flags&SLAB_RECLAIM_ACCOUNT)add_zone_page_state(page_zone(page),
@@ -3150,18 +3283,22 @@ alloc_done:if(unlikely(!ac->avail)){intx;+force_grow:x=cache_grow(cachep,flags|GFP_THISNODE,node,NULL);/* cache_grow can reenable interrupts, then ac could change. */ac=cpu_cache_get(cachep);-if(!x&&ac->avail==0)/* no objects in sight? abort */++/* no objects in sight? abort */+if(!x&&(ac->avail==0||force_refill))returnNULL;if(!ac->avail)/* objects refilled by interrupt? */gotoretry;}ac->touched=1;-returnac->entry[--ac->avail];++returnac_get_obj(cachep,ac,flags,force_refill);}staticinlinevoidcache_alloc_debugcheck_before(structkmem_cache*cachep,
@@ -2200,6 +2213,16 @@ redo:gotonew_slab;}+/*+*Byrights,weshouldbesearchingforaslabpagethatwas+*PFMEMALLOCbutrightnow,wearelosingthepfmemalloc+*informationwhenthepageleavestheper-cpuallocator+*/+if(unlikely(!pfmemalloc_match(c,gfpflags))){+deactivate_slab(s,c);+gotonew_slab;+}+/* must check again c->freelist in case of cpu migration or IRQ */object=c->freelist;if(object)
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Christoph Lameter <hidden> Date: 2012-02-10 21:01:44
On Fri, 10 Feb 2012, Mel Gorman wrote:
I have an updated version of this 02/15 patch below. It passed testing
and is a lot less invasive than the previous release. As you suggested,
it uses page flags and the bulk of the complexity is only executed if
someone is using network-backed storage.
Hmmm.. hmm... Still modifies the hotpaths of the allocators for a
pretty exotic feature.
quoted
On top of that you want to add
special code in various subsystems to also do that over the network.
Sigh. I think we agreed a while back that we want to limit the amount of
I/O triggered from reclaim paths?
Specifically we wanted to reduce or stop page reclaim calling ->writepage()
for file-backed pages because it generated awful IO patterns and deep
call stacks. We still write anonymous pages from page reclaim because we
do not have a dedicated thread for writing to swap. It is expected that
the call stack for writing to network storage would be less than a
filesystem.
quoted
AFAICT many filesystems do not support
writeout from reclaim anymore because of all the issues that arise at that
level.
NBD is a block device so filesystem restrictions like you mention do not
apply. In NFS, the direct_IO paths are used to write pages not
->writepage so again the restriction does not apply.
Block devices are a little simpler ok. But it is still not a desirable
thing to do (just think about raid and other complex filesystems that may
also have to do allocations).I do not think that block device writers
code with the VM in mind. In the case of network devices as block devices
we have a pretty serious problem since the network subsystem is certainly
not designed to be called from VM reclaim code that may be triggered
arbitrarily from deeply nested other code in the kernel. Implementing
something like this invites breakage all over the place to show up.
Modification to hotpath. That could be fixed here by forcing pfmemalloc
(like debug allocs) to always go to the slow path and checking in there
instead. Just keep c->freelist == NULL.
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Christoph Lameter <hidden> Date: 2012-02-10 22:08:04
Proposal for a patch for slub to move the pfmemalloc handling out of the
fastpath by simply not assigning a per cpu slab when pfmemalloc processing
is going on.
Subject: [slub] Fix so that no mods are required for the fast path
Remove the check for pfmemalloc from the alloc hotpath and put the logic after
the election of a new per cpu slab.
For a pfmemalloc page do not use the fast path but force use of the slow
path (which is also used for the debug case).
Signed-off-by: Christoph Lameter <redacted>
---
mm/slub.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
Index: linux-2.6/mm/slub.c
===================================================================
@@ -2273,11 +2273,12 @@ new_slab:}}-if(likely(!kmem_cache_debug(s)))+if(likely(!kmem_cache_debug(s)&&pfmemalloc_match(c,gfpflags)))gotoload_freelist;+/* Only entered in the debug case */-if(!alloc_debug_processing(s,c->page,object,addr))+if(kmem_cache_debug(s)&&!alloc_debug_processing(s,c->page,object,addr))gotonew_slab;/* Slab failed checks. Next slab needed */c->freelist=get_freepointer(s,object);
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
On Fri, Feb 10, 2012 at 04:07:57PM -0600, Christoph Lameter wrote:
Proposal for a patch for slub to move the pfmemalloc handling out of the
fastpath by simply not assigning a per cpu slab when pfmemalloc processing
is going on.
Subject: [slub] Fix so that no mods are required for the fast path
Remove the check for pfmemalloc from the alloc hotpath and put the logic after
the election of a new per cpu slab.
For a pfmemalloc page do not use the fast path but force use of the slow
path (which is also used for the debug case).
Signed-off-by: Christoph Lameter <redacted>
This weakens pfmemalloc processing in the following way
1. Process that is performance network swap calls __slab_alloc.
pfmemalloc_match is true so the freelist is loaded and
c->freelist is now pointing to a pfmemalloc page
2. Process that is attempting normal allocations calls slab_alloc,
finds the pfmemalloc page on the freelist and uses it because it
did not check pfmemalloc_match()
The patch allows non-pfmemalloc allocations to use pfmemalloc pages with
the kmalloc slabs being the most vunerable caches on the grounds they are
most likely to have a mix of pfmemalloc and !pfmemalloc requests. Patch
14 will still protect the system as processes will get throttled if the
pfmemalloc reserves get depleted so performance will not degrade as smoothly.
Assuming this passes testing, I'll add the patch to the series with the
information above included in the changelog.
Thanks Christoph.
--
Mel Gorman
SUSE Labs
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
On Fri, Feb 10, 2012 at 03:01:37PM -0600, Christoph Lameter wrote:
On Fri, 10 Feb 2012, Mel Gorman wrote:
quoted
I have an updated version of this 02/15 patch below. It passed testing
and is a lot less invasive than the previous release. As you suggested,
it uses page flags and the bulk of the complexity is only executed if
someone is using network-backed storage.
Hmmm.. hmm... Still modifies the hotpaths of the allocators for a
pretty exotic feature.
quoted
quoted
On top of that you want to add
special code in various subsystems to also do that over the network.
Sigh. I think we agreed a while back that we want to limit the amount of
I/O triggered from reclaim paths?
Specifically we wanted to reduce or stop page reclaim calling ->writepage()
for file-backed pages because it generated awful IO patterns and deep
call stacks. We still write anonymous pages from page reclaim because we
do not have a dedicated thread for writing to swap. It is expected that
the call stack for writing to network storage would be less than a
filesystem.
quoted
AFAICT many filesystems do not support
writeout from reclaim anymore because of all the issues that arise at that
level.
NBD is a block device so filesystem restrictions like you mention do not
apply. In NFS, the direct_IO paths are used to write pages not
->writepage so again the restriction does not apply.
Block devices are a little simpler ok. But it is still not a desirable
thing to do (just think about raid and other complex filesystems that may
also have to do allocations).
Swap IO is never desirable but it has to happen somehow and right now, we
only initiate swap IO from direct reclaim or kswapd. For complex filesystems,
it is mandatory if they are using direct_IO like I do for NFS that they
pin the necessary structures in advance to avoid any allocations in their
reclaim path. I do not expect RAID to be used over network-based swap files.
I do not think that block device writers
code with the VM in mind. In the case of network devices as block devices
we have a pretty serious problem since the network subsystem is certainly
not designed to be called from VM reclaim code that may be triggered
arbitrarily from deeply nested other code in the kernel. Implementing
something like this invites breakage all over the place to show up.
The whole point of the series is to allow the possibility of using
network-based swap devices starting with NBD and with NFS in the related
series. swap-over-NFS has been used for the last few years by enterprise
distros and while bugs do get reported, they are rare.
As the person that introduced this, I would support it in mainline for
NBD and NFS if it was merged.
@@ -1221,6 +1222,7 @@ void free_hot_cold_page(struct page *page, int cold)migratetype=get_pageblock_migratetype(page);set_page_private(page,migratetype);+page->pfmemalloc=false;local_irq_save(flags);if(unlikely(wasMlocked))free_page_mlock(page);
page allocator hotpaths affected.
I can remove these but then page->pfmemalloc has to be set on the allocation
side. It's a single write to a dirty cache line that is already local to
the processor. It's not measurable although I accept that the page
allocator paths could do with a diet in general.
Implying an additional cacheline use in critical slab paths?
This was the alternative to altering the slub structures.
Hopefully grouped with other variables already in cache.
This is the expectation. I considered tagging it read_mostly but didn't
at the time. I will now because it is genuinely expected to be
read-mostly and in the case where it is being written to, we're also
using network-based swap and the cost of a cache miss will be negligible
in comparison to swapping under memory pressure to a network.
The impact of ac_put_obj() is reduced in a later patch and becomes just
an additional read of a global variable. There was not an obvious way to
me to ensure pfmemalloc pages were not used for !pfmemalloc allocations
without having some sort of impact.
Modification to hotpath. That could be fixed here by forcing pfmemalloc
(like debug allocs) to always go to the slow path and checking in there
instead. Just keep c->freelist == NULL.
I picked up your patch for this, thanks.
--
Mel Gorman
SUSE Labs
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Fight unfair telecom internet charges in Canada: sign http://stopthemeter.ca/
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>