From: Chris Metcalf <hidden> Date: 2016-07-14 20:48:36
Here is a respin of the task-isolation patch set. This primarily
reflects feedback from Frederic and Peter Z.
Changes since v12:
- Rebased on v4.7-rc7.
- New default "strict" model for task isolation - tasks exit the
kernel from the initial prctl() to userspace, and can only legally
exit by calling prctl() again to turn off isolation. Any other
kernel entry results in a SIGKILL by default.
- New optional "relaxed" mode, where the application can receive some
signal other than SIGKILL, or no signal at all, when it re-enters
the kernel. Since by default task isolation is now strict, there is
no longer an additional "STRICT" mode, but rather a new "NOSIG" mode
that builds on top of the "USERSIG" support for setting a signal
other than SIGKILL to be delivered to the process. The "NOSIG" mode
also relaxes the required criteria for entering task isolation mode;
we just issue a warning if the affinity isn't set right, and we
don't fail with EAGAIN if the kernel isn't ready to stop the tick.
Running your task-isolation application in this "NOSIG" mode is also
necessary when debugging, since otherwise hitting breakpoints, etc.,
will cause a fatal signal to be sent to the process.
Frederic has suggested we might want to defer this functionality
until later, but (in addition to the debuggability aspect) there is
some thought that it might be useful for e.g. HPC, so I have just
broken out the additional semantics into a single separate patch at
the end of the series.
- Function naming has been changed and comments have been added to try
to clarify the role of the task-isolation reporting on kernel
entries that do NOT cause signals. This hopefully clarifies why we
only invoke the renamed task_isolation_quiet_exception() in a few
places, since all the other places generate signals anyway. [PeterZ]
- The task_isolation_debug() call now has an inline piece that checks
to see if the target is a task_isolation cpu before actually
calling. [PeterZ]
- In _task_isolation_debug(), we use the new task_struct_trylock()
call that is in linux-next now; for now I just have a static copy of
the function, which I will switch to using the version from
linux-next in the next rebasing. [PeterZ]
- We now pass a string describing the interrupt up from
task_isolation_debug() so there is more information on where the
interrupt came from beyond just the stack backtrace. [PeterZ]
- I added task_isolation_debug() hooks to smp_sched_reschedule() on
x86, which was missing before, and removed the hooks in the tile
send_IPI_*() routines, since there were already hooks in the
callers. Likewise I moved the hook for arm64 from the generic
smp_cross_call() routine to the only caller that wasn't already
hooked, smp_send_reschedule(). The commit message clarifies the
rationale for where hooks are placed.
- I moved the page fault reporting so that it only reports in the case
that we are not also sending a SIGSEGV/SIGBUS, for consistency with
other uses of task_isolation_quiet_exception().
The previous (v12) patch series is here:
https://lkml.kernel.org/g/1459877922-15512-1-git-send-email-cmetcalf@mellanox.com
This version of the patch series has been tested on arm64 and tilegx,
and build-tested on x86.
It remains true that the 1 Hz tick needs to be disabled for this
patch series to be able to achieve its primary goal of enabling
truly tick-free operation, but that is ongoing orthogonal work.
Frederick, do you have a sense of what is left to be done there?
I can certainly try to contribute to that effort as well.
The series is available at:
git://git.kernel.org/pub/scm/linux/kernel/git/cmetcalf/linux-tile.git dataplane
Chris Metcalf (12):
vmstat: add quiet_vmstat_sync function
vmstat: add vmstat_idle function
lru_add_drain_all: factor out lru_add_drain_needed
task_isolation: add initial support
task_isolation: track asynchronous interrupts
arch/x86: enable task isolation functionality
arm64: factor work_pending state machine to C
arch/arm64: enable task isolation functionality
arch/tile: enable task isolation functionality
arm, tile: turn off timer tick for oneshot_stopped state
task_isolation: support CONFIG_TASK_ISOLATION_ALL
task_isolation: add user-settable notification signal
Documentation/kernel-parameters.txt | 16 ++
arch/arm64/Kconfig | 1 +
arch/arm64/include/asm/thread_info.h | 5 +-
arch/arm64/kernel/entry.S | 12 +-
arch/arm64/kernel/ptrace.c | 15 +-
arch/arm64/kernel/signal.c | 42 +++-
arch/arm64/kernel/smp.c | 2 +
arch/arm64/mm/fault.c | 8 +-
arch/tile/Kconfig | 1 +
arch/tile/include/asm/thread_info.h | 4 +-
arch/tile/kernel/process.c | 9 +
arch/tile/kernel/ptrace.c | 7 +
arch/tile/kernel/single_step.c | 7 +
arch/tile/kernel/smp.c | 26 +--
arch/tile/kernel/time.c | 1 +
arch/tile/kernel/unaligned.c | 4 +
arch/tile/mm/fault.c | 13 +-
arch/tile/mm/homecache.c | 2 +
arch/x86/Kconfig | 1 +
arch/x86/entry/common.c | 18 +-
arch/x86/include/asm/thread_info.h | 2 +
arch/x86/kernel/smp.c | 2 +
arch/x86/kernel/traps.c | 3 +
arch/x86/mm/fault.c | 5 +
drivers/base/cpu.c | 18 ++
drivers/clocksource/arm_arch_timer.c | 2 +
include/linux/context_tracking_state.h | 6 +
include/linux/isolation.h | 73 +++++++
include/linux/sched.h | 3 +
include/linux/swap.h | 1 +
include/linux/tick.h | 2 +
include/linux/vmstat.h | 4 +
include/uapi/linux/prctl.h | 10 +
init/Kconfig | 37 ++++
kernel/Makefile | 1 +
kernel/fork.c | 3 +
kernel/irq_work.c | 5 +-
kernel/isolation.c | 337 +++++++++++++++++++++++++++++++++
kernel/sched/core.c | 42 ++++
kernel/signal.c | 15 ++
kernel/smp.c | 6 +-
kernel/softirq.c | 33 ++++
kernel/sys.c | 9 +
kernel/time/tick-sched.c | 36 ++--
mm/swap.c | 15 +-
mm/vmstat.c | 19 ++
46 files changed, 827 insertions(+), 56 deletions(-)
create mode 100644 include/linux/isolation.h
create mode 100644 kernel/isolation.c
--
2.7.2
From: Chris Metcalf <hidden> Date: 2016-07-14 20:48:50
The existing nohz_full mode is designed as a "soft" isolation mode
that makes tradeoffs to minimize userspace interruptions while
still attempting to avoid overheads in the kernel entry/exit path,
to provide 100% kernel semantics, etc.
However, some applications require a "hard" commitment from the
kernel to avoid interruptions, in particular userspace device driver
style applications, such as high-speed networking code.
This change introduces a framework to allow applications
to elect to have the "hard" semantics as needed, specifying
prctl(PR_SET_TASK_ISOLATION, PR_TASK_ISOLATION_ENABLE) to do so.
Subsequent commits will add additional flags and additional
semantics.
The kernel must be built with the new TASK_ISOLATION Kconfig flag
to enable this mode, and the kernel booted with an appropriate
task_isolation=CPULIST boot argument, which enables nohz_full and
isolcpus as well. The "task_isolation" state is then indicated by
setting a new task struct field, task_isolation_flag, to the value
passed by prctl(), and also setting a TIF_TASK_ISOLATION bit in
thread_info flags. When task isolation is enabled for a task, and it
is returning to userspace on a task isolation core, it calls the
new task_isolation_ready() / task_isolation_enter() routines to
take additional actions to help the task avoid being interrupted
in the future.
The task_isolation_ready() call is invoked when TIF_TASK_ISOLATION is
set in prepare_exit_to_usermode() or its architectural equivalent,
and forces the loop to retry if the system is not ready. It is
called with interrupts disabled and inspects the kernel state
to determine if it is safe to return into an isolated state.
In particular, if it sees that the scheduler tick is still enabled,
it reports that it is not yet safe.
Each time through the loop of TIF work to do, if TIF_TASK_ISOLATION
is set, we call the new task_isolation_enter() routine. This
takes any actions that might avoid a future interrupt to the core,
such as a worker thread being scheduled that could be quiesced now
(e.g. the vmstat worker) or a future IPI to the core to clean up some
state that could be cleaned up now (e.g. the mm lru per-cpu cache).
In addition, it reqeusts rescheduling if the scheduler dyntick is
still running.
Once the task has returned to userspace after issuing the prctl(),
if it enters the kernel again via system call, page fault, or any
of a number of other synchronous traps, the kernel will kill it
with SIGKILL. For system calls, this test is performed immediately
before the SECCOMP test and causes the syscall to return immediately
with ENOSYS.
To allow the state to be entered and exited, the syscall checking
test ignores the prctl() syscall so that we can clear the bit again
later, and ignores exit/exit_group to allow exiting the task without
a pointless signal killing you as you try to do so.
A new /sys/devices/system/cpu/task_isolation pseudo-file is added,
parallel to the comparable nohz_full file.
Separate patches that follow provide these changes for x86, tile,
and arm64.
Signed-off-by: Chris Metcalf <redacted>
---
Documentation/kernel-parameters.txt | 8 ++
drivers/base/cpu.c | 18 +++
include/linux/isolation.h | 60 ++++++++++
include/linux/sched.h | 3 +
include/linux/tick.h | 2 +
include/uapi/linux/prctl.h | 5 +
init/Kconfig | 27 +++++
kernel/Makefile | 1 +
kernel/fork.c | 3 +
kernel/isolation.c | 217 ++++++++++++++++++++++++++++++++++++
kernel/signal.c | 8 ++
kernel/sys.c | 9 ++
kernel/time/tick-sched.c | 36 +++---
13 files changed, 384 insertions(+), 13 deletions(-)
create mode 100644 include/linux/isolation.h
create mode 100644 kernel/isolation.c
@@ -3892,6 +3892,14 @@ bytes respectively. Such letter suffixes can also be entirely omitted. neutralize any effect of /proc/sys/kernel/sysrq. Useful for debugging.+ task_isolation= [KNL]+ In kernels built with CONFIG_TASK_ISOLATION=y, set+ the specified list of CPUs where cpus will be able+ to use prctl(PR_SET_TASK_ISOLATION) to set up task+ isolation mode. Setting this boot flag implicitly+ also sets up nohz_full and isolcpus mode for the+ listed set of cpus.+ tcpmhash_entries= [KNL,NET] Set the number of tcp_metrics_hash slots. Default value is 8192 or 16384 depending on total
@@ -0,0 +1,60 @@+/*+*Taskisolationrelatedglobalfunctions+*/+#ifndef _LINUX_ISOLATION_H+#define _LINUX_ISOLATION_H++#include<linux/tick.h>+#include<linux/prctl.h>++#ifdef CONFIG_TASK_ISOLATION++/* cpus that are configured to support task isolation */+externcpumask_var_ttask_isolation_map;++externinttask_isolation_init(void);++staticinlinebooltask_isolation_possible(intcpu)+{+returntask_isolation_map!=NULL&&+cpumask_test_cpu(cpu,task_isolation_map);+}++externinttask_isolation_set(unsignedintflags);++externbooltask_isolation_ready(void);+externvoidtask_isolation_enter(void);++staticinlinevoidtask_isolation_set_flags(structtask_struct*p,+unsignedintflags)+{+p->task_isolation_flags=flags;++if(flags&PR_TASK_ISOLATION_ENABLE)+set_tsk_thread_flag(p,TIF_TASK_ISOLATION);+else+clear_tsk_thread_flag(p,TIF_TASK_ISOLATION);+}++externinttask_isolation_syscall(intnr);++/* Report on exceptions that don't cause a signal for the user process. */+externvoid_task_isolation_quiet_exception(constchar*fmt,...);+#define task_isolation_quiet_exception(fmt, ...) \+do{\+if(current_thread_info()->flags&_TIF_TASK_ISOLATION)\+_task_isolation_quiet_exception(fmt,##__VA_ARGS__);\+}while(0)++#else+staticinlinevoidtask_isolation_init(void){}+staticinlinebooltask_isolation_possible(intcpu){returnfalse;}+staticinlinebooltask_isolation_ready(void){returntrue;}+staticinlinevoidtask_isolation_enter(void){}+externinlinevoidtask_isolation_set_flags(structtask_struct*p,+unsignedintflags){}+staticinlineinttask_isolation_syscall(intnr){return0;}+staticinlinevoidtask_isolation_quiet_exception(constchar*fmt,...){}+#endif++#endif
@@ -1918,6 +1918,9 @@ struct task_struct {#ifdef CONFIG_MMUstructtask_struct*oom_reaper_list;#endif+#ifdef CONFIG_TASK_ISOLATION+unsignedinttask_isolation_flags;+#endif/* CPU-specific state of this task */structthread_structthread;/*
@@ -783,6 +783,33 @@ config RCU_EXPEDITE_BOOTendmenu# "RCU Subsystem"+configHAVE_ARCH_TASK_ISOLATION+bool++configTASK_ISOLATION+bool"Provide hard CPU isolation from the kernel on demand"+depends onNO_HZ_FULL&&HAVE_ARCH_TASK_ISOLATION+help+Allowuserspaceprocessestoplacethemselvesontask_isolation+coresandrunprctl(PR_SET_TASK_ISOLATION)to"isolate"+themselvesfromthekernel.Priortoreturningtouserspace,+isolatedtaskswillarrangethatnofuturekernel+activitywillinterruptthetaskwhilethetaskisrunning+inuserspace.Bydefault,attemptingtore-enterthekernel+whileinthismodewillcausethetasktobeterminated+withasignal;youmustexplicitlyuseprctl()todisable+taskisolationbeforeresumingnormaluseofthekernel.++This"hard"isolationfromthekernelisrequiredfor+userspacetasksthatarerunninghardreal-timetasksin+userspace,suchasa10Gbitnetworkdriverinuserspace.+Withoutthisoption,butwithNO_HZ_FULLenabled,thekernel+willmakeabest-faith,"soft"efforttoshieldasingleuserspace+processfrominterrupts,butmakesnoguarantees.++Youshouldsay"N"unlessyouareintendingtoruna+high-performanceuserspacedriverorsimilartask.+configBUILD_BIN2Cbooldefaultn
@@ -1535,6 +1536,8 @@ static struct task_struct *copy_process(unsigned long clone_flags,#endifclear_all_latency_tracing(p);+task_isolation_set_flags(p,0);+/* ok, now we should be set up.. */p->pid=pid_nr(pid);if(clone_flags&CLONE_THREAD){
@@ -0,0 +1,217 @@+/*+*linux/kernel/isolation.c+*+*Implementationfortaskisolation.+*+*DistributedunderGPLv2.+*/++#include<linux/mm.h>+#include<linux/swap.h>+#include<linux/vmstat.h>+#include<linux/isolation.h>+#include<linux/syscalls.h>+#include<asm/unistd.h>+#include<asm/syscall.h>+#include"time/tick-sched.h"++cpumask_var_ttask_isolation_map;+staticboolsaw_boot_arg;++/*+*Isolationrequiresbothnohzandisolcpussupportfromthescheduler.+*Weprovideabootflagthatenablesbothfornow,andwhichwecan+*addotherfunctionalitytoovertimeifneeded.Notethatjust+*specifying"nohz_full=... isolcpus=..."doesnotenabletaskisolation.+*/+staticint__inittask_isolation_setup(char*str)+{+saw_boot_arg=true;++alloc_bootmem_cpumask_var(&task_isolation_map);+if(cpulist_parse(str,task_isolation_map)<0){+pr_warn("task_isolation: Incorrect cpumask '%s'\n",str);+return1;+}++return1;+}+__setup("task_isolation=",task_isolation_setup);++int__inittask_isolation_init(void)+{+/* For offstack cpumask, ensure we allocate an empty cpumask early. */+if(!saw_boot_arg){+zalloc_cpumask_var(&task_isolation_map,GFP_KERNEL);+return0;+}++/*+*Addourtask_isolationcpustonohz_fullandisolcpus.Note+*thatwearecalledrelativelyearlyinboot,fromtick_init();+*atthispointneithernohz_fullnorisolcpushasbeenused+*toconfigurethesystem,butisolcpushasbeenallocated+*alreadyinsched_init().+*/+tick_nohz_full_add_cpus(task_isolation_map);+cpumask_or(cpu_isolated_map,cpu_isolated_map,task_isolation_map);++return0;+}++/*+*Getasnapshotofwhether,atthismoment,itwouldbepossibleto+*stopthetick.Thistestnormallyrequiresinterruptsdisabledsince+*theconditioncanchangeifaninterruptisdelivered.However,in+*thiscaseweareusingitinanadvisorycapacitytoseeifthere+*isanythingobviouslyindicatingthatthetaskisolation+*preconditionshavenotbeenmet,soit'sOKthatinprincipleit+*mightnotstillbetruelaterintheprctl()syscallpath.+*/+staticboolcan_stop_my_full_tick_now(void)+{+boolret;++local_irq_disable();+ret=can_stop_my_full_tick();+local_irq_enable();+returnret;+}++/*+*Thisroutinecontrolswhetherwecanenabletask-isolationmode.+*Thetaskmustbeaffinitizedtoasingletask_isolationcore,or+*elsewereturnEINVAL.And,itmustbeatleaststaticallyableto+*stopthenohz_fulltick(e.g.,nootherschedulabletaskscurrently+*running,noPOSIXcputimerscurrentlysetup,etc.);ifnot,we+*returnEAGAIN.+*/+inttask_isolation_set(unsignedintflags)+{+if(flags!=0){+if(cpumask_weight(tsk_cpus_allowed(current))!=1||+!task_isolation_possible(raw_smp_processor_id())){+/* Invalid task affinity setting. */+return-EINVAL;+}+if(!can_stop_my_full_tick_now()){+/* System not yet ready for task isolation. */+return-EAGAIN;+}+}++task_isolation_set_flags(current,flags);+return0;+}++/*+*Intaskisolationmodewetrytoreturntouserspaceonlyafter+*attemptingtomakesurewewon'tbeinterruptedagain.Thistest+*isrunwithinterruptsdisabledtotestthateverythingweneed+*tobetrueistruebeforewecanreturntouserspace.+*/+booltask_isolation_ready(void)+{+WARN_ON_ONCE(!irqs_disabled());++return(!lru_add_drain_needed(smp_processor_id())&&+vmstat_idle()&&+tick_nohz_tick_stopped());+}++/*+*Eachtimewetrytoprepareforreturntouserspaceinaprocess+*withtaskisolationenabled,werunthiscodetoquiescewhatever+*subsystemswecanreadilyquiescetoavoidlaterinterrupts.+*/+voidtask_isolation_enter(void)+{+WARN_ON_ONCE(irqs_disabled());++/* Drain the pagevecs to avoid unnecessary IPI flushes later. */+lru_add_drain();++/* Quieten the vmstat worker so it won't interrupt us. */+quiet_vmstat_sync();++/*+*Requestreschedulingunlessweareinfulldynticksmode.+*Wewouldeventuallygetpre-emptedwithoutthis,andif+*there'sanothertaskwaiting,itwouldrun;butby+*explicitlyrequestingthereschedule,wemayreducethe+*latency.Wecoulddirectlycallschedule()hereaswell,+*butsinceourcalleristhestandardplacewhereschedule()+*iscalled,wedefertothecaller.+*+*Amoresubstantiveapproachherewouldbetouseastruct+*completionhereexplicitly,andcompleteitwhenweshut+*downdynticks,butsincewepresumablyhavenothingbetter+*todoonthiscoreanyway,justspinningseemsplausible.+*/+if(!tick_nohz_tick_stopped())+set_tsk_need_resched(current);+}++staticvoidtask_isolation_deliver_signal(structtask_struct*task,+constchar*buf)+{+siginfo_tinfo={};++info.si_signo=SIGKILL;++/*+*Reportonthefactthatisolationwasviolatedforthetask.+*Itmaynotbethetask'sfault(e.g.aTLBflushfromanother+*core)butwearenotblamingit,justreportingthatitlost+*itsisolationstatus.+*/+pr_warn("%s/%d: task_isolation mode lost due to %s\n",+task->comm,task->pid,buf);++/* Turn off task isolation mode to avoid further isolation callbacks. */+task_isolation_set_flags(task,0);++send_sig_info(info.si_signo,&info,task);+}++/*+*Thisroutineiscalledfromanyuserspaceexceptionthatdoesn't+*otherwisetriggerasignaltotheuserprocess(e.g.simplepagefault).+*/+void_task_isolation_quiet_exception(constchar*fmt,...)+{+structtask_struct*task=current;+va_listargs;+charbuf[100];++/* RCU should have been enabled prior to this point. */+RCU_LOCKDEP_WARN(!rcu_is_watching(),"kernel entry without RCU");++va_start(args,fmt);+vsnprintf(buf,sizeof(buf),fmt,args);+va_end(args);++task_isolation_deliver_signal(task,buf);+}++/*+*Thisroutineiscalledfromsyscallentry(withthesyscallnumber+*passedin),andpreventsmostsyscallsfromexecutingandraisesa+*signaltonotifytheprocess.+*/+inttask_isolation_syscall(intsyscall)+{+charbuf[20];++if(syscall==__NR_prctl||+syscall==__NR_exit||+syscall==__NR_exit_group)+return0;++snprintf(buf,sizeof(buf),"syscall %d",syscall);+task_isolation_deliver_signal(current,buf);++syscall_set_return_value(current,current_pt_regs(),+-ERESTARTNOINTR,-1);+return-1;+}
--
2.7.2
--
To unsubscribe, send a message with 'unsubscribe linux-mm' in
the body to majordomo@kvack.org. For more info on Linux MM,
see: http://www.linux-mm.org/ .
Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
From: Chris Metcalf <hidden> Date: 2016-07-14 20:49:37
By default, if a task in task isolation mode re-enters the kernel,
it is terminated with SIGKILL. With this commit, the application
can choose what signal to receive on a task isolation violation
by invoking prctl() with PR_TASK_ISOLATION_ENABLE, or'ing in the
PR_TASK_ISOLATION_USERSIG bit, and setting the specific requested
signal by or'ing in PR_TASK_ISOLATION_SET_SIG(sig).
This mode allows for catching the notification signal; for example,
in a production environment, it might be helpful to log information
to the application logging mechanism before exiting. Or, the
application might choose to re-enable task isolation and return to
continue execution.
As a special case, the user may set the signal to 0, which means
that no signal will be delivered. In this mode, the application
may freely enter the kernel for syscalls and synchronous exceptions
such as page faults, but each time it will be held in the kernel
before returning to userspace until the kernel has quiesced timer
ticks or other potential future interruptions, just like it does
on return from the initial prctl() call. Note that in this mode,
the task can be migrated away from its initial task_isolation core,
and if it is migrated to a non-isolated core it will lose task
isolation until it is migrated back to an isolated core.
In addition, in this mode we no longer require the affinity to
be set correctly on entry (though we warn on the console if it's
not right), and we don't bother to notify the user that the kernel
isn't ready to quiesce either (since we'll presumably be in and
out of the kernel multiple times with task isolation enabled anyway).
The PR_TASK_ISOLATION_NOSIG define is provided as a convenience
wrapper to express this semantic.
Signed-off-by: Chris Metcalf <redacted>
---
include/uapi/linux/prctl.h | 5 ++++
kernel/isolation.c | 62 ++++++++++++++++++++++++++++++++++++++--------
2 files changed, 56 insertions(+), 11 deletions(-)
@@ -85,6 +85,15 @@ static bool can_stop_my_full_tick_now(void)returnret;}+/* Get the signal number that will be sent for a particular set of flag bits. */+staticinttask_isolation_sig(intflags)+{+if(flags&PR_TASK_ISOLATION_USERSIG)+returnPR_TASK_ISOLATION_GET_SIG(flags);+else+returnSIGKILL;+}+/**Thisroutinecontrolswhetherwecanenabletask-isolationmode.*Thetaskmustbeaffinitizedtoasingletask_isolationcore,or
@@ -92,16 +101,30 @@ static bool can_stop_my_full_tick_now(void)*stopthenohz_fulltick(e.g.,nootherschedulabletaskscurrently*running,noPOSIXcputimerscurrentlysetup,etc.);ifnot,we*returnEAGAIN.+*+*Ifwewillnotbestrictlyenforcingkernelre-entrywithasignal,+*wejustgenerateawarningprintkifthereisabadaffinityset+*onentry(sinceafterallyoucanalwayschangeitagainafteryou+*callprctl)andwedon'tbotherfailingtheprctlwith-EAGAIN+*sinceweassumeyouwillgoinandoutofkernelmodeanyway.*/inttask_isolation_set(unsignedintflags){if(flags!=0){+intsig=task_isolation_sig(flags);+if(cpumask_weight(tsk_cpus_allowed(current))!=1||!task_isolation_possible(raw_smp_processor_id())){/* Invalid task affinity setting. */-return-EINVAL;+if(sig)+return-EINVAL;+else+pr_warn("%s/%d: enabling non-signalling task isolation\n"+"and not bound to a single task isolation core\n",+current->comm,current->pid);}-if(!can_stop_my_full_tick_now()){++if(sig&&!can_stop_my_full_tick_now()){/* System not yet ready for task isolation. */return-EAGAIN;}
@@ -175,7 +198,10 @@ static void task_isolation_deliver_signal(struct task_struct *task,pr_warn("%s/%d: task_isolation mode lost due to %s\n",task->comm,task->pid,buf);-/* Turn off task isolation mode to avoid further isolation callbacks. */+/*+*Turnofftaskisolationmodetoavoidfurtherisolationcallbacks.+*Itcanchoosetore-enabletaskisolationmodeinthesignalhandler.+*/task_isolation_set_flags(task,0);send_sig_info(info.si_signo,&info,task);
@@ -190,15 +216,20 @@ void _task_isolation_quiet_exception(const char *fmt, ...)structtask_struct*task=current;va_listargs;charbuf[100];+intsig;/* RCU should have been enabled prior to this point. */RCU_LOCKDEP_WARN(!rcu_is_watching(),"kernel entry without RCU");+sig=task_isolation_sig(task->task_isolation_flags);+if(sig==0)+return;+va_start(args,fmt);vsnprintf(buf,sizeof(buf),fmt,args);va_end(args);-task_isolation_deliver_signal(task,buf);+task_isolation_deliver_signal(task,buf,sig);}/*
From: Andy Lutomirski <luto@amacapital.net> Date: 2016-07-14 21:03:47
On Thu, Jul 14, 2016 at 1:48 PM, Chris Metcalf [off-list ref] wrote:
Here is a respin of the task-isolation patch set. This primarily
reflects feedback from Frederic and Peter Z.
I still think this is the wrong approach, at least at this point. The
first step should be to instrument things if necessary and fix the
obvious cases where the kernel gets entered asynchronously. Only once
there's a credible reason to believe it can work well should any form
of strictness be applied.
As an example, enough vmalloc/vfree activity will eventually cause
flush_tlb_kernel_range to be called and *boom*, there goes your shiny
production dataplane application. Once virtually mapped kernel stacks
happen, the frequency with which this happens will only increase.
On very brief inspection, __kmem_cache_shutdown will be a problem on
some workloads as well.
--Andy
From: Chris Metcalf <hidden> Date: 2016-07-14 22:38:21
On 7/14/2016 5:03 PM, Andy Lutomirski wrote:
On Thu, Jul 14, 2016 at 1:48 PM, Chris Metcalf [off-list ref] wrote:
quoted
Here is a respin of the task-isolation patch set. This primarily
reflects feedback from Frederic and Peter Z.
I still think this is the wrong approach, at least at this point. The
first step should be to instrument things if necessary and fix the
obvious cases where the kernel gets entered asynchronously.
Note, however, that the task_isolation_debug mode is a very convenient
way of discovering what is going on when things do go wrong for task isolation.
Only once
there's a credible reason to believe it can work well should any form
of strictness be applied.
I'm not sure what criteria you need for this, though. Certainly we've been
shipping our version of task isolation to customers since 2008, and there
are quite a few customer applications in production that are working well.
I'd argue that's a credible reason.
As an example, enough vmalloc/vfree activity will eventually cause
flush_tlb_kernel_range to be called and *boom*, there goes your shiny
production dataplane application.
Well, that's actually a refinement that I did not inflict on this patch series.
In our code base, we have a hook for kernel TLB flushes that defers such
flushes for cores that are running in userspace, because, after all, they
don't yet care about such flushes. Instead, we atomically set a flag that
is checked on entry to the kernel, and that causes the TLB flush to occur
at that point.
On very brief inspection, __kmem_cache_shutdown will be a problem on
some workloads as well.
That looks like it should be amenable to a version of the same fix I pushed
upstream in 5fbc461636c32efd ("mm: make lru_add_drain_all() selective").
You would basically check which cores have non-empty caches, and only
interrupt those cores. For extra credit, you empty the cache on your local cpu
when you are entering task isolation mode. Now you don't get interrupted.
To be fair, I've never seen this particular path cause an interruption. And I
think this speaks to the fact that there really can't be a black and white
decision about when you have removed enough possible interrupt paths.
It really does depend on what else is running on your machine in addition
to the task isolation code, and that will vary from application to application.
And, as the kernel evolves, new ways of interrupting task isolation cores
will get added and need to be dealt with. There really isn't a perfect time
you can wait for and then declare that all the asynchronous entry cases
have been dealt with and now things are safe for task isolation.
--
Chris Metcalf, Mellanox Technologies
http://www.mellanox.com
From: Christoph Lameter <hidden> Date: 2016-07-18 00:42:58
On Thu, 14 Jul 2016, Andy Lutomirski wrote:
As an example, enough vmalloc/vfree activity will eventually cause
flush_tlb_kernel_range to be called and *boom*, there goes your shiny
production dataplane application. Once virtually mapped kernel stacks
happen, the frequency with which this happens will only increase.
But then vmalloc/vfre activity is not to be expected if user space only is
running. Since the kernel is not active and this affects kernel address
space only it could be deferred. Such events will cause OS activity that
causes a number of high latency events but then the system will quiet down
again.
On very brief inspection, __kmem_cache_shutdown will be a problem on
some workloads as well.
These are all corner cases that can be worked on over time if they are
significant. The main issue here is to reduce the obvious and relatively
frequent causes for ticks and allow easier detection of events that cause
tick activity.
From: Andy Lutomirski <luto@amacapital.net> Date: 2016-07-18 22:12:06
On Thu, Jul 14, 2016 at 2:22 PM, Chris Metcalf [off-list ref] wrote:
On 7/14/2016 5:03 PM, Andy Lutomirski wrote:
quoted
On Thu, Jul 14, 2016 at 1:48 PM, Chris Metcalf [off-list ref]
wrote:
quoted
Here is a respin of the task-isolation patch set. This primarily
reflects feedback from Frederic and Peter Z.
I still think this is the wrong approach, at least at this point. The
first step should be to instrument things if necessary and fix the
obvious cases where the kernel gets entered asynchronously.
Note, however, that the task_isolation_debug mode is a very convenient
way of discovering what is going on when things do go wrong for task
isolation.
quoted
Only once
there's a credible reason to believe it can work well should any form
of strictness be applied.
I'm not sure what criteria you need for this, though. Certainly we've been
shipping our version of task isolation to customers since 2008, and there
are quite a few customer applications in production that are working well.
I'd argue that's a credible reason.
quoted
As an example, enough vmalloc/vfree activity will eventually cause
flush_tlb_kernel_range to be called and *boom*, there goes your shiny
production dataplane application.
Well, that's actually a refinement that I did not inflict on this patch
series.
Submit it separately, perhaps?
The "kill the process if it goofs" think while there are known goofs
in the kernel, apparently with patches written but unsent, seems
questionable.
From: Chris Metcalf <hidden> Date: 2016-07-19 03:25:30
On 7/18/2016 6:11 PM, Andy Lutomirski wrote:
quoted
quoted
As an example, enough vmalloc/vfree activity will eventually cause
flush_tlb_kernel_range to be called and*boom*, there goes your shiny
production dataplane application.
Well, that's actually a refinement that I did not inflict on this patch
series.
Submit it separately, perhaps?
The "kill the process if it goofs" thing while there are known goofs
in the kernel, apparently with patches written but unsent, seems
questionable.
Sure, that's a good idea.
I think what I will plan to do is, once the patch series is accepted into
some tree, return to this piece. I'll have to go back and look at the internal
Tilera version of this code, since we have diverged quite a ways from that
in the 13 versions of the patch series, but my memory is that the kernel TLB
flush management was the only substantial piece of additional code not in
the initial batch of changes. The extra requirement is the need to have a
hook very early on in the kernel entry path that you can hook in all paths;
arm64 has the ct_user_exit macro and tile has the finish_interrupt_save macro,
but I'm not sure there's something equivalent on x86 to catch all entries.
It's worth noting that the typical target application for task isolation, though
(at least in our experience) is a pretty dedicated machine, with the primary
application running in task isolation mode almost all of the time, and so
you are generally in pretty good control of all aspects of the system, including
whether or not you are generating kernel TLB flushes from your non task
isolation cores. So I would argue the kernel TLB flush management piece is
an improvement to, not a requirement for, the main patch series.
--
Chris Metcalf, Mellanox Technologies
http://www.mellanox.com
From: Chris Metcalf <hidden> Date: 2016-07-21 14:07:35
On 7/20/2016 10:04 PM, Christoph Lameter wrote:
We are trying to test the patchset on x86 and are getting strange
backtraces and aborts. It seems that the cpu before the cpu we are running
on creates an irq_work event that causes a latency event on the next cpu.
This is weird. Is there a new round robin IPI feature in the kernel that I
am not aware of?
This seems to be from your clocksource declaring itself to be
unstable, and then scheduling work to safely remove that timer.
I haven't looked at this code before (in kernel/time/clocksource.c
under CONFIG_CLOCKSOURCE_WATCHDOG) since the timers on
arm64 and tile aren't unstable. Is it possible to boot your machine
with a stable clocksource?
From: Christoph Lameter <hidden> Date: 2016-07-22 02:20:41
On Thu, 21 Jul 2016, Chris Metcalf wrote:
On 7/20/2016 10:04 PM, Christoph Lameter wrote:
unstable, and then scheduling work to safely remove that timer.
I haven't looked at this code before (in kernel/time/clocksource.c
under CONFIG_CLOCKSOURCE_WATCHDOG) since the timers on
arm64 and tile aren't unstable. Is it possible to boot your machine
with a stable clocksource?
It already as a stable clocksource. Sorry but that was one of the criteria
for the server when we ordered them. Could this be clock adjustments?
From: Chris Metcalf <hidden> Date: 2016-07-22 14:24:22
On 7/21/2016 10:20 PM, Christoph Lameter wrote:
On Thu, 21 Jul 2016, Chris Metcalf wrote:
quoted
On 7/20/2016 10:04 PM, Christoph Lameter wrote:
unstable, and then scheduling work to safely remove that timer.
I haven't looked at this code before (in kernel/time/clocksource.c
under CONFIG_CLOCKSOURCE_WATCHDOG) since the timers on
arm64 and tile aren't unstable. Is it possible to boot your machine
with a stable clocksource?
It already as a stable clocksource. Sorry but that was one of the criteria
for the server when we ordered them. Could this be clock adjustments?
We probably need to get clock folks to jump in on this thread!
Maybe it's disabling some built-in unstable clock just as part of
falling back to using the better, stable clock that you also have?
So maybe there's a way of just disabling that clocksource from the
get-go instead of having it be marked unstable later.
If you run the test again after this storm of unstable marking, does
it all happen again? Or is it a persistent state in the kernel?
If so, maybe you can just arrange to get to that state before starting
your application's task-isolation code.
Or, if you think it's clock adjustments, perhaps running your test with
ntpd disabled would make it work better?
--
Chris Metcalf, Mellanox Technologies
http://www.mellanox.com
From: Christoph Lameter <hidden> Date: 2016-07-25 16:35:36
On Fri, 22 Jul 2016, Chris Metcalf wrote:
quoted
It already as a stable clocksource. Sorry but that was one of the criteria
for the server when we ordered them. Could this be clock adjustments?
We probably need to get clock folks to jump in on this thread!
Guess so. I will have a look at this when I get some time again.
Maybe it's disabling some built-in unstable clock just as part of
falling back to using the better, stable clock that you also have?
So maybe there's a way of just disabling that clocksource from the
get-go instead of having it be marked unstable later.
This is a standard Dell server. No clocksources are marked as unstable as
far as I can tell.
If you run the test again after this storm of unstable marking, does
it all happen again? Or is it a persistent state in the kernel?
This happens anytime we try to run with prctl().
I hope to get some more detail once I get some time to look at this. But
this is likely an x86 specific problem.
From: Christoph Lameter <hidden> Date: 2016-07-27 13:55:33
On Mon, 25 Jul 2016, Christoph Lameter wrote:
Guess so. I will have a look at this when I get some time again.
Ok so the problem is the clocksource_watchdog() function in
kernel/time/clocksource.c. This function is active if
CONFIG_CLOCKSOURCE_WATCHDOG is defined. It will check the timesources of
each processor for being within bounds and then reschedule itself on the
next one.
The purpose of the function seems to be to determine *if* a clocksource is
unstable. It does not mean that the clocksource *is* unstable.
The critical piece of code is this:
/*
* Cycle through CPUs to check if the CPUs stay synchronized
* to each other.
*/
next_cpu = cpumask_next(raw_smp_processor_id(), cpu_online_mask);
if (next_cpu >= nr_cpu_ids)
next_cpu = cpumask_first(cpu_online_mask);
watchdog_timer.expires += WATCHDOG_INTERVAL;
add_timer_on(&watchdog_timer, next_cpu);
Should we just cycle through the cpus that are not isolated? Otherwise we
need to have some means to check the clocksources for accuracy remotely
(probably impossible for TSC etc).
The WATCHDOG_INTERVAL is 1 second so this causes an interrupt every
second.
Note that we are running with the patch that removes the 1 HZ mininum time
tick. With an older kernel code base (redhat) we can keep the kernel quiet
for minutes. The clocksource watchdog causes timers to fire again.
From: Christoph Lameter <hidden> Date: 2016-07-27 15:23:18
On Wed, 27 Jul 2016, Chris Metcalf wrote:
quoted
Should we just cycle through the cpus that are not isolated? Otherwise we
need to have some means to check the clocksources for accuracy remotely
(probably impossible for TSC etc).
That sounds like the right idea - use the housekeeping cpu mask instead of the
cpu online mask. Should be a straightforward patch; do you want to do that
and test it in your configuration, and I'll include it in the next spin of the
patch series?
Sadly housekeeping_mask is defined the following way:
static inline const struct cpumask *housekeeping_cpumask(void)
{
#ifdef CONFIG_NO_HZ_FULL
if (tick_nohz_full_enabled())
return housekeeping_mask;
#endif
return cpu_possible_mask;
}
Why is it not returning cpu_online_mask?
From: Christoph Lameter <hidden> Date: 2016-07-27 15:31:36
Ok here is a possible patch that explicitly checks for housekeeping cpus:
Subject: clocksource: Do not schedule watchdog on isolated or NOHZ cpus
watchdog checks can only run on housekeeping capable cpus. Otherwise
we will be generating noise that we would like to avoid on the isolated
processors.
Signed-off-by: Christoph Lameter <redacted>
Index: linux/kernel/time/clocksource.c
===================================================================
From: Chris Metcalf <hidden> Date: 2016-07-27 17:06:34
On 7/27/2016 11:31 AM, Christoph Lameter wrote:
quoted hunk
Ok here is a possible patch that explicitly checks for housekeeping cpus:
Subject: clocksource: Do not schedule watchdog on isolated or NOHZ cpus
watchdog checks can only run on housekeeping capable cpus. Otherwise
we will be generating noise that we would like to avoid on the isolated
processors.
Signed-off-by: Christoph Lameter <redacted>
Index: linux/kernel/time/clocksource.c
===================================================================
How about using cpumask_next_and(raw_smp_processor_id(), cpu_online_mask,
housekeeping_cpumask()), likewise cpumask_first_and()? Does that work?
Note that you should also cpumask_first_and() in clocksource_start_watchdog(),
just to be complete.
Hopefully the init code runs after tick_init(). It seems like that's probably true.
--
Chris Metcalf, Mellanox Technologies
http://www.mellanox.com
From: Christoph Lameter <hidden> Date: 2016-07-27 18:56:33
On Wed, 27 Jul 2016, Chris Metcalf wrote:
How about using cpumask_next_and(raw_smp_processor_id(), cpu_online_mask,
housekeeping_cpumask()), likewise cpumask_first_and()? Does that work?
Ok here is V2:
Subject: clocksource: Do not schedule watchdog on isolated or NOHZ cpus V2
watchdog checks can only run on housekeeping capable cpus. Otherwise
we will be generating noise that we would like to avoid on the isolated
processors.
Signed-off-by: Christoph Lameter <redacted>
Index: linux/kernel/time/clocksource.c
===================================================================
From: Christoph Lameter <hidden> Date: 2016-07-27 19:53:52
On Wed, 27 Jul 2016, Chris Metcalf wrote:
Looks good. Did you omit the equivalent fix in clocksource_start_watchdog()
on purpose? For now I just took your change, but tweaked it to add the
equivalent diff with cpumask_first_and() there.
Can the watchdog be started on an isolated cpu at all? I would expect that
the code would start a watchdog only on a housekeeping cpu.
From: Chris Metcalf <hidden> Date: 2016-07-27 20:04:50
On 7/27/2016 2:56 PM, Christoph Lameter wrote:
quoted hunk
On Wed, 27 Jul 2016, Chris Metcalf wrote:
quoted
How about using cpumask_next_and(raw_smp_processor_id(), cpu_online_mask,
housekeeping_cpumask()), likewise cpumask_first_and()? Does that work?
Ok here is V2:
Subject: clocksource: Do not schedule watchdog on isolated or NOHZ cpus V2
watchdog checks can only run on housekeeping capable cpus. Otherwise
we will be generating noise that we would like to avoid on the isolated
processors.
Signed-off-by: Christoph Lameter <redacted>
Index: linux/kernel/time/clocksource.c
===================================================================
Looks good. Did you omit the equivalent fix in clocksource_start_watchdog()
on purpose? For now I just took your change, but tweaked it to add the
equivalent diff with cpumask_first_and() there.
--
Chris Metcalf, Mellanox Technologies
http://www.mellanox.com
From: Chris Metcalf <hidden> Date: 2016-07-27 21:31:55
On 7/27/2016 3:53 PM, Christoph Lameter wrote:
On Wed, 27 Jul 2016, Chris Metcalf wrote:
quoted
Looks good. Did you omit the equivalent fix in clocksource_start_watchdog()
on purpose? For now I just took your change, but tweaked it to add the
equivalent diff with cpumask_first_and() there.
Can the watchdog be started on an isolated cpu at all? I would expect that
the code would start a watchdog only on a housekeeping cpu.
The code just starts the watchdog initially on the first online cpu.
In principle you could have configured that as an isolated cpu, so
without any change to that code, you'd interrupt that cpu.
I guess another way to slice it would be to start the watchdog on the
current core. But just using the same idiom as in clocksource_watchdog()
seems cleanest to me.
I added your patch to the series and pushed it up (along with adding your
Tested-by to the x86 enablement commit). It's still based on 4.6 so I'll need
to rebase it once the merge window closes.
--
Chris Metcalf, Mellanox Technologies
http://www.mellanox.com
From: Chris Metcalf <hidden> Date: 2016-07-28 08:11:46
On 7/27/2016 9:55 AM, Christoph Lameter wrote:
The critical piece of code is this:
/*
* Cycle through CPUs to check if the CPUs stay synchronized
* to each other.
*/
next_cpu = cpumask_next(raw_smp_processor_id(), cpu_online_mask);
if (next_cpu >= nr_cpu_ids)
next_cpu = cpumask_first(cpu_online_mask);
watchdog_timer.expires += WATCHDOG_INTERVAL;
add_timer_on(&watchdog_timer, next_cpu);
Should we just cycle through the cpus that are not isolated? Otherwise we
need to have some means to check the clocksources for accuracy remotely
(probably impossible for TSC etc).
That sounds like the right idea - use the housekeeping cpu mask instead of the
cpu online mask. Should be a straightforward patch; do you want to do that
and test it in your configuration, and I'll include it in the next spin of the
patch series?
Thanks for your testing!
--
Chris Metcalf, Mellanox Technologies
http://www.mellanox.com
From: Francis Giraldeau <hidden> Date: 2016-07-29 18:31:55
I tested this patch on 4.7 and confirm that irq_work does not occurs anymore on
the isolated cpu. Thanks!
I don't know of any utility to test the task isolation feature, so I started
one:
https://github.com/giraldeau/taskisol
The script exp.sh runs the taskisol to test five different conditions, but some
behavior is not the one I would expect.
At startup, it does:
- register a custom signal handler for SIGUSR1
- sched_setaffinity() on CPU 1, which is isolated
- mlockall(MCL_CURRENT) to prevent undesired page faults
The default strict mode is set with:
prctl(PR_SET_TASK_ISOLATION, PR_TASK_ISOLATION_ENABLE)
And then, the syscall write() is called. From previous discussion, the SIGKILL
should be sent, but it does not occur. When instead of calling write() we force
a page fault, then the SIGKILL is correctly sent.
When instead a custom signal handler SIGUSR1:
prctl(PR_SET_TASK_ISOLATION, PR_TASK_ISOLATION_USERSIG |
PR_TASK_ISOLATION_SET_SIG(SIGUSR1)
The signal is never delivered, either when the syscall is issued nor when the
page fault occurs.
I can confirm that, if two taskisol are created on the same CPU, the second one
fails with Resource temporarily unavailable, so that's fine.
I can add more test cases depending on your comments, such as the TLB events
triggered by another thread on a non-isolated core. But maybe there is already
a test suite?
Francis
2016-07-27 15:58 GMT-04:00 Chris Metcalf [off-list ref]:
On 7/27/2016 3:53 PM, Christoph Lameter wrote:
quoted
On Wed, 27 Jul 2016, Chris Metcalf wrote:
quoted
Looks good. Did you omit the equivalent fix in
clocksource_start_watchdog()
on purpose? For now I just took your change, but tweaked it to add the
equivalent diff with cpumask_first_and() there.
Can the watchdog be started on an isolated cpu at all? I would expect that
the code would start a watchdog only on a housekeeping cpu.
The code just starts the watchdog initially on the first online cpu.
In principle you could have configured that as an isolated cpu, so
without any change to that code, you'd interrupt that cpu.
I guess another way to slice it would be to start the watchdog on the
current core. But just using the same idiom as in clocksource_watchdog()
seems cleanest to me.
I added your patch to the series and pushed it up (along with adding your
Tested-by to the x86 enablement commit). It's still based on 4.6 so I'll
need
to rebase it once the merge window closes.
--
Chris Metcalf, Mellanox Technologies
http://www.mellanox.com
From: Chris Metcalf <hidden> Date: 2016-07-29 21:20:06
On 7/29/2016 2:31 PM, Francis Giraldeau wrote:
I tested this patch on 4.7 and confirm that irq_work does not occurs anymore on
the isolated cpu. Thanks!
Great! Let me know if you'd like me to add your Tested-by in the patch series.
I don't know of any utility to test the task isolation feature, so I started
one:
https://github.com/giraldeau/taskisol
The script exp.sh runs the taskisol to test five different conditions, but some
behavior is not the one I would expect.
At startup, it does:
- register a custom signal handler for SIGUSR1
- sched_setaffinity() on CPU 1, which is isolated
- mlockall(MCL_CURRENT) to prevent undesired page faults
The default strict mode is set with:
prctl(PR_SET_TASK_ISOLATION, PR_TASK_ISOLATION_ENABLE)
And then, the syscall write() is called. From previous discussion, the SIGKILL
should be sent, but it does not occur. When instead of calling write() we force
a page fault, then the SIGKILL is correctly sent.
This looks like it may be a bug in the x86-specific part of the kernel support.
On tilegx and arm64, running your test does the right thing:
# ./taskisol default syscall
taskisol run
taskisol/1855: task_isolation mode lost due to syscall 64
Killed
I think the x86 support doesn't properly return right away from a bad
syscall. The patch below should fix that; can you try it? However, it's
not clear to me why the signal isn't getting delivered. Perhaps you can
try adding some tracing to the syscall_trace_enter() path and see if we're
actually running this code as expected? Thank you! :-)
@@ -90,8 +90,10 @@ unsigned long syscall_trace_enter_phase1(struct pt_regs *regs, u32 arch)/* In isolation mode, we may prevent the syscall from running. */if(work&_TIF_TASK_ISOLATION){-if(task_isolation_syscall(regs->orig_ax)==-1)-return-1;+if(task_isolation_syscall(regs->orig_ax)==-1){+regs->orig_ax=-1;+return0;+}work&=~_TIF_TASK_ISOLATION;}
I updated my dataplane branch on kernel.org with this fix.
When instead a custom signal handler SIGUSR1:
prctl(PR_SET_TASK_ISOLATION, PR_TASK_ISOLATION_USERSIG |
PR_TASK_ISOLATION_SET_SIG(SIGUSR1)
The signal is never delivered, either when the syscall is issued nor when the
page fault occurs.
This is a bug in your test program. Try again with this fix:
--- a/taskisol.c+++ b/taskisol.c
@@ -79,8 +79,9 @@ int main(int argc, char *argv[])*TheprogramcompleteswhenusingUSERSIG,*butactuallynosignalisdelivered*/-if(strcmp(argv[1],"signal")==0){-if(prctl(PR_SET_TASK_ISOLATION,PR_TASK_ISOLATION_USERSIG|+elseif(strcmp(argv[1],"signal")==0){+if(prctl(PR_SET_TASK_ISOLATION,PR_TASK_ISOLATION_ENABLE|+PR_TASK_ISOLATION_USERSIG|PR_TASK_ISOLATION_SET_SIG(SIGUSR1))<0){perror("prctl sigusr");return-1;
The prctl() API is intended to be one-shot, i.e. you set all the state you
want with a single prctl(). The next call to prctl() will reset the state
to whatever you specify (including if you don't specify "enable").
(Also, as a side note, I'd expect your Makefile to invoke $(CC) for taskisol,
not $(CXX) - there doesn't seem to be any actual C++ in the program.)
I can confirm that, if two taskisol are created on the same CPU, the second one
fails with Resource temporarily unavailable, so that's fine.
I can add more test cases depending on your comments, such as the TLB events
triggered by another thread on a non-isolated core. But maybe there is already
a test suite?
The appended code is what I've been using as a test harness. It passes on
tilegx and arm64. No guarantees as to production-level code quality :-)
#define _GNU_SOURCE
#include <stdio.h>
#include <unistd.h>
#include <stdlib.h>
#include <fcntl.h>
#include <assert.h>
#include <string.h>
#include <errno.h>
#include <sched.h>
#include <pthread.h>
#include <sys/wait.h>
#include <sys/mman.h>
#include <sys/time.h>
#include <sys/prctl.h>
#ifndef PR_SET_TASK_ISOLATION // Not in system headers yet?
# define PR_SET_TASK_ISOLATION 48
# define PR_GET_TASK_ISOLATION 49
# define PR_TASK_ISOLATION_ENABLE (1 << 0)
# define PR_TASK_ISOLATION_USERSIG (1 << 1)
# define PR_TASK_ISOLATION_SET_SIG(sig) (((sig) & 0x7f) << 8)
# define PR_TASK_ISOLATION_GET_SIG(bits) (((bits) >> 8) & 0x7f)
# define PR_TASK_ISOLATION_NOSIG \
(PR_TASK_ISOLATION_USERSIG | PR_TASK_ISOLATION_SET_SIG(0))
#endif
// The cpu we are using for isolation tests.
static int task_isolation_cpu;
// Overall status, maintained as tests run.
static int exit_status = EXIT_SUCCESS;
// Set affinity to a single cpu.
int set_my_cpu(int cpu)
{
cpu_set_t set;
CPU_ZERO(&set);
CPU_SET(cpu, &set);
return sched_setaffinity(0, sizeof(cpu_set_t), &set);
}
// Run a child process in task isolation mode and report its status.
// The child does mlockall() and moves itself to the task isolation cpu.
// It then runs SETUP_FUNC (if specified), calls prctl(PR_SET_TASK_ISOLATION, )
// with FLAGS (if non-zero), and then invokes TEST_FUNC and exits
// with its status.
static int run_test(void (*setup_func)(), int (*test_func)(), int flags)
{
fflush(stdout);
int pid = fork();
assert(pid >= 0);
if (pid != 0) {
// In parent; wait for child and return its status.
int status;
waitpid(pid, &status, 0);
return status;
}
// In child.
int rc = mlockall(MCL_CURRENT);
assert(rc == 0);
rc = set_my_cpu(task_isolation_cpu);
assert(rc == 0);
if (setup_func)
setup_func();
if (flags) {
int rc;
do
rc = prctl(PR_SET_TASK_ISOLATION, flags);
while (rc != 0 && errno == EAGAIN);
if (rc != 0) {
printf("couldn't enable isolation (%d): FAIL\n", errno);
exit(EXIT_FAILURE);
}
}
rc = test_func();
exit(rc);
}
// Run a test and ensure it is killed with SIGKILL by default,
// for whatever misdemeanor is committed in TEST_FUNC.
// Also test it with SIGUSR1 as well to make sure that works.
static void test_killed(const char *testname, void (*setup_func)(),
int (*test_func)())
{
int status = run_test(setup_func, test_func, PR_TASK_ISOLATION_ENABLE);
if (WIFSIGNALED(status) && WTERMSIG(status) == SIGKILL) {
printf("%s: OK\n", testname);
} else {
printf("%s: FAIL (%#x)\n", testname, status);
exit_status = EXIT_FAILURE;
}
status = run_test(setup_func, test_func,
PR_TASK_ISOLATION_ENABLE | PR_TASK_ISOLATION_USERSIG |
PR_TASK_ISOLATION_SET_SIG(SIGUSR1));
if (WIFSIGNALED(status) && WTERMSIG(status) == SIGUSR1) {
printf("%s (SIGUSR1): OK\n", testname);
} else {
printf("%s (SIGUSR1): FAIL (%#x)\n", testname, status);
exit_status = EXIT_FAILURE;
}
}
// Run a test and make sure it exits with success.
static void test_ok(const char *testname, void (*setup_func)(),
int (*test_func)())
{
int status = run_test(setup_func, test_func, PR_TASK_ISOLATION_ENABLE);
if (status == EXIT_SUCCESS) {
printf("%s: OK\n", testname);
} else {
printf("%s: FAIL (%#x)\n", testname, status);
exit_status = EXIT_FAILURE;
}
}
// Run a test with no signals and make sure it exits with success.
static void test_nosig(const char *testname, void (*setup_func)(),
int (*test_func)())
{
int status =
run_test(setup_func, test_func,
PR_TASK_ISOLATION_ENABLE | PR_TASK_ISOLATION_NOSIG);
if (status == EXIT_SUCCESS) {
printf("%s: OK\n", testname);
} else {
printf("%s: FAIL (%#x)\n", testname, status);
exit_status = EXIT_FAILURE;
}
}
// Mapping address passed from setup function to test function.
static char *fault_file_mapping;
// mmap() a file in so we can test touching an unmapped page.
static void setup_fault(void)
{
char fault_file[] = "/tmp/isolation_XXXXXX";
int fd = mkstemp(fault_file);
assert(fd >= 0);
int rc = ftruncate(fd, getpagesize());
assert(rc == 0);
fault_file_mapping = mmap(NULL, getpagesize(), PROT_READ | PROT_WRITE,
MAP_SHARED, fd, 0);
assert(fault_file_mapping != MAP_FAILED);
close(fd);
unlink(fault_file);
}
// Now touch the unmapped page (and be killed).
static int do_fault(void)
{
*fault_file_mapping = 1;
return EXIT_FAILURE;
}
// Make a syscall (and be killed).
static int do_syscall(void)
{
write(STDOUT_FILENO, "goodbye, world\n", 13);
return EXIT_FAILURE;
}
// Turn isolation back off and don't be killed.
static int do_syscall_off(void)
{
prctl(PR_SET_TASK_ISOLATION, 0);
write(STDOUT_FILENO, "==> hello, world\n", 17);
return EXIT_SUCCESS;
}
// If we're not getting a signal, make sure we can do multiple system calls.
static int do_syscall_multi(void)
{
write(STDOUT_FILENO, "==> hello, world 1\n", 19);
write(STDOUT_FILENO, "==> hello, world 2\n", 19);
return EXIT_SUCCESS;
}
#ifdef __aarch64__
/* ARM64 uses tlbi instructions so doesn't need to interrupt the remote core. */
static void test_munmap(void) {}
#else
// Fork a thread that will munmap() after a short while.
// It will deliver a TLB flush to the task isolation core.
static void *start_munmap(void *p)
{
usleep(500000); // 0.5s
munmap(p, getpagesize());
return 0;
}
static void setup_munmap(void)
{
// First, go back to cpu 0 and allocate some memory.
set_my_cpu(0);
void *p = mmap(0, getpagesize(), PROT_READ|PROT_WRITE,
MAP_ANONYMOUS|MAP_POPULATE|MAP_PRIVATE, 0, 0);
assert(p != MAP_FAILED);
// Now fire up a thread that will wait half a second on cpu 0
// and then munmap the mapping.
pthread_t thr;
int rc = pthread_create(&thr, NULL, start_munmap, p);
assert(rc == 0);
// Back to the task-isolation cpu.
set_my_cpu(task_isolation_cpu);
}
// Global variable to avoid the compiler outsmarting us.
volatile int munmap_spin;
static int do_munmap(void)
{
while (munmap_spin < 1000000000)
++munmap_spin;
return EXIT_FAILURE;
}
static void test_munmap(void)
{
test_killed("test_munmap", setup_munmap, do_munmap);
}
#endif
#ifdef __tilegx__
// Make an unaligned access (and be killed).
// Only for tilegx, since other platforms don't do in-kernel fixups.
static int
do_unaligned(void)
{
static int buf[2];
volatile int* addr = (volatile int *)((char *)buf + 1);
*addr;
asm("nop");
return EXIT_FAILURE;
}
static void test_unaligned(void)
{
test_killed("test_unaligned", NULL, do_unaligned);
}
#else
static void test_unaligned(void) {}
#endif
// Fork a process that will spin annoyingly on the same core
// for a second. Since prctl() won't work if this task is actively
// running, we following this handshake sequence:
//
// 1. Child (in setup_quiesce, here) starts up, sets state 1 to let the
// parent know it's running, and starts doing short sleeps waiting on a
// state change.
// 2. Parent (in do_quiesce, below) starts up, spins waiting for state 1,
// then spins waiting on prctl() to succeed. At that point it is in
// isolation mode and the child is completing its most recent sleep.
// Now, as soon as the parent is scheduled out, it won't schedule back
// in until the child stops spinning.
// 3. Child sees the state change to 2, sets it to 3, and starts spinning
// waiting for a second to elapse, at which point it exits.
// 4. Parent spins waiting for the state to get to 3, then makes one
// syscall. This should take about a second even though the child
// was spinning for a whole second after changing the state to 3.
volatile int *statep, *childstate;
struct timeval quiesce_start, quiesce_end;
int child_pid;
static void setup_quiesce(void)
{
// First, go back to cpu 0 and allocate some shared memory.
set_my_cpu(0);
statep = mmap(0, getpagesize(), PROT_READ|PROT_WRITE,
MAP_ANONYMOUS|MAP_SHARED, 0, 0);
assert(statep != MAP_FAILED);
childstate = statep + 1;
gettimeofday(&quiesce_start, NULL);
// Fork and fault in all memory in both.
child_pid = fork();
assert(child_pid >= 0);
if (child_pid == 0)
*childstate = 1;
int rc = mlockall(MCL_CURRENT);
assert(rc == 0);
if (child_pid != 0) {
set_my_cpu(task_isolation_cpu);
return;
}
// In child. Wait until parent notifies us that it has completed
// its prctl, then jump to its cpu and let it know.
*childstate = 2;
while (*statep == 0)
;
*childstate = 3;
// printf("child: jumping to cpu %d\n", task_isolation_cpu);
set_my_cpu(task_isolation_cpu);
// printf("child: jumped to cpu %d\n", task_isolation_cpu);
*statep = 2;
*childstate = 4;
// Now we are competing for the runqueue on task_isolation_cpu.
// Spin for one second to ensure the parent gets caught in kernel space.
struct timeval start, tv;
gettimeofday(&start, NULL);
while (1) {
gettimeofday(&tv, NULL);
double time = (tv.tv_sec - start.tv_sec) +
(tv.tv_usec - start.tv_usec) / 1000000.0;
if (time >= 0.5)
exit(0);
}
}
static int do_quiesce(void)
{
double time;
int rc;
rc = prctl(PR_SET_TASK_ISOLATION,
PR_TASK_ISOLATION_ENABLE | PR_TASK_ISOLATION_NOSIG);
if (rc != 0) {
prctl(PR_SET_TASK_ISOLATION, 0);
printf("prctl failed: rc %d", rc);
goto fail;
}
*statep = 1;
// Wait for child to come disturb us.
while (*statep == 1) {
gettimeofday(&quiesce_end, NULL);
time = (quiesce_end.tv_sec - quiesce_start.tv_sec) +
(quiesce_end.tv_usec - quiesce_start.tv_usec)/1000000.0;
if (time > 0.1 && *statep == 1) {
prctl(PR_SET_TASK_ISOLATION, 0);
printf("timed out at %gs in child migrate loop (%d)\n",
time, *childstate);
char buf[100];
sprintf(buf, "cat /proc/%d/stack", child_pid);
system(buf);
goto fail;
}
}
assert(*statep == 2);
// At this point the child is spinning, so any interrupt will keep us
// in kernel space. Make a syscall to make sure it happens at least
// once during the second that the child is spinning.
kill(0, 0);
gettimeofday(&quiesce_end, NULL);
prctl(PR_SET_TASK_ISOLATION, 0);
time = (quiesce_end.tv_sec - quiesce_start.tv_sec) +
(quiesce_end.tv_usec - quiesce_start.tv_usec) / 1000000.0;
if (time < 0.4 || time > 0.6) {
printf("expected 1s wait after quiesce: was %g\n", time);
goto fail;
}
kill(child_pid, SIGKILL);
return EXIT_SUCCESS;
fail:
kill(child_pid, SIGKILL);
return EXIT_FAILURE;
}
int main(int argc, char **argv)
{
/* How many seconds to wait after running the other tests? */
double waittime;
if (argc == 1)
waittime = 10;
else if (argc == 2)
waittime = strtof(argv[1], NULL);
else {
printf("syntax: isolation [seconds]\n");
exit(EXIT_FAILURE);
}
/* Test that the /sys device is present and pick a cpu. */
FILE *f = fopen("/sys/devices/system/cpu/task_isolation", "r");
if (f == NULL) {
printf("/sys device: FAIL\n");
exit(EXIT_FAILURE);
}
char buf[100];
char *result = fgets(buf, sizeof(buf), f);
assert(result == buf);
fclose(f);
char *end;
task_isolation_cpu = strtol(buf, &end, 10);
assert(end != buf);
assert(*end == ',' || *end == '-' || *end == '\n');
assert(task_isolation_cpu >= 0);
printf("/sys device : OK\n");
// Test to see if with no mask set, we fail.
if (prctl(PR_SET_TASK_ISOLATION, PR_TASK_ISOLATION_ENABLE) == 0 ||
errno != EINVAL) {
printf("prctl unaffinitized: FAIL\n");
exit_status = EXIT_FAILURE;
} else {
printf("prctl unaffinitized: OK\n");
}
// Or if affinitized to the wrong cpu.
set_my_cpu(0);
if (prctl(PR_SET_TASK_ISOLATION, PR_TASK_ISOLATION_ENABLE) == 0 ||
errno != EINVAL) {
printf("prctl on cpu 0: FAIL\n");
exit_status = EXIT_FAILURE;
} else {
printf("prctl on cpu 0: OK\n");
}
// Run the tests.
test_killed("test_fault", setup_fault, do_fault);
test_killed("test_syscall", NULL, do_syscall);
test_munmap();
test_unaligned();
test_ok("test_off", NULL, do_syscall_off);
test_nosig("test_multi", NULL, do_syscall_multi);
test_nosig("test_quiesce", setup_quiesce, do_quiesce);
// Exit failure if any test failed.
if (exit_status != EXIT_SUCCESS)
return exit_status;
// Wait for however long was requested on the command line.
// Note that this requires a vDSO implementation of gettimeofday();
// if it's not available, we could just spin a fixed number of
// iterations instead.
struct timeval start, tv;
gettimeofday(&start, NULL);
while (1) {
gettimeofday(&tv, NULL);
double time = (tv.tv_sec - start.tv_sec) +
(tv.tv_usec - start.tv_usec) / 1000000.0;
if (time >= waittime)
break;
}
return EXIT_SUCCESS;
}
--
Chris Metcalf, Mellanox Technologies
http://www.mellanox.com
On Wed, Jul 27, 2016 at 08:55:28AM -0500, Christoph Lameter wrote:
On Mon, 25 Jul 2016, Christoph Lameter wrote:
quoted
Guess so. I will have a look at this when I get some time again.
Ok so the problem is the clocksource_watchdog() function in
kernel/time/clocksource.c. This function is active if
CONFIG_CLOCKSOURCE_WATCHDOG is defined. It will check the timesources of
each processor for being within bounds and then reschedule itself on the
next one.
The purpose of the function seems to be to determine *if* a clocksource is
unstable. It does not mean that the clocksource *is* unstable.
The critical piece of code is this:
/*
* Cycle through CPUs to check if the CPUs stay synchronized
* to each other.
*/
next_cpu = cpumask_next(raw_smp_processor_id(), cpu_online_mask);
if (next_cpu >= nr_cpu_ids)
next_cpu = cpumask_first(cpu_online_mask);
watchdog_timer.expires += WATCHDOG_INTERVAL;
add_timer_on(&watchdog_timer, next_cpu);
Should we just cycle through the cpus that are not isolated? Otherwise we
need to have some means to check the clocksources for accuracy remotely
(probably impossible for TSC etc).
The WATCHDOG_INTERVAL is 1 second so this causes an interrupt every
second.
Note that we are running with the patch that removes the 1 HZ mininum time
tick. With an older kernel code base (redhat) we can keep the kernel quiet
for minutes. The clocksource watchdog causes timers to fire again.
I had similar issues, this seems to happen when the tsc is considered not reliable
(which doesn't necessarily mean unstable. I think it has to do with some x86 CPU feature
flag).
IIRC, this _has_ to execute on all online CPUs because every TSCs of running CPUs
are concerned.
I personally override that with passing the tsc=reliable kernel parameter. Of course
use it at your own risk.
But eventually I don't think we can offline that to housekeeping only CPUs.
From: Chris Metcalf <hidden> Date: 2016-08-10 23:01:05
On 8/10/2016 6:16 PM, Frederic Weisbecker wrote:
On Wed, Jul 27, 2016 at 08:55:28AM -0500, Christoph Lameter wrote:
quoted
On Mon, 25 Jul 2016, Christoph Lameter wrote:
quoted
Guess so. I will have a look at this when I get some time again.
Ok so the problem is the clocksource_watchdog() function in
kernel/time/clocksource.c. This function is active if
CONFIG_CLOCKSOURCE_WATCHDOG is defined. It will check the timesources of
each processor for being within bounds and then reschedule itself on the
next one.
The purpose of the function seems to be to determine *if* a clocksource is
unstable. It does not mean that the clocksource *is* unstable.
The critical piece of code is this:
/*
* Cycle through CPUs to check if the CPUs stay synchronized
* to each other.
*/
next_cpu = cpumask_next(raw_smp_processor_id(), cpu_online_mask);
if (next_cpu >= nr_cpu_ids)
next_cpu = cpumask_first(cpu_online_mask);
watchdog_timer.expires += WATCHDOG_INTERVAL;
add_timer_on(&watchdog_timer, next_cpu);
Should we just cycle through the cpus that are not isolated? Otherwise we
need to have some means to check the clocksources for accuracy remotely
(probably impossible for TSC etc).
The WATCHDOG_INTERVAL is 1 second so this causes an interrupt every
second.
Note that we are running with the patch that removes the 1 HZ mininum time
tick. With an older kernel code base (redhat) we can keep the kernel quiet
for minutes. The clocksource watchdog causes timers to fire again.
I had similar issues, this seems to happen when the tsc is considered not reliable
(which doesn't necessarily mean unstable. I think it has to do with some x86 CPU feature
flag).
IIRC, this _has_ to execute on all online CPUs because every TSCs of running CPUs
are concerned.
I personally override that with passing the tsc=reliable kernel parameter. Of course
use it at your own risk.
But eventually I don't think we can offline that to housekeeping only CPUs.
Maybe the eventual model here is that as task-isolation cores
re-enter the kernel, they catch a hook that tells them to go
call the unreliable-tsc stuff and see what the state of it is.
This would be the same hook that we could use to defer
kernel TLB flushes, also.
The hard part is that on some platforms it may be fairly
intrusive to get all the hooks in. Arm64 has a nice consistent
set of assembly routines to enter the kernel, which is how they
manage the context_tracking as well, but I fear that x86 may
have a lot more.
--
Chris Metcalf, Mellanox Technologies
http://www.mellanox.com
From: Peter Zijlstra <peterz@infradead.org> Date: 2016-08-11 08:27:45
On Fri, Jul 22, 2016 at 08:50:44AM -0400, Chris Metcalf wrote:
On 7/21/2016 10:20 PM, Christoph Lameter wrote:
quoted
On Thu, 21 Jul 2016, Chris Metcalf wrote:
quoted
On 7/20/2016 10:04 PM, Christoph Lameter wrote:
unstable, and then scheduling work to safely remove that timer.
I haven't looked at this code before (in kernel/time/clocksource.c
under CONFIG_CLOCKSOURCE_WATCHDOG) since the timers on
arm64 and tile aren't unstable. Is it possible to boot your machine
with a stable clocksource?
It already as a stable clocksource. Sorry but that was one of the criteria
for the server when we ordered them. Could this be clock adjustments?
We probably need to get clock folks to jump in on this thread!
Boot with: tsc=reliable, this disables the watchdog.
We (sadly) have to have this thing running on most x86 because TSC, even
if initially stable, can do weird things once its running.
We have seen:
- SMI
- hotplug
- suspend
- multi-socket
mess up the TSC, even if it was deemed 'good' at boot time.
If you _know_ your TSC to be solid, boot with tsc=reliable and be happy.
From: Peter Zijlstra <peterz@infradead.org> Date: 2016-08-11 08:40:17
On Thu, Aug 11, 2016 at 12:16:58AM +0200, Frederic Weisbecker wrote:
I had similar issues, this seems to happen when the tsc is considered not reliable
(which doesn't necessarily mean unstable. I think it has to do with some x86 CPU feature
flag).
Right, as per the other email, in general we cannot know/assume the TSC
to be working as intended :/
IIRC, this _has_ to execute on all online CPUs because every TSCs of running CPUs
are concerned.
With modern Intel we could run it on one CPU per package I think, but at
the same time, too much in NOHZ_FULL assumes the TSC is indeed sane so
it doesn't make sense to me to keep the watchdog running, when it
triggers it would also have to kill all NOHZ_FULL stuff, which would
probably bring the entire machine down..
Arguably we should issue a boot time warning if NOHZ_FULL is configured
and the TSC watchdog is running.
I personally override that with passing the tsc=reliable kernel
parameter. Of course use it at your own risk.
Yes, that is (sadly) our only option. Manually assert our hardware is
solid under the intended workload and then manually disabling the
watchdog.
On Thu, Aug 11, 2016 at 10:40:02AM +0200, Peter Zijlstra wrote:
On Thu, Aug 11, 2016 at 12:16:58AM +0200, Frederic Weisbecker wrote:
quoted
I had similar issues, this seems to happen when the tsc is considered not reliable
(which doesn't necessarily mean unstable. I think it has to do with some x86 CPU feature
flag).
Right, as per the other email, in general we cannot know/assume the TSC
to be working as intended :/
Yeah, I remember you explained me that a little while ago.
quoted
IIRC, this _has_ to execute on all online CPUs because every TSCs of running CPUs
are concerned.
With modern Intel we could run it on one CPU per package I think, but at
the same time, too much in NOHZ_FULL assumes the TSC is indeed sane so
it doesn't make sense to me to keep the watchdog running, when it
triggers it would also have to kill all NOHZ_FULL stuff, which would
probably bring the entire machine down..
Arguably we should issue a boot time warning if NOHZ_FULL is configured
and the TSC watchdog is running.
That's a very good idea! We do that when tsc is unstable but indeed we can't
seriously run NOHZ_FULL on a non-reliable tsc.
I'll take care of that warning.
quoted
I personally override that with passing the tsc=reliable kernel
parameter. Of course use it at your own risk.
Yes, that is (sadly) our only option. Manually assert our hardware is
solid under the intended workload and then manually disabling the
watchdog.
Right, I'll tell about that in the warning.
Thanks for those details!
From: Paul E. McKenney <hidden> Date: 2016-08-11 22:29:17
On Thu, Aug 11, 2016 at 10:40:02AM +0200, Peter Zijlstra wrote:
On Thu, Aug 11, 2016 at 12:16:58AM +0200, Frederic Weisbecker wrote:
quoted
I had similar issues, this seems to happen when the tsc is considered not reliable
(which doesn't necessarily mean unstable. I think it has to do with some x86 CPU feature
flag).
Right, as per the other email, in general we cannot know/assume the TSC
to be working as intended :/
quoted
IIRC, this _has_ to execute on all online CPUs because every TSCs of running CPUs
are concerned.
With modern Intel we could run it on one CPU per package I think, but at
the same time, too much in NOHZ_FULL assumes the TSC is indeed sane so
it doesn't make sense to me to keep the watchdog running, when it
triggers it would also have to kill all NOHZ_FULL stuff, which would
probably bring the entire machine down..
Well, you -could- force a very low priority CPU-bound task to run on
all nohz_full CPUs. Not necessarily a good idea, but a relatively
non-intrusive response to that particular error condition.
Thanx, Paul
Arguably we should issue a boot time warning if NOHZ_FULL is configured
and the TSC watchdog is running.
quoted
I personally override that with passing the tsc=reliable kernel
parameter. Of course use it at your own risk.
Yes, that is (sadly) our only option. Manually assert our hardware is
solid under the intended workload and then manually disabling the
watchdog.
From: Christoph Lameter <hidden> Date: 2016-08-11 23:02:39
On Thu, 11 Aug 2016, Paul E. McKenney wrote:
quoted
With modern Intel we could run it on one CPU per package I think, but at
the same time, too much in NOHZ_FULL assumes the TSC is indeed sane so
it doesn't make sense to me to keep the watchdog running, when it
triggers it would also have to kill all NOHZ_FULL stuff, which would
probably bring the entire machine down..
Well, you -could- force a very low priority CPU-bound task to run on
all nohz_full CPUs. Not necessarily a good idea, but a relatively
non-intrusive response to that particular error condition.
Given that we want the cpu only to run the user task I would think that is
not a good idea.
From: Paul E. McKenney <hidden> Date: 2016-08-11 23:47:39
On Thu, Aug 11, 2016 at 06:02:34PM -0500, Christoph Lameter wrote:
On Thu, 11 Aug 2016, Paul E. McKenney wrote:
quoted
quoted
With modern Intel we could run it on one CPU per package I think, but at
the same time, too much in NOHZ_FULL assumes the TSC is indeed sane so
it doesn't make sense to me to keep the watchdog running, when it
triggers it would also have to kill all NOHZ_FULL stuff, which would
probably bring the entire machine down..
Well, you -could- force a very low priority CPU-bound task to run on
all nohz_full CPUs. Not necessarily a good idea, but a relatively
non-intrusive response to that particular error condition.
Given that we want the cpu only to run the user task I would think that is
not a good idea.
Heh! The only really good idea is for clocks to be reliably in sync.
But if they go out of sync, what do you want to do instead?
Thanx, Paul
From: Paul E. McKenney <hidden> Date: 2016-08-12 16:19:31
On Fri, Aug 12, 2016 at 04:26:13PM +0200, Frederic Weisbecker wrote:
On Fri, Aug 12, 2016 at 09:23:13AM -0500, Christoph Lameter wrote:
quoted
On Thu, 11 Aug 2016, Paul E. McKenney wrote:
quoted
Heh! The only really good idea is for clocks to be reliably in sync.
But if they go out of sync, what do you want to do instead?
For a NOHZ task? Write a message to the syslog and reenable tick.
Fair enough! Kicking off a low-priority task would achieve the latter
but not necessarily the former. And of course assumes that the worker
thread is at real-time priority with various scheduler anti-starvation
features disabled.
Indeed, a strong clocksource is a requirement for a full tickless machine.
On Fri, Aug 12, 2016 at 09:19:19AM -0700, Paul E. McKenney wrote:
On Fri, Aug 12, 2016 at 04:26:13PM +0200, Frederic Weisbecker wrote:
quoted
On Fri, Aug 12, 2016 at 09:23:13AM -0500, Christoph Lameter wrote:
quoted
On Thu, 11 Aug 2016, Paul E. McKenney wrote:
quoted
Heh! The only really good idea is for clocks to be reliably in sync.
But if they go out of sync, what do you want to do instead?
For a NOHZ task? Write a message to the syslog and reenable tick.
Fair enough! Kicking off a low-priority task would achieve the latter
but not necessarily the former. And of course assumes that the worker
thread is at real-time priority with various scheduler anti-starvation
features disabled.
quoted
Indeed, a strong clocksource is a requirement for a full tickless machine.
No disagrement here! ;-)
I have a bot in my mind that randomly posts obvious statements about nohz_full
here and then :-)
From: Chris Metcalf <hidden> Date: 2016-08-15 15:04:28
On 8/11/2016 7:58 AM, Frederic Weisbecker wrote:
quoted
Arguably we should issue a boot time warning if NOHZ_FULL is configured
quoted
and the TSC watchdog is running.
That's a very good idea! We do that when tsc is unstable but indeed we can't
seriously run NOHZ_FULL on a non-reliable tsc.
I'll take care of that warning.
Thanks. So I will drop Christoph's patch to run the TSC watchdog on just
housekeeping cores and we will rely on the "boot time warning" instead.
--
Chris Metcalf, Mellanox Technologies
http://www.mellanox.com