From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:55:30
[ Sorry for the duplicate, I tried to cancel sending, but quilt decided
to send anyway :-p ]
Dear RT Folks,
This is the RT stable review cycle of patch 4.9.65-rt57-rc1.
Please scream at me if I messed something up. Please test the patches too.
The -rc release will be uploaded to kernel.org and will be deleted when
the final release is out. This is just a review release (or release candidate).
The pre-releases will not be pushed to the git repository, only the
final release is.
If all goes well, this patch will be converted to the next main release
on 12/5/2017.
Enjoy,
-- Steve
To build 4.9.65-rt57-rc1 directly, the following patches should be applied:
http://www.kernel.org/pub/linux/kernel/v4.x/linux-4.9.tar.xzhttp://www.kernel.org/pub/linux/kernel/v4.x/patch-4.9.65.xzhttp://www.kernel.org/pub/linux/kernel/projects/rt/4.9/patch-4.9.65-rt57-rc1.patch.xz
You can also build from 4.9.65-rt56 by applying the incremental patch:
http://www.kernel.org/pub/linux/kernel/projects/rt/4.9/incr/patch-4.9.65-rt56-rt57-rc1.patch.xz
Changes from 4.9.65-rt56:
---
Alex Shi (1):
cpu_pm: replace raw_notifier to atomic_notifier
Mike Galbraith (2):
rtmutex: Fix lock stealing logic
kernel/hrtimer/hotplug: don't wake ktimersoftd while holding the hrtimer base lock
Sebastian Andrzej Siewior (10):
Revert "fs: jbd2: pull your plug when waiting for space"
posixtimer: init timer only with CONFIG_POSIX_TIMERS enabled
PM / CPU: replace raw_notifier with atomic_notifier (fixup)
kernel/hrtimer: migrate deferred timer on CPU down
net: take the tcp_sk_lock lock with BH disabled
kernel/hrtimer: don't wakeup a process while holding the hrtimer base lock
Bluetooth: avoid recursive locking in hci_send_to_channel()
iommu/amd: Use raw_cpu_ptr() instead of get_cpu_ptr() for ->flush_queue
rt/locking: allow recursive local_trylock()
locking/rtmutex: don't drop the wait_lock twice
Steven Rostedt (VMware) (2):
Revert "memcontrol: Prevent scheduling while atomic in cgroup code"
Linux 4.9.65-rt57-rc1
----
drivers/iommu/amd_iommu.c | 4 +--
fs/jbd2/checkpoint.c | 2 --
include/linux/init_task.h | 2 +-
include/linux/locallock.h | 9 ++++++
kernel/cpu_pm.c | 50 +++++++++-----------------------
kernel/locking/rtmutex.c | 74 +++++++++++++++++++++++------------------------
kernel/time/hrtimer.c | 35 ++++++++++++++++------
localversion-rt | 2 +-
mm/memcontrol.c | 13 ++++-----
net/bluetooth/hci_sock.c | 17 +++++++----
net/ipv4/tcp_ipv4.c | 8 ++---
11 files changed, 108 insertions(+), 108 deletions(-)
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:50:33
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
Lockdep may complain about an unsafe locking scenario:
| CPU0 CPU1
| ---- ----
| lock((tcp_sk_lock).lock);
| lock(&per_cpu(local_softirq_locks[i], __cpu).lock);
| lock((tcp_sk_lock).lock);
| lock(&per_cpu(local_softirq_locks[i], __cpu).lock);
in the call paths:
do_current_softirqs -> tcp_v4_send_ack()
vs
tcp_v4_send_reset -> do_current_softirqs().
This should not happen since local_softirq_locks is per CPU. Reversing
the order makes lockdep happy.
Reported-by: Jacek Konieczny <redacted>
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
net/ipv4/tcp_ipv4.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:50:34
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
hrtimers, which were deferred to the softirq context, and expire between
softirq shutdown and hrtimer migration are dangling around. If the CPU
goes back up the list head will be initialized and this corrupts the
timer's list. It will remain unnoticed until a hrtimer_cancel().
This moves those timers so they will expire.
Cc: stable-rt@vger.kernel.org
Reported-by: Mike Galbraith <redacted>
Tested-by: Mike Galbraith <redacted>
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
kernel/time/hrtimer.c | 5 +++++
1 file changed, 5 insertions(+)
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:50:37
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Mike Galbraith <redacted>
kernel/hrtimer: don't wakeup a process while holding the hrtimer base lock
missed a path, namely hrtimers_dead_cpu() -> migrate_hrtimer_list(). Defer
raising softirq until after base lock has been released there as well.
Signed-off-by: Mike Galbraith <redacted>
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
kernel/time/hrtimer.c | 19 +++++++++++++------
1 file changed, 13 insertions(+), 6 deletions(-)
@@ -1837,7 +1837,7 @@ int hrtimers_prepare_cpu(unsigned int cpu)#ifdef CONFIG_HOTPLUG_CPU-staticvoidmigrate_hrtimer_list(structhrtimer_clock_base*old_base,+staticintmigrate_hrtimer_list(structhrtimer_clock_base*old_base,structhrtimer_clock_base*new_base){structhrtimer*timer;
@@ -1891,13 +1895,16 @@ int hrtimers_dead_cpu(unsigned int scpu)raw_spin_lock_nested(&old_base->lock,SINGLE_DEPTH_NESTING);for(i=0;i<HRTIMER_MAX_CLOCK_BASES;i++){-migrate_hrtimer_list(&old_base->clock_base[i],-&new_base->clock_base[i]);+raise|=migrate_hrtimer_list(&old_base->clock_base[i],+&new_base->clock_base[i]);}raw_spin_unlock(&old_base->lock);raw_spin_unlock(&new_base->lock);+if(raise)+raise_softirq_irqoff(HRTIMER_SOFTIRQ);+/* Check, if we got expired work to do */__hrtimer_peek_ahead_timers();local_irq_enable();
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:50:38
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
We must not wake any process (and thus acquire the pi->lock) while
holding the hrtimer's base lock. This does not happen usually because
the hrtimer-callback is invoked in IRQ-context and so
raise_softirq_irqoff() does not wakeup a process.
However during CPU-hotplug it might get called from hrtimers_dead_cpu()
which would wakeup the thread immediately.
Reported-by: Mike Galbraith <redacted>
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
kernel/time/hrtimer.c | 15 ++++++++++-----
1 file changed, 10 insertions(+), 5 deletions(-)
@@ -1523,7 +1523,7 @@ void hrtimer_interrupt(struct clock_event_device *dev)*/cpu_base->expires_next.tv64=KTIME_MAX;-__hrtimer_run_queues(cpu_base,now);+raise=__hrtimer_run_queues(cpu_base,now);/* Reevaluate the clock bases for the next expiry */expires_next=__hrtimer_get_next_event(cpu_base);
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:50:40
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Alex Shi <redacted>
This patch replace a rwlock and raw notifier by atomic notifier which
protected by spin_lock and rcu.
The first to reason to have this replace is due to a 'scheduling while
atomic' bug of RT kernel on arm/arm64 platform. On arm/arm64, rwlock
cpu_pm_notifier_lock in cpu_pm cause a potential schedule after irq
disable in idle call chain:
cpu_startup_entry
cpu_idle_loop
local_irq_disable()
cpuidle_idle_call
call_cpuidle
cpuidle_enter
cpuidle_enter_state
->enter :arm_enter_idle_state
cpu_pm_enter/exit
CPU_PM_CPU_IDLE_ENTER
read_lock(&cpu_pm_notifier_lock); <-- sleep in idle
__rt_spin_lock();
schedule();
The kernel panic is here:
[ 4.609601] BUG: scheduling while atomic: swapper/1/0/0x00000002
[ 4.609608] [<ffff0000086fae70>] arm_enter_idle_state+0x18/0x70
[ 4.609614] Modules linked in:
[ 4.609615] [<ffff0000086f9298>] cpuidle_enter_state+0xf0/0x218
[ 4.609620] [<ffff0000086f93f8>] cpuidle_enter+0x18/0x20
[ 4.609626] Preemption disabled at:
[ 4.609627] [<ffff0000080fa234>] call_cpuidle+0x24/0x40
[ 4.609635] [<ffff000008882fa4>] schedule_preempt_disabled+0x1c/0x28
[ 4.609639] [<ffff0000080fa49c>] cpu_startup_entry+0x154/0x1f8
[ 4.609645] [<ffff00000808e004>] secondary_start_kernel+0x15c/0x1a0
Daniel Lezcano said this notification is needed on arm/arm64 platforms.
Sebastian suggested using atomic_notifier instead of rwlock, which is not
only removing the sleeping in idle, but also getting better latency
improvement.
This patch passed Fengguang's 0day testing.
Signed-off-by: Alex Shi <redacted>
Cc: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
Cc: Thomas Gleixner <redacted>
Cc: Anders Roxell <redacted>
Cc: Rik van Riel <redacted>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Rafael J. Wysocki <redacted>
Cc: Daniel Lezcano <redacted>
Cc: linux-rt-users <redacted>
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
kernel/cpu_pm.c | 43 ++++++-------------------------------------
1 file changed, 6 insertions(+), 37 deletions(-)
@@ -47,14 +46,7 @@ static int cpu_pm_notify(enum cpu_pm_event event, int nr_to_call, int *nr_calls)*/intcpu_pm_register_notifier(structnotifier_block*nb){-unsignedlongflags;-intret;--write_lock_irqsave(&cpu_pm_notifier_lock,flags);-ret=raw_notifier_chain_register(&cpu_pm_notifier_chain,nb);-write_unlock_irqrestore(&cpu_pm_notifier_lock,flags);--returnret;+returnatomic_notifier_chain_register(&cpu_pm_notifier_chain,nb);}EXPORT_SYMBOL_GPL(cpu_pm_register_notifier);
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:50:42
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
Mart reported a deadlock in -RT in the call path:
hci_send_monitor_ctrl_event() -> hci_send_to_channel()
because both functions acquire the same read lock hci_sk_list.lock. This
is also a mainline issue because the qrwlock implementation is writer
fair (the traditional rwlock implementation is reader biased).
To avoid the deadlock there is now __hci_send_to_channel() which expects
the readlock to be held.
Cc: Marcel Holtmann <marcel@holtmann.org>
Cc: Johan Hedberg <redacted>
Cc: rt-stable@vger.kernel.org
Fixes: 38ceaa00d02d ("Bluetooth: Add support for sending MGMT commands and events to monitor")
Reported-by: Mart van de Wege <redacted>
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
net/bluetooth/hci_sock.c | 17 +++++++++++------
1 file changed, 11 insertions(+), 6 deletions(-)
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:50:44
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Sebastian Andrzej Siewior <redacted>
get_cpu_ptr() disabled preemption and returns the ->flush_queue object
of the current CPU. raw_cpu_ptr() does the same except that it not
disable preemption which means the scheduler can move it to another CPU
after it obtained the per-CPU object.
In this case this is not bad because the data structure itself is
protected with a spin_lock. This change shouldn't matter however on RT
it does because the sleeping lock can't be accessed with disabled
preemption.
Cc: rt-stable-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
Cc: Joerg Roedel <redacted>
Cc: iommu-cunTk1MwBs9QetFLy7KEm3xJsTq8ys+cHZ5vskTnxNA@public.gmane.org
Reported-by: Vinod Adhikary <redacted>
Signed-off-by: Sebastian Andrzej Siewior <redacted>
---
drivers/iommu/amd_iommu.c | 4 +---
1 file changed, 1 insertion(+), 3 deletions(-)
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:50:46
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
required for following networking patch which does recursive try-lock.
While at it, add the !RT version of it because it did not yet exist.
Cc: rt-stable@vger.kernel.org
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
include/linux/locallock.h | 9 +++++++++
1 file changed, 9 insertions(+)
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:50:47
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
The original patch changed betwen its posting and what finally went into
Rafael's tree so here is the delta.
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
kernel/cpu_pm.c | 7 +++++++
1 file changed, 7 insertions(+)
@@ -28,8 +28,15 @@ static int cpu_pm_notify(enum cpu_pm_event event, int nr_to_call, int *nr_calls){intret;+/*+*__atomic_notifier_call_chainhasaRCUreadcriticalsection,which+*couldbedisfunctionalincpuidle.CopyRCU_NONIDLEcodetolet+*RCUknowthis.+*/+rcu_irq_enter_irqson();ret=__atomic_notifier_call_chain(&cpu_pm_notifier_chain,event,NULL,nr_to_call,nr_calls);+rcu_irq_exit_irqson();returnnotifier_to_errno(ret);}
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:50:50
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
In v4.11 it is possible to disable the posix timers and so we must not
attempt to initialize the task_struct on RT with !POSIX_TIMERS.
This patch does so.
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
include/linux/init_task.h | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:50:51
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Mike Galbraith <redacted>
1. When trying to acquire an rtmutex, we first try to grab it without
queueing the waiter, and explicitly check for that initial attempt
in the !waiter path of __try_to_take_rt_mutex(). Checking whether
the lock taker is top waiter before allowing a steal attempt in that
path is a thinko: the lock taker has not yet blocked.
2. It seems wrong to change the definition of rt_mutex_waiter_less()
to mean less or perhaps equal when we have an rt_mutex_waiter_equal().
Remove the thinko, restore rt_mutex_waiter_less(), implement and use
rt_mutex_steal() based upon rt_mutex_waiter_less/equal(), moving all
qualification criteria into the function itself.
Reviewed-by: Steven Rostedt (VMware) <rostedt@goodmis.org>
Signed-off-by: Mike Galbraith <redacted>
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
kernel/locking/rtmutex.c | 73 ++++++++++++++++++++++++------------------------
1 file changed, 36 insertions(+), 37 deletions(-)
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:50:53
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
Since the futex rework, __rt_mutex_start_proxy_lock() does no longer
acquire the wait_lock so it must not drop it. Otherwise the lock is not
only unlocked twice but also the preemption counter is underflown.
It is okay to remove that line because this function does not disable
interrupts nor does it acquire the ->wait_lock. The caller does this so it is
wrong do it here (after the futex rework).
Cc: rt-stable@vger.kernel.org #v4.9.18-rt14+
Reported-by: Gusenleitner Klaus <redacted>
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
kernel/locking/rtmutex.c | 1 -
1 file changed, 1 deletion(-)
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:55:27
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
This reverts commit "fs: jbd2: pull your plug when waiting for space".
This was a duct-tape fix which shouldn't be needed since commit
"locking/rt-mutex: fix deadlock in device mapper / block-IO".
Cc: stable@vger.kernel.org
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
Signed-off-by: Steven Rostedt (VMware) <rostedt@goodmis.org>
---
fs/jbd2/checkpoint.c | 2 --
1 file changed, 2 deletions(-)
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 15:55:29
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: "Steven Rostedt (VMware)" <rostedt@goodmis.org>
The commit "memcontrol: Prevent scheduling while atomic in cgroup code"
fixed this issue:
refill_stock()
get_cpu_var()
drain_stock()
res_counter_uncharge()
res_counter_uncharge_until()
spin_lock() <== boom
But commit 3e32cb2e0a12b ("mm: memcontrol: lockless page counters") replaced
the calls to res_counter_uncharge() in drain_stock() to the lockless
function page_counter_uncharge(). There is no more spin lock there and no
more reason to have that local lock.
Cc: <redacted>
Reported-by: Haiyang HY1 Tan <redacted>
Signed-off-by: Steven Rostedt (VMware) <rostedt@goodmis.org>
[bigeasy: That upstream commit appeared in v3.19 and the patch in
question in v3.18.7-rt2 and v3.18 seems still to be maintained. So I
guess that v3.18 would need the locallocks that we are about to remove
here. I am not sure if any earlier versions have the patch
backported.
The stable tag here is because Haiyang reported (and debugged) a crash
in 4.4-RT with this patch applied (which has get_cpu_light() instead
the locallocks it gained in v4.9-RT).
https://lkml.kernel.org/r/05AA4EC5C6EC1D48BE2CDCFF3AE0B8A637F78A15@CNMAILEX04.lenovo.com
]
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
mm/memcontrol.c | 13 ++++++-------
1 file changed, 6 insertions(+), 7 deletions(-)
From: Steven Rostedt <rostedt@goodmis.org> Date: 2017-12-01 17:35:42
On Fri, 01 Dec 2017 10:50:03 -0500
Steven Rostedt [off-list ref] wrote:
4.9.65-rt57-rc1 stable review patch.
If anyone has any objections, please let me know.
------------------
From: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
In v4.11 it is possible to disable the posix timers and so we must not
attempt to initialize the task_struct on RT with !POSIX_TIMERS.
This patch does so.
Hmm, I may have been too greedy in pulling in this patch. I'm going to
remove it from the list, and post a -rc2.
-- Steve
quoted hunk
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
include/linux/init_task.h | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)