Thread (23 messages) 23 messages, 6 authors, 2017-02-15

RE: [PATCH] genhd: Do not hold event lock when scheduling workqueue elements

From: Dexuan Cui <decui@microsoft.com>
Date: 2017-02-07 02:23:13
Also in: lkml

From: linux-block-owner@vger.kernel.org [mailto:linux-block-
owner@vger.kernel.org] On Behalf Of Dexuan Cui
Sent: Friday, February 3, 2017 20:23
To: Hannes Reinecke <hare@suse.com>; Bart Van Assche
[off-list ref]; hare@suse.de; axboe@kernel.dk
Cc: hch@lst.de; linux-kernel@vger.kernel.org; linux-block@vger.kernel.org=
;
jth@kernel.org
Subject: RE: [PATCH] genhd: Do not hold event lock when scheduling workqu=
eue
elements
=20
quoted
From: linux-kernel-owner@vger.kernel.org [mailto:linux-kernel-
owner@vger.kernel.org] On Behalf Of Hannes Reinecke
Sent: Wednesday, February 1, 2017 00:15
To: Bart Van Assche <redacted>; hare@suse.de;
axboe@kernel.dk
Cc: hch@lst.de; linux-kernel@vger.kernel.org; linux-block@vger.kernel.o=
rg;
quoted
jth@kernel.org
Subject: Re: [PATCH] genhd: Do not hold event lock when scheduling
workqueue
quoted
elements

On 01/31/2017 01:31 AM, Bart Van Assche wrote:
quoted
On Wed, 2017-01-18 at 10:48 +0100, Hannes Reinecke wrote:
quoted
@@ -1488,26 +1487,13 @@ static unsigned long
disk_events_poll_jiffies(struct gendisk *disk)
quoted
quoted
 void disk_block_events(struct gendisk *disk)
 {
        struct disk_events *ev =3D disk->ev;
-       unsigned long flags;
-       bool cancel;

        if (!ev)
                return;

-       /*
-        * Outer mutex ensures that the first blocker completes canc=
eling
quoted
quoted
quoted
-        * the event work before further blockers are allowed to fin=
ish.
quoted
quoted
quoted
-        */
-       mutex_lock(&ev->block_mutex);
-
-       spin_lock_irqsave(&ev->lock, flags);
-       cancel =3D !ev->block++;
-       spin_unlock_irqrestore(&ev->lock, flags);
-
-       if (cancel)
+       if (atomic_inc_return(&ev->block) =3D=3D 1)
                cancel_delayed_work_sync(&disk->ev->dwork);

-       mutex_unlock(&ev->block_mutex);
 }
Hello Hannes,

I have already encountered a few times a deadlock that was caused by =
the
quoted
quoted
event checking code so I agree with you that it would be a big step f=
orward
quoted
quoted
if such deadlocks wouldn't occur anymore. However, this patch realize=
s a
quoted
quoted
change that has not been described in the patch description, namely t=
hat
quoted
quoted
disk_block_events() calls are no longer serialized. Are you sure it i=
s safe
quoted
quoted
to drop the serialization of disk_block_events() calls?
Well, this whole synchronization stuff it a bit weird; I so totally fai=
l
quoted
to see the rationale for it.
But anyway, once we've converted ev->block to atomics I _think_ the
mutex_lock can remain; will be checking.

Cheers,

Hannes
--
=20
Hi, I think I got the same calltrace with today's linux-next (next-201702=
03).
=20
The issue happened every time when my Linux virtual machine booted and
Hannes's patch could NOT help.
=20
The calltrace is pasted below.
=20
-- Dexuan
=20
Any news on this thread?

The issue is still blocking Linux from booting up normally in my test. :-(

Have we identified the faulty patch?
If so, at least I can try to revert it to boot up.

Thanks,
-- Dexuan
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help