Re: [RFC PATCH v4 4/4] md: raid456 add nowait support

39 messages, 5 authors, 2022-01-02 · open the first message on its own page

Re: [RFC PATCH v4 4/4] md: raid456 add nowait support

From: Song Liu <song@kernel.org>
Date: 2021-12-14 01:11:36

On Mon, Dec 13, 2021 at 4:53 PM Vishal Verma [off-list ref] wrote:
[...]
What kernel base are you using for your patches?

These were based out of for-5.16-tag (037c50bfb)
Please rebase on top of md-next branch from here:

https://git.kernel.org/pub/scm/linux/kernel/git/song/md.git

Thanks,
Song

Re: [RFC PATCH v4 4/4] md: raid456 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-14 01:13:01

On 12/13/21 6:11 PM, Song Liu wrote:
On Mon, Dec 13, 2021 at 4:53 PM Vishal Verma [off-list ref] wrote:
[...]
quoted
What kernel base are you using for your patches?

These were based out of for-5.16-tag (037c50bfb)
Please rebase on top of md-next branch from here:

https://git.kernel.org/pub/scm/linux/kernel/git/song/md.git

Thanks,
Song
Ack, will do!

Re: [RFC PATCH v4 4/4] md: raid456 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-14 15:30:36

On 12/13/21 6:12 PM, Vishal Verma wrote:
On 12/13/21 6:11 PM, Song Liu wrote:
quoted
On Mon, Dec 13, 2021 at 4:53 PM Vishal Verma 
[off-list ref] wrote:
[...]
quoted
What kernel base are you using for your patches?

These were based out of for-5.16-tag (037c50bfb)
Please rebase on top of md-next branch from here:

https://git.kernel.org/pub/scm/linux/kernel/git/song/md.git

Thanks,
Song
Ack, will do!
After rebasing to md-next branch and re-running the tests 100% W, 100% 
R, 70%R30%W with both io_uring and libaio I don't see any issue. Thank you!

Re: [RFC PATCH v4 4/4] md: raid456 add nowait support

From: Song Liu <song@kernel.org>
Date: 2021-12-14 17:08:56

On Tue, Dec 14, 2021 at 7:30 AM Vishal Verma [off-list ref] wrote:

On 12/13/21 6:12 PM, Vishal Verma wrote:
quoted
On 12/13/21 6:11 PM, Song Liu wrote:
quoted
On Mon, Dec 13, 2021 at 4:53 PM Vishal Verma
[off-list ref] wrote:
[...]
quoted
What kernel base are you using for your patches?

These were based out of for-5.16-tag (037c50bfb)
Please rebase on top of md-next branch from here:

https://git.kernel.org/pub/scm/linux/kernel/git/song/md.git

Thanks,
Song
Ack, will do!
After rebasing to md-next branch and re-running the tests 100% W, 100%
R, 70%R30%W with both io_uring and libaio I don't see any issue. Thank you!
That's great! Please address all the feedback and submit v5.

Thanks,
Song

Re: [RFC PATCH v4 4/4] md: raid456 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-14 18:09:34

On 12/14/21 10:08 AM, Song Liu wrote:
On Tue, Dec 14, 2021 at 7:30 AM Vishal Verma [off-list ref] wrote:
quoted
On 12/13/21 6:12 PM, Vishal Verma wrote:
quoted
On 12/13/21 6:11 PM, Song Liu wrote:
quoted
On Mon, Dec 13, 2021 at 4:53 PM Vishal Verma
[off-list ref] wrote:
[...]
quoted
What kernel base are you using for your patches?

These were based out of for-5.16-tag (037c50bfb)
Please rebase on top of md-next branch from here:

https://git.kernel.org/pub/scm/linux/kernel/git/song/md.git

Thanks,
Song
Ack, will do!
After rebasing to md-next branch and re-running the tests 100% W, 100%
R, 70%R30%W with both io_uring and libaio I don't see any issue. Thank you!
That's great! Please address all the feedback and submit v5.

Thanks,
Song
Yup, will do later today or tomorrow. I need to test raid10 with similar 
cases and not 100% sure about discard case.

[PATCH v5 1/4] md: add support for REQ_NOWAIT

From: Vishal Verma <hidden>
Date: 2021-12-15 06:09:25

commit 021a24460dc2 ("block: add QUEUE_FLAG_NOWAIT") added support
for checking whether a given bdev supports handling of REQ_NOWAIT or not.
Since then commit 6abc49468eea ("dm: add support for REQ_NOWAIT and enable
it for linear target") added support for REQ_NOWAIT for dm. This uses
a similar approach to incorporate REQ_NOWAIT for md based bios.

This patch was tested using t/io_uring tool within FIO. A nvme drive
was partitioned into 2 partitions and a simple raid 0 configuration
/dev/md0 was created.

md0 : active raid0 nvme4n1p1[1] nvme4n1p2[0]
  937423872 blocks super 1.2 512k chunks

Before patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

38396   38396 pts/2    00:00:00 io_uring
38396   38397 pts/2    00:00:15 io_uring
38396   38398 pts/2    00:00:13 iou-wrk-38397

We can see iou-wrk-38397 io worker thread created which gets created
when io_uring sees that the underlying device (/dev/md0 in this case)
doesn't support nowait.

After patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

38341   38341 pts/2    00:10:22 io_uring
38341   38342 pts/2    00:10:37 io_uring

After running this patch, we don't see any io worker thread
being created which indicated that io_uring saw that the
underlying device does support nowait. This is the exact behaviour
noticed on a dm device which also supports nowait.

For all the other raid personalities except raid0, we would need
to train pieces which involves make_request fn in order for them
to correctly handle REQ_NOWAIT.

Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/md.c | 21 +++++++++++++++++++++
 1 file changed, 21 insertions(+)
diff --git a/drivers/md/md.c b/drivers/md/md.c
index 7fbf6f0ac01b..5b4c28e0e1db 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -419,6 +419,12 @@ void md_handle_request(struct mddev *mddev, struct bio *bio)
 	if (is_suspended(mddev, bio)) {
 		DEFINE_WAIT(__wait);
 		for (;;) {
+			/* Bail out if REQ_NOWAIT is set for the bio */
+			if (bio->bi_opf & REQ_NOWAIT) {
+				rcu_read_unlock();
+				bio_wouldblock_error(bio);
+				return;
+			}
 			prepare_to_wait(&mddev->sb_wait, &__wait,
 					TASK_UNINTERRUPTIBLE);
 			if (!is_suspended(mddev, bio))
@@ -5787,6 +5793,7 @@ int md_run(struct mddev *mddev)
 	int err;
 	struct md_rdev *rdev;
 	struct md_personality *pers;
+	bool nowait = true;
 
 	if (list_empty(&mddev->disks))
 		/* cannot run an array with no devices.. */
@@ -5857,8 +5864,13 @@ int md_run(struct mddev *mddev)
 			}
 		}
 		sysfs_notify_dirent_safe(rdev->sysfs_state);
+		nowait = nowait && blk_queue_nowait(bdev_get_queue(rdev->bdev));
 	}
 
+	/* Set the NOWAIT flags if all underlying devices support it */
+	if (nowait)
+		blk_queue_flag_set(QUEUE_FLAG_NOWAIT, mddev->queue);
+
 	if (!bioset_initialized(&mddev->bio_set)) {
 		err = bioset_init(&mddev->bio_set, BIO_POOL_SIZE, 0, BIOSET_NEED_BVECS);
 		if (err)
@@ -7002,6 +7014,15 @@ static int hot_add_disk(struct mddev *mddev, dev_t dev)
 	set_bit(MD_SB_CHANGE_DEVS, &mddev->sb_flags);
 	if (!mddev->thread)
 		md_update_sb(mddev, 1);
+	/*
+	 * If the new disk does not support REQ_NOWAIT,
+	 * disable on the whole MD.
+	 */
+	if (!blk_queue_nowait(bdev_get_queue(rdev->bdev))) {
+		pr_info("%s: Disabling nowait because %s does not support nowait\n",
+			mdname(mddev), bdevname(rdev->bdev, b));
+		blk_queue_flag_clear(QUEUE_FLAG_NOWAIT, mddev->queue);
+	}
 	/*
 	 * Kick recovery, maybe this spare has to be added to the
 	 * array immediately.
-- 
2.17.1

[PATCH v5 2/4] md: raid1 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-15 06:09:29

This adds nowait support to the RAID1 driver. It makes RAID1 driver
return with EAGAIN for situations where it could wait for eg:

- Waiting for the barrier,
- Array got frozen,
- Too many pending I/Os to be queued.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/raid1.c | 74 +++++++++++++++++++++++++++++++++++-----------
 1 file changed, 57 insertions(+), 17 deletions(-)
diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c
index 7dc8026cf6ee..727d31de5694 100644
--- a/drivers/md/raid1.c
+++ b/drivers/md/raid1.c
@@ -929,8 +929,9 @@ static void lower_barrier(struct r1conf *conf, sector_t sector_nr)
 	wake_up(&conf->wait_barrier);
 }
 
-static void _wait_barrier(struct r1conf *conf, int idx)
+static bool _wait_barrier(struct r1conf *conf, int idx, bool nowait)
 {
+	bool ret = true;
 	/*
 	 * We need to increase conf->nr_pending[idx] very early here,
 	 * then raise_barrier() can be blocked when it waits for
@@ -961,7 +962,7 @@ static void _wait_barrier(struct r1conf *conf, int idx)
 	 */
 	if (!READ_ONCE(conf->array_frozen) &&
 	    !atomic_read(&conf->barrier[idx]))
-		return;
+		return ret;
 
 	/*
 	 * After holding conf->resync_lock, conf->nr_pending[idx]
@@ -979,18 +980,27 @@ static void _wait_barrier(struct r1conf *conf, int idx)
 	 */
 	wake_up(&conf->wait_barrier);
 	/* Wait for the barrier in same barrier unit bucket to drop. */
-	wait_event_lock_irq(conf->wait_barrier,
-			    !conf->array_frozen &&
-			     !atomic_read(&conf->barrier[idx]),
-			    conf->resync_lock);
+	if (conf->array_frozen || atomic_read(&conf->barrier[idx])) {
+		/* Return false when nowait flag is set */
+		if (nowait)
+			ret = false;
+		else {
+			wait_event_lock_irq(conf->wait_barrier,
+					!conf->array_frozen &&
+					!atomic_read(&conf->barrier[idx]),
+					conf->resync_lock);
+		}
+	}
 	atomic_inc(&conf->nr_pending[idx]);
 	atomic_dec(&conf->nr_waiting[idx]);
 	spin_unlock_irq(&conf->resync_lock);
+	return ret;
 }
 
-static void wait_read_barrier(struct r1conf *conf, sector_t sector_nr)
+static bool wait_read_barrier(struct r1conf *conf, sector_t sector_nr, bool nowait)
 {
 	int idx = sector_to_idx(sector_nr);
+	bool ret = true;
 
 	/*
 	 * Very similar to _wait_barrier(). The difference is, for read
@@ -1002,7 +1012,7 @@ static void wait_read_barrier(struct r1conf *conf, sector_t sector_nr)
 	atomic_inc(&conf->nr_pending[idx]);
 
 	if (!READ_ONCE(conf->array_frozen))
-		return;
+		return ret;
 
 	spin_lock_irq(&conf->resync_lock);
 	atomic_inc(&conf->nr_waiting[idx]);
@@ -1013,19 +1023,27 @@ static void wait_read_barrier(struct r1conf *conf, sector_t sector_nr)
 	 */
 	wake_up(&conf->wait_barrier);
 	/* Wait for array to be unfrozen */
-	wait_event_lock_irq(conf->wait_barrier,
-			    !conf->array_frozen,
-			    conf->resync_lock);
+	if (conf->array_frozen || atomic_read(&conf->barrier[idx])) {
+		if (nowait)
+			/* Return false when nowait flag is set */
+			ret = false;
+		else {
+			wait_event_lock_irq(conf->wait_barrier,
+					!conf->array_frozen,
+					conf->resync_lock);
+		}
+	}
 	atomic_inc(&conf->nr_pending[idx]);
 	atomic_dec(&conf->nr_waiting[idx]);
 	spin_unlock_irq(&conf->resync_lock);
+	return ret;
 }
 
-static void wait_barrier(struct r1conf *conf, sector_t sector_nr)
+static bool wait_barrier(struct r1conf *conf, sector_t sector_nr, bool nowait)
 {
 	int idx = sector_to_idx(sector_nr);
 
-	_wait_barrier(conf, idx);
+	return _wait_barrier(conf, idx, nowait);
 }
 
 static void _allow_barrier(struct r1conf *conf, int idx)
@@ -1236,7 +1254,11 @@ static void raid1_read_request(struct mddev *mddev, struct bio *bio,
 	 * Still need barrier for READ in case that whole
 	 * array is frozen.
 	 */
-	wait_read_barrier(conf, bio->bi_iter.bi_sector);
+	if (!wait_read_barrier(conf, bio->bi_iter.bi_sector,
+				bio->bi_opf & REQ_NOWAIT)) {
+		bio_wouldblock_error(bio);
+		return;
+	}
 
 	if (!r1_bio)
 		r1_bio = alloc_r1bio(mddev, bio);
@@ -1336,6 +1358,10 @@ static void raid1_write_request(struct mddev *mddev, struct bio *bio,
 		     bio->bi_iter.bi_sector, bio_end_sector(bio))) {
 
 		DEFINE_WAIT(w);
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return;
+		}
 		for (;;) {
 			prepare_to_wait(&conf->wait_barrier,
 					&w, TASK_IDLE);
@@ -1353,17 +1379,26 @@ static void raid1_write_request(struct mddev *mddev, struct bio *bio,
 	 * thread has put up a bar for new requests.
 	 * Continue immediately if no resync is active currently.
 	 */
-	wait_barrier(conf, bio->bi_iter.bi_sector);
+	if (!wait_barrier(conf, bio->bi_iter.bi_sector,
+				bio->bi_opf & REQ_NOWAIT)) {
+		bio_wouldblock_error(bio);
+		return;
+	}
 
 	r1_bio = alloc_r1bio(mddev, bio);
 	r1_bio->sectors = max_write_sectors;
 
 	if (conf->pending_count >= max_queued_requests) {
 		md_wakeup_thread(mddev->thread);
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return;
+		}
 		raid1_log(mddev, "wait queued");
 		wait_event(conf->wait_barrier,
 			   conf->pending_count < max_queued_requests);
 	}
+
 	/* first select target devices under rcu_lock and
 	 * inc refcount on their rdev.  Record them by setting
 	 * bios[x] to bio
@@ -1458,9 +1493,14 @@ static void raid1_write_request(struct mddev *mddev, struct bio *bio,
 				rdev_dec_pending(conf->mirrors[j].rdev, mddev);
 		r1_bio->state = 0;
 		allow_barrier(conf, bio->bi_iter.bi_sector);
+
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return;
+		}
 		raid1_log(mddev, "wait rdev %d blocked", blocked_rdev->raid_disk);
 		md_wait_for_blocked_rdev(blocked_rdev, mddev);
-		wait_barrier(conf, bio->bi_iter.bi_sector);
+		wait_barrier(conf, bio->bi_iter.bi_sector, false);
 		goto retry_write;
 	}
 
@@ -1687,7 +1727,7 @@ static void close_sync(struct r1conf *conf)
 	int idx;
 
 	for (idx = 0; idx < BARRIER_BUCKETS_NR; idx++) {
-		_wait_barrier(conf, idx);
+		_wait_barrier(conf, idx, false);
 		_allow_barrier(conf, idx);
 	}
 
-- 
2.17.1

[PATCH v5 3/4] md: raid10 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-15 06:09:31

This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

- Waiting for the barrier,
- Too many pending I/Os to be queued,
- Reshape operation,
- Discard operation.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/raid10.c | 57 +++++++++++++++++++++++++++++++++++----------
 1 file changed, 45 insertions(+), 12 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..f6c73987e9ac 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,11 +952,18 @@ static void lower_barrier(struct r10conf *conf)
 	wake_up(&conf->wait_barrier);
 }
 
-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
 {
 	spin_lock_irq(&conf->resync_lock);
 	if (conf->barrier) {
 		struct bio_list *bio_list = current->bio_list;
+
+		/* Return false when nowait flag is set */
+		if (nowait) {
+			spin_unlock_irq(&conf->resync_lock);
+			return false;
+		}
+
 		conf->nr_waiting++;
 		/* Wait for the barrier to drop.
 		 * However if there are already pending
@@ -988,6 +995,7 @@ static void wait_barrier(struct r10conf *conf)
 	}
 	atomic_inc(&conf->nr_pending);
 	spin_unlock_irq(&conf->resync_lock);
+	return true;
 }
 
 static void allow_barrier(struct r10conf *conf)
@@ -1101,17 +1109,25 @@ static void raid10_unplug(struct blk_plug_cb *cb, bool from_schedule)
 static void regular_request_wait(struct mddev *mddev, struct r10conf *conf,
 				 struct bio *bio, sector_t sectors)
 {
-	wait_barrier(conf);
+	/* Bail out if REQ_NOWAIT is set for the bio */
+	if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+		bio_wouldblock_error(bio);
+		return;
+	}
 	while (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery) &&
 	    bio->bi_iter.bi_sector < conf->reshape_progress &&
 	    bio->bi_iter.bi_sector + sectors > conf->reshape_progress) {
 		raid10_log(conf->mddev, "wait reshape");
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return;
+		}
 		allow_barrier(conf);
 		wait_event(conf->wait_barrier,
 			   conf->reshape_progress <= bio->bi_iter.bi_sector ||
 			   conf->reshape_progress >= bio->bi_iter.bi_sector +
 			   sectors);
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 	}
 }
 
@@ -1179,7 +1195,7 @@ static void raid10_read_request(struct mddev *mddev, struct bio *bio,
 		bio_chain(split, bio);
 		allow_barrier(conf);
 		submit_bio_noacct(bio);
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 		bio = split;
 		r10_bio->master_bio = bio;
 		r10_bio->sectors = max_sectors;
@@ -1338,7 +1354,7 @@ static void wait_blocked_dev(struct mddev *mddev, struct r10bio *r10_bio)
 		raid10_log(conf->mddev, "%s wait rdev %d blocked",
 				__func__, blocked_rdev->raid_disk);
 		md_wait_for_blocked_rdev(blocked_rdev, mddev);
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 		goto retry_wait;
 	}
 }
@@ -1357,6 +1373,11 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
 					    bio_end_sector(bio)))) {
 		DEFINE_WAIT(w);
 		for (;;) {
+			/* Bail out if REQ_NOWAIT is set for the bio */
+			if (bio->bi_opf & REQ_NOWAIT) {
+				bio_wouldblock_error(bio);
+				return;
+			}
 			prepare_to_wait(&conf->wait_barrier,
 					&w, TASK_IDLE);
 			if (!md_cluster_ops->area_resyncing(mddev, WRITE,
@@ -1381,6 +1402,10 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
 			      BIT(MD_SB_CHANGE_DEVS) | BIT(MD_SB_CHANGE_PENDING));
 		md_wakeup_thread(mddev->thread);
 		raid10_log(conf->mddev, "wait reshape metadata");
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return;
+		}
 		wait_event(mddev->sb_wait,
 			   !test_bit(MD_SB_CHANGE_PENDING, &mddev->sb_flags));
 
@@ -1390,6 +1415,10 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
 	if (conf->pending_count >= max_queued_requests) {
 		md_wakeup_thread(mddev->thread);
 		raid10_log(mddev, "wait queued");
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return;
+		}
 		wait_event(conf->wait_barrier,
 			   conf->pending_count < max_queued_requests);
 	}
@@ -1482,7 +1511,7 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
 		bio_chain(split, bio);
 		allow_barrier(conf);
 		submit_bio_noacct(bio);
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 		bio = split;
 		r10_bio->master_bio = bio;
 	}
@@ -1607,7 +1636,11 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
 	if (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery))
 		return -EAGAIN;
 
-	wait_barrier(conf);
+	if (bio->bi_opf & REQ_NOWAIT) {
+		bio_wouldblock_error(bio);
+		return 0;
+	}
+	wait_barrier(conf, false);
 
 	/*
 	 * Check reshape again to avoid reshape happens after checking
@@ -1649,7 +1682,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
 		allow_barrier(conf);
 		/* Resend the fist split part */
 		submit_bio_noacct(split);
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 	}
 	div_u64_rem(bio_end, stripe_size, &remainder);
 	if (remainder) {
@@ -1660,7 +1693,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
 		/* Resend the second split part */
 		submit_bio_noacct(bio);
 		bio = split;
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 	}
 
 	bio_start = bio->bi_iter.bi_sector;
@@ -1816,7 +1849,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
 		end_disk_offset += geo->stride;
 		atomic_inc(&first_r10bio->remaining);
 		raid_end_discard_bio(r10_bio);
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 		goto retry_discard;
 	}
 
@@ -2011,7 +2044,7 @@ static void print_conf(struct r10conf *conf)
 
 static void close_sync(struct r10conf *conf)
 {
-	wait_barrier(conf);
+	wait_barrier(conf, false);
 	allow_barrier(conf);
 
 	mempool_exit(&conf->r10buf_pool);
@@ -4819,7 +4852,7 @@ static sector_t reshape_request(struct mddev *mddev, sector_t sector_nr,
 	if (need_flush ||
 	    time_after(jiffies, conf->reshape_checkpoint + 10*HZ)) {
 		/* Need to update reshape_position in metadata */
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 		mddev->reshape_position = conf->reshape_progress;
 		if (mddev->reshape_backwards)
 			mddev->curr_resync_completed = raid10_size(mddev, 0, 0)
-- 
2.17.1

[PATCH v5 4/4] md: raid456 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-15 06:09:35

Returns EAGAIN in case the raid456 driver would block
waiting for situations like:

- Reshape operation,
- Discard operation.

Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/raid5.c | 20 ++++++++++++++++++++
 1 file changed, 20 insertions(+)
diff --git a/drivers/md/raid5.c b/drivers/md/raid5.c
index 1240a5c16af8..b505e4cec777 100644
--- a/drivers/md/raid5.c
+++ b/drivers/md/raid5.c
@@ -5715,6 +5715,11 @@ static void make_discard_request(struct mddev *mddev, struct bio *bi)
 		set_bit(R5_Overlap, &sh->dev[sh->pd_idx].flags);
 		if (test_bit(STRIPE_SYNCING, &sh->state)) {
 			raid5_release_stripe(sh);
+			/* Bail out if REQ_NOWAIT is set */
+			if (bi->bi_opf & REQ_NOWAIT) {
+				bio_wouldblock_error(bi);
+				return;
+			}
 			schedule();
 			goto again;
 		}
@@ -5727,6 +5732,11 @@ static void make_discard_request(struct mddev *mddev, struct bio *bi)
 				set_bit(R5_Overlap, &sh->dev[d].flags);
 				spin_unlock_irq(&sh->stripe_lock);
 				raid5_release_stripe(sh);
+				/* Bail out if REQ_NOWAIT is set */
+				if (bi->bi_opf & REQ_NOWAIT) {
+					bio_wouldblock_error(bi);
+					return;
+				}
 				schedule();
 				goto again;
 			}
@@ -5820,6 +5830,16 @@ static bool raid5_make_request(struct mddev *mddev, struct bio * bi)
 	bi->bi_next = NULL;
 
 	md_account_bio(mddev, &bi);
+	/* Bail out if REQ_NOWAIT is set */
+	if ((bi->bi_opf & REQ_NOWAIT) &&
+	    (conf->reshape_progress != MaxSector) &&
+	    (mddev->reshape_backwards
+	    ? (logical_sector > conf->reshape_progress && logical_sector <= conf->reshape_safe)
+	    : (logical_sector >= conf->reshape_safe && logical_sector < conf->reshape_progress))) {
+		bio_wouldblock_error(bi);
+		return true;
+	}
+
 	prepare_to_wait(&conf->wait_for_overlap, &w, TASK_UNINTERRUPTIBLE);
 	for (; logical_sector < last_sector; logical_sector += RAID5_STRIPE_SECTORS(conf)) {
 		int previous;
-- 
2.17.1

Re: [PATCH v5 1/4] md: add support for REQ_NOWAIT

From: Song Liu <song@kernel.org>
Date: 2021-12-15 20:02:24

On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma [off-list ref] wrote:
[...]
quoted hunk
Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/md.c | 21 +++++++++++++++++++++
 1 file changed, 21 insertions(+)
diff --git a/drivers/md/md.c b/drivers/md/md.c
index 7fbf6f0ac01b..5b4c28e0e1db 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -419,6 +419,12 @@ void md_handle_request(struct mddev *mddev, struct bio *bio)
        if (is_suspended(mddev, bio)) {
                DEFINE_WAIT(__wait);
                for (;;) {
+                       /* Bail out if REQ_NOWAIT is set for the bio */
+                       if (bio->bi_opf & REQ_NOWAIT) {
+                               rcu_read_unlock();
+                               bio_wouldblock_error(bio);
+                               return;
+                       }
I moved this part to before the for (;;) loop. And applied to md-next.

Thanks,
Song

Re: [PATCH v5 2/4] md: raid1 add nowait support

From: Song Liu <song@kernel.org>
Date: 2021-12-15 20:33:51

On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma [off-list ref] wrote:
This adds nowait support to the RAID1 driver. It makes RAID1 driver
return with EAGAIN for situations where it could wait for eg:

- Waiting for the barrier,
- Array got frozen,
- Too many pending I/Os to be queued.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.
Please see some detailed comments below. But a general and more important
question: were you able to trigger these conditions (path that lead to
bio_wouldblock_error) in the tests?

Ideally, we should test all these conditions. If something is really
hard to trigger,
please highlight that in the commit log, so that I can run more tests on them.

Thanks,
Song
quoted hunk
Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/raid1.c | 74 +++++++++++++++++++++++++++++++++++-----------
 1 file changed, 57 insertions(+), 17 deletions(-)
diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c
index 7dc8026cf6ee..727d31de5694 100644
--- a/drivers/md/raid1.c
+++ b/drivers/md/raid1.c
@@ -929,8 +929,9 @@ static void lower_barrier(struct r1conf *conf, sector_t sector_nr)
        wake_up(&conf->wait_barrier);
 }

-static void _wait_barrier(struct r1conf *conf, int idx)
+static bool _wait_barrier(struct r1conf *conf, int idx, bool nowait)
 {
+       bool ret = true;
        /*
         * We need to increase conf->nr_pending[idx] very early here,
         * then raise_barrier() can be blocked when it waits for
@@ -961,7 +962,7 @@ static void _wait_barrier(struct r1conf *conf, int idx)
         */
        if (!READ_ONCE(conf->array_frozen) &&
            !atomic_read(&conf->barrier[idx]))
-               return;
+               return ret;

        /*
         * After holding conf->resync_lock, conf->nr_pending[idx]
@@ -979,18 +980,27 @@ static void _wait_barrier(struct r1conf *conf, int idx)
         */
        wake_up(&conf->wait_barrier);
        /* Wait for the barrier in same barrier unit bucket to drop. */
-       wait_event_lock_irq(conf->wait_barrier,
-                           !conf->array_frozen &&
-                            !atomic_read(&conf->barrier[idx]),
-                           conf->resync_lock);
+       if (conf->array_frozen || atomic_read(&conf->barrier[idx])) {
Do we really need this check?
+               /* Return false when nowait flag is set */
+               if (nowait)
+                       ret = false;
+               else {
+                       wait_event_lock_irq(conf->wait_barrier,
+                                       !conf->array_frozen &&
+                                       !atomic_read(&conf->barrier[idx]),
+                                       conf->resync_lock);
+               }
+       }
        atomic_inc(&conf->nr_pending[idx]);
Were you able to trigger the condition in the tests? I think we should
only increase
nr_pending for ret == true. Otherwise, we will leak a nr_pending.
quoted hunk
        atomic_dec(&conf->nr_waiting[idx]);
        spin_unlock_irq(&conf->resync_lock);
+       return ret;
 }

-static void wait_read_barrier(struct r1conf *conf, sector_t sector_nr)
+static bool wait_read_barrier(struct r1conf *conf, sector_t sector_nr, bool nowait)
 {
        int idx = sector_to_idx(sector_nr);
+       bool ret = true;

        /*
         * Very similar to _wait_barrier(). The difference is, for read
@@ -1002,7 +1012,7 @@ static void wait_read_barrier(struct r1conf *conf, sector_t sector_nr)
        atomic_inc(&conf->nr_pending[idx]);

        if (!READ_ONCE(conf->array_frozen))
-               return;
+               return ret;

        spin_lock_irq(&conf->resync_lock);
        atomic_inc(&conf->nr_waiting[idx]);
@@ -1013,19 +1023,27 @@ static void wait_read_barrier(struct r1conf *conf, sector_t sector_nr)
         */
        wake_up(&conf->wait_barrier);
        /* Wait for array to be unfrozen */
-       wait_event_lock_irq(conf->wait_barrier,
-                           !conf->array_frozen,
-                           conf->resync_lock);
+       if (conf->array_frozen || atomic_read(&conf->barrier[idx])) {
I guess we don't need this either. Also, the condition there is not identical
to wait_barrier (no need to check conf->barrier[idx]).
+               if (nowait)
+                       /* Return false when nowait flag is set */
+                       ret = false;
+               else {
+                       wait_event_lock_irq(conf->wait_barrier,
+                                       !conf->array_frozen,
+                                       conf->resync_lock);
+               }
+       }
        atomic_inc(&conf->nr_pending[idx]);
ditto on nr_pending.
quoted hunk
        atomic_dec(&conf->nr_waiting[idx]);
        spin_unlock_irq(&conf->resync_lock);
+       return ret;
 }

-static void wait_barrier(struct r1conf *conf, sector_t sector_nr)
+static bool wait_barrier(struct r1conf *conf, sector_t sector_nr, bool nowait)
 {
        int idx = sector_to_idx(sector_nr);

-       _wait_barrier(conf, idx);
+       return _wait_barrier(conf, idx, nowait);
 }

 static void _allow_barrier(struct r1conf *conf, int idx)
@@ -1236,7 +1254,11 @@ static void raid1_read_request(struct mddev *mddev, struct bio *bio,
         * Still need barrier for READ in case that whole
         * array is frozen.
         */
-       wait_read_barrier(conf, bio->bi_iter.bi_sector);
+       if (!wait_read_barrier(conf, bio->bi_iter.bi_sector,
+                               bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }

        if (!r1_bio)
                r1_bio = alloc_r1bio(mddev, bio);
@@ -1336,6 +1358,10 @@ static void raid1_write_request(struct mddev *mddev, struct bio *bio,
                     bio->bi_iter.bi_sector, bio_end_sector(bio))) {

                DEFINE_WAIT(w);
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                for (;;) {
                        prepare_to_wait(&conf->wait_barrier,
                                        &w, TASK_IDLE);
@@ -1353,17 +1379,26 @@ static void raid1_write_request(struct mddev *mddev, struct bio *bio,
         * thread has put up a bar for new requests.
         * Continue immediately if no resync is active currently.
         */
-       wait_barrier(conf, bio->bi_iter.bi_sector);
+       if (!wait_barrier(conf, bio->bi_iter.bi_sector,
+                               bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }

        r1_bio = alloc_r1bio(mddev, bio);
        r1_bio->sectors = max_write_sectors;

        if (conf->pending_count >= max_queued_requests) {
                md_wakeup_thread(mddev->thread);
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
I think we need to fix conf->nr_pending before returning.
quoted hunk
+                       return;
+               }
                raid1_log(mddev, "wait queued");
                wait_event(conf->wait_barrier,
                           conf->pending_count < max_queued_requests);
        }
+
        /* first select target devices under rcu_lock and
         * inc refcount on their rdev.  Record them by setting
         * bios[x] to bio
@@ -1458,9 +1493,14 @@ static void raid1_write_request(struct mddev *mddev, struct bio *bio,
                                rdev_dec_pending(conf->mirrors[j].rdev, mddev);
                r1_bio->state = 0;
                allow_barrier(conf, bio->bi_iter.bi_sector);
+
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                raid1_log(mddev, "wait rdev %d blocked", blocked_rdev->raid_disk);
                md_wait_for_blocked_rdev(blocked_rdev, mddev);
-               wait_barrier(conf, bio->bi_iter.bi_sector);
+               wait_barrier(conf, bio->bi_iter.bi_sector, false);
                goto retry_write;
        }
@@ -1687,7 +1727,7 @@ static void close_sync(struct r1conf *conf)
        int idx;

        for (idx = 0; idx < BARRIER_BUCKETS_NR; idx++) {
-               _wait_barrier(conf, idx);
+               _wait_barrier(conf, idx, false);
                _allow_barrier(conf, idx);
        }

--
2.17.1

Re: [PATCH v5 3/4] md: raid10 add nowait support

From: Song Liu <song@kernel.org>
Date: 2021-12-15 20:42:31

On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma [off-list ref] wrote:
quoted hunk
This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

- Waiting for the barrier,
- Too many pending I/Os to be queued,
- Reshape operation,
- Discard operation.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/raid10.c | 57 +++++++++++++++++++++++++++++++++++----------
 1 file changed, 45 insertions(+), 12 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..f6c73987e9ac 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,11 +952,18 @@ static void lower_barrier(struct r10conf *conf)
        wake_up(&conf->wait_barrier);
 }

-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
 {
        spin_lock_irq(&conf->resync_lock);
        if (conf->barrier) {
                struct bio_list *bio_list = current->bio_list;
+
+               /* Return false when nowait flag is set */
+               if (nowait) {
+                       spin_unlock_irq(&conf->resync_lock);
+                       return false;
+               }
+
                conf->nr_waiting++;
                /* Wait for the barrier to drop.
                 * However if there are already pending
@@ -988,6 +995,7 @@ static void wait_barrier(struct r10conf *conf)
        }
        atomic_inc(&conf->nr_pending);
        spin_unlock_irq(&conf->resync_lock);
+       return true;
 }

 static void allow_barrier(struct r10conf *conf)
@@ -1101,17 +1109,25 @@ static void raid10_unplug(struct blk_plug_cb *cb, bool from_schedule)
 static void regular_request_wait(struct mddev *mddev, struct r10conf *conf,
                                 struct bio *bio, sector_t sectors)
 {
-       wait_barrier(conf);
+       /* Bail out if REQ_NOWAIT is set for the bio */
+       if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }
I think we also need regular_request_wait to return bool and handle it properly.

Thanks,
Song
quoted hunk
        while (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery) &&
            bio->bi_iter.bi_sector < conf->reshape_progress &&
            bio->bi_iter.bi_sector + sectors > conf->reshape_progress) {
                raid10_log(conf->mddev, "wait reshape");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                allow_barrier(conf);
                wait_event(conf->wait_barrier,
                           conf->reshape_progress <= bio->bi_iter.bi_sector ||
                           conf->reshape_progress >= bio->bi_iter.bi_sector +
                           sectors);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
        }
 }
@@ -1179,7 +1195,7 @@ static void raid10_read_request(struct mddev *mddev, struct bio *bio,
                bio_chain(split, bio);
                allow_barrier(conf);
                submit_bio_noacct(bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                bio = split;
                r10_bio->master_bio = bio;
                r10_bio->sectors = max_sectors;
@@ -1338,7 +1354,7 @@ static void wait_blocked_dev(struct mddev *mddev, struct r10bio *r10_bio)
                raid10_log(conf->mddev, "%s wait rdev %d blocked",
                                __func__, blocked_rdev->raid_disk);
                md_wait_for_blocked_rdev(blocked_rdev, mddev);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                goto retry_wait;
        }
 }
@@ -1357,6 +1373,11 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
                                            bio_end_sector(bio)))) {
                DEFINE_WAIT(w);
                for (;;) {
+                       /* Bail out if REQ_NOWAIT is set for the bio */
+                       if (bio->bi_opf & REQ_NOWAIT) {
+                               bio_wouldblock_error(bio);
+                               return;
+                       }
                        prepare_to_wait(&conf->wait_barrier,
                                        &w, TASK_IDLE);
                        if (!md_cluster_ops->area_resyncing(mddev, WRITE,
@@ -1381,6 +1402,10 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
                              BIT(MD_SB_CHANGE_DEVS) | BIT(MD_SB_CHANGE_PENDING));
                md_wakeup_thread(mddev->thread);
                raid10_log(conf->mddev, "wait reshape metadata");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                wait_event(mddev->sb_wait,
                           !test_bit(MD_SB_CHANGE_PENDING, &mddev->sb_flags));
@@ -1390,6 +1415,10 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
        if (conf->pending_count >= max_queued_requests) {
                md_wakeup_thread(mddev->thread);
                raid10_log(mddev, "wait queued");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                wait_event(conf->wait_barrier,
                           conf->pending_count < max_queued_requests);
        }
@@ -1482,7 +1511,7 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
                bio_chain(split, bio);
                allow_barrier(conf);
                submit_bio_noacct(bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                bio = split;
                r10_bio->master_bio = bio;
        }
@@ -1607,7 +1636,11 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
        if (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery))
                return -EAGAIN;

-       wait_barrier(conf);
+       if (bio->bi_opf & REQ_NOWAIT) {
+               bio_wouldblock_error(bio);
+               return 0;
+       }
+       wait_barrier(conf, false);

        /*
         * Check reshape again to avoid reshape happens after checking
@@ -1649,7 +1682,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
                allow_barrier(conf);
                /* Resend the fist split part */
                submit_bio_noacct(split);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
        }
        div_u64_rem(bio_end, stripe_size, &remainder);
        if (remainder) {
@@ -1660,7 +1693,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
                /* Resend the second split part */
                submit_bio_noacct(bio);
                bio = split;
-               wait_barrier(conf);
+               wait_barrier(conf, false);
        }

        bio_start = bio->bi_iter.bi_sector;
@@ -1816,7 +1849,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
                end_disk_offset += geo->stride;
                atomic_inc(&first_r10bio->remaining);
                raid_end_discard_bio(r10_bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                goto retry_discard;
        }
@@ -2011,7 +2044,7 @@ static void print_conf(struct r10conf *conf)

 static void close_sync(struct r10conf *conf)
 {
-       wait_barrier(conf);
+       wait_barrier(conf, false);
        allow_barrier(conf);

        mempool_exit(&conf->r10buf_pool);
@@ -4819,7 +4852,7 @@ static sector_t reshape_request(struct mddev *mddev, sector_t sector_nr,
        if (need_flush ||
            time_after(jiffies, conf->reshape_checkpoint + 10*HZ)) {
                /* Need to update reshape_position in metadata */
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                mddev->reshape_position = conf->reshape_progress;
                if (mddev->reshape_backwards)
                        mddev->curr_resync_completed = raid10_size(mddev, 0, 0)
--
2.17.1

Re: [PATCH v5 2/4] md: raid1 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-15 22:20:11

On 12/15/21 1:33 PM, Song Liu wrote:
On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma [off-list ref] wrote:
quoted
This adds nowait support to the RAID1 driver. It makes RAID1 driver
return with EAGAIN for situations where it could wait for eg:

- Waiting for the barrier,
- Array got frozen,
- Too many pending I/Os to be queued.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.
Please see some detailed comments below. But a general and more important
question: were you able to trigger these conditions (path that lead to
bio_wouldblock_error) in the tests?

Ideally, we should test all these conditions. If something is really
hard to trigger,
please highlight that in the commit log, so that I can run more tests on them.

Thanks,
Song
quoted
Signed-off-by: Vishal Verma <redacted>
---
  drivers/md/raid1.c | 74 +++++++++++++++++++++++++++++++++++-----------
  1 file changed, 57 insertions(+), 17 deletions(-)
diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c
index 7dc8026cf6ee..727d31de5694 100644
--- a/drivers/md/raid1.c
+++ b/drivers/md/raid1.c
@@ -929,8 +929,9 @@ static void lower_barrier(struct r1conf *conf, sector_t sector_nr)
         wake_up(&conf->wait_barrier);
  }

-static void _wait_barrier(struct r1conf *conf, int idx)
+static bool _wait_barrier(struct r1conf *conf, int idx, bool nowait)
  {
+       bool ret = true;
         /*
          * We need to increase conf->nr_pending[idx] very early here,
          * then raise_barrier() can be blocked when it waits for
@@ -961,7 +962,7 @@ static void _wait_barrier(struct r1conf *conf, int idx)
          */
         if (!READ_ONCE(conf->array_frozen) &&
             !atomic_read(&conf->barrier[idx]))
-               return;
+               return ret;

         /*
          * After holding conf->resync_lock, conf->nr_pending[idx]
@@ -979,18 +980,27 @@ static void _wait_barrier(struct r1conf *conf, int idx)
          */
         wake_up(&conf->wait_barrier);
         /* Wait for the barrier in same barrier unit bucket to drop. */
-       wait_event_lock_irq(conf->wait_barrier,
-                           !conf->array_frozen &&
-                            !atomic_read(&conf->barrier[idx]),
-                           conf->resync_lock);
+       if (conf->array_frozen || atomic_read(&conf->barrier[idx])) {
Do we really need this check?
This was done when looking at the wait_event_lock_irq conditions.
I am not very sure about this.
quoted
+               /* Return false when nowait flag is set */
+               if (nowait)
+                       ret = false;
+               else {
+                       wait_event_lock_irq(conf->wait_barrier,
+                                       !conf->array_frozen &&
+                                       !atomic_read(&conf->barrier[idx]),
+                                       conf->resync_lock);
+               }
+       }
         atomic_inc(&conf->nr_pending[idx]);
Were you able to trigger the condition in the tests? I think we should
only increase
nr_pending for ret == true. Otherwise, we will leak a nr_pending.
No I wasn't able to. Makes sense about nr_pending. Thanks for catching.
quoted
         atomic_dec(&conf->nr_waiting[idx]);
         spin_unlock_irq(&conf->resync_lock);
+       return ret;
  }

-static void wait_read_barrier(struct r1conf *conf, sector_t sector_nr)
+static bool wait_read_barrier(struct r1conf *conf, sector_t sector_nr, bool nowait)
  {
         int idx = sector_to_idx(sector_nr);
+       bool ret = true;

         /*
          * Very similar to _wait_barrier(). The difference is, for read
@@ -1002,7 +1012,7 @@ static void wait_read_barrier(struct r1conf *conf, sector_t sector_nr)
         atomic_inc(&conf->nr_pending[idx]);

         if (!READ_ONCE(conf->array_frozen))
-               return;
+               return ret;

         spin_lock_irq(&conf->resync_lock);
         atomic_inc(&conf->nr_waiting[idx]);
@@ -1013,19 +1023,27 @@ static void wait_read_barrier(struct r1conf *conf, sector_t sector_nr)
          */
         wake_up(&conf->wait_barrier);
         /* Wait for array to be unfrozen */
-       wait_event_lock_irq(conf->wait_barrier,
-                           !conf->array_frozen,
-                           conf->resync_lock);
+       if (conf->array_frozen || atomic_read(&conf->barrier[idx])) {
I guess we don't need this either. Also, the condition there is not identical
to wait_barrier (no need to check conf->barrier[idx]).
OK
quoted
+               if (nowait)
+                       /* Return false when nowait flag is set */
+                       ret = false;
+               else {
+                       wait_event_lock_irq(conf->wait_barrier,
+                                       !conf->array_frozen,
+                                       conf->resync_lock);
+               }
+       }
         atomic_inc(&conf->nr_pending[idx]);
ditto on nr_pending.
OK
quoted
         atomic_dec(&conf->nr_waiting[idx]);
         spin_unlock_irq(&conf->resync_lock);
+       return ret;
  }

-static void wait_barrier(struct r1conf *conf, sector_t sector_nr)
+static bool wait_barrier(struct r1conf *conf, sector_t sector_nr, bool nowait)
  {
         int idx = sector_to_idx(sector_nr);

-       _wait_barrier(conf, idx);
+       return _wait_barrier(conf, idx, nowait);
  }

  static void _allow_barrier(struct r1conf *conf, int idx)
@@ -1236,7 +1254,11 @@ static void raid1_read_request(struct mddev *mddev, struct bio *bio,
          * Still need barrier for READ in case that whole
          * array is frozen.
          */
-       wait_read_barrier(conf, bio->bi_iter.bi_sector);
+       if (!wait_read_barrier(conf, bio->bi_iter.bi_sector,
+                               bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }

         if (!r1_bio)
                 r1_bio = alloc_r1bio(mddev, bio);
@@ -1336,6 +1358,10 @@ static void raid1_write_request(struct mddev *mddev, struct bio *bio,
                      bio->bi_iter.bi_sector, bio_end_sector(bio))) {

                 DEFINE_WAIT(w);
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                 for (;;) {
                         prepare_to_wait(&conf->wait_barrier,
                                         &w, TASK_IDLE);
@@ -1353,17 +1379,26 @@ static void raid1_write_request(struct mddev *mddev, struct bio *bio,
          * thread has put up a bar for new requests.
          * Continue immediately if no resync is active currently.
          */
-       wait_barrier(conf, bio->bi_iter.bi_sector);
+       if (!wait_barrier(conf, bio->bi_iter.bi_sector,
+                               bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }

         r1_bio = alloc_r1bio(mddev, bio);
         r1_bio->sectors = max_write_sectors;

         if (conf->pending_count >= max_queued_requests) {
                 md_wakeup_thread(mddev->thread);
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
I think we need to fix conf->nr_pending before returning.
OK, this one I am not sure. You mean dec conf->nr_pending?
quoted
+                       return;
+               }
                 raid1_log(mddev, "wait queued");
                 wait_event(conf->wait_barrier,
                            conf->pending_count < max_queued_requests);
         }
+
         /* first select target devices under rcu_lock and
          * inc refcount on their rdev.  Record them by setting
          * bios[x] to bio
@@ -1458,9 +1493,14 @@ static void raid1_write_request(struct mddev *mddev, struct bio *bio,
                                 rdev_dec_pending(conf->mirrors[j].rdev, mddev);
                 r1_bio->state = 0;
                 allow_barrier(conf, bio->bi_iter.bi_sector);
+
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                 raid1_log(mddev, "wait rdev %d blocked", blocked_rdev->raid_disk);
                 md_wait_for_blocked_rdev(blocked_rdev, mddev);
-               wait_barrier(conf, bio->bi_iter.bi_sector);
+               wait_barrier(conf, bio->bi_iter.bi_sector, false);
                 goto retry_write;
         }
@@ -1687,7 +1727,7 @@ static void close_sync(struct r1conf *conf)
         int idx;

         for (idx = 0; idx < BARRIER_BUCKETS_NR; idx++) {
-               _wait_barrier(conf, idx);
+               _wait_barrier(conf, idx, false);
                 _allow_barrier(conf, idx);
         }

--
2.17.1

Re: [PATCH v5 3/4] md: raid10 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-15 22:20:58

On 12/15/21 1:42 PM, Song Liu wrote:
On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma [off-list ref] wrote:
quoted
This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

- Waiting for the barrier,
- Too many pending I/Os to be queued,
- Reshape operation,
- Discard operation.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
  drivers/md/raid10.c | 57 +++++++++++++++++++++++++++++++++++----------
  1 file changed, 45 insertions(+), 12 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..f6c73987e9ac 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,11 +952,18 @@ static void lower_barrier(struct r10conf *conf)
         wake_up(&conf->wait_barrier);
  }

-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
  {
         spin_lock_irq(&conf->resync_lock);
         if (conf->barrier) {
                 struct bio_list *bio_list = current->bio_list;
+
+               /* Return false when nowait flag is set */
+               if (nowait) {
+                       spin_unlock_irq(&conf->resync_lock);
+                       return false;
+               }
+
                 conf->nr_waiting++;
                 /* Wait for the barrier to drop.
                  * However if there are already pending
@@ -988,6 +995,7 @@ static void wait_barrier(struct r10conf *conf)
         }
         atomic_inc(&conf->nr_pending);
         spin_unlock_irq(&conf->resync_lock);
+       return true;
  }

  static void allow_barrier(struct r10conf *conf)
@@ -1101,17 +1109,25 @@ static void raid10_unplug(struct blk_plug_cb *cb, bool from_schedule)
  static void regular_request_wait(struct mddev *mddev, struct r10conf *conf,
                                  struct bio *bio, sector_t sectors)
  {
-       wait_barrier(conf);
+       /* Bail out if REQ_NOWAIT is set for the bio */
+       if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }
I think we also need regular_request_wait to return bool and handle it properly.

Thanks,
Song
Ack, will fix it. Thanks!
quoted
         while (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery) &&
             bio->bi_iter.bi_sector < conf->reshape_progress &&
             bio->bi_iter.bi_sector + sectors > conf->reshape_progress) {
                 raid10_log(conf->mddev, "wait reshape");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                 allow_barrier(conf);
                 wait_event(conf->wait_barrier,
                            conf->reshape_progress <= bio->bi_iter.bi_sector ||
                            conf->reshape_progress >= bio->bi_iter.bi_sector +
                            sectors);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
         }
  }
@@ -1179,7 +1195,7 @@ static void raid10_read_request(struct mddev *mddev, struct bio *bio,
                 bio_chain(split, bio);
                 allow_barrier(conf);
                 submit_bio_noacct(bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 bio = split;
                 r10_bio->master_bio = bio;
                 r10_bio->sectors = max_sectors;
@@ -1338,7 +1354,7 @@ static void wait_blocked_dev(struct mddev *mddev, struct r10bio *r10_bio)
                 raid10_log(conf->mddev, "%s wait rdev %d blocked",
                                 __func__, blocked_rdev->raid_disk);
                 md_wait_for_blocked_rdev(blocked_rdev, mddev);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 goto retry_wait;
         }
  }
@@ -1357,6 +1373,11 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
                                             bio_end_sector(bio)))) {
                 DEFINE_WAIT(w);
                 for (;;) {
+                       /* Bail out if REQ_NOWAIT is set for the bio */
+                       if (bio->bi_opf & REQ_NOWAIT) {
+                               bio_wouldblock_error(bio);
+                               return;
+                       }
                         prepare_to_wait(&conf->wait_barrier,
                                         &w, TASK_IDLE);
                         if (!md_cluster_ops->area_resyncing(mddev, WRITE,
@@ -1381,6 +1402,10 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
                               BIT(MD_SB_CHANGE_DEVS) | BIT(MD_SB_CHANGE_PENDING));
                 md_wakeup_thread(mddev->thread);
                 raid10_log(conf->mddev, "wait reshape metadata");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                 wait_event(mddev->sb_wait,
                            !test_bit(MD_SB_CHANGE_PENDING, &mddev->sb_flags));
@@ -1390,6 +1415,10 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
         if (conf->pending_count >= max_queued_requests) {
                 md_wakeup_thread(mddev->thread);
                 raid10_log(mddev, "wait queued");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                 wait_event(conf->wait_barrier,
                            conf->pending_count < max_queued_requests);
         }
@@ -1482,7 +1511,7 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
                 bio_chain(split, bio);
                 allow_barrier(conf);
                 submit_bio_noacct(bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 bio = split;
                 r10_bio->master_bio = bio;
         }
@@ -1607,7 +1636,11 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
         if (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery))
                 return -EAGAIN;

-       wait_barrier(conf);
+       if (bio->bi_opf & REQ_NOWAIT) {
+               bio_wouldblock_error(bio);
+               return 0;
+       }
+       wait_barrier(conf, false);

         /*
          * Check reshape again to avoid reshape happens after checking
@@ -1649,7 +1682,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
                 allow_barrier(conf);
                 /* Resend the fist split part */
                 submit_bio_noacct(split);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
         }
         div_u64_rem(bio_end, stripe_size, &remainder);
         if (remainder) {
@@ -1660,7 +1693,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
                 /* Resend the second split part */
                 submit_bio_noacct(bio);
                 bio = split;
-               wait_barrier(conf);
+               wait_barrier(conf, false);
         }

         bio_start = bio->bi_iter.bi_sector;
@@ -1816,7 +1849,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
                 end_disk_offset += geo->stride;
                 atomic_inc(&first_r10bio->remaining);
                 raid_end_discard_bio(r10_bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 goto retry_discard;
         }
@@ -2011,7 +2044,7 @@ static void print_conf(struct r10conf *conf)

  static void close_sync(struct r10conf *conf)
  {
-       wait_barrier(conf);
+       wait_barrier(conf, false);
         allow_barrier(conf);

         mempool_exit(&conf->r10buf_pool);
@@ -4819,7 +4852,7 @@ static sector_t reshape_request(struct mddev *mddev, sector_t sector_nr,
         if (need_flush ||
             time_after(jiffies, conf->reshape_checkpoint + 10*HZ)) {
                 /* Need to update reshape_position in metadata */
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 mddev->reshape_position = conf->reshape_progress;
                 if (mddev->reshape_backwards)
                         mddev->curr_resync_completed = raid10_size(mddev, 0, 0)
--
2.17.1

Re: [PATCH v5 3/4] md: raid10 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-16 00:30:32

On 12/15/21 3:20 PM, Vishal Verma wrote:
On 12/15/21 1:42 PM, Song Liu wrote:
quoted
On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma 
[off-list ref] wrote:
quoted
This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

- Waiting for the barrier,
- Too many pending I/Os to be queued,
- Reshape operation,
- Discard operation.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
  drivers/md/raid10.c | 57 
+++++++++++++++++++++++++++++++++++----------
  1 file changed, 45 insertions(+), 12 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..f6c73987e9ac 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,11 +952,18 @@ static void lower_barrier(struct r10conf *conf)
         wake_up(&conf->wait_barrier);
  }

-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
  {
         spin_lock_irq(&conf->resync_lock);
         if (conf->barrier) {
                 struct bio_list *bio_list = current->bio_list;
+
+               /* Return false when nowait flag is set */
+               if (nowait) {
+ spin_unlock_irq(&conf->resync_lock);
+                       return false;
+               }
+
                 conf->nr_waiting++;
                 /* Wait for the barrier to drop.
                  * However if there are already pending
@@ -988,6 +995,7 @@ static void wait_barrier(struct r10conf *conf)
         }
         atomic_inc(&conf->nr_pending);
         spin_unlock_irq(&conf->resync_lock);
+       return true;
  }

  static void allow_barrier(struct r10conf *conf)
@@ -1101,17 +1109,25 @@ static void raid10_unplug(struct blk_plug_cb 
*cb, bool from_schedule)
  static void regular_request_wait(struct mddev *mddev, struct 
r10conf *conf,
                                  struct bio *bio, sector_t sectors)
  {
-       wait_barrier(conf);
+       /* Bail out if REQ_NOWAIT is set for the bio */
+       if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }
I think we also need regular_request_wait to return bool and handle 
it properly.

Thanks,
Song
Ack, will fix it. Thanks!
Ran into this while running with io_uring. With the current v5 (raid10 
patch) on top of md-next branch.
./t/io_uring -a 0 -d 256 </dev/raid10>

It didn't trigger with aio (-a 1)

[  248.128661] BUG: kernel NULL pointer dereference, address: 
00000000000000b8
[  248.135628] #PF: supervisor read access in kernel mode
[  248.140762] #PF: error_code(0x0000) - not-present page
[  248.145903] PGD 0 P4D 0
[  248.148443] Oops: 0000 [#1] PREEMPT SMP NOPTI
[  248.152800] CPU: 49 PID: 9461 Comm: io_uring Kdump: loaded Not 
tainted 5.16.0-rc3+ #2
[  248.160629] Hardware name: Dell Inc. PowerEdge R650xs/0PPTY2, BIOS 
1.3.8 08/31/2021
[  248.168279] RIP: 0010:raid10_end_read_request+0x74/0x140 [raid10]
[  248.174373] Code: 48 60 48 8b 58 58 48 c1 e2 05 49 03 55 08 48 89 4a 
10 40 84 f6 75 48 f0 41 80 4c 24 18 01 4c 89 e7 e8 e0 b8 ff ff 49 8b 4d 
00 <48> 8b 83 b8 00 00 00 f0 ff 8b f0 00 00 00 0f 94 c2 a8 01 74 04 84
[  248.193120] RSP: 0018:ffffb1c38d598ce8 EFLAGS: 00010086
[  248.198344] RAX: ffff8e5da2a1a100 RBX: 0000000000000000 RCX: 
ffff8e5d89747000
[  248.205479] RDX: 000000008040003a RSI: 0000000080400039 RDI: 
ffff8e1e00044900
[  248.212611] RBP: ffffb1c38d598d30 R08: 0000000000000000 R09: 
0000000000000001
[  248.219744] R10: ffff8e5da2a1ae00 R11: 000000411bab9000 R12: 
ffff8e5da2a1ae00
[  248.226877] R13: ffff8e5d8973fc00 R14: 0000000000000000 R15: 
0000000000001000
[  248.234009] FS:  00007fc26b07d700(0000) GS:ffff8e9c6e600000(0000) 
knlGS:0000000000000000
[  248.242096] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  248.247843] CR2: 00000000000000b8 CR3: 00000040b25d4005 CR4: 
0000000000770ee0
[  248.254973] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 
0000000000000000
[  248.262107] DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7: 
0000000000000400
[  248.269240] PKRU: 55555554
[  248.271953] Call Trace:
[  248.274406]  <IRQ>
[  248.276425]  bio_endio+0xf6/0x170
[  248.279743]  blk_update_request+0x12d/0x470
[  248.283931]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.288809]  blk_mq_end_request_batch+0x76/0x490
[  248.293429]  ? dma_direct_unmap_sg+0xdd/0x1a0
[  248.297786]  ? smp_call_function_single_async+0x46/0x70
[  248.303015]  ? mempool_kfree+0xe/0x10
[  248.306680]  ? mempool_kfree+0xe/0x10
[  248.310345]  nvme_pci_complete_batch+0x26/0xb0
[  248.314792]  nvme_irq+0x298/0x2f0
[  248.318110]  ? nvme_unmap_data+0xf0/0xf0
[  248.322038]  __handle_irq_event_percpu+0x3f/0x190
[  248.326744]  handle_irq_event_percpu+0x33/0x80
[  248.331190]  handle_irq_event+0x39/0x60
[  248.335028]  handle_edge_irq+0xbe/0x1e0
[  248.338869]  __common_interrupt+0x6b/0x110
[  248.342967]  common_interrupt+0xbd/0xe0
[  248.346808]  </IRQ>
[  248.348912]  <TASK>
[  248.351018]  asm_common_interrupt+0x1e/0x40
[  248.355206] RIP: 0010:_raw_spin_unlock_irqrestore+0x1e/0x37
[  248.360780] Code: 02 5d c3 0f 1f 44 00 00 5d c3 66 90 0f 1f 44 00 00 
55 48 89 e5 c6 07 00 0f 1f 40 00 f7 c6 00 02 00 00 74 01 fb bf 01 00 00 
00 <e8> ed 8e 5b ff 65 8b 05 66 7e 52 78 85 c0 74 02 5d c3 0f 1f 44 00

[  248.379525] RSP: 0018:ffffb1c3a429b958 EFLAGS: 00000206
[  248.384749] RAX: 0000000000000001 RBX: ffff8e5d8973fd08 RCX: 
ffff8e5d8973fd10
[  248.391884] RDX: 0000000000000001 RSI: 0000000000000246 RDI: 
0000000000000001
[  248.399017] RBP: ffffb1c3a429b958 R08: 0000000000000000 R09: 
ffffb1c3a429b970
[  248.406148] R10: 0000000000000c00 R11: 0000000000000001 R12: 
0000000000000001
[  248.413280] R13: 0000000000000246 R14: 0000000000000000 R15: 
0000000000000003
[  248.420415]  __wake_up_common_lock+0x8a/0xc0
[  248.424686]  __wake_up+0x13/0x20
[  248.427919]  raid10_make_request+0x101/0x170 [raid10]
[  248.432971]  md_handle_request+0x179/0x1e0
[  248.437071]  ? submit_bio_checks+0x1f6/0x5a0
[  248.441345]  md_submit_bio+0x6d/0xa0
[  248.444924]  __submit_bio+0x94/0x140
[  248.448504]  submit_bio_noacct+0xe1/0x2a0
[  248.452515]  submit_bio+0x48/0x120
[  248.455923]  blkdev_direct_IO+0x220/0x540
[  248.459935]  ? __fsnotify_parent+0xff/0x330
[  248.464121]  ? __fsnotify_parent+0x10f/0x330
[  248.468393]  ? common_interrupt+0x73/0xe0
[  248.472408]  generic_file_read_iter+0xa5/0x160
[  248.476852]  blkdev_read_iter+0x38/0x70
[  248.480693]  io_read+0x119/0x420
[  248.483923]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.488805]  ? blk_mq_end_request_batch+0x378/0x490
[  248.493684]  io_issue_sqe+0x7ec/0x19c0
[  248.497436]  ? io_req_prep+0x6a9/0xe60
[  248.501190]  io_submit_sqes+0x2a0/0x9f0
[  248.505030]  ? __fget_files+0x6a/0x90
[  248.508693]  __x64_sys_io_uring_enter+0x1da/0x8c0
[  248.513401]  do_syscall_64+0x38/0x90
[  248.516979]  entry_SYSCALL_64_after_hwframe+0x44/0xae
[  248.522033] RIP: 0033:0x7fc26b19b89d
[  248.525611] Code: 00 c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa 
48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 
05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d c3 f5 0c 00 f7 d8 64 89 01 48
[  248.544360] RSP: 002b:00007fc26b07ce98 EFLAGS: 00000246 ORIG_RAX: 
00000000000001aa
[  248.551925] RAX: ffffffffffffffda RBX: 00007fc26b3f2fc0 RCX: 
00007fc26b19b89d
[  248.559058] RDX: 0000000000000020 RSI: 0000000000000020 RDI: 
0000000000000004
[  248.566189] RBP: 0000000000000020 R08: 0000000000000000 R09: 
0000000000000000
[  248.573322] R10: 0000000000000001 R11: 0000000000000246 R12: 
00005623a4b7a2a0
[  248.580456] R13: 0000000000000020 R14: 0000000000000020 R15: 
0000000000000020
[  248.587591]  </TASK>
quoted
quoted
         while (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery) &&
             bio->bi_iter.bi_sector < conf->reshape_progress &&
             bio->bi_iter.bi_sector + sectors > 
conf->reshape_progress) {
                 raid10_log(conf->mddev, "wait reshape");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                 allow_barrier(conf);
                 wait_event(conf->wait_barrier,
                            conf->reshape_progress <= 
bio->bi_iter.bi_sector ||
                            conf->reshape_progress >= 
bio->bi_iter.bi_sector +
                            sectors);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
         }
  }
@@ -1179,7 +1195,7 @@ static void raid10_read_request(struct mddev 
*mddev, struct bio *bio,
                 bio_chain(split, bio);
                 allow_barrier(conf);
                 submit_bio_noacct(bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 bio = split;
                 r10_bio->master_bio = bio;
                 r10_bio->sectors = max_sectors;
@@ -1338,7 +1354,7 @@ static void wait_blocked_dev(struct mddev 
*mddev, struct r10bio *r10_bio)
                 raid10_log(conf->mddev, "%s wait rdev %d blocked",
                                 __func__, blocked_rdev->raid_disk);
                 md_wait_for_blocked_rdev(blocked_rdev, mddev);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 goto retry_wait;
         }
  }
@@ -1357,6 +1373,11 @@ static void raid10_write_request(struct mddev 
*mddev, struct bio *bio,
bio_end_sector(bio)))) {
                 DEFINE_WAIT(w);
                 for (;;) {
+                       /* Bail out if REQ_NOWAIT is set for the bio */
+                       if (bio->bi_opf & REQ_NOWAIT) {
+                               bio_wouldblock_error(bio);
+                               return;
+                       }
prepare_to_wait(&conf->wait_barrier,
                                         &w, TASK_IDLE);
                         if (!md_cluster_ops->area_resyncing(mddev, 
WRITE,
@@ -1381,6 +1402,10 @@ static void raid10_write_request(struct mddev 
*mddev, struct bio *bio,
                               BIT(MD_SB_CHANGE_DEVS) | 
BIT(MD_SB_CHANGE_PENDING));
                 md_wakeup_thread(mddev->thread);
                 raid10_log(conf->mddev, "wait reshape metadata");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                 wait_event(mddev->sb_wait,
                            !test_bit(MD_SB_CHANGE_PENDING, 
&mddev->sb_flags));
@@ -1390,6 +1415,10 @@ static void raid10_write_request(struct mddev 
*mddev, struct bio *bio,
         if (conf->pending_count >= max_queued_requests) {
                 md_wakeup_thread(mddev->thread);
                 raid10_log(mddev, "wait queued");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                 wait_event(conf->wait_barrier,
                            conf->pending_count < max_queued_requests);
         }
@@ -1482,7 +1511,7 @@ static void raid10_write_request(struct mddev 
*mddev, struct bio *bio,
                 bio_chain(split, bio);
                 allow_barrier(conf);
                 submit_bio_noacct(bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 bio = split;
                 r10_bio->master_bio = bio;
         }
@@ -1607,7 +1636,11 @@ static int raid10_handle_discard(struct mddev 
*mddev, struct bio *bio)
         if (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery))
                 return -EAGAIN;

-       wait_barrier(conf);
+       if (bio->bi_opf & REQ_NOWAIT) {
+               bio_wouldblock_error(bio);
+               return 0;
+       }
+       wait_barrier(conf, false);

         /*
          * Check reshape again to avoid reshape happens after checking
@@ -1649,7 +1682,7 @@ static int raid10_handle_discard(struct mddev 
*mddev, struct bio *bio)
                 allow_barrier(conf);
                 /* Resend the fist split part */
                 submit_bio_noacct(split);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
         }
         div_u64_rem(bio_end, stripe_size, &remainder);
         if (remainder) {
@@ -1660,7 +1693,7 @@ static int raid10_handle_discard(struct mddev 
*mddev, struct bio *bio)
                 /* Resend the second split part */
                 submit_bio_noacct(bio);
                 bio = split;
-               wait_barrier(conf);
+               wait_barrier(conf, false);
         }

         bio_start = bio->bi_iter.bi_sector;
@@ -1816,7 +1849,7 @@ static int raid10_handle_discard(struct mddev 
*mddev, struct bio *bio)
                 end_disk_offset += geo->stride;
                 atomic_inc(&first_r10bio->remaining);
                 raid_end_discard_bio(r10_bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 goto retry_discard;
         }
@@ -2011,7 +2044,7 @@ static void print_conf(struct r10conf *conf)
  static void close_sync(struct r10conf *conf)
  {
-       wait_barrier(conf);
+       wait_barrier(conf, false);
         allow_barrier(conf);

         mempool_exit(&conf->r10buf_pool);
@@ -4819,7 +4852,7 @@ static sector_t reshape_request(struct mddev 
*mddev, sector_t sector_nr,
         if (need_flush ||
             time_after(jiffies, conf->reshape_checkpoint + 10*HZ)) {
                 /* Need to update reshape_position in metadata */
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 mddev->reshape_position = conf->reshape_progress;
                 if (mddev->reshape_backwards)
                         mddev->curr_resync_completed = 
raid10_size(mddev, 0, 0)
-- 
2.17.1

Re: [PATCH v5 3/4] md: raid10 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-16 16:41:03

On 12/15/21 5:30 PM, Vishal Verma wrote:
On 12/15/21 3:20 PM, Vishal Verma wrote:
quoted
On 12/15/21 1:42 PM, Song Liu wrote:
quoted
On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma 
[off-list ref] wrote:
quoted
This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

- Waiting for the barrier,
- Too many pending I/Os to be queued,
- Reshape operation,
- Discard operation.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
  drivers/md/raid10.c | 57 
+++++++++++++++++++++++++++++++++++----------
  1 file changed, 45 insertions(+), 12 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..f6c73987e9ac 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,11 +952,18 @@ static void lower_barrier(struct r10conf *conf)
         wake_up(&conf->wait_barrier);
  }

-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
  {
         spin_lock_irq(&conf->resync_lock);
         if (conf->barrier) {
                 struct bio_list *bio_list = current->bio_list;
+
+               /* Return false when nowait flag is set */
+               if (nowait) {
+ spin_unlock_irq(&conf->resync_lock);
+                       return false;
+               }
+
                 conf->nr_waiting++;
                 /* Wait for the barrier to drop.
                  * However if there are already pending
@@ -988,6 +995,7 @@ static void wait_barrier(struct r10conf *conf)
         }
         atomic_inc(&conf->nr_pending);
         spin_unlock_irq(&conf->resync_lock);
+       return true;
  }

  static void allow_barrier(struct r10conf *conf)
@@ -1101,17 +1109,25 @@ static void raid10_unplug(struct 
blk_plug_cb *cb, bool from_schedule)
  static void regular_request_wait(struct mddev *mddev, struct 
r10conf *conf,
                                  struct bio *bio, sector_t sectors)
  {
-       wait_barrier(conf);
+       /* Bail out if REQ_NOWAIT is set for the bio */
+       if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }
I think we also need regular_request_wait to return bool and handle 
it properly.

Thanks,
Song
Ack, will fix it. Thanks!
Ran into this while running with io_uring. With the current v5 (raid10 
patch) on top of md-next branch.
./t/io_uring -a 0 -d 256 </dev/raid10>

It didn't trigger with aio (-a 1)

[  248.128661] BUG: kernel NULL pointer dereference, address: 
00000000000000b8
[  248.135628] #PF: supervisor read access in kernel mode
[  248.140762] #PF: error_code(0x0000) - not-present page
[  248.145903] PGD 0 P4D 0
[  248.148443] Oops: 0000 [#1] PREEMPT SMP NOPTI
[  248.152800] CPU: 49 PID: 9461 Comm: io_uring Kdump: loaded Not 
tainted 5.16.0-rc3+ #2
[  248.160629] Hardware name: Dell Inc. PowerEdge R650xs/0PPTY2, BIOS 
1.3.8 08/31/2021
[  248.168279] RIP: 0010:raid10_end_read_request+0x74/0x140 [raid10]
[  248.174373] Code: 48 60 48 8b 58 58 48 c1 e2 05 49 03 55 08 48 89 
4a 10 40 84 f6 75 48 f0 41 80 4c 24 18 01 4c 89 e7 e8 e0 b8 ff ff 49 
8b 4d 00 <48> 8b 83 b8 00 00 00 f0 ff 8b f0 00 00 00 0f 94 c2 a8 01 74 
04 84
[  248.193120] RSP: 0018:ffffb1c38d598ce8 EFLAGS: 00010086
[  248.198344] RAX: ffff8e5da2a1a100 RBX: 0000000000000000 RCX: 
ffff8e5d89747000
[  248.205479] RDX: 000000008040003a RSI: 0000000080400039 RDI: 
ffff8e1e00044900
[  248.212611] RBP: ffffb1c38d598d30 R08: 0000000000000000 R09: 
0000000000000001
[  248.219744] R10: ffff8e5da2a1ae00 R11: 000000411bab9000 R12: 
ffff8e5da2a1ae00
[  248.226877] R13: ffff8e5d8973fc00 R14: 0000000000000000 R15: 
0000000000001000
[  248.234009] FS:  00007fc26b07d700(0000) GS:ffff8e9c6e600000(0000) 
knlGS:0000000000000000
[  248.242096] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  248.247843] CR2: 00000000000000b8 CR3: 00000040b25d4005 CR4: 
0000000000770ee0
[  248.254973] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 
0000000000000000
[  248.262107] DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7: 
0000000000000400
[  248.269240] PKRU: 55555554
[  248.271953] Call Trace:
[  248.274406]  <IRQ>
[  248.276425]  bio_endio+0xf6/0x170
[  248.279743]  blk_update_request+0x12d/0x470
[  248.283931]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.288809]  blk_mq_end_request_batch+0x76/0x490
[  248.293429]  ? dma_direct_unmap_sg+0xdd/0x1a0
[  248.297786]  ? smp_call_function_single_async+0x46/0x70
[  248.303015]  ? mempool_kfree+0xe/0x10
[  248.306680]  ? mempool_kfree+0xe/0x10
[  248.310345]  nvme_pci_complete_batch+0x26/0xb0
[  248.314792]  nvme_irq+0x298/0x2f0
[  248.318110]  ? nvme_unmap_data+0xf0/0xf0
[  248.322038]  __handle_irq_event_percpu+0x3f/0x190
[  248.326744]  handle_irq_event_percpu+0x33/0x80
[  248.331190]  handle_irq_event+0x39/0x60
[  248.335028]  handle_edge_irq+0xbe/0x1e0
[  248.338869]  __common_interrupt+0x6b/0x110
[  248.342967]  common_interrupt+0xbd/0xe0
[  248.346808]  </IRQ>
[  248.348912]  <TASK>
[  248.351018]  asm_common_interrupt+0x1e/0x40
[  248.355206] RIP: 0010:_raw_spin_unlock_irqrestore+0x1e/0x37
[  248.360780] Code: 02 5d c3 0f 1f 44 00 00 5d c3 66 90 0f 1f 44 00 
00 55 48 89 e5 c6 07 00 0f 1f 40 00 f7 c6 00 02 00 00 74 01 fb bf 01 
00 00 00 <e8> ed 8e 5b ff 65 8b 05 66 7e 52 78 85 c0 74 02 5d c3 0f 1f 
44 00

[  248.379525] RSP: 0018:ffffb1c3a429b958 EFLAGS: 00000206
[  248.384749] RAX: 0000000000000001 RBX: ffff8e5d8973fd08 RCX: 
ffff8e5d8973fd10
[  248.391884] RDX: 0000000000000001 RSI: 0000000000000246 RDI: 
0000000000000001
[  248.399017] RBP: ffffb1c3a429b958 R08: 0000000000000000 R09: 
ffffb1c3a429b970
[  248.406148] R10: 0000000000000c00 R11: 0000000000000001 R12: 
0000000000000001
[  248.413280] R13: 0000000000000246 R14: 0000000000000000 R15: 
0000000000000003
[  248.420415]  __wake_up_common_lock+0x8a/0xc0
[  248.424686]  __wake_up+0x13/0x20
[  248.427919]  raid10_make_request+0x101/0x170 [raid10]
[  248.432971]  md_handle_request+0x179/0x1e0
[  248.437071]  ? submit_bio_checks+0x1f6/0x5a0
[  248.441345]  md_submit_bio+0x6d/0xa0
[  248.444924]  __submit_bio+0x94/0x140
[  248.448504]  submit_bio_noacct+0xe1/0x2a0
[  248.452515]  submit_bio+0x48/0x120
[  248.455923]  blkdev_direct_IO+0x220/0x540
[  248.459935]  ? __fsnotify_parent+0xff/0x330
[  248.464121]  ? __fsnotify_parent+0x10f/0x330
[  248.468393]  ? common_interrupt+0x73/0xe0
[  248.472408]  generic_file_read_iter+0xa5/0x160
[  248.476852]  blkdev_read_iter+0x38/0x70
[  248.480693]  io_read+0x119/0x420
[  248.483923]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.488805]  ? blk_mq_end_request_batch+0x378/0x490
[  248.493684]  io_issue_sqe+0x7ec/0x19c0
[  248.497436]  ? io_req_prep+0x6a9/0xe60
[  248.501190]  io_submit_sqes+0x2a0/0x9f0
[  248.505030]  ? __fget_files+0x6a/0x90
[  248.508693]  __x64_sys_io_uring_enter+0x1da/0x8c0
[  248.513401]  do_syscall_64+0x38/0x90
[  248.516979]  entry_SYSCALL_64_after_hwframe+0x44/0xae
[  248.522033] RIP: 0033:0x7fc26b19b89d
[  248.525611] Code: 00 c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e 
fa 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 
08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d c3 f5 0c 00 f7 d8 64 89 
01 48
[  248.544360] RSP: 002b:00007fc26b07ce98 EFLAGS: 00000246 ORIG_RAX: 
00000000000001aa
[  248.551925] RAX: ffffffffffffffda RBX: 00007fc26b3f2fc0 RCX: 
00007fc26b19b89d
[  248.559058] RDX: 0000000000000020 RSI: 0000000000000020 RDI: 
0000000000000004
[  248.566189] RBP: 0000000000000020 R08: 0000000000000000 R09: 
0000000000000000
[  248.573322] R10: 0000000000000001 R11: 0000000000000246 R12: 
00005623a4b7a2a0
[  248.580456] R13: 0000000000000020 R14: 0000000000000020 R15: 
0000000000000020
[  248.587591]  </TASK>
It seems this issue is triggering even when just using "md: add support 
for REQ_NOWAIT" patch running t/io_uring against a raid10 volume with 
very high iodepth (256).
quoted
quoted
quoted
         while (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery) &&
             bio->bi_iter.bi_sector < conf->reshape_progress &&
             bio->bi_iter.bi_sector + sectors > 
conf->reshape_progress) {
                 raid10_log(conf->mddev, "wait reshape");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                 allow_barrier(conf);
                 wait_event(conf->wait_barrier,
                            conf->reshape_progress <= 
bio->bi_iter.bi_sector ||
                            conf->reshape_progress >= 
bio->bi_iter.bi_sector +
                            sectors);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
         }
  }
@@ -1179,7 +1195,7 @@ static void raid10_read_request(struct mddev 
*mddev, struct bio *bio,
                 bio_chain(split, bio);
                 allow_barrier(conf);
                 submit_bio_noacct(bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 bio = split;
                 r10_bio->master_bio = bio;
                 r10_bio->sectors = max_sectors;
@@ -1338,7 +1354,7 @@ static void wait_blocked_dev(struct mddev 
*mddev, struct r10bio *r10_bio)
                 raid10_log(conf->mddev, "%s wait rdev %d blocked",
                                 __func__, blocked_rdev->raid_disk);
                 md_wait_for_blocked_rdev(blocked_rdev, mddev);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 goto retry_wait;
         }
  }
@@ -1357,6 +1373,11 @@ static void raid10_write_request(struct 
mddev *mddev, struct bio *bio,
bio_end_sector(bio)))) {
                 DEFINE_WAIT(w);
                 for (;;) {
+                       /* Bail out if REQ_NOWAIT is set for the 
bio */
+                       if (bio->bi_opf & REQ_NOWAIT) {
+                               bio_wouldblock_error(bio);
+                               return;
+                       }
prepare_to_wait(&conf->wait_barrier,
                                         &w, TASK_IDLE);
                         if (!md_cluster_ops->area_resyncing(mddev, 
WRITE,
@@ -1381,6 +1402,10 @@ static void raid10_write_request(struct 
mddev *mddev, struct bio *bio,
                               BIT(MD_SB_CHANGE_DEVS) | 
BIT(MD_SB_CHANGE_PENDING));
                 md_wakeup_thread(mddev->thread);
                 raid10_log(conf->mddev, "wait reshape metadata");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                 wait_event(mddev->sb_wait,
                            !test_bit(MD_SB_CHANGE_PENDING, 
&mddev->sb_flags));
@@ -1390,6 +1415,10 @@ static void raid10_write_request(struct 
mddev *mddev, struct bio *bio,
         if (conf->pending_count >= max_queued_requests) {
                 md_wakeup_thread(mddev->thread);
                 raid10_log(mddev, "wait queued");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                 wait_event(conf->wait_barrier,
                            conf->pending_count < 
max_queued_requests);
         }
@@ -1482,7 +1511,7 @@ static void raid10_write_request(struct mddev 
*mddev, struct bio *bio,
                 bio_chain(split, bio);
                 allow_barrier(conf);
                 submit_bio_noacct(bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 bio = split;
                 r10_bio->master_bio = bio;
         }
@@ -1607,7 +1636,11 @@ static int raid10_handle_discard(struct 
mddev *mddev, struct bio *bio)
         if (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery))
                 return -EAGAIN;

-       wait_barrier(conf);
+       if (bio->bi_opf & REQ_NOWAIT) {
+               bio_wouldblock_error(bio);
+               return 0;
+       }
+       wait_barrier(conf, false);

         /*
          * Check reshape again to avoid reshape happens after 
checking
@@ -1649,7 +1682,7 @@ static int raid10_handle_discard(struct mddev 
*mddev, struct bio *bio)
                 allow_barrier(conf);
                 /* Resend the fist split part */
                 submit_bio_noacct(split);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
         }
         div_u64_rem(bio_end, stripe_size, &remainder);
         if (remainder) {
@@ -1660,7 +1693,7 @@ static int raid10_handle_discard(struct mddev 
*mddev, struct bio *bio)
                 /* Resend the second split part */
                 submit_bio_noacct(bio);
                 bio = split;
-               wait_barrier(conf);
+               wait_barrier(conf, false);
         }

         bio_start = bio->bi_iter.bi_sector;
@@ -1816,7 +1849,7 @@ static int raid10_handle_discard(struct mddev 
*mddev, struct bio *bio)
                 end_disk_offset += geo->stride;
atomic_inc(&first_r10bio->remaining);
                 raid_end_discard_bio(r10_bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 goto retry_discard;
         }
@@ -2011,7 +2044,7 @@ static void print_conf(struct r10conf *conf)
  static void close_sync(struct r10conf *conf)
  {
-       wait_barrier(conf);
+       wait_barrier(conf, false);
         allow_barrier(conf);

         mempool_exit(&conf->r10buf_pool);
@@ -4819,7 +4852,7 @@ static sector_t reshape_request(struct mddev 
*mddev, sector_t sector_nr,
         if (need_flush ||
             time_after(jiffies, conf->reshape_checkpoint + 10*HZ)) {
                 /* Need to update reshape_position in metadata */
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                 mddev->reshape_position = conf->reshape_progress;
                 if (mddev->reshape_backwards)
                         mddev->curr_resync_completed = 
raid10_size(mddev, 0, 0)
-- 
2.17.1

Re: [PATCH v5 3/4] md: raid10 add nowait support

From: Jens Axboe <axboe@kernel.dk>
Date: 2021-12-16 16:42:50

On 12/15/21 5:30 PM, Vishal Verma wrote:
On 12/15/21 3:20 PM, Vishal Verma wrote:
quoted
On 12/15/21 1:42 PM, Song Liu wrote:
quoted
On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma 
[off-list ref] wrote:
quoted
This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

- Waiting for the barrier,
- Too many pending I/Os to be queued,
- Reshape operation,
- Discard operation.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
  drivers/md/raid10.c | 57 
+++++++++++++++++++++++++++++++++++----------
  1 file changed, 45 insertions(+), 12 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..f6c73987e9ac 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,11 +952,18 @@ static void lower_barrier(struct r10conf *conf)
         wake_up(&conf->wait_barrier);
  }

-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
  {
         spin_lock_irq(&conf->resync_lock);
         if (conf->barrier) {
                 struct bio_list *bio_list = current->bio_list;
+
+               /* Return false when nowait flag is set */
+               if (nowait) {
+ spin_unlock_irq(&conf->resync_lock);
+                       return false;
+               }
+
                 conf->nr_waiting++;
                 /* Wait for the barrier to drop.
                  * However if there are already pending
@@ -988,6 +995,7 @@ static void wait_barrier(struct r10conf *conf)
         }
         atomic_inc(&conf->nr_pending);
         spin_unlock_irq(&conf->resync_lock);
+       return true;
  }

  static void allow_barrier(struct r10conf *conf)
@@ -1101,17 +1109,25 @@ static void raid10_unplug(struct blk_plug_cb 
*cb, bool from_schedule)
  static void regular_request_wait(struct mddev *mddev, struct 
r10conf *conf,
                                  struct bio *bio, sector_t sectors)
  {
-       wait_barrier(conf);
+       /* Bail out if REQ_NOWAIT is set for the bio */
+       if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }
I think we also need regular_request_wait to return bool and handle 
it properly.

Thanks,
Song
Ack, will fix it. Thanks!
Ran into this while running with io_uring. With the current v5 (raid10 
patch) on top of md-next branch.
./t/io_uring -a 0 -d 256 </dev/raid10>

It didn't trigger with aio (-a 1)

[  248.128661] BUG: kernel NULL pointer dereference, address: 
00000000000000b8
[  248.135628] #PF: supervisor read access in kernel mode
[  248.140762] #PF: error_code(0x0000) - not-present page
[  248.145903] PGD 0 P4D 0
[  248.148443] Oops: 0000 [#1] PREEMPT SMP NOPTI
[  248.152800] CPU: 49 PID: 9461 Comm: io_uring Kdump: loaded Not 
tainted 5.16.0-rc3+ #2
[  248.160629] Hardware name: Dell Inc. PowerEdge R650xs/0PPTY2, BIOS 
1.3.8 08/31/2021
[  248.168279] RIP: 0010:raid10_end_read_request+0x74/0x140 [raid10]
[  248.174373] Code: 48 60 48 8b 58 58 48 c1 e2 05 49 03 55 08 48 89 4a 
10 40 84 f6 75 48 f0 41 80 4c 24 18 01 4c 89 e7 e8 e0 b8 ff ff 49 8b 4d 
00 <48> 8b 83 b8 00 00 00 f0 ff 8b f0 00 00 00 0f 94 c2 a8 01 74 04 84
[  248.193120] RSP: 0018:ffffb1c38d598ce8 EFLAGS: 00010086
[  248.198344] RAX: ffff8e5da2a1a100 RBX: 0000000000000000 RCX: 
ffff8e5d89747000
[  248.205479] RDX: 000000008040003a RSI: 0000000080400039 RDI: 
ffff8e1e00044900
[  248.212611] RBP: ffffb1c38d598d30 R08: 0000000000000000 R09: 
0000000000000001
[  248.219744] R10: ffff8e5da2a1ae00 R11: 000000411bab9000 R12: 
ffff8e5da2a1ae00
[  248.226877] R13: ffff8e5d8973fc00 R14: 0000000000000000 R15: 
0000000000001000
[  248.234009] FS:  00007fc26b07d700(0000) GS:ffff8e9c6e600000(0000) 
knlGS:0000000000000000
[  248.242096] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  248.247843] CR2: 00000000000000b8 CR3: 00000040b25d4005 CR4: 
0000000000770ee0
[  248.254973] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 
0000000000000000
[  248.262107] DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7: 
0000000000000400
[  248.269240] PKRU: 55555554
[  248.271953] Call Trace:
[  248.274406]  <IRQ>
[  248.276425]  bio_endio+0xf6/0x170
[  248.279743]  blk_update_request+0x12d/0x470
[  248.283931]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.288809]  blk_mq_end_request_batch+0x76/0x490
[  248.293429]  ? dma_direct_unmap_sg+0xdd/0x1a0
[  248.297786]  ? smp_call_function_single_async+0x46/0x70
[  248.303015]  ? mempool_kfree+0xe/0x10
[  248.306680]  ? mempool_kfree+0xe/0x10
[  248.310345]  nvme_pci_complete_batch+0x26/0xb0
[  248.314792]  nvme_irq+0x298/0x2f0
[  248.318110]  ? nvme_unmap_data+0xf0/0xf0
[  248.322038]  __handle_irq_event_percpu+0x3f/0x190
[  248.326744]  handle_irq_event_percpu+0x33/0x80
[  248.331190]  handle_irq_event+0x39/0x60
[  248.335028]  handle_edge_irq+0xbe/0x1e0
[  248.338869]  __common_interrupt+0x6b/0x110
[  248.342967]  common_interrupt+0xbd/0xe0
[  248.346808]  </IRQ>
[  248.348912]  <TASK>
[  248.351018]  asm_common_interrupt+0x1e/0x40
[  248.355206] RIP: 0010:_raw_spin_unlock_irqrestore+0x1e/0x37
[  248.360780] Code: 02 5d c3 0f 1f 44 00 00 5d c3 66 90 0f 1f 44 00 00 
55 48 89 e5 c6 07 00 0f 1f 40 00 f7 c6 00 02 00 00 74 01 fb bf 01 00 00 
00 <e8> ed 8e 5b ff 65 8b 05 66 7e 52 78 85 c0 74 02 5d c3 0f 1f 44 00

[  248.379525] RSP: 0018:ffffb1c3a429b958 EFLAGS: 00000206
[  248.384749] RAX: 0000000000000001 RBX: ffff8e5d8973fd08 RCX: 
ffff8e5d8973fd10
[  248.391884] RDX: 0000000000000001 RSI: 0000000000000246 RDI: 
0000000000000001
[  248.399017] RBP: ffffb1c3a429b958 R08: 0000000000000000 R09: 
ffffb1c3a429b970
[  248.406148] R10: 0000000000000c00 R11: 0000000000000001 R12: 
0000000000000001
[  248.413280] R13: 0000000000000246 R14: 0000000000000000 R15: 
0000000000000003
[  248.420415]  __wake_up_common_lock+0x8a/0xc0
[  248.424686]  __wake_up+0x13/0x20
[  248.427919]  raid10_make_request+0x101/0x170 [raid10]
[  248.432971]  md_handle_request+0x179/0x1e0
[  248.437071]  ? submit_bio_checks+0x1f6/0x5a0
[  248.441345]  md_submit_bio+0x6d/0xa0
[  248.444924]  __submit_bio+0x94/0x140
[  248.448504]  submit_bio_noacct+0xe1/0x2a0
[  248.452515]  submit_bio+0x48/0x120
[  248.455923]  blkdev_direct_IO+0x220/0x540
[  248.459935]  ? __fsnotify_parent+0xff/0x330
[  248.464121]  ? __fsnotify_parent+0x10f/0x330
[  248.468393]  ? common_interrupt+0x73/0xe0
[  248.472408]  generic_file_read_iter+0xa5/0x160
[  248.476852]  blkdev_read_iter+0x38/0x70
[  248.480693]  io_read+0x119/0x420
[  248.483923]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.488805]  ? blk_mq_end_request_batch+0x378/0x490
[  248.493684]  io_issue_sqe+0x7ec/0x19c0
[  248.497436]  ? io_req_prep+0x6a9/0xe60
[  248.501190]  io_submit_sqes+0x2a0/0x9f0
[  248.505030]  ? __fget_files+0x6a/0x90
[  248.508693]  __x64_sys_io_uring_enter+0x1da/0x8c0
[  248.513401]  do_syscall_64+0x38/0x90
[  248.516979]  entry_SYSCALL_64_after_hwframe+0x44/0xae
[  248.522033] RIP: 0033:0x7fc26b19b89d
[  248.525611] Code: 00 c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa 
48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 
05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d c3 f5 0c 00 f7 d8 64 89 01 48
[  248.544360] RSP: 002b:00007fc26b07ce98 EFLAGS: 00000246 ORIG_RAX: 
00000000000001aa
[  248.551925] RAX: ffffffffffffffda RBX: 00007fc26b3f2fc0 RCX: 
00007fc26b19b89d
[  248.559058] RDX: 0000000000000020 RSI: 0000000000000020 RDI: 
0000000000000004
[  248.566189] RBP: 0000000000000020 R08: 0000000000000000 R09: 
0000000000000000
[  248.573322] R10: 0000000000000001 R11: 0000000000000246 R12: 
00005623a4b7a2a0
[  248.580456] R13: 0000000000000020 R14: 0000000000000020 R15: 
0000000000000020
[  248.587591]  </TASK>
Do you have:

commit 75feae73a28020e492fbad2323245455ef69d687
Author: Pavel Begunkov [off-list ref]
Date:   Tue Dec 7 20:16:36 2021 +0000

    block: fix single bio async DIO error handling

in your tree?

-- 
Jens Axboe

Re: [PATCH v5 3/4] md: raid10 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-16 16:45:22

On 12/16/21 9:42 AM, Jens Axboe wrote:
On 12/15/21 5:30 PM, Vishal Verma wrote:
quoted
On 12/15/21 3:20 PM, Vishal Verma wrote:
quoted
On 12/15/21 1:42 PM, Song Liu wrote:
quoted
On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma
[off-list ref] wrote:
quoted
This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

- Waiting for the barrier,
- Too many pending I/Os to be queued,
- Reshape operation,
- Discard operation.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
   drivers/md/raid10.c | 57
+++++++++++++++++++++++++++++++++++----------
   1 file changed, 45 insertions(+), 12 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..f6c73987e9ac 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,11 +952,18 @@ static void lower_barrier(struct r10conf *conf)
          wake_up(&conf->wait_barrier);
   }

-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
   {
          spin_lock_irq(&conf->resync_lock);
          if (conf->barrier) {
                  struct bio_list *bio_list = current->bio_list;
+
+               /* Return false when nowait flag is set */
+               if (nowait) {
+ spin_unlock_irq(&conf->resync_lock);
+                       return false;
+               }
+
                  conf->nr_waiting++;
                  /* Wait for the barrier to drop.
                   * However if there are already pending
@@ -988,6 +995,7 @@ static void wait_barrier(struct r10conf *conf)
          }
          atomic_inc(&conf->nr_pending);
          spin_unlock_irq(&conf->resync_lock);
+       return true;
   }

   static void allow_barrier(struct r10conf *conf)
@@ -1101,17 +1109,25 @@ static void raid10_unplug(struct blk_plug_cb
*cb, bool from_schedule)
   static void regular_request_wait(struct mddev *mddev, struct
r10conf *conf,
                                   struct bio *bio, sector_t sectors)
   {
-       wait_barrier(conf);
+       /* Bail out if REQ_NOWAIT is set for the bio */
+       if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }
I think we also need regular_request_wait to return bool and handle
it properly.

Thanks,
Song
Ack, will fix it. Thanks!
Ran into this while running with io_uring. With the current v5 (raid10
patch) on top of md-next branch.
./t/io_uring -a 0 -d 256 </dev/raid10>

It didn't trigger with aio (-a 1)

[  248.128661] BUG: kernel NULL pointer dereference, address:
00000000000000b8
[  248.135628] #PF: supervisor read access in kernel mode
[  248.140762] #PF: error_code(0x0000) - not-present page
[  248.145903] PGD 0 P4D 0
[  248.148443] Oops: 0000 [#1] PREEMPT SMP NOPTI
[  248.152800] CPU: 49 PID: 9461 Comm: io_uring Kdump: loaded Not
tainted 5.16.0-rc3+ #2
[  248.160629] Hardware name: Dell Inc. PowerEdge R650xs/0PPTY2, BIOS
1.3.8 08/31/2021
[  248.168279] RIP: 0010:raid10_end_read_request+0x74/0x140 [raid10]
[  248.174373] Code: 48 60 48 8b 58 58 48 c1 e2 05 49 03 55 08 48 89 4a
10 40 84 f6 75 48 f0 41 80 4c 24 18 01 4c 89 e7 e8 e0 b8 ff ff 49 8b 4d
00 <48> 8b 83 b8 00 00 00 f0 ff 8b f0 00 00 00 0f 94 c2 a8 01 74 04 84
[  248.193120] RSP: 0018:ffffb1c38d598ce8 EFLAGS: 00010086
[  248.198344] RAX: ffff8e5da2a1a100 RBX: 0000000000000000 RCX:
ffff8e5d89747000
[  248.205479] RDX: 000000008040003a RSI: 0000000080400039 RDI:
ffff8e1e00044900
[  248.212611] RBP: ffffb1c38d598d30 R08: 0000000000000000 R09:
0000000000000001
[  248.219744] R10: ffff8e5da2a1ae00 R11: 000000411bab9000 R12:
ffff8e5da2a1ae00
[  248.226877] R13: ffff8e5d8973fc00 R14: 0000000000000000 R15:
0000000000001000
[  248.234009] FS:  00007fc26b07d700(0000) GS:ffff8e9c6e600000(0000)
knlGS:0000000000000000
[  248.242096] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  248.247843] CR2: 00000000000000b8 CR3: 00000040b25d4005 CR4:
0000000000770ee0
[  248.254973] DR0: 0000000000000000 DR1: 0000000000000000 DR2:
0000000000000000
[  248.262107] DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7:
0000000000000400
[  248.269240] PKRU: 55555554
[  248.271953] Call Trace:
[  248.274406]  <IRQ>
[  248.276425]  bio_endio+0xf6/0x170
[  248.279743]  blk_update_request+0x12d/0x470
[  248.283931]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.288809]  blk_mq_end_request_batch+0x76/0x490
[  248.293429]  ? dma_direct_unmap_sg+0xdd/0x1a0
[  248.297786]  ? smp_call_function_single_async+0x46/0x70
[  248.303015]  ? mempool_kfree+0xe/0x10
[  248.306680]  ? mempool_kfree+0xe/0x10
[  248.310345]  nvme_pci_complete_batch+0x26/0xb0
[  248.314792]  nvme_irq+0x298/0x2f0
[  248.318110]  ? nvme_unmap_data+0xf0/0xf0
[  248.322038]  __handle_irq_event_percpu+0x3f/0x190
[  248.326744]  handle_irq_event_percpu+0x33/0x80
[  248.331190]  handle_irq_event+0x39/0x60
[  248.335028]  handle_edge_irq+0xbe/0x1e0
[  248.338869]  __common_interrupt+0x6b/0x110
[  248.342967]  common_interrupt+0xbd/0xe0
[  248.346808]  </IRQ>
[  248.348912]  <TASK>
[  248.351018]  asm_common_interrupt+0x1e/0x40
[  248.355206] RIP: 0010:_raw_spin_unlock_irqrestore+0x1e/0x37
[  248.360780] Code: 02 5d c3 0f 1f 44 00 00 5d c3 66 90 0f 1f 44 00 00
55 48 89 e5 c6 07 00 0f 1f 40 00 f7 c6 00 02 00 00 74 01 fb bf 01 00 00
00 <e8> ed 8e 5b ff 65 8b 05 66 7e 52 78 85 c0 74 02 5d c3 0f 1f 44 00

[  248.379525] RSP: 0018:ffffb1c3a429b958 EFLAGS: 00000206
[  248.384749] RAX: 0000000000000001 RBX: ffff8e5d8973fd08 RCX:
ffff8e5d8973fd10
[  248.391884] RDX: 0000000000000001 RSI: 0000000000000246 RDI:
0000000000000001
[  248.399017] RBP: ffffb1c3a429b958 R08: 0000000000000000 R09:
ffffb1c3a429b970
[  248.406148] R10: 0000000000000c00 R11: 0000000000000001 R12:
0000000000000001
[  248.413280] R13: 0000000000000246 R14: 0000000000000000 R15:
0000000000000003
[  248.420415]  __wake_up_common_lock+0x8a/0xc0
[  248.424686]  __wake_up+0x13/0x20
[  248.427919]  raid10_make_request+0x101/0x170 [raid10]
[  248.432971]  md_handle_request+0x179/0x1e0
[  248.437071]  ? submit_bio_checks+0x1f6/0x5a0
[  248.441345]  md_submit_bio+0x6d/0xa0
[  248.444924]  __submit_bio+0x94/0x140
[  248.448504]  submit_bio_noacct+0xe1/0x2a0
[  248.452515]  submit_bio+0x48/0x120
[  248.455923]  blkdev_direct_IO+0x220/0x540
[  248.459935]  ? __fsnotify_parent+0xff/0x330
[  248.464121]  ? __fsnotify_parent+0x10f/0x330
[  248.468393]  ? common_interrupt+0x73/0xe0
[  248.472408]  generic_file_read_iter+0xa5/0x160
[  248.476852]  blkdev_read_iter+0x38/0x70
[  248.480693]  io_read+0x119/0x420
[  248.483923]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.488805]  ? blk_mq_end_request_batch+0x378/0x490
[  248.493684]  io_issue_sqe+0x7ec/0x19c0
[  248.497436]  ? io_req_prep+0x6a9/0xe60
[  248.501190]  io_submit_sqes+0x2a0/0x9f0
[  248.505030]  ? __fget_files+0x6a/0x90
[  248.508693]  __x64_sys_io_uring_enter+0x1da/0x8c0
[  248.513401]  do_syscall_64+0x38/0x90
[  248.516979]  entry_SYSCALL_64_after_hwframe+0x44/0xae
[  248.522033] RIP: 0033:0x7fc26b19b89d
[  248.525611] Code: 00 c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa
48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f
05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d c3 f5 0c 00 f7 d8 64 89 01 48
[  248.544360] RSP: 002b:00007fc26b07ce98 EFLAGS: 00000246 ORIG_RAX:
00000000000001aa
[  248.551925] RAX: ffffffffffffffda RBX: 00007fc26b3f2fc0 RCX:
00007fc26b19b89d
[  248.559058] RDX: 0000000000000020 RSI: 0000000000000020 RDI:
0000000000000004
[  248.566189] RBP: 0000000000000020 R08: 0000000000000000 R09:
0000000000000000
[  248.573322] R10: 0000000000000001 R11: 0000000000000246 R12:
00005623a4b7a2a0
[  248.580456] R13: 0000000000000020 R14: 0000000000000020 R15:
0000000000000020
[  248.587591]  </TASK>
Do you have:

commit 75feae73a28020e492fbad2323245455ef69d687
Author: Pavel Begunkov [off-list ref]
Date:   Tue Dec 7 20:16:36 2021 +0000

     block: fix single bio async DIO error handling

in your tree?
Nope. I will get it in and test. Thanks!

Re: [PATCH v5 3/4] md: raid10 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-16 18:14:27

On 12/16/21 9:42 AM, Jens Axboe wrote:
On 12/15/21 5:30 PM, Vishal Verma wrote:
quoted
On 12/15/21 3:20 PM, Vishal Verma wrote:
quoted
On 12/15/21 1:42 PM, Song Liu wrote:
quoted
On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma
[off-list ref] wrote:
quoted
This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

- Waiting for the barrier,
- Too many pending I/Os to be queued,
- Reshape operation,
- Discard operation.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
   drivers/md/raid10.c | 57
+++++++++++++++++++++++++++++++++++----------
   1 file changed, 45 insertions(+), 12 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..f6c73987e9ac 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,11 +952,18 @@ static void lower_barrier(struct r10conf *conf)
          wake_up(&conf->wait_barrier);
   }

-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
   {
          spin_lock_irq(&conf->resync_lock);
          if (conf->barrier) {
                  struct bio_list *bio_list = current->bio_list;
+
+               /* Return false when nowait flag is set */
+               if (nowait) {
+ spin_unlock_irq(&conf->resync_lock);
+                       return false;
+               }
+
                  conf->nr_waiting++;
                  /* Wait for the barrier to drop.
                   * However if there are already pending
@@ -988,6 +995,7 @@ static void wait_barrier(struct r10conf *conf)
          }
          atomic_inc(&conf->nr_pending);
          spin_unlock_irq(&conf->resync_lock);
+       return true;
   }

   static void allow_barrier(struct r10conf *conf)
@@ -1101,17 +1109,25 @@ static void raid10_unplug(struct blk_plug_cb
*cb, bool from_schedule)
   static void regular_request_wait(struct mddev *mddev, struct
r10conf *conf,
                                   struct bio *bio, sector_t sectors)
   {
-       wait_barrier(conf);
+       /* Bail out if REQ_NOWAIT is set for the bio */
+       if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }
I think we also need regular_request_wait to return bool and handle
it properly.

Thanks,
Song
Ack, will fix it. Thanks!
Ran into this while running with io_uring. With the current v5 (raid10
patch) on top of md-next branch.
./t/io_uring -a 0 -d 256 </dev/raid10>

It didn't trigger with aio (-a 1)

[  248.128661] BUG: kernel NULL pointer dereference, address:
00000000000000b8
[  248.135628] #PF: supervisor read access in kernel mode
[  248.140762] #PF: error_code(0x0000) - not-present page
[  248.145903] PGD 0 P4D 0
[  248.148443] Oops: 0000 [#1] PREEMPT SMP NOPTI
[  248.152800] CPU: 49 PID: 9461 Comm: io_uring Kdump: loaded Not
tainted 5.16.0-rc3+ #2
[  248.160629] Hardware name: Dell Inc. PowerEdge R650xs/0PPTY2, BIOS
1.3.8 08/31/2021
[  248.168279] RIP: 0010:raid10_end_read_request+0x74/0x140 [raid10]
[  248.174373] Code: 48 60 48 8b 58 58 48 c1 e2 05 49 03 55 08 48 89 4a
10 40 84 f6 75 48 f0 41 80 4c 24 18 01 4c 89 e7 e8 e0 b8 ff ff 49 8b 4d
00 <48> 8b 83 b8 00 00 00 f0 ff 8b f0 00 00 00 0f 94 c2 a8 01 74 04 84
[  248.193120] RSP: 0018:ffffb1c38d598ce8 EFLAGS: 00010086
[  248.198344] RAX: ffff8e5da2a1a100 RBX: 0000000000000000 RCX:
ffff8e5d89747000
[  248.205479] RDX: 000000008040003a RSI: 0000000080400039 RDI:
ffff8e1e00044900
[  248.212611] RBP: ffffb1c38d598d30 R08: 0000000000000000 R09:
0000000000000001
[  248.219744] R10: ffff8e5da2a1ae00 R11: 000000411bab9000 R12:
ffff8e5da2a1ae00
[  248.226877] R13: ffff8e5d8973fc00 R14: 0000000000000000 R15:
0000000000001000
[  248.234009] FS:  00007fc26b07d700(0000) GS:ffff8e9c6e600000(0000)
knlGS:0000000000000000
[  248.242096] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  248.247843] CR2: 00000000000000b8 CR3: 00000040b25d4005 CR4:
0000000000770ee0
[  248.254973] DR0: 0000000000000000 DR1: 0000000000000000 DR2:
0000000000000000
[  248.262107] DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7:
0000000000000400
[  248.269240] PKRU: 55555554
[  248.271953] Call Trace:
[  248.274406]  <IRQ>
[  248.276425]  bio_endio+0xf6/0x170
[  248.279743]  blk_update_request+0x12d/0x470
[  248.283931]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.288809]  blk_mq_end_request_batch+0x76/0x490
[  248.293429]  ? dma_direct_unmap_sg+0xdd/0x1a0
[  248.297786]  ? smp_call_function_single_async+0x46/0x70
[  248.303015]  ? mempool_kfree+0xe/0x10
[  248.306680]  ? mempool_kfree+0xe/0x10
[  248.310345]  nvme_pci_complete_batch+0x26/0xb0
[  248.314792]  nvme_irq+0x298/0x2f0
[  248.318110]  ? nvme_unmap_data+0xf0/0xf0
[  248.322038]  __handle_irq_event_percpu+0x3f/0x190
[  248.326744]  handle_irq_event_percpu+0x33/0x80
[  248.331190]  handle_irq_event+0x39/0x60
[  248.335028]  handle_edge_irq+0xbe/0x1e0
[  248.338869]  __common_interrupt+0x6b/0x110
[  248.342967]  common_interrupt+0xbd/0xe0
[  248.346808]  </IRQ>
[  248.348912]  <TASK>
[  248.351018]  asm_common_interrupt+0x1e/0x40
[  248.355206] RIP: 0010:_raw_spin_unlock_irqrestore+0x1e/0x37
[  248.360780] Code: 02 5d c3 0f 1f 44 00 00 5d c3 66 90 0f 1f 44 00 00
55 48 89 e5 c6 07 00 0f 1f 40 00 f7 c6 00 02 00 00 74 01 fb bf 01 00 00
00 <e8> ed 8e 5b ff 65 8b 05 66 7e 52 78 85 c0 74 02 5d c3 0f 1f 44 00

[  248.379525] RSP: 0018:ffffb1c3a429b958 EFLAGS: 00000206
[  248.384749] RAX: 0000000000000001 RBX: ffff8e5d8973fd08 RCX:
ffff8e5d8973fd10
[  248.391884] RDX: 0000000000000001 RSI: 0000000000000246 RDI:
0000000000000001
[  248.399017] RBP: ffffb1c3a429b958 R08: 0000000000000000 R09:
ffffb1c3a429b970
[  248.406148] R10: 0000000000000c00 R11: 0000000000000001 R12:
0000000000000001
[  248.413280] R13: 0000000000000246 R14: 0000000000000000 R15:
0000000000000003
[  248.420415]  __wake_up_common_lock+0x8a/0xc0
[  248.424686]  __wake_up+0x13/0x20
[  248.427919]  raid10_make_request+0x101/0x170 [raid10]
[  248.432971]  md_handle_request+0x179/0x1e0
[  248.437071]  ? submit_bio_checks+0x1f6/0x5a0
[  248.441345]  md_submit_bio+0x6d/0xa0
[  248.444924]  __submit_bio+0x94/0x140
[  248.448504]  submit_bio_noacct+0xe1/0x2a0
[  248.452515]  submit_bio+0x48/0x120
[  248.455923]  blkdev_direct_IO+0x220/0x540
[  248.459935]  ? __fsnotify_parent+0xff/0x330
[  248.464121]  ? __fsnotify_parent+0x10f/0x330
[  248.468393]  ? common_interrupt+0x73/0xe0
[  248.472408]  generic_file_read_iter+0xa5/0x160
[  248.476852]  blkdev_read_iter+0x38/0x70
[  248.480693]  io_read+0x119/0x420
[  248.483923]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.488805]  ? blk_mq_end_request_batch+0x378/0x490
[  248.493684]  io_issue_sqe+0x7ec/0x19c0
[  248.497436]  ? io_req_prep+0x6a9/0xe60
[  248.501190]  io_submit_sqes+0x2a0/0x9f0
[  248.505030]  ? __fget_files+0x6a/0x90
[  248.508693]  __x64_sys_io_uring_enter+0x1da/0x8c0
[  248.513401]  do_syscall_64+0x38/0x90
[  248.516979]  entry_SYSCALL_64_after_hwframe+0x44/0xae
[  248.522033] RIP: 0033:0x7fc26b19b89d
[  248.525611] Code: 00 c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa
48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f
05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d c3 f5 0c 00 f7 d8 64 89 01 48
[  248.544360] RSP: 002b:00007fc26b07ce98 EFLAGS: 00000246 ORIG_RAX:
00000000000001aa
[  248.551925] RAX: ffffffffffffffda RBX: 00007fc26b3f2fc0 RCX:
00007fc26b19b89d
[  248.559058] RDX: 0000000000000020 RSI: 0000000000000020 RDI:
0000000000000004
[  248.566189] RBP: 0000000000000020 R08: 0000000000000000 R09:
0000000000000000
[  248.573322] R10: 0000000000000001 R11: 0000000000000246 R12:
00005623a4b7a2a0
[  248.580456] R13: 0000000000000020 R14: 0000000000000020 R15:
0000000000000020
[  248.587591]  </TASK>
Do you have:

commit 75feae73a28020e492fbad2323245455ef69d687
Author: Pavel Begunkov [off-list ref]
Date:   Tue Dec 7 20:16:36 2021 +0000

     block: fix single bio async DIO error handling

in your tree?
Hmm, it is still triggering..

Re: [PATCH v5 3/4] md: raid10 add nowait support

From: Jens Axboe <axboe@kernel.dk>
Date: 2021-12-16 18:49:09

On 12/16/21 9:45 AM, Vishal Verma wrote:
On 12/16/21 9:42 AM, Jens Axboe wrote:
quoted
On 12/15/21 5:30 PM, Vishal Verma wrote:
quoted
On 12/15/21 3:20 PM, Vishal Verma wrote:
quoted
On 12/15/21 1:42 PM, Song Liu wrote:
quoted
On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma
[off-list ref] wrote:
quoted
This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

- Waiting for the barrier,
- Too many pending I/Os to be queued,
- Reshape operation,
- Discard operation.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
   drivers/md/raid10.c | 57
+++++++++++++++++++++++++++++++++++----------
   1 file changed, 45 insertions(+), 12 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..f6c73987e9ac 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,11 +952,18 @@ static void lower_barrier(struct r10conf *conf)
          wake_up(&conf->wait_barrier);
   }

-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
   {
          spin_lock_irq(&conf->resync_lock);
          if (conf->barrier) {
                  struct bio_list *bio_list = current->bio_list;
+
+               /* Return false when nowait flag is set */
+               if (nowait) {
+ spin_unlock_irq(&conf->resync_lock);
+                       return false;
+               }
+
                  conf->nr_waiting++;
                  /* Wait for the barrier to drop.
                   * However if there are already pending
@@ -988,6 +995,7 @@ static void wait_barrier(struct r10conf *conf)
          }
          atomic_inc(&conf->nr_pending);
          spin_unlock_irq(&conf->resync_lock);
+       return true;
   }

   static void allow_barrier(struct r10conf *conf)
@@ -1101,17 +1109,25 @@ static void raid10_unplug(struct blk_plug_cb
*cb, bool from_schedule)
   static void regular_request_wait(struct mddev *mddev, struct
r10conf *conf,
                                   struct bio *bio, sector_t sectors)
   {
-       wait_barrier(conf);
+       /* Bail out if REQ_NOWAIT is set for the bio */
+       if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }
I think we also need regular_request_wait to return bool and handle
it properly.

Thanks,
Song
Ack, will fix it. Thanks!
Ran into this while running with io_uring. With the current v5 (raid10
patch) on top of md-next branch.
./t/io_uring -a 0 -d 256 </dev/raid10>

It didn't trigger with aio (-a 1)

[  248.128661] BUG: kernel NULL pointer dereference, address:
00000000000000b8
[  248.135628] #PF: supervisor read access in kernel mode
[  248.140762] #PF: error_code(0x0000) - not-present page
[  248.145903] PGD 0 P4D 0
[  248.148443] Oops: 0000 [#1] PREEMPT SMP NOPTI
[  248.152800] CPU: 49 PID: 9461 Comm: io_uring Kdump: loaded Not
tainted 5.16.0-rc3+ #2
[  248.160629] Hardware name: Dell Inc. PowerEdge R650xs/0PPTY2, BIOS
1.3.8 08/31/2021
[  248.168279] RIP: 0010:raid10_end_read_request+0x74/0x140 [raid10]
[  248.174373] Code: 48 60 48 8b 58 58 48 c1 e2 05 49 03 55 08 48 89 4a
10 40 84 f6 75 48 f0 41 80 4c 24 18 01 4c 89 e7 e8 e0 b8 ff ff 49 8b 4d
00 <48> 8b 83 b8 00 00 00 f0 ff 8b f0 00 00 00 0f 94 c2 a8 01 74 04 84
[  248.193120] RSP: 0018:ffffb1c38d598ce8 EFLAGS: 00010086
[  248.198344] RAX: ffff8e5da2a1a100 RBX: 0000000000000000 RCX:
ffff8e5d89747000
[  248.205479] RDX: 000000008040003a RSI: 0000000080400039 RDI:
ffff8e1e00044900
[  248.212611] RBP: ffffb1c38d598d30 R08: 0000000000000000 R09:
0000000000000001
[  248.219744] R10: ffff8e5da2a1ae00 R11: 000000411bab9000 R12:
ffff8e5da2a1ae00
[  248.226877] R13: ffff8e5d8973fc00 R14: 0000000000000000 R15:
0000000000001000
[  248.234009] FS:  00007fc26b07d700(0000) GS:ffff8e9c6e600000(0000)
knlGS:0000000000000000
[  248.242096] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  248.247843] CR2: 00000000000000b8 CR3: 00000040b25d4005 CR4:
0000000000770ee0
[  248.254973] DR0: 0000000000000000 DR1: 0000000000000000 DR2:
0000000000000000
[  248.262107] DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7:
0000000000000400
[  248.269240] PKRU: 55555554
[  248.271953] Call Trace:
[  248.274406]  <IRQ>
[  248.276425]  bio_endio+0xf6/0x170
[  248.279743]  blk_update_request+0x12d/0x470
[  248.283931]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.288809]  blk_mq_end_request_batch+0x76/0x490
[  248.293429]  ? dma_direct_unmap_sg+0xdd/0x1a0
[  248.297786]  ? smp_call_function_single_async+0x46/0x70
[  248.303015]  ? mempool_kfree+0xe/0x10
[  248.306680]  ? mempool_kfree+0xe/0x10
[  248.310345]  nvme_pci_complete_batch+0x26/0xb0
[  248.314792]  nvme_irq+0x298/0x2f0
[  248.318110]  ? nvme_unmap_data+0xf0/0xf0
[  248.322038]  __handle_irq_event_percpu+0x3f/0x190
[  248.326744]  handle_irq_event_percpu+0x33/0x80
[  248.331190]  handle_irq_event+0x39/0x60
[  248.335028]  handle_edge_irq+0xbe/0x1e0
[  248.338869]  __common_interrupt+0x6b/0x110
[  248.342967]  common_interrupt+0xbd/0xe0
[  248.346808]  </IRQ>
[  248.348912]  <TASK>
[  248.351018]  asm_common_interrupt+0x1e/0x40
[  248.355206] RIP: 0010:_raw_spin_unlock_irqrestore+0x1e/0x37
[  248.360780] Code: 02 5d c3 0f 1f 44 00 00 5d c3 66 90 0f 1f 44 00 00
55 48 89 e5 c6 07 00 0f 1f 40 00 f7 c6 00 02 00 00 74 01 fb bf 01 00 00
00 <e8> ed 8e 5b ff 65 8b 05 66 7e 52 78 85 c0 74 02 5d c3 0f 1f 44 00

[  248.379525] RSP: 0018:ffffb1c3a429b958 EFLAGS: 00000206
[  248.384749] RAX: 0000000000000001 RBX: ffff8e5d8973fd08 RCX:
ffff8e5d8973fd10
[  248.391884] RDX: 0000000000000001 RSI: 0000000000000246 RDI:
0000000000000001
[  248.399017] RBP: ffffb1c3a429b958 R08: 0000000000000000 R09:
ffffb1c3a429b970
[  248.406148] R10: 0000000000000c00 R11: 0000000000000001 R12:
0000000000000001
[  248.413280] R13: 0000000000000246 R14: 0000000000000000 R15:
0000000000000003
[  248.420415]  __wake_up_common_lock+0x8a/0xc0
[  248.424686]  __wake_up+0x13/0x20
[  248.427919]  raid10_make_request+0x101/0x170 [raid10]
[  248.432971]  md_handle_request+0x179/0x1e0
[  248.437071]  ? submit_bio_checks+0x1f6/0x5a0
[  248.441345]  md_submit_bio+0x6d/0xa0
[  248.444924]  __submit_bio+0x94/0x140
[  248.448504]  submit_bio_noacct+0xe1/0x2a0
[  248.452515]  submit_bio+0x48/0x120
[  248.455923]  blkdev_direct_IO+0x220/0x540
[  248.459935]  ? __fsnotify_parent+0xff/0x330
[  248.464121]  ? __fsnotify_parent+0x10f/0x330
[  248.468393]  ? common_interrupt+0x73/0xe0
[  248.472408]  generic_file_read_iter+0xa5/0x160
[  248.476852]  blkdev_read_iter+0x38/0x70
[  248.480693]  io_read+0x119/0x420
[  248.483923]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.488805]  ? blk_mq_end_request_batch+0x378/0x490
[  248.493684]  io_issue_sqe+0x7ec/0x19c0
[  248.497436]  ? io_req_prep+0x6a9/0xe60
[  248.501190]  io_submit_sqes+0x2a0/0x9f0
[  248.505030]  ? __fget_files+0x6a/0x90
[  248.508693]  __x64_sys_io_uring_enter+0x1da/0x8c0
[  248.513401]  do_syscall_64+0x38/0x90
[  248.516979]  entry_SYSCALL_64_after_hwframe+0x44/0xae
[  248.522033] RIP: 0033:0x7fc26b19b89d
[  248.525611] Code: 00 c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa
48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f
05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d c3 f5 0c 00 f7 d8 64 89 01 48
[  248.544360] RSP: 002b:00007fc26b07ce98 EFLAGS: 00000246 ORIG_RAX:
00000000000001aa
[  248.551925] RAX: ffffffffffffffda RBX: 00007fc26b3f2fc0 RCX:
00007fc26b19b89d
[  248.559058] RDX: 0000000000000020 RSI: 0000000000000020 RDI:
0000000000000004
[  248.566189] RBP: 0000000000000020 R08: 0000000000000000 R09:
0000000000000000
[  248.573322] R10: 0000000000000001 R11: 0000000000000246 R12:
00005623a4b7a2a0
[  248.580456] R13: 0000000000000020 R14: 0000000000000020 R15:
0000000000000020
[  248.587591]  </TASK>
Do you have:

commit 75feae73a28020e492fbad2323245455ef69d687
Author: Pavel Begunkov [off-list ref]
Date:   Tue Dec 7 20:16:36 2021 +0000

     block: fix single bio async DIO error handling

in your tree?
Nope. I will get it in and test. Thanks!
Might be worth re-running with KASAN enabled in your config to see if
that triggers anything.

-- 
Jens Axboe

Re: [PATCH v5 3/4] md: raid10 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-16 19:40:33

On 12/16/21 11:49 AM, Jens Axboe wrote:
On 12/16/21 9:45 AM, Vishal Verma wrote:
quoted
On 12/16/21 9:42 AM, Jens Axboe wrote:
quoted
On 12/15/21 5:30 PM, Vishal Verma wrote:
quoted
On 12/15/21 3:20 PM, Vishal Verma wrote:
quoted
On 12/15/21 1:42 PM, Song Liu wrote:
quoted
On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma
[off-list ref] wrote:
quoted
This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

- Waiting for the barrier,
- Too many pending I/Os to be queued,
- Reshape operation,
- Discard operation.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
    drivers/md/raid10.c | 57
+++++++++++++++++++++++++++++++++++----------
    1 file changed, 45 insertions(+), 12 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..f6c73987e9ac 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,11 +952,18 @@ static void lower_barrier(struct r10conf *conf)
           wake_up(&conf->wait_barrier);
    }

-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
    {
           spin_lock_irq(&conf->resync_lock);
           if (conf->barrier) {
                   struct bio_list *bio_list = current->bio_list;
+
+               /* Return false when nowait flag is set */
+               if (nowait) {
+ spin_unlock_irq(&conf->resync_lock);
+                       return false;
+               }
+
                   conf->nr_waiting++;
                   /* Wait for the barrier to drop.
                    * However if there are already pending
@@ -988,6 +995,7 @@ static void wait_barrier(struct r10conf *conf)
           }
           atomic_inc(&conf->nr_pending);
           spin_unlock_irq(&conf->resync_lock);
+       return true;
    }

    static void allow_barrier(struct r10conf *conf)
@@ -1101,17 +1109,25 @@ static void raid10_unplug(struct blk_plug_cb
*cb, bool from_schedule)
    static void regular_request_wait(struct mddev *mddev, struct
r10conf *conf,
                                    struct bio *bio, sector_t sectors)
    {
-       wait_barrier(conf);
+       /* Bail out if REQ_NOWAIT is set for the bio */
+       if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }
I think we also need regular_request_wait to return bool and handle
it properly.

Thanks,
Song
Ack, will fix it. Thanks!
Ran into this while running with io_uring. With the current v5 (raid10
patch) on top of md-next branch.
./t/io_uring -a 0 -d 256 </dev/raid10>

It didn't trigger with aio (-a 1)

[  248.128661] BUG: kernel NULL pointer dereference, address:
00000000000000b8
[  248.135628] #PF: supervisor read access in kernel mode
[  248.140762] #PF: error_code(0x0000) - not-present page
[  248.145903] PGD 0 P4D 0
[  248.148443] Oops: 0000 [#1] PREEMPT SMP NOPTI
[  248.152800] CPU: 49 PID: 9461 Comm: io_uring Kdump: loaded Not
tainted 5.16.0-rc3+ #2
[  248.160629] Hardware name: Dell Inc. PowerEdge R650xs/0PPTY2, BIOS
1.3.8 08/31/2021
[  248.168279] RIP: 0010:raid10_end_read_request+0x74/0x140 [raid10]
[  248.174373] Code: 48 60 48 8b 58 58 48 c1 e2 05 49 03 55 08 48 89 4a
10 40 84 f6 75 48 f0 41 80 4c 24 18 01 4c 89 e7 e8 e0 b8 ff ff 49 8b 4d
00 <48> 8b 83 b8 00 00 00 f0 ff 8b f0 00 00 00 0f 94 c2 a8 01 74 04 84
[  248.193120] RSP: 0018:ffffb1c38d598ce8 EFLAGS: 00010086
[  248.198344] RAX: ffff8e5da2a1a100 RBX: 0000000000000000 RCX:
ffff8e5d89747000
[  248.205479] RDX: 000000008040003a RSI: 0000000080400039 RDI:
ffff8e1e00044900
[  248.212611] RBP: ffffb1c38d598d30 R08: 0000000000000000 R09:
0000000000000001
[  248.219744] R10: ffff8e5da2a1ae00 R11: 000000411bab9000 R12:
ffff8e5da2a1ae00
[  248.226877] R13: ffff8e5d8973fc00 R14: 0000000000000000 R15:
0000000000001000
[  248.234009] FS:  00007fc26b07d700(0000) GS:ffff8e9c6e600000(0000)
knlGS:0000000000000000
[  248.242096] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  248.247843] CR2: 00000000000000b8 CR3: 00000040b25d4005 CR4:
0000000000770ee0
[  248.254973] DR0: 0000000000000000 DR1: 0000000000000000 DR2:
0000000000000000
[  248.262107] DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7:
0000000000000400
[  248.269240] PKRU: 55555554
[  248.271953] Call Trace:
[  248.274406]  <IRQ>
[  248.276425]  bio_endio+0xf6/0x170
[  248.279743]  blk_update_request+0x12d/0x470
[  248.283931]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.288809]  blk_mq_end_request_batch+0x76/0x490
[  248.293429]  ? dma_direct_unmap_sg+0xdd/0x1a0
[  248.297786]  ? smp_call_function_single_async+0x46/0x70
[  248.303015]  ? mempool_kfree+0xe/0x10
[  248.306680]  ? mempool_kfree+0xe/0x10
[  248.310345]  nvme_pci_complete_batch+0x26/0xb0
[  248.314792]  nvme_irq+0x298/0x2f0
[  248.318110]  ? nvme_unmap_data+0xf0/0xf0
[  248.322038]  __handle_irq_event_percpu+0x3f/0x190
[  248.326744]  handle_irq_event_percpu+0x33/0x80
[  248.331190]  handle_irq_event+0x39/0x60
[  248.335028]  handle_edge_irq+0xbe/0x1e0
[  248.338869]  __common_interrupt+0x6b/0x110
[  248.342967]  common_interrupt+0xbd/0xe0
[  248.346808]  </IRQ>
[  248.348912]  <TASK>
[  248.351018]  asm_common_interrupt+0x1e/0x40
[  248.355206] RIP: 0010:_raw_spin_unlock_irqrestore+0x1e/0x37
[  248.360780] Code: 02 5d c3 0f 1f 44 00 00 5d c3 66 90 0f 1f 44 00 00
55 48 89 e5 c6 07 00 0f 1f 40 00 f7 c6 00 02 00 00 74 01 fb bf 01 00 00
00 <e8> ed 8e 5b ff 65 8b 05 66 7e 52 78 85 c0 74 02 5d c3 0f 1f 44 00

[  248.379525] RSP: 0018:ffffb1c3a429b958 EFLAGS: 00000206
[  248.384749] RAX: 0000000000000001 RBX: ffff8e5d8973fd08 RCX:
ffff8e5d8973fd10
[  248.391884] RDX: 0000000000000001 RSI: 0000000000000246 RDI:
0000000000000001
[  248.399017] RBP: ffffb1c3a429b958 R08: 0000000000000000 R09:
ffffb1c3a429b970
[  248.406148] R10: 0000000000000c00 R11: 0000000000000001 R12:
0000000000000001
[  248.413280] R13: 0000000000000246 R14: 0000000000000000 R15:
0000000000000003
[  248.420415]  __wake_up_common_lock+0x8a/0xc0
[  248.424686]  __wake_up+0x13/0x20
[  248.427919]  raid10_make_request+0x101/0x170 [raid10]
[  248.432971]  md_handle_request+0x179/0x1e0
[  248.437071]  ? submit_bio_checks+0x1f6/0x5a0
[  248.441345]  md_submit_bio+0x6d/0xa0
[  248.444924]  __submit_bio+0x94/0x140
[  248.448504]  submit_bio_noacct+0xe1/0x2a0
[  248.452515]  submit_bio+0x48/0x120
[  248.455923]  blkdev_direct_IO+0x220/0x540
[  248.459935]  ? __fsnotify_parent+0xff/0x330
[  248.464121]  ? __fsnotify_parent+0x10f/0x330
[  248.468393]  ? common_interrupt+0x73/0xe0
[  248.472408]  generic_file_read_iter+0xa5/0x160
[  248.476852]  blkdev_read_iter+0x38/0x70
[  248.480693]  io_read+0x119/0x420
[  248.483923]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.488805]  ? blk_mq_end_request_batch+0x378/0x490
[  248.493684]  io_issue_sqe+0x7ec/0x19c0
[  248.497436]  ? io_req_prep+0x6a9/0xe60
[  248.501190]  io_submit_sqes+0x2a0/0x9f0
[  248.505030]  ? __fget_files+0x6a/0x90
[  248.508693]  __x64_sys_io_uring_enter+0x1da/0x8c0
[  248.513401]  do_syscall_64+0x38/0x90
[  248.516979]  entry_SYSCALL_64_after_hwframe+0x44/0xae
[  248.522033] RIP: 0033:0x7fc26b19b89d
[  248.525611] Code: 00 c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa
48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f
05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d c3 f5 0c 00 f7 d8 64 89 01 48
[  248.544360] RSP: 002b:00007fc26b07ce98 EFLAGS: 00000246 ORIG_RAX:
00000000000001aa
[  248.551925] RAX: ffffffffffffffda RBX: 00007fc26b3f2fc0 RCX:
00007fc26b19b89d
[  248.559058] RDX: 0000000000000020 RSI: 0000000000000020 RDI:
0000000000000004
[  248.566189] RBP: 0000000000000020 R08: 0000000000000000 R09:
0000000000000000
[  248.573322] R10: 0000000000000001 R11: 0000000000000246 R12:
00005623a4b7a2a0
[  248.580456] R13: 0000000000000020 R14: 0000000000000020 R15:
0000000000000020
[  248.587591]  </TASK>
Do you have:

commit 75feae73a28020e492fbad2323245455ef69d687
Author: Pavel Begunkov [off-list ref]
Date:   Tue Dec 7 20:16:36 2021 +0000

      block: fix single bio async DIO error handling

in your tree?
Nope. I will get it in and test. Thanks!
Might be worth re-running with KASAN enabled in your config to see if
that triggers anything.
Got this:
[  739.336669] CPU: 63 PID: 10373 Comm: io_uring Kdump: loaded Not 
tainted 5.16.0-rc3+ #8
[  739.344583] Hardware name: Dell Inc. PowerEdge R650xs/0PPTY2, BIOS 
1.3.8 08/31/2021
[  739.352236] Call Trace:
[  739.354687]  <IRQ>
[  739.356705]  dump_stack_lvl+0x38/0x49
[  739.360381]  print_address_description.constprop.0+0x28/0x150
[  739.366136]  ? 0xffffffffc046f3df
[  739.369455]  kasan_report.cold+0x82/0xdb
[  739.373383]  ? 0xffffffffc046f3df
[  739.376700]  __asan_load8+0x69/0x90
[  739.380194]  0xffffffffc046f3df
[  739.383339]  ? 0xffffffffc046f340
[  739.386659]  ? blkcg_iolatency_done_bio+0x26/0x390
[  739.391461]  ? __rcu_read_unlock+0x5b/0x270
[  739.395655]  ? kmem_cache_alloc+0x143/0x460
[  739.399841]  ? mempool_alloc_slab+0x17/0x20
[  739.404027]  ? bio_uninit+0x6c/0xf0
[  739.407522]  bio_endio+0x27f/0x2a0
[  739.410926]  blk_update_request+0x1e8/0x750
[  739.415112]  blk_mq_end_request_batch+0x10b/0x9b0
[  739.419818]  ? blk_mq_end_request+0x460/0x460
[  739.424179]  ? kfree+0xa0/0x400
[  739.427322]  ? mempool_kfree+0xe/0x10
[  739.430989]  ? generic_file_write_iter+0xf0/0xf0
[  739.435609]  ? dma_unmap_page_attrs+0x15f/0x2c0
[  739.440144]  nvme_pci_complete_batch+0x34/0x160
[  739.444684]  ? blk_mq_complete_request_remote+0x1ca/0x2d0
[  739.450084]  nvme_irq+0x5fa/0x630
[  739.453404]  ? nvme_timeout+0x370/0x370
[  739.457242]  ? nvme_unmap_data+0x1e0/0x1e0
[  739.461340]  ? __kasan_check_write+0x14/0x20
[  739.465614]  ? credit_entropy_bits.constprop.0+0x76/0x190
[  739.471015]  ? nvme_timeout+0x370/0x370
[  739.474853]  __handle_irq_event_percpu+0x69/0x260
[  739.479568]  handle_irq_event_percpu+0x70/0xf0
[  739.484015]  ? __handle_irq_event_percpu+0x260/0x260
[  739.488981]  ? __kasan_check_write+0x14/0x20
[  739.493261]  ? _raw_spin_lock+0x88/0xe0
[  739.497109]  ? _raw_spin_lock_irqsave+0xf0/0xf0
[  739.501643]  handle_irq_event+0x5a/0x90
[  739.505480]  handle_edge_irq+0x148/0x320
[  739.509407]  __common_interrupt+0x75/0x130
[  739.513514]  common_interrupt+0xae/0xd0
[  739.517354]  </IRQ>
[  739.519460]  <TASK>
[  739.521567]  asm_common_interrupt+0x1e/0x40
[  739.525753] RIP: 0010:__asan_store8+0x37/0x90
[  739.530111] Code: 4f 48 b8 ff ff ff ff ff 7f ff ff 48 39 c7 76 40 48 
8d 47 07 48 89 c2 83 e2 07 48 83 fa 07 75 18 48 ba 00 00 00 00 00 fc ff 
df <48> c1 e8 03 0f b6 04 10 84 c0 75 2b 5d c3 48 be 00 00 00 00 00 fc
[  739.548865] RSP: 0018:ffffc900307c7100 EFLAGS: 00000246
[  739.554092] RAX: ffffc900307c7237 RBX: ffffc900307c7780 RCX: 
ffffffff82c9b461
[  739.561227] RDX: dffffc0000000000 RSI: ffffc900307c7790 RDI: 
ffffc900307c7230
[  739.568358] RBP: ffffc900307c7100 R08: ffffffff82c9b202 R09: 
ffff8882d96c0000
[  739.575490] R10: ffffc900307c7257 R11: fffff520060f8e4a R12: 
0000000082c9b201
[  739.582624] R13: 0000000000000000 R14: ffffc900307c7238 R15: 
ffffc900307c71e8
[  739.589760]  ? update_stack_state+0x22/0x2c0
[  739.594038]  ? update_stack_state+0x281/0x2c0
[  739.598398]  update_stack_state+0x281/0x2c0
[  739.602585]  unwind_next_frame.part.0+0xe0/0x360
[  739.607204]  ? bio_alloc_bioset+0x223/0x2f0
[  739.611389]  ? create_prof_cpu_mask+0x30/0x30
[  739.615749]  ? mempool_alloc_slab+0x17/0x20
[  739.619935]  unwind_next_frame+0x23/0x30
[  739.623860]  arch_stack_walk+0x88/0xf0
[  739.627613]  ? bio_alloc_bioset+0x223/0x2f0
[  739.631800]  stack_trace_save+0x94/0xc0
[  739.635640]  ? filter_irq_stacks+0x70/0x70
[  739.639738]  ? blk_mq_put_tag+0x80/0x80
[  739.643576]  ? _raw_spin_unlock_irqrestore+0x23/0x40
[  739.648545]  ? __wake_up_common_lock+0xfd/0x150
[  739.653084]  kasan_save_stack+0x26/0x60
[  739.656924]  ? kasan_save_stack+0x26/0x60
[  739.660938]  ? __kasan_slab_alloc+0x6d/0x90
[  739.665123]  ? kmem_cache_alloc+0x143/0x460
[  739.669308]  ? mempool_alloc_slab+0x17/0x20
[  739.673494]  ? mempool_alloc+0xef/0x280
[  739.677333]  ? bio_alloc_bioset+0x223/0x2f0
[  739.681519]  ? blk_mq_rq_ctx_init.isra.0+0x28a/0x3c0
[  739.686489]  ? __blk_mq_alloc_requests+0x655/0x680
[  739.691288]  ? blkcg_iolatency_throttle+0x5d/0x760
[  739.696081]  ? bio_to_wbt_flags+0x47/0xf0
[  739.700093]  ? update_io_ticks+0x5e/0xd0
[  739.704021]  ? preempt_count_sub+0x18/0xc0
[  739.708120]  ? __kasan_check_read+0x11/0x20
[  739.712305]  ? blk_mq_submit_bio+0x740/0xce0
[  739.716580]  ? blk_mq_try_issue_list_directly+0x1b0/0x1b0
[  739.721987]  ? kasan_poison+0x3c/0x50
[  739.725650]  ? kasan_unpoison+0x28/0x50
[  739.729491]  __kasan_slab_alloc+0x6d/0x90
[  739.733503]  kmem_cache_alloc+0x143/0x460
[  739.737518]  mempool_alloc_slab+0x17/0x20
[  739.741529]  mempool_alloc+0xef/0x280
[  739.745194]  ? mempool_free+0x170/0x170
[  739.749035]  ? mempool_destroy+0x30/0x30
[  739.752961]  ? __fsnotify_update_child_dentry_flags.part.0+0x170/0x170
[  739.759498]  bio_alloc_bioset+0x223/0x2f0
[  739.763517]  ? __this_cpu_preempt_check+0x13/0x20
[  739.768223]  ? bvec_alloc+0xd0/0xd0
[  739.771715]  ? __fsnotify_parent+0x1ed/0x590
[  739.775988]  ? do_direct_IO+0x150/0x1880
[  739.779916]  ? submit_bio+0xb0/0x220
[  739.783496]  bio_alloc_kiocb+0x185/0x1c0
[  739.787430]  blkdev_direct_IO+0x114/0x400
[  739.791441]  generic_file_read_iter+0x152/0x250
[  739.795974]  blkdev_read_iter+0x84/0xd0
[  739.799815]  io_read+0x1ec/0x770
[  739.803056]  ? __rcu_read_unlock+0x5b/0x270
[  739.807240]  ? io_setup_async_rw+0x270/0x270
[  739.811515]  ? __sbq_wake_up+0x2d/0x1b0
[  739.815352]  ? __rcu_read_unlock+0x5b/0x270
[  739.819537]  ? sbitmap_queue_clear+0xc9/0xe0
[  739.823813]  ? blk_queue_exit+0x35/0x90
[  739.827653]  ? __blk_mq_free_request+0x111/0x160
[  739.832280]  io_issue_sqe+0xcac/0x27f0
[  739.836031]  ? blk_mq_free_plug_rqs+0x3f/0x50
[  739.840394]  ? io_poll_add.isra.0+0x290/0x290
[  739.844760]  ? io_req_prep+0xcc2/0x1bb0
[  739.848598]  ? io_submit_sqes+0x43b/0x1260
[  739.852701]  io_submit_sqes+0x5e5/0x1260
[  739.856634]  ? io_do_iopoll+0x561/0x720
[  739.860474]  ? io_wq_submit_work+0x230/0x230
[  739.864746]  ? __kasan_check_write+0x14/0x20
[  739.869016]  ? mutex_lock+0x8f/0xe0
[  739.872510]  ? __mutex_lock_slowpath+0x20/0x20
[  739.876956]  ? __rcu_read_unlock+0x5b/0x270
[  739.881144]  __x64_sys_io_uring_enter+0x367/0xef0
[  739.885859]  ? io_submit_sqes+0x1260/0x1260
[  739.890044]  ? __this_cpu_preempt_check+0x13/0x20
[  739.894749]  ? xfd_validate_state+0x3c/0xd0
[  739.898936]  ? __schedule+0x5be/0x10c0
[  739.902687]  ? restore_fpregs_from_fpstate+0xa2/0x170
[  739.907741]  ? kernel_fpu_begin_mask+0x170/0x170
[  739.912362]  ? debug_smp_processor_id+0x17/0x20
[  739.916903]  ? debug_smp_processor_id+0x17/0x20
[  739.921434]  ? fpregs_assert_state_consistent+0x5f/0x70
[  739.926662]  ? exit_to_user_mode_prepare+0x4b/0x1e0
[  739.931549]  do_syscall_64+0x38/0x90
[  739.935129]  entry_SYSCALL_64_after_hwframe+0x44/0xae
[  739.940182] RIP: 0033:0x7f7345d7d89d
[  739.943759] Code: 00 c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa 
48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 
05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d c3 f5 0c 00 f7 d8 64 89 01 48
[  739.962513] RSP: 002b:00007f7345c5ee98 EFLAGS: 00000246 ORIG_RAX: 
00000000000001aa
[  739.970081] RAX: ffffffffffffffda RBX: 00007f7345fd2fc0 RCX: 
00007f7345d7d89d
[  739.977216] RDX: 0000000000000000 RSI: 0000000000000020 RDI: 
0000000000000004
[  739.984349] RBP: 0000000000000020 R08: 0000000000000000 R09: 
0000000000000000
[  739.991481] R10: 0000000000000000 R11: 0000000000000246 R12: 
000055be1d11e2a0
[  739.998612] R13: 0000000000000020 R14: 0000000000000000 R15: 
0000000000000020
[  740.005749]  </TASK>
[  740.007945]
[  740.009444] Allocated by task 10373:
[  740.013084]
[  740.014583] Freed by task 10373:
[  740.017879]
[  740.019378] The buggy address belongs to the object at ffff88c1be016e00
[  740.019378]  which belongs to the cache kmalloc-256 of size 256
[  740.031886] The buggy address is located 40 bytes inside of
[  740.031886]  256-byte region [ffff88c1be016e00, ffff88c1be016f00)
[  740.043534] The buggy address belongs to the page:
[  740.048345]
[  740.049840] Memory state around the buggy address:
[  740.054636]  ffff88c1be016d00: fc fc fc fc fc fc fc fc fc fc fc fc fc 
fc fc fc
[  740.061854]  ffff88c1be016d80: fc fc fc fc fc fc fc fc fc fc fc fc fc 
fc fc fc
[  740.069074] >ffff88c1be016e00: fa fb fb fb fb fb fb fb fb fb fb fb fb 
fb fb fb
[  740.076294]                                   ^
[  740.080827]  ffff88c1be016e80: fb fb fb fb fb fb fb fb fb fb fb fb fb 
fb fb fb
[  740.088045]  ffff88c1be016f00: fc fc fc fc fc fc fc fc fc fc fc fc fc 
fc fc fc
[  740.095265] 
==================================================================
[  740.102497] kernel BUG at mm/slub.c:379!
[  740.106431] invalid opcode: 0000 [#1] PREEMPT SMP KASAN NOPTI

Re: [PATCH v5 3/4] md: raid10 add nowait support

From: Song Liu <song@kernel.org>
Date: 2021-12-16 20:19:11

On Thu, Dec 16, 2021 at 11:40 AM Vishal Verma [off-list ref] wrote:

On 12/16/21 11:49 AM, Jens Axboe wrote:
quoted
On 12/16/21 9:45 AM, Vishal Verma wrote:
quoted
On 12/16/21 9:42 AM, Jens Axboe wrote:
quoted
On 12/15/21 5:30 PM, Vishal Verma wrote:
quoted
On 12/15/21 3:20 PM, Vishal Verma wrote:
quoted
On 12/15/21 1:42 PM, Song Liu wrote:
quoted
On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma
[off-list ref] wrote:
quoted
This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

- Waiting for the barrier,
- Too many pending I/Os to be queued,
- Reshape operation,
- Discard operation.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
    drivers/md/raid10.c | 57
+++++++++++++++++++++++++++++++++++----------
    1 file changed, 45 insertions(+), 12 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..f6c73987e9ac 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,11 +952,18 @@ static void lower_barrier(struct r10conf *conf)
           wake_up(&conf->wait_barrier);
    }

-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
    {
           spin_lock_irq(&conf->resync_lock);
           if (conf->barrier) {
                   struct bio_list *bio_list = current->bio_list;
+
+               /* Return false when nowait flag is set */
+               if (nowait) {
+ spin_unlock_irq(&conf->resync_lock);
+                       return false;
+               }
+
                   conf->nr_waiting++;
                   /* Wait for the barrier to drop.
                    * However if there are already pending
@@ -988,6 +995,7 @@ static void wait_barrier(struct r10conf *conf)
           }
           atomic_inc(&conf->nr_pending);
           spin_unlock_irq(&conf->resync_lock);
+       return true;
    }

    static void allow_barrier(struct r10conf *conf)
@@ -1101,17 +1109,25 @@ static void raid10_unplug(struct blk_plug_cb
*cb, bool from_schedule)
    static void regular_request_wait(struct mddev *mddev, struct
r10conf *conf,
                                    struct bio *bio, sector_t sectors)
    {
-       wait_barrier(conf);
+       /* Bail out if REQ_NOWAIT is set for the bio */
+       if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }
I think we also need regular_request_wait to return bool and handle
it properly.

Thanks,
Song
Ack, will fix it. Thanks!
Ran into this while running with io_uring. With the current v5 (raid10
patch) on top of md-next branch.
./t/io_uring -a 0 -d 256 </dev/raid10>

It didn't trigger with aio (-a 1)

[  248.128661] BUG: kernel NULL pointer dereference, address:
00000000000000b8
[  248.135628] #PF: supervisor read access in kernel mode
[  248.140762] #PF: error_code(0x0000) - not-present page
[  248.145903] PGD 0 P4D 0
[  248.148443] Oops: 0000 [#1] PREEMPT SMP NOPTI
[  248.152800] CPU: 49 PID: 9461 Comm: io_uring Kdump: loaded Not
tainted 5.16.0-rc3+ #2
[  248.160629] Hardware name: Dell Inc. PowerEdge R650xs/0PPTY2, BIOS
1.3.8 08/31/2021
[  248.168279] RIP: 0010:raid10_end_read_request+0x74/0x140 [raid10]
[  248.174373] Code: 48 60 48 8b 58 58 48 c1 e2 05 49 03 55 08 48 89 4a
10 40 84 f6 75 48 f0 41 80 4c 24 18 01 4c 89 e7 e8 e0 b8 ff ff 49 8b 4d
00 <48> 8b 83 b8 00 00 00 f0 ff 8b f0 00 00 00 0f 94 c2 a8 01 74 04 84
[  248.193120] RSP: 0018:ffffb1c38d598ce8 EFLAGS: 00010086
[  248.198344] RAX: ffff8e5da2a1a100 RBX: 0000000000000000 RCX:
ffff8e5d89747000
[  248.205479] RDX: 000000008040003a RSI: 0000000080400039 RDI:
ffff8e1e00044900
[  248.212611] RBP: ffffb1c38d598d30 R08: 0000000000000000 R09:
0000000000000001
[  248.219744] R10: ffff8e5da2a1ae00 R11: 000000411bab9000 R12:
ffff8e5da2a1ae00
[  248.226877] R13: ffff8e5d8973fc00 R14: 0000000000000000 R15:
0000000000001000
[  248.234009] FS:  00007fc26b07d700(0000) GS:ffff8e9c6e600000(0000)
knlGS:0000000000000000
[  248.242096] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  248.247843] CR2: 00000000000000b8 CR3: 00000040b25d4005 CR4:
0000000000770ee0
[  248.254973] DR0: 0000000000000000 DR1: 0000000000000000 DR2:
0000000000000000
[  248.262107] DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7:
0000000000000400
[  248.269240] PKRU: 55555554
[  248.271953] Call Trace:
[  248.274406]  <IRQ>
[  248.276425]  bio_endio+0xf6/0x170
[  248.279743]  blk_update_request+0x12d/0x470
[  248.283931]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.288809]  blk_mq_end_request_batch+0x76/0x490
[  248.293429]  ? dma_direct_unmap_sg+0xdd/0x1a0
[  248.297786]  ? smp_call_function_single_async+0x46/0x70
[  248.303015]  ? mempool_kfree+0xe/0x10
[  248.306680]  ? mempool_kfree+0xe/0x10
[  248.310345]  nvme_pci_complete_batch+0x26/0xb0
[  248.314792]  nvme_irq+0x298/0x2f0
[  248.318110]  ? nvme_unmap_data+0xf0/0xf0
[  248.322038]  __handle_irq_event_percpu+0x3f/0x190
[  248.326744]  handle_irq_event_percpu+0x33/0x80
[  248.331190]  handle_irq_event+0x39/0x60
[  248.335028]  handle_edge_irq+0xbe/0x1e0
[  248.338869]  __common_interrupt+0x6b/0x110
[  248.342967]  common_interrupt+0xbd/0xe0
[  248.346808]  </IRQ>
[  248.348912]  <TASK>
[  248.351018]  asm_common_interrupt+0x1e/0x40
[  248.355206] RIP: 0010:_raw_spin_unlock_irqrestore+0x1e/0x37
[  248.360780] Code: 02 5d c3 0f 1f 44 00 00 5d c3 66 90 0f 1f 44 00 00
55 48 89 e5 c6 07 00 0f 1f 40 00 f7 c6 00 02 00 00 74 01 fb bf 01 00 00
00 <e8> ed 8e 5b ff 65 8b 05 66 7e 52 78 85 c0 74 02 5d c3 0f 1f 44 00

[  248.379525] RSP: 0018:ffffb1c3a429b958 EFLAGS: 00000206
[  248.384749] RAX: 0000000000000001 RBX: ffff8e5d8973fd08 RCX:
ffff8e5d8973fd10
[  248.391884] RDX: 0000000000000001 RSI: 0000000000000246 RDI:
0000000000000001
[  248.399017] RBP: ffffb1c3a429b958 R08: 0000000000000000 R09:
ffffb1c3a429b970
[  248.406148] R10: 0000000000000c00 R11: 0000000000000001 R12:
0000000000000001
[  248.413280] R13: 0000000000000246 R14: 0000000000000000 R15:
0000000000000003
[  248.420415]  __wake_up_common_lock+0x8a/0xc0
[  248.424686]  __wake_up+0x13/0x20
[  248.427919]  raid10_make_request+0x101/0x170 [raid10]
[  248.432971]  md_handle_request+0x179/0x1e0
[  248.437071]  ? submit_bio_checks+0x1f6/0x5a0
[  248.441345]  md_submit_bio+0x6d/0xa0
[  248.444924]  __submit_bio+0x94/0x140
[  248.448504]  submit_bio_noacct+0xe1/0x2a0
[  248.452515]  submit_bio+0x48/0x120
[  248.455923]  blkdev_direct_IO+0x220/0x540
[  248.459935]  ? __fsnotify_parent+0xff/0x330
[  248.464121]  ? __fsnotify_parent+0x10f/0x330
[  248.468393]  ? common_interrupt+0x73/0xe0
[  248.472408]  generic_file_read_iter+0xa5/0x160
[  248.476852]  blkdev_read_iter+0x38/0x70
[  248.480693]  io_read+0x119/0x420
[  248.483923]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.488805]  ? blk_mq_end_request_batch+0x378/0x490
[  248.493684]  io_issue_sqe+0x7ec/0x19c0
[  248.497436]  ? io_req_prep+0x6a9/0xe60
[  248.501190]  io_submit_sqes+0x2a0/0x9f0
[  248.505030]  ? __fget_files+0x6a/0x90
[  248.508693]  __x64_sys_io_uring_enter+0x1da/0x8c0
[  248.513401]  do_syscall_64+0x38/0x90
[  248.516979]  entry_SYSCALL_64_after_hwframe+0x44/0xae
[  248.522033] RIP: 0033:0x7fc26b19b89d
[  248.525611] Code: 00 c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa
48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f
05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d c3 f5 0c 00 f7 d8 64 89 01 48
[  248.544360] RSP: 002b:00007fc26b07ce98 EFLAGS: 00000246 ORIG_RAX:
00000000000001aa
[  248.551925] RAX: ffffffffffffffda RBX: 00007fc26b3f2fc0 RCX:
00007fc26b19b89d
[  248.559058] RDX: 0000000000000020 RSI: 0000000000000020 RDI:
0000000000000004
[  248.566189] RBP: 0000000000000020 R08: 0000000000000000 R09:
0000000000000000
[  248.573322] R10: 0000000000000001 R11: 0000000000000246 R12:
00005623a4b7a2a0
[  248.580456] R13: 0000000000000020 R14: 0000000000000020 R15:
0000000000000020
[  248.587591]  </TASK>
Do you have:

commit 75feae73a28020e492fbad2323245455ef69d687
Author: Pavel Begunkov [off-list ref]
Date:   Tue Dec 7 20:16:36 2021 +0000

      block: fix single bio async DIO error handling

in your tree?
Nope. I will get it in and test. Thanks!
Might be worth re-running with KASAN enabled in your config to see if
that triggers anything.
Got this:
[  739.336669] CPU: 63 PID: 10373 Comm: io_uring Kdump: loaded Not
tainted 5.16.0-rc3+ #8
[  739.344583] Hardware name: Dell Inc. PowerEdge R650xs/0PPTY2, BIOS
1.3.8 08/31/2021
[  739.352236] Call Trace:
[  739.354687]  <IRQ>
[  739.356705]  dump_stack_lvl+0x38/0x49
[  739.360381]  print_address_description.constprop.0+0x28/0x150
[  739.366136]  ? 0xffffffffc046f3df
[  739.369455]  kasan_report.cold+0x82/0xdb
[  739.373383]  ? 0xffffffffc046f3df
[  739.376700]  __asan_load8+0x69/0x90
[  739.380194]  0xffffffffc046f3df
[  739.383339]  ? 0xffffffffc046f340
[  739.386659]  ? blkcg_iolatency_done_bio+0x26/0x390
[  739.391461]  ? __rcu_read_unlock+0x5b/0x270
[  739.395655]  ? kmem_cache_alloc+0x143/0x460
[  739.399841]  ? mempool_alloc_slab+0x17/0x20
[  739.404027]  ? bio_uninit+0x6c/0xf0
[  739.407522]  bio_endio+0x27f/0x2a0
[  739.410926]  blk_update_request+0x1e8/0x750
[  739.415112]  blk_mq_end_request_batch+0x10b/0x9b0
[  739.419818]  ? blk_mq_end_request+0x460/0x460
[  739.424179]  ? kfree+0xa0/0x400
[  739.427322]  ? mempool_kfree+0xe/0x10
[  739.430989]  ? generic_file_write_iter+0xf0/0xf0
[  739.435609]  ? dma_unmap_page_attrs+0x15f/0x2c0
[  739.440144]  nvme_pci_complete_batch+0x34/0x160
[  739.444684]  ? blk_mq_complete_request_remote+0x1ca/0x2d0
[  739.450084]  nvme_irq+0x5fa/0x630
[  739.453404]  ? nvme_timeout+0x370/0x370
[  739.457242]  ? nvme_unmap_data+0x1e0/0x1e0
[  739.461340]  ? __kasan_check_write+0x14/0x20
[  739.465614]  ? credit_entropy_bits.constprop.0+0x76/0x190
[  739.471015]  ? nvme_timeout+0x370/0x370
[  739.474853]  __handle_irq_event_percpu+0x69/0x260
[  739.479568]  handle_irq_event_percpu+0x70/0xf0
[  739.484015]  ? __handle_irq_event_percpu+0x260/0x260
[  739.488981]  ? __kasan_check_write+0x14/0x20
[  739.493261]  ? _raw_spin_lock+0x88/0xe0
[  739.497109]  ? _raw_spin_lock_irqsave+0xf0/0xf0
[  739.501643]  handle_irq_event+0x5a/0x90
[  739.505480]  handle_edge_irq+0x148/0x320
[  739.509407]  __common_interrupt+0x75/0x130
[  739.513514]  common_interrupt+0xae/0xd0
[  739.517354]  </IRQ>
[  739.519460]  <TASK>
[  739.521567]  asm_common_interrupt+0x1e/0x40
[  739.525753] RIP: 0010:__asan_store8+0x37/0x90
[  739.530111] Code: 4f 48 b8 ff ff ff ff ff 7f ff ff 48 39 c7 76 40 48
8d 47 07 48 89 c2 83 e2 07 48 83 fa 07 75 18 48 ba 00 00 00 00 00 fc ff
df <48> c1 e8 03 0f b6 04 10 84 c0 75 2b 5d c3 48 be 00 00 00 00 00 fc
[  739.548865] RSP: 0018:ffffc900307c7100 EFLAGS: 00000246
[  739.554092] RAX: ffffc900307c7237 RBX: ffffc900307c7780 RCX:
ffffffff82c9b461
[  739.561227] RDX: dffffc0000000000 RSI: ffffc900307c7790 RDI:
ffffc900307c7230
[  739.568358] RBP: ffffc900307c7100 R08: ffffffff82c9b202 R09:
ffff8882d96c0000
[  739.575490] R10: ffffc900307c7257 R11: fffff520060f8e4a R12:
0000000082c9b201
[  739.582624] R13: 0000000000000000 R14: ffffc900307c7238 R15:
ffffc900307c71e8
[  739.589760]  ? update_stack_state+0x22/0x2c0
[  739.594038]  ? update_stack_state+0x281/0x2c0
[  739.598398]  update_stack_state+0x281/0x2c0
[  739.602585]  unwind_next_frame.part.0+0xe0/0x360
[  739.607204]  ? bio_alloc_bioset+0x223/0x2f0
[  739.611389]  ? create_prof_cpu_mask+0x30/0x30
[  739.615749]  ? mempool_alloc_slab+0x17/0x20
[  739.619935]  unwind_next_frame+0x23/0x30
[  739.623860]  arch_stack_walk+0x88/0xf0
[  739.627613]  ? bio_alloc_bioset+0x223/0x2f0
[  739.631800]  stack_trace_save+0x94/0xc0
[  739.635640]  ? filter_irq_stacks+0x70/0x70
[  739.639738]  ? blk_mq_put_tag+0x80/0x80
[  739.643576]  ? _raw_spin_unlock_irqrestore+0x23/0x40
[  739.648545]  ? __wake_up_common_lock+0xfd/0x150
[  739.653084]  kasan_save_stack+0x26/0x60
[  739.656924]  ? kasan_save_stack+0x26/0x60
[  739.660938]  ? __kasan_slab_alloc+0x6d/0x90
[  739.665123]  ? kmem_cache_alloc+0x143/0x460
[  739.669308]  ? mempool_alloc_slab+0x17/0x20
[  739.673494]  ? mempool_alloc+0xef/0x280
[  739.677333]  ? bio_alloc_bioset+0x223/0x2f0
[  739.681519]  ? blk_mq_rq_ctx_init.isra.0+0x28a/0x3c0
[  739.686489]  ? __blk_mq_alloc_requests+0x655/0x680
[  739.691288]  ? blkcg_iolatency_throttle+0x5d/0x760
[  739.696081]  ? bio_to_wbt_flags+0x47/0xf0
[  739.700093]  ? update_io_ticks+0x5e/0xd0
[  739.704021]  ? preempt_count_sub+0x18/0xc0
[  739.708120]  ? __kasan_check_read+0x11/0x20
[  739.712305]  ? blk_mq_submit_bio+0x740/0xce0
[  739.716580]  ? blk_mq_try_issue_list_directly+0x1b0/0x1b0
[  739.721987]  ? kasan_poison+0x3c/0x50
[  739.725650]  ? kasan_unpoison+0x28/0x50
[  739.729491]  __kasan_slab_alloc+0x6d/0x90
[  739.733503]  kmem_cache_alloc+0x143/0x460
[  739.737518]  mempool_alloc_slab+0x17/0x20
[  739.741529]  mempool_alloc+0xef/0x280
[  739.745194]  ? mempool_free+0x170/0x170
[  739.749035]  ? mempool_destroy+0x30/0x30
[  739.752961]  ? __fsnotify_update_child_dentry_flags.part.0+0x170/0x170
[  739.759498]  bio_alloc_bioset+0x223/0x2f0
[  739.763517]  ? __this_cpu_preempt_check+0x13/0x20
[  739.768223]  ? bvec_alloc+0xd0/0xd0
[  739.771715]  ? __fsnotify_parent+0x1ed/0x590
[  739.775988]  ? do_direct_IO+0x150/0x1880
[  739.779916]  ? submit_bio+0xb0/0x220
[  739.783496]  bio_alloc_kiocb+0x185/0x1c0
[  739.787430]  blkdev_direct_IO+0x114/0x400
[  739.791441]  generic_file_read_iter+0x152/0x250
[  739.795974]  blkdev_read_iter+0x84/0xd0
[  739.799815]  io_read+0x1ec/0x770
[  739.803056]  ? __rcu_read_unlock+0x5b/0x270
[  739.807240]  ? io_setup_async_rw+0x270/0x270
[  739.811515]  ? __sbq_wake_up+0x2d/0x1b0
[  739.815352]  ? __rcu_read_unlock+0x5b/0x270
[  739.819537]  ? sbitmap_queue_clear+0xc9/0xe0
[  739.823813]  ? blk_queue_exit+0x35/0x90
[  739.827653]  ? __blk_mq_free_request+0x111/0x160
[  739.832280]  io_issue_sqe+0xcac/0x27f0
[  739.836031]  ? blk_mq_free_plug_rqs+0x3f/0x50
[  739.840394]  ? io_poll_add.isra.0+0x290/0x290
[  739.844760]  ? io_req_prep+0xcc2/0x1bb0
[  739.848598]  ? io_submit_sqes+0x43b/0x1260
[  739.852701]  io_submit_sqes+0x5e5/0x1260
[  739.856634]  ? io_do_iopoll+0x561/0x720
[  739.860474]  ? io_wq_submit_work+0x230/0x230
[  739.864746]  ? __kasan_check_write+0x14/0x20
[  739.869016]  ? mutex_lock+0x8f/0xe0
[  739.872510]  ? __mutex_lock_slowpath+0x20/0x20
[  739.876956]  ? __rcu_read_unlock+0x5b/0x270
[  739.881144]  __x64_sys_io_uring_enter+0x367/0xef0
[  739.885859]  ? io_submit_sqes+0x1260/0x1260
[  739.890044]  ? __this_cpu_preempt_check+0x13/0x20
[  739.894749]  ? xfd_validate_state+0x3c/0xd0
[  739.898936]  ? __schedule+0x5be/0x10c0
[  739.902687]  ? restore_fpregs_from_fpstate+0xa2/0x170
[  739.907741]  ? kernel_fpu_begin_mask+0x170/0x170
[  739.912362]  ? debug_smp_processor_id+0x17/0x20
[  739.916903]  ? debug_smp_processor_id+0x17/0x20
[  739.921434]  ? fpregs_assert_state_consistent+0x5f/0x70
[  739.926662]  ? exit_to_user_mode_prepare+0x4b/0x1e0
[  739.931549]  do_syscall_64+0x38/0x90
[  739.935129]  entry_SYSCALL_64_after_hwframe+0x44/0xae
[  739.940182] RIP: 0033:0x7f7345d7d89d
[  739.943759] Code: 00 c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa
48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f
05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d c3 f5 0c 00 f7 d8 64 89 01 48
[  739.962513] RSP: 002b:00007f7345c5ee98 EFLAGS: 00000246 ORIG_RAX:
00000000000001aa
[  739.970081] RAX: ffffffffffffffda RBX: 00007f7345fd2fc0 RCX:
00007f7345d7d89d
[  739.977216] RDX: 0000000000000000 RSI: 0000000000000020 RDI:
0000000000000004
[  739.984349] RBP: 0000000000000020 R08: 0000000000000000 R09:
0000000000000000
[  739.991481] R10: 0000000000000000 R11: 0000000000000246 R12:
000055be1d11e2a0
[  739.998612] R13: 0000000000000020 R14: 0000000000000000 R15:
0000000000000020
[  740.005749]  </TASK>
[  740.007945]
[  740.009444] Allocated by task 10373:
[  740.013084]
[  740.014583] Freed by task 10373:
[  740.017879]
[  740.019378] The buggy address belongs to the object at ffff88c1be016e00
[  740.019378]  which belongs to the cache kmalloc-256 of size 256
[  740.031886] The buggy address is located 40 bytes inside of
[  740.031886]  256-byte region [ffff88c1be016e00, ffff88c1be016f00)
[  740.043534] The buggy address belongs to the page:
[  740.048345]
[  740.049840] Memory state around the buggy address:
[  740.054636]  ffff88c1be016d00: fc fc fc fc fc fc fc fc fc fc fc fc fc
fc fc fc
[  740.061854]  ffff88c1be016d80: fc fc fc fc fc fc fc fc fc fc fc fc fc
fc fc fc
[  740.069074] >ffff88c1be016e00: fa fb fb fb fb fb fb fb fb fb fb fb fb
fb fb fb
[  740.076294]                                   ^
[  740.080827]  ffff88c1be016e80: fb fb fb fb fb fb fb fb fb fb fb fb fb
fb fb fb
[  740.088045]  ffff88c1be016f00: fc fc fc fc fc fc fc fc fc fc fc fc fc
fc fc fc
[  740.095265]
==================================================================
[  740.102497] kernel BUG at mm/slub.c:379!
[  740.106431] invalid opcode: 0000 [#1] PREEMPT SMP KASAN NOPTI
What's the exact command line that triggers this? I am not able to
trigger it with
either fio or t/io_uring.

Song

Re: [PATCH v5 3/4] md: raid10 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-16 20:38:00

On 12/16/21 1:18 PM, Song Liu wrote:
On Thu, Dec 16, 2021 at 11:40 AM Vishal Verma [off-list ref] wrote:
quoted
On 12/16/21 11:49 AM, Jens Axboe wrote:
quoted
On 12/16/21 9:45 AM, Vishal Verma wrote:
quoted
On 12/16/21 9:42 AM, Jens Axboe wrote:
quoted
On 12/15/21 5:30 PM, Vishal Verma wrote:
quoted
On 12/15/21 3:20 PM, Vishal Verma wrote:
quoted
On 12/15/21 1:42 PM, Song Liu wrote:
quoted
On Tue, Dec 14, 2021 at 10:09 PM Vishal Verma
[off-list ref] wrote:
quoted
This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

- Waiting for the barrier,
- Too many pending I/Os to be queued,
- Reshape operation,
- Discard operation.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
     drivers/md/raid10.c | 57
+++++++++++++++++++++++++++++++++++----------
     1 file changed, 45 insertions(+), 12 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..f6c73987e9ac 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,11 +952,18 @@ static void lower_barrier(struct r10conf *conf)
            wake_up(&conf->wait_barrier);
     }

-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
     {
            spin_lock_irq(&conf->resync_lock);
            if (conf->barrier) {
                    struct bio_list *bio_list = current->bio_list;
+
+               /* Return false when nowait flag is set */
+               if (nowait) {
+ spin_unlock_irq(&conf->resync_lock);
+                       return false;
+               }
+
                    conf->nr_waiting++;
                    /* Wait for the barrier to drop.
                     * However if there are already pending
@@ -988,6 +995,7 @@ static void wait_barrier(struct r10conf *conf)
            }
            atomic_inc(&conf->nr_pending);
            spin_unlock_irq(&conf->resync_lock);
+       return true;
     }

     static void allow_barrier(struct r10conf *conf)
@@ -1101,17 +1109,25 @@ static void raid10_unplug(struct blk_plug_cb
*cb, bool from_schedule)
     static void regular_request_wait(struct mddev *mddev, struct
r10conf *conf,
                                     struct bio *bio, sector_t sectors)
     {
-       wait_barrier(conf);
+       /* Bail out if REQ_NOWAIT is set for the bio */
+       if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return;
+       }
I think we also need regular_request_wait to return bool and handle
it properly.

Thanks,
Song
Ack, will fix it. Thanks!
Ran into this while running with io_uring. With the current v5 (raid10
patch) on top of md-next branch.
./t/io_uring -a 0 -d 256 </dev/raid10>

It didn't trigger with aio (-a 1)

[  248.128661] BUG: kernel NULL pointer dereference, address:
00000000000000b8
[  248.135628] #PF: supervisor read access in kernel mode
[  248.140762] #PF: error_code(0x0000) - not-present page
[  248.145903] PGD 0 P4D 0
[  248.148443] Oops: 0000 [#1] PREEMPT SMP NOPTI
[  248.152800] CPU: 49 PID: 9461 Comm: io_uring Kdump: loaded Not
tainted 5.16.0-rc3+ #2
[  248.160629] Hardware name: Dell Inc. PowerEdge R650xs/0PPTY2, BIOS
1.3.8 08/31/2021
[  248.168279] RIP: 0010:raid10_end_read_request+0x74/0x140 [raid10]
[  248.174373] Code: 48 60 48 8b 58 58 48 c1 e2 05 49 03 55 08 48 89 4a
10 40 84 f6 75 48 f0 41 80 4c 24 18 01 4c 89 e7 e8 e0 b8 ff ff 49 8b 4d
00 <48> 8b 83 b8 00 00 00 f0 ff 8b f0 00 00 00 0f 94 c2 a8 01 74 04 84
[  248.193120] RSP: 0018:ffffb1c38d598ce8 EFLAGS: 00010086
[  248.198344] RAX: ffff8e5da2a1a100 RBX: 0000000000000000 RCX:
ffff8e5d89747000
[  248.205479] RDX: 000000008040003a RSI: 0000000080400039 RDI:
ffff8e1e00044900
[  248.212611] RBP: ffffb1c38d598d30 R08: 0000000000000000 R09:
0000000000000001
[  248.219744] R10: ffff8e5da2a1ae00 R11: 000000411bab9000 R12:
ffff8e5da2a1ae00
[  248.226877] R13: ffff8e5d8973fc00 R14: 0000000000000000 R15:
0000000000001000
[  248.234009] FS:  00007fc26b07d700(0000) GS:ffff8e9c6e600000(0000)
knlGS:0000000000000000
[  248.242096] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  248.247843] CR2: 00000000000000b8 CR3: 00000040b25d4005 CR4:
0000000000770ee0
[  248.254973] DR0: 0000000000000000 DR1: 0000000000000000 DR2:
0000000000000000
[  248.262107] DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7:
0000000000000400
[  248.269240] PKRU: 55555554
[  248.271953] Call Trace:
[  248.274406]  <IRQ>
[  248.276425]  bio_endio+0xf6/0x170
[  248.279743]  blk_update_request+0x12d/0x470
[  248.283931]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.288809]  blk_mq_end_request_batch+0x76/0x490
[  248.293429]  ? dma_direct_unmap_sg+0xdd/0x1a0
[  248.297786]  ? smp_call_function_single_async+0x46/0x70
[  248.303015]  ? mempool_kfree+0xe/0x10
[  248.306680]  ? mempool_kfree+0xe/0x10
[  248.310345]  nvme_pci_complete_batch+0x26/0xb0
[  248.314792]  nvme_irq+0x298/0x2f0
[  248.318110]  ? nvme_unmap_data+0xf0/0xf0
[  248.322038]  __handle_irq_event_percpu+0x3f/0x190
[  248.326744]  handle_irq_event_percpu+0x33/0x80
[  248.331190]  handle_irq_event+0x39/0x60
[  248.335028]  handle_edge_irq+0xbe/0x1e0
[  248.338869]  __common_interrupt+0x6b/0x110
[  248.342967]  common_interrupt+0xbd/0xe0
[  248.346808]  </IRQ>
[  248.348912]  <TASK>
[  248.351018]  asm_common_interrupt+0x1e/0x40
[  248.355206] RIP: 0010:_raw_spin_unlock_irqrestore+0x1e/0x37
[  248.360780] Code: 02 5d c3 0f 1f 44 00 00 5d c3 66 90 0f 1f 44 00 00
55 48 89 e5 c6 07 00 0f 1f 40 00 f7 c6 00 02 00 00 74 01 fb bf 01 00 00
00 <e8> ed 8e 5b ff 65 8b 05 66 7e 52 78 85 c0 74 02 5d c3 0f 1f 44 00

[  248.379525] RSP: 0018:ffffb1c3a429b958 EFLAGS: 00000206
[  248.384749] RAX: 0000000000000001 RBX: ffff8e5d8973fd08 RCX:
ffff8e5d8973fd10
[  248.391884] RDX: 0000000000000001 RSI: 0000000000000246 RDI:
0000000000000001
[  248.399017] RBP: ffffb1c3a429b958 R08: 0000000000000000 R09:
ffffb1c3a429b970
[  248.406148] R10: 0000000000000c00 R11: 0000000000000001 R12:
0000000000000001
[  248.413280] R13: 0000000000000246 R14: 0000000000000000 R15:
0000000000000003
[  248.420415]  __wake_up_common_lock+0x8a/0xc0
[  248.424686]  __wake_up+0x13/0x20
[  248.427919]  raid10_make_request+0x101/0x170 [raid10]
[  248.432971]  md_handle_request+0x179/0x1e0
[  248.437071]  ? submit_bio_checks+0x1f6/0x5a0
[  248.441345]  md_submit_bio+0x6d/0xa0
[  248.444924]  __submit_bio+0x94/0x140
[  248.448504]  submit_bio_noacct+0xe1/0x2a0
[  248.452515]  submit_bio+0x48/0x120
[  248.455923]  blkdev_direct_IO+0x220/0x540
[  248.459935]  ? __fsnotify_parent+0xff/0x330
[  248.464121]  ? __fsnotify_parent+0x10f/0x330
[  248.468393]  ? common_interrupt+0x73/0xe0
[  248.472408]  generic_file_read_iter+0xa5/0x160
[  248.476852]  blkdev_read_iter+0x38/0x70
[  248.480693]  io_read+0x119/0x420
[  248.483923]  ? sbitmap_queue_clear_batch+0xc7/0x110
[  248.488805]  ? blk_mq_end_request_batch+0x378/0x490
[  248.493684]  io_issue_sqe+0x7ec/0x19c0
[  248.497436]  ? io_req_prep+0x6a9/0xe60
[  248.501190]  io_submit_sqes+0x2a0/0x9f0
[  248.505030]  ? __fget_files+0x6a/0x90
[  248.508693]  __x64_sys_io_uring_enter+0x1da/0x8c0
[  248.513401]  do_syscall_64+0x38/0x90
[  248.516979]  entry_SYSCALL_64_after_hwframe+0x44/0xae
[  248.522033] RIP: 0033:0x7fc26b19b89d
[  248.525611] Code: 00 c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa
48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f
05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d c3 f5 0c 00 f7 d8 64 89 01 48
[  248.544360] RSP: 002b:00007fc26b07ce98 EFLAGS: 00000246 ORIG_RAX:
00000000000001aa
[  248.551925] RAX: ffffffffffffffda RBX: 00007fc26b3f2fc0 RCX:
00007fc26b19b89d
[  248.559058] RDX: 0000000000000020 RSI: 0000000000000020 RDI:
0000000000000004
[  248.566189] RBP: 0000000000000020 R08: 0000000000000000 R09:
0000000000000000
[  248.573322] R10: 0000000000000001 R11: 0000000000000246 R12:
00005623a4b7a2a0
[  248.580456] R13: 0000000000000020 R14: 0000000000000020 R15:
0000000000000020
[  248.587591]  </TASK>
Do you have:

commit 75feae73a28020e492fbad2323245455ef69d687
Author: Pavel Begunkov [off-list ref]
Date:   Tue Dec 7 20:16:36 2021 +0000

       block: fix single bio async DIO error handling

in your tree?
Nope. I will get it in and test. Thanks!
Might be worth re-running with KASAN enabled in your config to see if
that triggers anything.
Got this:
[  739.336669] CPU: 63 PID: 10373 Comm: io_uring Kdump: loaded Not
tainted 5.16.0-rc3+ #8
[  739.344583] Hardware name: Dell Inc. PowerEdge R650xs/0PPTY2, BIOS
1.3.8 08/31/2021
[  739.352236] Call Trace:
[  739.354687]  <IRQ>
[  739.356705]  dump_stack_lvl+0x38/0x49
[  739.360381]  print_address_description.constprop.0+0x28/0x150
[  739.366136]  ? 0xffffffffc046f3df
[  739.369455]  kasan_report.cold+0x82/0xdb
[  739.373383]  ? 0xffffffffc046f3df
[  739.376700]  __asan_load8+0x69/0x90
[  739.380194]  0xffffffffc046f3df
[  739.383339]  ? 0xffffffffc046f340
[  739.386659]  ? blkcg_iolatency_done_bio+0x26/0x390
[  739.391461]  ? __rcu_read_unlock+0x5b/0x270
[  739.395655]  ? kmem_cache_alloc+0x143/0x460
[  739.399841]  ? mempool_alloc_slab+0x17/0x20
[  739.404027]  ? bio_uninit+0x6c/0xf0
[  739.407522]  bio_endio+0x27f/0x2a0
[  739.410926]  blk_update_request+0x1e8/0x750
[  739.415112]  blk_mq_end_request_batch+0x10b/0x9b0
[  739.419818]  ? blk_mq_end_request+0x460/0x460
[  739.424179]  ? kfree+0xa0/0x400
[  739.427322]  ? mempool_kfree+0xe/0x10
[  739.430989]  ? generic_file_write_iter+0xf0/0xf0
[  739.435609]  ? dma_unmap_page_attrs+0x15f/0x2c0
[  739.440144]  nvme_pci_complete_batch+0x34/0x160
[  739.444684]  ? blk_mq_complete_request_remote+0x1ca/0x2d0
[  739.450084]  nvme_irq+0x5fa/0x630
[  739.453404]  ? nvme_timeout+0x370/0x370
[  739.457242]  ? nvme_unmap_data+0x1e0/0x1e0
[  739.461340]  ? __kasan_check_write+0x14/0x20
[  739.465614]  ? credit_entropy_bits.constprop.0+0x76/0x190
[  739.471015]  ? nvme_timeout+0x370/0x370
[  739.474853]  __handle_irq_event_percpu+0x69/0x260
[  739.479568]  handle_irq_event_percpu+0x70/0xf0
[  739.484015]  ? __handle_irq_event_percpu+0x260/0x260
[  739.488981]  ? __kasan_check_write+0x14/0x20
[  739.493261]  ? _raw_spin_lock+0x88/0xe0
[  739.497109]  ? _raw_spin_lock_irqsave+0xf0/0xf0
[  739.501643]  handle_irq_event+0x5a/0x90
[  739.505480]  handle_edge_irq+0x148/0x320
[  739.509407]  __common_interrupt+0x75/0x130
[  739.513514]  common_interrupt+0xae/0xd0
[  739.517354]  </IRQ>
[  739.519460]  <TASK>
[  739.521567]  asm_common_interrupt+0x1e/0x40
[  739.525753] RIP: 0010:__asan_store8+0x37/0x90
[  739.530111] Code: 4f 48 b8 ff ff ff ff ff 7f ff ff 48 39 c7 76 40 48
8d 47 07 48 89 c2 83 e2 07 48 83 fa 07 75 18 48 ba 00 00 00 00 00 fc ff
df <48> c1 e8 03 0f b6 04 10 84 c0 75 2b 5d c3 48 be 00 00 00 00 00 fc
[  739.548865] RSP: 0018:ffffc900307c7100 EFLAGS: 00000246
[  739.554092] RAX: ffffc900307c7237 RBX: ffffc900307c7780 RCX:
ffffffff82c9b461
[  739.561227] RDX: dffffc0000000000 RSI: ffffc900307c7790 RDI:
ffffc900307c7230
[  739.568358] RBP: ffffc900307c7100 R08: ffffffff82c9b202 R09:
ffff8882d96c0000
[  739.575490] R10: ffffc900307c7257 R11: fffff520060f8e4a R12:
0000000082c9b201
[  739.582624] R13: 0000000000000000 R14: ffffc900307c7238 R15:
ffffc900307c71e8
[  739.589760]  ? update_stack_state+0x22/0x2c0
[  739.594038]  ? update_stack_state+0x281/0x2c0
[  739.598398]  update_stack_state+0x281/0x2c0
[  739.602585]  unwind_next_frame.part.0+0xe0/0x360
[  739.607204]  ? bio_alloc_bioset+0x223/0x2f0
[  739.611389]  ? create_prof_cpu_mask+0x30/0x30
[  739.615749]  ? mempool_alloc_slab+0x17/0x20
[  739.619935]  unwind_next_frame+0x23/0x30
[  739.623860]  arch_stack_walk+0x88/0xf0
[  739.627613]  ? bio_alloc_bioset+0x223/0x2f0
[  739.631800]  stack_trace_save+0x94/0xc0
[  739.635640]  ? filter_irq_stacks+0x70/0x70
[  739.639738]  ? blk_mq_put_tag+0x80/0x80
[  739.643576]  ? _raw_spin_unlock_irqrestore+0x23/0x40
[  739.648545]  ? __wake_up_common_lock+0xfd/0x150
[  739.653084]  kasan_save_stack+0x26/0x60
[  739.656924]  ? kasan_save_stack+0x26/0x60
[  739.660938]  ? __kasan_slab_alloc+0x6d/0x90
[  739.665123]  ? kmem_cache_alloc+0x143/0x460
[  739.669308]  ? mempool_alloc_slab+0x17/0x20
[  739.673494]  ? mempool_alloc+0xef/0x280
[  739.677333]  ? bio_alloc_bioset+0x223/0x2f0
[  739.681519]  ? blk_mq_rq_ctx_init.isra.0+0x28a/0x3c0
[  739.686489]  ? __blk_mq_alloc_requests+0x655/0x680
[  739.691288]  ? blkcg_iolatency_throttle+0x5d/0x760
[  739.696081]  ? bio_to_wbt_flags+0x47/0xf0
[  739.700093]  ? update_io_ticks+0x5e/0xd0
[  739.704021]  ? preempt_count_sub+0x18/0xc0
[  739.708120]  ? __kasan_check_read+0x11/0x20
[  739.712305]  ? blk_mq_submit_bio+0x740/0xce0
[  739.716580]  ? blk_mq_try_issue_list_directly+0x1b0/0x1b0
[  739.721987]  ? kasan_poison+0x3c/0x50
[  739.725650]  ? kasan_unpoison+0x28/0x50
[  739.729491]  __kasan_slab_alloc+0x6d/0x90
[  739.733503]  kmem_cache_alloc+0x143/0x460
[  739.737518]  mempool_alloc_slab+0x17/0x20
[  739.741529]  mempool_alloc+0xef/0x280
[  739.745194]  ? mempool_free+0x170/0x170
[  739.749035]  ? mempool_destroy+0x30/0x30
[  739.752961]  ? __fsnotify_update_child_dentry_flags.part.0+0x170/0x170
[  739.759498]  bio_alloc_bioset+0x223/0x2f0
[  739.763517]  ? __this_cpu_preempt_check+0x13/0x20
[  739.768223]  ? bvec_alloc+0xd0/0xd0
[  739.771715]  ? __fsnotify_parent+0x1ed/0x590
[  739.775988]  ? do_direct_IO+0x150/0x1880
[  739.779916]  ? submit_bio+0xb0/0x220
[  739.783496]  bio_alloc_kiocb+0x185/0x1c0
[  739.787430]  blkdev_direct_IO+0x114/0x400
[  739.791441]  generic_file_read_iter+0x152/0x250
[  739.795974]  blkdev_read_iter+0x84/0xd0
[  739.799815]  io_read+0x1ec/0x770
[  739.803056]  ? __rcu_read_unlock+0x5b/0x270
[  739.807240]  ? io_setup_async_rw+0x270/0x270
[  739.811515]  ? __sbq_wake_up+0x2d/0x1b0
[  739.815352]  ? __rcu_read_unlock+0x5b/0x270
[  739.819537]  ? sbitmap_queue_clear+0xc9/0xe0
[  739.823813]  ? blk_queue_exit+0x35/0x90
[  739.827653]  ? __blk_mq_free_request+0x111/0x160
[  739.832280]  io_issue_sqe+0xcac/0x27f0
[  739.836031]  ? blk_mq_free_plug_rqs+0x3f/0x50
[  739.840394]  ? io_poll_add.isra.0+0x290/0x290
[  739.844760]  ? io_req_prep+0xcc2/0x1bb0
[  739.848598]  ? io_submit_sqes+0x43b/0x1260
[  739.852701]  io_submit_sqes+0x5e5/0x1260
[  739.856634]  ? io_do_iopoll+0x561/0x720
[  739.860474]  ? io_wq_submit_work+0x230/0x230
[  739.864746]  ? __kasan_check_write+0x14/0x20
[  739.869016]  ? mutex_lock+0x8f/0xe0
[  739.872510]  ? __mutex_lock_slowpath+0x20/0x20
[  739.876956]  ? __rcu_read_unlock+0x5b/0x270
[  739.881144]  __x64_sys_io_uring_enter+0x367/0xef0
[  739.885859]  ? io_submit_sqes+0x1260/0x1260
[  739.890044]  ? __this_cpu_preempt_check+0x13/0x20
[  739.894749]  ? xfd_validate_state+0x3c/0xd0
[  739.898936]  ? __schedule+0x5be/0x10c0
[  739.902687]  ? restore_fpregs_from_fpstate+0xa2/0x170
[  739.907741]  ? kernel_fpu_begin_mask+0x170/0x170
[  739.912362]  ? debug_smp_processor_id+0x17/0x20
[  739.916903]  ? debug_smp_processor_id+0x17/0x20
[  739.921434]  ? fpregs_assert_state_consistent+0x5f/0x70
[  739.926662]  ? exit_to_user_mode_prepare+0x4b/0x1e0
[  739.931549]  do_syscall_64+0x38/0x90
[  739.935129]  entry_SYSCALL_64_after_hwframe+0x44/0xae
[  739.940182] RIP: 0033:0x7f7345d7d89d
[  739.943759] Code: 00 c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa
48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f
05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d c3 f5 0c 00 f7 d8 64 89 01 48
[  739.962513] RSP: 002b:00007f7345c5ee98 EFLAGS: 00000246 ORIG_RAX:
00000000000001aa
[  739.970081] RAX: ffffffffffffffda RBX: 00007f7345fd2fc0 RCX:
00007f7345d7d89d
[  739.977216] RDX: 0000000000000000 RSI: 0000000000000020 RDI:
0000000000000004
[  739.984349] RBP: 0000000000000020 R08: 0000000000000000 R09:
0000000000000000
[  739.991481] R10: 0000000000000000 R11: 0000000000000246 R12:
000055be1d11e2a0
[  739.998612] R13: 0000000000000020 R14: 0000000000000000 R15:
0000000000000020
[  740.005749]  </TASK>
[  740.007945]
[  740.009444] Allocated by task 10373:
[  740.013084]
[  740.014583] Freed by task 10373:
[  740.017879]
[  740.019378] The buggy address belongs to the object at ffff88c1be016e00
[  740.019378]  which belongs to the cache kmalloc-256 of size 256
[  740.031886] The buggy address is located 40 bytes inside of
[  740.031886]  256-byte region [ffff88c1be016e00, ffff88c1be016f00)
[  740.043534] The buggy address belongs to the page:
[  740.048345]
[  740.049840] Memory state around the buggy address:
[  740.054636]  ffff88c1be016d00: fc fc fc fc fc fc fc fc fc fc fc fc fc
fc fc fc
[  740.061854]  ffff88c1be016d80: fc fc fc fc fc fc fc fc fc fc fc fc fc
fc fc fc
[  740.069074] >ffff88c1be016e00: fa fb fb fb fb fb fb fb fb fb fb fb fb
fb fb fb
[  740.076294]                                   ^
[  740.080827]  ffff88c1be016e80: fb fb fb fb fb fb fb fb fb fb fb fb fb
fb fb fb
[  740.088045]  ffff88c1be016f00: fc fc fc fc fc fc fc fc fc fc fc fc fc
fc fc fc
[  740.095265]
==================================================================
[  740.102497] kernel BUG at mm/slub.c:379!
[  740.106431] invalid opcode: 0000 [#1] PREEMPT SMP KASAN NOPTI
What's the exact command line that triggers this? I am not able to
trigger it with
either fio or t/io_uring.

Song
I only had 1 nvme so was creating 4 partitions on it and creating a 
raid10 and doing:

mdadm -C /dev/md10 -l 10 -n 4 /dev/nvme4n1p1 /dev/nvme4n1p2 
/dev/nvme4n1p3 /dev/nvme4n1p4
./t/io_uring /dev/md10-d 256 -p 0 -a 0 -r 100

on top of commit: c14704e1cb556 (md-next branch) + "md: add support for 
REQ_NOWAIT" patch
Also, applied the commit (75feae73a28) Jens pointed earlier today.

Re: [PATCH v5 3/4] md: raid10 add nowait support

From: Song Liu <song@kernel.org>
Date: 2021-12-16 23:50:35

On Thu, Dec 16, 2021 at 12:38 PM Vishal Verma [off-list ref] wrote:
[...]
quoted
quoted
[  740.106431] invalid opcode: 0000 [#1] PREEMPT SMP KASAN NOPTI
What's the exact command line that triggers this? I am not able to
trigger it with
either fio or t/io_uring.

Song
I only had 1 nvme so was creating 4 partitions on it and creating a
raid10 and doing:

mdadm -C /dev/md10 -l 10 -n 4 /dev/nvme4n1p1 /dev/nvme4n1p2
/dev/nvme4n1p3 /dev/nvme4n1p4
./t/io_uring /dev/md10-d 256 -p 0 -a 0 -r 100

on top of commit: c14704e1cb556 (md-next branch) + "md: add support for
REQ_NOWAIT" patch
Also, applied the commit (75feae73a28) Jens pointed earlier today.
I am able to trigger the following error. I will look into it.

Thanks,
Song

[ 1583.149004] ==================================================================
[ 1583.150100] BUG: KASAN: use-after-free in raid10_end_read_request+0x91/0x310
[ 1583.151042] Read of size 8 at addr ffff888160a1c928 by task io_uring/1165
[ 1583.152016]
[ 1583.152247] CPU: 0 PID: 1165 Comm: io_uring Not tainted 5.16.0-rc3+ #660
[ 1583.153159] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996),
BIOS 1.13.0-2.module_el8.4.0+547+a85d02ba 04/01/2014
[ 1583.154572] Call Trace:
[ 1583.155005]  <IRQ>
[ 1583.155338]  dump_stack_lvl+0x44/0x57
[ 1583.155950]  print_address_description.constprop.8.cold.17+0x12/0x339
[ 1583.156969]  ? raid10_end_read_request+0x91/0x310
[ 1583.157578]  ? raid10_end_read_request+0x91/0x310
[ 1583.158272]  kasan_report.cold.18+0x83/0xdf
[ 1583.158889]  ? raid10_end_read_request+0x91/0x310
[ 1583.159554]  raid10_end_read_request+0x91/0x310
[ 1583.160201]  ? raid10_resize+0x270/0x270
[ 1583.160724]  ? bio_uninit+0xc7/0x1e0
[ 1583.161274]  blk_update_request+0x21f/0x810
[ 1583.161893]  blk_mq_end_request_batch+0x11c/0xa70
[ 1583.162497]  ? blk_mq_end_request+0x460/0x460
[ 1583.163204]  ? nvme_complete_batch_req+0x12/0x30
[ 1583.163888]  nvme_irq+0x6ad/0x6f0
[ 1583.164354]  ? io_queue_count_set+0xe0/0xe0
[ 1583.164980]  ? nvme_unmap_data+0x1e0/0x1e0
[ 1583.165504]  ? rcu_read_lock_bh_held+0xb0/0xb0
[ 1583.166149]  ? io_queue_count_set+0xe0/0xe0
[ 1583.166721]  __handle_irq_event_percpu+0x79/0x440
[ 1583.167446]  handle_irq_event_percpu+0x6f/0xe0
[ 1583.168101]  ? __handle_irq_event_percpu+0x440/0x440
[ 1583.168734]  ? lock_contended+0x6e0/0x6e0
[ 1583.169349]  ? do_raw_spin_unlock+0xa2/0x130
[ 1583.169961]  handle_irq_event+0x54/0x90
[ 1583.170442]  handle_edge_irq+0x121/0x300
[ 1583.171012]  __common_interrupt+0x7d/0x170
[ 1583.171538]  common_interrupt+0xa0/0xc0
[ 1583.172103]  </IRQ>
[ 1583.172389]  <TASK>

[PATCH v6 1/4] md: add support for REQ_NOWAIT

From: Vishal Verma <hidden>
Date: 2021-12-21 20:06:39

commit 021a24460dc2 ("block: add QUEUE_FLAG_NOWAIT") added support
for checking whether a given bdev supports handling of REQ_NOWAIT or not.
Since then commit 6abc49468eea ("dm: add support for REQ_NOWAIT and enable
it for linear target") added support for REQ_NOWAIT for dm. This uses
a similar approach to incorporate REQ_NOWAIT for md based bios.

This patch was tested using t/io_uring tool within FIO. A nvme drive
was partitioned into 2 partitions and a simple raid 0 configuration
/dev/md0 was created.

md0 : active raid0 nvme4n1p1[1] nvme4n1p2[0]
      937423872 blocks super 1.2 512k chunks

Before patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

  38396   38396 pts/2    00:00:00 io_uring
  38396   38397 pts/2    00:00:15 io_uring
  38396   38398 pts/2    00:00:13 iou-wrk-38397

We can see iou-wrk-38397 io worker thread created which gets created
when io_uring sees that the underlying device (/dev/md0 in this case)
doesn't support nowait.

After patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

  38341   38341 pts/2    00:10:22 io_uring
  38341   38342 pts/2    00:10:37 io_uring

After running this patch, we don't see any io worker thread
being created which indicated that io_uring saw that the
underlying device does support nowait. This is the exact behaviour
noticed on a dm device which also supports nowait.

For all the other raid personalities except raid0, we would need
to train pieces which involves make_request fn in order for them
to correctly handle REQ_NOWAIT.

Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/md.c | 20 ++++++++++++++++++++
 1 file changed, 20 insertions(+)
diff --git a/drivers/md/md.c b/drivers/md/md.c
index 5111ed966947..ccd296aa9641 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -418,6 +418,11 @@ void md_handle_request(struct mddev *mddev, struct bio *bio)
 	rcu_read_lock();
 	if (is_suspended(mddev, bio)) {
 		DEFINE_WAIT(__wait);
+		/* Bail out if REQ_NOWAIT is set for the bio */
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return;
+		}
 		for (;;) {
 			prepare_to_wait(&mddev->sb_wait, &__wait,
 					TASK_UNINTERRUPTIBLE);
@@ -5792,6 +5797,7 @@ int md_run(struct mddev *mddev)
 	int err;
 	struct md_rdev *rdev;
 	struct md_personality *pers;
+	bool nowait = true;
 
 	if (list_empty(&mddev->disks))
 		/* cannot run an array with no devices.. */
@@ -5862,8 +5868,13 @@ int md_run(struct mddev *mddev)
 			}
 		}
 		sysfs_notify_dirent_safe(rdev->sysfs_state);
+		nowait = nowait && blk_queue_nowait(bdev_get_queue(rdev->bdev));
 	}
 
+	/* Set the NOWAIT flags if all underlying devices support it */
+	if (nowait)
+		blk_queue_flag_set(QUEUE_FLAG_NOWAIT, mddev->queue);
+
 	if (!bioset_initialized(&mddev->bio_set)) {
 		err = bioset_init(&mddev->bio_set, BIO_POOL_SIZE, 0, BIOSET_NEED_BVECS);
 		if (err)
@@ -7007,6 +7018,15 @@ static int hot_add_disk(struct mddev *mddev, dev_t dev)
 	set_bit(MD_SB_CHANGE_DEVS, &mddev->sb_flags);
 	if (!mddev->thread)
 		md_update_sb(mddev, 1);
+	/*
+	 * If the new disk does not support REQ_NOWAIT,
+	 * disable on the whole MD.
+	 */
+	if (!blk_queue_nowait(bdev_get_queue(rdev->bdev))) {
+		pr_info("%s: Disabling nowait because %s does not support nowait\n",
+			mdname(mddev), bdevname(rdev->bdev, b));
+		blk_queue_flag_clear(QUEUE_FLAG_NOWAIT, mddev->queue);
+	}
 	/*
 	 * Kick recovery, maybe this spare has to be added to the
 	 * array immediately.
-- 
2.17.1

[PATCH v6 2/4] md: raid1 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-21 20:06:44

This adds nowait support to the RAID1 driver. It makes RAID1 driver
return with EAGAIN for situations where it could wait for eg:

  - Waiting for the barrier,
  - Too many pending I/Os to be queued.

wait_barrier() fn is modified to return bool to support error for
wait barriers. It returns true in case of wait or if wait is not
required and returns false if wait was required but not performed
to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/raid1.c | 83 +++++++++++++++++++++++++++++++++++-----------
 1 file changed, 64 insertions(+), 19 deletions(-)
diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c
index 7dc8026cf6ee..e488671bb563 100644
--- a/drivers/md/raid1.c
+++ b/drivers/md/raid1.c
@@ -929,8 +929,9 @@ static void lower_barrier(struct r1conf *conf, sector_t sector_nr)
 	wake_up(&conf->wait_barrier);
 }
 
-static void _wait_barrier(struct r1conf *conf, int idx)
+static bool _wait_barrier(struct r1conf *conf, int idx, bool nowait)
 {
+	bool ret = true;
 	/*
 	 * We need to increase conf->nr_pending[idx] very early here,
 	 * then raise_barrier() can be blocked when it waits for
@@ -961,7 +962,7 @@ static void _wait_barrier(struct r1conf *conf, int idx)
 	 */
 	if (!READ_ONCE(conf->array_frozen) &&
 	    !atomic_read(&conf->barrier[idx]))
-		return;
+		return ret;
 
 	/*
 	 * After holding conf->resync_lock, conf->nr_pending[idx]
@@ -979,18 +980,29 @@ static void _wait_barrier(struct r1conf *conf, int idx)
 	 */
 	wake_up(&conf->wait_barrier);
 	/* Wait for the barrier in same barrier unit bucket to drop. */
-	wait_event_lock_irq(conf->wait_barrier,
-			    !conf->array_frozen &&
-			     !atomic_read(&conf->barrier[idx]),
-			    conf->resync_lock);
-	atomic_inc(&conf->nr_pending[idx]);
+
+	/* Return false when nowait flag is set */
+	if (nowait)
+		ret = false;
+	else {
+		wait_event_lock_irq(conf->wait_barrier,
+				!conf->array_frozen &&
+				!atomic_read(&conf->barrier[idx]),
+				conf->resync_lock);
+	}
+
+	/* Only increment nr_pending when we wait */
+	if (ret)
+		atomic_inc(&conf->nr_pending[idx]);
 	atomic_dec(&conf->nr_waiting[idx]);
 	spin_unlock_irq(&conf->resync_lock);
+	return ret;
 }
 
-static void wait_read_barrier(struct r1conf *conf, sector_t sector_nr)
+static bool wait_read_barrier(struct r1conf *conf, sector_t sector_nr, bool nowait)
 {
 	int idx = sector_to_idx(sector_nr);
+	bool ret = true;
 
 	/*
 	 * Very similar to _wait_barrier(). The difference is, for read
@@ -1002,7 +1014,7 @@ static void wait_read_barrier(struct r1conf *conf, sector_t sector_nr)
 	atomic_inc(&conf->nr_pending[idx]);
 
 	if (!READ_ONCE(conf->array_frozen))
-		return;
+		return ret;
 
 	spin_lock_irq(&conf->resync_lock);
 	atomic_inc(&conf->nr_waiting[idx]);
@@ -1013,19 +1025,30 @@ static void wait_read_barrier(struct r1conf *conf, sector_t sector_nr)
 	 */
 	wake_up(&conf->wait_barrier);
 	/* Wait for array to be unfrozen */
-	wait_event_lock_irq(conf->wait_barrier,
-			    !conf->array_frozen,
-			    conf->resync_lock);
-	atomic_inc(&conf->nr_pending[idx]);
+
+	/* Return false when nowait flag is set */
+	if (nowait)
+		/* Return false when nowait flag is set */
+		ret = false;
+	else {
+		wait_event_lock_irq(conf->wait_barrier,
+				!conf->array_frozen,
+				conf->resync_lock);
+	}
+
+	/* Only increment nr_pending when we wait */
+	if (ret)
+		atomic_inc(&conf->nr_pending[idx]);
 	atomic_dec(&conf->nr_waiting[idx]);
 	spin_unlock_irq(&conf->resync_lock);
+	return ret;
 }
 
-static void wait_barrier(struct r1conf *conf, sector_t sector_nr)
+static bool wait_barrier(struct r1conf *conf, sector_t sector_nr, bool nowait)
 {
 	int idx = sector_to_idx(sector_nr);
 
-	_wait_barrier(conf, idx);
+	return _wait_barrier(conf, idx, nowait);
 }
 
 static void _allow_barrier(struct r1conf *conf, int idx)
@@ -1236,7 +1259,11 @@ static void raid1_read_request(struct mddev *mddev, struct bio *bio,
 	 * Still need barrier for READ in case that whole
 	 * array is frozen.
 	 */
-	wait_read_barrier(conf, bio->bi_iter.bi_sector);
+	if (!wait_read_barrier(conf, bio->bi_iter.bi_sector,
+				bio->bi_opf & REQ_NOWAIT)) {
+		bio_wouldblock_error(bio);
+		return;
+	}
 
 	if (!r1_bio)
 		r1_bio = alloc_r1bio(mddev, bio);
@@ -1336,6 +1363,10 @@ static void raid1_write_request(struct mddev *mddev, struct bio *bio,
 		     bio->bi_iter.bi_sector, bio_end_sector(bio))) {
 
 		DEFINE_WAIT(w);
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return;
+		}
 		for (;;) {
 			prepare_to_wait(&conf->wait_barrier,
 					&w, TASK_IDLE);
@@ -1353,17 +1384,26 @@ static void raid1_write_request(struct mddev *mddev, struct bio *bio,
 	 * thread has put up a bar for new requests.
 	 * Continue immediately if no resync is active currently.
 	 */
-	wait_barrier(conf, bio->bi_iter.bi_sector);
+	if (!wait_barrier(conf, bio->bi_iter.bi_sector,
+				bio->bi_opf & REQ_NOWAIT)) {
+		bio_wouldblock_error(bio);
+		return;
+	}
 
 	r1_bio = alloc_r1bio(mddev, bio);
 	r1_bio->sectors = max_write_sectors;
 
 	if (conf->pending_count >= max_queued_requests) {
 		md_wakeup_thread(mddev->thread);
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return;
+		}
 		raid1_log(mddev, "wait queued");
 		wait_event(conf->wait_barrier,
 			   conf->pending_count < max_queued_requests);
 	}
+
 	/* first select target devices under rcu_lock and
 	 * inc refcount on their rdev.  Record them by setting
 	 * bios[x] to bio
@@ -1458,9 +1498,14 @@ static void raid1_write_request(struct mddev *mddev, struct bio *bio,
 				rdev_dec_pending(conf->mirrors[j].rdev, mddev);
 		r1_bio->state = 0;
 		allow_barrier(conf, bio->bi_iter.bi_sector);
+
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return;
+		}
 		raid1_log(mddev, "wait rdev %d blocked", blocked_rdev->raid_disk);
 		md_wait_for_blocked_rdev(blocked_rdev, mddev);
-		wait_barrier(conf, bio->bi_iter.bi_sector);
+		wait_barrier(conf, bio->bi_iter.bi_sector, false);
 		goto retry_write;
 	}
 
@@ -1687,7 +1732,7 @@ static void close_sync(struct r1conf *conf)
 	int idx;
 
 	for (idx = 0; idx < BARRIER_BUCKETS_NR; idx++) {
-		_wait_barrier(conf, idx);
+		_wait_barrier(conf, idx, false);
 		_allow_barrier(conf, idx);
 	}
 
-- 
2.17.1

[PATCH v6 3/4] md: raid10 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-21 20:06:47

This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

  - Waiting for the barrier,
  - Too many pending I/Os to be queued,
  - Reshape operation,
  - Discard operation.

wait_barrier() and regular_request_wait() fn are modified to return bool
to support error for wait barriers. They returns true in case of wait
or if wait is not required and returns false if wait was required
but not performed to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/raid10.c | 90 +++++++++++++++++++++++++++++++--------------
 1 file changed, 62 insertions(+), 28 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..7ceae00e863e 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,8 +952,9 @@ static void lower_barrier(struct r10conf *conf)
 	wake_up(&conf->wait_barrier);
 }
 
-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
 {
+	bool ret = true;
 	spin_lock_irq(&conf->resync_lock);
 	if (conf->barrier) {
 		struct bio_list *bio_list = current->bio_list;
@@ -968,26 +969,33 @@ static void wait_barrier(struct r10conf *conf)
 		 * count down.
 		 */
 		raid10_log(conf->mddev, "wait barrier");
-		wait_event_lock_irq(conf->wait_barrier,
-				    !conf->barrier ||
-				    (atomic_read(&conf->nr_pending) &&
-				     bio_list &&
-				     (!bio_list_empty(&bio_list[0]) ||
-				      !bio_list_empty(&bio_list[1]))) ||
-				     /* move on if recovery thread is
-				      * blocked by us
-				      */
-				     (conf->mddev->thread->tsk == current &&
-				      test_bit(MD_RECOVERY_RUNNING,
-					       &conf->mddev->recovery) &&
-				      conf->nr_queued > 0),
-				    conf->resync_lock);
+		/* Return false when nowait flag is set */
+		if (nowait)
+			ret = false;
+		else
+			wait_event_lock_irq(conf->wait_barrier,
+					    !conf->barrier ||
+					    (atomic_read(&conf->nr_pending) &&
+					     bio_list &&
+					     (!bio_list_empty(&bio_list[0]) ||
+					      !bio_list_empty(&bio_list[1]))) ||
+					     /* move on if recovery thread is
+					      * blocked by us
+					      */
+					     (conf->mddev->thread->tsk == current &&
+					      test_bit(MD_RECOVERY_RUNNING,
+						       &conf->mddev->recovery) &&
+					      conf->nr_queued > 0),
+					    conf->resync_lock);
 		conf->nr_waiting--;
 		if (!conf->nr_waiting)
 			wake_up(&conf->wait_barrier);
 	}
-	atomic_inc(&conf->nr_pending);
+	/* Only increment nr_pending when we wait */
+	if (ret)
+		atomic_inc(&conf->nr_pending);
 	spin_unlock_irq(&conf->resync_lock);
+	return ret;
 }
 
 static void allow_barrier(struct r10conf *conf)
@@ -1098,21 +1106,30 @@ static void raid10_unplug(struct blk_plug_cb *cb, bool from_schedule)
  * currently.
  * 2. If IO spans the reshape position.  Need to wait for reshape to pass.
  */
-static void regular_request_wait(struct mddev *mddev, struct r10conf *conf,
+static bool regular_request_wait(struct mddev *mddev, struct r10conf *conf,
 				 struct bio *bio, sector_t sectors)
 {
-	wait_barrier(conf);
+	/* Bail out if REQ_NOWAIT is set for the bio */
+	if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+		bio_wouldblock_error(bio);
+		return false;
+	}
 	while (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery) &&
 	    bio->bi_iter.bi_sector < conf->reshape_progress &&
 	    bio->bi_iter.bi_sector + sectors > conf->reshape_progress) {
 		raid10_log(conf->mddev, "wait reshape");
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return false;
+		}
 		allow_barrier(conf);
 		wait_event(conf->wait_barrier,
 			   conf->reshape_progress <= bio->bi_iter.bi_sector ||
 			   conf->reshape_progress >= bio->bi_iter.bi_sector +
 			   sectors);
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 	}
+	return true;
 }
 
 static void raid10_read_request(struct mddev *mddev, struct bio *bio,
@@ -1179,7 +1196,7 @@ static void raid10_read_request(struct mddev *mddev, struct bio *bio,
 		bio_chain(split, bio);
 		allow_barrier(conf);
 		submit_bio_noacct(bio);
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 		bio = split;
 		r10_bio->master_bio = bio;
 		r10_bio->sectors = max_sectors;
@@ -1338,7 +1355,7 @@ static void wait_blocked_dev(struct mddev *mddev, struct r10bio *r10_bio)
 		raid10_log(conf->mddev, "%s wait rdev %d blocked",
 				__func__, blocked_rdev->raid_disk);
 		md_wait_for_blocked_rdev(blocked_rdev, mddev);
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 		goto retry_wait;
 	}
 }
@@ -1356,6 +1373,11 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
 					    bio->bi_iter.bi_sector,
 					    bio_end_sector(bio)))) {
 		DEFINE_WAIT(w);
+		/* Bail out if REQ_NOWAIT is set for the bio */
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return;
+		}
 		for (;;) {
 			prepare_to_wait(&conf->wait_barrier,
 					&w, TASK_IDLE);
@@ -1381,6 +1403,10 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
 			      BIT(MD_SB_CHANGE_DEVS) | BIT(MD_SB_CHANGE_PENDING));
 		md_wakeup_thread(mddev->thread);
 		raid10_log(conf->mddev, "wait reshape metadata");
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return;
+		}
 		wait_event(mddev->sb_wait,
 			   !test_bit(MD_SB_CHANGE_PENDING, &mddev->sb_flags));
 
@@ -1390,6 +1416,10 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
 	if (conf->pending_count >= max_queued_requests) {
 		md_wakeup_thread(mddev->thread);
 		raid10_log(mddev, "wait queued");
+		if (bio->bi_opf & REQ_NOWAIT) {
+			bio_wouldblock_error(bio);
+			return;
+		}
 		wait_event(conf->wait_barrier,
 			   conf->pending_count < max_queued_requests);
 	}
@@ -1482,7 +1512,7 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
 		bio_chain(split, bio);
 		allow_barrier(conf);
 		submit_bio_noacct(bio);
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 		bio = split;
 		r10_bio->master_bio = bio;
 	}
@@ -1607,7 +1637,11 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
 	if (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery))
 		return -EAGAIN;
 
-	wait_barrier(conf);
+	if (bio->bi_opf & REQ_NOWAIT) {
+		bio_wouldblock_error(bio);
+		return 0;
+	}
+	wait_barrier(conf, false);
 
 	/*
 	 * Check reshape again to avoid reshape happens after checking
@@ -1649,7 +1683,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
 		allow_barrier(conf);
 		/* Resend the fist split part */
 		submit_bio_noacct(split);
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 	}
 	div_u64_rem(bio_end, stripe_size, &remainder);
 	if (remainder) {
@@ -1660,7 +1694,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
 		/* Resend the second split part */
 		submit_bio_noacct(bio);
 		bio = split;
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 	}
 
 	bio_start = bio->bi_iter.bi_sector;
@@ -1816,7 +1850,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
 		end_disk_offset += geo->stride;
 		atomic_inc(&first_r10bio->remaining);
 		raid_end_discard_bio(r10_bio);
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 		goto retry_discard;
 	}
 
@@ -2011,7 +2045,7 @@ static void print_conf(struct r10conf *conf)
 
 static void close_sync(struct r10conf *conf)
 {
-	wait_barrier(conf);
+	wait_barrier(conf, false);
 	allow_barrier(conf);
 
 	mempool_exit(&conf->r10buf_pool);
@@ -4819,7 +4853,7 @@ static sector_t reshape_request(struct mddev *mddev, sector_t sector_nr,
 	if (need_flush ||
 	    time_after(jiffies, conf->reshape_checkpoint + 10*HZ)) {
 		/* Need to update reshape_position in metadata */
-		wait_barrier(conf);
+		wait_barrier(conf, false);
 		mddev->reshape_position = conf->reshape_progress;
 		if (mddev->reshape_backwards)
 			mddev->curr_resync_completed = raid10_size(mddev, 0, 0)
-- 
2.17.1

[PATCH v6 4/4] md: raid456 add nowait support

From: Vishal Verma <hidden>
Date: 2021-12-21 20:06:51

Returns EAGAIN in case the raid456 driver would block
waiting for situations like:

  - Reshape operation,
  - Discard operation.

Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/raid5.c | 20 ++++++++++++++++++++
 1 file changed, 20 insertions(+)
diff --git a/drivers/md/raid5.c b/drivers/md/raid5.c
index 9c1a5877cf9f..d9647c384820 100644
--- a/drivers/md/raid5.c
+++ b/drivers/md/raid5.c
@@ -5715,6 +5715,11 @@ static void make_discard_request(struct mddev *mddev, struct bio *bi)
 		set_bit(R5_Overlap, &sh->dev[sh->pd_idx].flags);
 		if (test_bit(STRIPE_SYNCING, &sh->state)) {
 			raid5_release_stripe(sh);
+			/* Bail out if REQ_NOWAIT is set */
+			if (bi->bi_opf & REQ_NOWAIT) {
+				bio_wouldblock_error(bi);
+				return;
+			}
 			schedule();
 			goto again;
 		}
@@ -5727,6 +5732,11 @@ static void make_discard_request(struct mddev *mddev, struct bio *bi)
 				set_bit(R5_Overlap, &sh->dev[d].flags);
 				spin_unlock_irq(&sh->stripe_lock);
 				raid5_release_stripe(sh);
+				/* Bail out if REQ_NOWAIT is set */
+				if (bi->bi_opf & REQ_NOWAIT) {
+					bio_wouldblock_error(bi);
+					return;
+				}
 				schedule();
 				goto again;
 			}
@@ -5820,6 +5830,16 @@ static bool raid5_make_request(struct mddev *mddev, struct bio * bi)
 	bi->bi_next = NULL;
 
 	md_account_bio(mddev, &bi);
+	/* Bail out if REQ_NOWAIT is set */
+	if ((bi->bi_opf & REQ_NOWAIT) &&
+	    (conf->reshape_progress != MaxSector) &&
+	    (mddev->reshape_backwards
+	    ? (logical_sector > conf->reshape_progress && logical_sector <= conf->reshape_safe)
+	    : (logical_sector >= conf->reshape_safe && logical_sector < conf->reshape_progress))) {
+		bio_wouldblock_error(bi);
+		return true;
+	}
+
 	prepare_to_wait(&conf->wait_for_overlap, &w, TASK_UNINTERRUPTIBLE);
 	for (; logical_sector < last_sector; logical_sector += RAID5_STRIPE_SECTORS(conf)) {
 		int previous;
-- 
2.17.1

Re: [PATCH v6 4/4] md: raid456 add nowait support

From: John Stoffel <hidden>
Date: 2021-12-21 22:02:47

quoted
quoted
quoted
quoted
"Vishal" == Vishal Verma [off-list ref] writes:
Vishal> Returns EAGAIN in case the raid456 driver would block
Vishal> waiting for situations like:

Vishal>   - Reshape operation,
Vishal>   - Discard operation.

Vishal> Signed-off-by: Vishal Verma [off-list ref]

Are there any performance implications with this patch set?  I didn't
see any discussion in the patch set (v6) and I was just wondering what
this buys us?  Your patch 1/4 talks about using fio as a test, but
there's no mention of whether it's now faster or slower.

John

Re: [PATCH v6 1/4] md: add support for REQ_NOWAIT

From: Jens Axboe <axboe@kernel.dk>
Date: 2021-12-22 16:06:52

On 12/21/21 1:06 PM, Vishal Verma wrote:
commit 021a24460dc2 ("block: add QUEUE_FLAG_NOWAIT") added support
for checking whether a given bdev supports handling of REQ_NOWAIT or not.
Since then commit 6abc49468eea ("dm: add support for REQ_NOWAIT and enable
it for linear target") added support for REQ_NOWAIT for dm. This uses
a similar approach to incorporate REQ_NOWAIT for md based bios.

This patch was tested using t/io_uring tool within FIO. A nvme drive
was partitioned into 2 partitions and a simple raid 0 configuration
/dev/md0 was created.

md0 : active raid0 nvme4n1p1[1] nvme4n1p2[0]
      937423872 blocks super 1.2 512k chunks

Before patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

  38396   38396 pts/2    00:00:00 io_uring
  38396   38397 pts/2    00:00:15 io_uring
  38396   38398 pts/2    00:00:13 iou-wrk-38397

We can see iou-wrk-38397 io worker thread created which gets created
when io_uring sees that the underlying device (/dev/md0 in this case)
doesn't support nowait.

After patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

  38341   38341 pts/2    00:10:22 io_uring
  38341   38342 pts/2    00:10:37 io_uring

After running this patch, we don't see any io worker thread
being created which indicated that io_uring saw that the
underlying device does support nowait. This is the exact behaviour
noticed on a dm device which also supports nowait.

For all the other raid personalities except raid0, we would need
to train pieces which involves make_request fn in order for them
to correctly handle REQ_NOWAIT.
1-4 look fine to me now:

Reviewed-by: Jens Axboe <axboe@kernel.dk>

-- 
Jens Axboe

Re: [PATCH v6 3/4] md: raid10 add nowait support

From: Song Liu <song@kernel.org>
Date: 2021-12-22 23:58:33

On Tue, Dec 21, 2021 at 12:06 PM Vishal Verma [off-list ref] wrote:
quoted hunk
This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

  - Waiting for the barrier,
  - Too many pending I/Os to be queued,
  - Reshape operation,
  - Discard operation.

wait_barrier() and regular_request_wait() fn are modified to return bool
to support error for wait barriers. They returns true in case of wait
or if wait is not required and returns false if wait was required
but not performed to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/raid10.c | 90 +++++++++++++++++++++++++++++++--------------
 1 file changed, 62 insertions(+), 28 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..7ceae00e863e 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,8 +952,9 @@ static void lower_barrier(struct r10conf *conf)
        wake_up(&conf->wait_barrier);
 }

-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
 {
+       bool ret = true;
        spin_lock_irq(&conf->resync_lock);
        if (conf->barrier) {
                struct bio_list *bio_list = current->bio_list;
@@ -968,26 +969,33 @@ static void wait_barrier(struct r10conf *conf)
                 * count down.
                 */
                raid10_log(conf->mddev, "wait barrier");
-               wait_event_lock_irq(conf->wait_barrier,
-                                   !conf->barrier ||
-                                   (atomic_read(&conf->nr_pending) &&
-                                    bio_list &&
-                                    (!bio_list_empty(&bio_list[0]) ||
-                                     !bio_list_empty(&bio_list[1]))) ||
-                                    /* move on if recovery thread is
-                                     * blocked by us
-                                     */
-                                    (conf->mddev->thread->tsk == current &&
-                                     test_bit(MD_RECOVERY_RUNNING,
-                                              &conf->mddev->recovery) &&
-                                     conf->nr_queued > 0),
-                                   conf->resync_lock);
+               /* Return false when nowait flag is set */
+               if (nowait)
+                       ret = false;
+               else
+                       wait_event_lock_irq(conf->wait_barrier,
+                                           !conf->barrier ||
+                                           (atomic_read(&conf->nr_pending) &&
+                                            bio_list &&
+                                            (!bio_list_empty(&bio_list[0]) ||
+                                             !bio_list_empty(&bio_list[1]))) ||
+                                            /* move on if recovery thread is
+                                             * blocked by us
+                                             */
+                                            (conf->mddev->thread->tsk == current &&
+                                             test_bit(MD_RECOVERY_RUNNING,
+                                                      &conf->mddev->recovery) &&
+                                             conf->nr_queued > 0),
+                                           conf->resync_lock);
                conf->nr_waiting--;
                if (!conf->nr_waiting)
                        wake_up(&conf->wait_barrier);
        }
-       atomic_inc(&conf->nr_pending);
+       /* Only increment nr_pending when we wait */
+       if (ret)
+               atomic_inc(&conf->nr_pending);
        spin_unlock_irq(&conf->resync_lock);
+       return ret;
 }

 static void allow_barrier(struct r10conf *conf)
@@ -1098,21 +1106,30 @@ static void raid10_unplug(struct blk_plug_cb *cb, bool from_schedule)
  * currently.
  * 2. If IO spans the reshape position.  Need to wait for reshape to pass.
  */
-static void regular_request_wait(struct mddev *mddev, struct r10conf *conf,
+static bool regular_request_wait(struct mddev *mddev, struct r10conf *conf,
                                 struct bio *bio, sector_t sectors)
This doesn't sound right: regular_request_wait() is called in two
places. But we are
not checking the return value in either of them.

Song
[...]

Re: [PATCH v6 1/4] md: add support for REQ_NOWAIT

From: Song Liu <song@kernel.org>
Date: 2021-12-23 01:22:57

On Tue, Dec 21, 2021 at 12:06 PM Vishal Verma [off-list ref] wrote:
quoted hunk
commit 021a24460dc2 ("block: add QUEUE_FLAG_NOWAIT") added support
for checking whether a given bdev supports handling of REQ_NOWAIT or not.
Since then commit 6abc49468eea ("dm: add support for REQ_NOWAIT and enable
it for linear target") added support for REQ_NOWAIT for dm. This uses
a similar approach to incorporate REQ_NOWAIT for md based bios.

This patch was tested using t/io_uring tool within FIO. A nvme drive
was partitioned into 2 partitions and a simple raid 0 configuration
/dev/md0 was created.

md0 : active raid0 nvme4n1p1[1] nvme4n1p2[0]
      937423872 blocks super 1.2 512k chunks

Before patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

  38396   38396 pts/2    00:00:00 io_uring
  38396   38397 pts/2    00:00:15 io_uring
  38396   38398 pts/2    00:00:13 iou-wrk-38397

We can see iou-wrk-38397 io worker thread created which gets created
when io_uring sees that the underlying device (/dev/md0 in this case)
doesn't support nowait.

After patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

  38341   38341 pts/2    00:10:22 io_uring
  38341   38342 pts/2    00:10:37 io_uring

After running this patch, we don't see any io worker thread
being created which indicated that io_uring saw that the
underlying device does support nowait. This is the exact behaviour
noticed on a dm device which also supports nowait.

For all the other raid personalities except raid0, we would need
to train pieces which involves make_request fn in order for them
to correctly handle REQ_NOWAIT.

Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/md.c | 20 ++++++++++++++++++++
 1 file changed, 20 insertions(+)
diff --git a/drivers/md/md.c b/drivers/md/md.c
index 5111ed966947..ccd296aa9641 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -418,6 +418,11 @@ void md_handle_request(struct mddev *mddev, struct bio *bio)
        rcu_read_lock();
        if (is_suspended(mddev, bio)) {
                DEFINE_WAIT(__wait);
+               /* Bail out if REQ_NOWAIT is set for the bio */
+               if (bio->bi_opf & REQ_NOWAIT) {
We need rcu_read_unlock() here.

Re: [PATCH v6 3/4] md: raid10 add nowait support

From: Song Liu <song@kernel.org>
Date: 2021-12-23 01:47:41

On Tue, Dec 21, 2021 at 12:06 PM Vishal Verma [off-list ref] wrote:
quoted hunk
This adds nowait support to the RAID10 driver. Very similar to
raid1 driver changes. It makes RAID10 driver return with EAGAIN
for situations where it could wait for eg:

  - Waiting for the barrier,
  - Too many pending I/Os to be queued,
  - Reshape operation,
  - Discard operation.

wait_barrier() and regular_request_wait() fn are modified to return bool
to support error for wait barriers. They returns true in case of wait
or if wait is not required and returns false if wait was required
but not performed to support nowait.

Signed-off-by: Vishal Verma <redacted>
---
 drivers/md/raid10.c | 90 +++++++++++++++++++++++++++++++--------------
 1 file changed, 62 insertions(+), 28 deletions(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index dde98f65bd04..7ceae00e863e 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -952,8 +952,9 @@ static void lower_barrier(struct r10conf *conf)
        wake_up(&conf->wait_barrier);
 }

-static void wait_barrier(struct r10conf *conf)
+static bool wait_barrier(struct r10conf *conf, bool nowait)
 {
+       bool ret = true;
        spin_lock_irq(&conf->resync_lock);
        if (conf->barrier) {
                struct bio_list *bio_list = current->bio_list;
@@ -968,26 +969,33 @@ static void wait_barrier(struct r10conf *conf)
                 * count down.
                 */
                raid10_log(conf->mddev, "wait barrier");
-               wait_event_lock_irq(conf->wait_barrier,
-                                   !conf->barrier ||
-                                   (atomic_read(&conf->nr_pending) &&
-                                    bio_list &&
-                                    (!bio_list_empty(&bio_list[0]) ||
-                                     !bio_list_empty(&bio_list[1]))) ||
-                                    /* move on if recovery thread is
-                                     * blocked by us
-                                     */
-                                    (conf->mddev->thread->tsk == current &&
-                                     test_bit(MD_RECOVERY_RUNNING,
-                                              &conf->mddev->recovery) &&
-                                     conf->nr_queued > 0),
-                                   conf->resync_lock);
+               /* Return false when nowait flag is set */
+               if (nowait)
+                       ret = false;
+               else
+                       wait_event_lock_irq(conf->wait_barrier,
+                                           !conf->barrier ||
+                                           (atomic_read(&conf->nr_pending) &&
+                                            bio_list &&
+                                            (!bio_list_empty(&bio_list[0]) ||
+                                             !bio_list_empty(&bio_list[1]))) ||
+                                            /* move on if recovery thread is
+                                             * blocked by us
+                                             */
+                                            (conf->mddev->thread->tsk == current &&
+                                             test_bit(MD_RECOVERY_RUNNING,
+                                                      &conf->mddev->recovery) &&
+                                             conf->nr_queued > 0),
+                                           conf->resync_lock);
                conf->nr_waiting--;
                if (!conf->nr_waiting)
                        wake_up(&conf->wait_barrier);
        }
-       atomic_inc(&conf->nr_pending);
+       /* Only increment nr_pending when we wait */
+       if (ret)
+               atomic_inc(&conf->nr_pending);
        spin_unlock_irq(&conf->resync_lock);
+       return ret;
 }

 static void allow_barrier(struct r10conf *conf)
@@ -1098,21 +1106,30 @@ static void raid10_unplug(struct blk_plug_cb *cb, bool from_schedule)
  * currently.
  * 2. If IO spans the reshape position.  Need to wait for reshape to pass.
  */
-static void regular_request_wait(struct mddev *mddev, struct r10conf *conf,
+static bool regular_request_wait(struct mddev *mddev, struct r10conf *conf,
                                 struct bio *bio, sector_t sectors)
 {
-       wait_barrier(conf);
+       /* Bail out if REQ_NOWAIT is set for the bio */
+       if (!wait_barrier(conf, bio->bi_opf & REQ_NOWAIT)) {
+               bio_wouldblock_error(bio);
+               return false;
+       }
        while (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery) &&
            bio->bi_iter.bi_sector < conf->reshape_progress &&
            bio->bi_iter.bi_sector + sectors > conf->reshape_progress) {
                raid10_log(conf->mddev, "wait reshape");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return false;
+               }
                allow_barrier(conf);
                wait_event(conf->wait_barrier,
                           conf->reshape_progress <= bio->bi_iter.bi_sector ||
                           conf->reshape_progress >= bio->bi_iter.bi_sector +
                           sectors);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
        }
+       return true;
 }

 static void raid10_read_request(struct mddev *mddev, struct bio *bio,
@@ -1179,7 +1196,7 @@ static void raid10_read_request(struct mddev *mddev, struct bio *bio,
                bio_chain(split, bio);
                allow_barrier(conf);
                submit_bio_noacct(bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                bio = split;
                r10_bio->master_bio = bio;
                r10_bio->sectors = max_sectors;
@@ -1338,7 +1355,7 @@ static void wait_blocked_dev(struct mddev *mddev, struct r10bio *r10_bio)
                raid10_log(conf->mddev, "%s wait rdev %d blocked",
                                __func__, blocked_rdev->raid_disk);
                md_wait_for_blocked_rdev(blocked_rdev, mddev);
I think we need more handling here.
quoted hunk
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                goto retry_wait;
        }
 }
@@ -1356,6 +1373,11 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
                                            bio->bi_iter.bi_sector,
                                            bio_end_sector(bio)))) {
                DEFINE_WAIT(w);
+               /* Bail out if REQ_NOWAIT is set for the bio */
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                for (;;) {
                        prepare_to_wait(&conf->wait_barrier,
                                        &w, TASK_IDLE);
@@ -1381,6 +1403,10 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
                              BIT(MD_SB_CHANGE_DEVS) | BIT(MD_SB_CHANGE_PENDING));
                md_wakeup_thread(mddev->thread);
                raid10_log(conf->mddev, "wait reshape metadata");
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                wait_event(mddev->sb_wait,
                           !test_bit(MD_SB_CHANGE_PENDING, &mddev->sb_flags));
@@ -1390,6 +1416,10 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
        if (conf->pending_count >= max_queued_requests) {
                md_wakeup_thread(mddev->thread);
                raid10_log(mddev, "wait queued");
We need the check before logging "wait queued".
quoted hunk
+               if (bio->bi_opf & REQ_NOWAIT) {
+                       bio_wouldblock_error(bio);
+                       return;
+               }
                wait_event(conf->wait_barrier,
                           conf->pending_count < max_queued_requests);
        }
@@ -1482,7 +1512,7 @@ static void raid10_write_request(struct mddev *mddev, struct bio *bio,
                bio_chain(split, bio);
                allow_barrier(conf);
                submit_bio_noacct(bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                bio = split;
                r10_bio->master_bio = bio;
        }
@@ -1607,7 +1637,11 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
        if (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery))
                return -EAGAIN;

-       wait_barrier(conf);
+       if (bio->bi_opf & REQ_NOWAIT) {
+               bio_wouldblock_error(bio);
+               return 0;
Shall we return -EAGAIN here?
quoted hunk
+       }
+       wait_barrier(conf, false);

        /*
         * Check reshape again to avoid reshape happens after checking
@@ -1649,7 +1683,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
                allow_barrier(conf);
                /* Resend the fist split part */
                submit_bio_noacct(split);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
        }
        div_u64_rem(bio_end, stripe_size, &remainder);
        if (remainder) {
@@ -1660,7 +1694,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
                /* Resend the second split part */
                submit_bio_noacct(bio);
                bio = split;
-               wait_barrier(conf);
+               wait_barrier(conf, false);
        }

        bio_start = bio->bi_iter.bi_sector;
@@ -1816,7 +1850,7 @@ static int raid10_handle_discard(struct mddev *mddev, struct bio *bio)
                end_disk_offset += geo->stride;
                atomic_inc(&first_r10bio->remaining);
                raid_end_discard_bio(r10_bio);
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                goto retry_discard;
        }
@@ -2011,7 +2045,7 @@ static void print_conf(struct r10conf *conf)

 static void close_sync(struct r10conf *conf)
 {
-       wait_barrier(conf);
+       wait_barrier(conf, false);
        allow_barrier(conf);

        mempool_exit(&conf->r10buf_pool);
@@ -4819,7 +4853,7 @@ static sector_t reshape_request(struct mddev *mddev, sector_t sector_nr,
        if (need_flush ||
            time_after(jiffies, conf->reshape_checkpoint + 10*HZ)) {
                /* Need to update reshape_position in metadata */
-               wait_barrier(conf);
+               wait_barrier(conf, false);
                mddev->reshape_position = conf->reshape_progress;
                if (mddev->reshape_backwards)
                        mddev->curr_resync_completed = raid10_size(mddev, 0, 0)
--
2.17.1

Re: [PATCH v6 1/4] md: add support for REQ_NOWAIT

From: Song Liu <song@kernel.org>
Date: 2021-12-23 02:57:50

On Tue, Dec 21, 2021 at 12:06 PM Vishal Verma [off-list ref] wrote:
commit 021a24460dc2 ("block: add QUEUE_FLAG_NOWAIT") added support
for checking whether a given bdev supports handling of REQ_NOWAIT or not.
Since then commit 6abc49468eea ("dm: add support for REQ_NOWAIT and enable
it for linear target") added support for REQ_NOWAIT for dm. This uses
a similar approach to incorporate REQ_NOWAIT for md based bios.

This patch was tested using t/io_uring tool within FIO. A nvme drive
was partitioned into 2 partitions and a simple raid 0 configuration
/dev/md0 was created.

md0 : active raid0 nvme4n1p1[1] nvme4n1p2[0]
      937423872 blocks super 1.2 512k chunks

Before patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

  38396   38396 pts/2    00:00:00 io_uring
  38396   38397 pts/2    00:00:15 io_uring
  38396   38398 pts/2    00:00:13 iou-wrk-38397

We can see iou-wrk-38397 io worker thread created which gets created
when io_uring sees that the underlying device (/dev/md0 in this case)
doesn't support nowait.

After patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

  38341   38341 pts/2    00:10:22 io_uring
  38341   38342 pts/2    00:10:37 io_uring

After running this patch, we don't see any io worker thread
being created which indicated that io_uring saw that the
underlying device does support nowait. This is the exact behaviour
noticed on a dm device which also supports nowait.

For all the other raid personalities except raid0, we would need
to train pieces which involves make_request fn in order for them
to correctly handle REQ_NOWAIT.

Signed-off-by: Vishal Verma <redacted>
I have made some changes and applied the set to md-next. However,
I think we don't yet have enough test coverage. Please continue testing
the code and send fixes on top of it. Based on the test results, we will
see whether we can ship it in the next merge window.

Note, md-next branch doesn't have [1], so we need to cherry-pick it
for testing.

Thanks,
Song

[1] a08ed9aae8a3 ("block: fix double bio queue when merging in cached
request path")

Re: [PATCH v6 1/4] md: add support for REQ_NOWAIT

From: Vishal Verma <hidden>
Date: 2021-12-23 03:08:14

On 12/22/21 7:57 PM, Song Liu wrote:
On Tue, Dec 21, 2021 at 12:06 PM Vishal Verma [off-list ref] wrote:
quoted
commit 021a24460dc2 ("block: add QUEUE_FLAG_NOWAIT") added support
for checking whether a given bdev supports handling of REQ_NOWAIT or not.
Since then commit 6abc49468eea ("dm: add support for REQ_NOWAIT and enable
it for linear target") added support for REQ_NOWAIT for dm. This uses
a similar approach to incorporate REQ_NOWAIT for md based bios.

This patch was tested using t/io_uring tool within FIO. A nvme drive
was partitioned into 2 partitions and a simple raid 0 configuration
/dev/md0 was created.

md0 : active raid0 nvme4n1p1[1] nvme4n1p2[0]
       937423872 blocks super 1.2 512k chunks

Before patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

   38396   38396 pts/2    00:00:00 io_uring
   38396   38397 pts/2    00:00:15 io_uring
   38396   38398 pts/2    00:00:13 iou-wrk-38397

We can see iou-wrk-38397 io worker thread created which gets created
when io_uring sees that the underlying device (/dev/md0 in this case)
doesn't support nowait.

After patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

   38341   38341 pts/2    00:10:22 io_uring
   38341   38342 pts/2    00:10:37 io_uring

After running this patch, we don't see any io worker thread
being created which indicated that io_uring saw that the
underlying device does support nowait. This is the exact behaviour
noticed on a dm device which also supports nowait.

For all the other raid personalities except raid0, we would need
to train pieces which involves make_request fn in order for them
to correctly handle REQ_NOWAIT.

Signed-off-by: Vishal Verma <redacted>
I have made some changes and applied the set to md-next. However,
I think we don't yet have enough test coverage. Please continue testing
the code and send fixes on top of it. Based on the test results, we will
see whether we can ship it in the next merge window.

Note, md-next branch doesn't have [1], so we need to cherry-pick it
for testing.

Thanks,
Song

[1] a08ed9aae8a3 ("block: fix double bio queue when merging in cached
request path")
Great, and I agree will continue testing this.

Just saw you already addressing some silly things
I missed in v6. Sorry about that.

Thank you!

Re: [PATCH v6 1/4] md: add support for REQ_NOWAIT

From: Christoph Hellwig <hch@infradead.org>
Date: 2021-12-23 08:36:50

Please post the new series in a new thread.

Re: [PATCH v6 4/4] md: raid456 add nowait support

From: Song Liu <song@kernel.org>
Date: 2021-12-25 02:14:20

On Tue, Dec 21, 2021 at 12:06 PM Vishal Verma [off-list ref] wrote:
Returns EAGAIN in case the raid456 driver would block
waiting for situations like:

  - Reshape operation,
  - Discard operation.

Signed-off-by: Vishal Verma <redacted>
I think we will need the following fix for raid456:

============================ 8< ============================
diff --git i/drivers/md/raid5.c w/drivers/md/raid5.c
index 6ab22f29dacd..55d372ce3300 100644
--- i/drivers/md/raid5.c
+++ w/drivers/md/raid5.c
@@ -5717,6 +5717,7 @@ static void make_discard_request(struct mddev
*mddev, struct bio *bi)
                        raid5_release_stripe(sh);
                        /* Bail out if REQ_NOWAIT is set */
                        if (bi->bi_opf & REQ_NOWAIT) {
+                               finish_wait(&conf->wait_for_overlap, &w);
                                bio_wouldblock_error(bi);
                                return;
                        }
@@ -5734,6 +5735,7 @@ static void make_discard_request(struct mddev
*mddev, struct bio *bi)
                                raid5_release_stripe(sh);
                                /* Bail out if REQ_NOWAIT is set */
                                if (bi->bi_opf & REQ_NOWAIT) {
+
finish_wait(&conf->wait_for_overlap, &w);
                                        bio_wouldblock_error(bi);
                                        return;
                                }
@@ -5829,7 +5831,6 @@ static bool raid5_make_request(struct mddev
*mddev, struct bio * bi)
        last_sector = bio_end_sector(bi);
        bi->bi_next = NULL;

-       md_account_bio(mddev, &bi);
        /* Bail out if REQ_NOWAIT is set */
        if ((bi->bi_opf & REQ_NOWAIT) &&
            (conf->reshape_progress != MaxSector) &&
@@ -5837,9 +5838,11 @@ static bool raid5_make_request(struct mddev
*mddev, struct bio * bi)
            ? (logical_sector > conf->reshape_progress &&
logical_sector <= conf->reshape_safe)
            : (logical_sector >= conf->reshape_safe && logical_sector
< conf->reshape_progress))) {
                bio_wouldblock_error(bi);
+               if (rw == WRITE)
+                       md_write_end(mddev);
                return true;
        }
-
+       md_account_bio(mddev, &bi);
        prepare_to_wait(&conf->wait_for_overlap, &w, TASK_UNINTERRUPTIBLE);
        for (; logical_sector < last_sector; logical_sector +=
RAID5_STRIPE_SECTORS(conf)) {
                int previous;

============================ 8< ============================

Vishal, please try to trigger all these conditions (including raid1,
raid10) and make sure
they work properly.

For example, I triggered raid5 reshape and used something like the
following to make
sure the logic is triggered:
diff --git a/drivers/md/raid5.c b/drivers/md/raid5.c
index 55d372ce3300..e79de48a0027 100644
--- a/drivers/md/raid5.c
+++ b/drivers/md/raid5.c
@@ -5840,6 +5840,11 @@ static bool raid5_make_request(struct mddev
*mddev, struct bio * bi)
                bio_wouldblock_error(bi);
                if (rw == WRITE)
                        md_write_end(mddev);
+               {
+                       static int count = 0;
+                       if (count++ < 10)
+                               pr_info("%s REQ_NOWAIT return\n", __func__);
+               }
                return true;
        }
        md_account_bio(mddev, &bi);

Thanks,
Song

Re: [PATCH v6 1/4] md: add support for REQ_NOWAIT

From: Song Liu <song@kernel.org>
Date: 2022-01-02 00:12:06

On Wed, Dec 22, 2021 at 6:57 PM Song Liu [off-list ref] wrote:
On Tue, Dec 21, 2021 at 12:06 PM Vishal Verma [off-list ref] wrote:
quoted
commit 021a24460dc2 ("block: add QUEUE_FLAG_NOWAIT") added support
for checking whether a given bdev supports handling of REQ_NOWAIT or not.
Since then commit 6abc49468eea ("dm: add support for REQ_NOWAIT and enable
it for linear target") added support for REQ_NOWAIT for dm. This uses
a similar approach to incorporate REQ_NOWAIT for md based bios.

This patch was tested using t/io_uring tool within FIO. A nvme drive
was partitioned into 2 partitions and a simple raid 0 configuration
/dev/md0 was created.

md0 : active raid0 nvme4n1p1[1] nvme4n1p2[0]
      937423872 blocks super 1.2 512k chunks

Before patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

  38396   38396 pts/2    00:00:00 io_uring
  38396   38397 pts/2    00:00:15 io_uring
  38396   38398 pts/2    00:00:13 iou-wrk-38397

We can see iou-wrk-38397 io worker thread created which gets created
when io_uring sees that the underlying device (/dev/md0 in this case)
doesn't support nowait.

After patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

  38341   38341 pts/2    00:10:22 io_uring
  38341   38342 pts/2    00:10:37 io_uring

After running this patch, we don't see any io worker thread
being created which indicated that io_uring saw that the
underlying device does support nowait. This is the exact behaviour
noticed on a dm device which also supports nowait.

For all the other raid personalities except raid0, we would need
to train pieces which involves make_request fn in order for them
to correctly handle REQ_NOWAIT.

Signed-off-by: Vishal Verma <redacted>
I have made some changes and applied the set to md-next. However,
I think we don't yet have enough test coverage. Please continue testing
the code and send fixes on top of it. Based on the test results, we will
see whether we can ship it in the next merge window.

Note, md-next branch doesn't have [1], so we need to cherry-pick it
for testing.
I went through all these changes again and tested many (but not all)
cases. The latest version is available in md-next branch.

Vishal, please run tests on this version and send fixes if anything
is broken.

Thanks,
Song

Re: [PATCH v6 1/4] md: add support for REQ_NOWAIT

From: Vishal Verma <hidden>
Date: 2022-01-02 02:08:56

On 1/1/22 5:11 PM, Song Liu wrote:
On Wed, Dec 22, 2021 at 6:57 PM Song Liu [off-list ref] wrote:
quoted
On Tue, Dec 21, 2021 at 12:06 PM Vishal Verma [off-list ref] wrote:
quoted
commit 021a24460dc2 ("block: add QUEUE_FLAG_NOWAIT") added support
for checking whether a given bdev supports handling of REQ_NOWAIT or not.
Since then commit 6abc49468eea ("dm: add support for REQ_NOWAIT and enable
it for linear target") added support for REQ_NOWAIT for dm. This uses
a similar approach to incorporate REQ_NOWAIT for md based bios.

This patch was tested using t/io_uring tool within FIO. A nvme drive
was partitioned into 2 partitions and a simple raid 0 configuration
/dev/md0 was created.

md0 : active raid0 nvme4n1p1[1] nvme4n1p2[0]
       937423872 blocks super 1.2 512k chunks

Before patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

   38396   38396 pts/2    00:00:00 io_uring
   38396   38397 pts/2    00:00:15 io_uring
   38396   38398 pts/2    00:00:13 iou-wrk-38397

We can see iou-wrk-38397 io worker thread created which gets created
when io_uring sees that the underlying device (/dev/md0 in this case)
doesn't support nowait.

After patch:

$ ./t/io_uring /dev/md0 -p 0 -a 0 -d 1 -r 100

Running top while the above runs:

$ ps -eL | grep $(pidof io_uring)

   38341   38341 pts/2    00:10:22 io_uring
   38341   38342 pts/2    00:10:37 io_uring

After running this patch, we don't see any io worker thread
being created which indicated that io_uring saw that the
underlying device does support nowait. This is the exact behaviour
noticed on a dm device which also supports nowait.

For all the other raid personalities except raid0, we would need
to train pieces which involves make_request fn in order for them
to correctly handle REQ_NOWAIT.

Signed-off-by: Vishal Verma <redacted>
I have made some changes and applied the set to md-next. However,
I think we don't yet have enough test coverage. Please continue testing
the code and send fixes on top of it. Based on the test results, we will
see whether we can ship it in the next merge window.

Note, md-next branch doesn't have [1], so we need to cherry-pick it
for testing.
I went through all these changes again and tested many (but not all)
cases. The latest version is available in md-next branch.

Vishal, please run tests on this version and send fixes if anything
is broken.

Thanks,
Song
Thanks Song. This latest version looks good!
And yes, will report out if I notice any issues or anything.
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help