Thread (18 messages) 18 messages, 4 authors, 12d ago

Re: [PATCH v6 1/2] md: Don't set MD_BROKEN for RAID1 and RAID10 when using FailFast

From: Martin Wilck <hidden>
Date: 2026-09-24 17:59:38
Also in: lkml

On Thu, 2026-09-24 at 16:27 +0000, Kenta Akagi wrote:

On 2026/09/23 17:29, Martin Wilck wrote:
quoted
Hi Kenta,

On Wed, 2026-09-23 at 04:08 +0000, Kenta Akagi wrote:
quoted

Hi Martin,

I have not given up on it, but I have not managed to post v7 yet.
I still think failfast should be usable even in setups like that.

It has been a while, but I intend to resume work on it.
My thinking is that, in the fastfail case, code to prevent total
failure could be placed in the end-IO code code path, e.g. by
attempting a retry directly from the md layer when the last rdev
fails,
instead of setting failing the device. But I haven't thought it
through.
Hi Martin,

I may be misunderstanding your suggestion, but Neil's original
failfast 
implementation already retries a failed failfast I/O to the last
rdev. 
But a later change introduced a regression, which I intend to fix.
Ah OK, I wasn't aware of that.
So the sequence should be:

1. The failfast bios to all mirrored rdevs fail.
2. The first rdev is marked faulty because its bio failed.
3. The other rdev is now the last, so it is not marked faulty.
4. Since the last rdev remains usable, its bio error handler 
   retries the I/O without failfast.
Yes, that makes sense to me.

Martin

-- 
Dr. Martin Wilck [off-list ref]
SUSE Software Solutions Germany GmbH, Frankenstr. 146, 90461 Nürnberg,
Germany
Geschäftsführer: Stefan Gaiser, Jochen Jaser, Abhinav Puri (HRB
36809,AG Nürnberg)
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help