Re: Requesting replace mode for changing a disk

From: Sandeep K Sinha <hidden>
Date: 2009-05-13 04:08:44

"Leslie Rhorer" [off-list ref] writes:
quoted
This is one of many things proposed occasionally here, no real
objection, sometimes loud support, but no one actually *does* the code.
	At the risk of being a metoo, I would really love this feature.
quoted
You have described the problem exactly, and the solution is still to do
it manually. But you don't need to fail the drive long term, if you can
stop the array for a few moments. You stop the array, remove the suspect
drive,
Um, how, exactly?  That is to say, after stopping the array, how does one
remove the drive?  From the next step in your suggestion, it doesn't seem
tome you are talking about physically removing the drive, so how does one
remove a drive from a stopped array for this purpose?  I didn't think that
either

	mdadm -r <drive> <array>
or

	mdadm -f <drive> <array>

could be used on a stopped array.  Am I mistaken?
quoted
create a raid1 of the suspect drive marked write-mostly and the
new spare,
But doesn't creating the array with the drive wipe the contents?  If so, it
doesn't seem to me this provides much redundancy.
quoted
then add the raid1 in place of the suspect drive.
Before starting the array?  If so, how?  Or should one do an assemble
including the newly minted RAID1?  I thought mdadm would take the newly
added drive to be blank, even if it isn't.
quoted
For any
chunks present on the new drive the reads will go there, reducing
Huh?  Are you saying any read which finds one chunk missing will
automatically write back the missing data (doing a spot rebuild), or
something else?
Yes, this is one of concepts that used while using a mirrored configuration.
I am not sure though, that even md has it. Even a read fails on one of
the part of the mirror
then it is served from the mirror. Also, a write is sent back to keep
the mirror in sync.
This should ideally happen, as if a read fails on the device which is
being mirrored, you are
aware that there is something wrong and you should try to recover from
it silently.

quoted
access, while data is copied from the old to the new in resync, and
See my query above.  It seems to me you are saying the RAID1 can be created
without wiping the drive.
quoted
writes still go to the old suspect drive so if the new drive fails you
are no worse off.
I think I would expect the old drive to be more likely to fail than the new.
quoted
When the raid1 is clean you stop the main array and
back the suspect drive out.
OK, basically the same question.  How does one disassemble the RAID1 array
without wiping the data on the new drive?
I think he ment this:

mdadm --stop /dev/md0
mdadm --build /dev/md9 --chunk=64k --level=1 --raid-devices=2 /dev/suspect /dev/new
mdadm --assemble /dev/md0 /dev/md9 /dev/other ...
MfG
       Goswin
Now the point here is that is it so simple to offline an array and
just assimilate it back with a different raid level.
IMO, may be on desktop environment it is acceptable but not when
deployed on a larger scale.

What I believe here is that there should be an option to fail to a
disk gradually. This can be invoked the moment you start getting to
know that your disk is faulty and is soon going to be blown away.
The first approach is to start copying the contents of that disk to a
spare disk. And once completed, the faulty disk should be replaced
with the newly constructed disk.

In the course of construction, write must go on both the devices and
read must be served from the original faulty disk.
I am not sure whether we have a similar mechanism implemented in md or
not, but this will sound better for a long term goal.

If not I believe failing should be provided with this extra
functionality of "immediate" or "gradual" failure. As both should
ideally start reconstruction.

Comments please.


--
Regards,
Sandeep.






“To learn is to change. Education is a process that changes the learner.”
--
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help