Thread (7 messages) 7 messages, 5 authors, 2009-08-26

Re: Raid 5 - not clean and then a failure.

From: Jon Hardcastle <hidden>
Date: 2009-08-26 14:19:51

Can a bitmap be easily removed? I might give it ago if it can.

I am never sure how thorough these checks are. Are they read/write, or just read? for example. I make of point of doing read/write badblocks checks with e2fck -cc when I do run them (not the automatic ones tho - dunno how) but that only checks that partition, which is on LVM, which is on RAID so WHO KNOWS what underlying drives are being checked.

I have before now, dismantled the array and run read/write badblocks directly on the constituent drives so at least smart is aware of them and although i aim to do this once every six months, I think I have actually done it only 1nce in the 2 year life of the array.

-----------------------
N: Jon Hardcastle
E: Jon@eHardcastle.com
'Do not worry about tomorrow, for tomorrow will bring worries of its own.'
-----------------------

--- On Wed, 26/8/09, Ryan Wagoner <rswagoner@gmail.com> wrote:
From: Ryan Wagoner <redacted>
Subject: Re: Raid 5 - not clean and then a failure.
To: "Goswin von Brederlow" <redacted>
Cc: Jon@ehardcastle.com, linux-raid@vger.kernel.org
Date: Wednesday, 26 August, 2009, 3:14 PM
Wouldn't weekly RAID consistency
checks reveal a bad block before you
had a failure that required the need to do a full resync?
It only
takes 3 hours to resync my 3 x 1TB drives and having a
bitmap would
reduce the performance. I've never had to have a resync in
the year
I've had the array up. I just wonder if the performance
drawback is
worth having the bitmap to save a possible resync once
every couple
years. Or are the RAID consistency checks not reliable
enough to
prevent more errors during a resync?

Ryan

On Wed, Aug 26, 2009 at 7:18 AM, Goswin von Brederlow[off-list ref]
wrote:
quoted
Jon Hardcastle [off-list ref]
writes:
quoted
quoted
Guys,

I have been having some problems with my arrays
that I think i have nailed down to a pci controller (well I
say that - it is always the drives connected to *a*
controller but I have tried 2!) anyway the latest saga is i
was trying some new kernel options last night - which didn't
work.
quoted
quoted
But when i booted up again this morning it said
one of the drives was in an inconsistent state (not sure of
the *exact* error message). I then kicked off an add of the
drive and it started syncing. It got about 5% in and then
the second drive in on that controller complained and the
array failed.
quoted
quoted
Is there any hope for my data? If i get a good
controller in there will the resync continue? can I try and
tell it to assume the drives are good (which they ought to
be)?
quoted
quoted
Please help!
The inconsistency is probably just a block here or
there and I'm
quoted
assuming none of your drives actualy failed. So
99.9999% of your data
quoted
should be there. Just rebooting might actualy just get
your raid back
quoted
(to syncing). If not then you have to force reassembly
from the drives
quoted
with the newest serials. That will give you some data
corruption,
quoted
whatever was writing when the controler gave errors.
Worst case you
quoted
have to recreate the raid with --assume-clean.

I recommend adding a bitmap to the raid. That way a
wrongfully failed
quoted
drive can be resynced in a matter of minutes instead
of hours or
quoted
days. Makes it way less likely another error occurs
during resync.
quoted
MfG
       Goswin
--
To unsubscribe from this list: send the line
"unsubscribe linux-raid" in
quoted
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
--
To unsubscribe from this list: send the line "unsubscribe
linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

      
--
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help