Re: Root Drive Mirroring and LVM.

7 messages, 6 authors, 2004-02-06 · open the first message on its own page

Re: Root Drive Mirroring and LVM.

From: Neil Brown <hidden>
Date: 2004-02-04 23:07:32

On Wednesday February 4, linas@austin.ibm.com wrote:
On Wed, Feb 04, 2004 at 05:32:01PM -0500, Mark Hahn wrote:
quoted
quoted
So what's wrong with autodetect, again?
my guess is: in-kernel autodetect is the problem.  
out-of-kernel detection can be much smarter, 
and can be more easily tested/replaced.
Hm, yes, that makes sense.

Good :-)

Just to flesh out my thoughts a bit more:

 If the root filesystem is on an MD array, then I see the process of
 assembling that md array as quite similar to the process of finding
 the device for the root filesystem.

 We don't expect the kernel, or anyone else, to scan all devices
 looking for something that looks like a root filesystem, and loading
 that.  Rather we tell the kernel or boot loader exactly where to find
 the root filesystem.  And if the root filesystem moves, we get to
 explicitly tell the boot loader where it is (root=/dev/hdc1 or
 whatever).
 Assembling the root device should be handled the same way.  We tell
 the boot loader/kernel where to expect it, but can over-ride that if
 we need to:
   md=0,/dev/hdc1,/dev/hde4

 All other md arrays can, and so should, be assembled by code running
 out of the root filesystem.   This could be some program that
 assembles anything it finds after scanning all devices, or something
 a bit more focused, but it should be controllable by the sysadmin.

 It is true that in-kernel auto-detect can be controlled by fiddling
 with partition types, but the problem is that it runs *before* the
 root filesystem is mounted and so could conceivably confuse the
 assembly of the root device (if e.g. you plugged in some other device
 that also claimed to be part of /dev/md0, and it got scanned before
 your real root device).

NeilBrown

badblock handling

From: Donghui Wen <hidden>
Date: 2004-02-05 04:45:36

Hi,
     I am running a server with Linux software-raid (3ware controller). But
from time to time,
a disk is kicked out by md. This will happen when 3ware card reports a
unrecovered read error
for a sector. But when I run raidhotadd, the disk can added back to raid
with no problem.
     So my questions are:
        (1) Will md  kick out one disk if it find out ONE bad block?
        (2) Is it possible to set up a threshold, only the amount of bad
blocks pass this threshold,
            the disk will be kicked out.
        (3) Is it possible to remap the bad blocks to some spare blocks
automatically on the fly?

Thanks!

Donghui

RE: badblock handling

From: Guy <hidden>
Date: 2004-02-05 05:26:46

I have had the same problems.  Once there is a bad block on the disk, you
can't read it.  However, if you overwrite the bad block with new data (or
the same data) the disk drive will relocate the bad block to a spare block.
Once relocated, the disk is good again.  This auto relocation of bad blocks
(or sectors) has been around since IDE disks.  Not sure all IDE disk, but
all that I know about.

What I think should happen is:
If Read error,
  re-create missing data,
  re-write the data to the "bad" disk.
  If write fails,
    fail the disk,
  else,
    go on with life.

In the past 2 years I have had to remove and re-add failed disks about 5-10
times.  I never had to replace a bad disk.

If something like the above could be implemented, it would save a lot of
effort.

Also, if the bad block table is almost full, then fail the drive regardless.
Maybe a threshold that can be set.  Like 90% full.

Guy

-----Original Message-----
From: linux-raid-owner@vger.kernel.org
[mailto:linux-raid-owner@vger.kernel.org] On Behalf Of Donghui Wen
Sent: Wednesday, February 04, 2004 11:46 PM
To: linux-raid@vger.kernel.org
Cc: Neil Brown
Subject: badblock handling

Hi,
     I am running a server with Linux software-raid (3ware controller). But
from time to time,
a disk is kicked out by md. This will happen when 3ware card reports a
unrecovered read error
for a sector. But when I run raidhotadd, the disk can added back to raid
with no problem.
     So my questions are:
        (1) Will md  kick out one disk if it find out ONE bad block?
        (2) Is it possible to set up a threshold, only the amount of bad
blocks pass this threshold,
            the disk will be kicked out.
        (3) Is it possible to remap the bad blocks to some spare blocks
automatically on the fly?

Thanks!

Donghui

-
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

Re: badblock handling

From: Martin Dohmen <hidden>
Date: 2004-02-05 07:35:26

I have seen the same problem, but I have also noticed that if I use the 3ware cards own raid5 instead of md things work diffrently, the only thing that seems to happen when a disk develops a bad block is this in the log:

3w-xxxx: scsi0: AEN: ERROR: Drive error: Port #6.
3w-xxxx: scsi0: AEN: WARNING: Sector repair occurred: Port #6.

I do not now if 3ware sets some block asside when the raid is created for use when errors occure or if they justs writes the block agan and the disk does the realocation by it self.

The downside of using 3ware raid5 is that performance is quite much lower than with md, with 8 disks I get 90MB/s with md and 40MB/s with 3ware in writespeed.


--On den 4 februari 2004 20:45 -0800 Donghui Wen [off-list ref] wrote:
Hi,
     I am running a server with Linux software-raid (3ware controller).
But from time to time,
a disk is kicked out by md. This will happen when 3ware card reports a
unrecovered read error
for a sector. But when I run raidhotadd, the disk can added back to raid
with no problem.
     So my questions are:
        (1) Will md  kick out one disk if it find out ONE bad block?
        (2) Is it possible to set up a threshold, only the amount of bad
blocks pass this threshold,
            the disk will be kicked out.
        (3) Is it possible to remap the bad blocks to some spare blocks
automatically on the fly?

Thanks!

Donghui

-
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html


Med vänlig hälsning
TRIPNET AB

Martin Dohmen
________________________________________________________________________
Tripnet AB                Besöksadress:      Telefon:  031-725 25 00
Box 5071                  Åvägen 42          Telefax:  031-725 25 01
402 22  GÖTEBORG          GÖTEBORG           Direkt:   031-725 25 11
http://www.tripnet.se     dohmen@tripnet.se  Mobil:    0733-58 25 11

-
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

Re: Root Drive Mirroring and LVM.

From: Joe Pruett <hidden>
Date: 2004-02-05 18:04:26

On Thu, 5 Feb 2004, Neil Brown wrote:
 We don't expect the kernel, or anyone else, to scan all devices
 looking for something that looks like a root filesystem, and loading
 that.  Rather we tell the kernel or boot loader exactly where to find
 the root filesystem.  And if the root filesystem moves, we get to
 explicitly tell the boot loader where it is (root=/dev/hdc1 or
 whatever).
what about:
root=LABEL=/

Re: badblock handling

From: Donghui Wen <hidden>
Date: 2004-02-06 04:43:47

I read some posts before says 3ware's hardware raid is doing bad block
remapping, which is transparent to file system. I was wondering if md is
doing
the same way.

I am trying hardware raid 1+0 now, even it will lost half of the disk space,
but if it is stable and fast, I will go for it.

Donghui

----- Original Message ----- 
From: "Martin Dohmen" <redacted>
To: "'Donghui Wen'" <redacted>;
[off-list ref]
Sent: Wednesday, February 04, 2004 11:35 PM
Subject: Re: badblock handling


I have seen the same problem, but I have also noticed that if I use the
3ware cards own raid5 instead of md things work diffrently, the only thing
that seems to happen when a disk develops a bad block is this in the log:

3w-xxxx: scsi0: AEN: ERROR: Drive error: Port #6.
3w-xxxx: scsi0: AEN: WARNING: Sector repair occurred: Port #6.

I do not now if 3ware sets some block asside when the raid is created for
use when errors occure or if they justs writes the block agan and the disk
does the realocation by it self.

The downside of using 3ware raid5 is that performance is quite much lower
than with md, with 8 disks I get 90MB/s with md and 40MB/s with 3ware in
writespeed.


--On den 4 februari 2004 20:45 -0800 Donghui Wen
[off-list ref] wrote:
Hi,
     I am running a server with Linux software-raid (3ware controller).
But from time to time,
a disk is kicked out by md. This will happen when 3ware card reports a
unrecovered read error
for a sector. But when I run raidhotadd, the disk can added back to raid
with no problem.
     So my questions are:
        (1) Will md  kick out one disk if it find out ONE bad block?
        (2) Is it possible to set up a threshold, only the amount of bad
blocks pass this threshold,
            the disk will be kicked out.
        (3) Is it possible to remap the bad blocks to some spare blocks
automatically on the fly?

Thanks!

Donghui

-
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html


Med vänlig hälsning
TRIPNET AB

Martin Dohmen
________________________________________________________________________
Tripnet AB                Besöksadress:      Telefon:  031-725 25 00
Box 5071                  Åvägen 42          Telefax:  031-725 25 01
402 22  GÖTEBORG          GÖTEBORG           Direkt:   031-725 25 11
http://www.tripnet.se     dohmen@tripnet.se  Mobil:    0733-58 25 11

-
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

-
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

Re: badblock handling

From: Holger Kiehl <hidden>
Date: 2004-02-06 08:54:33

On Wed, 4 Feb 2004, Donghui Wen wrote:
Hi,
     I am running a server with Linux software-raid (3ware controller). But
from time to time,
a disk is kicked out by md. This will happen when 3ware card reports a
unrecovered read error
for a sector. But when I run raidhotadd, the disk can added back to raid
with no problem.
     So my questions are:
        (1) Will md  kick out one disk if it find out ONE bad block?
        (2) Is it possible to set up a threshold, only the amount of bad
blocks pass this threshold,
            the disk will be kicked out.
        (3) Is it possible to remap the bad blocks to some spare blocks
automatically on the fly?
I am not an expert on this and I am only guessing, so please someone
correct me if I say something wrong.

The md block device gets an error from the underneath scsi/ide/usb/firewire/
network layer. It does not tell exactly what type of error has occured,
the best thing for md to do is to kick out this device.

Holger
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help