Neil Brown said to send this message to linux-scsi, so here it is.
Please help.
Thanks,
Guy
On Thursday January 29, bugzilla@watkins-home.com wrote:
As you can see in the log, the write error recovered with auto
reallocation!
As I understand it, this is a normal event with today's disks.
I don't think the disk should have been considered failed.
Comments please?
You need to talk to linux-scsi about this.
The scsi subsystem told the raid subsystem that there was an error, so
the raid subsystem stopped using the device.
If the write error was recovered, scsi shouldn't have reported an
error to raid.
NeilBrown
Thanks,
Guy
The spare disk resynced just fine.,A I never knew for over 24 hours!
This is cool stuff!
Jan 27 12:44:06 watkins kernel: SCSI disk error : host 2 channel 0 id 4
lun
0 return code = 8000002
Jan 27 12:44:06 watkins kernel: Info fld=0x7e5c81, Deferred sd08:71: sense
key Recovered Error
Jan 27 12:44:06 watkins kernel: Additional sense indicates Write error -
recovered with auto reallocation
Jan 27 12:44:06 watkins kernel:,A I/O error: dev 08:71, sector 8280704
Jan 27 12:44:06 watkins kernel: raid5: Disk failure on sdh1, disabling
device. Operation continuing on 13 devices
Jan 27 12:44:06 watkins kernel: md: updating md2 RAID superblock on device
Jan 27 12:44:06 watkins kernel: md: sdc1 [events: 00000009]<6>(write)
sdc1's
sb offset: 17767744
Jan 27 12:44:06 watkins kernel: md: recovery thread got woken up ...
Jan 27 12:44:06 watkins kernel: md2: resyncing spare disk sdc1 to replace
failed disk
Jan 27 12:44:06 watkins kernel: RAID5 conf printout:
Jan 27 12:44:06 watkins kernel:,A --- rd:14 wd:13 fd:1
-
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
-
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Sorry about the re-post, but no comments after almost 2 days.
-----Original Message-----
From: linux-scsi-owner@vger.kernel.org
[mailto:linux-scsi-owner@vger.kernel.org] On Behalf Of Guy
Sent: Thursday, January 29, 2004 12:21 AM
To: linux-scsi@vger.kernel.org
Subject: Recovered disk error caused disk to go offline.
Neil Brown said to send this message to linux-scsi, so here it is.
Please help.
Thanks,
Guy
On Thursday January 29, bugzilla@watkins-home.com wrote:
As you can see in the log, the write error recovered with auto
reallocation!
As I understand it, this is a normal event with today's disks.
I don't think the disk should have been considered failed.
Comments please?
You need to talk to linux-scsi about this.
The scsi subsystem told the raid subsystem that there was an error, so
the raid subsystem stopped using the device.
If the write error was recovered, scsi shouldn't have reported an
error to raid.
NeilBrown
Thanks,
Guy
The spare disk resynced just fine.,A I never knew for over 24 hours!
This is cool stuff!
Jan 27 12:44:06 watkins kernel: SCSI disk error : host 2 channel 0 id 4
lun
0 return code = 8000002
Jan 27 12:44:06 watkins kernel: Info fld=0x7e5c81, Deferred sd08:71: sense
key Recovered Error
Jan 27 12:44:06 watkins kernel: Additional sense indicates Write error -
recovered with auto reallocation
Jan 27 12:44:06 watkins kernel:,A I/O error: dev 08:71, sector 8280704
Jan 27 12:44:06 watkins kernel: raid5: Disk failure on sdh1, disabling
device. Operation continuing on 13 devices
Jan 27 12:44:06 watkins kernel: md: updating md2 RAID superblock on device
Jan 27 12:44:06 watkins kernel: md: sdc1 [events: 00000009]<6>(write)
sdc1's
sb offset: 17767744
Jan 27 12:44:06 watkins kernel: md: recovery thread got woken up ...
Jan 27 12:44:06 watkins kernel: md2: resyncing spare disk sdc1 to replace
failed disk
Jan 27 12:44:06 watkins kernel: RAID5 conf printout:
Jan 27 12:44:06 watkins kernel:,A --- rd:14 wd:13 fd:1
-
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
-
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
-
To unsubscribe from this list: send the line "unsubscribe linux-scsi" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
On Fri, 2004-01-30 at 14:02, Guy wrote:
Sorry about the re-post, but no comments after almost 2 days.
Recovered disc error processing has been in the sd driver for nearly two
years now. What kernel version was this?
James
RedHat 9.0
From uname -a:
Linux watkins 2.4.20-8smp #1 SMP Thu Mar 13 17:45:54 EST 2003 i686 i686 i386
GNU/Linux
Guy
-----Original Message-----
From: linux-scsi-owner@vger.kernel.org
[mailto:linux-scsi-owner@vger.kernel.org] On Behalf Of James Bottomley
Sent: Sunday, February 01, 2004 10:25 AM
To: Guy
Cc: SCSI Mailing List
Subject: RE: Recovered disk error caused disk to go offline.
On Fri, 2004-01-30 at 14:02, Guy wrote:
Sorry about the re-post, but no comments after almost 2 days.
Recovered disc error processing has been in the sd driver for nearly two
years now. What kernel version was this?
James
-
To unsubscribe from this list: send the line "unsubscribe linux-scsi" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
On Sun, 2004-02-01 at 11:10, Guy wrote:
RedHat 9.0
quoted
From uname -a:
Linux watkins 2.4.20-8smp #1 SMP Thu Mar 13 17:45:54 EST 2003 i686 i686 i386
GNU/Linux
Aha, for a vendor kernel you need to file a bugzilla with redhat.
Alternatively, if you can reproduce it with the latest kernel.org 2.4 or
2.6 kernel, we can diagnose it.
James