Crash and rebuild...

From: Neil Bortnak <hidden>
Date: 2006-02-01 16:05:56

Hi guys,

So recently I had a hard drive go down with some unusual behaviour I
thought I'd report. Since this is a production machine and I can't
really replicate it, I've tired to be as detailed about the situation as
I can.

In a nutshell, after some odd errors (which I think originated in the
SATA code) the array degraded. I let it rebuild to a hot spare, but upon
reboot it started rebuilding again even though the spare checked out as
ok. I let rebuild again, but upon reboot it rebuilt again. This
behaviour occured spanned 2.6.12-ck3s and 2.6.15.1 (after the first
weird rebuild, I did a kernel upgrade thinking the bug may have been
fixed). I ended up having to swap the hot-spare to the old drives
position on the SATA controller (i.e. put sdo where sdj used to be).
Everything was groovy from then on out.

This array is normally dormant (very light duty). I was doing some heavy
io when the errors came up, but I think this may have been the result of
a bug in the SATA stack, because after the kernel upgrade the drives
were a *LOT* quieter. A kernel upgrade shouldn't really do that.

In any case, I have included the relevant dmesgs and a --examine and
--detail for all drives as soon as the second (2.6.15.1) weird rebuild
started. If you need any more info I'll do my best to provide it, but I
thought I should at least report this.

Neil

P.S. I realise in the files below, md3 is also degraded. sdg died after
the first rebuild but before the reboot, due to me running my program
again (md5s of all files on the array) and subsequently tripping the
alleged SATA bug.

---

dmesg below
--detail and --examine files attached along with an unhappy mdstat

Random Info:
CPU - AMD Athlon(TM) XP 2500+
Memory: 512M
SATA cards: 3 x SATA 3114 (md3 and md4)
Drives: 6 x Maxtor 6Y200M0 (md3) and 6 x Maxtor 7L300S0 (md4)


In this dmesg, sdo should *not* have been booted out of the array.

md: autorun ...
md: considering sdo1 ...
md:  adding sdo1 ...
md:  adding sdn1 ...
md:  adding sdm1 ...
md:  adding sdl1 ...
md:  adding sdk1 ...
md:  adding sdj1 ...
md:  adding sdi1 ...
md: sdh1 has different UUID to sdo1
md: sdg1 has different UUID to sdo1
md: sdf1 has different UUID to sdo1
md: sde1 has different UUID to sdo1
md: sdd1 has different UUID to sdo1
md: sdc1 has different UUID to sdo1
md: sdb3 has different UUID to sdo1
md: sdb2 has different UUID to sdo1
md: sdb1 has different UUID to sdo1
md: sda3 has different UUID to sdo1
md: sda2 has different UUID to sdo1
md: sda1 has different UUID to sdo1
devfs_mk_dev: could not append to parent for md/4
md: created md4
md: bind<sdi1>
md: bind<sdj1>
md: bind<sdk1>
md: bind<sdl1>
md: bind<sdm1>
md: bind<sdn1>
md: export_rdev(sdo1)
md: running: <sdn1><sdm1><sdl1><sdk1><sdj1><sdi1>
md: kicking non-fresh sdj1 from array!
md: unbind<sdj1>
md: export_rdev(sdj1)
raid5: device sdn1 operational as raid disk 5
raid5: device sdm1 operational as raid disk 0
raid5: device sdl1 operational as raid disk 1
raid5: device sdk1 operational as raid disk 2
raid5: device sdi1 operational as raid disk 4
raid5: allocated 6290kB for md4
raid5: raid level 5 set md4 active with 5 out of 6 devices, algorithm 2
 --- rd:6 wd:5 fd:1
 disk 0, o:1, dev:sdm1
 disk 1, o:1, dev:sdl1
 disk 2, o:1, dev:sdk1
 disk 4, o:1, dev:sdi1
 disk 5, o:1, dev:sdn1

Attachments

Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help