Crash and rebuild...
From: Neil Bortnak <hidden>
Date: 2006-02-01 16:05:56
Hi guys, So recently I had a hard drive go down with some unusual behaviour I thought I'd report. Since this is a production machine and I can't really replicate it, I've tired to be as detailed about the situation as I can. In a nutshell, after some odd errors (which I think originated in the SATA code) the array degraded. I let it rebuild to a hot spare, but upon reboot it started rebuilding again even though the spare checked out as ok. I let rebuild again, but upon reboot it rebuilt again. This behaviour occured spanned 2.6.12-ck3s and 2.6.15.1 (after the first weird rebuild, I did a kernel upgrade thinking the bug may have been fixed). I ended up having to swap the hot-spare to the old drives position on the SATA controller (i.e. put sdo where sdj used to be). Everything was groovy from then on out. This array is normally dormant (very light duty). I was doing some heavy io when the errors came up, but I think this may have been the result of a bug in the SATA stack, because after the kernel upgrade the drives were a *LOT* quieter. A kernel upgrade shouldn't really do that. In any case, I have included the relevant dmesgs and a --examine and --detail for all drives as soon as the second (2.6.15.1) weird rebuild started. If you need any more info I'll do my best to provide it, but I thought I should at least report this. Neil P.S. I realise in the files below, md3 is also degraded. sdg died after the first rebuild but before the reboot, due to me running my program again (md5s of all files on the array) and subsequently tripping the alleged SATA bug. --- dmesg below --detail and --examine files attached along with an unhappy mdstat Random Info: CPU - AMD Athlon(TM) XP 2500+ Memory: 512M SATA cards: 3 x SATA 3114 (md3 and md4) Drives: 6 x Maxtor 6Y200M0 (md3) and 6 x Maxtor 7L300S0 (md4) In this dmesg, sdo should *not* have been booted out of the array. md: autorun ... md: considering sdo1 ... md: adding sdo1 ... md: adding sdn1 ... md: adding sdm1 ... md: adding sdl1 ... md: adding sdk1 ... md: adding sdj1 ... md: adding sdi1 ... md: sdh1 has different UUID to sdo1 md: sdg1 has different UUID to sdo1 md: sdf1 has different UUID to sdo1 md: sde1 has different UUID to sdo1 md: sdd1 has different UUID to sdo1 md: sdc1 has different UUID to sdo1 md: sdb3 has different UUID to sdo1 md: sdb2 has different UUID to sdo1 md: sdb1 has different UUID to sdo1 md: sda3 has different UUID to sdo1 md: sda2 has different UUID to sdo1 md: sda1 has different UUID to sdo1 devfs_mk_dev: could not append to parent for md/4 md: created md4 md: bind<sdi1> md: bind<sdj1> md: bind<sdk1> md: bind<sdl1> md: bind<sdm1> md: bind<sdn1> md: export_rdev(sdo1) md: running: <sdn1><sdm1><sdl1><sdk1><sdj1><sdi1> md: kicking non-fresh sdj1 from array! md: unbind<sdj1> md: export_rdev(sdj1) raid5: device sdn1 operational as raid disk 5 raid5: device sdm1 operational as raid disk 0 raid5: device sdl1 operational as raid disk 1 raid5: device sdk1 operational as raid disk 2 raid5: device sdi1 operational as raid disk 4 raid5: allocated 6290kB for md4 raid5: raid level 5 set md4 active with 5 out of 6 devices, algorithm 2 --- rd:6 wd:5 fd:1 disk 0, o:1, dev:sdm1 disk 1, o:1, dev:sdl1 disk 2, o:1, dev:sdk1 disk 4, o:1, dev:sdi1 disk 5, o:1, dev:sdn1
Attachments
- crash-mdadm-detail.txt [text/plain] 4294 bytes · preview
- crash-mdadm-examine.txt [text/plain] 15257 bytes · preview
- crash-mdstat.txt [text/plain] 603 bytes · preview