[PATCH] Online RAID-5 resizing

DORMANTno replies

From: Steinar H. Gunderson <hidden>
Date: 2005-09-20 14:33:46

(Please Cc me on any replies, I'm not subscribed)

Hi,

Attached is a patch (against 2.6.12) for adding online RAID-5 resize
capabilities to Linux' RAID code. It needs to changes to mdadm (I've only
tested with mdadm 1.12.0, though), you can just do

  mdadm --add /dev/md1 /dev/hd[eg]1
  mdadm --grow /dev/md1 -n 4

and it will restripe /dev/md1; you can still use the volume just fine
during the expand process. (cat /proc/mdstat to get the progress; it will
look like a regular sync, and when the restripe is done the volume will
suddenly get larger and do a regular sync of the new parts.)

The patch is quite rough -- it's my first trip ever into the md code, the
block layer or really kernel code in general, so expect subtle race
conditions and problems here and there. :-) That being said, it seems to be
quite stable on my (SMP) test system now -- I would really take backups
before testing it, though! You have been warned :-)

Things still to do, off the top of my head:

- It's RAID-5 only; I don't really use RAID-0, and RAID-6 would probably be
  more complex.
- It supports only growing, not shrinking. (Not sure if I really care about
  fixing this one.)
- It leaks memory; it doesn't properly free up the old stripes etc. at the
  end of the resize. (This also makes it impossible to do a grow and then
  another grow without stopping and starting the volumes.)
- There is absolutely no crash recovery -- this shouldn't be so hard to do
  (just update the superblock every time, with some progress meter, and
  restart from that spot in case of a crash), but I have no knowledge of the
  on-disk superblock format at all, so some help would be appreciated here.
  Also, I'm not really sure what happens if it encounters a bad block during
  the restripe.
- It's quite slow; on my test system with old IDE disks, it achieves about
  1MB/sec. One could probably make a speed/memory tradeoff here, and move
  more chunks at a time instead of just one by one; I'm a bit concerned
  about the implications of the kernel allocating something like 64MB in one
  go, though :-)
  
Comments, patches, fixes etc. would be greatly appreciated. (Again, remember
to Cc me, I'm not on the list.)

/* Steinar */
-- 
Homepage: http://www.sesse.net/

Attachments

Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help