Thread (2 messages) 2 messages, 2 authors, 2017-05-23

Re: 4.11.2: reshape raid5 -> raid6 atop bcache deadlocks at start on md_attr_store / raid5_make_request

From: NeilBrown <hidden>
Date: 2017-05-23 01:20:04

On Mon, May 22 2017, Nix wrote:
On 22 May 2017, Wols Lists verbalised:
quoted
aiui, provided you're using a v1 superblock, the data offset means there
is spare space on the drives precisely for the purpose (one of then at
least) of keeping a backup. So the reshape will start reshaping into the
spare space and eliminate the need for the backup - the new version of
the stripe will be safely written before the space occupied by the old
stripe is required.
But it's only a few KiB by default! The amount of seeking needed to
reshape with such a small intermediate would be fairly horrific. (It was
bad enough as it was: the reshape of 7TiB took more than two days,
running at under 15MiB/s, though the component drives can all handle
220MiB/s easily. The extra time was spent seeking to and from the
backup, it seems.)
If the space before were "only a few KiB", it wouldn't be used.
You need at least 1 full stripe, typically more.
Current mdadm leaves several megabytes I think.
quoted
Typically, the new layout will move the data offset, pushing the data
block up or down.
Will it? Why? You'd need to move the data offset quite a long way to
make enough space for the 500MiB backup file I saw being created during
my last reshape.

(Though, for me, this is somewhat redundant:

    Data Offset : 262144 sectors

because I set a data offset specifically to allow reshaping without
having to move anything. Is the idea that reshaping that adds data
spindles will move the data offset such that it is (still) on a chunk or
stripe multiple? That's neat, if so, and means I wasted 128MiB on this,
uh, 12TiB array. OK I'm not terribly blown away by this, particularly
given that I'm wasting the same again inside the bcache partition for
the same reason: I'm sure mdadm won't move *that* data offset.)
Data offset is always moved by a multiple of the chunk size.
When I create a 12-device raid5 on 1TB devices, then examine one of them,
it says:

    Data Offset : 262144 sectors
   Unused Space : before=262064 sectors, after=0 sectors
     Chunk Size : 512K

so there is 130Megabytes of space per device, enough for 255 chunks.
When mdadm moves the Data Offset to allow a reshape to happen without a
backup file, it aims to use half the available space.  So it would use
about 60Meg in about 120 chunks or 720Meg total across all devices.
This is more than the 500MiB backup file you saw.

NeilBrown
-- 
NULL && (void)

Attachments

Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help