Re: [PATCH v2 0/4] btrfs: sysfs: set / query btrfs stripe size
From: Josef Bacik <josef@toxicpanda.com>
Date: 2021-10-28 13:43:55
On Wed, Oct 27, 2021 at 01:14:37PM -0700, Stefan Roesch wrote:
Motivation:
The btrfs allocator is currently not ideal for all workloads. It tends
to suffer from overallocating data block groups and underallocating
metadata block groups. This results in filesystems becoming read-only
even though there is plenty of "free" space.
This is naturally confusing and distressing to users.
Patches:
1) Store the stripe and chunk size in the btrfs_space_info structure
2) Add a sysfs entry to expose the above information
3) Add a sysfs entry to force a space allocation
4) Increase the default size of the metadata chunk allocation to 5GB
for volumes greater than 50GB.
Testing:
A new test is being added to the xfstest suite. For reference the
corresponding patch has the title:
[PATCH] btrfs: Test chunk allocation with different sizes
In addition also manual testing has been performed.
- Run xfstests with the changes and the new test. It does not
show new diffs.
- Test with storage devices 10G, 20G, 30G, 50G, 60G
- Default allocation
- Increase of chunk size
- If the stripe size is > the free space, it allocates
free space - 1MB. The 1MB is left as free space.
- If the device has a storage size > 50G, it uses a 5GB
chunk size for new allocations.
Stefan Roesch (4):
btrfs: store stripe size and chunk size in space-info struct.
btrfs: expose stripe and chunk size in sysfs.
btrfs: add force_chunk_alloc sysfs entry to force allocation
btrfs: increase metadata alloc size to 5GB for volumes > 50GBSorry, I had this thought previously but it got lost when I started doing the actual code review. We have conflated stripe size and chunk size here, and unfortunately "stripe size" means different things to different people. What you are actually trying to do here is to allow us to allocate a larger logical chunk size. In terms of how this works out in the code you are changing the correct thing, generally the stripe_size is what dictates the actual block group chunk size we end up with at the end. But this is sort of confusing when it comes to the interface, because people are going to think it means something different. Instead we should name the sysfs file chunk_size, and then keep the code you have the way it is, just with the new name. That way it's clear to the user that they're changing how large of a chunk we're allocating at any given time. Make that change, and I have a few other code comments, and then that should be good. Thanks, Josef