Not yet ready for the integration. As I need to introduce
-o no_read_mirror_policy instead of -o read_mirror_policy=-<devid>
to reset the policy as in 3/3. But I am sending this early so that
we can use it for btrfs/161 in the ML, and this patch-set is stable
enough for the testing.
Anand Jain (3):
btrfs: add mount option read_mirror_policy
btrfs: add read_mirror_policy parameter devid
btrfs: read_mirror_policy ability to reset
fs/btrfs/ctree.h | 2 ++
fs/btrfs/super.c | 44 ++++++++++++++++++++++++++++++++++++++++++++
fs/btrfs/volumes.c | 18 +++++++++++++++++-
fs/btrfs/volumes.h | 7 +++++++
4 files changed, 70 insertions(+), 1 deletion(-)
--
2.7.0
In case of RAID1 and RAID10 devices are mirror-ed, a read IO can
pick any device for reading. This choice of picking a device for
reading should be configurable. In short not one policy would
satisfy all types of workload and configs.
So before we add more policies, this patch-set makes existing
$pid policy configurable from the mount option.
For example..
mount -o read_mirror_policy=pid (which is also default)
Signed-off-by: Anand Jain <redacted>
---
fs/btrfs/ctree.h | 2 ++
fs/btrfs/super.c | 10 ++++++++++
fs/btrfs/volumes.c | 8 +++++++-
fs/btrfs/volumes.h | 5 +++++
4 files changed, 24 insertions(+), 1 deletion(-)
Adds the mount option:
mount -o read_mirror_policy=<devid>
To set the devid of the device which should be used for read. That
means all the normal reads will go to that particular device only.
This also helps testing and gives a better control for the test
scripts including mount context reads.
Signed-off-by: Anand Jain <redacted>
---
fs/btrfs/super.c | 21 +++++++++++++++++++++
fs/btrfs/volumes.c | 10 ++++++++++
fs/btrfs/volumes.h | 2 ++
3 files changed, 33 insertions(+)
On Wed, May 16, 2018 at 06:03:56PM +0800, Anand Jain wrote:
quoted
Not yet ready for the integration. As I need to introduce
-o no_read_mirror_policy instead of -o read_mirror_policy=-<devid>
Mount option is mostly likely not the right interface for setting such
options, as usual.
I am ok to make it ioctl for the final. What do you think?
But to reproduce the bug posted in
Btrfs: fix the corruption by reading stale btree blocks
It needs to be a mount option, as randomly the pid can
still pick the disk specified in the mount option.
-Anand
--
To unsubscribe from this list: send the line "unsubscribe linux-btrfs" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
From: Austin S. Hemmelgarn <hidden> Date: 2018-05-17 12:25:46
On 2018-05-16 22:32, Anand Jain wrote:
On 05/17/2018 06:35 AM, David Sterba wrote:
quoted
On Wed, May 16, 2018 at 06:03:56PM +0800, Anand Jain wrote:
quoted
Not yet ready for the integration. As I need to introduce
-o no_read_mirror_policy instead of -o read_mirror_policy=-<devid>
Mount option is mostly likely not the right interface for setting such
options, as usual.
I am ok to make it ioctl for the final. What do you think?
But to reproduce the bug posted in
Btrfs: fix the corruption by reading stale btree blocks
It needs to be a mount option, as randomly the pid can
still pick the disk specified in the mount option.
Personally, I'd vote for filesystem property (thus handled through the standard `btrfs property` command) that can be overridden by a mount option. With that approach, no new tool (or change to an existing tool) would be needed, existing volumes could be converted to use it in a backwards compatible manner (old kernels would just ignore the property), and you could still have the behavior you want in tests (and in theory it could easily be adapted to be a per-subvolume setting if we ever get per-subvolume chunk profile support).
Of course, I'd actually like to see most of the mount options available as filesystem level properties with the option to override through mount options, but that's a lot more ambitious of an undertaking.
From: Jeff Mahoney <hidden> Date: 2018-05-17 14:46:13
On 5/16/18 6:35 PM, David Sterba wrote:
On Wed, May 16, 2018 at 06:03:56PM +0800, Anand Jain wrote:
quoted
Not yet ready for the integration. As I need to introduce
-o no_read_mirror_policy instead of -o read_mirror_policy=-<devid>
Mount option is mostly likely not the right interface for setting such
options, as usual.
I've seen a few alternate suggestions in the thread. I suppose the real
question is: what and where is the intended persistence for this choice?
A mount option gets it via fstab. How would a user be expected to set
it consistently via ioctl on each mount? Properties could work, but
there's more discussion needed there. Personally, I like the property
idea since it could conceivably be used on a per-file basis.
-Jeff
--
Jeff Mahoney
SUSE Labs
From: Jeff Mahoney <hidden> Date: 2018-05-17 14:46:38
On 5/17/18 8:25 AM, Austin S. Hemmelgarn wrote:
On 2018-05-16 22:32, Anand Jain wrote:
quoted
On 05/17/2018 06:35 AM, David Sterba wrote:
quoted
On Wed, May 16, 2018 at 06:03:56PM +0800, Anand Jain wrote:
quoted
Not yet ready for the integration. As I need to introduce
-o no_read_mirror_policy instead of -o read_mirror_policy=-<devid>
Mount option is mostly likely not the right interface for setting such
options, as usual.
I am ok to make it ioctl for the final. What do you think?
But to reproduce the bug posted in
Btrfs: fix the corruption by reading stale btree blocks
It needs to be a mount option, as randomly the pid can
still pick the disk specified in the mount option.
Personally, I'd vote for filesystem property (thus handled through the
standard `btrfs property` command) that can be overridden by a mount
option. With that approach, no new tool (or change to an existing tool)
would be needed, existing volumes could be converted to use it in a
backwards compatible manner (old kernels would just ignore the
property), and you could still have the behavior you want in tests (and
in theory it could easily be adapted to be a per-subvolume setting if we
ever get per-subvolume chunk profile support).
Properties are a combination of interfaces presented through a single
command. Although the kernel API would allow a direct-to-property
interface via the btrfs.* extended attributes, those are currently
limited to a single inode. The label property is set via ioctl and
stored in the superblock. The read-only subvolume property is also set
by ioctl but stored in the root flags.
As it stands, every property is explicitly defined in the tools, so any
addition would require tools changes. This is a bigger discussion,
though. We *could* use the xattr interface to access per-root or
fs-global properties, but we'd need to define that interface.
btrfs_listxattr could get interesting, though I suppose we could
simplify it by only allowing the per-subvolume and fs-global operations
on root inodes.
-Jeff
--
Jeff Mahoney
SUSE Labs
From: Austin S. Hemmelgarn <hidden> Date: 2018-05-17 15:30:25
On 2018-05-17 10:46, Jeff Mahoney wrote:
On 5/16/18 6:35 PM, David Sterba wrote:
quoted
On Wed, May 16, 2018 at 06:03:56PM +0800, Anand Jain wrote:
quoted
Not yet ready for the integration. As I need to introduce
-o no_read_mirror_policy instead of -o read_mirror_policy=-<devid>
Mount option is mostly likely not the right interface for setting such
options, as usual.
I've seen a few alternate suggestions in the thread. I suppose the real
question is: what and where is the intended persistence for this choice?
A mount option gets it via fstab. How would a user be expected to set
it consistently via ioctl on each mount? Properties could work, but
there's more discussion needed there. Personally, I like the property
idea since it could conceivably be used on a per-file basis.
For the specific proposed use case (the tests), it probably doesn't need to be persistent beyond mount options.
However, this also allows for a trivial configuration using a slow storage device to provide redundancy for a fast storage device of the same size, which is potentially very useful for some people. In that case, I can see most people who would be using it wanting it to follow the filesystem regardless of what context it's being mounted in (for example, it shouldn't need an extra option if mounted from a recovery environment or if it's moved to another system).
Most of my reason for recommending properties is that filesystem level properties appear to be the best thing BTRFS has to store per-volume configuration that's supposed to stay with the volume, despite not really being used for that even though there are quite a few mount options that are logical candidates for this type of thing (for example, the `ssd` options, `metadata_ratio`, and `max_inline` all make more logical sense as a property of the volume, not the mount).
Thanks Austin and Jeff for the suggestion.
I am not particularly a fan of mount option either mainly because
those options aren't persistent and host independent luns will
have tough time to have them synchronize manually.
Properties are better as it is persistent. And we can apply this
read_mirror_policy property on the fsid object.
But if we are talking about the properties then it can be stored
as extended attributes or ondisk key value pair, and I am doubt
if ondisk key value pair will get a nod.
I can explore the extended attribute approach but appreciate more
comments.
-Anand
On 05/17/2018 10:46 PM, Jeff Mahoney wrote:
On 5/17/18 8:25 AM, Austin S. Hemmelgarn wrote:
quoted
On 2018-05-16 22:32, Anand Jain wrote:
quoted
On 05/17/2018 06:35 AM, David Sterba wrote:
quoted
On Wed, May 16, 2018 at 06:03:56PM +0800, Anand Jain wrote:
quoted
Not yet ready for the integration. As I need to introduce
-o no_read_mirror_policy instead of -o read_mirror_policy=-<devid>
Mount option is mostly likely not the right interface for setting such
options, as usual.
I am ok to make it ioctl for the final. What do you think?
But to reproduce the bug posted in
Btrfs: fix the corruption by reading stale btree blocks
It needs to be a mount option, as randomly the pid can
still pick the disk specified in the mount option.
Personally, I'd vote for filesystem property (thus handled through the
standard `btrfs property` command) that can be overridden by a mount
option. With that approach, no new tool (or change to an existing tool)
would be needed, existing volumes could be converted to use it in a
backwards compatible manner (old kernels would just ignore the
property), and you could still have the behavior you want in tests (and
in theory it could easily be adapted to be a per-subvolume setting if we
ever get per-subvolume chunk profile support).
Properties are a combination of interfaces presented through a single
command. Although the kernel API would allow a direct-to-property
interface via the btrfs.* extended attributes, those are currently
limited to a single inode. The label property is set via ioctl and
stored in the superblock. The read-only subvolume property is also set
by ioctl but stored in the root flags.
As it stands, every property is explicitly defined in the tools, so any
addition would require tools changes. This is a bigger discussion,
though. We *could* use the xattr interface to access per-root or
fs-global properties, but we'd need to define that interface.
btrfs_listxattr could get interesting, though I suppose we could
simplify it by only allowing the per-subvolume and fs-global operations
on root inodes.
-Jeff
From: Austin S. Hemmelgarn <hidden> Date: 2018-05-18 12:36:53
On 2018-05-18 04:06, Anand Jain wrote:
Thanks Austin and Jeff for the suggestion.
I am not particularly a fan of mount option either mainly because
those options aren't persistent and host independent luns will
have tough time to have them synchronize manually.
Properties are better as it is persistent. And we can apply this
read_mirror_policy property on the fsid object.
But if we are talking about the properties then it can be stored
as extended attributes or ondisk key value pair, and I am doubt
if ondisk key value pair will get a nod.
I can explore the extended attribute approach but appreciate more
comments.
Hmm, thinking a bit further, might it be easier to just keep this as a mount option, and add something that lets you embed default mount options in the volume in a free-form manner? Then, you could set this persistently there, and could specify any others you want too. Doing that would also give very well defined behavior for exactly when changes would apply (the next time you mount or remount the volume), though handling of whether or not an option came from there or was specified on the command-line might be a bit complicated.
On 05/17/2018 10:46 PM, Jeff Mahoney wrote:
quoted
On 5/17/18 8:25 AM, Austin S. Hemmelgarn wrote:
quoted
On 2018-05-16 22:32, Anand Jain wrote:
quoted
On 05/17/2018 06:35 AM, David Sterba wrote:
quoted
On Wed, May 16, 2018 at 06:03:56PM +0800, Anand Jain wrote:
quoted
Not yet ready for the integration. As I need to introduce
-o no_read_mirror_policy instead of -o read_mirror_policy=-<devid>
Mount option is mostly likely not the right interface for setting such
options, as usual.
I am ok to make it ioctl for the final. What do you think?
But to reproduce the bug posted in
Btrfs: fix the corruption by reading stale btree blocks
It needs to be a mount option, as randomly the pid can
still pick the disk specified in the mount option.
Personally, I'd vote for filesystem property (thus handled through the
standard `btrfs property` command) that can be overridden by a mount
option. With that approach, no new tool (or change to an existing tool)
would be needed, existing volumes could be converted to use it in a
backwards compatible manner (old kernels would just ignore the
property), and you could still have the behavior you want in tests (and
in theory it could easily be adapted to be a per-subvolume setting if we
ever get per-subvolume chunk profile support).
Properties are a combination of interfaces presented through a single
command. Although the kernel API would allow a direct-to-property
interface via the btrfs.* extended attributes, those are currently
limited to a single inode. The label property is set via ioctl and
stored in the superblock. The read-only subvolume property is also set
by ioctl but stored in the root flags.
As it stands, every property is explicitly defined in the tools, so any
addition would require tools changes. This is a bigger discussion,
though. We *could* use the xattr interface to access per-root or
fs-global properties, but we'd need to define that interface.
btrfs_listxattr could get interesting, though I suppose we could
simplify it by only allowing the per-subvolume and fs-global operations
on root inodes.
-Jeff
From: Steven Davies <hidden> Date: 2019-01-21 12:25:52
On 2018-05-16 11:03, Anand Jain wrote:
Going back to an old patchset I was testing this weekend:
Adds the mount option:
mount -o read_mirror_policy=<devid>
To set the devid of the device which should be used for read. That
means all the normal reads will go to that particular device only.
This also helps testing and gives a better control for the test
scripts including mount context reads.
Signed-off-by: Anand Jain <redacted>
Not an expert, but might this be overwritten with another state? e.g. BTRFS_DEV_STATE_REPLACE_TGT. The device would then lose its READ_MIRROR flag and the fs would always use the first device for reading.
+ break;
+ }
I noticed that it's possible to pass this option multiple times at mount, which sets multiple devices as read mirrors. While that doesn't do anything harmful, the code below will only use the first device. It may be worth at least mentioning this in documentation.
Why set preferred_mirror again? The only effect of re-setting it will be to use the lowest devid (1) rather than the highest (2). Is there any difference? Either way it should never happen, because the above code traps it (except when the BTRFS_DEV_STATE_READ_MIRROR flag has been changed as mentioned above).
From Goffredo's comment[1] on Timofey's similar effect patch, if it becomes possible to have more mirrors in a RAID1/RAID10 fs then this code will need to be updated to test dev_state for more than two devices. Would it be sensible to implement this as a for loop straight away?
quoted hunk
+ break; case BTRFS_READ_MIRROR_DEFAULT: case BTRFS_READ_MIRROR_BY_PID: default:
On 2018-05-16 11:03, Anand Jain wrote:
Going back to an old patchset I was testing this weekend:
quoted
Adds the mount option:
mount -o read_mirror_policy=<devid>
To set the devid of the device which should be used for read. That
means all the normal reads will go to that particular device only.
This also helps testing and gives a better control for the test
scripts including mount context reads.
Signed-off-by: Anand Jain <redacted>
Not an expert, but might this be overwritten with another state? e.g. BTRFS_DEV_STATE_REPLACE_TGT. The device would then lose its READ_MIRROR flag and the fs would always use the first device for reading.
It won't it is defined as bitmap.
quoted
+ break;
+ }
I noticed that it's possible to pass this option multiple times at mount, which sets multiple devices as read mirrors. While that doesn't do anything harmful, the code below will only use the first device. It may be worth at least mentioning this in documentation.
There were few feedback if read_mirror_policy should be a mount
option or a sysfs or a property. IMO property is better as it would
be persistent. In sysfs and mount-option, the user or a config file
has to remember. Will fix.
Why set preferred_mirror again? The only effect of re-setting it will be to use the lowest devid (1) rather than the highest (2). Is there any difference? Either way it should never happen, because the above code traps it (except when the BTRFS_DEV_STATE_READ_MIRROR flag has been changed as mentioned above).
Code at [*] above does ++preferred_mirror. So the following
preferred_mirror = first;
is not redundant.
From Goffredo's comment[1] on Timofey's similar effect patch, if it becomes possible to have more mirrors in a RAID1/RAID10 fs then this code will need to be updated to test dev_state for more than two devices. Would it be sensible to implement this as a for loop straight away?
Right. A loop is better, will add.
quoted
+ break;
case BTRFS_READ_MIRROR_DEFAULT:
case BTRFS_READ_MIRROR_BY_PID:
default:
From: Steven Davies <hidden> Date: 2019-01-22 14:28:56
On 2019-01-22 13:43, Anand Jain wrote:
On 01/21/2019 07:56 PM, Steven Davies wrote:
quoted
On 2018-05-16 11:03, Anand Jain wrote:
quoted
quoted
+ break;
+ }
I noticed that it's possible to pass this option multiple times at mount, which sets multiple devices as read mirrors. While that doesn't do anything harmful, the code below will only use the first device. It may be worth at least mentioning this in documentation.
There were few feedback if read_mirror_policy should be a mount
option or a sysfs or a property. IMO property is better as it would
be persistent. In sysfs and mount-option, the user or a config file
has to remember. Will fix.
I agree. Would/could this be set at mkfs time? It would then be similar to how mdadm's --write-mostly is set up.
Why set preferred_mirror again? The only effect of re-setting it will be to use the lowest devid (1) rather than the highest (2). Is there any difference? Either way it should never happen, because the above code traps it (except when the BTRFS_DEV_STATE_READ_MIRROR flag has been changed as mentioned above).
Code at [*] above does ++preferred_mirror. So the following
preferred_mirror = first;
is not redundant.
Yes, but preferred_mirror will be first + num_stripes; meaning that without that line the last stripe is preferred and with it the first stripe is preferred. It only changes the fallback case from reading from devid 2 to devid 1, which barely matters. Perhaps it would be even better to fall back to pid % num_stripes in this case anyway?
quoted
From Goffredo's comment[1] on Timofey's similar effect patch, if it becomes possible to have more mirrors in a RAID1/RAID10 fs then this code will need to be updated to test dev_state for more than two devices. Would it be sensible to implement this as a for loop straight away?
Right. A loop is better, will add.
Thinking more about this, if it becomes possible to have more than two devices in a RAID1 or part-RAID10 then do we also need to consider that there may be more than one READ_MIRROR drive? e.g. slow+fast+fast drives in a RAID1 with this approach would result in all chunks being read from the first drive with READ_MIRROR set if both chunks are on the fast drives. This is where Timofey's queue length patch in [1] would become useful - in any case, that is an existing problem and probably deserving of an entirely new patch.