From: Albert Cahalan <hidden> Date: 2007-06-13 05:45:29
Neat! It's great to see somebody else waking up to the idea that
storage media is NOT to be trusted.
Judging by the design paper, it looks like your structs have some
alignment problems.
The usual wishlist:
* inode-to-pathnames mapping
* a subvolume that is a single file (disk image, database, etc.)
* directory indexes to better support Wine and Samba
* secure delete via destruction of per-file or per-block random crypto keys
* fast (seekless) access to normal-sized SE Linux data
* atomic creation of copy-on-write directory trees
* immutable bits like UFS has
* hole punch ability
* insert/delete ability (add/remove a chunk in the middle of a file)
From: Chris Mason <hidden> Date: 2007-06-13 12:04:01
On Wed, Jun 13, 2007 at 01:45:28AM -0400, Albert Cahalan wrote:
Neat! It's great to see somebody else waking up to the idea that
storage media is NOT to be trusted.
Judging by the design paper, it looks like your structs have some
alignment problems.
Actual defs are all packed, but I may still shuffle around the structs
to optimize alignment. The keys are fixed, although I may make the u32
in the middle smaller.
The usual wishlist:
* inode-to-pathnames mapping
This one I'll code, it will help with inode link count verification. I
want to be able to detect at run time that an inode with a link count of
zero is still actually in a directory. So there will be back pointers
from the inode to the directory.
Also, the incremental backup code will be able to walk the btree to find
inodes that have changed, and the backpointers will help make a list of
file names that need to be rsync'd or whatever.
* a subvolume that is a single file (disk image, database, etc.)
subvolumes can be made that have a single file in them, but they have to
be directories right now. Doing otherwise would complicate mounts and
other management tools (inside the btree, it doesn't really matter).
* directory indexes to better support Wine and Samba
* secure delete via destruction of per-file or per-block random crypto keys
I'd rather keep secure delete as a userland problem (or a layered FS
problem). When you take backups and other copies of the file into
account, it's a bigger problem than btrfs wants to tackle right now.
* fast (seekless) access to normal-sized SE Linux data
acls and xattrs will adjacent to the inode in the tree. Most of the
time it'll be seekless.
* atomic creation of copy-on-write directory trees
Do you mean something more fine grained than the current snapshotting
system?
* immutable bits like UFS has
I'll do the ext2 chattr calls.
* hole punch ability
Hole punching isn't harder or easier in btrfs than most other
filesystems that support holes. It's largely a VM issue.
* insert/delete ability (add/remove a chunk in the middle of a file)
The disk format makes this O(extent records past the chunk). It's
possible to code but it would not be optimized.
-chris
From: Albert Cahalan <hidden> Date: 2007-06-13 16:14:41
On 6/13/07, Chris Mason [off-list ref] wrote:
On Wed, Jun 13, 2007 at 01:45:28AM -0400, Albert Cahalan wrote:
quoted
The usual wishlist:
* inode-to-pathnames mapping
This one I'll code, it will help with inode link count verification. I
want to be able to detect at run time that an inode with a link count of
zero is still actually in a directory. So there will be back pointers
from the inode to the directory.
Great, but fsck improvement wasn't on my mind. This is
a desirable feature for the NFS server, and for regular users.
Think about a backup program trying to maintain hard links.
Also, the incremental backup code will be able to walk the btree to find
inodes that have changed, and the backpointers will help make a list of
file names that need to be rsync'd or whatever.
quoted
* a subvolume that is a single file (disk image, database, etc.)
subvolumes can be made that have a single file in them, but they have to
be directories right now. Doing otherwise would complicate mounts and
other management tools (inside the btree, it doesn't really matter).
Bummer. As I understand it, ZFS provides this. :-)
quoted
* directory indexes to better support Wine and Samba
* secure delete via destruction of per-file or per-block random crypto keys
I'd rather keep secure delete as a userland problem (or a layered FS
problem). When you take backups and other copies of the file into
account, it's a bigger problem than btrfs wants to tackle right now.
It can't be a userland problem if you allow disk blocks to move.
Volume resizing, logging/journalling, etc. -- they combine to make
the userland solution essentially impossible. (one could wipe the
whole partition, or maybe fill ALL space on the volume)
I think it needs to be per-extent.
At each level in the btree, you place a randomly generated key
for the more leafward nodes. This means that secure deletion is
merely the act of wiping the key... which can itself occur by
wiping the key of the more rootward node.
quoted
* atomic creation of copy-on-write directory trees
Do you mean something more fine grained than the current snapshotting
system?
I believe so. Example: I have a linux-2.6 directory. It's not
a mount point or anything special like that. I want to copy
it to a new directory called wip, without actually copying
all the blocks. To all the normal POSIX API stuff, this copy
should look like the result of "cp -a", not hard links.
quoted
* insert/delete ability (add/remove a chunk in the middle of a file)
The disk format makes this O(extent records past the chunk). It's
possible to code but it would not be optimized.
That's understandable, but note that Reiserfs can support this.
From: Chris Mason <hidden> Date: 2007-06-13 17:00:42
On Wed, Jun 13, 2007 at 12:14:40PM -0400, Albert Cahalan wrote:
On 6/13/07, Chris Mason [off-list ref] wrote:
quoted
On Wed, Jun 13, 2007 at 01:45:28AM -0400, Albert Cahalan wrote:
quoted
quoted
The usual wishlist:
* inode-to-pathnames mapping
This one I'll code, it will help with inode link count verification. I
want to be able to detect at run time that an inode with a link count of
zero is still actually in a directory. So there will be back pointers
from the inode to the directory.
Great, but fsck improvement wasn't on my mind. This is
a desirable feature for the NFS server, and for regular users.
Think about a backup program trying to maintain hard links.
Sure, it'll be there either way ;)
quoted
Also, the incremental backup code will be able to walk the btree to find
inodes that have changed, and the backpointers will help make a list of
file names that need to be rsync'd or whatever.
quoted
* a subvolume that is a single file (disk image, database, etc.)
subvolumes can be made that have a single file in them, but they have to
be directories right now. Doing otherwise would complicate mounts and
other management tools (inside the btree, it doesn't really matter).
Bummer. As I understand it, ZFS provides this. :-)
Grin, when the pain of typing cd subvol is btrfs' biggest worry, I'll be
doing very well.
quoted
quoted
* directory indexes to better support Wine and Samba
* secure delete via destruction of per-file or per-block random crypto
keys
I'd rather keep secure delete as a userland problem (or a layered FS
problem). When you take backups and other copies of the file into
account, it's a bigger problem than btrfs wants to tackle right now.
It can't be a userland problem if you allow disk blocks to move.
Volume resizing, logging/journalling, etc. -- they combine to make
the userland solution essentially impossible. (one could wipe the
whole partition, or maybe fill ALL space on the volume)
Right about here is where I would insert a long story about ecryptfs, or
encryption solutions that happen all in userland. At any rate, it is
outside the scope of v1.0, even though I definitely agree it is an
important problem for some people.
quoted
quoted
* atomic creation of copy-on-write directory trees
Do you mean something more fine grained than the current snapshotting
system?
I believe so. Example: I have a linux-2.6 directory. It's not
a mount point or anything special like that. I want to copy
it to a new directory called wip, without actually copying
all the blocks. To all the normal POSIX API stuff, this copy
should look like the result of "cp -a", not hard links.
This would be a snapshot, which has to be done on a subvolume right now.
It is not as nice as being able to pick a random directory, but I've
only been able to get this far by limiting the feature scope
significantly. What I did do was make subvolumes very cheap...just make
a bunch of them.
Keep in mind that if you implement a cow directory tree without a
snapshot, and you don't want to duplicate any blocks in the cow, you're
going to have fun with inode numbers.
-chris
From: Albert Cahalan <hidden> Date: 2007-06-14 06:59:23
On 6/13/07, Chris Mason [off-list ref] wrote:
On Wed, Jun 13, 2007 at 12:14:40PM -0400, Albert Cahalan wrote:
quoted
On 6/13/07, Chris Mason [off-list ref] wrote:
quoted
On Wed, Jun 13, 2007 at 01:45:28AM -0400, Albert Cahalan wrote:
quoted
quoted
quoted
* secure delete via destruction of per-file or per-block random crypto
keys
I'd rather keep secure delete as a userland problem (or a layered FS
problem). When you take backups and other copies of the file into
account, it's a bigger problem than btrfs wants to tackle right now.
It can't be a userland problem if you allow disk blocks to move.
Volume resizing, logging/journalling, etc. -- they combine to make
the userland solution essentially impossible. (one could wipe the
whole partition, or maybe fill ALL space on the volume)
Right about here is where I would insert a long story about ecryptfs, or
encryption solutions that happen all in userland. At any rate, it is
outside the scope of v1.0, even though I definitely agree it is an
important problem for some people.
I'm sure you do have a nice long story, and I'm sure it seems
correct, but there is something not quite right about the add-on
hacks.
BTW, I'm suggesting that this be about deletion, not protection
of data you wish to keep. It covers more than just file bodies.
It covers inode data, block allocations, etc.
quoted
quoted
quoted
* atomic creation of copy-on-write directory trees
Do you mean something more fine grained than the current snapshotting
system?
I believe so. Example: I have a linux-2.6 directory. It's not
a mount point or anything special like that. I want to copy
it to a new directory called wip, without actually copying
all the blocks. To all the normal POSIX API stuff, this copy
should look like the result of "cp -a", not hard links.
This would be a snapshot, which has to be done on a subvolume right now.
It is not as nice as being able to pick a random directory, but I've
only been able to get this far by limiting the feature scope
significantly. What I did do was make subvolumes very cheap...just make
a bunch of them.
Can a regular user create and use a subvolume? If not, then
this doesn't work. (if so, then I have other concerns...)
From: Chris Mason <hidden> Date: 2007-06-14 12:34:13
On Thu, Jun 14, 2007 at 02:59:23AM -0400, Albert Cahalan wrote:
On 6/13/07, Chris Mason [off-list ref] wrote:
[ secure deletion in btrfs ]
quoted
Right about here is where I would insert a long story about ecryptfs, or
encryption solutions that happen all in userland. At any rate, it is
outside the scope of v1.0, even though I definitely agree it is an
important problem for some people.
I'm sure you do have a nice long story, and I'm sure it seems
correct, but there is something not quite right about the add-on
hacks.
BTW, I'm suggesting that this be about deletion, not protection
of data you wish to keep. It covers more than just file bodies.
It covers inode data, block allocations, etc.
Sorry, it's still way outside the scope of v1.0.
quoted
quoted
quoted
quoted
* atomic creation of copy-on-write directory trees
Do you mean something more fine grained than the current snapshotting
system?
I believe so. Example: I have a linux-2.6 directory. It's not
a mount point or anything special like that. I want to copy
it to a new directory called wip, without actually copying
all the blocks. To all the normal POSIX API stuff, this copy
should look like the result of "cp -a", not hard links.
This would be a snapshot, which has to be done on a subvolume right now.
It is not as nice as being able to pick a random directory, but I've
only been able to get this far by limiting the feature scope
significantly. What I did do was make subvolumes very cheap...just make
a bunch of them.
Can a regular user create and use a subvolume? If not, then
this doesn't work. (if so, then I have other concerns...)
That's the long term goal, but I'll have to reorganize things such that
subvolumes created by a user can all fall under sane accounting.
-chris