Re: [GIT] Bcache version 12
From: LuVar <hidden>
Date: 2011-10-01 15:19:57
Also in:
linux-bcache, lkml
Hi here. ----- "Dan J Williams" [off-list ref] wrote:
On Fri, Sep 30, 2011 at 12:14 AM, Kent Overstreet [off-list ref] wrote:quoted
quoted
quoted
Cache devices have a basically identical superblock as backingdevicesquoted
quoted
quoted
though, and some of the registration code is shared, but cachedevicesquoted
quoted
quoted
don't correspond to any block devices.Just like a raid0 is a virtual creation from two block devices?Orquoted
quoted
some other meaning of "don't correspond"?No. Remember, you can hang multiple backing devices off a cache. Each backing device shows up as as a new block device - i.e. ifyou'requoted
caching /dev/sdb, you now use it as /dev/bcache0. But the SSD doesn't belong to any of those /dev/bcacheN devices.So to clarify I read that as "it belongs to all of them". The ssd (/dev/sda, for example) can cache the contents of N block devices, and to get to the cached version of each of those you go through /dev/bcache[0..N]. The problem you perceive is that an md device requires a 1:1 mapping of member devices to md devices. So if we had /dev/sda and /dev/sdb in a cache configuration (/dev/md0) your concern is that if we simultaneously wanted a /dev/md1 that caches /dev/sda and /dev/sdc that md would not be able to handle it. Is that the right interpretation? I assume /dev/sda in the example would have some bcache-logical partitions to delineate the /dev/sdb and /dev/sdc cache data? Which sounds similar to the logical partitions md handles now for external metadata. I'm not proposing that cache-state metadata could be handled in userspace it's too integral to the i/o path, just pointing out that having /dev/sda be a member of both /dev/md0 and /dev/md1 is possible.quoted
quoted
quoted
A cache set is a set of cache devices - i.e. SSDs. The primary motivitation for cache sets (as distinct from just caches) is tohavequoted
quoted
quoted
the ability to mirror only dirty data, and not clean data. i.e. if you're doing writeback caching of a raid6, your ssd isnow aquoted
quoted
quoted
single point of failure. You could use raid1 SSDs, but most ofthe dataquoted
quoted
quoted
in the cache is clean, so you don't need to mirror that... justthequoted
quoted
quoted
dirty data....but you only incur that "mirror clean data" penalty once, andthenquoted
quoted
it's just a normal raid1 mirroring writes, right?No idea what you mean.../dev/md1 is a slow raid5 and /dev/md0 is a raid1 of two ssds. Once /dev/md0 is synced the only mirror traffic is for incoming cache-dirtying writes and cache-clean read allocations. We agree about incoming dirty-data, but you are saying you don't want to mirror read allocations?
Just one visualization of my understand of bcache set with mirroring only dirty data: http://147.175.167.212/~luvar/bcache/bcacheSSDset.png . If I am not wrong, read alocations are for example green and blue data. Dirty allocations is red one and it should be mirrored across all ssds in mirror set to provide ssd fail security. On the other hand, greed, blue... data are backed up on raid6 and it is nod needed to mirror them across ssd set. They should be only on one ssd to provide read speedup. Hmmm (sci-fi), if read allocations (not dirty data) will be mirrored in ssds set, they could be used to improve cache read speed, sacrificing some ssd space. It would be great if cache algorithm can mark really hot data to be mirrored for speed reading...
quoted
quoted
See, if these things were just md devices multiple cache devicewouldquoted
quoted
already be "done", or at least on its way by just stacking mddevices.quoted
quoted
Where "done" is probably an oversimplification.No, it really wouldn't save us anything. If all we wanted to do was mirror everything, there'd be no point in implementing multiplecachequoted
device support, and you'd just use bcache on top of md. We're implementing something completely new! You read what I said about only mirroring dirty data... right?I did but I guess I did not fully grok it.quoted
quoted
quoted
quoted
In any case it certainly could be modelled in md - and if themodelling werequoted
quoted
quoted
quoted
not elegant (e.g. even device numbers for backing devices, odddevice numbersquoted
quoted
quoted
quoted
for cache devices) we could "fix" md to make it more elegant.But we've no reason to create block devices for caches or have a1:1quoted
quoted
quoted
mapping - that'd be a serious step backwards in functionality.I don't follow that... there's nothing that prevents havingmultiplequoted
quoted
superblocks per cache array.Multiple... superblocks? Do you mean partitioning up the cache, ordoquoted
you mean creating multiple block devices for a cache? Either wayit's aquoted
silly hack.quoted
A couple reasons I'm probing the md angle. 1/ Since the backing devices are md devices it would be nice ifallquoted
quoted
the user space assembly logic that has seeped into udev and dracut could be re-used for assembling bcache devices. As it stands itseemsquoted
quoted
bcache relies on in-kernel auto-assembly, which md has discouraged with the v1 superblock.md was doing in kernel probing, which bcache does not do. Whatbcache isquoted
doing is centralizing all the code that touches the on disk superblock/metadata. You want to change something in the superblock-quoted
you just have to tell the kernel to do it for you. Otherwise notonlyquoted
would there be duplication of code, it'd be impossible to do safely without races or the userspace code screwing something up; only the kernel knows and controls the state of everything.Makes sense but there is a difference between the metadata that specifies the configuration and the metadata that tracks the state of the cache. If that distinction is made then userspace can tell the kernel to run a block cache of blockdevA and blockdevB and the kernel only needs to handle the cache state metadata.quoted
Or do you expect the ext4 superblock to be managed in normaloperationquoted
by userspace tools?No.quoted
quoted
We even have nascent GUI support in gnome-disk-utility it would be nice to harness some of thatenablingquoted
quoted
momentum for this.I've got nothing against standardizing the userspace interfaces tomakequoted
life easier for things like gnome-disk-utility. Tell me what youwantquoted
and if it's sane I'll see about implementing it.That's the point, userspace has some knowledge of how to interrogate and manage md devices. A bcache device is brand new... maybe for good reason but that's what I'm trying to understand.quoted
quoted
2/ md supports multiple superblock formats and if you Google "ssd caching" you'll see that there may be other superblock formatsthatquoted
quoted
the Linux block-caching driver could be asked to support down the road. And wouldn't it be nice if bcache had at least the optiontoquoted
quoted
support the on-disk format of whatever dm-cache is doing?That's pure fantasy. That's like expecting the ext4 code to mount antfsquoted
filesystem!No, there's portions of what bcache does that are similar to what md does. Do we need to invent new multiple-device handling infrastructure for a block device driver? But we are quickly approaching the "show me the code" portion of this discussion, so I need to go do more reading of bcache.quoted
There's a lot more to bcache's metadata than a superblock, there'saquoted
journal and a full b-tree. A cache is going to need an index ofsomequoted
kind.Yes, but that can be independent of the configuration metadata. -- Dan -- To unsubscribe from this list: send the line "unsubscribe linux-bcache" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html