Re: [RFC] nanoseconds vs. the epoch

Andreas Dilger <[email protected]>
Newsgroups gmane.comp.file-systems.ext2.devel
Message-ID <[email protected]>
Stephen,
could you please reply with your requirements for this reserved space?
We want to get moving with the nanosecond timestamps.

On Jun 05, 2006  15:02 -0600, Andreas Dilger wrote:
> On Jun 05, 2006  14:32 +0100, Stephen C. Tweedie wrote:
> > On Fri, 2006-06-02 at 11:57 -0600, Andreas Dilger wrote:
> > > The only other thing this patch is missing (which isn't totally needed)
> > > is to potentially update existing inodes to have a larger i_extra_isize
> > > for holding the new timestamps.  This implies potentially moving the 
> > > EAs further out, which adds some complexity.
> > 
> > I think we need a long, hard think about that whole issue before we go
> > much further down this path.  I'd rather reserve a chunk of space in the
> > extra_size region in advance so that we can assign future fields
> > cleanlyu.  Because if we rely on growing the extra_size fields
> > dynamically, then there will always be easily-exploitable edge cases
> > where the inode simply has no room left (the current EA mechanism lets
> > you fill *both* the inode unused space *and* the entire EA block with
> > xattrs.)
> 
> While I partly agree, I also think it doesn't make a lot of sense to
> reserve too large a fixed space for the inodes if we don't have any
> idea what we will be using them for, because that can impact the
> performance of EA-using applications greatly if they have their EAs
> bumped to an external block.
> 
> How many fields were you thinking of reserving, and how would this
> impact e.g. selinux, ACLs (i.e. how many ACLs can be stored in the
> inode now, how many are normally used, etc)?  Hopefully, before we
> reserve too many fields for the future we will also have fixed the
> EA code to allow EAs to be stored more flexibly and this will stop
> being a DOS issue, and just a performance one (i.e. user EAs will
> be stored in slower external EA space if there is no fast EA space).
> 
> The current proposal is for i_ctime_extra, i_mtime_extra, i_atime_extra
> (12 bytes).  There was some mention of i_crtime (creation time) in the
> past (either 32 or 64 bits?), a checksum[*].
> 
> For Lustre we wouldn't (currently) be impacted by up to 20 bytes of
> new space usage (which is including the 12 bytes for nsec timestamps).
> Older filesystems were formatted with at least 256-byte inodes, and
> new Lustre filesystems are formatted with at least 512-byte inodes.
> Not that this is a deciding factor, but just a point of reference.
> 
> [*] probably 16 bits is enough for a 128-byte inode, if we are doing
>     per-inode checksums instead of for the whole block, so the padding
>     after i_extra_isize might be enough.

Cheers, Andreas
--
Andreas Dilger
Principal Software Engineer
Cluster File Systems, Inc.



-------------------------------------------------------------------------
Using Tomcat but need to do more? Need to support web services, security?
Get stuff done quickly with pre-integrated technology to make your job easier
Download IBM WebSphere Application Server v.1.0.1 based on Apache Geronimo
http://sel.as-us.falkag.net/sel?cmd=lnk&kid=120709&bid=263057&dat=121642
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.