Re: Progressing RFC errata for RFC 5661
Rick Macklem <[email protected]>
| Newsgroups | gmane.ietf.nfsv4 |
|---|---|
| Message-ID | <YT1PR01MB35935EBFB2ED93141B18FB80DD890@YT1PR01MB3593.CANPRD01.PROD.OUTLOOK.COM> |
Rick Macklem wrote: >Mkrtchyan, Tigran wrote: >>----- Original Message ----- >>> From: "Rick Macklem" <[email protected]> >>> To: "Trond Myklebust" <[email protected]> >>> Cc: "Tigran Mkrtchyan" <[email protected]>, "Dave Noveck" <[email protected]>, "Magnus Westerlund" >>> <[email protected]>, "NFSv4" <[email protected]> >>> Sent: Thursday, September 19, 2019 6:07:32 AM >>> Subject: Re: [nfsv4] Progressing RFC errata for RFC 5661 >> >>> Trond Myklebust wrote: >>>>Rick, >>>> >>>>That errata predates most of the Linux pNFS client implementation. We >>>>wrote the implementation to conform to the errata. >>>> >>>>So no. It's not a bug. It's a deliberate design based on a decision >>>>that was discussed in the IETF WG, on the mailing list >>>> >>>>https://mailarchive.ietf.org/arch/msg/nfsv4/_KTtO6uz-MvRoStbhPuOXWZr6yI >>> This actually appears to be a discussion related to the offset and length >>> arguments for LayoutCommit, but... >>>> >>>>and in a special session of the IETF: >>>> >>>>https://mailarchive.ietf.org/arch/msg/nfsv4/Rpw9XCwCARxfaU4ym5L2TauV6ao >>> Ok. This was long before I got around to implementing it, so I wouldn't have >>> understood the implications. >>> --> I would have been interested in hearing the rationale behind not doing >>> LayoutCommit for FILE_SYNC4 writes, since it seems to me that RFC-5661 >>> had gotten it right when it required them. >> >>In case of a cluster filesystem back-end, FILE_SYNC4 on write indicates that >>MDS already >>has the correct file attributes. An extra LAYOUTCOMMIT will introduce >>additional overhead. >>I can imagine a HPC workload where client talks to DS over InfiniBand and >>Ethernet to >>MDS. An extra LAYOUTCOMMIT will drop write throughput. >Ok, I'll assume you have a server which needs LayoutCommit for UNSTABLE4 >writes, but doesn't need one for FILE_SYNC4 writes. >(If the server never needs LayoutCommit, it can simply do what I believe > the Netapp filer does, which is reply NFS4ERR_NOTSUPP.) > >I agree that this could result in extra overhead, but at least for some cases >the client could put the LayoutCommit in the same compound as something >like Close (or Commit if the server does commit-through-mds), which would >avoid an extra RPC RTT. >Presumably the server would know it didn't need to do anything and could >just reply NFS_OK for the operation in the FILE_SYNC4 case? > >But, yes, this is a case where using the Errata might improve performance. > >For NFSv4.2, I think a new recommended attribute could be added, so that >a server could indicate to a client when it wanted a LayoutCommit. >(I think the NFSv4.2 versioning rules would allow this to be added?) >I'm not sure how the client would "handshake" with the server, acknowledging >that it understood the new attribute, though? Duh. I realized that the "handshake" would simply be the client getting the attribute. >It is tempting to add NFL4_UFLG_FILE_SYNC4_LAYOUT_COMMIT, but that >would probably cause problems for extant implementations that don't >expect the flag bit (returning an error instead of ignoring it, etc). rick > > As I said, the FreeBSD server can handle this case, it just results in a lot of > overhead synchronizing Size, Change, Time_Modify between MDS and DS > whenever a RW layout is issued to a client for the file. > > Thanks for pointing this out, rick > > On Wed, 18 Sep 2019 at 12:39, Rick Macklem <[email protected]> wrote: >> >> Mkrtchyan, Tigran wrote: >> [stuff snipped] >> >Hi Rick, >> > >> >here is the public link to errata >> > >> >https://www.rfc-editor.org/errata/eid2751 >> > >> >Tigran. >> Thanks Tigran. >> >> Ok, so now that I've read it I have to admit I think it is rewriting the RFC >> to conform with what the Linux client does. >> >> I think this para. from Sec. 13.10 of RFC-5661 is clear: >> The NFSv4.1 protocol only provides close-to-open file data cache >> semantics; meaning that when the file is closed, all modified data is >> written to the server. When a subsequent OPEN of the file is done, >> the change attribute is inspected for a difference from a cached >> value for the change attribute. For the case above, this means that >> a LAYOUTCOMMIT will be done at close (along with the data WRITEs) and >> will update the file's size and change attribute. Access from >> another client after that point will result in the appropriate size >> being returned. >> >> It states "will be done". It doesn't say anything about UNSTABLE4 vs FILE_SYNC4. >> (I think most POSIX-like clients would consider the fsync(2) syscall to require >> the same treatment as "close" above, but that is a POSIX-specific client issue.) >> I can see the argument that, since there is no "must" in the statement, that a >> client can choose not to do this, but that would also imply that the client will >> need to live with the consequences of it. >> >> I think the second sentence of the first para. of the errata is bogus: >> For file layouts, WRITEs to a Data Server that return a stable_how4 value of >> FILE_SYNC4 guarantee that data and file system metadata are on stable >> storage. This means that a LAYOUTCOMMIT is not needed in order to make the >> data and metadata visible to the metadata server and other clients. >> >> Why? >> The FILE_SYNC4 was returned by the DS. This would imply the DS >> has committed data and metadata to stable storage on the DS. >> However, I am not aware of anything in RFC-5661 that would imply that >> the Size, Time_Modify and Change attributes or anything else must have >> been updated or in stable storage on the MDS at this time. >> >> If a server does not require LayoutCommit operations for correct behaviour >> then it can simply reply NFS4ERR_NOTSUPP (as I believe the Netapp filer >> does) and the client then no longer needs to do them. >> >> If there is somewhere in RFC-5661 that it is stated that LayoutCommits are >> not required when the DS replies FILE_SYNC4, then I missed it and there >> is a problem with RFC-5661 that needs to be addressed. >> >> Otherwise, sorry, but it seems that the bug is in the Linux client >> implementation >> and not RFC-5661. >> >> Is there a File Layout pNFS server implementation where the DSs return >> FILE_SYNC4 that will break if the client does a LayoutCommit for this case? >> (If so, then something may need to be done.) >> >> rick >> >> _______________________________________________ nfsv4 mailing list [email protected] https://www.ietf.org/mailman/listinfo/nfsv4 _______________________________________________ nfsv4 mailing list [email protected] https://www.ietf.org/mailman/listinfo/nfsv4