Re: Progressing RFC errata for RFC 5661

Rick Macklem <[email protected]>
Newsgroups gmane.ietf.nfsv4
Message-ID <YT1PR01MB35935EBFB2ED93141B18FB80DD890@YT1PR01MB3593.CANPRD01.PROD.OUTLOOK.COM>
Rick Macklem wrote:
>Mkrtchyan, Tigran wrote:
>>----- Original Message -----
>>> From: "Rick Macklem" <[email protected]>
>>> To: "Trond Myklebust" <[email protected]>
>>> Cc: "Tigran Mkrtchyan" <[email protected]>, "Dave Noveck" <[email protected]>, "Magnus Westerlund"
>>> <[email protected]>, "NFSv4" <[email protected]>
>>> Sent: Thursday, September 19, 2019 6:07:32 AM
>>> Subject: Re: [nfsv4] Progressing RFC errata for RFC 5661
>>
>>> Trond Myklebust wrote:
>>>>Rick,
>>>>
>>>>That errata predates most of the Linux pNFS client implementation. We
>>>>wrote the implementation to conform to the errata.
>>>>
>>>>So no. It's not a bug. It's a deliberate design based on a decision
>>>>that was discussed in the IETF WG, on the mailing list
>>>>
>>>>https://mailarchive.ietf.org/arch/msg/nfsv4/_KTtO6uz-MvRoStbhPuOXWZr6yI
>>> This actually appears to be a discussion related to the offset and length
>>> arguments for LayoutCommit, but...
>>>>
>>>>and in a special session of the IETF:
>>>>
>>>>https://mailarchive.ietf.org/arch/msg/nfsv4/Rpw9XCwCARxfaU4ym5L2TauV6ao
>>> Ok. This was long before I got around to implementing it, so I wouldn't have
>>> understood the implications.
>>> --> I would have been interested in hearing the rationale behind not doing
>>>      LayoutCommit for FILE_SYNC4 writes, since it seems to me that RFC-5661
>>>      had gotten it right when it required them.
>>
>>In case of a cluster filesystem back-end, FILE_SYNC4 on write indicates that >>MDS already
>>has the correct file attributes. An extra LAYOUTCOMMIT will introduce >>additional overhead.
>>I can imagine a HPC workload where client talks to DS over InfiniBand and >>Ethernet to
>>MDS. An extra LAYOUTCOMMIT will drop write throughput.
>Ok, I'll assume you have a server which needs LayoutCommit for UNSTABLE4
>writes, but doesn't need one for FILE_SYNC4 writes.
>(If the server never needs LayoutCommit, it can simply do what I believe
> the Netapp filer does, which is reply NFS4ERR_NOTSUPP.)
>
>I agree that this could result in extra overhead, but at least for some cases
>the client could put the LayoutCommit in the same compound as something
>like Close (or Commit if the server does commit-through-mds), which would
>avoid an extra RPC RTT.
>Presumably the server would know it didn't need to do anything and could
>just reply NFS_OK for the operation in the FILE_SYNC4 case?
>
>But, yes, this is a case where using the Errata might improve performance.
>
>For NFSv4.2, I think a new recommended attribute could be added, so that
>a server could indicate to a client when it wanted a LayoutCommit.
>(I think the NFSv4.2 versioning rules would allow this to be added?)
>I'm not sure how the client would "handshake" with the server, acknowledging
>that it understood the new attribute, though?
Duh. I realized that the "handshake" would simply be the client getting the
attribute.

>It is tempting to add NFL4_UFLG_FILE_SYNC4_LAYOUT_COMMIT, but that
>would probably cause problems for extant implementations that don't
>expect the flag bit (returning an error instead of ignoring it, etc).

rick



>
> As I said, the FreeBSD server can handle this case, it just results in a lot of
> overhead synchronizing Size, Change, Time_Modify between MDS and DS
> whenever a RW layout is issued to a client for the file.
>
> Thanks for pointing this out, rick
>
> On Wed, 18 Sep 2019 at 12:39, Rick Macklem <[email protected]> wrote:
>>
>> Mkrtchyan, Tigran wrote:
>> [stuff snipped]
>> >Hi Rick,
>> >
>> >here is the public link to errata
>> >
>> >https://www.rfc-editor.org/errata/eid2751
>> >
>> >Tigran.
>> Thanks Tigran.
>>
>> Ok, so now that I've read it I have to admit I think it is rewriting the RFC
>> to conform with what the Linux client does.
>>
>> I think this para. from Sec. 13.10 of RFC-5661 is clear:
>>    The NFSv4.1 protocol only provides close-to-open file data cache
>>    semantics; meaning that when the file is closed, all modified data is
>>    written to the server.  When a subsequent OPEN of the file is done,
>>    the change attribute is inspected for a difference from a cached
>>    value for the change attribute.  For the case above, this means that
>>    a LAYOUTCOMMIT will be done at close (along with the data WRITEs) and
>>    will update the file's size and change attribute.  Access from
>>    another client after that point will result in the appropriate size
>>    being returned.
>>
>> It states "will be done". It doesn't say anything about UNSTABLE4 vs FILE_SYNC4.
>> (I think most POSIX-like clients would consider the fsync(2) syscall to require
>>  the same treatment as "close" above, but that is a POSIX-specific client issue.)
>> I can see the argument that, since there is no "must" in the statement, that a
>> client can choose not to do this, but that would also imply that the client will
>> need to live with the consequences of it.
>>
>> I think the second sentence of the first para. of the errata is bogus:
>> For file layouts, WRITEs to a Data Server that return a stable_how4 value of
>> FILE_SYNC4 guarantee that data and file system metadata are on stable
>> storage.  This means that a LAYOUTCOMMIT is not needed in order to make the
>> data and metadata visible to the metadata server and other clients.
>>
>> Why?
>> The FILE_SYNC4 was returned by the DS. This would imply the DS
>> has committed data and metadata to stable storage on the DS.
>> However, I am not aware of anything in RFC-5661 that would imply that
>> the Size, Time_Modify and Change attributes or anything else must have
>> been updated or in stable storage on the MDS at this time.
>>
>> If a server does not require LayoutCommit operations for correct behaviour
>> then it can simply reply NFS4ERR_NOTSUPP (as I believe the Netapp filer
>> does) and the client then no longer needs to do them.
>>
>> If there is somewhere in RFC-5661 that it is stated that LayoutCommits are
>> not required when the DS replies FILE_SYNC4, then I missed it and there
>> is a problem with RFC-5661 that needs to be addressed.
>>
>> Otherwise, sorry, but it seems that the bug is in the Linux client
>> implementation
>> and not RFC-5661.
>>
>> Is there a File Layout pNFS server implementation where the DSs return
>> FILE_SYNC4 that will break if the client does a LayoutCommit for this case?
>> (If so, then something may need to be done.)
>>
>> rick
>>
>>

_______________________________________________
nfsv4 mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/nfsv4

_______________________________________________
nfsv4 mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/nfsv4
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.