Re: [nfsv4] Is NFSv4.2's clone_blksize per-file or per-file-system?

Rick Macklem <[email protected]> Sat, 9 Aug 2025 14:02:30 -0700
Newsgroups gmane.linux.nfs,gmane.ietf.nfsv4
Message-ID <CAM5tNy4kPWfPHHRVr712AG=g5wJ+fThG9KFX_9JoT85seTSE=g@mail.gmail.com>
On Sat, Aug 9, 2025 at 1:12=E2=80=AFPM David Noveck <[email protected]> =
wrote:
>
>
>
> On Friday, August 8, 2025, Rick Macklem <[email protected]> wrote:
>>
>> On Fri, Aug 8, 2025 at 8:38=E2=80=AFPM Trond Myklebust <[email protected]=
m> wrote:
>> >
>> >
>> >
>> > On Fri, Aug 8, 2025 at 9:47=E2=80=AFPM Rick Macklem <rick.macklem@gmai=
l.com> wrote:
>> >>
>> >> Hi,
>> >>
>> >> I'm looking at RFC7862 and I cannot find where it
>> >> states if the clone_blksize attribute is per-file or
>> >> per-file-system.
>> >>
>> >> If it is not in the RFC, which do others think it is?
>
>
>  Before you told us about ZFS,  I would have assumed per-fs.
>
> Given the uncertainty in the spec, you may wind up dealing clients that a=
ssume it is per-fs.
>
> Although this is not a  catastrophe, you might want to file an errata rep=
ort explaining the negative consequences of assuming this is per-fs. It won=
't get into a spec for a long while but it does provide as much warning as =
you can right now .
>
>
>
>>
>> >> (Or maybe, if you have implemented CLONE,
>> >> which does your implementation assume?)
>> >>
>> >> In case you are wondering why I am asking,
>> >> it turns out that files in a ZFS volume can have
>> >> different block sizes. (It can be changed after the
>> >> file system is created.)
>
>
> The guy who allowed that probably thinks it's a helpful feature.  Sigh!
It's not just a feature change after creation, it turns out to be based
on file size as well.  A small file gets 512 and a larger one gets a full r=
ecord
(128K on my test system).

And, yes, block cloning requires alignment with 512bytes or 128Kbytes
depending on the file.

I can return 128K for clone_blksize and that will (sub-optimally) handle
the 512byte case, but I think it is also possible to increase the record
size from 128K-> after the file system has files in it.

I'll take a look at the Linux client to try and see if/how it uses
clone_blksize.  I need to decide if I should always return 128K
(or whatever the full recordsize is) or 512 for the small files.

Thanks for the comments, rick

>
>> >>
>
>
>>
>> >> Thanks, rick
>> >>
>> >
>> > Yes, but since ZFS only supports filesystem level snapshots, and not a=
ctual file cloning, does that matter to anything?
>> ZFS now has a feature it calls block cloning, which does clone file rang=
es.
>> (It was only added recently. I do not know if the Linux port uses it yet=
?)
>>
>> rick
>>
>> >
>> > Cheers
>> >   Trond
>>
>> _______________________________________________
>> nfsv4 mailing list -- [email protected]
>> To unsubscribe send an email to [email protected]