Re: nfs server issues
Rick Macklem <[email protected]> Wed, 8 Jul 2026 07:57:42 -0700
| Newsgroups | gmane.os.freebsd.devel.file-systems |
|---|---|
| Message-ID | <CAM5tNy5nW2vacSeukmcp-o1qn_R25VgfSUAnSbjeaUYERwPiBg@mail.gmail.com> |
On Wed, Jul 8, 2026 at 2:28 AM Richard Purdie <[email protected]> wrote: > > On Tue, 2026-07-07 at 15:54 -0700, Michael Halstead wrote: > > On Tue, Jul 7, 2026 at 1:52 AM Richard Purdie <[email protected]> wrote: > > > On Sat, 2026-07-04 at 12:41 -0700, Rick Macklem wrote: > > > > On Sat, Jul 4, 2026 at 1:05 AM Richard Purdie > > > > <[email protected]> wrote: > > > > > > > > > > nfsstat -E -s > > > > > Server Info: > > > > > Getattr Setattr Lookup Readlink Read Write > > > > > 17354100116 1674996998 5136769057 40544 18429670397 986443058 > > > > > Create Remove Rename Link Symlink Mkdir > > > > > 0 98646613 119865835 17071853 0 0 > > > > > Rmdir Readdir RdirPlus Access Mknod Fsstat > > > > > 0 1655544711 0 4129535844 0 0 > > > > > FSinfo pathConf Commit LookupP SetClId SetClIdCf > > > > > 0 0 8723436 0 0 0 > > > > One more thing to note. The # of Commits is much smaller than the # of Writes > > > > on the top line. This "hints" that a lot of the Writes are being done File_sync. > > > > (To check that, you'd need to capture packets and look at them in wireshark.) > > > > > > I've captured data > > at https://valkyrie.yocto.io/pub/non-release/mtest/nfs_capture2.pcap. > > zst from a fresh reboot running the testlock.py for about 100MB of > > data and umounting the filesystems. Please let me know if this > > reveals the issue or what else I can offer. > > Thanks Michael. I think that didn't have enough parallel test scripts > running. This is one I just captured with more parallelism (4 clients), > just tracing one of them: > > https://valkyrie.yocto.io/pub/non-release/mtest/rp_capture.pcap.zst > > This one has a lot of NFS4ERR_STALE in it which I suspect is what we > were looking for... Yep. This one definitely shows RPCs that would do Close failing when PutFH replies NFS4ERR_STALE. (I looked at the previous one and couldn't see any problems in it, although the previous one does have the weird case where the same Lock is acquired twice. Both succeed because it is the same client and lock_owner. I am going to check to see what the RFCs say w.r.t. duplicate lock operations. Like Open, the Lock ops were modelled after what Windows wants, but were intended to handle POSIX lock semantics as well. Of course, no one implements mandatory locking and the protocol doesn't even allow the client to figure out which the server is doing, but that's another story.;-) I do think I can come up with a simple solution for this: - Since the operations that get rid of state all require a CFH, which must be set by PutFH.. I cannot see any reason to keep state once the PutFH is returning NFS4ERR_STALE for the file. --> So, I think I can just free all the state for that FH at that time. (The state is keyed on FH and not a vnode, so the PutFH has the required FH.) --> I'll post a patch here later to-day. Hopefully you guys can run a kernel from patched sources to test it. Thanks for the packet trace and help tracking this down, rick > > Cheers, > > Richard