Re: nfs server issues
Richard Purdie <[email protected]> Wed, 08 Jul 2026 10:58:07 +0100
| Newsgroups | gmane.os.freebsd.devel.file-systems |
|---|---|
| Message-ID | <b8a347c3688155df0a404d749e2021df5cc5146f.camel@linuxfoundation.org> |
On Tue, 2026-07-07 at 06:54 -0700, Rick Macklem wrote: > On Tue, Jul 7, 2026 at 1:52 AM Richard Purdie > <[email protected]> wrote: > > > > On Sat, 2026-07-04 at 12:41 -0700, Rick Macklem wrote: > > > On Sat, Jul 4, 2026 at 1:05 AM Richard Purdie > > > <[email protected]> wrote: > > > > > > > > nfsstat -E -s > > > > Server Info: > > > > Getattr Setattr Lookup Readlink Read Write > > > > 17354100116 1674996998 5136769057 40544 18429670397 986443058 > > > > Create Remove Rename Link Symlink Mkdir > > > > 0 98646613 119865835 17071853 0 0 > > > > Rmdir Readdir RdirPlus Access Mknod Fsstat > > > > 0 1655544711 0 4129535844 0 0 > > > > FSinfo pathConf Commit LookupP SetClId SetClIdCf > > > > 0 0 8723436 0 0 0 > > > One more thing to note. The # of Commits is much smaller than the # of Writes > > > on the top line. This "hints" that a lot of the Writes are being done File_sync. > > > (To check that, you'd need to capture packets and look at them in wireshark.) > > > > > > If a lot of the Writes are File_sync (which means the server must commit the > > > data/metadata changes to stable storage before replying), doing the "mirrored > > > pair of storage devices dedicated to the ZIL" could be what you need to get > > > writes to work well. (As I mentioned, the "cheat" alternative is to > > > set "sync=disabled", > > > but that runs a risk of data loss when the NFS server crashes/reboots.) > > > > I was able to find a way to increase the Open/Lock counts on the > > server. Most of the locking we do on the clients is from a tool called > > Bitbake and the locking is isolated to specific function. I maintain > > Bitbake and wrote/maintain the locking code over the <too many> years. > > I copied and pasted the relevant code straight from bitbake into a test > > script: > > > > https://valkyrie.yocto.io/pub/non-release/rptest/testlock.py > > > > (code is from > > https://git.openembedded.org/bitbake/tree/lib/bb/utils.py) > > > > Obviously it could easily be simplified more but I wanted to test our > > actual code. The lock file name/path is hardcoded at the end of the > > script but easily changed. > > > > I ran that script on three of the NFS clients for around 5 minutes. On > > the server, the counts went from: > > > > nfsdumpstate | egrep 'fd01:172:16::242:2157|fd01:172:16::12|fd01:172:16:1::11 ' > > CB 2 6307 55 55 0 0 fd01:172:16::242:2157 4c696e7578204e465376342e322064656269616e31322d766b2d31 > > CB 3 6543 77 77 0 0 fd01:172:16::12 4c696e7578204e465376342e322064656269616e31322d766b2d33 > > CB 2 6571 83 83 0 0 fd01:172:16:1::11 4c696e7578204e465376342e322064656269616e31322d766b2d32 > > > > to: > > > > nfsdumpstate | egrep 'fd01:172:16::242:2157|fd01:172:16::12|fd01:172:16:1::11 ' > > CB 2 6462 65 65 0 0 fd01:172:16::242:2157 4c696e7578204e465376342e322064656269616e31322d766b2d31 > > CB 3 6679 83 83 0 0 fd01:172:16::12 4c696e7578204e465376342e322064656269616e31322d766b2d33 > > CB 2 6691 88 88 0 0 fd01:172:16:1::11 4c696e7578204e465376342e322064656269616e31322d766b2d32 > > > > I've then stopped the scripts and waited and the counts just stayed > > there. > > > > I'd be very interested if someone else could reproduce that and if so, > > explain what is happening and if it is an issue? > The next step would be to run a test like the above while capturing packets. > # tcpdump -s 0 -w out.pcap host <nfs-client> > on the server should do it. > > Then pull out.pcap into wireshark and look for LockU and Close operations. > (If you want, you can stick out.pcap somewhere and I can look at it.) I've shared one in reply to Michael's message. > If the operations are failing (due to ESTALE, for example), then I could > consider allowing these operations to succeed even when the file no > longer exists. Both LockU and Close take a CFH (current file handle) > as arguments, which suggests a valid file handle is required (per the RFC) > but I can see the argument that "the file is gone, so allow it anyhow". It does look like they're failing with ESTALE to me but it is after a PUTFH and from your comments in a different reply, that might complicate things. The client is definitely trying to close them. > I will also make the generic statement... > NFS is not (and will never be, from what I've seen) a POSIX compliant > file system (due to limitations in the protocol design). As such, some POSIX > applications will not work well over NFS. I'm not that concerned whether it is POSIX or not, I'd just prefer the server not to leak file handles as it appears to be doing at the moment! :) > > I can't prove the write hangs are related to the counts but it does > > seem there is some issue there regardless as they don't make sense. > > As I noted before, 15K Opens should not have any real effect. They are > just small in-memory data structures on linked lists anchored in hash tables. > (They might result in slightly slower response, due to the time taken for a > linked list traversal. Making the hash tables larger should resolve that and > only results in somewhat bigger tables, which is not a problem, given the > amount of memory in modern systems.) > --> To give you a clue, long ago when I developed this stuff, I tested 100 > NFS clients against a server with only a couple of Gbytes of ram > (yes, a 32bit x86 around year 2000). I got it up to 500K Opens and it > still worked). Would the leaking opens/locks run us into vfs.nfsd.v4statelimit ? If we did hit that, could that somehow stop writes but not reads? It does look a bit like the opens/locks may be causing the cpu load issue too and whilst increasing the hash table sizes may help with that (and we should do that), if we can avoid the overhead entirely... Cheers, Richard