Re: nfs server issues
Richard Purdie <[email protected]> Sat, 04 Jul 2026 09:01:29 +0100
| Newsgroups | gmane.os.freebsd.devel.file-systems |
|---|---|
| Message-ID | <8055a84cc650253e1347545a18eed22ad83d8686.camel@linuxfoundation.org> |
On Fri, 2026-07-03 at 20:16 -0700, Rick Macklem wrote: > On Fri, Jul 3, 2026 at 10:00 AM Richard Purdie > <[email protected]> wrote: > > > > Hi, > > > > I was hoping someone might be able to give me some pointers on how to > > debug an NFS issue we keep running into. I asked on #freebsd and they > > suggested I should send an email. > > > > We have a FreeBSD 14.4 NFS server which we connect to with Linux nfs > > clients using NFS 4.2. The clients are many different Linux distros > > (Alma, Fedora, Ubuntu, Debian, OpenSUSE) of differing versions. > > > > We reboot the clients on Monday, by Friday, the clients start locking > > up showing "Remote I/O error" for something like "echo xxx > > > /nfs/testfile". Reads work, writes don't and it affects all of them > > eventually. Having things break just before or at the weekend is > > getting annoying! > > > > I was able to get one of the clients working again by terminating all > > the processes using the mount point and then remounting. We can > > obviously restart the nfs server but that just buys time until it > > happens again eventually. > > > > On the server, nfsd is using 5580% CPU and there appear to be very high > > open file (10-15k) and lock (100s) counts. The output from nfsdumpstate > > shows that: > > > > Flags OpenOwner Open LockOwner Lock Deleg OldDeleg Clientaddr ClientID > > CB 1 14837 208 208 0 0 fd01:172:16::1a 4c696e7578204e465376342e32207562756e7475323430342d766b2d31 > > CB 1 10945 136 136 0 0 fd01:172:16::242:2154 4c696e7578204e465376342e3220616c6d61382d766b2d31 > > CB 1 11006 157 157 0 0 fd01:172:16::14 4c696e7578204e465376342e322064656269616e31322d766b2d35 > > CB 1 10330 127 127 0 0 2a01:4f9:3070:2b44::2 4c696e7578204e465376342e32207562756e7475323230342d766b2d32 > > CB 1 14515 205 205 0 0 fd01:172:16::242:2162 4c696e7578204e465376342e322064656269616e31312d766b2d33 > > CB 1 13184 261 261 0 0 fd01:172:16::242:2152 4c696e7578204e465376342e32207562756e7475323531302d766b2d31 > > CB 1 14713 264 264 0 0 fd01:172:16:1::11 4c696e7578204e465376342e322064656269616e31322d766b2d32 > > CB 1 8275 123 123 0 0 fd01:172:16::18 4c696e7578204e465376342e3220616c6d61382d766b2d32 > > CB 1 14778 165 165 0 0 fd01:172:16::242:2158 4c696e7578204e465376342e322064656269616e31312d766b2d31 > > CB 1 2525 16 16 0 0 fd01:172:16::16 4c696e7578204e465376342e32207562756e7475323230342d766b2d61726d32 > > CB 1 13050 163 163 0 0 fd01:172:16::2:18 4c696e7578204e465376342e322064656269616e31322d766b2d38 > > CB 1 11290 155 155 0 0 fd01:172:16::242:2157 4c696e7578204e465376342e322064656269616e31322d766b2d31 > > CB 0 0 0 0 0 0 2a01:4f9:3081:38ea::2 4c696e7578204e465376342e3220706572662d64656269616e31322d766b > > CB 1 15116 191 191 0 0 fd01:172:16::1:26 4c696e7578204e465376342e32207562756e7475323430342d766b2d33 > > CB 1 9935 181 181 0 0 fd01:172:16::242:2150 4c696e7578204e465376342e32206665646f726134332d766b2d32 > > CB 1 14844 183 183 0 0 fd01:172:16::12 4c696e7578204e465376342e322064656269616e31322d766b2d33 > > CB 1 22771 356 356 0 0 fd01:172:16::37 4c696e7578204e465376342e32206f70656e737573653135362d766b2d31 > > CB 1 1 0 0 0 0 2a01:4f9:3b:4ec5::2 4c696e7578204e465376342e3220706572662d616c6d61382d766b > > CB 1 11093 163 163 0 0 fd01:172:16::242:2160 4c696e7578204e465376342e3220616c6d61392d766b2d32 > > CB 1 15780 179 179 0 0 fd01:172:16::1:25 4c696e7578204e465376342e32207562756e7475323430342d766b2d32 > > CB 1 16558 333 333 0 0 fd01:172:16::13 4c696e7578204e465376342e322064656269616e31322d766b2d34 > > CB 1 6297 93 93 0 0 2a01:4f9:3090:14cc::2 4c696e7578204e465376342e32206665646f726134342d766b2d31 > > CB 1 13180 188 188 0 0 fd01:172:16::242:2155 4c696e7578204e465376342e3220726f636b79382d766b2d31 > > CB 1 11373 196 196 0 0 fd01:172:16::19 4c696e7578204e465376342e3220726f636b79392d766b2d32 > > CB 1 14881 195 195 0 0 fd01:172:16::2:19 4c696e7578204e465376342e322064656269616e31322d766b2d39 > > CB 1 4068 63 63 0 0 fd01:172:16::38 4c696e7578204e465376342e32207562756e7475323630342d766b2d61726d31 > > CB 1 11171 157 157 0 0 fd01:172:16::1:28 4c696e7578204e465376342e32207562756e7475323230342d766b2d34 > > CB 1 14374 226 226 0 0 fd01:172:16::15 4c696e7578204e465376342e322064656269616e31322d766b2d36 > > CB 1 14584 290 290 0 0 fd01:172:16::1:27 4c696e7578204e465376342e32207562756e7475323230342d766b2d33 > > CB 1 19905 516 516 0 0 fd01:172:16::242:2143 4c696e7578204e465376342e32206f70656e737573653136302d766b2d312e796f63746f2e696f > > CB 1 0 0 0 0 0 2a01:4f9:3071:1625::2 4c696e7578204e465376342e32207562756e7475323630342d766b2d31 > > CB 1 13805 214 214 0 0 fd01:172:16::242:2153 4c696e7578204e465376342e32207562756e7475323530342d766b2d31 > > CB 1 11157 125 125 0 0 fd01:172:16::242:2165 4c696e7578204e465376342e32207562756e7475323230342d766b2d31 > > CB 1 6668 103 103 0 0 fd01:172:16:1::33 4c696e7578204e465376342e322064656269616e31332d766b2d61726d31 > > CB 1 8055 105 105 0 0 2a01:4f9:3071:1624::2 4c696e7578204e465376342e32207562756e7475323630342d766b2d32 > > CB 1 2669 23 23 0 0 fd01:172:16::36 4c696e7578204e465376342e32207562756e7475323430342d766b2d61726d32 > > CB 1 10060 554 554 0 0 fd01:172:16:1::29 4c696e7578204e465376342e322064656269616e31332d766b2d31 > > CB 1 11182 532 532 0 0 fd01:172:16:1::242:2159 4c696e7578204e465376342e322064656269616e31332d766b2d32 > > CB 1 11565 195 195 0 0 2a01:4f9:3081:33de::2 4c696e7578204e465376342e322073747265616d392d766b31 > > CB 1 16998 338 338 0 0 fd01:172:16::242:2161 4c696e7578204e465376342e32206f70656e737573653136302d766b2d32 > > CB 1 3597 36 36 0 0 fd01:172:16::35 4c696e7578204e465376342e32207562756e7475323430342d766b2d61726d31 > > CB 1 7276 188 188 0 0 2a01:4f9:3051:510f::2 4c696e7578204e465376342e322073747265616d31302d766b2d61726d31 > > CB 1 15294 210 210 0 0 fd01:172:16::2:17 4c696e7578204e465376342e322064656269616e31322d766b2d37 > > CB 1 9599 192 192 0 0 fd01:172:16::222:2160 4c696e7578204e465376342e3220616c6d6131302d766b2d31 > > > > The open file and lock counts are odd as there shouldn't be any! I also > > noticed some duplicate ClientIDs which I wasn't sure was an issue or > > not. One of the zero counts lines is the client I remounted on. > > > > I was hoping I could find out which files were open/locked but I > > couldn't work out if there is a way to do that. > > > > Any pointers on how to try and debug the issue further, or what the > > issue might be would be welcome. > First off, you didn't tell us anything about your server, so filling in some > information might help. > > - What kind of storage are you using (jbod connected, hardware raid or ??) > and what file system (ZFS, UFS or ??). > (For example, if you are using ZFS and the pool is getting more that 60% > full, you can expect terrible write performance from what little I know about > ZFS. There's a bunch of lore w.r.t. ZFS configuration and people who run > it in production mode should be able to help much more than I. > Some examples: Always set the ashift to at lease 12, so that storage > devices that pretend to have 512byte sectors but actually have 4K sectors > work well. Either live with a risk of data loss when the server crashes by > setting sync=disabled or put the ZIL on a mirrored pair of dedicated storage > devices with good write performance. Don't use hardware raid. > Avoid deduplication (I'm not sure if this is still the recommendation?). > (If you use ZFS and haven't yet read "man zfsprops", now is the time.) We're using zfs which is about 45% full, I'm not 100% sure of the hardware config but it is setup to be redundant as we need to avoid data loss. Michael (cc'd) can fill in more details there if it helps. There are web services also running off the system and those are all working fine, as are local filesystem accesses. Most of the time NFS clients would be very read heavy and the clients are all on a dedicated 10GBit network for speed. The read hits are large bursts of data (GBs) as fast as the clients can get it. > UFS - Far less to worry about. > Any other file system type (don't do it unless you really know what you into). > > If you've enabled delegations, try turning them off. > # sysctl vfs.nfsd.issue_delegations=0 I checked and those are already turned off. > (FreeBSD-14 is pretty old and that means that delegations are less likely > to work ok compared with an up-to-date FreeBSD version.) Are there any fixes in new versions relevant to this kind of problem? > As for the Opens/Locks. They remain until the client(s) unlock or close them. > (So, no idea why you are seeing this.) To be honest, 15K opens should just > be noise. If you are seeing 100000 of them, then I'd suspect it might cause > trouble. If all the processes on the client have exited, that should mean everything should be closed/unlocked? Or is there some pattern which allows an open/lock to leak? It certainly seems like there is some kind of leaking going on but where I don't know, hence the question I started with! I don't have a good feel for which kinds of numbers are problematic either which doesn't help. > On the NFS server... > # ps axHl > ps.txt > - and then look in ps.txt. First off, since you seem to be runniing with > high CPU, look for the lines with lots of time accrued and the ones > with a STAT of R (those are the CPU hogs). If the "nfsd:server.." threads > are running up their TIME, then it might be caused by the # of Opens/Locks > or it might be that you don't have enough of them. As kernel threads, they > don't use a lot of resources and the default of 8 per CPU is very low. > (You can use the "-n" option for the nfsd, set by nfs_server_flags in > /etc/rc.conf to bump the number up. Unless the server is very small > hardware, having a few hundred of them shouldn't be a problem.) > ("top" can also be useful to see where the CPU is being used up and > whether or not the server is becoming memory constrained.) We've had to reset things to get services back online but I'll take a look next time this happens, thanks. I was seeing the high CPU in top. > If you are using ZFS, do definitely want advice from the ZFS folk. > (Although FreeBSD discourages cross-posting, you might try a > post on freebsd-stable@. Garrett Wollman reads that one and he > runs about 20 FreeBSD NFS servers at MIT, for example.) > > I'm kinda the NFS guy, but haven't run any production stuff in > quite a few years, so I can't really help with performance related > stuff. Thanks for the reply, I guess at least doesn't appear to be anything obvious! Cheers, Richard