Re: [Ms-nfs41-client-devel] FreeBSD 16.0-CURRENT nfsd causes kernel panic when ms-nfs41-client runs WinZFS from file

Aurélien Couderc <[email protected]> Sat, 27 Jun 2026 17:13:32 +0200
Newsgroups gmane.os.freebsd.devel.hackers
Message-ID <CA+1jF5pr9DN0tZbVYoX7octCZxeowaVFH0ySJ3wgrymb2fAZeQ@mail.gmail.com>
On Fri, Jun 26, 2026 at 5:38 PM Rick Macklem <[email protected]> wrote:
>
> On Fri, Jun 26, 2026 at 8:20 AM Rick Macklem <[email protected]> wrote:
> >
> > On Fri, Jun 26, 2026 at 2:51 AM Aurélien Couderc
> > <[email protected]> wrote:
> > >
> > > On Wed, Jun 24, 2026 at 8:07 PM Ryan Libby <[email protected]> wrote:
> > > >
> > > > On Wed, Jun 24, 2026 at 10:00 AM Roland Mainz <[email protected]> wrote:
> > > > >
> > > > > On Tue, Jun 16, 2026 at 3:03 PM Rick Macklem <[email protected]> wrote:
> > > > > >
> > > > > > On Tue, Jun 16, 2026 at 1:42 AM Aurélien Couderc
> > > > > > <[email protected]> wrote:
> > > > > > >
> > > > > > > FreeBSD NFSD causes repeated kernel panics with ms-nfs41-client, but
> > > > > > > only if we run a Windows OpenZFS zpool from a NFS filesystem.
> > > > > > >
> > > > > > > Steps to use OpenZFS on NFS (I've send the instructions to Roland):
> > > > > > > cd nfs_subdir_on_freebsd26
> > > > > > > truncate -s $((512*1024*1024*1024)) poolfile1.bin
> > > > > > > chmod a+rw poolfile1.bin
> > > > > > > zpool.exe create -o ashift=15 -o autotrim=on -O
> > > > > > > casesensitivity=insensitive -O driveletter=Z nfspool1
> > > > > > > "\\\\?\\$(cygpath -w "$PWD/poolfile1.bin")"
> > > > > > >
> > > > > > > The ZFS filesystem on Windows works for a while until the FreeBSD
> > > > > > > kernel drops into the kernel debugger prompt, warning that NFSD had
> > > > > > > become unresponsive.
> > > > > > >
> > > > > > > How can I debug this?
> > > > > > Capture packets until the FreeBSD server hangs. On the FreeBSD
> > > > > > server:
> > > > > > # tcpdump -s 0 -w out.pcap host <windows-client>
> > > > > > --> Hopefully the hang doesn't result in a huge out.pcap, but..
> > > > > >
> > > > > > Create a bug report on bugs.freebsd.org and either attach
> > > > > > out.pcap to it or email me w.r.t. where I can grab it.
> > > > > >
> > > > > > I need to be able to look at out.pcap in wireshark.
> > > > > >
> > > > > > Once the FreeBSD server appears hung..
> > > > > > - If you can get onto it before it panics.
> > > > > > # ps axHl > proc.txt
> > > > > > # procstat -kk -a > stat.txt
> > > > > > and attach both proc.txt and stat.txt to the bug report as well.
> > > > >
> > > > > I can reproduce the issue, including the panic. Long before the
> > > > > FreeBSD kernel panic the Windows NFSv4.2 client (ms-nfs41-client) can
> > > > > mount but not cd into the nfs shares and hang forever.
> > > > > Sometimes the TCP connection gets dropped, but reconnect also does hang.
> > > > >
> > > > > Steps to reproduce on the ms-nfs41-client side:
> > > > > 1. OpenZFS Installation:
> > > > > ---- snip ----
> > > > > wget 'https://github.com/openzfsonwindows/openzfs/releases/download/zfswin-2.4.1rc12/OpenZFSOnWindows-debug-2.4.1rc12.exe'
> > > > > chmod a+x OpenZFSOnWindows-debug-2.4.1rc12.exe
> > > > > ./OpenZFSOnWindows-debug-2.4.1rc12
> > > > > <reboot>
> > > > > ---- snip ----
> > > > >
> > > > > 2. Mount filesystem on Windows machine:
> > > > > Mount nfs share via $ /sbin/nfs_globalmount -o rw 'F:'
> > > > > 'nfs://11.22.33.44//path/to/mountdoom' ...
> > > > >
> > > > > 3. Create ZFS pool and assign it the drive letter "U:":
> > > > > ---- snip ----
> > > > > # cd into NFS subdir
> > > > > truncate -s $((128*1024*1024*1024)) poolfile1.bin
> > > > > chmod a+rw poolfile1.bin
> > > > > zpool.exe create -o ashift=15 -o autotrim=on -O xattr=dir -O
> > > > > casesensitivity=insensitive -O driveletter=U nfspool1
> > > > > "\\\\?\\$(cygpath -w "$PWD/poolfile1.bin")"
> > > > > ---- snip ----
> > > > >
> > > > > 4. Do some stuff on U:
> > > > >
> > > > > 5. Unmount+remove pool:
> > > > > ---- snip ----
> > > > > zpool export nfspool1
> > > > > ---- snip ----
> > > > >
> > > > > 6. Import pool and mount it again:
> > > > > ---- snip ----
> > > > > # cd into NFS subdir
> > > > > zpool.exe import -d "$(cygpath -w "$PWD")" nfspool1
> > > > > ---- snip ----
> > > > >
> > > > > Repeat steps 4-6 until it hangs in importing the pool (step 6) - a cd
> > > > > into the nfs dir should hang then too.
> > > > >
> > > > > I've attached the requested "proc.txt" and "stat.txt" from my FreeBSD
> > > > > 16.0 machine (uname -a == 'FreeBSD freebsd16002 16.0-CURRENT FreeBSD
> > > > > 16.0-CURRENT main-n285769-1d24638d3e88 GENERIC amd64
> > > > > ') to this email, and I uploaded a kernel dump to
> > > > > http://nrubsig.org/freebsd16002_nfs_kerneldeadlockpanicwithnfs20260624.tar.bz2
> > > > > (this will be deleted in 32 days) ... IMO it would be nice if someone
> > > > > could post a kernel stacktrace here...
> > > > >
> > > > > ----
> > > > >
> > > > > Bye,
> > > > > Roland
> > > > > --
> > > > >   __ .  . __
> > > > >  (o.\ \/ /.o) [email protected]
> > > > >   \__\/\/__/  MPEG specialist, C&&JAVA&&Sun&&Unix programmer
> > > > >   /O /==\ O\  TEL +49 641 3992797
> > > > >  (;O/ \/ \O;)
> > > >
> > > > Drive-by suggestion here.
> > > >
> > > > I don't know how to debug your vmcore without the symbols.  If you
> > > > have them you could make them available.
> > > >
> > > > Alternatively, if you can repro at will, here's a suggestion to get
> > > > some text-based debugging.
> > > >
> > > > First, configure textdump(4):
> > > > sysctl debug.ddb.capture.bufsize=$(sysctl -n debug.ddb.capture.maxbufsize)
> > > > ddb script "kdb.enter.default=textdump set; capture on; show page;
> > > > show pageq; ps; show allchains; show lockedvnods; alltrace; textdump
> > > > dump; reset"
> > > >
> > > > That should arm a textdump to collect certain information and reset.
> > > > Then run your repro.  You should be able to either wait for the panic
> > > > or trigger the dump and reboot yourself once the system is wedged with
> > > > sysctl debug.kdb.enter=1
> > > >
> > > > After reboot you should see /var/crash/textdump.tar.*.  That should
> > > > aid in debugging without needing your kernel symbols.
> > >
> > > Does that mean that the vmcore.0 file Roland provided is useless? Rick? Roland?
> > I haven't looked at the core yet. (I will later to-day.)
> >
> > Could you try turning off delegations in the NFS server?
> > # sysctl vfs.nfsd.issue_delegations=0
> >
> > A lot of threads are waiting on a vnode lock while one thread is trying to
> > CB_RECALL. (Getting rid of delegations will get rid of the CB_RECALLs.)
> Btw, at a glance, the hang appears to be one nfsd thread waiting for a callback
> session slot, while holding a vnode lock.
> --> This could be the client not being responsive to callbacks
> (failing to reply,
>      resulting in the slots not being free'd or ???).
> It looks like the server should set a timeout on waiting for an
> available callback
> slot, so that the nfsd thread will fail to do the callback instead of
> hanging the
> system while trying to do one.

Maybe this is a "netlock", i.e. deadlock over network between client
and server, as the libtirpc sarcastically says "Now go hang yourself."

If I access files on a ZFS filesystem exported from FreeBSD I often
get a NFS4ERR_DELAY (why?). What happens if a CB callback is sent
while the client is still in a NFS4ERR_DELAY loop?

Aurélien
-- 
Aurélien Couderc <[email protected]>
Big Data/Data mining expert, chess enthusiast