Re: [Ms-nfs41-client-devel] FreeBSD 16.0-CURRENT nfsd causes kernel panic when ms-nfs41-client runs WinZFS from file
Aurélien Couderc <[email protected]> Fri, 26 Jun 2026 11:50:50 +0200
| Newsgroups | gmane.os.freebsd.devel.hackers |
|---|---|
| Message-ID | <CA+1jF5ojmcP_OkxeMhKtY3amogLvqCRO4jYx_JE_uNXgjmFG=Q@mail.gmail.com> |
On Wed, Jun 24, 2026 at 8:07 PM Ryan Libby <[email protected]> wrote: > > On Wed, Jun 24, 2026 at 10:00 AM Roland Mainz <[email protected]> wrote: > > > > On Tue, Jun 16, 2026 at 3:03 PM Rick Macklem <[email protected]> wrote: > > > > > > On Tue, Jun 16, 2026 at 1:42 AM Aurélien Couderc > > > <[email protected]> wrote: > > > > > > > > FreeBSD NFSD causes repeated kernel panics with ms-nfs41-client, but > > > > only if we run a Windows OpenZFS zpool from a NFS filesystem. > > > > > > > > Steps to use OpenZFS on NFS (I've send the instructions to Roland): > > > > cd nfs_subdir_on_freebsd26 > > > > truncate -s $((512*1024*1024*1024)) poolfile1.bin > > > > chmod a+rw poolfile1.bin > > > > zpool.exe create -o ashift=15 -o autotrim=on -O > > > > casesensitivity=insensitive -O driveletter=Z nfspool1 > > > > "\\\\?\\$(cygpath -w "$PWD/poolfile1.bin")" > > > > > > > > The ZFS filesystem on Windows works for a while until the FreeBSD > > > > kernel drops into the kernel debugger prompt, warning that NFSD had > > > > become unresponsive. > > > > > > > > How can I debug this? > > > Capture packets until the FreeBSD server hangs. On the FreeBSD > > > server: > > > # tcpdump -s 0 -w out.pcap host <windows-client> > > > --> Hopefully the hang doesn't result in a huge out.pcap, but.. > > > > > > Create a bug report on bugs.freebsd.org and either attach > > > out.pcap to it or email me w.r.t. where I can grab it. > > > > > > I need to be able to look at out.pcap in wireshark. > > > > > > Once the FreeBSD server appears hung.. > > > - If you can get onto it before it panics. > > > # ps axHl > proc.txt > > > # procstat -kk -a > stat.txt > > > and attach both proc.txt and stat.txt to the bug report as well. > > > > I can reproduce the issue, including the panic. Long before the > > FreeBSD kernel panic the Windows NFSv4.2 client (ms-nfs41-client) can > > mount but not cd into the nfs shares and hang forever. > > Sometimes the TCP connection gets dropped, but reconnect also does hang. > > > > Steps to reproduce on the ms-nfs41-client side: > > 1. OpenZFS Installation: > > ---- snip ---- > > wget 'https://github.com/openzfsonwindows/openzfs/releases/download/zfswin-2.4.1rc12/OpenZFSOnWindows-debug-2.4.1rc12.exe' > > chmod a+x OpenZFSOnWindows-debug-2.4.1rc12.exe > > ./OpenZFSOnWindows-debug-2.4.1rc12 > > <reboot> > > ---- snip ---- > > > > 2. Mount filesystem on Windows machine: > > Mount nfs share via $ /sbin/nfs_globalmount -o rw 'F:' > > 'nfs://11.22.33.44//path/to/mountdoom' ... > > > > 3. Create ZFS pool and assign it the drive letter "U:": > > ---- snip ---- > > # cd into NFS subdir > > truncate -s $((128*1024*1024*1024)) poolfile1.bin > > chmod a+rw poolfile1.bin > > zpool.exe create -o ashift=15 -o autotrim=on -O xattr=dir -O > > casesensitivity=insensitive -O driveletter=U nfspool1 > > "\\\\?\\$(cygpath -w "$PWD/poolfile1.bin")" > > ---- snip ---- > > > > 4. Do some stuff on U: > > > > 5. Unmount+remove pool: > > ---- snip ---- > > zpool export nfspool1 > > ---- snip ---- > > > > 6. Import pool and mount it again: > > ---- snip ---- > > # cd into NFS subdir > > zpool.exe import -d "$(cygpath -w "$PWD")" nfspool1 > > ---- snip ---- > > > > Repeat steps 4-6 until it hangs in importing the pool (step 6) - a cd > > into the nfs dir should hang then too. > > > > I've attached the requested "proc.txt" and "stat.txt" from my FreeBSD > > 16.0 machine (uname -a == 'FreeBSD freebsd16002 16.0-CURRENT FreeBSD > > 16.0-CURRENT main-n285769-1d24638d3e88 GENERIC amd64 > > ') to this email, and I uploaded a kernel dump to > > http://nrubsig.org/freebsd16002_nfs_kerneldeadlockpanicwithnfs20260624.tar.bz2 > > (this will be deleted in 32 days) ... IMO it would be nice if someone > > could post a kernel stacktrace here... > > > > ---- > > > > Bye, > > Roland > > -- > > __ . . __ > > (o.\ \/ /.o) [email protected] > > \__\/\/__/ MPEG specialist, C&&JAVA&&Sun&&Unix programmer > > /O /==\ O\ TEL +49 641 3992797 > > (;O/ \/ \O;) > > Drive-by suggestion here. > > I don't know how to debug your vmcore without the symbols. If you > have them you could make them available. > > Alternatively, if you can repro at will, here's a suggestion to get > some text-based debugging. > > First, configure textdump(4): > sysctl debug.ddb.capture.bufsize=$(sysctl -n debug.ddb.capture.maxbufsize) > ddb script "kdb.enter.default=textdump set; capture on; show page; > show pageq; ps; show allchains; show lockedvnods; alltrace; textdump > dump; reset" > > That should arm a textdump to collect certain information and reset. > Then run your repro. You should be able to either wait for the panic > or trigger the dump and reboot yourself once the system is wedged with > sysctl debug.kdb.enter=1 > > After reboot you should see /var/crash/textdump.tar.*. That should > aid in debugging without needing your kernel symbols. Does that mean that the vmcore.0 file Roland provided is useless? Rick? Roland? Aurélien -- Aurélien Couderc <[email protected]> Big Data/Data mining expert, chess enthusiast