Re: kern/60457: Hang during shutdown/reboot (with zfs root with failed component)
"Andrius V via gnats" <[email protected]>
| Newsgroups | gmane.os.netbsd.bugs |
|---|---|
| Message-ID | <[email protected]> |
The following reply was made to PR kern/60457; it has been noted by GNATS. From: Andrius V <[email protected]> To: [email protected] Cc: [email protected], [email protected] Subject: Re: kern/60457: Hang during shutdown/reboot (with zfs root with failed component) Date: Sat, 18 Jul 2026 14:18:21 +0300 On Fri, Jul 17, 2026 at 10:55=E2=80=AFPM [email protected] via gnats <[email protected]> wrote: > > >Number: 60457 > >Category: kern > >Synopsis: Hang during shutdown/reboot (with zfs root with failed c= omponent) > >Confidential: no > >Severity: serious > >Priority: high > >Responsible: kern-bug-people > >State: open > >Class: sw-bug > >Submitter-Id: net > >Arrival-Date: Fri Jul 17 19:55:00 +0000 2026 > >Originator: Jeff Rizzo > >Release: 11.0_RC6 > >Organization: > >Environment: > NetBSD cherimoya.tastylime.net 11.0_RC6 NetBSD 11.0_RC6 (GENERIC) #6: Fri= Jul 17 18:46:30 UTC 2026 [email protected]:/tank/netbsd/netbsd-11/ob= j/sys/arch/amd64/compile/GENERIC amd64 > >Description: > This machine consistently gets stuck sometime _after_ hw.cnmagic is effec= tive (I can't get into ddb) and _before_ everything is undone for shutdown,= and it will indefinitely (for at least an hour, anyway) hang: > > > Jul 17 18:56:59 [ 156.7362082] syncing disks... done > [ 214.3661561] ipmi0: normal state on 'FAN6_OUTLET' > [ 214.4161559] dk11 at sd2 (c3440c02-2fec-11f1-8000-000000000000) deleted > [ 214.4261558] dk10 at sd2 (c3440c01-2fec-11f1-8000-000000000000) deleted > [ 214.4261558] sd2: detached > [ 214.4361555] sd1: detached > [ 214.4361555] dk1 at sd0 (NBImgRoot) deleted > [ 214.4361555] dk0 at sd0 (EFI system) deleted > [ 214.4461557] sd0: detached > [ 214.4461557] scsibus2: detached > [ 214.4461557] scsibus1: detached > [ 214.4561557] scsibus0: detached > [ 214.4561557] dk7 at wd2 (zfs-rpool-wd2) deleted > [ 214.4661550] dk6 at wd2 (rf-component-wd2) deleted > [ 214.4661550] wd2: detached > [ 214.4661550] uhub5: detached > [ 214.4761557] uhub4: detached > [ 214.4761557] uhub3: detached > [ 214.4761557] uhub2: detached > [ 214.4861557] uhub0: detached > [ 214.4861557] igphy1: detached > lagg0: [ 214.4861557] disable distributing on wm1 > [ 244.3661277] ipmi0: normal state on 'FAN6_INLET' > [ 244.3761278] ipmi0: normal state on 'FAN5_OUTLET' > [ 244.3761278] ipmi0: normal state on 'FAN5_INLET' > [ 1354.3650987] ipmi0: normal state on 'FAN4_OUTLET' > [ 1354.3750988] ipmi0: normal state on 'FAN3_OUTLET' > [ 1894.3645934] ipmi0: critical under limit on 'FAN4_OUTLET' > [ 1894.3745935] ipmi0: critical under limit on 'FAN3_OUTLET' > > ...the messages from ipmi0 continue, but no more disks detach. > > I have other machines with zfs root which don't exhibit this problem, but= this machine currently has a failed component (and working systems don't h= ave failed components): > > NAME STATE READ WRITE CKSUM > rpool DEGRADED 0 0 0 > raidz2-0 DEGRADED 0 0 0 > dk3 ONLINE 0 0 0 > dk5 ONLINE 0 0 0 > 13185259948679916708 OFFLINE 0 0 0 was /dev/dk= 7 > dk9 ONLINE 0 0 0 > > Here's what disks look like after a boot: > > [ 5.387244] dk0 at sd0: "EFI system", 262144 blocks at 2048, type: ms= dos > [ 5.403522] dk1 at sd0: "NBImgRoot", 976502784 blocks at 264192, type= : ffs > [ 5.696224] dk2 at wd0: "rf-component-wd0", 4194304 blocks at 128, ty= pe: raidframe > [ 5.713521] dk3 at wd0: "zfs-rpool-wd0", 31247564632 blocks at 419443= 2, type: zfs > [ 5.835644] dk4 at wd1: "rf-component-wd1", 4194304 blocks at 128, ty= pe: raidframe > [ 5.844158] dk5 at wd1: "zfs-rpool-wd1", 31247564632 blocks at 419443= 2, type: zfs > [ 5.963521] dk6 at wd2: "rf-component-wd2", 4194304 blocks at 128, ty= pe: raidframe > [ 5.973521] dk7 at wd2: "zfs-rpool-wd2", 31247564632 blocks at 419443= 2, type: zfs > [ 6.093521] dk8 at wd3: "cherimoya-swap-wd3", 4194304 blocks at 128, = type: swap > [ 6.103521] dk9 at wd3: "zfs-rpool-wd3", 31247564632 blocks at 419443= 2, type: zfs > [ 7.627989] dk10 at sd2: "c3440c01-2fec-11f1-8000-000000000000", 2621= 44 blocks at 2048, type: msdos > [ 7.653664] dk11 at sd2: "c3440c02-2fec-11f1-8000-000000000000", 4853= 760 blocks at 264192, type: ffs > [ 8.203517] raid0: Components: /dev/dk2 /dev/dk4 > > (wd2 has a ton of errors and is currently awaiting replacement) > > I'd love suggestions on more troubleshooting steps - I can't get into DDB= at the point of hang; I could sprinkle debugging printfs but so > far they've just told me that x86_reset() isn't reached... > >How-To-Repeat: > boot 11.0_RC6 with zfs root as described in the netbsd wiki (pivot root) > > shutdown -r now > > get as far as detaching some-but-not-all dk(4) devices > > wait indefinitely for IPMI power reset to finish reboot > >Fix: > none given > Hi, Since, I see lagg(4) in the logs, I think it may cause hang, not zfs. I have a similar problem on my NAS system. If it is not being brought down first, shutdown or reboot hangs, I've just never filed a bug report for that. If you have this option, try to bring down lagg with ifconfig lagg0 down and reboot. Regards, Andrius V