https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297620
Bug ID: 297620
Summary: cxl: enabling TOE causes Fatal trap 12: page fault
while in kernel mode
Product: Base System
Version: 15.0-RELEASE
Hardware: Any
OS: Any
Status: New
Severity: Affects Only Me
Priority: ---
Component: kern
Assignee: [email protected]
Reporter: [email protected]
CC: [email protected]
using a T540-CR, configured with 16 VFs, which are attached to the host.
hardware is an AMD Ryzen 9 5950X with a B450 chipset, one port of the T540 is
connected via a DAC to a MikroTik switch.
t5nex0: <Chelsio T540-CR> mem
0xfa300000-0xfa37ffff,0xf9000000-0xf9ffffff,0xfac04000-0xfac05fff at device 0.4
on pci1
t5nex0: Disabled No Snoop/Relaxed Ordering on pcib1
cxl0: <port 0> on t5nex0
cxl0: Ethernet address: 00:07:43:3f:e7:60
cxl0: 16 txq, 8 rxq (NIC); 8 txq (TOE), 2 rxq (TOE)
cxl1: <port 1> on t5nex0
cxl1: Ethernet address: 00:07:43:3f:e7:68
cxl1: 16 txq, 8 rxq (NIC); 8 txq (TOE), 2 rxq (TOE)
cxl2: <port 2> on t5nex0
cxl2: Ethernet address: 00:07:43:3f:e7:70
cxl2: 16 txq, 8 rxq (NIC); 8 txq (TOE), 2 rxq (TOE)
cxl3: <port 3> on t5nex0
cxl3: Ethernet address: 00:07:43:3f:e7:78
cxl3: 16 txq, 8 rxq (NIC); 8 txq (TOE), 2 rxq (TOE)
t5nex0: PCIe gen3 x4, 4 ports, 42 MSI-X interrupts, 140 eq, 41 iq
t5vf0: <Chelsio T540-CR VF> at device 0.8 on pci1
cxlv0: <port 0> on t5vf0
cxlv0: Ethernet address: 06:44:3f:e7:60:00
cxlv0: 2 txq, 1 rxq (NIC)
t5vf0: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf1: <Chelsio T540-CR VF> at device 0.12 on pci1
cxlv1: <port 0> on t5vf1
cxlv1: Ethernet address: 06:44:3f:e7:60:01
cxlv1: 2 txq, 1 rxq (NIC)
t5vf1: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf2: <Chelsio T540-CR VF> at device 0.16 on pci1
cxlv2: <port 0> on t5vf2
cxlv2: Ethernet address: 06:44:3f:e7:60:02
cxlv2: 2 txq, 1 rxq (NIC)
t5vf2: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf3: <Chelsio T540-CR VF> at device 0.20 on pci1
cxlv3: <port 0> on t5vf3
cxlv3: Ethernet address: 06:44:3f:e7:60:03
cxlv3: 2 txq, 1 rxq (NIC)
t5vf3: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf4: <Chelsio T540-CR VF> at device 0.24 on pci1
cxlv4: <port 0> on t5vf4
cxlv4: Ethernet address: 06:44:3f:e7:60:04
cxlv4: 2 txq, 1 rxq (NIC)
t5vf4: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf5: <Chelsio T540-CR VF> at device 0.28 on pci1
cxlv5: <port 0> on t5vf5
cxlv5: Ethernet address: 06:44:3f:e7:60:05
cxlv5: 2 txq, 1 rxq (NIC)
t5vf5: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf6: <Chelsio T540-CR VF> at device 0.32 on pci1
cxlv6: <port 0> on t5vf6
cxlv6: Ethernet address: 06:44:3f:e7:60:06
cxlv6: 2 txq, 1 rxq (NIC)
t5vf6: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf7: <Chelsio T540-CR VF> at device 0.36 on pci1
cxlv7: <port 0> on t5vf7
cxlv7: Ethernet address: 06:44:3f:e7:60:07
cxlv7: 2 txq, 1 rxq (NIC)
t5vf7: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf8: <Chelsio T540-CR VF> at device 0.40 on pci1
cxlv8: <port 0> on t5vf8
cxlv8: Ethernet address: 06:44:3f:e7:60:08
cxlv8: 2 txq, 1 rxq (NIC)
t5vf8: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf9: <Chelsio T540-CR VF> at device 0.44 on pci1
cxlv9: <port 0> on t5vf9
cxlv9: Ethernet address: 06:44:3f:e7:60:09
cxlv9: 2 txq, 1 rxq (NIC)
t5vf9: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf10: <Chelsio T540-CR VF> at device 0.48 on pci1
cxlv10: <port 0> on t5vf10
cxlv10: Ethernet address: 06:44:3f:e7:60:0a
cxlv10: 2 txq, 1 rxq (NIC)
t5vf10: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf11: <Chelsio T540-CR VF> at device 0.52 on pci1
cxlv11: <port 0> on t5vf11
cxlv11: Ethernet address: 06:44:3f:e7:60:0b
cxlv11: 2 txq, 1 rxq (NIC)
t5vf11: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf12: <Chelsio T540-CR VF> at device 0.56 on pci1
cxlv12: <port 0> on t5vf12
cxlv12: Ethernet address: 06:44:3f:e7:60:0c
cxlv12: 2 txq, 1 rxq (NIC)
t5vf12: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf13: <Chelsio T540-CR VF> at device 0.60 on pci1
cxlv13: <port 0> on t5vf13
cxlv13: Ethernet address: 06:44:3f:e7:60:0d
cxlv13: 2 txq, 1 rxq (NIC)
t5vf13: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf14: <Chelsio T540-CR VF> at device 0.64 on pci1
cxlv14: <port 0> on t5vf14
cxlv14: Ethernet address: 06:44:3f:e7:60:0e
cxlv14: 2 txq, 1 rxq (NIC)
t5vf14: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
t5vf15: <Chelsio T540-CR VF> at device 0.68 on pci1
cxlv15: <port 0> on t5vf15
cxlv15: Ethernet address: 06:44:3f:e7:60:0f
cxlv15: 2 txq, 1 rxq (NIC)
t5vf15: 1 ports, 2 MSI-X interrupts, 4 eq, 2 iq
loader configuration:
t5fw_cfg_load=YES
if_cxl_load=YES
if_cxgbev_load=YES
hw.cxgbe.tx_vm_wr=1
t4_tom_load=YES
interface configuration:
cxl0: flags=1008843<UP,BROADCAST,RUNNING,SIMPLEX,MULTICAST,LOWER_UP> metric 0
mtu 1500
options=6ec07bb<RXCSUM,TXCSUM,VLAN_MTU,VLAN_HWTAGGING,JUMBO_MTU,VLAN_HWCSUM,TSO4,TSO6,LRO,VLAN_HWTSO,LINKSTATE,RXCSUM_IPV6,TXCSUM_IPV6,HWSTATS,HWRXTSTMP,MEXTPG>
ether 00:07:43:3f:e7:60
inet 90.155.69.72/32 broadcast 90.155.69.72
inet6 fe80::207:43ff:fe3f:e760%cxl0/64 scopeid 0x1
inet6 2001:8b0:aab5:b003::2/64
media: Ethernet 10Gbase-Twinax <full-duplex,rxpause,txpause>
status: active
nd6 options=21<PERFORMNUD,AUTO_LINKLOCAL>
i tried to enable TOE on the PF and download a file to test performance, and
the system immediately panicked:
root@hemlock:~ # ifconfig cxl0 toe
root@hemlock:~ # curl -o/dev/null https://www.le-fay.org/files/bigfile.1000m
Fatal trap 12: page fault while in kernel mode
cpuid = 8; apic id = 08
fault virtual address = 0x0
fault code = supervisor read instruction, page not present
instruction pointer = 0x20:0x0
stack pointer = 0x28:0xfffffe019276ec08
frame pointer = 0x28:0xfffffe019276ec90
code segment = base 0x0, limit 0xfffff, type 0x1b
= DPL 0, pres 1, long 1, def32 0, gran 1
processor eflags = interrupt enabled, resume, IOPL = 0
current process = 82401 (curl)
rdi: fffff80033d83800 rsi: 0000000000000004 rdx: fffff80746833000
rcx: fffff80746833000 r8: 0000000000000040 r9: fffffe002d921500
rax: ffffffff82c240b0 rbx: fffff80746833000 rbp: fffffe019276ec90
r10: fffffe0170050c90 r11: 0000000000000000 r12: 0000000000000002
r13: fffffe019276eca8 r14: 0000000000000000 r15: 0000000000000000
trap number = 12
panic: page fault
cpuid = 8
time = 1786992180
KDB: stack backtrace:
#0 0xffffffff80be24c3 at kdb_backtrace+0x63
#1 0xffffffff80b94a76 at vpanic+0x136
#2 0xffffffff80b94933 at panic+0x43
#3 0xffffffff810b0312 at trap_pfault+0x3d2
#4 0xffffffff810861a8 at calltrap+0x8
#5 0xffffffff80c0751d at kern_poll+0xfd
#6 0xffffffff80c073b0 at sys_poll+0x50
#7 0xffffffff810b0c71 at amd64_syscall+0x131
#8 0xffffffff81086a9b at fast_syscall_common+0xf8
Uptime: 34m32s
Dumping 2662 out of 32645 MB:..1%..11%..21%..31%..41%..51%..61%..71%..81%..91%
backtrace from the active thread does not look very interesting:
(kgdb) bt
#0 __curthread () at /usr/src/sys/amd64/include/pcpu_aux.h:57
#1 doadump (textdump=<optimized out>) at /usr/src/sys/kern/kern_shutdown.c:399
#2 0xffffffff80b945ee in kern_reboot (howto=260) at
/usr/src/sys/kern/kern_shutdown.c:519
#3 0xffffffff80b94b07 in vpanic (fmt=0xffffffff8120a73a "%s",
ap=ap@entry=0xfffffe019276eac0) at /usr/src/sys/kern/kern_shutdown.c:974
#4 0xffffffff80b94933 in panic (fmt=<unavailable>) at
/usr/src/sys/kern/kern_shutdown.c:887
#5 0xffffffff810b0312 in trap_fatal (frame=<optimized out>, eva=<optimized
out>) at /usr/src/sys/amd64/amd64/trap.c:969
#6 0xffffffff810b0312 in trap_pfault (frame=0xfffffe019276eb40,
usermode=false, signo=<optimized out>, ucode=<optimized out>)
#7 <signal handler called>
#8 0x0000000000000000 in ?? ()
#9 0xffffffff80c07a4b in fo_poll (fp=0xfffff80033d83800,
fp@entry=0xfffff80746833000, events=4, events@entry=0,
active_cred=0xfffff80746833000, active_cred@entry=0x0, td=0xfffff80746833000)
at /usr/src/sys/sys/file.h:386
#10 pollscan (td=0xfffff80746833000, fds=0xfffffe019276eca8, nfd=3) at
/usr/src/sys/kern/sys_generic.c:1793
#11 kern_poll_kfds (td=td@entry=0xfffff80746833000,
kfds=kfds@entry=0xfffffe019276eca0, nfds=nfds@entry=3,
tsp=tsp@entry=0xfffffe019276edf0, uset=uset@entry=0x0) at
/usr/src/sys/kern/sys_generic.c:1582
#12 0xffffffff80c0751d in kern_poll (td=0xfffff80746833000, ufds=0x8209f47b0,
nfds=3, tsp=0xfffffe019276edf0, set=set@entry=0x0) at
/usr/src/sys/kern/sys_generic.c:1663
#13 0xffffffff80c073b0 in sys_poll (td=0xfffff80033d83800, uap=<optimized out>)
at /usr/src/sys/kern/sys_generic.c:1533
#14 0xffffffff810b0c71 in syscallenter (td=0xfffff80746833000) at
/usr/src/sys/amd64/amd64/../../kern/subr_syscall.c:193
#15 amd64_syscall (td=0xfffff80746833000, traced=0) at
/usr/src/sys/amd64/amd64/trap.c:1208
#16 <signal handler called>
#17 0x000000082898103a in ?? ()
Backtrace stopped: Cannot access memory at address 0x8209f46a8
i can provide the kernel + core privately if needed.
--
You are receiving this mail because:
You are the assignee for the bug.
lmpx.com only provides a reader for public news (NNTP) servers. It is not
affiliated with the servers or forums shown here and is not responsible for
the content of articles, which is written by their respective authors.