SS20 MP memory issue

Gordon Zaft <[email protected]>
Newsgroups gmane.os.netbsd.ports.sparc64
Message-ID <CAGuhNT0KfgAxxXGRyR=Y5csMS0=b=kioX86BYsUA7ATB3wbn2A@mail.gmail.com>
I have an SS20 with twin SM71s and 512MB RAM.

I built an MP kernel (7.1_RC1).  Following that and building a bunch of
packages I see this error in dmesg:

NetBSD 7.1_RC1 (SYLVESTER) #0: Sun Feb 12 18:51:42 MST 2017
        root@wilbur:/usr/obj/sys/arch/sparc/compile/SYLVESTER
total memory = 511 MB
avail memory = 496 MB
kern.module.path=/stand/sparc/7.1/modules
timecounter: Timecounters tick every 10.000 msec
bootpath: /iommu@f,e0000000/sbus@f,e0001000/espdma@f,400000/esp@f
,800000/sd@1,0
mainbus0 (root): SUNW,SPARCstation-20: hostid 72c6d7b3
cpu0 at mainbus0: mid 8: TMS390Z50 v0 or TMS390Z55 @ 75 MHz, on-chip FPU
cpu0: physical 20K instruction (64 b/l), 16K data (32 b/l), 1024K external
(32 b/l): cache enabled
cpu1 at mainbus0: mid 10: TMS390Z50 v0 or TMS390Z55 @ 75 MHz, on-chip FPU
cpu1: physical 20K instruction (64 b/l), 16K data (32 b/l), 1024K external
(32 b/l): cache enabled
sx0 at mainbus0 ioaddr 0x80000000
sx0: architecture rev. 27 chip rev. 0
obio0 at mainbus0
clock0 at obio0 slot 0 offset 0x200000: mk48t08
timer0 at obio0 slot 0 offset 0x300000: delay constant 35, frequency =
2000000 Hz
timer: limit 0 shift 9 mask 3fffff
timecounter: Timecounter "timer-counter" frequency 2000000 Hz quality 100
zs0 at obio0 slot 0 offset 0x100000 level 12 softpri 6
zstty0 at zs0 channel 0 (console i/o)
zstty1 at zs0 channel 1
zs1 at obio0 slot 0 offset 0x0 level 12 softpri 6
zstty4 at zs1 channel 0
kbd0 at zstty4
zstty5 at zs1 channel 1
ms0 at zstty5
wsmouse0 at ms0 mux 0
fdc0 at obio0 slot 0 offset 0x700000 level 11: no drives attached
auxreg0 at obio0 slot 0 offset 0x800000
power0 at obio0 slot 0 offset 0xa01000 level 2
iommu0 at mainbus0 ioaddr 0xe0000000: version 0x3/0x1, page-size 4096,
range 64MB
sbus0 at iommu0: clock = 25 MHz
dma0 at sbus0 slot 15 offset 0x400000: DMA rev 2
esp0 at dma0 slot 15 offset 0x800000 level 4: ESP200, 40MHz, SCSI ID 7
scsibus0 at esp0: 8 targets, 8 luns per target
ledma0 at sbus0 slot 15 offset 0x400010: DMA rev 2
le0 at ledma0 slot 15 offset 0xc00000 level 6: address 08:00:20:c6:d7:b3
le0: 8 receive buffers, 2 transmit buffers
bpp0 at sbus0 slot 15 offset 0x4800000 level 2 (ipl 3): DMA rev 2
dbri0 at sbus0 slot 14 offset 0x10000 level 9: rev e
cgsix0 at sbus0 slot 2 offset 0x0 level 9: SUNW,501-2325, 1152 x 900, rev 11
cgsix0: attached to /dev/fb0
cgsix0: framebuffer size: 1 MB
wsdisplay1 at cgsix0 kbdmux 1
wsmux1: connecting to wsdisplay1
eccmemctl0 at mainbus0 ioaddr 0x0: version 0x0/0x2
timecounter: Timecounter "clockinterrupt" frequency 100 Hz quality 0
cpu0: booting secondary processors: cpu1
scsibus0: waiting 2 seconds for devices to settle...
wskbd0 at kbd0 mux 1
sd0 at scsibus0 target 1 lun 0: <SEAGATE, ST32430W SUN2.1G, 0666> disk fixed
sd0: 2049 MB, 3992 cyl, 9 head, 116 sec, 512 bytes/sect x 4197405 sectors
sd0: sync (100.00ns offset 15), 8-bit (10.000MB/s) transfers, tagged
queueing
sd1 at scsibus0 target 3 lun 0: <COMPAQ, BB018135B5, B013> disk fixed
sd1: 17365 MB, 7001 cyl, 20 head, 254 sec, 512 bytes/sect x 35565080 sectors
sd1: sync (100.00ns offset 15), 8-bit (10.000MB/s) transfers, tagged
queueing
kbd0: reset failed
wskbd0: connecting to wsdisplay1
cd0 at scsibus0 target 6 lun 0: <TOSHIBA, XM-4101TASUNSLCD, 1084> cdrom
removable
cd0: async, 8-bit transfers
Kernelized RAIDframe activated
dbri0: speakerbox detected
dbri0: cs4215 rev E found at offset 8
audio0 at dbri0: full duplex, playback, capture, mmap
root on sd0a dumps on sd0b
root file system type: ffs
cpu0: NMI: system interrupts: 0x10000000<VME=0x0,SBUS=0x0,M>
memory error:
        EFSR: 0x8a31<CE,DW=0x3,SYNDROME=0x8a>
        MBus transaction:
0x87ffcd50<VAH=0x0,TYPE=0x5,SIZE=0x5,C,VA=0xff,MID=0x8>
        address: 0x01eb3ba80
        module location: J0304
cpu0: NMI: system interrupts: 0x10000000<VME=0x0,SBUS=0x0,M>
memory error:
        EFSR: 0x8a31<CE,DW=0x3,SYNDROME=0x8a>
        MBus transaction:
0xa7ffcd50<VAH=0x0,TYPE=0x5,SIZE=0x5,C,VA=0xff,MID=0xa>
        address: 0x01eb3ba80
        module location: J0304
cpu0: NMI: system interrupts: 0x10000000<VME=0x0,SBUS=0x0,M>
memory error:
        EFSR: 0x8a21<CE,DW=0x2,SYNDROME=0x8a>
        MBus transaction:
0xa7ffcd50<VAH=0x0,TYPE=0x5,SIZE=0x5,C,VA=0xff,MID=0xa>
        address: 0x01eb3ba8c
        module location: J0304


I didn't see these errors while running a kernel build etc. with a
single-cpu kernel (stock 7.0.1).

Should I replace the J0304 memory module?  Or is this a software issue?

G


-- 
Gordon Zaft
Province 35 Governor
Phi Mu Alpha Sinfonia Fraternity
[email protected]
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.