Re: drbd breaks druing syncing openssi 1.9.1
John Hughes <[email protected]>
| Newsgroups | gmane.linux.cluster.ssic.user |
|---|---|
| Message-ID | <[email protected]> |
stefan wrote:
> I set up a debian- sarge- openssi- drbd- ha- cluster with openssi 1.9.1.
>
> All is running well till I tried some failovers. The first failover was ok.
> node1 was the primary, power off node2.
> start node 2 again- > syncing
> all node are up
> power off node1
> then the problem beginns.
>
> I see that node 1 losing connection 3 times and after third lost connection it
> gives a kdb- error. Both nodes are dead!
> doing all again, starting node 2 as primary it gives a kdb and node1 stand
> still with kdb- error.
>
> Also booting over PXE gives same errors:
>
Well, I've managed to arrive at a similar problem on 1.9.3, got the
system up with both nodes working,
crashed node 2 a couple of times, worked, then crashed node 1, node 2
took over with no problems, but when I restarted node 1 it got part of
the way then timed out:
This is a CI/OpenSSI kernel.
This Cluster Node: 1
Potential Initnode(s): 1:192.168.64.1,2:192.168.64.2
ipcnameserver ready completed
Name server registered with clms
ipcname_read completed
drbd: module cleanup done.
modprobe -k drbd minor_count=1
drbd: initialised. Version: 0.7.24 (api:79/proto:74)
drbd: SVN Revision: 2875 build by john@node1, 2007-09-06 17:57:32
drbd: registered as block device major 147
Starting DRBD resource:
drbd0: resync bitmap: bits=487975 words=15250
drbd0: size = 1906 MB (1951897 KB)
drbd0: 0 KB marked out-of-sync by on disk bit-map.
drbd0: Found 6 transactions (80 active extents) in activity log.
drbd0: Marked additional 304 MB as out-of-sync based on AL.
drbd0: drbdsetup [66052]: cstate Unconfigured --> StandAlone
drbd0: drbdsetup [66055]: cstate StandAlone --> Unconnected
drbd0: drbd0_receiver [66056]: cstate Unconnected --> WFConnection
drbd0: Registering drbd0 with CLMS subsystem
# WARNING: Do not type 'yes' while waiting for DRBD connection
# unless you know what you are doing! You have been warned!
# The only exception is when setting up DRBD first time.
#
drbd0: drbd0_receiver [66056]: cstate WFConnection --> WFReportParams
drbd0: Handshake successful: DRBD Network Protocol version 74
drbd0: Connection established.
drbd0: I am(S): 1:00000002:00000001:00000005:00000001:10
drbd0: Peer(P): 1:00000002:00000001:00000005:00000002:10
drbd0: drbd0_receiver [66056]: cstate WFReportParams --> WFBitMapT
drbd0: Secondary/Unknown --> Secondary/Primary
drbd0: drbd0_receiver [66056]: cstate WFBitMapT --> SyncTarget
drbd0: Resync started as SyncTarget (need to sync 311296 KB [77824 bits set]).
Attempting to pivot_root
Running post-root cluster initialization
warning: can't open /proc/mounts: No such file or directory
Mounting a tmpfs over /dev...done.
Creating initial device nodes...done.
Unmounting /sys
Unmounting /proc
Starting init
/etc/init.d/rc.nodeup 1 running
drbd0: PingAck did not arrive in time.
drbd0: drbd0_asender [66063]: cstate SyncTarget --> NetworkFailure
drbd0: asender terminated
drbd0: drbd0_receiver [66056]: cstate NetworkFailure --> BrokenPipe
drbd0: short read expecting header on sock: r=-512
drbd0: worker terminated
drbd0: drbd0_receiver [66056]: cstate BrokenPipe --> Unconnected
drbd0: Connection lost.
drbd0: drbd0_receiver [66056]: cstate Unconnected --> WFConnection
Kernel panic - not syncing: Lost network connection to all potential root nodes!Instruction(i) breakpoint #0 at 0xc0127340 (adjusted)
0xc0127340 panic_hook: int3
Entering kdb (current=0xcebce950, pid 66029) on processor 0 due to Breakpoint @ 0xc0127340
[0]kdb>
-------------------------------------------------------------------------
This SF.net email is sponsored by: Microsoft
Defy all challenges. Microsoft(R) Visual Studio 2005.
http://clk.atdmt.com/MRT/go/vse0120000070mrt/direct/01/