Ethernet down ?

"Shi, Josh (Unix Admin - PCS)" <[email protected]>
Newsgroups gmane.linux.highavailability.ultramonkey
Message-ID <284CD042B4ACD4118EAE0008C786DC51244323D7@cc02-exchange.systemax.com>
Running Ultramonkey 3 on RHEL AS 3.0 Update 4 ( 2.4.21-27.EL)
It has run fine for months. But it failed over a few days ago.
Still try to find it out why.

There are info in ha-log and messages

[root@pdmail1man log]# tail ha-log
heartbeat: 2005/09/29_23:24:25 info: Current arena value: 135168
heartbeat: 2005/09/29_23:24:25 info: MSG stats: 0/8646457 ms age 770
[pid1561/HBREAD]
heartbeat: 2005/09/29_23:24:25 info: ha_malloc stats: 0/172929140  252/0
[pid1561/HBREAD]
heartbeat: 2005/09/29_23:24:25 info: RealMalloc stats: 1468 total malloc
bytes. pid [1561/HBREAD]
heartbeat: 2005/09/29_23:24:25 info: Current arena value: 135168
heartbeat: 2005/09/29_23:24:25 info: These are nothing to worry about.
heartbeat: 2005/09/30_01:26:45 info: Link pdmail2man:eth0 up.
heartbeat: 2005/09/30_01:26:55 info: Link pdmail2man:eth0 dead.
heartbeat: 2005/09/30_02:44:09 info: Link pdmail2man:eth0 up.
heartbeat: 2005/09/30_02:44:19 info: Link pdmail2man:eth0 dead.

but ping is fine
[root@pdmail1man log]# ping pdmail2man
PING pdmail2man (172.16.20.190) 56(84) bytes of data.
64 bytes from pdmail2man (172.16.20.190): icmp_seq=0 ttl=64 time=0.082 ms
64 bytes from pdmail2man (172.16.20.190): icmp_seq=1 ttl=64 time=0.080 ms

on pdmail1man, /var/log/messages
Sep 29 23:24:25 pdmail1man heartbeat[1544]: info: Current arena value:
135168
Sep 29 23:24:25 pdmail1man heartbeat[1544]: info: These are nothing to worry
about.
Sep 30 01:26:45 pdmail1man heartbeat[1544]: info: Link pdmail2man:eth0 up.
Sep 30 01:26:45 pdmail1man ipfail[1573]: info: Link Status update: Link
pdmail2man/eth0 now has status up
Sep 30 01:26:55 pdmail1man heartbeat[1544]: info: Link pdmail2man:eth0 dead.
Sep 30 01:26:55 pdmail1man ipfail[1573]: info: Link Status update: Link
pdmail2man/eth0 now has status dead
Sep 30 01:26:55 pdmail1man ipfail[1573]: info: Asking other side for ping
node count.
Sep 30 01:26:55 pdmail1man ipfail[1573]: info: Checking remote count of ping
nodes.
Sep 30 01:26:55 pdmail1man ipfail[1573]: info: No giveup timer to abort.
Sep 30 02:44:09 pdmail1man heartbeat[1544]: info: Link pdmail2man:eth0 up.
Sep 30 02:44:09 pdmail1man ipfail[1573]: info: Link Status update: Link
pdmail2man/eth0 now has status up
Sep 30 02:44:19 pdmail1man heartbeat[1544]: info: Link pdmail2man:eth0 dead.
Sep 30 02:44:19 pdmail1man ipfail[1573]: info: Link Status update: Link
pdmail2man/eth0 now has status dead
Sep 30 02:44:19 pdmail1man ipfail[1573]: info: Asking other side for ping
node count.
Sep 30 02:44:19 pdmail1man ipfail[1573]: info: Checking remote count of ping
nodes.
Sep 30 02:44:19 pdmail1man ipfail[1573]: info: No giveup timer to abort.
Sep 30 04:20:05 pdmail1man ipfail[1573]: info: Ping node count is balanced.

on pdmail2man, /var/log/ha-log
heartbeat: 2005/09/29_12:25:02 info: Link 172.16.20.1:172.16.20.1 up.
heartbeat: 2005/09/29_12:25:02 info: Status update for node 172.16.20.1:
status ping
heartbeat: 2005/09/29_12:25:02 info: Local status now set to: 'active'
heartbeat: 2005/09/29_12:25:02 info: Starting child client
"/usr/lib/heartbeat/ipfail " (500,500)
heartbeat: 2005/09/29_12:25:02 info: Starting "/usr/lib/heartbeat/ipfail "
as uid 500  gid 500 (pid 1819)
heartbeat: 2005/09/29_12:25:02 info: remote resource transition completed.
heartbeat: 2005/09/29_12:25:02 info: remote resource transition completed.
heartbeat: 2005/09/29_12:25:02 info: Local Resource acquisition completed.
(none)
heartbeat: 2005/09/29_12:25:02 info: Initial resource acquisition complete
(T_RESOURCES(them))
heartbeat: 2005/09/30_04:20:06 info: Link pdmail1man:eth0 up.
heartbeat: 2005/09/30_04:20:16 info: Link pdmail1man:eth0 dead.
on pdmail2man, /var/log/messages
Sep 29 12:28:08 pdmail2man kernel: mtrr: type mismatch for fc000000,800000
old: uncachable new: write-combining
Sep 29 12:28:08 pdmail2man kernel: mtrr: type mismatch for fc000000,800000
old: uncachable new: write-combining
Sep 29 13:05:01 pdmail2man ipfail[1819]: info: Ping node count is balanced.
Sep 30 01:27:06 pdmail2man ipfail[1819]: info: Ping node count is balanced.
Sep 30 02:44:30 pdmail2man ipfail[1819]: info: Ping node count is balanced.
Sep 30 04:20:06 pdmail2man heartbeat[1790]: info: Link pdmail1man:eth0 up.
Sep 30 04:20:06 pdmail2man ipfail[1819]: info: Link Status update: Link
pdmail1man/eth0 now has status up
Sep 30 04:20:16 pdmail2man heartbeat[1790]: info: Link pdmail1man:eth0 dead.
Sep 30 04:20:16 pdmail2man ipfail[1819]: info: Link Status update: Link
pdmail1man/eth0 now has status dead
Sep 30 04:20:16 pdmail2man ipfail[1819]: info: Asking other side for ping
node count.
Sep 30 04:20:16 pdmail2man ipfail[1819]: info: Checking remote count of ping
nodes.
Sep 30 04:20:16 pdmail2man ipfail[1819]: info: No giveup timer to abort.


Here is /etc/ha.d/ha.cf

debugfile /var/log/ha-debug
logfile /var/log/ha-log
logfacility     local0
keepalive 1
deadtime 10
warntime 10
initdead 30
auto_failback off
udpport 694
bcast   eth0 eth2               # Linux
mcast eth2 225.0.0.1 694 1 0
node pdmail1man
node pdmail2man
ping 172.16.20.1
respawn hacluster /usr/lib/heartbeat/ipfail


-- 
Ultra Monkey - http://www.ultramonkey.org/
To UNSUBSCRIBE, email to [email protected], with a body:
unsubscribe ultramonkey-users [email protected]
where "[email protected]" is YOUR email address.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.