Ethernet down ?
"Shi, Josh (Unix Admin - PCS)" <[email protected]>
| Newsgroups | gmane.linux.highavailability.ultramonkey |
|---|---|
| Message-ID | <284CD042B4ACD4118EAE0008C786DC51244323D7@cc02-exchange.systemax.com> |
Running Ultramonkey 3 on RHEL AS 3.0 Update 4 ( 2.4.21-27.EL) It has run fine for months. But it failed over a few days ago. Still try to find it out why. There are info in ha-log and messages [root@pdmail1man log]# tail ha-log heartbeat: 2005/09/29_23:24:25 info: Current arena value: 135168 heartbeat: 2005/09/29_23:24:25 info: MSG stats: 0/8646457 ms age 770 [pid1561/HBREAD] heartbeat: 2005/09/29_23:24:25 info: ha_malloc stats: 0/172929140 252/0 [pid1561/HBREAD] heartbeat: 2005/09/29_23:24:25 info: RealMalloc stats: 1468 total malloc bytes. pid [1561/HBREAD] heartbeat: 2005/09/29_23:24:25 info: Current arena value: 135168 heartbeat: 2005/09/29_23:24:25 info: These are nothing to worry about. heartbeat: 2005/09/30_01:26:45 info: Link pdmail2man:eth0 up. heartbeat: 2005/09/30_01:26:55 info: Link pdmail2man:eth0 dead. heartbeat: 2005/09/30_02:44:09 info: Link pdmail2man:eth0 up. heartbeat: 2005/09/30_02:44:19 info: Link pdmail2man:eth0 dead. but ping is fine [root@pdmail1man log]# ping pdmail2man PING pdmail2man (172.16.20.190) 56(84) bytes of data. 64 bytes from pdmail2man (172.16.20.190): icmp_seq=0 ttl=64 time=0.082 ms 64 bytes from pdmail2man (172.16.20.190): icmp_seq=1 ttl=64 time=0.080 ms on pdmail1man, /var/log/messages Sep 29 23:24:25 pdmail1man heartbeat[1544]: info: Current arena value: 135168 Sep 29 23:24:25 pdmail1man heartbeat[1544]: info: These are nothing to worry about. Sep 30 01:26:45 pdmail1man heartbeat[1544]: info: Link pdmail2man:eth0 up. Sep 30 01:26:45 pdmail1man ipfail[1573]: info: Link Status update: Link pdmail2man/eth0 now has status up Sep 30 01:26:55 pdmail1man heartbeat[1544]: info: Link pdmail2man:eth0 dead. Sep 30 01:26:55 pdmail1man ipfail[1573]: info: Link Status update: Link pdmail2man/eth0 now has status dead Sep 30 01:26:55 pdmail1man ipfail[1573]: info: Asking other side for ping node count. Sep 30 01:26:55 pdmail1man ipfail[1573]: info: Checking remote count of ping nodes. Sep 30 01:26:55 pdmail1man ipfail[1573]: info: No giveup timer to abort. Sep 30 02:44:09 pdmail1man heartbeat[1544]: info: Link pdmail2man:eth0 up. Sep 30 02:44:09 pdmail1man ipfail[1573]: info: Link Status update: Link pdmail2man/eth0 now has status up Sep 30 02:44:19 pdmail1man heartbeat[1544]: info: Link pdmail2man:eth0 dead. Sep 30 02:44:19 pdmail1man ipfail[1573]: info: Link Status update: Link pdmail2man/eth0 now has status dead Sep 30 02:44:19 pdmail1man ipfail[1573]: info: Asking other side for ping node count. Sep 30 02:44:19 pdmail1man ipfail[1573]: info: Checking remote count of ping nodes. Sep 30 02:44:19 pdmail1man ipfail[1573]: info: No giveup timer to abort. Sep 30 04:20:05 pdmail1man ipfail[1573]: info: Ping node count is balanced. on pdmail2man, /var/log/ha-log heartbeat: 2005/09/29_12:25:02 info: Link 172.16.20.1:172.16.20.1 up. heartbeat: 2005/09/29_12:25:02 info: Status update for node 172.16.20.1: status ping heartbeat: 2005/09/29_12:25:02 info: Local status now set to: 'active' heartbeat: 2005/09/29_12:25:02 info: Starting child client "/usr/lib/heartbeat/ipfail " (500,500) heartbeat: 2005/09/29_12:25:02 info: Starting "/usr/lib/heartbeat/ipfail " as uid 500 gid 500 (pid 1819) heartbeat: 2005/09/29_12:25:02 info: remote resource transition completed. heartbeat: 2005/09/29_12:25:02 info: remote resource transition completed. heartbeat: 2005/09/29_12:25:02 info: Local Resource acquisition completed. (none) heartbeat: 2005/09/29_12:25:02 info: Initial resource acquisition complete (T_RESOURCES(them)) heartbeat: 2005/09/30_04:20:06 info: Link pdmail1man:eth0 up. heartbeat: 2005/09/30_04:20:16 info: Link pdmail1man:eth0 dead. on pdmail2man, /var/log/messages Sep 29 12:28:08 pdmail2man kernel: mtrr: type mismatch for fc000000,800000 old: uncachable new: write-combining Sep 29 12:28:08 pdmail2man kernel: mtrr: type mismatch for fc000000,800000 old: uncachable new: write-combining Sep 29 13:05:01 pdmail2man ipfail[1819]: info: Ping node count is balanced. Sep 30 01:27:06 pdmail2man ipfail[1819]: info: Ping node count is balanced. Sep 30 02:44:30 pdmail2man ipfail[1819]: info: Ping node count is balanced. Sep 30 04:20:06 pdmail2man heartbeat[1790]: info: Link pdmail1man:eth0 up. Sep 30 04:20:06 pdmail2man ipfail[1819]: info: Link Status update: Link pdmail1man/eth0 now has status up Sep 30 04:20:16 pdmail2man heartbeat[1790]: info: Link pdmail1man:eth0 dead. Sep 30 04:20:16 pdmail2man ipfail[1819]: info: Link Status update: Link pdmail1man/eth0 now has status dead Sep 30 04:20:16 pdmail2man ipfail[1819]: info: Asking other side for ping node count. Sep 30 04:20:16 pdmail2man ipfail[1819]: info: Checking remote count of ping nodes. Sep 30 04:20:16 pdmail2man ipfail[1819]: info: No giveup timer to abort. Here is /etc/ha.d/ha.cf debugfile /var/log/ha-debug logfile /var/log/ha-log logfacility local0 keepalive 1 deadtime 10 warntime 10 initdead 30 auto_failback off udpport 694 bcast eth0 eth2 # Linux mcast eth2 225.0.0.1 694 1 0 node pdmail1man node pdmail2man ping 172.16.20.1 respawn hacluster /usr/lib/heartbeat/ipfail -- Ultra Monkey - http://www.ultramonkey.org/ To UNSUBSCRIBE, email to [email protected], with a body: unsubscribe ultramonkey-users [email protected] where "[email protected]" is YOUR email address.