Re: database problems
"Daniel Berg" <[email protected]> Wed, 03 Jul 2002 02:34:07 -0500
| Newsgroups | gmane.linux.failsafe |
|---|---|
| Message-ID | <[email protected]> |
Sorry for the intrusion. I think I'm getting pretty close to gettig this to work but there is however still some problems with the database I think. When I start ha-services both nodes start their ha-processes but only one of them, the local node shows the status "up". In the cmsd-log there is a message that confuses me a bit. It says that the checksum of CDB differs between the nodes and therefore the second node gets marked down. When I do cdbutil gettree \# on both nodes it looks like they contain almost the same information, only the structure of the dump differs between the nodes. I've really tried hard fixing this myself by reading the manual and "trial and error"-method, but no luck so I turn to you. I need to get this working by the end of this week, so I hope you can help me. here are some of the logs that looks "funny": cmsd-log: Thu Jun 27 08:35:45.273 <I0 ha_cmsd cms 1308:0 cmsd_recv_timeout.c:114> Forcing new membership computation Thu Jun 27 08:35:45.273 <I0 ha_cmsd cms 1308:0 cmsd_newconf.c:183> Using new config info (subtree checksums = 0x583a0d87e2caa763:0xffd22bcfeefe53ad:0x78e637bdc4466e62 combined checksum = 0x583a0d87a9b898f9) Thu Jun 27 08:35:45.274 <I0 ha_cmsd cms 1308:0 cmsd_config.c:646> Begin configuration. Thu Jun 27 08:35:45.274 <I0 ha_cmsd cms 1308:0 cmsd_config.c:652> Node tore with nodeid 1 is enabled in cluster. Thu Jun 27 08:35:45.274 <I0 ha_cmsd cms 1308:0 cmsd_config.c:652> Node tyson with nodeid 2 is enabled in cluster. Thu Jun 27 08:35:45.274 <I0 ha_cmsd cms 1308:0 cmsd_config.c:656> The tie breaker node is tore. Thu Jun 27 08:35:45.274 <I0 ha_cmsd cms 1308:0 cmsd_config.c:658> Node timeout is 15000 msecs. Thu Jun 27 08:35:45.274 <I0 ha_cmsd cms 1308:0 cmsd_config.c:660> Heartbeat period is 1000 msecs. Thu Jun 27 08:35:45.274 <I0 ha_cmsd cms 1308:0 cmsd_config.c:663> Cluster is in normal mode. Thu Jun 27 08:35:45.274 <I0 ha_cmsd cms 1308:0 cmsd_config.c:665> End configuration. Thu Jun 27 08:35:45.275 <I0 ha_cmsd cms 1308:0 cmsd_state.c:628> cmsd state change from monitor to leader Thu Jun 27 08:35:45.275 <N ha_cmsd cms 1308:0 cmsd_memb.c:795> Confirmed Membership: sqn 3 G_sqn = 3, ack false node tore [1] : UP incarnation 2 age 3:0 node tyson [2] : DOWN incarnation 0 age 0:0 Thu Jun 27 08:35:45.918 <W ha_cmsd cms 1308:0 cmsd_bcast.c:142> LAST MESSAGE IN THE cms SUBSYSTEM REPEATED 2 TIMES Thu Jun 27 08:35:45.918 <W ha_cmsd cms 1308:0 cmsd_bcast.c:142> Message (from node tyson:2) with a different CDB checksum local checksums = 0x583a0d87e2caa763:0xffd22bcfeefe53ad:0x78e637bdc4466e62 remote checksums = 0x583a0d87e2caa763:0xffd22bcfeefe53ad:0x5425d582a85a5b52, Rejecting message ... Thu Jun 27 08:35:45.919 <W ha_cmsd cms 1308:0 cmsd_bcast.c:142> Message (from node tyson:2) with a different CDB checksum local checksums = 0x583a0d87e2caa763:0xffd22bcfeefe53ad:0x78e637bdc4466e62 remote checksums = 0x583a0d87e2caa763:0xffd22bcfeefe53ad:0x5425d582a85a5b52, Rejecting message ... Thu Jun 27 08:35:46.283 <N ha_cmsd cms 1308:0 cmsd_memb.c:795> Confirmed Membership: sqn 3 G_sqn = 3, ack false node tore [1] : UP incarnation 2 age 3:0 node tyson [2] : DOWN incarnation 0 age 0:0 Thu Jun 27 08:35:46.928 <W ha_cmsd cms 1308:0 cmsd_bcast.c:142> LAST MESSAGE IN THE cms SUBSYSTEM REPEATED 3 TIMES Thu Jun 27 08:35:46.928 <W ha_cmsd cms 1308:0 cmsd_bcast.c:142> Message (from node tyson:2) with a different CDB checksum local checksums = 0x583a0d87e2caa763:0xffd22bcfeefe53ad:0x78e637bdc4466e62 remote checksums = 0x583a0d87e2caa763:0xffd22bcfeefe53ad:0x5425d582a85a5b52, Rejecting message ... cad-log: Tue Jul 2 08:35:35.230 <cad 1894:1024> ccacdb_poll called Tue Jul 2 08:35:35.230 <cad 1894:1024> ccicdb_poll called Tue Jul 2 08:35:35.230 <cad 1894:1024> cfscdb_poll called Tue Jul 2 08:35:35.230 <cad 1894:1024> ccms_poll called Tue Jul 2 08:35:35.230 <cad 1894:1024> cfs_poll called Tue Jul 2 08:35:35.230 <cad 1894:1024> ccamail_poll called Tue Jul 2 08:35:35.230 <cam_casmail 1894:1024> ccamail_cam_register: called. Tue Jul 2 08:35:35.230 <cam_casmail 1894:1024> ccamail_send_mail() called. Tue Jul 2 08:35:35.230 <cam_casmail 1894:1024> Cluster Perfekt, no of msgs 0. Tue Jul 2 08:35:35.231 <cad 1906:7176> cfs_fs_connect: fs_cam_register failed with error FailSafe is not ready to accept admin requests. Tue Jul 2 08:35:35.231 <cad 1906:7176> cfs_fs_work: failsafe state 1 Tue Jul 2 08:35:35.231 <cad 1906:7176> cfs_main: waiting to accept requests Tue Jul 2 08:35:35.232 <cam_cms 1905:6151> ccms_cms_connect: cms_register failed with error CI_IPCERR_NOSERVER Tue Jul 2 08:35:35.232 <cam_cms 1905:6151> ccms_main: waiting to accept requests Tue Jul 2 08:35:35.232 <cam_cicdb 1903:4101> ccicdb_main: waiting to accept requests Tue Jul 2 08:35:35.232 <cam_cascdb 1902:3076> ccacdb_main: waiting to accept requests Tue Jul 2 08:35:35.233 <cam_srm 1901:2051> csrm_connect_srmd: srm_register failed with error CI_IPCERR_NOSERVER Tue Jul 2 08:35:35.233 <cam_srm 1901:2051> csrm_srm_work: srm state 1 Tue Jul 2 08:35:35.233 <cam_srm 1901:2051> csrm_main: waiting to accept requests cdbd-log (this one looks ok I think): Tue Jul 2 07:47:06.956 tore cdbd - fs2d_initiate_new_quorum initiating new quorum at request of machine 2 Tue Jul 2 07:47:06.961 tore cdbd - Copying global CDB to node tyson (2) Tue Jul 2 07:47:08.079 tore cdbd - New quorum: cluster id: 0x00000000.0x3d1c7aef30700258, master: 1, sequence: 17, member count: 2, members: 2, 1 Tue Jul 2 07:47:08.079 tore cdbd - New quorum: cluster id: 0x00000000.0x3d1c7aef30700258, master: 1, sequence: 17, member count: 2, members: 2, 1 Tue Jul 2 07:47:08.079 tore cdbd - Ready and valid new quorum: cluster id: 0x00000000.0x3d1c7aef30700258, master: 1, sequence: 17, member count: 2, members: 2, 1 Tue Jul 2 07:47:09.314 tore cdbd - Checking quorum with 2 members for any unknown members. Tue Jul 2 07:47:09.314 tore cdbd - All quorum member machines are known to us. Tue Jul 2 07:52:22.989 tore cdbd - CDB on node tyson (2) marked current in fs2d_copy_cdb_to_machine Tue Jul 2 07:52:24.019 tore cdbd - Successfully copied global CDB to node tyson (297) *But sometimes I get this and I don't know what it means, though it seems like it gets going after a while* Tue Jul 2 07:46:58.857 tore cdbd - RPC machine register: rejecting quorum from machine tyson due to that machine not responding to our poll attempts Tue Jul 2 07:47:00.264 tore cdbd - RPC machine register: rejecting quorum from machine tyson due to that machine not responding to our poll attempts Tue Jul 2 07:47:02.190 tore cdbd - Machine tyson.perfekt has started polling. Tue Jul 2 07:47:06.482 tore cdbd - Requesting CDB sync to machine tyson (2) Tue Jul 2 07:47:06.565 tore cdbd - Pool machine tyson (2) responding Hoping for some clearence in this matter. Thanks in advance/ Daniel ------------------------------------------- Daniel Berg Väktargatan 44B nb 754 22 Uppsala tel: 018-25 75 30 mob: 070-634 5556 -- __________________________________________________________ Sign-up for your own FREE Personalized E-mail at Mail.com http://www.mail.com/?sr=signup Save up to $160 by signing up for NetZero Platinum Internet service. http://www.netzero.net/?refcd=N2P0602NEP8