[jgroups-dev] Concurrent startup of many servers(Coordinator is changed)

Kuwon Kang <[email protected]> Wed, 17 Nov 2010 15:57:36 +0900
Newsgroups gmane.comp.java.javagroups.devel
Message-ID <[email protected]>
Thanks for your help.

We have 14nodes in a GMS on 7 servers(this mean each server has 2nodes for
failover).

Normally If we need to upgrade our engine in production mode, we stop the
secondary engines for each server concurrently(while this primary engines
provides service) and release them, start all secondary engines in
concurrent.

After this steps, we upgrade primary engines too for each server in
concurrent.

While we are doing these steps, there are some problem.

Each server has many error and warn messages.
Aftera ll we lost a coordinator(One of secondary agent must be a coord, but
at the end of whole upgrades, *the coord is changed from one of secondary
engines to one of primary engines*.

I snipped messages from few servers.

I will explain the problems according to the steps in flow.

*1. *Stop 7 secondary engines in each server(After this there is a coord in
primary engines) -> *2. *Release secondary engines -> *3.* Startup all
secondary engines concurrently in each server->*4.* Stop 7 primary engines
in each server(after this there is a coord in secondary
engines->*5.*Release primary engines ->
*6.* Startup all primary engines concurrently in each server->
*Result: *According
to the JGroups view installation, One of secondary engine must be a coord
but we lost the coord in secondary engines(One of primary engines became a
coord).


*-- Server FLPCOBA1(This is one of primary engines).*
WARN  [11/17 14:22:15] FLPCOBA1-43250: dropped message from FLPCOA01-51680
(not in xmit_table), keys are [FLPCOA06-31629, FLPCOBA1-43250],
view=[FLPCOBA1-43250|1] [FLPCOBA1-43250, FLPCOA06-31629]
WARN  [11/17 14:22:15] FLPCOBA1-43250: dropped message from FLPCOA03-65117
(not in xmit_table), keys are [FLPCOA06-31629, FLPCOBA1-43250],
view=[FLPCOBA1-43250|1] [FLPCOBA1-43250, FLPCOA06-31629]
WARN  [11/17 14:22:15] FLPCOBA1-43250: dropped message from FLPCOA02-64862
(not in xmit_table), keys are [FLPCOA06-31629, FLPCOBA1-43250],
view=[FLPCOBA1-43250|1] [FLPCOBA1-43250, FLPCOA06-31629]
WARN  [11/17 14:22:15] FLPCOBA1-43250: dropped message from FLPCOA06-46079
(not in xmit_table), keys are [FLPCOA06-31629, FLPCOBA1-43250],
view=[FLPCOBA1-43250|1] [FLPCOBA1-43250, FLPCOA06-31629]
WARN  [11/17 14:22:15] FLPCOBA1-43250: dropped message from FLPCOA04-34052
(not in xmit_table), keys are [FLPCOA06-31629, FLPCOBA1-43250],
view=[FLPCOBA1-43250|1] [FLPCOBA1-43250, FLPCOA06-31629]
WARN  [11/17 14:22:15] FLPCOBA1-43250: dropped message from FLPCOBA1-61893
(not in xmit_table), keys are [FLPCOA06-31629, FLPCOBA1-43250],
view=[FLPCOBA1-43250|1] [FLPCOBA1-43250, FLPCOA06-31629]
WARN  [11/17 14:22:15] FLPCOBA1-43250: dropped message from FLPCOA05-20663
(not in xmit_table), keys are [FLPCOA06-31629, FLPCOBA1-43250],
view=[FLPCOBA1-43250|1] [FLPCOBA1-43250, FLPCOA06-31629]
ERROR [11/17 14:22:18] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:18] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:19] failed fetching digests from subpartition members;
dropping merge response
WARN  [11/17 14:22:20] FLPCOBA1-43250: merge leader did not get data from
all partition coordinators [FLPCOBA1-61893, FLPCOA01-51680, FLPCOA03-59531,
FLPCOA03-65117, FLPCOBA1-43250, FLPCOA01-37784, FLPCOA04-34052,
FLPCOA05-43926, FLPCOA04-37935, FLPCOBA1-45204, FLPCOA02-59383,
FLPCOA02-64862, FLPCOA05-20663, FLPCOA06-46079], merge is cancelled
ERROR [11/17 14:22:21] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:22] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:23] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:24] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:25] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:27] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:28] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:29] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:30] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:32] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:33] failed fetching digests from subpartition members;
dropping merge response
INFO  [11/17 14:22:34] VSAM File Service has been requested with following
options:page No = 1, pageSize = 200000, startRowNumber = 1, startPosition =
0, totalPageCount = 1, totalRowNumber = 0,  totalCount = 0, parameter =
/batch_sam/SHARE/log/sli/co/ra/mntrplc/PRAMNT3004_CFG_2010.11.17.04.00.04.315_RAContNoDupRemv.log,
additional parameters = {}
ERROR [11/17 14:22:34] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:35] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:36] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:38] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:39] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:40] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:41] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:42] failed fetching digests from subpartition members;
dropping merge response
INFO  [11/17 14:22:43] VSAM File Service has been requested with following
options:page No = 1, pageSize = 200000, startRowNumber = 1, startPosition =
0, totalPageCount = 1, totalRowNumber = 0,  totalCount = 0, parameter =
/batch_sam/SHARE/log/sli/co/ra/mntrplc/PRAMNT3004_CFG_2010.11.17.04.00.04.315_RAContNoDupRemv.log,
additional parameters = {}
ERROR [11/17 14:22:44] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:45] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:46] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:47] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:48] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:50] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:51] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:52] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:53] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:55] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:56] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:57] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:58] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:22:59] failed fetching digests from subpartition members;
dropping merge response
weblogic@FLPCOBA1:/LOG/sli_batch/tomcat-co-1 > cat console.log |grep 14:23
INFO  [11/17 14:14:23] SAM File Service has been requested with following
options:page No = 1, pageSize = 100, startRowNumber = 0, startPosition = 0,
totalPageCount = 0, totalRowNumber = 0,  totalCount = 0, parameter =
/batch_sam/SAM/work/mc/P.MCRN2011D3.05024D, additional parameters = {}
ERROR [11/17 14:23:01] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:23:02] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:23:03] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:23:04] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:23:05] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:23:07] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:23:08] failed fetching digests from subpartition members;
dropping merge response
ERROR [11/17 14:23:09] failed fetching digests from subpartition members;
dropping merge response
WARN  [11/17 14:23:10] FLPCOBA1-43250: dropped message from FLPCOA03-65117
(not in xmit_table), keys are [FLPCOA06-31629, FLPCOBA1-43250],
view=[FLPCOBA1-43250|1] [FLPCOBA1-43250, FLPCOA06-31629]
INFO  [11/17 14:23:27] Member view: [FLPCOBA1-43250, FLPCOA06-46079,
FLPCOA06-31629, FLPCOA03-65117, FLPCOA02-64862, FLPCOA03-13929,
 FLPCOBA1-61893, FLPCOA01-48349, FLPCOA02-40798, FLPCOA04-34052,
FLPCOA05-20663, FLPCOA05-20126, FLPCOA04-42687, FLPCOA01-51680]

*-- Server FLPCOA05*
WARN  [11/17 14:23:15] there was more than 1 candidate for coordinator:
{FLPCOBA1-43250=1, FLPCOA03-65117=4}
INFO  [11/17 14:23:15] Member view: [FLPCOA03-65117, FLPCOA04-34052,
FLPCOA01-51680, FLPCOBA1-61893, FLPCOA05-20663, FLPCOA06-46079,
 FLPCOA02-64862, FLPCOA04-42687, FLPCOA02-40798, FLPCOA05-20126,
FLPCOA03-13929]
WARN  [11/17 14:23:26] FLPCOA05-20126: dropped message from FLPCOBA1-43250
(not in xmit_table), keys are [FLPCOA03-13929, FLPCOA01-4
8349, FLPCOA04-42687, FLPCOA05-20126, FLPCOBA1-61893, FLPCOA01-51680,
FLPCOA03-65117, FLPCOA04-34052, FLPCOA02-40798, FLPCOA02-64862
, FLPCOA05-20663, FLPCOA06-46079], view=[FLPCOA03-65117|91] [FLPCOA03-65117,
FLPCOA04-34052, FLPCOA01-51680, FLPCOBA1-61893, FLPCOA0
5-20663, FLPCOA06-46079, FLPCOA02-64862, FLPCOA04-42687, FLPCOA02-40798,
FLPCOA05-20126, FLPCOA03-13929, FLPCOA01-48349]
WARN  [11/17 14:23:27] FLPCOA05-20126: dropped message from FLPCOBA1-43250
(not in xmit_table), keys are [FLPCOA03-13929, FLPCOA01-4
8349, FLPCOA04-42687, FLPCOA05-20126, FLPCOBA1-61893, FLPCOA01-51680,
FLPCOA03-65117, FLPCOA04-34052, FLPCOA02-40798, FLPCOA02-64862
, FLPCOA05-20663, FLPCOA06-46079], view=[FLPCOA03-65117|91] [FLPCOA03-65117,
FLPCOA04-34052, FLPCOA01-51680, FLPCOBA1-61893, FLPCOA0
5-20663, FLPCOA06-46079, FLPCOA02-64862, FLPCOA04-42687, FLPCOA02-40798,
FLPCOA05-20126, FLPCOA03-13929, FLPCOA01-48349]
INFO  [11/17 14:23:27] Member view: [FLPCOBA1-43250, FLPCOA06-46079,
FLPCOA06-31629, FLPCOA03-65117, FLPCOA02-64862, FLPCOA03-13929,
 FLPCOBA1-61893, FLPCOA01-48349, FLPCOA02-40798, FLPCOA04-34052,
FLPCOA05-20663, FLPCOA05-20126, FLPCOA04-42687, FLPCOA01-51680]
INFO  [11/17 14:23:27] My rank -> 11
WARN  [11/17 14:23:03] join(FLPCOA05-20126) sent to FLPCOA06-64573 timed out
(after 3000 ms), retrying
WARN  [11/17 14:23:06] join(FLPCOA05-20126) sent to FLPCOA06-64573 timed out
(after 3000 ms), retrying
WARN  [11/17 14:23:09] join(FLPCOA05-20126) sent to FLPCOA06-64573 timed out
(after 3000 ms), retrying
WARN  [11/17 14:23:12] join(FLPCOA05-20126) sent to FLPCOBA1-43250 timed out
(after 3000 ms), retrying
WARN  [11/17 14:23:15] join(FLPCOA05-20126) sent to FLPCOA01-37784 timed out
(after 3000 ms), retrying
WARN  [11/17 14:23:15] there was more than 1 candidate for coordinator:
{FLPCOBA1-43250=1, FLPCOA03-65117=4}
INFO  [11/17 14:23:15] Member view: [FLPCOA03-65117, FLPCOA04-34052,
FLPCOA01-51680, FLPCOBA1-61893, FLPCOA05-20663, FLPCOA06-46079,
 FLPCOA02-64862, FLPCOA04-42687, FLPCOA02-40798, FLPCOA05-20126,
FLPCOA03-13929]
INFO  [11/17 14:23:15] My rank -> 9
INFO  [11/17 14:23:15] Clustered information will inquire to my
friend(pirmay or secondary)[FLPCOA05-20663]
INFO  [11/17 14:23:16] Received a result from FLPCOA05-20663, Command:
GetClusteredCondition
INFO  [11/17 14:23:16] Clustered information received from my friend(pirmay
or secondary)[server name=FLPCOA05-20663, blocking=false
, execution limit=100]
INFO  [11/17 14:24:14] Received a result from FLPCOA06-31629, Command:
GetClusteredCondition
INFO  [11/17 14:24:14] Clustered Condition received from
FLPCOA06-31629[server name=FLPCOA06-31629, blocking=false, execution limit=
100]
INFO  [11/17 14:24:14] Clustered Condition requested for FLPCOA03-65117
INFO  [11/17 14:24:14] Received a result from FLPCOA03-65117, Command:
GetClusteredCondition
INFO  [11/17 14:24:14] Clustered Condition received from
FLPCOA03-65117[server name=FLPCOA03-65117, blocking=false, execution limit=
100]
INFO  [11/17 14:24:14] Clustered Condition requested for FLPCOA02-64862
INFO  [11/17 14:24:15] Received a result from FLPCOA02-64862, Command:
GetClusteredCondition
INFO  [11/17 14:24:15] Clustered Condition received from
FLPCOA02-64862[server name=FLPCOA02-64862, blocking=false, execution limit=
100]
WARN  [11/17 14:23:03] join(FLPCOA05-20126) sent to FLPCOA06-64573 timed out
(after 3000 ms), retrying
WARN  [11/17 14:23:06] join(FLPCOA05-20126) sent to FLPCOA06-64573 timed out
(after 3000 ms), retrying
WARN  [11/17 14:23:09] join(FLPCOA05-20126) sent to FLPCOA06-64573 timed out
(after 3000 ms), retrying
WARN  [11/17 14:23:12] join(FLPCOA05-20126) sent to FLPCOBA1-43250 timed out
(after 3000 ms), retrying
WARN  [11/17 14:23:15] join(FLPCOA05-20126) sent to FLPCOA01-37784 timed out
(after 3000 ms), retrying
WARN  [11/17 14:23:15] there was more than 1 candidate for coordinator:
{FLPCOBA1-43250=1, FLPCOA03-65117=4}
INFO  [11/17 14:23:15] Member view: [FLPCOA03-65117, FLPCOA04-34052,
FLPCOA01-51680, FLPCOBA1-61893, FLPCOA05-20663, FLPCOA06-46079,
 FLPCOA02-64862, FLPCOA04-42687, FLPCOA02-40798, FLPCOA05-20126,
FLPCOA03-13929]
WARN  [11/17 14:23:26] FLPCOA05-20126: dropped message from FLPCOBA1-43250
(not in xmit_table), keys are [FLPCOA03-13929, FLPCOA01-4
8349, FLPCOA04-42687, FLPCOA05-20126, FLPCOBA1-61893, FLPCOA01-51680,
FLPCOA03-65117, FLPCOA04-34052, FLPCOA02-40798, FLPCOA02-64862
, FLPCOA05-20663, FLPCOA06-46079], view=[FLPCOA03-65117|91] [FLPCOA03-65117,
FLPCOA04-34052, FLPCOA01-51680, FLPCOBA1-61893, FLPCOA0
5-20663, FLPCOA06-46079, FLPCOA02-64862, FLPCOA04-42687, FLPCOA02-40798,
FLPCOA05-20126, FLPCOA03-13929, FLPCOA01-48349]
WARN  [11/17 14:23:27] FLPCOA05-20126: dropped message from FLPCOBA1-43250
(not in xmit_table), keys are [FLPCOA03-13929, FLPCOA01-4
8349, FLPCOA04-42687, FLPCOA05-20126, FLPCOBA1-61893, FLPCOA01-51680,
FLPCOA03-65117, FLPCOA04-34052, FLPCOA02-40798, FLPCOA02-64862
, FLPCOA05-20663, FLPCOA06-46079], view=[FLPCOA03-65117|91] [FLPCOA03-65117,
FLPCOA04-34052, FLPCOA01-51680, FLPCOBA1-61893, FLPCOA0
5-20663, FLPCOA06-46079, FLPCOA02-64862, FLPCOA04-42687, FLPCOA02-40798,
FLPCOA05-20126, FLPCOA03-13929, FLPCOA01-48349]
INFO  [11/17 14:23:27] Member view: [FLPCOBA1-43250, FLPCOA06-46079,
FLPCOA06-31629, FLPCOA03-65117, FLPCOA02-64862, FLPCOA03-13929,
 FLPCOBA1-61893, FLPCOA01-48349, FLPCOA02-40798, FLPCOA04-34052,
FLPCOA05-20663, FLPCOA05-20126, FLPCOA04-42687, FLPCOA01-51680]


*- Our configuration file for JGroups*
<!--^M
    TCP based stack, with flow control and message bundling. This is usually
used when IP^M
    multicasting cannot be used in a network, e.g. because it is disabled
(routers discard multicast).^M
    Note that TCP.bind_addr and TCPPING.initial_hosts should be set,
possibly via system properties, e.g.^M
    -Djgroups.bind_addr=192.168.5.2 and
-Djgroups.tcpping.initial_hosts=192.168.5.2[7800]^M
    author: Bela Ban^M
    version: $Id: tcp.xml,v 1.40 2009/12/18 09:28:30 belaban Exp $^M
-->^M
<config xmlns="urn:org:jgroups"^M
        xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"^M
        xsi:schemaLocation="urn:org:jgroups
http://www.jgroups.org/schema/JGroups-2.8.xsd">^M
    <TCP bind_port="7800"^M
         loopback="true"^M
         recv_buf_size="${tcp.recv_buf_size:20M}"^M
         send_buf_size="${tcp.send_buf_size:640K}"^M
         discard_incompatible_packets="true"^M
         max_bundle_size="64K"^M
         max_bundle_timeout="30"^M
         enable_bundling="true"^M
         use_send_queues="true"^M
         sock_conn_timeout="300"^M
         timer.num_threads="4"^M
         ^M
         thread_pool.enabled="true"^M
         thread_pool.min_threads="1"^M
         thread_pool.max_threads="25"^M
         thread_pool.keep_alive_time="5000"^M
         thread_pool.queue_enabled="false"^M
         thread_pool.queue_max_size="100"^M
         thread_pool.rejection_policy="discard"^M
^M
         oob_thread_pool.enabled="true"^M
         oob_thread_pool.min_threads="1"^M
         oob_thread_pool.max_threads="8"^M
         oob_thread_pool.keep_alive_time="5000"^M
         oob_thread_pool.queue_enabled="false"^M
         oob_thread_pool.queue_max_size="100"^M
         oob_thread_pool.rejection_policy="discard"/>^M
                         ^M
    <TCPPING timeout="3000"^M

initial_hosts="${jgroups.tcpping.initial_hosts:100.254.163.51[7800],
100.254.161.11[7800], 100.254.161.12[7800], 100.25
4.161.13[7800], 100.254.161.14[7800], 100.254.161.15[7800],
100.254.161.16[7800]}"^M
             port_range="2"^M
             num_initial_members="14"/>^M
 <MERGE2  min_interval="10000"^M
             max_interval="30000"/>^M
    <FD_SOCK/>^M
    <FD timeout="3000" max_tries="3" />^M
    <VERIFY_SUSPECT timeout="1500"  />^M
    <BARRIER />^M
    <pbcast.NAKACK^M
                   use_mcast_xmit="false" gc_lag="0"^M
                   retransmit_timeout="300,600,1200,2400,4800"^M
                   discard_delivered_msgs="true"/>^M
    <UNICAST timeout="300,600,1200" />^M
    <pbcast.STABLE stability_delay="1000" desired_avg_gossip="50000"^M
                   max_bytes="400K"/>^M
    <pbcast.GMS print_local_addr="true" join_timeout="3000"^M
                view_bundling="true"/>^M
    <FC max_credits="2M"^M
        min_threshold="0.10"/>^M
    <FRAG2 frag_size="60K"  />^M
    <pbcast.STREAMING_STATE_TRANSFER/>^M
</config>
-- 
Blessings.
Kuwon Kang
............................................................
IT specialist and architect.
Java technology engineer.
JavaEE architect/developer.
Spring Framework specialist.
Solution developer.

- Model Driven Architecture.
- Test Driven Development.
- Test Driven Software Design.
- Refactoring Oriented Development.
- Practical design and modeling.
- Robust  Software Design and Engineering.
............................................................
Prever,Inc.
 http://www.prever.co.kr/
Blog:
 http://josh.prever.co.kr/
M 010 9440 9090
O 070 8682 0790
............................
............................................................

God's love is eternal and leads us to live for heaven and His glory

------------------------------------------------------------------------------
Beautiful is writing same markup. Internet Explorer 9 supports
standards for HTML5, CSS3, SVG 1.1,  ECMAScript5, and DOM L2 & L3.
Spend less time writing and  rewriting code and more time creating great
experiences on the web. Be a part of the beta today
http://p.sf.net/sfu/msIE9-sfdev2dev

_______________________________________________
Javagroups-development mailing list