CVS update: JGroups/doc/design RELAY.png RELAY.fig RELAY.txt

"Bela Ban" <[email protected]> Mon, 8 Nov 2010 12:02:45 +0000
Newsgroups gmane.comp.java.javagroups.cvs
Message-ID <[email protected]>
  User: belaban 
  Date: 10/11/08 12:02:45

  Added:       doc/design RELAY.png RELAY.fig RELAY.txt
  Log:
  new design
  
  Revision  Changes    Path
  1.1                  JGroups/doc/design/RELAY.png
  
  	<<Binary file>>
  
  
  1.1                  JGroups/doc/design/RELAY.fig
  
  Index: RELAY.fig
  ===================================================================
  #FIG 3.2  Produced by xfig version 3.2.5
  Landscape
  Center
  Inches
  Letter  
  100.00
  Single
  -2
  1200 2
  6 3375 4800 6375 9075
  2 2 0 1 0 7 50 -1 -1 0.000 0 0 -1 0 0 5
  	 3375 4800 6375 4800 6375 9075 3375 9075 3375 4800
  2 1 0 1 0 7 50 -1 -1 0.000 0 0 -1 0 0 2
  	 3375 5400 6375 5400
  2 1 0 1 0 7 50 -1 -1 0.000 0 0 -1 0 0 2
  	 3375 6000 6375 6000
  2 1 0 1 0 7 50 -1 -1 0.000 0 0 -1 0 0 2
  	 3375 6675 6375 6675
  2 1 0 1 0 7 50 -1 -1 0.000 0 0 -1 0 0 2
  	 3375 7350 6375 7350
  2 1 0 1 0 7 50 -1 -1 0.000 0 0 -1 0 0 2
  	 3375 8025 6375 8025
  2 1 0 1 0 7 50 -1 -1 0.000 0 0 -1 0 0 2
  	 3375 8550 6375 8550
  4 0 0 50 -1 16 16 0.0000 4 195 885 4425 5175 RELAY\001
  -6
  6 1575 825 4575 4200
  6 3750 1950 4350 2550
  1 3 0 1 0 7 50 -1 -1 0.000 1 0.0000 4050 2250 270 270 4050 2250 4200 2475
  4 0 0 50 -1 16 16 0.0000 4 195 180 3975 2400 A\001
  -6
  6 2100 1500 2700 2100
  1 3 0 1 0 7 50 -1 -1 0.000 1 0.0000 2400 1800 270 270 2400 1800 2550 2025
  4 0 0 50 -1 16 16 0.0000 4 195 195 2325 1875 C\001
  -6
  6 2475 2850 3075 3450
  1 3 0 1 0 7 50 -1 -1 0.000 1 0.0000 2775 3150 270 270 2775 3150 2925 3375
  4 0 0 50 -1 16 16 0.0000 4 195 180 2700 3225 B\001
  -6
  1 3 0 1 0 7 50 -1 -1 0.000 1 0.0000 3069 2364 1479 1479 3069 2364 4344 3114
  4 0 0 50 -1 16 16 0.0000 4 240 450 2850 4125 udp\001
  -6
  6 10125 2700 10725 3300
  1 3 0 1 0 7 50 -1 -1 0.000 1 0.0000 10425 3000 270 270 10425 3000 10575 3225
  4 0 0 50 -1 16 16 0.0000 4 195 165 10350 3075 F\001
  -6
  6 10200 1200 10800 1800
  1 3 0 1 0 7 50 -1 -1 0.000 1 0.0000 10500 1500 270 270 10500 1500 10650 1725
  4 0 0 50 -1 16 16 0.0000 4 195 180 10425 1650 E\001
  -6
  6 8700 1950 9300 2550
  1 3 0 1 0 7 50 -1 -1 0.000 1 0.0000 9000 2250 270 270 9000 2250 9150 2475
  4 0 0 50 -1 16 16 0.0000 4 195 195 8925 2325 D\001
  -6
  1 3 0 1 0 7 50 -1 -1 0.000 1 0.0000 9969 2289 1479 1479 9969 2289 11244 3039
  2 1 0 1 0 7 50 -1 -1 0.000 0 0 -1 0 0 2
  	 2175 9450 8400 9450
  2 1 0 2 0 7 50 -1 -1 0.000 0 0 7 1 0 2
  	1 1 4.00 60.00 120.00
  	 5475 9450 5475 4050
  2 1 0 2 0 7 50 -1 -1 0.000 0 0 7 1 0 2
  	1 1 4.00 60.00 120.00
  	 5475 5175 7425 5175
  2 4 0 3 0 7 50 -1 -1 0.000 0 0 7 0 0 5
  	 9450 2775 3525 2775 3525 1800 9450 1800 9450 2775
  4 0 0 50 -1 18 18 0.0000 4 225 2445 1800 600 Data Center NYC\001
  4 0 0 50 -1 18 18 0.0000 4 225 2415 8775 525 Data Center SFO\001
  4 0 0 50 -1 16 16 0.0000 4 195 990 7350 9300 Network\001
  4 0 0 50 -1 16 16 0.0000 4 255 3465 6525 5025 Relaying to other data center\001
  4 0 0 50 -1 16 16 0.0000 4 240 1320 4875 3900 Application\001
  4 0 0 50 -1 16 16 0.0000 4 210 360 6300 2325 tcp\001
  4 0 0 50 -1 16 16 0.0000 4 240 450 9750 4050 udp\001
  
  
  
  1.1                  JGroups/doc/design/RELAY.txt
  
  Index: RELAY.txt
  ===================================================================
  
  RELAY - replication between data centers
  ========================================
  
  Author: Bela Ban
  Version: $Id: RELAY.txt,v 1.1 2010/11/08 12:02:45 belaban Exp $
  
  This is an enhanced version of DataCenterReplication.txt with the ability to send unicast messages and to provide views
  to the application, which list members of all local clusters.
  
  We have data centers with a local cluster each in New York (NYC) and San Francisco (SFO). The idea is to relay
  traffic from NYC to SFO, and vice versa.
  
  In case of a site failure of NYC, the state is available in SFO, and all clients can be switched over to SFO and
  continue working with (almost) up-to-date data. The failing over of clients to SFO is outside the scope of this
  proposal, and could be done for example by changing DNS entries, load balancers etc.
  
  The data centers in NYC and SFO are *completely autonomous local clusters*. There are no stability, flow control or
  retransmission messages exchanged between NYC and SFO. This is critical because we don't want the SFO cluster to block
  for example on waiting for credits from a node in the NYC cluster !
  
  For the example, we assume that each site uses a UDP based stack, and relaying between the sites uses a
  TCP based stack, see figure RELAY.png.
  
  There is a local cluster, based on UDP, at each site and one global cluster, based on TCP, which connects the
  two sites. Each coordinator of the local cluster is also a member of the global cluster, e.g. member E in NYC
  (assuming it is the coordinator) is also member X of the TCP cluster. This is called a *relay* member. A relay
  member is always member of the local and global cluster.
  
  A relay member has a UDP stack which additionally contains a protocol RELAY at the top (shown in the bottom part
  of the figure). RELAY has a JChannel which connects to the TCP group, but *only* when it is (or becomes) coordinator
  of the local cluster. The configuration of the TCP channel is done via a property in RELAY.
  
  Any *multicast* message (we don't relay unicast messages) that is received by RELAY traveling
  up the stack is sent via the TCP channel to the other site. When received there, the corresponding RELAY
  protocol changes the destination of the message to null (those are multicast messages after all) and leaves
  the src (which might point to X if sent from NYC), then it sends the message down the stack, where it will get
  multicast to all members of the local cluster (including the sender). When a response is received which
  points to any non-local address (e.g. X), RELAY simply drops it.
  
  When forwarding a message to the local cluster, RELAY adds a header. When it receives the multicast message it
  forwarded itself, and a header is present, it does *not* relay it back to the other site but simply drops it.
  Otherwise, we would have a cycle.
  
  When a coordinator crashes or leaves, the next-in-line becomes coordinator and activates the RELAY protocol,
  connecting to the TCP channel and starting to relay messages.
  
  However, if we receive messages from the local cluster while the coordinator has crashed and the new one hasn't taken
  over yet, we'd lose messages. Therefore, we need additional functionality in RELAY which buffers the last N messages
  (or M bytes, or for T seconds) and numbers all messages sent. This is done by the second-in-line.
  
  When there is a coordinator failover, the new coordinator communicates briefly with the other site to determine
  which was the highest message relayed by it. It then forwards buffered messages with lower numbers and removes the
  remaining messages in the buffer. During this replay, message relaying is suspended.
  
  Therefore, a relay has to handle 3 types of messages from the global (TCP) cluster:
   (1) Regular multicast messages
   (2) A message asking for the highest sequence number received from another relay, and the response to this
   (3) A message stating that the other side will go down gracefully (no need to replay buffered messages)
  
  
  Example walkthrough
  -------------------
  - C (in the NYC cluster, with coordinator E) multicasts a message
  - A, B, C, D and E receive the multicast
  - D (second-in-line) buffer the message (bounded buffer)
  - E is the relay. The byte buffer is extracted and a new message M is created. M's source is C, the dest is null
    (= send to all). Note that the original headers are *not* sent with M. If this is needed, we need to revisit.
  - X receives M, drops it (because it is the sender, determined by the header).
  - Y receives M, adds a RELAY header and sends it down the stack
  - T, U, V, W and S receive M and deliver it
  - Y does not relay M because M has a header
  - Should some member reply (to X), then RELAY at Y will drop the message
  
  
  Issues
  ------
  
  Cluster RPCs in JBossCache
  --------------------------
  - When we invoke a cluster RPC in JBossCache, the destination list is (for example) {A,C,D,E}
  - This list is sent with the RPC (as a header)
  - The receiver drops the request if its local address is not part of the destination list
  ==> This would cause all nodes in the SFO cluster to drop the cluster RPC request !
  
  Buddy Replication
  -----------------
  - When we replicate from B to C, the relay (e.g. E) doesn't know about this and will not replicate !
  - How do we determine to whom we should replicate in SFO if we use buddy replication ?
  
  Identity
  --------
  - What if we have members that have the same address in NYC and SFO ?
  - If we use TCP plus bin_port (7800), then every member starts at 7800
  - If we happen to use IP local addresses in both SFO and NYC (e.g. 192.168.1.1), then
    we run into this issue
  
  Reception of message sent by self
  ---------------------------------
  - Because we relay messages only on *reception*, NOT on sending, a relay might not receive it own message
  - This is possible when the relay (e.g. E) calls RpcDispatcher.callRemoteMethods() with a target list which excludes
    itself
  ==> Solution: catch the message when sending, *not* receiving !
  
  No, doesn't work ! The relay would have to be active on every node ! This is not feasible as this would require
  every node to have a TCP connection to NYC ! We still need to do it on reception not sending. However, messages
  cannot exclude the sender
  
  
  

------------------------------------------------------------------------------
The Next 800 Companies to Lead America's Growth: New Video Whitepaper
David G. Thomson, author of the best-selling book "Blueprint to a 
Billion" shares his insights and actions to help propel your 
business during the next growth cycle. Listen Now!
http://p.sf.net/sfu/SAP-dev2dev