RE: Questions on draft-ietf-vrrp-ipv4-timers-02.txt

"Hott, Robert W CIV B35-Branch" <[email protected]>
Newsgroups gmane.ietf.vrrp
Message-ID <1929B8C5B318524495727D8A241DAFB20147FDC7@NAEAMILLEX03VA.nadsusea.nads.navy.mil>
Don,
    I appreciate your thoughts. See my comments below.
 

-----Original Message-----
From: Don Provan [mailto:[email protected]]
Sent: Wednesday, March 29, 2006 13:48
To: Hott, Robert W CIV B35-Branch; 'Steve Bates'; [email protected]
Cc: Odonoghue, Karen F CIV B35-Branch
Subject: RE: [VRRP] Questions on draft-ietf-vrrp-ipv4-timers-02.txt


OK, please excuse me for doubting you, but I'm
afraid this response has kinda supported my fears.
The issues you're talking about -- and I note
turning on logging as a specific case -- should
have the effect of delaying the processing of each
individual packet including the one that should
have prevented the flap. That suggests that the
over all timeout is too small, *not* that too few
packets were transmitted.
 
But since you've obviously investigated this way
beyond anything I've done, perhaps you already have
the more detailed data that would convince me.
First, are you sure your tests were checking the
number of packets transmitted without changing the
length of the timeout. And, second, were you able
to confirm that flapping was caused by packets that
were lost and not by packets that were delayed.
 
I'm not arguing against being able to configure
the number of retransmissions, as long as everyone
agrees VRRP needs it. I just want to make sure we
understand (and the spec expresses) what this really
accomplishes. In my experience (which was, admittedly,
limited to my implementation), I found that limits
to the timeout were *always* caused by packet latency,
never, ever by packet loss. I have seen problems
caused by packet loss, but with normal intervals
as much as with small ones.


I have done some testing and I think I have seen situations where
the ability to have additional advertisements has prevented
flapping. I did not perform enough analysis to determine if having
more advertisements during an overall time period managed to get an
advertisement through a queue where fewer during the same
overall time period did not, OR if some of the advertisements
were dropped due to queue overload. I suspect that the issue
was latency related, thus getting an advertisement in the
queue sooner might help but would not guarantee that flapping
would not occur.
 
For your info, I ran several tests where I had a desire to keep
the failover time under .6 seconds (this is an example as not to
point fingers at a specific implementation). To do that,
with the standard implementation, an advertisement interval
of .2 seconds was used. Flapping could be made to occur
under network loading or intense activity on the Master
router (say a denial of service type of attack, logging, or
network management accesses). Using vendor specific
options for sending more advertisements, the
advertisement interval was lowered to .1 second and
a total of 6 could be missed prior to a Backup taking
over. The stability of the protocol was greatly increased.
 
Is latency the issue, as opposed to packets dropped. You
are probably correct, that it is. Either way, it is a problem.
If packets aren't lost, getting them in the queue sooner
did help. There are environments that have a specific failover
requirement. The flexibility to send more advertisements
over the same time interval appears to help the stability of
the protocol. As you stated below, Master routers can be
set to inappropriate values and flapping will occur when this
happens.

 
On another note, I think we agree that the lowest,
reliable timeout period depends on many factors
in the implementation and in the environment; it
isn't something that can be defined by the
protocol. I wonder if, at these rates, we can or
should add some rules for "flapping recovery."
When a backup inappropriate takes over the VR,
the correct master knows it. I wonder if we should
add something for a master to announce that it was
inappropriately replaced and set a higher timeout
to avoid future mistakes? Just a thought off the
top of my head....
 

Yes, we do agree! You are right, that the real Master
does know when it has been replaced by another. I
could envision a configuration option that would allow
the real Master to alter its timeout, but I would want
the ability to prevent this action (raising the
timeout) too. I certainly think that the real Master
should report the potential flapping.

 
-don 
 

Thanks for your insight! I appreciate you perspective. My view is from
the end user and your view is not always obvious to me. Thanks again.
 
Bob Hott
Robert (Bob) W. Hott 
NSWC-DD 
Code B35, Bldg. 1500A/122A 
17320 Dahlgren Road 
Dahlgren, VA 22448-5100 
540-653-1497 (W) 
540-653-8673 (FAX) 
[email protected] (E-mail)

_______________________________________________
vrrp mailing list
[email protected]
https://www1.ietf.org/mailman/listinfo/vrrp
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.