timeouts in draft-ietf-mboned-driad-amt-discovery (was: MBONED WG Call For Adoption: draft-jholland-mboned-driad-amt-discovery)

"Holland, Jake" <[email protected]> Fri, 25 Jan 2019 22:34:04 +0000
Newsgroups gmane.ietf.mboned
Message-ID <[email protected]>
Hi Mikael,

Thanks for your comments, they led me to realize I had left out some
important context. I've submitted a new version of the draft:
https://datatracker.ietf.org/doc/draft-ietf-mboned-driad-amt-discovery/

Unfortunately I'm not seeing the diff links from the prior version
in the html--is there a way to get those turned on?  (I assume I
screwed that up when submitting, but I checked and saw it there
for others?)

Anyway, the bulk of the changes are the new section 2.4, plus some
edits to the abstract, and a new next-to-last paragraph in section 2.1.

I've added an explanation that this doc is an update to the relay discovery
process in AMT.  (Which carries with it the defined behavior from RFC
7450: whenever the relay discovery process is restarted, if the gateway
learns a new result, it would send a teardown to the old, and membership
updates to the new relay.)

I've added a section with recommendations structured as non-normative
guidelines and considerations, since experience is still limited with
the protocol operating at scale, and different gateways will sometimes
operate in very different environments.

Among other things, the new guidelines try to address your question
about DNS TTL expiry, but please let me know if you notice any problems.

I expect (and tried to explain in the new text) that making poor choices
with these guidelines would impact quality of service for the tunneled
traffic to some degree, but the basic operation of the protocol should
be unaffected.

Regardless of the choices when following those guidelines, there's some
opportunity for disruption to the traffic, and it'll be relatively
similar-looking whenever a rediscovery event happens.  The guidelines
basically try to advise avoiding them when possible, within reason.

If the gateway is at a network ingest point and you need less
disruption, in practice you'd probably want to run redundant tunnels and
skew the rediscovery timing.  I didn't explicitly address that, do you
think it needs a mention? (And if so, can you suggest any references or
example text from other RFCs?  This seems a pretty normal network design
practice for all kinds of traffic, but I've rarely seen it spelled out.)

(If the gateway is embedded in an app or host, you probably just cry
softly instead.  Or maybe set the app settings to a fixed value if it's
flapping for you.)

Thanks and regards,
Jake

On 2019-01-21, 00:49, "Mikael Abrahamsson" <[email protected]> wrote:
    What I am missing currently from the draft is how to handle long-lived S,G 
    joins in relationship with these DNS lookups. If the DNS information has 
    3600 seconds TTL, what happens when this DNS information expires. Is the 
    AMT gateway supposed to ask again, what happens if it now receives other 
    information than it received before? Is it supposed to re-evaluate, cease 
    communication with the current set of AMT relays, and re-establish 
    communication with the new ones listed in that preferred order? Or just 
    keep going as long as the current AMT session seems fine?
    
    The words "cache", "lifetime" etc doesn't exist in the draft. I'd like 
    some text on this, even if there is no need to handle it. Then it should 
    explicitly say that lookups are only done at S,G join time or whatever 
    other conditions that might happen within the AMT state machine for error 
    handling/handshakes etc.
    
    -- 
    Mikael Abrahamsson    email: [email protected]
    
    _______________________________________________
    MBONED mailing list
    [email protected]
    https://www.ietf.org/mailman/listinfo/mboned
    

_______________________________________________
MBONED mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/mboned