RE: SIGTRAN Plugtest Day 3
"Erickson, Mark" <[email protected]>
| Newsgroups | gmane.ietf.sigtran |
|---|---|
| Message-ID | <[email protected]> |
Brian: A robust protocol's description must not only provide the elements to for each side to signal what happened, but then must also provide recovery mechanisms for faults. Granted this is an operational issue, but doesn't a protocol also provide procedures for other 'operational' issues (congestion for one)? The ASP, in this case, is not making ANY assumptions - it knows the far end is in the ASP-INACTIVE state, and simply probing to see if the far end is ready to change state as it has been commanded. The point is requiring local management to intervene. What does this mean - the M3UA user now has to implement a retry mechanism for automatic recovery? Treating issues like this one in such a fashion only leads interop issues (due to differences in interpretation of the protocol description) which detracts from the robustness of the protocol (which, BTW, is why we're here in the first place). Mark -----Original Message----- From: Brian F. G. Bidulock [mailto:[email protected]] Sent: Thursday, April 19, 2007 8:40 AM To: Erickson, Mark Cc: SIGTRAN Mailing List Subject: Re: [Sigtran] SIGTRAN Plugtest Day 3 Mark, Recovery from a configuration error is an operational issue and not a protocol issue. The protocol provides all of the protocol elements necessary to determine the cause of the error. Whether the ASP attempts recover automatically or not is not only an operational consideration, but one that is local to the ASP. In doing so, the ASP either makes assumptions about provisioned information or is privy to provisioned information using mechanisms outside the protocol. Retransmitting after after Tack can be inappropriate for some operational environments and certainly cannot be made as a blanket recommendation in a protocol specification. From a protocol perspective, it is sufficient to say that the SGP MUST respond. What the ASP does as a result of a refusal is a local operational consideration. Nevertheless, two suggestions (MAY) appear in RFC 4666 and other UA RFCs: retransmit, inform management. Should people consider this when designing their APS? Yes, indeed. Does it have a place as a full recommendation in a protocol specification? No. --brian Erickson, Mark wrote: (Thu, 19 Apr 2007 07:01:08) > Brian: > > I'll agree that the SG must answer to the ACTIVE request - that's not my > point. A failure to do would use the Tack timer to force the ASP to > retry. > > However, if an error is sent back (as you indicate, Management Blocking) > - which could indicate a configuration issue, and the configuration > issue is resolved, there is no automatic recovery mechanism other than > to either rely on the Tack timer or force the ASP to the DOWN state BY > the SG (if it could!!) - and the only way I can see the latter happening > is to force the association down, which is just wrong (consider if the > association is a member of another AS that is in a traffic handling > state). > > MARK > > -----Original Message----- > From: Brian F. G. Bidulock [mailto:[email protected]] > Sent: Thursday, April 19, 2007 7:36 AM > To: Erickson, Mark > Cc: SIGTRAN Mailing List > Subject: Re: [Sigtran] SIGTRAN Plugtest Day 3 > > Erickson,, > > Erickson, Mark wrote: > (Thu, 19 Apr 2007 03:43:11) > > Brian: > > > > It's not a question of SCTP. What if the first one is refused by the > SG > > for some reason - why shouldn't the ASP keep resending to try to bring > > the ASP into service? > > Well, if the first one is refused, then an Error message MUST be > returned identifying the reason (e.g. Management Blocking). If the ASP > considers this refusal to be of a transient nature, it is welcome to > issue another ASP Active attempt at any time (even immediately). If the > ASP would like to simply ignore such ERROR messages and fall back on a > Tack timer, it is also welcome to do so. However, I do not think that > it can be recommended that the timer be used when the Error message will > do and so the SHOULD keyword does not apply. > > Tack was intended for the case where _no_ reply comes from the SGP > within the timer. In this case either there is a significant delay on > the association (longer than Tack RTT) or the SGP does not comply with > the ASP Active procedures. Both indicate a more serious problem and > that is why we said that Layer Management could be informed instead of > retransmitting another ASP Active into a black hole. > > I don't think that recommending retransmission solves issues with SGP > that refuse to send Error messages or ASPs that refuse to process them. > > --brian > > > > > If the SGP has been configured such that it refuses the ASP-ACTIVE > > request, if an operator changes the configuration as the SGP to permit > > the SGP to accept the request and bring the ASP active, there is no > way > > to signal the ASP what has occurred at the SG, and both sides should > not > > have to rely on some mechanism that forces the ASP to return to a DOWN > > state (such as restarting the association) to cause the process to be > > restarted automatically or require some manual intervention at the ASP > > end. > > > > Mark > > > > -- > Brian F. G. Bidulock > [email protected] > http://www.openss7.org/ -- Brian F. G. Bidulock [email protected] http://www.openss7.org/