Re: Question related with DAUD Message

Andrew Booth <[email protected]>
Newsgroups gmane.ietf.sigtran
Message-ID <[email protected]>
Hi Brian,

Here's an attempt at clarification.

My previous comments were made while wearing my "paranoid ASP developer
hat".  While I'm wearing that hat I don't like to make too many
assumptions about the behaviour of the SG, where I can easily avoid it. 
Obviously, some bugs are not worth working around.

I find it interesting to note that the behaviour of an SG during
overload is not specified.  This is different than the SS7
specifications which explicitly tell you how to deal with many kinds of
congestion.  Maybe it's worth putting some recommendations together (or
maybe I missed the guidance that was there).  In the absence of such
guidance an ASP will have to be pretty conservative about its
assumptions, or delve into non-portability by discussing with SG
vendors.  That's a bit of an aside from this thread, so I'll leave it
for you to decide whether to spin it into a separate thread.  I'll just
say that I agree that delaying a DAVA is better than discarding it. 
Whether closing an association is better than discarding a DAVA seems to
me to depends a bit on the network architecture.  In some cases (such as
when we know the ASP will audit), it may be better to discard the DAVA. 
In other cases it's probably worse.

More comments inline.

Brian F. G. Bidulock wrote:
> Andrew,
>
> I am a little confused by your statements.  A couple of comments:
>
> Andrew Booth wrote:                        (Wed, 31 Jan 2007 11:17:54)
>   
>> I think the ASP can send DAUD in any way it sees fit, according to the RFC.
>>
>> The following comments are my own and are not specified in the RFC.
>>
>> I think assuming that state is always synchronized is a dangerous
>> assumption.  For instance, what if the SG is in overload and discards a
>> DAVA?
>>     
>
> SCTP is a reliable transport.  DAVAs don't go missing without
> loss of the association.
>
>   
>> What if there's a bug on one end or the other?
>>     
>
> Bugs can cause failures: serious bugs, serious failures.
> Rigorous software test and field testing comes to mind ;)
>   
Sure, but sometimes bugs slip through anyway.  Some types of bugs are
notoriously difficult to detect even with rigorous testing.  Bugs that
require an overload condition then a DAVA probably fall into that
category.  The first 100 times you do the test you might not have a full
buffer (of whatever kind) when the DAVA should go out.
>   
>> What if the ASP is connected to two SGPs and receives DUNA + DAVA from SGP1
>> and DUNA + association loss from SGP2, in that order?
>>     
>
> The concensus of the designers of the M3UA and SUA protocols
> long ago (at about m3ua-06) was that SNMM are to be interpreted
> by the ASP on a per-SG basis rather than a per-SGP basis.  For
> example, when an AS receives DUNA from any SGP in the SG, it
> means DUNA for the entire SG, not just the SGP that sent the
> message.
>   
Exactly.  So, in the race condition described above, the ASP is left
believing that the destination is inaccessible to the SG (incorrectly). 
This will persist until manual intervention, or the ASP audits.

Whatever the reason for a mismatch in destination state, it's easy to
improve your chances of recovery by periodically auditing.

>   
>> The main risk would probably be a missing DAVA, since a missing DUNA
>> would get discovered by a response DUNA if traffic is sent to the
>> destination.
>>     
>
> The easier test instead of DAUD is to have the SCCP-User send
> traffic.  ITU specs require a DUNA every 8 or 10 messages for
> N-PCSTATE.  
If the ASP has another SG available, the DAUD may be a better choice.

> Nevertheless, back to the first question, why would
> an SG "miss" a DAVA?
>
> I do not understand why the SG would be designed in such way
> that it gets a positive response to an SST and yet somehow
> "discards" the DAVA.
>
> If it is going to lose simple management state it certainly
> cannot provide service and the better course of action would be
> sending ASP Inactive Ack, ASP Down Ack, or dropping the SCTP
> association.
>
> You see, if it cannot issue the DAVA message, how can it expect
> to be able to handle data traffic for the AS?  If it cannot
> issue the DAVA message, how is it to issue a DAUD response?
>   
It might, or it might not.  I think the odds of recovery are better if
you send the DAUD.  For instance, possibly the nodal congestion will
have abated by the time the DAUD comes in.
>   
>> >From an operational standpoint I'd be careful sending DAUD with a big
>> wildcard, since it's difficult to estimate how many responses you might
>> get and how fast, so it's hard to avoid potential overloads.  Also, you
>> won't necessarily know when you have all the responses (if you care).
>>     
>
> If an SGP cannot handle sending DAUD responses for, say, 1024
> point codes, how could it expect to handle traffic for same?
> (E.g, 80 messages per second for each of the 1024 point codes.)
> Management message are rare compared to the data traffic and
> only represent a very small percentage of the engineered
> bandwidth use between SGP and ASP.  If there is insufficient
> bandwidth for an SGP to send management messages, there is
> insufficient bandwidth for normal operation.  In that case,
> "delaying" (I don't agree with discarding) a DAVA message is a
> good thing.
>   
I was concerned more with DAUD(*), which could require a lot more
responses.  Also, management traffic and user traffic have different
models.  Management traffic can be quite bursty, and can for short
periods outstrip the expected user traffic.

Does that make any more sense?
Andrew
>   
>> Other than that, it's more bandwidth efficient to send multiple affected
>> PCs in one message, but that's only a concern if you're running over
>> bandwidth constrained networks or auditing many PCs.  So, take your pick.
>>     
>
> Again, management messages is a very small fraction of the
> engineered bandwidth for data traffic.  If management messages
> cannot be exchanged, neither can normal traffic.  Both the SGP
> and ASP have all the protocol means necessary for indicating
> same (ASP Inactive procedures, ASP Down procedures, and SCTP
> communications lost procedures.)
>
> --brian
>
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.