AS2 Reliability and an issue for comment.

"Dale Moberg" <[email protected]> Tue, 7 Mar 2006 10:11:14 -0700
Newsgroups gmane.ietf.ediint
Message-ID <BA94BA2FA0B10041BD9E4E24D520C1BD01F6A1D1@mail1.cyclonecommerce.com>
This is a multi-part message in MIME format.

------_=_NextPart_001_01C6420A.23ACCB7F
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

There are several residual Internet-Drafts under review by the Ediint
vender and user community that are responding to industry requests for
standardization. Most of these Internet-Drafts are informational RFCs
and not "official" Ediint chartered efforts. Examples include drafts on
Compression, the Features header, Certificate Exchange messages (CEM),
filename transmission, multipart payload support, and Reliability for
AS2. The question here concerns a small but important point in the AS2
reliability draft.

=20

AS2 has an option for either synchronous or asynchronous MDNs. The issue
for comment is concerned with synchronous MDN mode.

=20

Most venders have attempted to provide for some recovery from network
and/or server failures, and also to protect their customers from
resource exhaustion. When synchronous MDNs are used to transfer large
amounts of business data with compression, digital signatures, and
encryption applied to that data, heavily loaded systems can take a large
amount of time to produce the MDN to send back in the HTTP response. The
HTTP connection then needs to be held open for an unpredictable amount
of time, using resources on both sides.=20

=20

Now, because it is possible for an AS2 application to become "hung" on
the server side, software engineers often build in a "timer" that closes
a connection after some period of time. Unfortunately, the timeout can
occur before the HTTP requester (client) has received the protocol's
HTTP response. In addition, sometimes various HTTP intermediaries
(tunnels/proxies/gateways/etc) may time out a connection along the path
from client to final HTTP server based on "inactivity," and again
prevent the completion of the HTTP protocol.=20

=20

These exceptional conditions may be tied to an exception handler that
retries the HTTP request with its large payload. More often than not,
this retry of a large payload to an ever increasingly loaded server is a
recipe for further failure (and retry).  Because AS2 payloads are
growing from the tens to hundreds of megabytes, and the AS2 traffic on
existing servers is growing, the "timeout/retry" spiral has become an
operational difficulty for AS2 systems that needs consideration.

=20

The AS2 specification does have a built-in solution for this
problem-asynchronous MDN mode. However, users have indicated an interest
in whether there is anything else that might be done to address the
timeout problem and make AS2 in synchronous MDN mode more reliable.

=20

One direction is to try to make the timeout interval value "flexible"
and adapt it intelligently. While both transmission time and payload
size are known to the sender, the receiver load (often the most critical
factor) is not known. So it becomes difficult to arrive at an
intelligent solution that will not sometimes be wrong, which tends to
not satisfy AS2 endusers.

=20

Another direction might be to prohibit timeouts. This solution would
remove protections against tying up resources (both on sender and
receiver sides) in the really exceptional situations of a hung or dead
thread/process that did not clean up with an appropriate HTTP status
code (5xx range). Again there would be resistance to the adoption of
this solution by developers and engineers.

=20

Another direction might be to prohibit retries when using synchronous
MDNs. This direction effectively gives up on AS2 reliability. When the
specific error condition is recoverable (server down, connection
refused, transient network error, server temporarily busy), then retry
can be a reasonable way to enhance automation and reduce the need for
operational intervention and special manual handling.

=20

If the basic problem of the "timeout/retry" spiral is that there is no
way to tell intermediaries or the client that there is forward progress
being made on completing the HTTP response, then one remaining direction
is to provide a forward progress indicator. The HTTP protocol does have
an option for providing this feature that takes advantage of the HTTP
response "100 continue" status. In other words, a HTTP server can be
configured to send a sequence of "100 continue" replies, and a HTTP 1.1
client is effectively instructed to wait for a reply in the success
range ("2xx") or possibly failure ("5xx"). [ 3xx and 4xx cases ignored
here for simplicity-these statuses should be given as an initial HTTP
response IMO.]  This solution does not magically create resources when
they are falling short but at least it does potentially avoid the
"retry/timeout" spiral.

=20

Recommending that AS2 reliability makes use of this "keep alive" or
"forward progress" indicator would mark a change in current operational
modes. It is to be expected that this capability would be marked by a
special feature value (or AS2-version number if the feature header is
not approved) to allow a smooth transition to interoperability. Also,
how frequently to send 100 continues and how to react to a stretch of
time without "100 continues"  are issues needing consensus from the
participants on this list. This is assuming that people support the
direction here proposed so stakeholders should let their views be known!

=20

=20

=20

=20

=20


------_=_NextPart_001_01C6420A.23ACCB7F
Content-Type: text/html;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

<html xmlns:o=3D"urn:schemas-microsoft-com:office:office" =
xmlns:w=3D"urn:schemas-microsoft-com:office:word" =
xmlns=3D"http://www.w3.org/TR/REC-html40">

<head>
<meta http-equiv=3DContent-Type content=3D"text/html; =
charset=3Dus-ascii">
<meta name=3DGenerator content=3D"Microsoft Word 11 (filtered medium)">
<style>
<!--
 /* Style Definitions */
 p.MsoNormal, li.MsoNormal, div.MsoNormal
	{margin:0in;
	margin-bottom:.0001pt;
	font-size:12.0pt;
	font-family:"Times New Roman";}
a:link, span.MsoHyperlink
	{color:blue;
	text-decoration:underline;}
a:visited, span.MsoHyperlinkFollowed
	{color:purple;
	text-decoration:underline;}
span.EmailStyle17
	{mso-style-type:personal-compose;
	font-family:Arial;
	color:windowtext;}
@page Section1
	{size:8.5in 11.0in;
	margin:1.0in 1.25in 1.0in 1.25in;}
div.Section1
	{page:Section1;}
-->
</style>

</head>

<body lang=3DEN-US link=3Dblue vlink=3Dpurple>

<div class=3DSection1>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'>There are several residual Internet-Drafts under =
review by
the Ediint vender and user community that are responding to industry =
requests
for standardization. Most of these Internet-Drafts are informational =
RFCs and
not &#8220;official&#8221; Ediint chartered efforts. Examples include =
drafts on
Compression, the Features header, Certificate Exchange messages (CEM), =
filename
transmission, multipart payload support, and Reliability for AS2. The =
question
here concerns a small but important point in the AS2 reliability =
draft.<o:p></o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'>AS2 has an option for either synchronous or =
asynchronous
MDNs. The issue for comment is concerned with synchronous MDN =
mode.<o:p></o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'>Most venders have attempted to provide for some =
recovery
from network and/or server failures, and also to protect their customers =
from
resource exhaustion. When synchronous MDNs are used to transfer large =
amounts
of business data with compression, digital signatures, and encryption =
applied
to that data, heavily loaded systems can take a large amount of time to =
produce
the MDN to send back in the HTTP response. The HTTP connection then =
needs to be
held open for an unpredictable amount of time, using resources on both =
sides. <o:p></o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'>Now, because it is possible for an AS2 application to =
become
&#8220;hung&#8221; on the server side, software engineers often build in =
a &#8220;timer&#8221;
that closes a connection after some period of time. Unfortunately, the =
timeout can
occur before the HTTP requester (client) has received the =
protocol&#8217;s HTTP
response. In addition, sometimes various HTTP intermediaries =
(tunnels/proxies/gateways/etc)
may time out a connection along the path from client to final HTTP =
server based
on &#8220;inactivity,&#8221; and again prevent the completion of the =
HTTP
protocol. <o:p></o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'>These exceptional conditions may be tied to an =
exception
handler that retries the HTTP request with its large payload. More often =
than
not, this retry of a large payload to an ever increasingly loaded server =
is a recipe
for further failure (and retry). &nbsp;Because AS2 payloads are growing =
from the
tens to hundreds of megabytes, and the AS2 traffic on existing servers =
is
growing, the &#8220;timeout/retry&#8221; spiral has become an =
operational
difficulty for AS2 systems that needs =
consideration.<o:p></o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'>The AS2 specification does have a built-in solution =
for this
problem&#8212;asynchronous MDN mode. However, users have indicated an =
interest
in whether there is anything else that might be done to address the =
timeout
problem and make AS2 in synchronous MDN mode more =
reliable.<o:p></o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'>One direction is to try to make the timeout interval =
value &#8220;flexible&#8221;
and adapt it intelligently. While both transmission time and payload =
size are
known to the sender, the receiver load (often the most critical factor) =
is not
known. So it becomes difficult to arrive at an intelligent solution that =
will
not sometimes be wrong, which tends to not satisfy AS2 =
endusers.<o:p></o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'>Another direction might be to prohibit timeouts. This
solution would remove protections against tying up resources (both on =
sender
and receiver sides) in the really exceptional situations of a hung or =
dead thread/process
that did not clean up with an appropriate HTTP status code (5xx range). =
Again
there would be resistance to the adoption of this solution by developers =
and
engineers.<o:p></o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'>Another direction might be to prohibit retries when =
using
synchronous MDNs. This direction effectively gives up on AS2 =
reliability. When
the specific error condition is recoverable (server down, connection =
refused,
transient network error, server temporarily busy), then retry can be a
reasonable way to enhance automation and reduce the need for operational
intervention and special manual handling.<o:p></o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'>If the basic problem of the =
&#8220;timeout/retry&#8221; spiral
is that there is no way to tell intermediaries or the client that there =
is
forward progress being made on completing the HTTP response, then one =
remaining
direction is to provide a forward progress indicator. The HTTP protocol =
does
have an option for providing this feature that takes advantage of the =
HTTP
response &#8220;100 continue&#8221; status. In other words, a HTTP =
server can
be configured to send a sequence of &#8220;100 continue&#8221; replies, =
and a
HTTP 1.1 client is effectively instructed to wait for a reply in the =
success
range (&#8220;2xx&#8221;) or possibly failure (&#8220;5xx&#8221;). [ 3xx =
and
4xx cases ignored here for simplicity&#8212;these statuses should be =
given as
an initial HTTP response IMO.] &nbsp;This solution does not magically =
create
resources when they are falling short but at least it does potentially =
avoid
the &#8220;retry/timeout&#8221; spiral.<o:p></o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'>Recommending that AS2 reliability makes use of this =
&#8220;keep
alive&#8221; or &#8220;forward progress&#8221; indicator would mark a =
change in
current operational modes. It is to be expected that this capability =
would be
marked by a special feature value (or AS2-version number if the feature =
header
is not approved) to allow a smooth transition to interoperability. Also, =
how frequently
to send 100 continues and how to react to a stretch of time without =
&#8220;100
continues&#8221; &nbsp;are issues needing consensus from the =
participants on
this list. This is assuming that people support the direction here =
proposed so
stakeholders should let their views be =
known!<o:p></o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

<p class=3DMsoNormal><font size=3D2 face=3DArial><span =
style=3D'font-size:10.0pt;
font-family:Arial'><o:p>&nbsp;</o:p></span></font></p>

</div>

</body>

</html>

------_=_NextPart_001_01C6420A.23ACCB7F--