Re: utf8 messages

Brandon Long <[email protected]> Thu, 14 Aug 2014 17:41:24 -0700
Newsgroups gmane.ietf.rfc822
Message-ID <CABa8R6sbFQHaP=YgrejJjUKJS+20BFP+kATZ+PrDnTgUhPMpHw@mail.gmail.com>
--===============7011750636945576961==
Content-Type: multipart/alternative; boundary=90e6ba614c64e9e5e90500a047db

--90e6ba614c64e9e5e90500a047db
Content-Type: text/plain; charset=UTF-8

On Wed, Aug 13, 2014 at 6:34 PM, Ned Freed <[email protected]> wrote:

> > Let me try one more time, since something isn't making it through.
>
> > I have three messages.  One message has an entirely 7bit header with 2047
> > encoded subject.  Another message is a 6532 message, with the subject in
> > utf8.  A third message is has a cp-1250 8bit subject.  There are two 8bit
> > bytes in the subject in both of the last two messages, and in the cp1250
> > case, those two bytes happen to also be a valid utf8 character.
>
> > We want to be able to parse all three of those and do so correctly.  We
> > know the third type is technically invalid, but we see millions of such
> > messages every day, dropping all of those would be a dis-service to our
> > users.  We currently see way more of such messages than we do of 6532
> > messages... though in practice, the most common charset now is utf-8, so
> I
> > guess those are now the same as 6532 messages that have leaked.
>
> I thought I understood the problem you were attempting to solve, but now
> I'm
> totally confused, because this seems to hqve nothing to do with additional
> labeling of legitimate EAI messages at all.
>

My point is that without a label, I can't tell the difference between the
6532 messages and the illegitimate messages, given just the message.

Brandon

--90e6ba614c64e9e5e90500a047db
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><br><div class=3D"gmail_extra"><br><br><div class=3D"gmail=
_quote">On Wed, Aug 13, 2014 at 6:34 PM, Ned Freed <span dir=3D"ltr">&lt;<a=
 href=3D"mailto:[email protected]" target=3D"_blank" class=3D"cremed">n=
[email protected]</a>&gt;</span> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><div class=3D"">&gt; Let me try one more tim=
e, since something isn&#39;t making it through.<br>
<br>
&gt; I have three messages.=C2=A0 One message has an entirely 7bit header w=
ith 2047<br>
&gt; encoded subject.=C2=A0 Another message is a 6532 message, with the sub=
ject in<br>
&gt; utf8.=C2=A0 A third message is has a cp-1250 8bit subject.=C2=A0 There=
 are two 8bit<br>
&gt; bytes in the subject in both of the last two messages, and in the cp12=
50<br>
&gt; case, those two bytes happen to also be a valid utf8 character.<br>
<br>
&gt; We want to be able to parse all three of those and do so correctly.=C2=
=A0 We<br>
&gt; know the third type is technically invalid, but we see millions of suc=
h<br>
&gt; messages every day, dropping all of those would be a dis-service to ou=
r<br>
&gt; users.=C2=A0 We currently see way more of such messages than we do of =
6532<br>
&gt; messages... though in practice, the most common charset now is utf-8, =
so I<br>
&gt; guess those are now the same as 6532 messages that have leaked.<br>
<br>
</div>I thought I understood the problem you were attempting to solve, but =
now I&#39;m<br>
totally confused, because this seems to hqve nothing to do with additional<=
br>
labeling of legitimate EAI messages at all.<br></blockquote><div><br></div>=
<div>My point is that without a label, I can&#39;t tell the difference betw=
een the 6532 messages and the illegitimate messages, given just the message=
.</div>
<div><br></div><div>Brandon</div></div></div></div>

--90e6ba614c64e9e5e90500a047db--


--===============7011750636945576961==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
ietf-822 mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/ietf-822

--===============7011750636945576961==--