Re: utf8 messages
Brandon Long <[email protected]> Thu, 14 Aug 2014 17:41:24 -0700
| Newsgroups | gmane.ietf.rfc822 |
|---|---|
| Message-ID | <CABa8R6sbFQHaP=YgrejJjUKJS+20BFP+kATZ+PrDnTgUhPMpHw@mail.gmail.com> |
--===============7011750636945576961== Content-Type: multipart/alternative; boundary=90e6ba614c64e9e5e90500a047db --90e6ba614c64e9e5e90500a047db Content-Type: text/plain; charset=UTF-8 On Wed, Aug 13, 2014 at 6:34 PM, Ned Freed <[email protected]> wrote: > > Let me try one more time, since something isn't making it through. > > > I have three messages. One message has an entirely 7bit header with 2047 > > encoded subject. Another message is a 6532 message, with the subject in > > utf8. A third message is has a cp-1250 8bit subject. There are two 8bit > > bytes in the subject in both of the last two messages, and in the cp1250 > > case, those two bytes happen to also be a valid utf8 character. > > > We want to be able to parse all three of those and do so correctly. We > > know the third type is technically invalid, but we see millions of such > > messages every day, dropping all of those would be a dis-service to our > > users. We currently see way more of such messages than we do of 6532 > > messages... though in practice, the most common charset now is utf-8, so > I > > guess those are now the same as 6532 messages that have leaked. > > I thought I understood the problem you were attempting to solve, but now > I'm > totally confused, because this seems to hqve nothing to do with additional > labeling of legitimate EAI messages at all. > My point is that without a label, I can't tell the difference between the 6532 messages and the illegitimate messages, given just the message. Brandon --90e6ba614c64e9e5e90500a047db Content-Type: text/html; charset=UTF-8 Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr"><br><div class=3D"gmail_extra"><br><br><div class=3D"gmail= _quote">On Wed, Aug 13, 2014 at 6:34 PM, Ned Freed <span dir=3D"ltr"><<a= href=3D"mailto:[email protected]" target=3D"_blank" class=3D"cremed">n= [email protected]</a>></span> wrote:<br> <blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p= x #ccc solid;padding-left:1ex"><div class=3D"">> Let me try one more tim= e, since something isn't making it through.<br> <br> > I have three messages.=C2=A0 One message has an entirely 7bit header w= ith 2047<br> > encoded subject.=C2=A0 Another message is a 6532 message, with the sub= ject in<br> > utf8.=C2=A0 A third message is has a cp-1250 8bit subject.=C2=A0 There= are two 8bit<br> > bytes in the subject in both of the last two messages, and in the cp12= 50<br> > case, those two bytes happen to also be a valid utf8 character.<br> <br> > We want to be able to parse all three of those and do so correctly.=C2= =A0 We<br> > know the third type is technically invalid, but we see millions of suc= h<br> > messages every day, dropping all of those would be a dis-service to ou= r<br> > users.=C2=A0 We currently see way more of such messages than we do of = 6532<br> > messages... though in practice, the most common charset now is utf-8, = so I<br> > guess those are now the same as 6532 messages that have leaked.<br> <br> </div>I thought I understood the problem you were attempting to solve, but = now I'm<br> totally confused, because this seems to hqve nothing to do with additional<= br> labeling of legitimate EAI messages at all.<br></blockquote><div><br></div>= <div>My point is that without a label, I can't tell the difference betw= een the 6532 messages and the illegitimate messages, given just the message= .</div> <div><br></div><div>Brandon</div></div></div></div> --90e6ba614c64e9e5e90500a047db-- --===============7011750636945576961== Content-Type: text/plain; charset="us-ascii" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit Content-Disposition: inline _______________________________________________ ietf-822 mailing list [email protected] https://www.ietf.org/mailman/listinfo/ietf-822 --===============7011750636945576961==--