Re: utf8 messages

Ned Freed <[email protected]> Thu, 14 Aug 2014 18:42:38 -0700 (PDT)
Newsgroups gmane.ietf.rfc822
Message-ID <[email protected]>
> On Wed, Aug 13, 2014 at 6:34 PM, Ned Freed <[email protected]> wrote:

> > > Let me try one more time, since something isn't making it through.
> >
> > > I have three messages.  One message has an entirely 7bit header with 2047
> > > encoded subject.  Another message is a 6532 message, with the subject in
> > > utf8.  A third message is has a cp-1250 8bit subject.  There are two 8bit
> > > bytes in the subject in both of the last two messages, and in the cp1250
> > > case, those two bytes happen to also be a valid utf8 character.
> >
> > > We want to be able to parse all three of those and do so correctly.  We
> > > know the third type is technically invalid, but we see millions of such
> > > messages every day, dropping all of those would be a dis-service to our
> > > users.  We currently see way more of such messages than we do of 6532
> > > messages... though in practice, the most common charset now is utf-8, so
> > I
> > > guess those are now the same as 6532 messages that have leaked.
> >
> > I thought I understood the problem you were attempting to solve, but now
> > I'm
> > totally confused, because this seems to hqve nothing to do with additional
> > labeling of legitimate EAI messages at all.
> >

> My point is that without a label, I can't tell the difference between the
> 6532 messages and the illegitimate messages, given just the message.

Yes, that's exactly what I thought. But given that, surely it makes more sense
to label the illegimate messages as such, especially since you're going to want
to set the EAI message bit for some of them, making it essentally an orthogonal
setting?

Did you even bother to read the rest of my response?

				Ned

_______________________________________________
ietf-822 mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/ietf-822