Re: Fwd: I-D ACTION:draft-hoffman-utf8headers-00.txt

Charles Lindsey <[email protected]> Thu, 01 Jan 2004 11:17:57 -0000
Newsgroups gmane.ietf.imaa
Message-ID <[email protected]>
On Wed, 31 Dec 2003 16:41:55 -0500, Martin Duerst <[email protected]> wrote:

> At 23:54 03/12/22 +0000, Charles Lindsey wrote:

>
>> In addition to that, SMTP is not the only mechanism for transporting 
>> email
>> (or netnews). There is UUCP. There is NNTP. ... Not all of these 
>> protocols will
>> want to implement a UTF-8-HEADERS extension.
>
> For X.400 and UUCP, my assumption would be that things would be
> downgraded anyway, which would mean to remove the header. Satellites
> are not a protocol, and carrier pigeons carry paper, where we don't
> even need UTF-8 :-). But in connection with NNTP, and for certain kinds
> of local processing (procmail,...), it would probably make sense.
> It may also ease implementation because it gives guidance for
> internal (mail spool) formats.

I don't think you would need to downgrade for UUCP, because it is already 
8-bit clean. But my point was that a message might happily wander around 
within one protocol (UUCP or NNTP) without anybody needing to care about 
the encoding or to check for "UTF-8-HEADERS". Then suddenly it arrives at 
a gateway into something else (e.g. SMTP or an IMAP store) where the 
distinction really matters. So the implementor of the gateway needs some 
quick way to discover whether this particular message needs special 
handling, and the presence of an extra header is probably the simplest way 
to do it.

That is also the reason why I don't like the "8:" header prefix. In some 
environments (notably Netnews) it would be much simpler to leave the 
headers in their present form (otherwise, all agents will have to learn to 
recognise a new set of headers which are really just synonyms for existing 
ones - that could be true of mail user agents too). The advantage of the 
special header is that agents that don't need to be aware of the 
distinction can just ignore it.

>> ...But to get random SMTP servers worldwide to upgrade will be a
>> hard slog, and it will only be the dedicated people who want to use the
>> facility who will have reason to apply the pressure to make it happen.
>
> I can see the 'political advantage' of such a header. But I don't see
> the relationship to server upgrade patterns.

My point was that a message may pass through several servers en route, and 
the intermediate ones are unlikely to be under the control of the end 
users (who are the ones who will actually benefit from having headers 
written in their own languages). But it is still desirable that those 
intermediate servers be upgraded so that UTF-8 stuff passes straight 
through them without unnecessary down- and up-gradings or, worse, 558 
bounces. Therefore, it is in our interests to make upgrading a server as 
simple and straightforward as possible, at least so far as stuff that is 
just passed through to other servers is concerned. That is why I spoke of 
a 'political advantage'.
>

> I think there are various ways to see this. You seem to be saying
> "we didn't get further than experimental for usefor, so better not
> try to get it for email". But I think it is better to see this as
> "usefor alone didn't make it, but email and usefor together should
> make it". Email carries a lot more weight within the IETF. The main
> issue with the UTF-8 extension for usefor only going to experimental,
> as far as I understand, was the interaction with email. This of course
> is gone once email is also moving towards UTF-8.

Yes, Email carrries more weight within IETF, and if that means this can be 
brought straight to standards track, then I would be delighted. But I am 
not so sure. It is a matter of timescale, and if an Experimental Protocol 
can get it in the field sooner, then that might be better. Again, it is a 
matter of politics - we should just go ahead, make our proposal, and then 
take soundings as to which way to play it.

>> So let me suggest a header so that UTF-8 users can mark their messages 
>> as
>> "unclean".
>
> I don't see anything 'unclean' in UTF-8.

You know that, and I know that, but some others out there don't. So maybe 
these messages need to go around waving handbells and shouting 'unclean', 
just so that other people can keep out of their way :-) .

> Allow me to start now: I think the name "Header-Transfer-Encoding"
> is problematic, because it will further increase confusion about
> the various encoding layers. Second, I very much think the
> distinction should be between US-ASCII and UTF-8, not 8bit and 7bit.

Yes, maybe we shoud just call it the "Foobar header" until we have decided 
exactly what it is to contain. As to whether the distinction is on the 
basis of "UTF-8" or of "8bit", there is just one problem, and that is the 
Chinese.

UTF-8 is official IETF policy. The French, the Scandinavians and even the 
Japanese will probably go along with it. But if you look at the 8bit 
headers that are already sloshing around the internet (and certainly in 
Usenet) you will observe that the code most commonly employed is some 
GBxxxx, and people are very reluctant to give up things that are "already 
working".

Yes, our standard will say that the code used in headers MUST be UTF-8, 
and other codes MUST NOT be used. That, sadly, is not sufficient to 
prevent it from happening. Which is why I suggest that our Foobar header 
should contain a possible handle to indicate other usages, though clearly 
that handle "MUST NOT be used".
>
>> Some people have doubts about including a language header. I put it 
>> there ........
>>
> I'm definitely very doubtful about this. There is already a
> Content-Language: header, and except for the odd case where all
> the headers are in one language, and the body in another, this
> parameter would not add anything.

It may be that the Content-Language header is sufficient. I was just 
pointing out that Bruce Lilly will be along presently, and that he has 
some IETF BCPs on his side :-( .

-- 
Charles H. Lindsey ---------At Home, doing my own thing------------------------
Tel: +44 161 436 6131 Fax: +44 161 436 6133   Web: http://www.cs.man.ac.uk/~chl
Email: [email protected]      Snail: 5 Clerewood Ave, CHEADLE, SK8 3JU, U.K.
PGP: 2C15F1A9      Fingerprint: 73 6D C2 51 93 A0 01 E7 65 E8 64 7E 14 A4 AB A5