Re: UUCP, etc., and SMTP/822/MIME mail (was: Re: I-D ACTION:draft-hoffman-utf8headers-00.txt)

[email protected] Thu, 01 Jan 2004 15:24:59 -0800 (PST)
Newsgroups gmane.ietf.imaa
Message-ID <[email protected]>
> > We _could_ have adopted a rigid, Unicode-only, rule a decade ago
> > rather than doing charset-specific tagging in text content types
> > and 2047 encodings.

> IIRC, unicode wasn't known to be stable at that time.  it certainly
> wasn't widely adopted, and it's hard to imagine that we could have
> gotten consensus on such a rule.  it's only after 10 years' experience
> with unicode that we have some confidence in its character repertoire,
> and we have even less experience with other aspects of it.

A decade ago isn't a particularly relevant date in the development of MIME --
at that point in time MIME was already a draft standard, making such a change
quite difficult to make.

A much more relevant date is November 18-22, 1991. This is the date of the last
IETF meeting prior to the approval of MIME as a proposed standard.
Realistically, this was the last point in time at which a change as major as
uncategorical endorsement of a single, universal charset  specification could
have been made. The actual MIME specifications were subsequently submitted to
the IESG for approval in January 1992.

According to the Unicode web site the complete specification of Unicode 1.0
wasn't published until June, 1992. (Amusingly enough, that was the same month
in which RFC 1341 appeared.) In November 1991 the universal charset situation
was far from clear: What what then called 10646 seemed to be on the way
out and Unicode seemed to be on the way in but no conclusions had been reached.

This led to the following text appearing in RFC 1341:

            NOTE:   Beyond  US-ASCII,  an  enormous   proliferation   of
            character  sets  is  possible. It is the opinion of the IETF
            working group that a large number of character sets is NOT a
            good  thing.   We would prefer to specify a single character
            set that can be used universally for representing all of the
            world's   languages   in  electronic  mail.   Unfortunately,
            existing practice in several communities seems to  point  to
            the  continued  use  of  multiple character sets in the near
            future.  For this reason, we define names for a small number
            of  character  sets  for  which  a  strong  constituent base
            exists.    It is our hope  that  ISO  10646  or  some  other
            effort  will  eventually define a single world character set
            which can then be specified for use in Internet mail, but in
            the  advance of that definition we cannot specify the use of
            ISO  10646,  Unicode,  or  any  other  character  set  whose
            definition is, as of this writing, incomplete.

Even with 20:20 hindsight I fail to see any other reasonable course of action
we could have taken at the time.

				Ned