Re: bo: suppress script Tibt

"Doug Ewell" <[email protected]> Fri, 23 Mar 2018 12:24:45 -0700
Newsgroups gmane.ietf.languages
Message-ID <20180323122445.665a7a7059d7ee80bb4d670165c8327d.34b0424dba.wbe@email03.godaddy.com>
Élie Roux wrote:

> Is there a reason why
>
> Suppress-Script: Tibt
>
> is not present in the bo description? What is the process for deciding
> when a Suppress-Script is relevant? What other Scripts is Tibetan
> written in? Does it mean that "bo" is supposed to be ambiguous?

To answer these questions, it helps to know the history of
Suppress-Script.

Before RFC 4646, what we now call "script subtags" had to be registered
one by one, as part of a whole tag. Generative script subtags were added
in 2006, with RFC 4646.

At the time, there was a great deal of concern about placing the script
subtag (if any) in front of the region subtag (if any). The concerns
were that tag producers would include the new subtags reflexively and
unnecessarily, and that existing processes which matched tags
simplistically by truncating them from the right would be unable to
match, say, "fr-Latn-CA" with "fr-CA".

The solution was to provide an advisory field, Suppress-Script, to
inform tag producers, "You don't normally need to specify script X for
this language; it's usually obvious."

The tricky part is trying to define "unnecessarily" in this context.
Clearly it doesn't make sense, for thousands of languages, to try to
determine whether each has a single "default" script, and what that
script is, and what percentage of usage makes it "unnecessary" to
specify that script in a tag, and in which contexts.

So somebody, I don't remember who, created a list of language subtags
derived from 639-1 and 639-2 (i.e. those most likely to be understood by
pre-4646 processes) for which these questions could be answered with
minimal controversy, and those were assigned Suppress-Script values.
Over the last 11 years, a few more have been added, case by case. Nobody
claims that the list is comprehensive, but nobody has the time,
expertise, or authority to make it comprehensive.

It's important to remember that Suppress-Script is only a suggestion.
The fact that 'fr' has a Suppress-Script of 'Latn' doesn't prevent
anyone from creating a tag "fr-Latn", and it doesn't change the behavior
of any matching process. Likewise, the fact that 'bo' does not currently
have a Suppress-Script does not mean it is necessary to write "bo-Tibt"
for proper language identification.

It certainly doesn't mean the Tibetan language is "ambiguous" as to how
it is written. (A tag consisting of just a language subtag is always
ambiguous; "fr" could always be used for French written in Arabic or
Tifinagh script.)

For many languages, especially a language that one knows well, it may
not be hard to determine that *that one language* is written in script X
99.999% of the time, and it would normally be unhelpful for tag
producers to specify script X in a tag for that language. So there is a
process, of sorts: someone makes the argument and provides data, and
occasionally a Suppress-Script is added.

The major concern is that it will become a slippery slope: someone else
will say, "Language B is written in a single script *at least* as
commonly as language A is, so you should add this Suppress-Script as
well," and someone else will make the claim for Languages C and D, and
then someone will suggest we go through the entire Registry to flesh
this all out, once and for all. The gains in tagging and searching
accuracy are probably not worth the effort of verifying these claims and
updating the Registry.

Looking at the first few entries in the Registry, you can see many
individual Suppress-Script opportunities: "Afar, that's pretty much
always written in Latin... Aragonese, always Latin... Bislama, Latin..."
 At some point it gets difficult to separate evidence of "always" from
impression, and the amount of work to verify and update outweighs the
practical benefit.

Now, as you no doubt are aware, both Classical and Standard Tibetan are
written pretty much exclusively in the Tibetan script (except for
transcriptions, which are a different matter). Also, Standard Tibetan
has over a million speakers and 'bo' has always been available in
language tags, going back to RFC 1766 and earlier. So there might be
legacy processes that would choke on "bo-Tibt" although the overuse of
unnecessary script subtags today is less than was predicted in 2006.

So my opinion is that it would seem reasonable to add a Suppress-Script
of 'Tibt' for 'bo'. But we need to stand guard over the floodgates, and
not to get caught up in trying to make the Suppress-Script entries
comprehensive and authoritative, because that's not what it, or we, are
here for.
 
 
--
Doug Ewell | Thornton, CO, US | ewellic.org


_______________________________________________
Ietf-languages mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/ietf-languages