Re: bo: suppress script Tibt
"Doug Ewell" <[email protected]> Fri, 23 Mar 2018 12:24:45 -0700
| Newsgroups | gmane.ietf.languages |
|---|---|
| Message-ID | <20180323122445.665a7a7059d7ee80bb4d670165c8327d.34b0424dba.wbe@email03.godaddy.com> |
Élie Roux wrote: > Is there a reason why > > Suppress-Script: Tibt > > is not present in the bo description? What is the process for deciding > when a Suppress-Script is relevant? What other Scripts is Tibetan > written in? Does it mean that "bo" is supposed to be ambiguous? To answer these questions, it helps to know the history of Suppress-Script. Before RFC 4646, what we now call "script subtags" had to be registered one by one, as part of a whole tag. Generative script subtags were added in 2006, with RFC 4646. At the time, there was a great deal of concern about placing the script subtag (if any) in front of the region subtag (if any). The concerns were that tag producers would include the new subtags reflexively and unnecessarily, and that existing processes which matched tags simplistically by truncating them from the right would be unable to match, say, "fr-Latn-CA" with "fr-CA". The solution was to provide an advisory field, Suppress-Script, to inform tag producers, "You don't normally need to specify script X for this language; it's usually obvious." The tricky part is trying to define "unnecessarily" in this context. Clearly it doesn't make sense, for thousands of languages, to try to determine whether each has a single "default" script, and what that script is, and what percentage of usage makes it "unnecessary" to specify that script in a tag, and in which contexts. So somebody, I don't remember who, created a list of language subtags derived from 639-1 and 639-2 (i.e. those most likely to be understood by pre-4646 processes) for which these questions could be answered with minimal controversy, and those were assigned Suppress-Script values. Over the last 11 years, a few more have been added, case by case. Nobody claims that the list is comprehensive, but nobody has the time, expertise, or authority to make it comprehensive. It's important to remember that Suppress-Script is only a suggestion. The fact that 'fr' has a Suppress-Script of 'Latn' doesn't prevent anyone from creating a tag "fr-Latn", and it doesn't change the behavior of any matching process. Likewise, the fact that 'bo' does not currently have a Suppress-Script does not mean it is necessary to write "bo-Tibt" for proper language identification. It certainly doesn't mean the Tibetan language is "ambiguous" as to how it is written. (A tag consisting of just a language subtag is always ambiguous; "fr" could always be used for French written in Arabic or Tifinagh script.) For many languages, especially a language that one knows well, it may not be hard to determine that *that one language* is written in script X 99.999% of the time, and it would normally be unhelpful for tag producers to specify script X in a tag for that language. So there is a process, of sorts: someone makes the argument and provides data, and occasionally a Suppress-Script is added. The major concern is that it will become a slippery slope: someone else will say, "Language B is written in a single script *at least* as commonly as language A is, so you should add this Suppress-Script as well," and someone else will make the claim for Languages C and D, and then someone will suggest we go through the entire Registry to flesh this all out, once and for all. The gains in tagging and searching accuracy are probably not worth the effort of verifying these claims and updating the Registry. Looking at the first few entries in the Registry, you can see many individual Suppress-Script opportunities: "Afar, that's pretty much always written in Latin... Aragonese, always Latin... Bislama, Latin..." At some point it gets difficult to separate evidence of "always" from impression, and the amount of work to verify and update outweighs the practical benefit. Now, as you no doubt are aware, both Classical and Standard Tibetan are written pretty much exclusively in the Tibetan script (except for transcriptions, which are a different matter). Also, Standard Tibetan has over a million speakers and 'bo' has always been available in language tags, going back to RFC 1766 and earlier. So there might be legacy processes that would choke on "bo-Tibt" although the overuse of unnecessary script subtags today is less than was predicted in 2006. So my opinion is that it would seem reasonable to add a Suppress-Script of 'Tibt' for 'bo'. But we need to stand guard over the floodgates, and not to get caught up in trying to make the Suppress-Script entries comprehensive and authoritative, because that's not what it, or we, are here for. -- Doug Ewell | Thornton, CO, US | ewellic.org _______________________________________________ Ietf-languages mailing list [email protected] https://www.ietf.org/mailman/listinfo/ietf-languages