Re: [FWD: RE: Philly (was: Re: Language Subtag Registration)]
"Doug Ewell" <[email protected]> Sun, 16 Sep 2018 11:49:19 -0600
| Newsgroups | gmane.ietf.languages |
|---|---|
| Message-ID | <B1EB4575E2D74E759ADABB47E3B83618@DougEwell> |
Michael, what are your thoughts on registering this subtag? Do you feel it is in scope for BCP 47? Should I post a proposed record to go with the registration form Nicholas posted on August 21? -- Doug Ewell | Thornton, CO, US | ewellic.org From: Nicholas Felker Sent: Saturday, September 15, 2018 18:00 To: Hugh Paterson Cc: ietf-languages ; Doug Ewell Subject: Re: [Ietf-languages] [FWD: RE: Philly (was: Re: Language Subtag Registration)] Language tagging for voice systems is useful not just in the ASR step, but in the reverse, generating the AI's text back to speech. In the sense that some languages are more common spoken, it makes sense that systems you talk to will respond in an informal, personalized way. On Tue, Sep 11, 2018 at 12:51 PM Hugh Paterson <[email protected]> wrote: In reply to Sebastian My one questions is: At what point is language tagging necessary for AI in the ASR task? It seems to me that in the ASR task that an AI response is the optimal way to go, but AI models need to reference something, and it doesn't always need to be language tags, it could be GIS tags or training models or other things. It is an entirely different matter if we are talking about the metadata needed in a linguistics/Institutional/ repository or linked data cloud to Identify a corpus of speech and map it to a model that an AI would reference during training. - Hugh On Tue, Sep 11, 2018 at 12:29 PM, Doug Ewell <[email protected]> wrote: Forwarding to the list on Sebastian's behalf. To subscribe to the ietf-languages list, visit https://www.ietf.org/mailman/listinfo/ietf-languages and follow the instructions under the heading "Subscribing to Ietf-languages". -- Doug Ewell | Thornton, CO, US | ewellic.org -------- Original Message -------- Subject: RE: [Ietf-languages] Philly (was: Re: Language Subtag Registration) From: Sebastian Velten Drude <[email protected]> Date: Tue, September 11, 2018 9:41 am To: Registrar ISO639-3 <mailto:[email protected]>, Doug Ewell <[email protected]>, ietf-languages <[email protected]> Cc: 'Christian Galinski' <mailto:[email protected]> Dear Doug, Melinda, all, Thanks for bringing me into this discussion (where is this discussion going on, how can I access and perhaps subscribe?). A few points in response to the questions below from my side. 1) Scientific / linguistic research on language variation is advancing quickly, and is making use of digital resources in written and oral form (recordings etc.). They are controlling for factors of linguistic variation much finer than major dialects. 2) As more and more of our interaction with machines is bound to be done ORALLY, being able to capture different dialects of major languages, and other kinds of linguistic varieties, such as sociolects, will soon become increasingly important, so that ideally one day the machines adapt to us, and we do not have to adapt our speech to them. 3) The framework for dealing with inner-language linguistic variation that we are working on will NOT provide all codes for all linguistic varieties – these are arguably almost infinite, but certainly in the hundreds of thousands (even if ignoring personal varieties). It just will provide a CONCEPTUAL AND PRINCIPAL FRAMEWORK on HOW to tag any linguistic variety that might be relevant. So we do not say that “Philadelphian English” should have code xyz, or any code for that matter, but that IF and WHEN somebody has a need to tag that, that it would be done in a way consistent with the tagging of other varieties (recognizing that this is a geographical variety (‘dialect’), different from, say, ‘African-American Vernacular English’, which is a sociolect, hence belonging to a different dimension, and recognizing that these two can be combined, to “dialect: Philadelphia; sociolect: African-American Vernacular English”. We hold that there is only a finite small number of dimensions, and that these need to be identified and applied consistently. 4) Once one admits that some tagging of not only regional or local varieties but also social varieties, medial varieties (spoken versus written, etc.) and varieties of other dimensions may be necessary, it becomes important to do that in a coherent fashion. Each regional variety can in principle have a number of sub-varieties. So where to stop? That depends on the use cases, and we are agnostic about which varieties will actually be registered and coded. 5) Only because there was no need to have a tag for “(Longchuan) Achang” (acn) thirty years ago, it did not mean that there is no need today, or will be no need tomorrow. Equally, that there may not be any need for tagging “Philadelphian English” today does not mean that it never will be relevant for nobody (see point (2)). So NO, I do NOT have “the vision of tagging all language varieties”, but the vision that if someone needs to tag some variety, that this is done in a principled and coherent way. When that need for tagging more than a few language varieties is here, and it is coming, we are better prepared and have a good framework to do so. That’s what we are aiming at to provide. I hope this is convincing enough? Christian in CC may have some points on this, too. Best wishes, Sebastian -- Sebastian Drude, Director The Vigdís International Centre for Multilingualism and Intercultural Understanding Veröld – hús Vigdísar • The University of Iceland • Brynjólfsgata 1 • 107 Reykjavík • Iceland Email: [email protected] • Phone: +354 525.4281 • Mobile: +354 695.6784 From: Registrar ISO639-3 [mailto:[email protected]] Sent: Tuesday, September 11, 2018 14:23 To: Sebastian Velten Drude <[email protected]> Subject: Fwd: [Ietf-languages] Philly (was: Re: Language Subtag Registration) Dear Sebastian, You are the one who has the vision to tag all varieties of language. Can you give Mr. Doug Ewell a good answer about why one would need to tag Philadelphian English as separate from standard American English (or Alabaman English)? Melinda ---------- Forwarded message ---------- From: Doug Ewell <[email protected]> Date: Mon, Sep 10, 2018 at 6:17 PM Subject: [Ietf-languages] Philly (was: Re: Language Subtag Registration) To: ietf-languages <[email protected]> Mark Davis wrote: > Now, I doubt that any significant number of people would use these > codes, much as I doubt that any significant number of people would use > en-philly. This highlights the question that I don't think has been answered about 'philly'. In BCP 47 we have what's known as a "taggable distinction," meaning a distinction between language varieties that a content producer or consumer would consider worthwhile to capture in a language tag. English vs. French is obviously a distinction that usually needs to be captured, which is why language tags exist in the first place. In many cases, "American English" and "British English" are also worth distinguishing, so we have region subtags. And so on. Is there a real-world situation in which someone would need to tag content as Philadelphia English, or search for content thus tagged, so as to distinguish it from content in non-Philadelphia English? Remember that language tags aren't really for identifying languages and varieties as such, but for identifying the language (including varieties) of actual content, so that users can find that content and so that appropriate resources, like spell-checkers, can be applied to that content. Even if the distinguishing features of Philadelphia English are interesting in their own right -- and no, it doesn't matter whether some of them are in evidence outside of Philadelphia proper -- they don't constitute a "taggable difference" unless there is actually a need to tag content differently. I'd be interested in hearing an argument that such a need exists. -- Doug Ewell | Thornton, CO, US | ewellic.org _______________________________________________ Ietf-languages mailing list [email protected] https://www.ietf.org/mailman/listinfo/ietf-languages