Re: TEI element for a grapheme?
Martin Holmes <[email protected]>
| Newsgroups | gmane.text.tei.general |
|---|---|
| Message-ID | <[email protected]> |
On 2016-09-27 04:03 PM, Piotr Bański wrote: > On 27/09/16 21:02, Martin Holmes wrote: >> This is a really interesting discussion, but I must confess that this: >> >> "Graphemes are sequences of one or more encoded characters that >> correspond to what users think of as characters." >> >> is one of the most unhelpful definitions imaginable. Who is a user, and >> who can tell what he or she might choose to think of as a "character"? > > That's your English bias. ;-) In Polish, we have a graph 'dotted z' and > a graph 'crossed z' (I'm sure Prof. Bień can quote their established > standardized names) which for 'users' (bah, take any plausible > definition) represent a single grapheme. Yes, this works nicely for a single language group using a single script. But there are so many cases which are not so simple. Consider this: :-) I typed three "characters" to create a combination sign. In my email client, and in my mind, they are distinct characters (as if I'd typed "Yay!"). In your client, the software might substitute U+1F600 GRINNING FACE, which is a single glyph, and which you might perceive as a single character; someone else might receive exactly what I sent, three glyphs, and yet still "think of" it as a "character" because they use emojis all the time. It just seems to me that any attempt at a definition which depends on the perception of a non-specific, presumably non-expert user is hardly a definition at all. Cheers, Martin > Similarly with the grapheme > realised phonetically as [w], in writing (it can have a tilde across 'l' > or over 'l'). Or the sound [t] in Russian, spelled with <m> or a 'little > (capital) T' in writing. This is free variation (context-independent, > across the population, but I guess you can also see it in the production > of the same individual). In English, 'plain z' and 'crossed z' can be > found in this sort of free variation, sometimes. > > You can also have complementary (context-dependent) distribution of > graphs representing the same grapheme, think of word-final vs. > non-word-final Greek sigma. (In English, in some contexts, word-initial > <g> and non-word-initial <dg> could qualify though I'm not sure I would > use this example in a 101-course, due to the diachronic variables that > have to be taken into consideration here.) > > Now, two remarks regarding Emmanuelle's original posting: > 1. I am still not sure how Emmanuelle wants to define 'grapheme'; I > think that a bunch of examples might help to solve this particular > issue, without having to delve too deep into philosophy of language > and/or semiotics. > > 2. <fs> is "feature structure" in TEI-speak. It can represent anything, > including phonemes. But it definitely is _not_ the "default encoding" > for phonemes. > > Best regards, > > Piotr > >> >> Cheers, >> Martin >> >> On 2016-09-27 11:21 AM, Janusz S. Bień wrote: >>> On Tue, Sep 27 2016 at 19:17 CEST, [email protected] writes: >>>> Hello, >>>> >>>> On 27/09/2016 18:32, Paul Schaffner wrote: >>>>> But I am probably misunderstanding. In my experience, all discussions >>>>> of characters, glyphs, graphs, and symbols end in a metaphysical >>>>> muddle. >>>> >>>> Unicode has a definition for all these terms. >>> >>> Really? What is the Unicode definition of graphs? Is there a Unicode >>> definition of symbols? >>> >>>> In the case of grapheme, >>> >>> In the case of grapheme the only Unicode definition I know of is that >>> from the glossary of Unicode terms (http://www.unicode.org/glossary/), >>> which I quote below. >>> >>>> the key definitions is that of _grapheme cluster_. >>> >>> Exactly. There is even a definition of extended grapheme cluster. The >>> term "grapheme" doesn't occur in the standard alone. >>> >>>> Please see >>>> <http://unicode.org/reports/tr29/#Grapheme_Cluster_Boundaries>, >>>> especially the part about Hangul and Devanagari. Reasoning over >>>> non-Latin scripts makes that discussion always more scientific and less >>>> metaphysical. :) >>> >>> The title of the quoted Unicode® Standard Annex #29 is "UNICODE TEXT >>> SEGMENTATION" and grapheme clusters are just fragment of some >>> specific texts. It's unclear for me whether they are clusters of >>> (undefined) grahemes or just graphemic clusters in some vague relation >>> to the non-Unicode meaning of the term "grapheme". >>> >>> Let me quote my recent posting to the Unicode mailing list in the thread >>> >>> http://www.unicode.org/mail-arch/unicode-ml/y2016-m09/0068.html >>> >>> entitled "graphemes": >>> >>> On Wed, Sep 21 2016 at 7:09 CEST, [email protected] writes: >>> >>> [...] >>> >>>> Let me remind the issues which started the thread: >>>> >>>> >>>> On Sun, Sep 18 2016 at 12:26 CEST, [email protected] writes: >>>>> Quote/Cytat - Christoph Päper <[email protected]> (pią, 16 >>>>> wrz 2016, 23:51:38): >>>>> >>>>>> Janusz S. Bień <[email protected]>: >>>>>>> >>>>>>> 1. Graphemes, if I understand correctly, are language dependent, … >>>>>> >>>>>> That’s true in linguistic terminology – well, at least within the >>>>>> more popular schools of thought –, but not in technical (i.e. >>>>>> Unicode) jargon. >>>> >>>> And what is "grapheme" in "technical (i.e. Unicode) jargon"? >>>> >>>>> >>>>> From the Unicode glossary: >>>>> >>>>> Grapheme. (1) A minimally distinctive unit of writing in the context >>>>> of a particular writing system.[...] (2) What a user thinks of as a >>>>> character. >>>>> >>>>> As for (2), cf. >>>>> >>>>> User-Perceived Character. What everyone thinks of as a character in >>>>> their script. >>>>> >>>>> So we have "a user" versus "everyone...in their script" - is the >>>>> difference intentional? Probably not. Anyway the definitions are >>>>> language/locale dependent. >>>> >>>> Does 'Grapheme' (2) make sense with "a (single?) user"? >>>> >>>> BTW, it is rather well know that the term "phoneme" was proposed first >>>> by a Polish linguist Jan Niecisław Ignacy Baudouin de Courtenay (13 >>>> March 1845 – 3 November 1929), cf. e.g >>>> https://en.wikipedia.org/wiki/Jan_Baudouin_de_Courtenay. It is much >>>> less know that he proposed also the term "grapheme". Let me quote >>>> Alexander Berg's "English Historical Linguistics vol. I" page 230 from >>>> Google Books: >>>> >>>> Since the introduction of the term grapheme by Baudouin de >>>> Courtenay in 1901 (Ruszkiewicz 1976:24-37, 1981 [1978], 20-34), >>>> it has been defined in various ways: >>>> >>>> [...] >>>> >>>> As can be seen from these quotatioms, the available definitions >>>> can be divided into two groups, corresponding to two main >>>> senses, >>>> and reflecting "conflicting linguistics views of the status of >>>> writing" (Henderson 1985:142): >>>> >>>> 1. a letter or cluster of letters referring to or >>>> corresponding with a >>>> single phoneme; >>>> >>>> 2. the minimal distinctive unit of a writing system. >>>> >>>> For me the first meaning (not mentioned at all in English Wikipedia) is >>>> the primary, i.e. more useful, meaning, as is has some practical >>>> applications e.g. for describing Polish hyphenation rules. >>> >>> Best regards >>> >>> Janusz >>> >>