Re: TEI element for a grapheme?
Paul Schaffner <[email protected]>
| Newsgroups | gmane.text.tei.general |
|---|---|
| Message-ID | <1475020454.1777962.739059329.7D07142D@webmail.messagingengine.com> |
I used to cite this made-up sentence (which nevertheless is quite possible in my world) to illustrate the well-known problems with 'ff'. "Whilst suffering in ffraunce, he thought to himselff of the proverb 'Pan darffo treiglo pob tre / Da yw edrych tuag adre.'" 'suFFering' and 'himselFF' use the doubled f to indicate a voiceless fricative. The former is the common spelling, the latter an uncommon one. But to the everyday user, 'ff' is not a character: 'f' is the character in question, and is the same 'f' as in 'oF' (where it represents a voiced sound). These are the same everyday users who think of 't' and 'h' as characters but 'th' as a combination of characters, not as a single character composed of graphemes. 'ffraunce' contains an old-fashioned use of 'ff' to represent an upper-case (capitalized) variant of "f". Are upper-case and lower-case letters the same character or not? They are not glyph variants in the usual sense. And if they *are* different characters, then is "ff" simply a glyph variant of "F"? The usual answer is 'yes.' /F/ is a character that may be represented either as grapheme F or as grapheme cluster ff. And of course the Welsh 'ff' is usually thought of (by its users!) as a character in its own right, quite distinct from 'f'. If we believe these users, we should regard 'f' as a grapheme, which when it appears singly represents the character /f/ and when it is doubled represents the character /ff/. In actual transcription, it is rare indeed to find someone who distinguishes character from glyph in any kind of consistent way: most, I think, fall back on the Latin alphabet and its extensions as the default character inventory. (The 'capital' use of 'ff' is an exception, since I think most transcribers would capture it as "F" -- not all, however.) pfs On Tue, Sep 27, 2016, at 19:24, Martin Holmes wrote: > On 2016-09-27 04:03 PM, Piotr Bański wrote: > > On 27/09/16 21:02, Martin Holmes wrote: > >> This is a really interesting discussion, but I must confess that this: > >> > >> "Graphemes are sequences of one or more encoded characters that > >> correspond to what users think of as characters." > >> > >> is one of the most unhelpful definitions imaginable. Who is a user, and > >> who can tell what he or she might choose to think of as a "character"? > > > > That's your English bias. ;-) In Polish, we have a graph 'dotted z' and > > a graph 'crossed z' (I'm sure Prof. Bień can quote their established > > standardized names) which for 'users' (bah, take any plausible > > definition) represent a single grapheme. > > Yes, this works nicely for a single language group using a single > script. But there are so many cases which are not so simple. Consider > this: > > :-) > > I typed three "characters" to create a combination sign. In my email > client, and in my mind, they are distinct characters (as if I'd typed > "Yay!"). In your client, the software might substitute U+1F600 GRINNING > FACE, which is a single glyph, and which you might perceive as a single > character; someone else might receive exactly what I sent, three glyphs, > and yet still "think of" it as a "character" because they use emojis all > the time. > > It just seems to me that any attempt at a definition which depends on > the perception of a non-specific, presumably non-expert user is hardly a > definition at all. > > Cheers, > Martin > > > Similarly with the grapheme > > realised phonetically as [w], in writing (it can have a tilde across 'l' > > or over 'l'). Or the sound [t] in Russian, spelled with <m> or a 'little > > (capital) T' in writing. This is free variation (context-independent, > > across the population, but I guess you can also see it in the production > > of the same individual). In English, 'plain z' and 'crossed z' can be > > found in this sort of free variation, sometimes. > > > > You can also have complementary (context-dependent) distribution of > > graphs representing the same grapheme, think of word-final vs. > > non-word-final Greek sigma. (In English, in some contexts, word-initial > > <g> and non-word-initial <dg> could qualify though I'm not sure I would > > use this example in a 101-course, due to the diachronic variables that > > have to be taken into consideration here.) > > > > Now, two remarks regarding Emmanuelle's original posting: > > 1. I am still not sure how Emmanuelle wants to define 'grapheme'; I > > think that a bunch of examples might help to solve this particular > > issue, without having to delve too deep into philosophy of language > > and/or semiotics. > > > > 2. <fs> is "feature structure" in TEI-speak. It can represent anything, > > including phonemes. But it definitely is _not_ the "default encoding" > > for phonemes. > > > > Best regards, > > > > Piotr > > > >> > >> Cheers, > >> Martin > >> > >> On 2016-09-27 11:21 AM, Janusz S. Bień wrote: > >>> On Tue, Sep 27 2016 at 19:17 CEST, [email protected] writes: > >>>> Hello, > >>>> > >>>> On 27/09/2016 18:32, Paul Schaffner wrote: > >>>>> But I am probably misunderstanding. In my experience, all discussions > >>>>> of characters, glyphs, graphs, and symbols end in a metaphysical > >>>>> muddle. > >>>> > >>>> Unicode has a definition for all these terms. > >>> > >>> Really? What is the Unicode definition of graphs? Is there a Unicode > >>> definition of symbols? > >>> > >>>> In the case of grapheme, > >>> > >>> In the case of grapheme the only Unicode definition I know of is that > >>> from the glossary of Unicode terms (http://www.unicode.org/glossary/), > >>> which I quote below. > >>> > >>>> the key definitions is that of _grapheme cluster_. > >>> > >>> Exactly. There is even a definition of extended grapheme cluster. The > >>> term "grapheme" doesn't occur in the standard alone. > >>> > >>>> Please see > >>>> <http://unicode.org/reports/tr29/#Grapheme_Cluster_Boundaries>, > >>>> especially the part about Hangul and Devanagari. Reasoning over > >>>> non-Latin scripts makes that discussion always more scientific and less > >>>> metaphysical. :) > >>> > >>> The title of the quoted Unicode® Standard Annex #29 is "UNICODE TEXT > >>> SEGMENTATION" and grapheme clusters are just fragment of some > >>> specific texts. It's unclear for me whether they are clusters of > >>> (undefined) grahemes or just graphemic clusters in some vague relation > >>> to the non-Unicode meaning of the term "grapheme". > >>> > >>> Let me quote my recent posting to the Unicode mailing list in the thread > >>> > >>> http://www.unicode.org/mail-arch/unicode-ml/y2016-m09/0068.html > >>> > >>> entitled "graphemes": > >>> > >>> On Wed, Sep 21 2016 at 7:09 CEST, [email protected] writes: > >>> > >>> [...] > >>> > >>>> Let me remind the issues which started the thread: > >>>> > >>>> > >>>> On Sun, Sep 18 2016 at 12:26 CEST, [email protected] writes: > >>>>> Quote/Cytat - Christoph Päper <[email protected]> (pią, 16 > >>>>> wrz 2016, 23:51:38): > >>>>> > >>>>>> Janusz S. Bień <[email protected]>: > >>>>>>> > >>>>>>> 1. Graphemes, if I understand correctly, are language dependent, … > >>>>>> > >>>>>> That’s true in linguistic terminology – well, at least within the > >>>>>> more popular schools of thought –, but not in technical (i.e. > >>>>>> Unicode) jargon. > >>>> > >>>> And what is "grapheme" in "technical (i.e. Unicode) jargon"? > >>>> > >>>>> > >>>>> From the Unicode glossary: > >>>>> > >>>>> Grapheme. (1) A minimally distinctive unit of writing in the context > >>>>> of a particular writing system.[...] (2) What a user thinks of as a > >>>>> character. > >>>>> > >>>>> As for (2), cf. > >>>>> > >>>>> User-Perceived Character. What everyone thinks of as a character in > >>>>> their script. > >>>>> > >>>>> So we have "a user" versus "everyone...in their script" - is the > >>>>> difference intentional? Probably not. Anyway the definitions are > >>>>> language/locale dependent. > >>>> > >>>> Does 'Grapheme' (2) make sense with "a (single?) user"? > >>>> > >>>> BTW, it is rather well know that the term "phoneme" was proposed first > >>>> by a Polish linguist Jan Niecisław Ignacy Baudouin de Courtenay (13 > >>>> March 1845 – 3 November 1929), cf. e.g > >>>> https://en.wikipedia.org/wiki/Jan_Baudouin_de_Courtenay. It is much > >>>> less know that he proposed also the term "grapheme". Let me quote > >>>> Alexander Berg's "English Historical Linguistics vol. I" page 230 from > >>>> Google Books: > >>>> > >>>> Since the introduction of the term grapheme by Baudouin de > >>>> Courtenay in 1901 (Ruszkiewicz 1976:24-37, 1981 [1978], 20-34), > >>>> it has been defined in various ways: > >>>> > >>>> [...] > >>>> > >>>> As can be seen from these quotatioms, the available definitions > >>>> can be divided into two groups, corresponding to two main > >>>> senses, > >>>> and reflecting "conflicting linguistics views of the status of > >>>> writing" (Henderson 1985:142): > >>>> > >>>> 1. a letter or cluster of letters referring to or > >>>> corresponding with a > >>>> single phoneme; > >>>> > >>>> 2. the minimal distinctive unit of a writing system. > >>>> > >>>> For me the first meaning (not mentioned at all in English Wikipedia) is > >>>> the primary, i.e. more useful, meaning, as is has some practical > >>>> applications e.g. for describing Polish hyphenation rules. > >>> > >>> Best regards > >>> > >>> Janusz > >>> > >> -- Paul Schaffner Digital Library Production Service [email protected] | http://www.umich.edu/~pfs/