Re: TEI element for a grapheme?

Paul Schaffner <[email protected]>
Newsgroups gmane.text.tei.general
Message-ID <1475020454.1777962.739059329.7D07142D@webmail.messagingengine.com>
I used to cite this made-up sentence (which nevertheless
is quite possible in my world) to illustrate the well-known
problems with 'ff'.

"Whilst suffering in ffraunce, he thought to himselff of 
the proverb 'Pan darffo treiglo pob tre / Da yw edrych tuag 
adre.'"

'suFFering' and 'himselFF' use the doubled f to 
indicate a voiceless fricative. The former is the 
common spelling, the latter an uncommon one.
But to the everyday user, 'ff' is not a character:
'f' is the character in question, and is the same
'f' as in 'oF' (where it represents a voiced sound).
These are the same everyday users who think
of 't' and 'h' as characters but 'th' as a combination
of characters, not as a single character composed
of graphemes.

'ffraunce' contains an old-fashioned use of 'ff'
to represent an upper-case (capitalized) variant
of "f". Are upper-case and lower-case letters the
same character or not? They are not glyph variants
in the usual sense. And if they *are* different 
characters, then is "ff" simply a glyph variant of "F"?
The usual answer is 'yes.'  /F/ is a character that
may be represented either as grapheme F or as
grapheme cluster ff.

And of course the Welsh 'ff' is usually thought of
(by its users!) as a character in its own right,
quite distinct from 'f'. If we believe these users,
we should regard 'f' as a grapheme, which when
it appears singly represents the character /f/ 
and when it is doubled represents the character 
/ff/.

In actual transcription, it is rare indeed to find
someone who distinguishes character from
glyph in any kind of consistent way: most, I think,
fall back on the Latin alphabet and its extensions
as the default character inventory. (The 'capital'
use of 'ff' is an exception, since I think most
transcribers would capture it as "F" -- not all,
however.)

pfs


On Tue, Sep 27, 2016, at 19:24, Martin Holmes wrote:
> On 2016-09-27 04:03 PM, Piotr Bański wrote:
> > On 27/09/16 21:02, Martin Holmes wrote:
> >> This is a really interesting discussion, but I must confess that this:
> >>
> >> "Graphemes are sequences of one or more encoded characters that
> >> correspond to what users think of as characters."
> >>
> >> is one of the most unhelpful definitions imaginable. Who is a user, and
> >> who can tell what he or she might choose to think of as a "character"?
> >
> > That's your English bias. ;-) In Polish, we have a graph 'dotted z' and
> > a graph 'crossed z' (I'm sure Prof. Bień can quote their established
> > standardized names) which for 'users' (bah, take any plausible
> > definition) represent a single grapheme.
> 
> Yes, this works nicely for a single language group using a single 
> script. But there are so many cases which are not so simple. Consider
> this:
> 
> :-)
> 
> I typed three "characters" to create a combination sign. In my email 
> client, and in my mind, they are distinct characters (as if I'd typed 
> "Yay!"). In your client, the software might substitute U+1F600 GRINNING 
> FACE, which is a single glyph, and which you might perceive as a single 
> character; someone else might receive exactly what I sent, three glyphs, 
> and yet still "think of" it as a "character" because they use emojis all 
> the time.
> 
> It just seems to me that any attempt at a definition which depends on 
> the perception of a non-specific, presumably non-expert user is hardly a 
> definition at all.
> 
> Cheers,
> Martin
> 
> > Similarly with the grapheme
> > realised phonetically as [w], in writing (it can have a tilde across 'l'
> > or over 'l'). Or the sound [t] in Russian, spelled with <m> or a 'little
> > (capital) T' in writing. This is free variation (context-independent,
> > across the population, but I guess you can also see it in the production
> > of the same individual). In English, 'plain z' and 'crossed z' can be
> > found in this sort of free variation, sometimes.
> >
> > You can also have complementary (context-dependent) distribution of
> > graphs representing the same grapheme, think of word-final vs.
> > non-word-final Greek sigma. (In English, in some contexts, word-initial
> > <g> and non-word-initial <dg> could qualify though I'm not sure I would
> > use this example in a 101-course, due to the diachronic variables that
> > have to be taken into consideration here.)
> >
> > Now, two remarks regarding Emmanuelle's original posting:
> > 1. I am still not sure how Emmanuelle wants to define 'grapheme'; I
> > think that a bunch of examples might help to solve this particular
> > issue, without having to delve too deep into philosophy of language
> > and/or semiotics.
> >
> > 2. <fs> is "feature structure" in TEI-speak. It can represent anything,
> > including phonemes. But it definitely is _not_ the "default encoding"
> > for phonemes.
> >
> > Best regards,
> >
> >   Piotr
> >
> >>
> >> Cheers,
> >> Martin
> >>
> >> On 2016-09-27 11:21 AM, Janusz S. Bień wrote:
> >>> On Tue, Sep 27 2016 at 19:17 CEST, [email protected] writes:
> >>>> Hello,
> >>>>
> >>>> On 27/09/2016 18:32, Paul Schaffner wrote:
> >>>>> But I am probably misunderstanding. In my experience, all discussions
> >>>>> of characters, glyphs, graphs, and symbols end in a metaphysical
> >>>>> muddle.
> >>>>
> >>>> Unicode has a definition for all these terms.
> >>>
> >>> Really? What is the Unicode definition of graphs? Is there a Unicode
> >>> definition of symbols?
> >>>
> >>>> In the case of grapheme,
> >>>
> >>> In the case of grapheme the only Unicode definition I know of is that
> >>> from the glossary of Unicode terms (http://www.unicode.org/glossary/),
> >>> which I quote below.
> >>>
> >>>> the key definitions is that of _grapheme cluster_.
> >>>
> >>> Exactly. There is even a definition of extended grapheme cluster. The
> >>> term "grapheme" doesn't occur in the standard alone.
> >>>
> >>>> Please see
> >>>> <http://unicode.org/reports/tr29/#Grapheme_Cluster_Boundaries>,
> >>>> especially the part about Hangul and Devanagari. Reasoning over
> >>>> non-Latin scripts makes that discussion always more scientific and less
> >>>> metaphysical. :)
> >>>
> >>> The title of the quoted Unicode® Standard Annex #29 is "UNICODE TEXT
> >>> SEGMENTATION" and grapheme clusters are just fragment of some
> >>> specific texts. It's unclear for me whether they are clusters of
> >>> (undefined) grahemes or just graphemic clusters in some vague relation
> >>> to the non-Unicode meaning of the term "grapheme".
> >>>
> >>> Let me quote my recent posting to the Unicode mailing list in the thread
> >>>
> >>> http://www.unicode.org/mail-arch/unicode-ml/y2016-m09/0068.html
> >>>
> >>> entitled "graphemes":
> >>>
> >>> On Wed, Sep 21 2016 at  7:09 CEST, [email protected] writes:
> >>>
> >>> [...]
> >>>
> >>>> Let me remind the issues which started the thread:
> >>>>
> >>>>
> >>>> On Sun, Sep 18 2016 at 12:26 CEST, [email protected] writes:
> >>>>> Quote/Cytat - Christoph Päper <[email protected]> (pią, 16
> >>>>> wrz 2016, 23:51:38):
> >>>>>
> >>>>>> Janusz S. Bień <[email protected]>:
> >>>>>>>
> >>>>>>> 1. Graphemes, if I understand correctly, are language dependent, …
> >>>>>>
> >>>>>> That’s true in linguistic terminology – well, at least within the
> >>>>>> more popular schools of thought –, but not in technical (i.e.
> >>>>>> Unicode) jargon.
> >>>>
> >>>> And what is "grapheme" in "technical (i.e. Unicode) jargon"?
> >>>>
> >>>>>
> >>>>> From the Unicode glossary:
> >>>>>
> >>>>> Grapheme. (1) A minimally distinctive unit of writing in the context
> >>>>> of a particular writing system.[...] (2) What a user thinks of as a
> >>>>> character.
> >>>>>
> >>>>> As for (2), cf.
> >>>>>
> >>>>> User-Perceived Character. What everyone thinks of as a character in
> >>>>> their script.
> >>>>>
> >>>>> So we have "a user" versus "everyone...in their script" - is the
> >>>>> difference intentional? Probably not. Anyway the definitions are
> >>>>> language/locale dependent.
> >>>>
> >>>> Does 'Grapheme' (2) make sense with "a (single?) user"?
> >>>>
> >>>> BTW, it is rather well know that the term "phoneme" was proposed first
> >>>> by a Polish linguist Jan Niecisław Ignacy Baudouin de Courtenay (13
> >>>> March 1845 – 3 November 1929), cf. e.g
> >>>> https://en.wikipedia.org/wiki/Jan_Baudouin_de_Courtenay.  It is much
> >>>> less know that he proposed also the term "grapheme". Let me quote
> >>>> Alexander Berg's "English Historical Linguistics vol. I" page 230 from
> >>>> Google Books:
> >>>>
> >>>>        Since the introduction of the term grapheme by Baudouin de
> >>>>        Courtenay in 1901 (Ruszkiewicz 1976:24-37, 1981 [1978], 20-34),
> >>>>        it has been defined in various ways:
> >>>>
> >>>>        [...]
> >>>>
> >>>>        As can be seen from these quotatioms, the available definitions
> >>>>        can be divided into two groups, corresponding to two main
> >>>> senses,
> >>>>        and reflecting "conflicting linguistics views of the status of
> >>>>        writing" (Henderson 1985:142):
> >>>>
> >>>>        1. a letter or cluster of letters referring to or
> >>>> corresponding with a
> >>>>        single phoneme;
> >>>>
> >>>>        2. the minimal distinctive unit of a writing system.
> >>>>
> >>>> For me the first meaning (not mentioned at all in English Wikipedia) is
> >>>> the primary, i.e. more useful, meaning, as is has some practical
> >>>> applications e.g. for describing Polish hyphenation rules.
> >>>
> >>> Best regards
> >>>
> >>> Janusz
> >>>
> >>
-- 
Paul Schaffner  Digital Library Production Service
[email protected] | http://www.umich.edu/~pfs/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.