Re: Tibetan Ewts variant proposal (and related questions)

Élie Roux <[email protected]>
Newsgroups gmane.ietf.languages
Message-ID <[email protected]>
> Just a question about version numbers and corpora. If you were to 
> imagine an evolution in the transcription system from 2.0 to 3.0
> would you imagine that the encoded texts using the transcription
> system would all be updated. Perhaps via a script?

Well, it's a bit hard for me to answer your question without knowing the
kind of languages and transliteration you're dealing with. In Tibetan
and Sanskrit, romanizations have been fairly stable and we have never
encountered a case where we would have had to update our data.

But I think we have a few problems that may be a bit similar to what you
describe (encoding of subtleties not matched by usual bcp-47 tags). One
example would be with Sanskrit: we could have texts encoded in 3
different ways (it can happen in all scripts and romanizations):
   * no space at all (like in old manuscripts)
   * "desandhied" words separated by space (like in modern Devanagari
editions)
   * words separated by space and word compounds separated with a dash
(-) when the sandhi allows it

this is a crucial indication for corpus analysis, but we infer it with
some heuristics and conventions... Actually as I write it I realize
private subtags would make sense...

Is it the kind of issues you're asking about?

Thank you,
-- 
Elie
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.