Re: Tibetan Ewts variant proposal (and related questions)
Élie Roux <[email protected]>
| Newsgroups | gmane.ietf.languages |
|---|---|
| Message-ID | <[email protected]> |
> Just a question about version numbers and corpora. If you were to > imagine an evolution in the transcription system from 2.0 to 3.0 > would you imagine that the encoded texts using the transcription > system would all be updated. Perhaps via a script? Well, it's a bit hard for me to answer your question without knowing the kind of languages and transliteration you're dealing with. In Tibetan and Sanskrit, romanizations have been fairly stable and we have never encountered a case where we would have had to update our data. But I think we have a few problems that may be a bit similar to what you describe (encoding of subtleties not matched by usual bcp-47 tags). One example would be with Sanskrit: we could have texts encoded in 3 different ways (it can happen in all scripts and romanizations): * no space at all (like in old manuscripts) * "desandhied" words separated by space (like in modern Devanagari editions) * words separated by space and word compounds separated with a dash (-) when the sandhi allows it this is a crucial indication for corpus analysis, but we infer it with some heuristics and conventions... Actually as I write it I realize private subtags would make sense... Is it the kind of issues you're asking about? Thank you, -- Elie