Re: SWI-Prolog 7.1.3
Anne Ogborn <[email protected]>
| Newsgroups | gmane.comp.ai.prolog.swi |
|---|---|
| Message-ID | <[email protected]> |
I don't know if I'm the person Richard's referring to, but I'm a) reasonably fluent in Hindi/Urdu b) was Apple's tech lead for the Language Kit Group in the period 95-98, including the time when we made the original Indian Language Kit. c) was *in* the meeting where a lot of the decisions about how Brahmic languages would be handled happened, in the context of WS1. Most of those decisions carried over into Unicode. I was fighting a major defect in the Gurmukhi page at the time (I lost, the defect's still there). I haven't looked at Unicode in quite a while, so I can't address it, but can say that Devanagari, the writing system of Hindi and Sanskrit, is even weirder than you make it out to be. For example, the common family name Sharma is written with the r on the end, because r's starting syllables are denoted by a curlycue mark above the descender line at the end of the syllable. Even worse, adjacent consonants are combined by this rule: let a be the number of consonants in the language (a is 25, off the top of my head). given 3 adjacent consonants try to find a special glyph representing all 3. If you fail, treat the first two as a combination of two via below rule and write the third to the right. given 2 adjacent consonants look up in an a by a table hwich of these subrules to use: a) use a special glyph b) write the second letter's glyph under the first c) write the left half of the first letter's glyph and the right half of the second letter's glyph in the right half d) write the second glyph in a distorted form inside the counter of the first glyph There are special rules if r is involved that are too complicated to go into here. The word thaththera (brassworker) is taller than it is wide using most hindi fonts. There's some politics of colonialism in designating Devanagari 'alphabetic'. At one point European scripts were considered superior by academics in the west because they were 'alphabetic' and supposedly more regular. Indian academics countered by declaring Devanagari alphabetic. In actual fact, Devanagari is far more regular than the Romaji alphabet (the one used for English). Children in N. India learn the alphabet from a square arrangement of characters that resembles a periodic table more than a list. But, while the script is made up of individual glyphs, they combine in a far more complex way than Romaji does. You clearly haven't gotten to Nastaliq yet. Nastaliq is written right to left (it looks like 'really runny' arabic) on a diagonal slant down to the left, without spaces. The 'slant' stops and returns to the baseline at the end of each phrase, which means that typesetting Nastaliq properly requires doing part-of-speech parsing of the content. Nor is this some obscure script only of interest to academics. Nastaliq is the script of Urdu, Farsi, Pashtun, and often used to write Punjabi, so it's used daily by somewhere around 400 million people. It's nasty enough to set competently that newspapers in these languages are usually photoset from hand calligraphy. Example: https://sites.google.com/site/alijsh/shekste_nastaliq.jpg If pure, Farsi-ized Urdu wasn't bad enough, real Urdu has lots of Sanskrit loanwords in it. Those have to be laid out by first figuring out how they'd be laid out in Devanagari, then transliterating into Nastaliq. If you're in the border area with China, ... lets not go there. Oh, if you intend to compose any poetry in Urdu, remember that it not only has to express lofty thoughts and sound nice on the tongue, but has to look nice on the page. It should be at once bawdy and express your love for God, have meter and rhyme, and lend itself to being bent into a calligraphic artwork. There's good reason why poets and mathematicians are often the same people in southwest asia! EUROPEAN PROLOG IMPLEMENTORS ARE THE WRONG PEOPLE TO DESIGN TEXT HANDLING OPERATIONS BY THEMSELVES. I'd certainly concur that competent non Romaji text handling needs to involve people who are specialists. At Apple's Text and International we hired one engineer who knew each major script script system. Indeed, I was laid off when Apple had a major 4400 person layoff. Little language kit group had to lay off one of it's two engineers. We'd just finished the Indian Language Kit, and were starting a rev of the Korean Language Kit. The two engineers were me and Red John Park, from Seoul... An *adequate* design team for string operations in the 21st century will include at least one person intimately familiar with CJK text processing, at least one person intimately familiar with Indic text processing, at least one person intimately familiar with Arabic text processing, none of these people needing deep understanding of Prolog, *and* at least one Prolog implementor. 8cD Actually, there's lots of gotchas even then. For example, if you're going to mix RDF support and Brahmic atoms, you have to be aware that colon means glottal stop (a letter) in Hindi, Marathi, etc. Or bar is the usual end of sentence character. The number digits are different. And people divide their numbers in different places - 123456789 is read as 12 crore 34 lakh 56 thousand 7 hundred 89. Aaaah! Now, just to keep you up at night... If you want to support Medu Neter you're going to have to do your layout on a paragraph level. The topic sentence of the paragraph radiates out in both directions horizontally and subsidiary sentences go vertically, or so they tell me. Fortunately it's a dead language. Mongolian, on the other hand.... -------------- next part -------------- HTML attachment scrubbed and removed