Re: [OOo-Hebrew] Spell checker on Windows
"Nadav Har'El" <[email protected]> Wed, 13 Sep 2006 12:42:46 +0300
| Newsgroups | gmane.comp.openoffice.hebrew |
|---|---|
| Message-ID | <[email protected]> |
On Wed, Sep 13, 2006, Yitzchak Gale wrote about "Re: [OOo-Hebrew] Spell checker on Windows": > Of course, part of the problem is that many wordsdo not have a universally > accepted "correct"spelling in the sense that people would considerother > variants as "wrong". There are different styles,and it is not the job of a > spelling checker to enforceone of them. I think I don't agree with you on the job of the spell-checker, and I'd like to solicit opinions on this from other people as well. It is true that the spell-checker should not enforce writing *style*. It should allow you to write using any register you want (from literary style to colloquial (spoken) style, and even slang), should allow you to use synonyms (several words with the same meaning) freely, and so on. The issue you are raising - of allowing to write each spoken word in many different ways - is not an issue of style, in my opinion. The problem is that if the spell-checker allows you to write each word in 2 different ways (for example), you can end up with a single document spelling half of the time like this, and half of the time like that. Even if the spell-checker remembers the variants it sees and only allows one per document, the inconsistency problem just moves to the next level: you end up with writing one word with an added yod and a very similar word without a yod (for example). These inconsistencies are not a style choice - they are plain and simple mistakes - something you'd not want to see in your writing, and something that most people would hope that the spell checker solves them. If you think that these problems are rare, think again. Being aware of this situation, I see them even in newspapers all the time. You see an article about "îéñéí" in the header, and "îñéí" a few sentences down. You see an article with îìâåú in one place and the (totally wrong) îéìâåú in another place. Another thing to ponder is why we Israelis have come to accept the situation that in any other language we know, each word (with very few exceptions) has one accepted spelling, and in Hebrew, we have one accepted spelling with niqqud - but when it comes to the way we normally spell - without niqqud - suddenly there is no "accepted" spelling, even though the Academia, and recent Hebrew dictionaries, all have a spelling standard. My guess is that there is just one reason for this: when I (and probably you) were in school, we were taught spelling only with niqqud. We were never taught rules on how to spell without niqqud, and were never tested on these rules. So we grew up thinking there are none. Is this a good situation? A bad situation? I don't know, but the big question is, what should a spell checker do when there are no rules on how to spell? :-) An amusing example on what happens when you don't pay attention to consistent spelling, and don't decide on a single spelling, is here: http://cs.anu.edu.au/~bdm/dilugim/abulafia.html It's a picture of one building with 4 signs with the same name - each spelled differently :-) As a more realistic example, try searching for "ãåâîä ãåâîà" in Google, and see thousands of pages which use the spelling "ãåâîä" in one sentence and "ãåâîà" in the next sentence. Wouldn't you like a spell-checker to save you from making such mistakes? > False positives are not such a big problem for me.However, Hspell does seem > to make some really bizarresuggestions sometimes. The correction algorithm is actually part of OpenOffice, not Hspell. What OpenOffice does is to suggest any word which has a small "edit distance" (difference in letters) from the wrong word. This might work well in English, but in Hebrew it gives dozens of ridiculous possibilities and sometimes you can't find the right ones from all the wrong ones. Hspell's own correction algorithm (in the "hspell" program, not OpenOffice) is much more conservative, and only tries to correct certain types of errors, resulting in much fewer suggestions. This has both a good side and a bad side - see discussion on http://ivrix.org.il/bugzilla/show_bug.cgi?id=16 > As you mentioned, Hebrew has different density thanEnglish. I'm not sure > what the solution is. A start mightbe to include by default in the > dictionary only words thateither: > - are recognized out of context by a significant proportionof speakers, or > - are not confused with any such word. We sort-of follow this guideline when we can: if we come across an obscure word that hardly nobody heard of, and can only be confused with another much more common word - we resist adding this word to the lexicon. > Unfortunately, my intuition says that any easy rules-basedapproach will be > unsatifying. Ideally, an empirical studyneeds to be done to find out what > usages are common(after defining what that means), and their frequency.Then > the spell checker should accept any commonusage, and order suggestions by > frequency. Do we do this "survey" for every single word, or for word classes? For example, if (hypothetically) we see that people write "âðéáä" with yod but "ùøôä" without, do we spell them like this, and loose any connection to the original spelling with niqqud (which was identical for both words)? If we do this survey by word classes, i.e., make one consistent decision for all words like âðáä, ùøôä, etc., then we end up exactly with "rules". This is what the Academia did. And like I said, most of their rules are actually followed by every modern spell-checker and dictionary - only very few of their rules (especially the one regarding kamats katan) is disputed. Almost all the "variants" you gave in your mail: æëøåï, çôùé, ôøåéé÷è, etc., are rarely disputed, and you won't find them in newspapers or dictionaries. Why would we want to allow these variants? > The study would probably need to combine datafrom a carefully chosen sample > of informants thatrepresent the population of users of the writtenlanguage. > Is any such data available? If not, this obviouslyis not too practical > without major funding and/orsupport of some linguistics department. Unfortunately, the data we've been collecting while working on the Hspell project - based on crawling web sites, news sites, and so on - is skewed by existing spell-checkers. Sometimes you don't really see what people wanted to write, but rather what their spell-checker "allowed" them to write. Nevertheless, I've been able to study many of the Academia spelling rules and compare them to the actual popular usage and to existing dictionaries. As I said, I've been writing a document about this, which explains Hspell's spelling standard. I'll try to have a very partial draft of this document ready in a few days, and post a link to it here. -- Nadav Har'El | Wednesday, Sep 13 2006, 20 Elul 5766 [email protected] |----------------------------------------- Phone +972-523-790466, ICQ 13349191 |Someone offered you a cute little quote http://nadav.harel.org.il |for your signature? JUST SAY NO! -- Hebrew OpenOffice Mailing List [email protected] To unsubscribe see: http://openoffice.org.il/mailman/listinfo/hebrew