Re: [OOo-Hebrew] Spell checker on Windows
"Yitzchak Gale" <[email protected]> Wed, 13 Sep 2006 11:20:27 +0300
| Newsgroups | gmane.comp.openoffice.hebrew |
|---|---|
| Message-ID | <[email protected]> |
Thanks to Dan and Nadav for great responses! I wrote: >> ...I think most people are quite satisfied with Word's spell >> checker. Nadav wrote: > When you say people are "quite satisfied" with Word's spell-checker, are > you sure you actually mean "satisfied" and not "used to"? There's a big > difference, you know. Could be. This is just the impression I get from many people I speak to. I myself have been mostly focused on OOo for a few years now. > ...How come it forces me to add a yud where > I was taught not to? The many variant of ktiv male and ktiv chaser are definitely the biggest issue, especially for false negatives. Of course, part of the problem is that many words do not have a universally accepted "correct" spelling in the sense that people would consider other variants as "wrong". There are different styles, and it is not the job of a spelling checker to enforce one of them. See below for some examples I found with a quick scan over a few documents. > How come if I type a random string of letters, it > accept it as a valid word? False positives are not such a big problem for me. However, Hspell does seem to make some really bizarre suggestions sometimes. As you mentioned, Hebrew has different density than English. I'm not sure what the solution is. A start might be to include by default in the dictionary only words that either: - are recognized out of context by a significant proportion of speakers, or - are not confused with any such word. > Anyway, like Dan explained, making Hspell use any other standard but the > Academia's is quite easy - perhaps as little as a day of work. But, where > are we going to get that "other standard"? I'm not aware of anybody > publishing a coherent standard other than the Academia. I'd be happy if > you, or other readers of this list, point us to such a standard, or can > volunteer to create a new one. I can help you get started: I am now > preparing a long document about the Academia's standard and how it applies > to Hspell; This standard contains dozens of cases and subcases, and like > I said most other standards you'll see, including Word's, will agree with > the Academia in 90% of the cases. Unfortunately, my intuition says that any easy rules-based approach will be unsatifying. Ideally, an empirical study needs to be done to find out what usages are common (after defining what that means), and their frequency. Then the spell checker should accept any common usage, and order suggestions by frequency. (It would be nice if the UI could then somehow group together variants when it makes suggestions.) The study would probably need to combine data from a carefully chosen sample of informants that represent the population of users of the written language. Is any such data available? If not, this obviously is not too practical without major funding and/or support of some linguistics department. Can we interest anyone in this? Is there any way to roughly approximate this, without doing all the work? OK, here are a few quick examles of words that Hspell rejects but that I feel should be accepted: Here are two legitimate words that are completely rejected: בביה"ס שמיעתית Here are some common variants that are rejected: זכרון פרוייקט דוגמא מוסיקה חויה תאור היתה מאופינת חפשית איתך איזור סינטזה להתיעץ איטית לעיתים -Yitzchak -- Hebrew OpenOffice Mailing List [email protected] To unsubscribe see: http://openoffice.org.il/mailman/listinfo/hebrew