Re: [OOo-Hebrew] Spell checker on Windows

"Yitzchak Gale" <[email protected]> Wed, 13 Sep 2006 11:20:27 +0300
Newsgroups gmane.comp.openoffice.hebrew
Message-ID <[email protected]>
Thanks to Dan and Nadav for great responses!

I wrote:
>> ...I think most people are quite satisfied with Word's spell
>> checker.

Nadav wrote:
> When you say people are "quite satisfied" with Word's spell-checker, are
> you sure you actually mean "satisfied" and not "used to"? There's a big
> difference, you know.

Could be. This is just the impression I get from many
people I speak to. I myself have been mostly focused
on OOo for a few years now.

> ...How come it forces me to add a yud where
> I was taught not to?

The many variant of ktiv male and ktiv chaser are
definitely the biggest issue, especially for false
negatives.

Of course, part of the problem is that many words
do not have a universally accepted "correct"
spelling in the sense that people would consider
other variants as "wrong". There are different styles,
and it is not the job of a spelling checker to enforce
one of them.

See below for some examples I found with a quick scan
over a few documents.

> How come if I type a random string of letters, it
> accept it as a valid word?

False positives are not such a big problem for me.
However, Hspell does seem to make some really bizarre
suggestions sometimes.

As you mentioned, Hebrew has different density than
English. I'm not sure what the solution is. A start might
be to include by default in the dictionary only words that
either:

- are recognized out of context by a significant proportion
of speakers, or

- are not confused with any such word.

> Anyway, like Dan explained, making Hspell use any other standard but the
> Academia's is quite easy - perhaps as little as a day of work. But, where
> are we going to get that "other standard"? I'm not aware of anybody
> publishing a coherent standard other than the Academia. I'd be happy if
> you, or other readers of this list, point us to such a standard, or can
> volunteer to create a new one. I can help you get started: I am now
> preparing a long document about the Academia's standard and how it applies
> to Hspell; This standard contains dozens of cases and subcases, and like
> I said most other standards you'll see, including Word's, will agree with
> the Academia in 90% of the cases.

Unfortunately, my intuition says that any easy rules-based
approach will be unsatifying. Ideally, an empirical study
needs to be done to find out what usages are common
(after defining what that means), and their frequency.
Then the spell checker should accept any common
usage, and order suggestions by frequency.

(It would be nice if the UI could then somehow group
together variants when it makes suggestions.)

The study would probably need to combine data
from a carefully chosen sample of informants that
represent the population of users of the written
language.

Is any such data available? If not, this obviously
is not too practical without major funding and/or
support of some linguistics department.

Can we interest anyone in this? Is there any
way to roughly approximate this, without doing
all the work?

OK, here are a few quick examles of words that
Hspell rejects but that I feel should be accepted:

Here are two legitimate words that are completely
rejected:

בביה"ס
שמיעתית

Here are some common variants that are rejected:

זכרון
פרוייקט
דוגמא
מוסיקה
חויה
תאור
היתה
מאופינת
חפשית
איתך
איזור
סינטזה
להתיעץ
איטית
לעיתים

-Yitzchak

--
Hebrew OpenOffice Mailing List [email protected]
To unsubscribe see: http://openoffice.org.il/mailman/listinfo/hebrew