Re: evolution and hebrew

Ron Artstein <artstein-FrESSTt7Abv7r6psnUbsSmZHpeb/A1Y/@public.gmane.org> Fri, 30 Sep 2005 11:40:22 +0300 (IDT)
Newsgroups gmane.linux.region.israel.ivrix.discuss
Message-ID <[email protected]>
On Fri, 30 Sep 2005, Tzafrir Cohen wrote:

> Again:
> 
> There is exactly one buggy mailer that sends ISO-8859-8 as 
> logical Hebrew: Evolution. Fix that.

Again: as a user, I don't care what mailer my correspondents use 
and if and when that mailer gets its bugs fixed. I want *my* client 
to display the mail in a way that I can read it. If incoming mail 
is readable in one client and unreadable in another, I'll use the 
former.

> There is also the "Hebrew support" in pine that can send Visual 
> Hebrew emails.

Very rare to receive such messages.

I have also received mail from pine in logical order with the 
header ISO-8859-8. Yes, it was the fault of the author, or whoever 
configured their client. So what? I still want to read it.

> I see no point in bending an exstablished standard and 
> over-guessing just because of one buggy (and not very common) 
> client. This will only allow others to follow.

I wouldn't call this bending the standard, but rather dealing with 
noisy input. I also don't accept the pedagogical argument. In my 
opinion, the reasons for adhering to standards in *interpreting* 
text have to do with efficiency, ease of maintenance, time and 
money. The price of strict adherence is usability, given the noisy 
world out there.

> > I bet a simple bigram analysis can determine with a very high 
> > degree of confidence whether a string of Hebrew characters is 
> > in visual or logical order.
> 
> Error-prone over-guessing.

Relying on headers is also error-prone: this is what started the 
whole discussion. 

As for how error-prone an n-gram analysis can be, consider the two 
lines of text that Tzafrir has just sent. These were very short, 
just one word and four words. It just so happened that one had 
the trigram #FM, the other the trigram MF#, with M = medial letter, 
F = final letter, # = non-letter. How hard is it to determine which 
line is in which order? Again, I haven't done a quantitative 
analysis, but determining text direction is exactly the kind of 
problem where an n-gram analysis can be *very* accurate, even for 
short texts. 

Taking this a step further, n-gram analysis is useful even for 
parts of text. Ever seen a search engine result page which mixes 
visual and logical ordering? A paragraph-based n-gram analysis can 
fix that. I'm talking about a level of text processing that has not 
yet been implemented in clients, as far as I am aware, but the 
reasons for this are again efficiency, time and money; if this were 
implemented, usability would go up.

-Ron.
----
Ivrix-discuss list. See http://ivrix.org.il.
To unsubscribe, please send mail to [email protected] with
only the following line in the message body (NOT SUBJECT!): unsubscribe