Re: UTF-8 vs ISO-8859-1 (Latin1)

Yusniel Hidalgo Delgado <[email protected]>
Newsgroups gmane.comp.web.mnogosearch.general
Message-ID <[email protected]>
Hi Alexander: 

>There is a small problem in mnoGoSearch: when a remote page clastarims to be 
>UTF-8 but it is in fact ISO-8859-1, then mnoGoSearch does not check 
>well-formedness of the text and tries to insert it into the database 
>as is, so as a result error happen on PostgreSQL level. 

Do you know some patch for this problem?. May be, I will can add this patch in the source code of the mnogosearch and rebuild it again. 

>Can you please send me the URL of the document which made indexer fail? 
>I will check if this is the case. 

Well, I am seeing the log file of postgresql and mnogosearch. I am searching what is the problem, but nothing for now. In my case, mnogosearch is running in local network. This LAN is not available from internet. 

>If you need only English and Spanish, then ISO-8859-1 is enough. 
> 
>However, if you also index documents with scripts other than Latin>, 
>e.g. Cyrillic, Hebrew, Arabic, Asian, then UTF-8 is the best choice. 

Thank you for your answer and recomendations. 




----- Mensaje original ----- 
De: "Alexander Barkov" <[email protected]> 
Para: "Yusniel Hidalgo Delgado" <[email protected]> 
CC: [email protected] 
Enviados: Viernes, 2 de Abril 2010 1:03:26 (GMT-0500) Auto-Detected 
Asunto: Re: UTF-8 vs ISO-8859-1 (Latin1) 

Hi Yusniel, 

Yusniel Hidalgo Delgado wrote: 
> Hi. I have a problem with my mnogoserach 3.3.9 and I need your help. My 
> database encode is UTF-8. Mnogoserach is configurated with the following 
> settings: 
> 
> LocalCharset: UTF-8 
> RemoteCharset: UTF-8 
> VaryLang: ¨es en¨ 
> 
> Everything was fine, but recently, I get the following error when 
> indexer -R is running: 
> 
> ¨PQexecPrepared: ERROR: secuencia de bytes no válida para 
> codificación «*UTF8*»: 0xf7848c93#012HINT 
> 
> I am using PostgreSQL Server 8.3. 

There is a small problem in mnoGoSearch: when a remote page claims to be 
UTF-8 but it is in fact ISO-8859-1, then mnoGoSearch does not check 
well-formedness of the text and tries to insert it into the database 
as is, so as a result error happen on PostgreSQL level. 

Can you please send me the URL of the document which made indexer fail? 
I will check if this is the case. 

> 
> PD: In my sites I have multiples language, English and Spanish for example. 
> 
> What charset is the best recomended?. UTF-8 or ISO-8859-1? Thanks for 
> your time. 
> 
> 

If you need only English and Spanish, then ISO-8859-1 is enough. 

However, if you also index documents with scripts other than Latin, 
e.g. Cyrillic, Hebrew, Arabic, Asian, then UTF-8 is the best choice. 


-- 
--------------------------------------------------------- 
Yusniel Hidalgo Delgado 
Brigade 4501 
University of Informatics Sciences 
http://www.uci.cu/ 
Linux User # 438033 
Blog: http://facultad15.uci.cu/dragoblogs/dragotux/ 
--------------------------------------------------------- 
"Nunca desistas de un sueño, sólo trata de ver las señales 
que te lleven a él, la posibilidad de realizarlo es lo que 
hace que la vida sea interesante. Recuerda que sólo una cosa 
lo vuelve imposible: el miedo a fracasar." 
---------------------------------------------------------
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.