Re: UTF-8 vs ISO-8859-1 (Latin1)
Yusniel Hidalgo Delgado <[email protected]>
| Newsgroups | gmane.comp.web.mnogosearch.general |
|---|---|
| Message-ID | <[email protected]> |
Hi Alexander: >There is a small problem in mnoGoSearch: when a remote page clastarims to be >UTF-8 but it is in fact ISO-8859-1, then mnoGoSearch does not check >well-formedness of the text and tries to insert it into the database >as is, so as a result error happen on PostgreSQL level. Do you know some patch for this problem?. May be, I will can add this patch in the source code of the mnogosearch and rebuild it again. >Can you please send me the URL of the document which made indexer fail? >I will check if this is the case. Well, I am seeing the log file of postgresql and mnogosearch. I am searching what is the problem, but nothing for now. In my case, mnogosearch is running in local network. This LAN is not available from internet. >If you need only English and Spanish, then ISO-8859-1 is enough. > >However, if you also index documents with scripts other than Latin>, >e.g. Cyrillic, Hebrew, Arabic, Asian, then UTF-8 is the best choice. Thank you for your answer and recomendations. ----- Mensaje original ----- De: "Alexander Barkov" <[email protected]> Para: "Yusniel Hidalgo Delgado" <[email protected]> CC: [email protected] Enviados: Viernes, 2 de Abril 2010 1:03:26 (GMT-0500) Auto-Detected Asunto: Re: UTF-8 vs ISO-8859-1 (Latin1) Hi Yusniel, Yusniel Hidalgo Delgado wrote: > Hi. I have a problem with my mnogoserach 3.3.9 and I need your help. My > database encode is UTF-8. Mnogoserach is configurated with the following > settings: > > LocalCharset: UTF-8 > RemoteCharset: UTF-8 > VaryLang: ¨es en¨ > > Everything was fine, but recently, I get the following error when > indexer -R is running: > > ¨PQexecPrepared: ERROR: secuencia de bytes no válida para > codificación «*UTF8*»: 0xf7848c93#012HINT > > I am using PostgreSQL Server 8.3. There is a small problem in mnoGoSearch: when a remote page claims to be UTF-8 but it is in fact ISO-8859-1, then mnoGoSearch does not check well-formedness of the text and tries to insert it into the database as is, so as a result error happen on PostgreSQL level. Can you please send me the URL of the document which made indexer fail? I will check if this is the case. > > PD: In my sites I have multiples language, English and Spanish for example. > > What charset is the best recomended?. UTF-8 or ISO-8859-1? Thanks for > your time. > > If you need only English and Spanish, then ISO-8859-1 is enough. However, if you also index documents with scripts other than Latin, e.g. Cyrillic, Hebrew, Arabic, Asian, then UTF-8 is the best choice. -- --------------------------------------------------------- Yusniel Hidalgo Delgado Brigade 4501 University of Informatics Sciences http://www.uci.cu/ Linux User # 438033 Blog: http://facultad15.uci.cu/dragoblogs/dragotux/ --------------------------------------------------------- "Nunca desistas de un sueño, sólo trata de ver las señales que te lleven a él, la posibilidad de realizarlo es lo que hace que la vida sea interesante. Recuerda que sólo una cosa lo vuelve imposible: el miedo a fracasar." ---------------------------------------------------------