RE: [RFI] octstr_recode
"Oded Arbel" <[email protected]>
| Newsgroups | gmane.comp.mobile.kannel.devel |
|---|---|
| Message-ID | <[email protected]> |
that doesn't solve the problem of using unicode : if that is the case, and you want to use non ISO-8859-1 keywords, then the config file has to be written using UTF-16, which is very ASCII unfriendly (possible will break configuration parsing in Kannel). -- Oded Arbel m-Wise Inc. [email protected] "...Linux, .... that could probably be ported to solar-powered calculators..." -- from PC Week NOS benchmark. > -----Original Message----- > From: Bruno David Rodrigues [mailto:[email protected]] > Sent: Wednesday, March 06, 2002 5:27 PM > To: Oded Arbel; Kannel-devel (E-mail) > Subject: Re: [RFI] octstr_recode > > > with mo-record = true, kannel will TRY to convert to iso-8859-1 only. > > If mo-record is false or kannel can't convert UCS2 to > iso-8859-1 (because > of a unknown char), kannel will behave just like before: text > will have the > ucs2 string, coding (%c) will have 3 and charset (%C) will > have UTF16-BE. > > > ----- Original Message ----- > From: "Oded Arbel" <[email protected]> > To: "Bruno David Rodrigues" <[email protected]>; "Kannel-devel > (E-mail)" <[email protected]> > Sent: Wednesday, March 06, 2002 3:03 PM > Subject: RE: [RFI] octstr_recode > > > Sorry - I didn't understand : with which setting it will > recode UCS2 to > utf-8 , with mo-record set to true or to false ? > > -- > Oded Arbel > m-Wise Inc. > [email protected] > > When you know absolutely nothing about the topic, make your > forecast by > asking a carefully selected probability sample of 300 others who don't > know the answer either. > -- Edgar R. Fiedler > > > > -----Original Message----- > > From: Bruno David Rodrigues [mailto:[email protected]] > > Sent: Wednesday, March 06, 2002 4:42 PM > > To: Kannel-devel (E-mail) > > Subject: Re: [RFI] octstr_recode > > > > > > Ok. Before we think about real unicode, as we usually inject > > iso-8859-1 when > > we are sending a message, if (and only if) you set > mo-recode = true in > > smsbox > > group, smsbox will try to recode the ucs2 to iso-8859-1 when > > it receive a > > message. > > > > As usually, if you don't want it, just don't set mo-recode > > and everything > > will > > be just like before. > > > > The patch was commited to cvs yesterday... > > > > ----- Original Message ----- > > From: "Oded Arbel" <[email protected]> > > To: "Andreas Fink" <[email protected]> > > Cc: <[email protected]> > > Sent: Wednesday, March 06, 2002 9:40 AM > > Subject: RE: [RFI] octstr_recode > > > > > > > >> Well what I'm trying to say is that Kannel should > > support Unicode > > > >receiving but > > > >> keyword matching in smsbox against unicode is quite a headake. > > > > > > > > not so - just convert it to utf-8 and strstr() it. > > > > > > > > > Yes but how do you type a keyword in utf-8 in the config > file if the > > > config file is plain ascii? > > > > Just type it using an utf-8 aware editor. > > > > > the target is to match non ascii > > > keywords. If the keyword would be ASCII then we woulndt have any > > > Unicode. > > > > of course you would - you forget that the ASCII chars in > ISO-8859 and > > single-byte chars in UTF-8 are bitwise identical. if you > > write a keyword > > using "low ISO-8859-1" (characters where only 7 bits are in > > use) and try > > to match it to the same word written in UTF-8, you'll have a > > match. OTOH > > - if you wrote the keyword in mutli-byte UTF-8, one would > > expect that to > > match this keyword one would have to send a SM in unicode. > > note: ASCII (American Standard Code for Information Interchange) is > > defined as 7 bit and often used as 8bit by adding a 0 bit > at the front > > of each char. > > > > i.e. - as long as you use on the "low" range (0-127) characters of > > ISO-8859-1 in the config file, it doesn't matter if its ASCII > > or UTF-8. > > it only matter if you attempt to use weirder stuff (not > sure where in > > ISO-8859-8 are the accented chracters, but I guess that some > > of these do > > fall in the "high" range). when that is the case- we should pick a > > coding and stick to it : I'd suggest to go with UTF-8, otherwise we > > either don't have unicode support, or we have a headache ;-) > > > > -- > > Oded Arbel > > m-Wise Inc. > > [email protected] > > > > "A Meltdown? One of those annoying buzzwords. We prefer to > think of it > > as an > > unrequested fission surplus!" > > -- Montegue Burns, (The Simpsons) > > > > > > > > > > > > > >