Re: Bug: extract_msgdata_part_by_coding / message split

Paco <[email protected]>
Newsgroups gmane.comp.mobile.kannel.devel
Message-ID <[email protected]>
Hello Stipe,

Not really, I know that dcs 8 and ucs-2 are not modified, however
sms_split() transforms utf-8 to gsm dropping all non-gsm characters, so when
the message get to the smsc driver it will be utf-8 but only with gsm
characters, so in case there's a smsc accepting latin1 encoding
(iso-8859-1), and we got (for example) a msg of type mt-push submitted via
smsbox http with latin1 encoding, the smsbox convert it to utf-8 (no
characters dropped) the bearerbox call sms_split() on the utf-8 msg,
sms_split transform the utf-8 msg into gsm (dropping non gsm characteres),
and back into utf-8 (without the non gsm characters of course), then the
resulting message (or messages) in utf-8 format (but without non gsm
characters) get to the smsc driver (i'm only using smpp), and if the
alt-charset is set to iso-8859-1 (latin1, same as how we got the original
mt-push msg), it's converted from utf-8 to iso-8859-1, however the resulting
message is not going to be the same as the one we got, because all the non
gsm characters where dropped...

So, my first impression was: "this is kind of pointless" and then Alex
explain that it's unlikely for the smsc to send non-gsm characters to the
mobile (ss7 network),  so I assume that the smsc, even if it gets the
message in iso-8859-1 will convert it to gsm, dropping non gsm characters in
the process, thrust making my first impression pointless :p so everything ok
:)

The only problem is that I think there's a bug in the smsc_smpp msg_to_pdu,
the conversion from utf-8 to gsm is only done when the msg->sms.charset is
set, and I think smsbox is not setting it, so for my example i'm sending
utf-8 to the smsc, please check previews email for details:

1. got http request
2. convert payload to utf-8
3. build mt-push msg
4. send to core
5. check for re-encoding to alt-charset, if not set, convert to gsm
6. send to smsc

piont 5 is failing for me, the code is:

if (msg->sms.charset != NULL && octstr_case_compare(msg->sms.charset,
octstr_imm("GSM-03.38")) != 0)
               charset_utf8_to_gsm(pdu->u.submit_sm.short_message);

msg->sms.charset should be set on poing 3, however it is not, also if msg->
sms.charset is different than utf8 (shouldn't be however) it does the
re-encoding of utf-8 to gsm.

so my solution is to either set msg->sms.charset to UTF-8 in point 3 or
modify the if on point 5 (msg_to_pdu()) or even both!...

can you confirm this bug?

On 5/19/07, Stipe Tolj <[email protected]> wrote:
>
> -----BEGIN PGP SIGNED MESSAGE-----
> Hash: SHA1
>
> Alexander Malysh wrote:
> >
> > yes, you are correct but even if operator accepts utf-8 it doesn't mean
> > that they will send any non gsm03.38 chars on SS7 (I assume GSM network
> > here).
>
> I think Paco means something different here.
>
> @Paco: do you mean lossing characters while we do interal utf-8
> processing,
> sms_split() is called and we still have an alt-charset (ie. DCS=0x08 for
> ucs-2
> in SMPP) to be used?
>
> Stipe
>
> - -------------------------------------------------------------------
> Kölner Landstrasse 419
> 40589 Düsseldorf, NRW, Germany
>
> tolj.org system architecture      Kannel Software Foundation (KSF)
> http://www.tolj.org/              http://www.kannel.org/
>
> mailto:st_{at}_tolj.org           mailto:stolj_{at}_kannel.org
> - -------------------------------------------------------------------
> -----BEGIN PGP SIGNATURE-----
> Version: GnuPG v1.4.7 (MingW32)
> Comment: Using GnuPG with Mozilla - http://enigmail.mozdev.org
>
> iD8DBQFGTwEH9ez0oeKvYs0RAmccAKCRM9OdRAr7U4JLyQc3ncl8kziXigCfVFEj
> M7v1grNJjUDxfVxV+8gr/mo=
> =S5mc
> -----END PGP SIGNATURE-----
>
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.