Re: FreeTDS and UTF-8 support

Steve Langasek <[email protected]>
Newsgroups gmane.comp.db.tds.freetds
Message-ID <[email protected]>
On Thu, Jul 24, 2008 at 09:16:38AM +0200, Sebastien FLAESCH wrote:
> >> We agree; we're just tripping over terminology.  By "multibyte characters"
> >> I meant pre-Unicode encodings, such as input to mbtowc(3).  I'm happy to
> >> support UTF-8.  

> > FWIW, UTF-8 is a multibyte character set, and is understood by mbtowc() from
> > glibc.

> That's true.

> I think FreeTDS could support all single-byte or multi-byte character sets
> if:
> 
> 1) The application (main) calls setlocale() to define the current charset
> for libc functions like mbtowc, wctomb, mblen (that would include UTF-8
> support).

> 2) Any (char*) string manipulation in FreeTDS lib uses mblen() to know the
> size in bytes of the characters (= no byte-based parsers/lexers).

> 3) Before sending on TDS protocol, convert from current charset to
> WideChar (UTF-16/UCS-2) with mbstowcs() and when receiving TDS responses,
> convert from WideChar to current charset with wcstombs().

So mbtowc() and wctomb() only convert between widechar and multibyte
representations *of the same charset*.  That makes them insufficient for
converting to "the current charset", because you always want your WC
representation to be UCS-2 but you want your mb representation to be in the
current locale, which may not be a Unicode locale.

Hence iconv().

> ...

>  From my understanding, a C application on UNIX can be written in:

> A- Pure ASCII (LC_ALL=POSIX)
> B- A specific single-byte charset (LC_ALL=fr_FR.iso88591)
> C- A specific multi-byte charset (LC_ALL=ja_JP.eucjp or en_US.utf8)
> D- Wide Char (UTF-16/UCS-2), using the wchar_t type.

I don't think you can be very effective writing a C application on Unix
using a widechar representation internally.  All string manipulation
functions in the standard library, and in other common libraries built on
top of them, are implemented using char* as the argument type.

> Maybe FreeTDS could work in wchar_t internally and do charset conversion
> at the ODBC API level? I think this is how MS ODBC drivers work... no?

This is how MS ODBC drivers work, because MS is wchar_t-centric.  Unix is
not.  Using wchar_t internally and having to either convert to char for
every string manipulation, or reimplement your own wchar-libc, == pain.

> I mean (casts are just there for understanding):

> SQLPrepare(hp, (char *) "my SQL string in current SBCD or MBCS charset",
> SQL_NTS);

> would be mapped to:

> SQLPrepareW(hp, (wchar_t *) "my SQL string in UCS-2", SQL_NTS);

> and just use the mbstowcs() standard C lib function to convert current
> charset to wide char...

> WARNING: wchar_t on UNIX is 4 bytes while UCS-2 and wchar_t on Windonws
> is 2 bytes.

Right, that means you can't even use mbtowc() to do the conversions, because
your wchar_t internal representation no longer matches the output format of
*either* of your APIs, so... what's the point in using it as the internal
representation?

-- 
Steve Langasek                   Give me a lever long enough and a Free OS
Debian Developer                   to set it on, and I can move the world.
Ubuntu Developer                                    http://www.debian.org/
[email protected]                                     [email protected]
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.