Re: Strings again
Mark Hahn <[email protected]> Fri, 30 Jul 2004 21:48:34 -0700
| Newsgroups | gmane.comp.lang.prothon.user |
|---|---|
| Organization | Hahn Creative Applications |
| Message-ID | <[email protected]> |
On Fri, 30 Jul 2004 19:41:15 -0700, Roger Binns wrote: > Mark Hahn wrote: >> I'll be honest. I know nothing about locale and it's issues. Can >> anyone get me started? Any urls for what Python does about locales? > > [Apologies for the delay in responding. I had to move house and wait > for the phone company to hook me up again] > > There are several locale related issues. The one most people > are aware of is formatting. For example some locales use > dots as thousand seperators and commas as the decimal point > in numbers. Another is dealing with dates where the US > uniquely uses M/D/Y and various other locales use the > other (more logical :-) orderings. > > Usually that sort of thing is handled by the system library. > On startup you do a setlocale or similarly named call, and > pick up the defaults from the system or the environment. > Calls to functions like scanf, printf and strftime then > all do the right thing. > > The other major locale issue is to do with string formatting. > For example doing a string sort is different based on the locale. > In particular think about accented characters such as the many > variations of the letter 'e' with marks in various directions > around it. If you had a whole bunch of words that all started > with those, the sort ordering is dependent on the locale. > > For example, in some locales è comes before é and in others > it is the other way round. > > Not remembering the locale can also destroy information. > For example the German à character is similar to 'ss'. > If you made the string uppercase and then lower case > again, the locale information would be useful to converting > the upper case characters into lower case again and deciding > whether or not to use it. > > The other area of issue is in the oriental character sets. > Unicode lumped together characters that looked similar > from Chinese, Japanese and Korean, but they aren't the same > character. > > If you cast your mind back 20 years you will remember how > most typewriters printed '1' (the digit one) and 'l' (lower > case L) identically. Imagine some character set committee > decided they were the same thing since they were rendered > identically. Imagine how annoyed you would get now :-) > If you remember the locale of the character then you won't > do harm since you will know which nation it originally came > from. > > The locale stuff gets even more complicated since they can > be combined. Not only do you have to distinguish between > French in France, but also French in Canada. Then you get > nice hybrids such as an American living in Japan. (Which > date format do you think they want and which language do > they want messages in?) > > Then you have to look at the scope of the locale information. > IIRC Win32 lets it be different per thread. Imagine a web > server sending stuff to clients all over the world to see > why this is necessary. UNIX is per process last time I > used (to great frustration when a server process is in a > different locale than the client you are talking to :-) > > Roger Thanks for the lengthy detailed reply. Can you give me any pointers to Python's implementation or support of locale?