Re: ... convert_unicode.c ...
David Relson <[email protected]>
| Newsgroups | gmane.mail.bogofilter.devel |
|---|---|
| Organization | Osage Software Systems, Inc. |
| Message-ID | <[email protected]> |
On Mon, 20 Jun 2005 13:35:19 +0200 Matthias Andree wrote: > David Relson <[email protected]> writes: > > > The question of the moment is what to do when iconv_open() fails. As > > you suggest we could just ignore the message. That seems like a bad > > idea as one could just add a dummy mime body section with a bogus > > charset and bogofilter would be disabled. Not good! > > Right you are - the question is what will mailers present to the user > with strange character sets? We should probably log these for a while to > obtain relevant information. Attached is a list of 76 charsets in this month's spam. I don't have time to see what iconv_open( "from_charset", "UTF-8" ) thinks of them -- have to head to work. > > > It would be better to turn off translation and simply parse whatever > > text is present. Translation will resume at the next > > "Content-Type: ... charset=" directive. True, some untranslated text > > would be passed through, but the impact would probably be minor. > > I'm a bit concerned about storing non-UTF-8 tokens in a database that > claims UTF-8 format. This is a can of worms we can avoid - like reading > the database back to show it to the user (we don't do that yet) fails > with EILSEQ or similar. The solutions I can think of are all hacks ignore tokens for invalid charsets (when scoring and registering) don't register tokens from messages with invalid charsets The present policy of accepting tokens from invalid charsets is relatively benign. Of course, when the bogus charset name is seen (after registration as spam), it has a very high spam score -- effectively a red flag! David _______________________________________________ Bogofilter-dev mailing list [email protected] http://www.bogofilter.org/mailman/listinfo/bogofilter-dev
charset.2005-06-Spam.txt
(text/plain, 1.4 KB)
charset= charset=big5 charset=%charset charset=cp-1252 charset=%custom_charset charset=default charset=euc charset=euc-kr charset=euc-kr[Çѱ¹¾î] charset=gb2312 charset=iso-0145-4 charset=iso-0151-6 charset=iso-0237-6 charset=iso-0408-1 charset=iso-0501-5 charset=iso-0950-0 charset=iso-1278-3 charset=iso-1326-6 charset=iso-1656-6 charset=iso-1846-3 charset=iso-2005-7 charset=iso-2022-jp charset=iso-2093-1 charset=iso-2286-9 charset=iso-2332-5 charset=iso-2377-5 charset=iso-2650-8 charset=iso-3227-5 charset=iso-3409-2 charset=iso-3499-2 charset=iso-3596-0 charset=iso-3700-4 charset=iso-4186-0 charset=iso-4334-5 charset=iso-4452-8 charset=iso-4667-5 charset=iso-4951-3 charset=iso-5111-2 charset=iso-5120-6 charset=iso-5bf1-d charset=iso-6114-1 charset=iso-6243-9 charset=iso-6437-4 charset=iso-6619-1 charset=iso-6775-9 charset=iso-7048-6 charset=iso-708f-6 charset=iso-7183-2 charset=iso-72cd-1 charset=iso-7521-9 charset=iso-7841-5 charset=iso-7896-9 charset=iso-8006-4 charset=iso-8064-1 charset=iso-8365-7 charset=iso-8468-9 charset=iso-8609-7 charset=iso-8723-7 charset=iso-8859-1 charset=iso-8859-1format=flowed charset=iso-8859-2 charset=iso-9017-7 charset=iso-9034-1 charset=iso-9234-0 charset=iso-9477-7 charset=iso-d198-1 charset=iso-e18d-c charset=ks_c_5601-1987 charset=.*$n*`$n.ddone charset=us-asci charset=-us-ascii charset=us-ascii charset=utf-8 charset=windows-1251 charset=windows-1252 charset=windows-1255