Re: CRM-114 customization for chocolate?

[email protected] Fri, 14 Oct 2011 08:35:54 -0400
Newsgroups gmane.mail.spam.crm114
Message-ID <[email protected]>
Martin Steigerwald <[email protected]> writes:

>> What I would like to have is a mailreaver type system that classifies 
>> messages into several categories, eg. personal messages, work 
>> messages, desired commercial messages, etc. and spam.  In this way I 
>> could handle the desired commercial messages at my own convenience 
>> (rather than having them interrupt my thoughts when they randomly 
>> show up).  Also spam/desired commercial errors would be much less 
>> bothersome.  I can think of a couple of different ways of doing this 
>> (including maintaining a whitelist of desired commercial senders, and 
>> flagging all undesired commercial e-mail as spam, or running a 
>> cascade of standard mailreaver filters; first separating commercial 
>> e-mail from everything else, and then dividing commercial e-mail into 
>> good and bad).  However I presume others have thought about more 
>> generalized classification of messages, and I realize than in the 
>> time I could implement this, someone more directly skilled could do 
>> 10x the product at much higher quality.
>
> as an idea. I thought CRM114 could actually replace manually setting up 
> mail filters that sort mail into folders. I just drag mails in folders and 
> CRM114 learns it.

Multi-class differentiation is a _much_ harder problem.  

We've tried all sorts of ideas on the 20-newsgroups test-set (which is
sort of the "world standard" for this kind of problem, thank you Jason
Rennie) and we get what everyone else gets - about 90% correct
placement (i.e. 85% lower bound, 90% is a publishable result, and
92% is the best reported by anyone anywhere to my knowledge)

Interestingly, it does not seem to matter what the algorithm is; 
you always get somewhere around 90% accuracy.

I'd love to come up with the way to get it up to 99% or even 97% but
I don't know what it is.  My gut tells me that you need something
like that in order to be truly useful.


> But as my last C programming has been ages ago and I would need to work 
> myself into CRM114 code and regexp handling and all the statistical stuff 
> in there - I decided not to even try to implement it by now ;).

It's a lovely idea and if you use libcrm114, then you do it all from C.

The only question is this: is it still worth doing with only a 90% 
accuracy?

(n.b. anybody want to work on a CRM114-to-C translator? Kinda 
like "cfront" was for C++.  

Call it C42C (CRM114-2-C) for now.  I'm thinking that might be fun
to work on too.

   - Bill


------------------------------------------------------------------------------
All the data continuously generated in your IT infrastructure contains a
definitive record of customers, application performance, security
threats, fraudulent activity and more. Splunk takes this data and makes
sense of it. Business sense. IT sense. Common sense.
http://p.sf.net/sfu/splunk-d2d-oct