Re: CRM-114 customization for chocolate?
[email protected] Fri, 14 Oct 2011 08:35:54 -0400
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
Martin Steigerwald <[email protected]> writes: >> What I would like to have is a mailreaver type system that classifies >> messages into several categories, eg. personal messages, work >> messages, desired commercial messages, etc. and spam. In this way I >> could handle the desired commercial messages at my own convenience >> (rather than having them interrupt my thoughts when they randomly >> show up). Also spam/desired commercial errors would be much less >> bothersome. I can think of a couple of different ways of doing this >> (including maintaining a whitelist of desired commercial senders, and >> flagging all undesired commercial e-mail as spam, or running a >> cascade of standard mailreaver filters; first separating commercial >> e-mail from everything else, and then dividing commercial e-mail into >> good and bad). However I presume others have thought about more >> generalized classification of messages, and I realize than in the >> time I could implement this, someone more directly skilled could do >> 10x the product at much higher quality. > > as an idea. I thought CRM114 could actually replace manually setting up > mail filters that sort mail into folders. I just drag mails in folders and > CRM114 learns it. Multi-class differentiation is a _much_ harder problem. We've tried all sorts of ideas on the 20-newsgroups test-set (which is sort of the "world standard" for this kind of problem, thank you Jason Rennie) and we get what everyone else gets - about 90% correct placement (i.e. 85% lower bound, 90% is a publishable result, and 92% is the best reported by anyone anywhere to my knowledge) Interestingly, it does not seem to matter what the algorithm is; you always get somewhere around 90% accuracy. I'd love to come up with the way to get it up to 99% or even 97% but I don't know what it is. My gut tells me that you need something like that in order to be truly useful. > But as my last C programming has been ages ago and I would need to work > myself into CRM114 code and regexp handling and all the statistical stuff > in there - I decided not to even try to implement it by now ;). It's a lovely idea and if you use libcrm114, then you do it all from C. The only question is this: is it still worth doing with only a 90% accuracy? (n.b. anybody want to work on a CRM114-to-C translator? Kinda like "cfront" was for C++. Call it C42C (CRM114-2-C) for now. I'm thinking that might be fun to work on too. - Bill ------------------------------------------------------------------------------ All the data continuously generated in your IT infrastructure contains a definitive record of customers, application performance, security threats, fraudulent activity and more. Splunk takes this data and makes sense of it. Business sense. IT sense. Common sense. http://p.sf.net/sfu/splunk-d2d-oct