Some questions setting up latest CRM114

Leonard Lin <[email protected]>
Newsgroups gmane.mail.spam.crm114
Message-ID <[email protected]>
I have some more free time recently and have been playing around w/  
the latest versions of CRM114 and had a couple questions: I know that  
if you use the reaver_cache you can avoid this, but if the cache is  
missing, does mailreaver.crm filter out X-CRM114 headers in messages  
or do I need to do that if I'm doing my own training?

i.e. I wrote a python script so that I could train complete sets via  
IMAP which is fine with my own clean corpus (in IMAP folders), but  
doesn't strip X-CRM114 headers at the moment: http://jabba.randomfoo.decenturl.com/crm114-trainer.py 
   -- is this something I can run (zapping everything now that I also  
have procmailed items also being move in there?

Also, I'm using the 'hyperspace unique' classifier (it seemed like it  
had a lot less unsures early on - is it more accurate overall?) w/ the  
default pR threshold of 10 (I'm using 20070810-BlameTheSegfault).   
I've trained 276 Ham and 172 Spam (in date order, ToE and on UNSUREs)  
- as of right now it looks like most new mail is still going in the  
unsure folder still - is this normal?  At what point does it usually  
level off?


.l




-------------------------------------------------------------------------
Check out the new SourceForge.net Marketplace.
It's the best place to buy or sell services for
just about anything Open Source.
http://ad.doubleclick.net/clk;164216239;13503038;w?http://sf.net/marketplace
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.