Re: features/tokens
Bill Yerazunis <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
From: Eugene Crosser <[email protected]> Bill Yerazunis wrote: > is there any way to see the features and/or tokens that are stored i= n the > .cfc-files? >=20 > If you mean the text strings that went into creating > those features and tokens, no, and intentionally not. >=20 > Because those tokens would represent snippets of plaintext > of your email, they represent a significant disclosure risk. Which reminds me about a thought that I've been mulling for a while. As (in the case of email) the corpus is basically the same for every part= icipant (we are talking about 90% of it being almost identical spam messages, rig= ht?), anyone can collect a dictionary of textstrings -> features. As there is o= nly so much textstrings, this will be a fairly complete list. Then, translating anyone's tokens back into text is trivial, even if the hash function was irreversible. Does is impose privacy issues? Am I missing something? No, you're not missing anything. That's the "dictionary attack" that I warn against, and that's why one should not consider one's statistics files as being cryptologically secured. They're secure against using "grep" and "strings". They are not secure to any sort of dedicated attack. - Bill Yerazunis ------------------------------------------------------------------------------ Open Source Business Conference (OSBC), March 24-25, 2009, San Francisco, CA -OSBC tackles the biggest issue in open source: Open Sourcing the Enterprise -Strategies to boost innovation and cut costs with open source participation -Receive a $600 discount off the registration fee with the source code: SFAD http://p.sf.net/sfu/XcvMzF8H