Re: Mentor needed
"Eric S. Johansson" <[email protected]> Sat, 16 Apr 2011 12:05:32 -0400
| Newsgroups | gmane.mail.spam.hashcash |
|---|---|
| Message-ID | <[email protected]> |
On 4/16/2011 9:40 AM, Steve Baker wrote: > But isn't the hashcash technique doomed to failure because: Forgive me if I seem little cranky but these questions are are equivalent of global warming deniers although I call them "geeks bad at math" yes, it can cut both ways and it's part of the territory. :-) > a) These days, spam is mostly emitted by botnets where the spammer has > vast compute reserves and doesn't have to pay for it. Hence this system > encourages botnets by giving spammers even more incentive to use them. 1) if you sit down and work through the numbers, you'll find that the masses been generated in the world today can be generated by a very small number of machines. Daily Spam volume / seconds in a day is Spam messages per second 10 billion messages per day works out to roughly 116 messages per second Divide by the number of messages per machine per second and you end up with a very small number of machines. what conclusion would I draw from this? I would say you're asking the wrong question. Instead of encouraging spammers use more machines, I would ask why don't we have more spam today since the number of bot nets machines greatly outnumber the number needed for generating the daily Spam volume. 2) I argue that the limitations we are seeing in spam volumes are economic rather technical in nature. The primary effect of a proof of work system against spammers is reduction of income through increasing opportunity cost. It takes longer to send the same volume of messages which reduces their income per minute. By increasing costs (i.e. sending one message every 60 seconds versus thousands of messages per second), we can significantly drop spammer revenue with respect to time. > Because it is a "white-listing" mechansm, it depends on "normal > people" to sign up to use it in vast numbers because until 100% of > people from whom I'd like to get email are using it (which includes > people I don't know), I can't use it as a means for automatically > discarding spam. Hence there is a chicken& egg issue - nobody wants to > waste CPU time creating tags until nearly everybody else is doing the > same thing - and until everybody is using it, it's not useful. Yeah and if we had started using it when hashcash was first invented, getting started wouldn't be a problem now. The problem with chicken and egg analogy is that if you use it, you never start anything. Look at all the attempts to improve http. We've ignored some pretty good improvements just because of the chicken and egg problem. In fact, htttp was an egg at one time. Just ask Gopher. Oh a minute, that was another egg. twitter? Facebook? Dig? Stumble upon? Ajax? SQL? UNIX? C++? Java? The only difference between those things and us are that they pushed harder on deployment for longer time frames and we explained benefit in mathematical and logical terms. each of them have outlasted the haters. If we had gotten a little more help from the Thunderbird community back in circa 2002, hash cash would have been present and and use invisibly, making it more useful today. for what it's worth, Spam assassin has had hash cash detection built into it for quite a few years and I have had twopenny blue operational since roughly the early 2000s as well. not a lot of deployment because of have problems with a content filter (CRM 114) generating bad results but, I'm living with it and it works. (And this message should have a stamp on it) Chicken and egg problem aside, it's not an all or nothing proposition. network effects play a big role. a proof of work system can be highly effective within a small community and as others become part of that community, it becomes even more effective. You don't need near universal adoption. How does supply Thunderbird? Well, Thunderbird is used by many people. If we are able to incorporate it into Thunderbird, then we very quickly get widespread adoption which means lots of people will gain benefit from it giving the adoption curve a boost. as I mentioned, spam assassin has hashcash support in place and therefore, we have n even wider range of installed receivers. Hashcash is an anonymous reputation tool that gives you a bit more information than a simple white list. the results of hash cash detection can be used to feed a white list or more usefully, reputation list along with other signals. when you have a reputation list, you can hold, accept or reject messages based on that reputation which is far more powerful than a simple white list. A wasting CPU time? I swear to God that is such a non-issue in a world where many people run flash. I have explained the concept of doing a time intensive calculation to help bypass spam filters to non-geeks over the years and there was not a single person who thinks that a proof of work postage stamp is a bad idea. We understand why it works, they understand the benefits, and most important, they are willing to do it without a blink of an eye. What someone says to you "right, I get it. It's why I don't advertise my business more. I can't afford the postage", you know your idea has a reasonable chance of success. I hope you've picked up by now that there is one additional benefit over all the other posters schemes which is that this is completely decentralized. It is strictly a sender-receiver relationship and there's no way to cheat (that we've found so far) > The most effective system of filtering by content works much better > because it requires no help from the sender of the email and it requires > the bad guy to "compute" content that is: > > * human-intelligable... > * ...and gets the message across to the victim effectively > * ...and gets past their content filtering Problem solved often enough to be profitable. you can tell this is true because we have spam. > ...is not just compute-intensive but also human-brain intensive. If > effort is to be spent - then let's use it to improve content filtering. content filtering is human brain intensive for everybody. How do I craft a message that won't be treated as spam by your filter? This question is not asked by just spammers but ordinary people. how many times have you seen the following message in a newsletter? "Please put my e-mail address to your white list so won't be trapped in your spam filters" How many times have people lost a mail list sign-up or notification message to the dumpster? Content filters have reached the limits of what they can do. The math behind them gets you only so close and then you are guaranteed to fail. Content filters give us a system which forces a person to look in their spam trap and look through all the messages they shouldn't have to see to find the one that they missed. How is that the most effective system? If you are willing to say "screw you" to anyone sending you messages and let the messages fall where they may, then yes, content filters are effective. But if you had a content filter with a proof of works stamp, the sender of the message can now do something to guarantee delivery or at least not lose the message the dumpster. If I attach a ninety second stamp to my message should receive a higher score to the content filter reducing the likelihood of being lost. This simple conjunction of filter and stamp is only the start of reputation modeling. It's also the limit of what you can do with a client-side filter and no cooperation from the upstream mail server but, it's a start. As we get past a couple the basics like incorporating proof of work systems into Thunderbird, some of us are looking at ways of limiting the damage of content filters by altering its conclusions based message source and reputation. In twopenny blue, I'm gathering data from message delivery, message contents, and user interaction. With any luck, I'll have the time to put in the reputation engine to analyze and act on this data. Which brings us back to the original request. We need a mentor. We are hoping someone will step up and help us. We have proven it works in our own experiments in the small with low penetration clients like mutt and server-side filters in postfix. Now we would like to experiment with it in the Thunderbird context. --- eric Speech recognition in use. It makes mistakes, I correct some.