Re: Drop Dynamic Querying - It only benefits Spammers these days

Bill Pringlemeir <[email protected]>
Newsgroups gmane.network.gnutella.devel
Message-ID <[email protected]>
>    Posted by: "Philippe Verdy" [email protected] verdy_p
>    Date: Mon Jan 29, 2007 8:52 pm ((PST))

>> http://credence-p2p.org/Credence_nsdi.pdf

> How do you know who has voted negatively against your files, so that you give

> them negative value for their votes against you?  Couldn't the identity of 
> voters be abused as well (fake or random identities?)
> Is that the fact that each identity used must first reach a sufficient 
> threshold of good votes for their votes to have importance?

> I see that this paper attempts to find methods under which nodes are 
> classified by behavior, but i am still doubtful. Is it really complex enough 
> to defeat attacks by "smart" automated spammers constantly computing their 
> threshold to remain in the statistics of "good" nodes?

Spammer voting against your files is a good thing.  It will make them
anti-correlated to you.  There are really three parties.  Legitimate content
providers, spammers and downloaders.  The Credence system *requires* that the
downloader has voted on several files.  The downloader will look at how he has
ranked files versus the spammer and legitimate content provider (of course who
is who is unknown).  In order for the spammer to look better they must
correlate better with all of the downloaders votes.

With host browsing, it might be possible for a spammer to query the downloader
for all of his votes on the existing files and get a higher correlation than
the legitimate source.  I don't know the Credence protocol well enough to know
a defense against this.  I think that the downloader is suppose to drop random
results as one defence; but there are always multiple spam machines.

>     Posted by: "Philippe Verdy" [email protected] verdy_p
>     Date: Mon Jan 29, 2007 3:44 pm ((PST))

> Remember that spammers can also vote massively AGAINST all their non-spamming
> competitors! And do you think that spammers would vote "honestly" ???

I think everyone would know that.  However, I don't see anything would be
better than the correlative method and adjust to new content automatically.  It
is partly based on Bayesian statitistics like many spam e-mail filters.

Spammers may try to invent content that is partially relavent, but contains
some trojan or ad inserted in a file.  This would be equivalent to the emails
that try to alter the bayesian filters by spewing many good words with spam
words.

Interestingly, if all spammers did what you said (mark other spammers as bad),
then I think that no spam would work if there are more than two spammers.  

The spammers must attack the downloaders file ranking to be successful afaik. 
This is the major weakness in this protocol; the requirement that users must
rank spam and non-spam (and not confuse spam with good content).  

Ie, I downloaded "Sick_of_you.mp3" and I don't like Gwar, so I vote against it.
 This would be a legitimate share, but the user could be confuse mis-labelled
content (or false search responses) versus content they like.  When exactly
does a file become so crappy it is garbage?  A few skips in an avi in *your*
player?  I think this is a great weakness of this protocol.  I might
under-estimate users though?

I believe this is why GTKG is using a developer only spam.txt list of spam
hashes and names.

At least it is an interesting problem.  As I don't know that much about
Gnutella, I was wonder if GUESS is a solution to the rare file problem.  It
seems that LimeWire has had this for a while, but not any other clients I know
of.  I think that this was one critisism of dynamic querying that started this
thread.  Is GUESS bad for some reason?

fwiw,
Bill Pringlemeir.



 
____________________________________________________________________________________
Cheap talk?
Check out Yahoo! Messenger's low PC-to-Phone call rates.
http://voice.yahoo.com
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.