My Gnutella spam reduction idea revised

"just_courtney_girl" <[email protected]>
Newsgroups gmane.network.gnutella.devel
Message-ID <[email protected]>
I just rejoned under a new address.  In a post
sometime back, I came up with a spam reduction
idea.  I had an idea for eliminating junk hits.
The problem is the On-The_Fly spam generators
that take the Gnutella search string and
create files with the search strings in the
name.

My suggestion at the time was to add or remove
characters from the search string and see what
comes back with that exact string, then search
again, blocking the addresses that returned
hits previously.  That idea was not very well
received, and I didn't really expected it to,
since that would be doubling bandwidth usage.

But as I got to thinking, some things dawned
on  me.  I believe that my idea of tricking the
spam generators into identifying themselves as
such still has merit.  The adaptive behavior
or the spam servers is clearly its downfall.
There just needs to be a better way to
implement it with enough benefits to make it
practical.  So I got to thinking on how we
could do just that.

Now what if I could introduce an anti-spam
idea that not only avoids consuming
bandwidth, but may actually conserve it?
As I was brainstorming, it dawned on me,
why not make this idea into a single step,
and make it an integral part of the searches?
So we can still chop off a letter from the
search terms, but we do the comparison on
the individual machine, not across the
network.  Here is an example:

Say I want to search for "voice training"
materials.  Normally, I may get results
like this:

voice_training_install.exe
voice training porn video.avi
voice training xxx.mpg
voice training - control breath and tone.mp3
[crack] voice training.zip
voice training sales affiliate program.doc
xxx voice training.mpeg

Now lets say I searched for voic trainin:

voic_trainin_install.exe
voic trainin porn video.avi
voic trainin xxx.mpg
voice training - control breath and tone.mp3
[crack] voic trainin.zip
voic trainin sales affiliate program.doc
xxx voic trainin.mpeg

Now, do you tell which is the intended result
and which is the spam?

So, Why not incorporate this into the search?
You search for the entire string, the search
engine searches for shorter strings, but only
returns the results that match the longer
strings?  This eliminates the need to query
twice as I once suggested.  This sniffs out
the adaptive spam generators without impacting
bandwidth in a negative manner.

The above in itself would be a major help and
help reduce the dynamic spam.  The dynamic
spam generators lie to us, so what is wrong
with lying to those clients and saying we
want something else?  Chopping letters off
before searching should still return correct
results while identifying which is spam.
Only on rare occasion will the blocked
results be a typo and not a spam generator.

Now, the above could be expanded upon.  Like
what if the same address gives mismatches for
multiple and unrelated searches?  What if
the same address pops up in searches on
other leaves of the same node?  Like how about
creating a spam information packet format?
A leaf can identify spam addresses and send a
warning packet to its hub/supernode.  It could
even add a priority rating to indicate if
that IP previously returned mismatches for
different searches.  The supernode could then
use the information to block all packets
related to that address for that particular
session.  Cacheing them for weeks at a time
would probably be useless since the spammers
probably rotate addresses.

Just some brainstorming.

                             Courtney
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.