Re: Creating a Query for a SHA-1 Hash

"Aaron Walkhouse" <[email protected]> Tue, 30 Dec 2008 03:02:49 -0000
Newsgroups gmane.network.gnutella.devel
Message-ID <[email protected]>
I use them daily and they still work as well as ever. In fact, they
are absolutely essential for the vital activity of tracking and
fighting the masses of spam, malware and worms that pollute the
network daily with literally millions of false filenames.  When I
search for the bad stuff by keyword I can find a small percentage of
it but it's merely a tiny fraction of the results obtained when I
browse infected hosts for more of the same infestations and search the
resulting hashes to reveal the remaining infected hosts or the
original sources.  Even if the criminals or worms actively try to
block my downloads they can't stop me from getting every last trace
from elsewhere on the network and accurately mapping their spread.

Like I said, they are so specific to individual files that there is no
possibility, read it again, _no_possibility_ that hash queries could
do harm to the network in any way.  It was more efficient than flood
searches in the gnutella 0.4 days and is even more efficient now that
we are all on 0.6 and high outdegree.  

Even during beta testing of the feature when deliberate flooding of
hundreds of queries per second at high TTLs was used to test it, the
network showed no signs of related distress at all even as it was
beginning to struggle to keep up with the phenomenal growth of the
day.  Since then they have been throttled to a pace slower than
keyword searches [8 second delay instead of 5] and nobody has seen fit
to touch it again because there have been no measurable effects at
all. The benefit to end users has been clear and undiminished even as
the whole network has completely transformed into the giant we enjoy
today.  It isn't broken, so nobody sees a need to fix it.

How much bandwidth do those queries actually cost?  
Have you actually measured it? 

Is it so small that it cannot be measured at all?  No?  Prove it.  

Is the bandwidth used so large that it crashes your own product or
knock it off the network?  If so, fix your own software first before
trying to get others to weaken theirs for your sake alone.  

If it is in between try to produce numbers instead of speculation or
baseless opinion.

There is no reasonable justification to eliminate such a valuable and
useful tool, especially that which has existed peacefully and usefully
on this open public network for years without ill effect while giving
real, tangible benefits to end users. Trying to save a few tenths of a
percent from the hundreds of megabytes a day our ultrapeers typically
spend in the course of normal operations should not come at the cost
of one of our most accurate and beneficial tools.

Instead of trying to get rid of it take a chance and test it in your
own software first for a few months.  Don't forget to be fair and use
the same limits established here back in the beginning, which is 8
seconds delay between user-directed queries and only one per hour
automatically, restricted to files which are stalled for lack of
working sources. When your download performance jumps to levels
comparable to BearShare and Shareaza you'll probably consider the half
day or so of coding time well spent on behalf of your end users.

By the way.  Don't forget that hashes are never, [I'll say it again,
never] added to query routing tables, not even by BearShare.  They
never have been and probably never will, not even as digests.  It was
never needed in the first place.  That makes them trivially easy to
add to any gnutella servent.  

Once a working DHT is up and running in LimeWire it will likely proxy
hash searches to it for the sake of the public as well, since gnutella
is still an open network with room for us all and LimeWire LLC still
feels that way as a company.  That means virtually everybody will be
using hashes cooperatively for the sake of the whole. Do you really
want to be left behind?