Re: Creating a Query for a SHA-1 Hash

"Aaron Walkhouse" <[email protected]> Mon, 05 Jan 2009 16:13:48 -0000
Newsgroups gmane.network.gnutella.devel
Message-ID <[email protected]>
--- In [email protected], Arne Babenhauserheide <arne_bab@...>
wrote:
> No. 
> 
> You have chosen to be left behind when you chose to stick to a dead
client. 
> 
> A program is dead as soon as noone develops it anymore, but the
environment 
> changes around it. 

Your personal definition of "dead" just doesn't cut it in the real
world.  Not if you consider BearShare already gone.  In fact, the
number of users on BearShare was still rising when the stats node
finally went offline a couple of years after Free Peers shut down and
the horizons on the ultrapeers are stronger right now than when the
stats froze.  

The fact is that BearShare was in the lead at the point development
stopped three and a half years ago and it is still so strong that it
is still considered an equal to LimeWire.  The b25 version with it's
beta testing controls and higher limits is still a top performer.


> At some point Bearshare will cease to interoperate with other
clients, because 
> the other clients will work on improving the protocol, and even though 
> Gnutella is _extremely_ conserative when it comes to protocol
changes, . such 
> changes will become necessary. 

As things stand right now that point will not occur for many years. 
Gnutella is pretty much settled in it's basic structure and the
enhancements in the works now won't radically change the basic
protocol enough to leave BearShare out.

 
> The Bearshare team brought that problem on you when they chose to
keep their 
> sources unfree. Don't blame it on others that they keep improving teh 
> protocol. 

Open-source elitism doesn't really support your argument.  
BearShare is still the fastest and most powerful gnutella node 
precisely because of the concentrated effort of a single 
well-coordinated company using a toolset optimized for Windows.

The simple fact that closed-source [and in this case, well-protected]
code is much more difficult for the spammers and anti-P2P to decipher
and exploit means that BearShare users are still the ones best
equipped to cut through the crap and find the files they really want,
even when somebody is trying very hard to take over searches with
their own decoys and spam.

 
> First: I paid attention to hash queries. I've been advocating to use
a more 
> efficient hash query mechanism for years, and to treat hash queries 
> differently from keyword queries, since they require completely
different 
> optimizations. 

Just because you didn't get what you wanted right away doesn't justify
weakening an existing mechanism by dropping other people's queries.
Think of the users, and not in the abstract sense of imposing progress
upon them against their will for their benefit.  Hash queries are in
use now, they still work, and you have it within your power to allow
the old way to proceed instead of blocking it.  Go right ahead and
develop the next step but don't get rid of hash queries as they are
now just because something better may come along soon.  Even if you
don't reply with hits (which has got to be the easiest feature to add
anyway), please pass those queries along to whatever ultrapeers you
have and don't worry about those theoretical bandwidth storms.


 
> Second: They did real life load evaluations. BearShare did them on QRP, 
> LimeWire on DQ, and each mechanism individually reduced the load for
keyword 
> queries by at least 90%. 
> 
> With out of band replies, hash queries create about equal load as
keyword 
> queries without QRP and DQ. 
> 
> With QRP and DQ each keyword query creates only about 1% of the load
of a Hash 
> query. 

Yet several years later, keyword queries still consume more than
thirty times as much bandwidth.  Go figure.  :P

Don't forget that hash queries typically result in much less hits, in
band or out, and the query hit traffic is 100% efficient as opposed to
keyword hit traffic which consists of everything that may be the one
result desired.

 
> And plase don't calculate the impact on an UP 3.5 hops away, but the
impact of 
> a single search on all nodes in the network added together. That's
what we 
> have to work with to optimize the network. 

Look at the whole network again, and this time as a whole and not as a
snapshot of a theoretical cross-section at the time of a single
search.  How much bandwidth has gone into hash queries, really?  A
theory may be in your head that hash queries make a big mess out there
but the facts speak for themselves.  It still works, not much real
bandwidth is being spent on them because not as many people use them,
and nobody anywhere can point to a single instance where they caused a
problem of ANY kind.

That a hash search can instantly reveal the many false filenames of a
fake, spam or trojan is alone worth it and the enhanced download
performance has never been in doubt.  Preserve that advantage by not
blocking the old version of the feature and letting it run unhindered
until you can properly replace it with DHT.  Even then, try to help
out older servents by maintaining all the features of gnutella 0.6
until it is completely supplanted by the next one.  To start
cherrypicking features and incrementally making your own software less
compatible before anyone else does will only hurt you and your users,
especially when the majority are not going along with you at the same
rate of progress.  In other words go ahead and advance but don't leave
the team and race ahead because that makes you seem less reliable and
compatible to the rest of us.