Re: Creating a Query for a SHA-1 Hash

"Aaron Walkhouse" <[email protected]> Sat, 03 Jan 2009 01:02:26 -0000
Newsgroups gmane.network.gnutella.devel
Message-ID <[email protected]>
--- In [email protected], Arne Babenhauserheide <arne_bab@...>
wrote:
>
> Am Dienstag 30 Dezember 2008 04:02:49 schrieb Aaron Walkhouse:
> > By the way.  Don't forget that hashes are never, [I'll say it again,
> > never] added to query routing tables, not even by BearShare.  
> > And this is exactly why they shouldn't be used. 

The fact remains that they are and still will be for years to come.

> A Hash query uses a pure flooding mechanism tempered by neither
QRP[1] nor 
> DQ[2], so almost any single Hash query consumes about 100 times the
bandwidth 
> of a regular keyword query. 
> 
> For your own bandwidth that's mostly irrelevant, since the other
nodes who 
> relay your request on the network pay most of the cost. 
> 
> So if everyone used only Hash queries, the bandwidth consumption at the 
> Ultrapeers would rise by two orders of magnitude and would likely
completely 
> bog down the network. 

That's a guess.  Nothing more.  If there had been a real problem with
the millions of BearShare users' SHA1 searches over the past few
years, don't you think somebody would have noticed?

Also, don't forget to discount extreme hypotheticals like the above.
Hash queries are still quite rare and cause no bandwidth problems in
the real world.  Even if they do become more popular in the future,
proxying by UPs off a DHT will eliminate the theoretical problem long
before it could become a real one and those old SHA1 searches will
still be a productive and valuable tool for all, 

 
> Hash queries can however be served extremely efficiently by a DHT
(which can't 
> do keyword queries very well, by the way). 
> 
> And chances are that LimeWire will tell other developers to
implement the DHT 
> and send Hash queries only over the DHT, since they want the network to 
> continue to evolve and not be bogged down by dead clients. 

BearShare isn't dead.  It works so well that people will be using it
for many years to come.  Don't bother trying to denigrate it.  :p

 
> [1]: Since Hashes aren't added to QRTs. The QRP reduces the neded
bandwidth by 
> more than 90%. 

So it does not save any bandwidth in this case and never did.


> [2]: Since Hashes return only very few results and DQ stops the
queries after 
> reaching at least a certain number of results. So Dynamic Querying
doesn't 
> reduce the amount of requests sent by Hash queries. DQ also saves
more than 
> 90% Bandwidth. 

Or rather, it does not save any bandwidth in this case and never did.

> These numbers come from Bearshare (QRT) and LimeWire (DQ) from the
time when 
> they introduced QRP and DQ. 

So, from an absence of evidence and a couple of old guesses made by
others who were actually working on other things entirely you presume
a 100 to one potential savings on a type of query you admittedly have
paid little to no attention to over the past how many years?  ;]

There's a reason hash queries were left to run in the old way.  They
didn't actually waste bandwidth as you have just assumed here.  Run
the math again and this time don't forget to run it in a high
outdegree model instead of the original flat 0.4 model with a TTL of
7.  Show the impact on the UPs 3.5 hops away from the source.