Re: Creating a Query for a SHA-1 Hash
"Aaron Walkhouse" <[email protected]> Tue, 30 Dec 2008 03:02:49 -0000
| Newsgroups | gmane.network.gnutella.devel |
|---|---|
| Message-ID | <[email protected]> |
I use them daily and they still work as well as ever. In fact, they are absolutely essential for the vital activity of tracking and fighting the masses of spam, malware and worms that pollute the network daily with literally millions of false filenames. When I search for the bad stuff by keyword I can find a small percentage of it but it's merely a tiny fraction of the results obtained when I browse infected hosts for more of the same infestations and search the resulting hashes to reveal the remaining infected hosts or the original sources. Even if the criminals or worms actively try to block my downloads they can't stop me from getting every last trace from elsewhere on the network and accurately mapping their spread. Like I said, they are so specific to individual files that there is no possibility, read it again, _no_possibility_ that hash queries could do harm to the network in any way. It was more efficient than flood searches in the gnutella 0.4 days and is even more efficient now that we are all on 0.6 and high outdegree. Even during beta testing of the feature when deliberate flooding of hundreds of queries per second at high TTLs was used to test it, the network showed no signs of related distress at all even as it was beginning to struggle to keep up with the phenomenal growth of the day. Since then they have been throttled to a pace slower than keyword searches [8 second delay instead of 5] and nobody has seen fit to touch it again because there have been no measurable effects at all. The benefit to end users has been clear and undiminished even as the whole network has completely transformed into the giant we enjoy today. It isn't broken, so nobody sees a need to fix it. How much bandwidth do those queries actually cost? Have you actually measured it? Is it so small that it cannot be measured at all? No? Prove it. Is the bandwidth used so large that it crashes your own product or knock it off the network? If so, fix your own software first before trying to get others to weaken theirs for your sake alone. If it is in between try to produce numbers instead of speculation or baseless opinion. There is no reasonable justification to eliminate such a valuable and useful tool, especially that which has existed peacefully and usefully on this open public network for years without ill effect while giving real, tangible benefits to end users. Trying to save a few tenths of a percent from the hundreds of megabytes a day our ultrapeers typically spend in the course of normal operations should not come at the cost of one of our most accurate and beneficial tools. Instead of trying to get rid of it take a chance and test it in your own software first for a few months. Don't forget to be fair and use the same limits established here back in the beginning, which is 8 seconds delay between user-directed queries and only one per hour automatically, restricted to files which are stalled for lack of working sources. When your download performance jumps to levels comparable to BearShare and Shareaza you'll probably consider the half day or so of coding time well spent on behalf of your end users. By the way. Don't forget that hashes are never, [I'll say it again, never] added to query routing tables, not even by BearShare. They never have been and probably never will, not even as digests. It was never needed in the first place. That makes them trivially easy to add to any gnutella servent. Once a working DHT is up and running in LimeWire it will likely proxy hash searches to it for the sake of the public as well, since gnutella is still an open network with room for us all and LimeWire LLC still feels that way as a company. That means virtually everybody will be using hashes cooperatively for the sake of the whole. Do you really want to be left behind?