Crawler problems
Steve Holdoway <[email protected]> Tue, 22 Nov 2011 09:45:02 +1300
| Newsgroups | gmane.org.user-groups.linux.new-zealand.general |
|---|---|
| Organization | Green Gecko Global Ltd. |
| Message-ID | <1321908302.17618.330.camel@steve-desktop> |
Came to work yesterday, and had 3 (geographically) separate web servers on their knees, looking like they were being DDoSed. Web services unusable, databases screaming, resources depleted. 2 VPSes and a dedicated server. It transpires that all of them are being heavily spidered by Baiduspider, Bingbot and Googlebot. After a load of analysis this is what I've done: 1. Baiduspider - drop at firewall. These are English speaking sites, so we don't need this. 2. Bingbot - the same. 3. Google - dropped the 5 most prolific crawlers. I did want to keep bing, but tell it to behave itself, but it seems to just ignore robots.txt and do what it wants to anyway. Google was an afterthought - I'd hoped I could get away with that but it was still slowing the site down - so I've not analyzed so closely. Has anyone else noticed this increased aggressiveness in crawling policy? Any ideas on how to control it?? Cheers, Steve -- Steve Holdoway BSc(Hons) MNZCS <[email protected]> http://www.greengecko.co.nz MSN: [email protected] Skype: sholdowa _______________________________________________ NZLUG mailing list [email protected] http://www.linux.net.nz/cgi-bin/mailman/listinfo/nzlug