Crawler problems

Steve Holdoway <[email protected]> Tue, 22 Nov 2011 09:45:02 +1300
Newsgroups gmane.org.user-groups.linux.new-zealand.general
Organization Green Gecko Global Ltd.
Message-ID <1321908302.17618.330.camel@steve-desktop>
Came to work yesterday, and had 3 (geographically) separate web servers
on their knees, looking like they were being DDoSed. Web services
unusable, databases screaming, resources depleted. 2 VPSes and a
dedicated server.

It transpires that all of them are being heavily spidered by
Baiduspider, Bingbot and Googlebot.

After a load of analysis this is what I've done:

1. Baiduspider - drop at firewall. These are English speaking sites, so
we don't need this.
2. Bingbot - the same.
3. Google - dropped the 5 most prolific crawlers.

I did want to keep bing, but tell it to behave itself, but it seems to
just ignore robots.txt and do what it wants to anyway.

Google was an afterthought - I'd hoped I could get away with that but it
was still slowing the site down - so I've not analyzed so closely.

Has anyone else noticed this increased aggressiveness in crawling
policy? Any ideas on how to control it??

Cheers,

Steve
-- 
Steve Holdoway BSc(Hons) MNZCS <[email protected]>
http://www.greengecko.co.nz
MSN: [email protected]
Skype: sholdowa

_______________________________________________
NZLUG mailing list [email protected]
http://www.linux.net.nz/cgi-bin/mailman/listinfo/nzlug