abcl.org/trac/* being spidered by multiple bots
Erik Huelsmann <[email protected]> Sun, 7 Jan 2024 11:51:35 +0100
| Newsgroups | gmane.editors.j.devel |
|---|---|
| Message-ID | <CACOoB6hBq5Jk7q_uRHbRX0EO+TFO7FiowXgTc=zYs-0HcyYeLA__47157.5310266175$1704624766$gmane$org@mail.gmail.com> |
--000000000000281f39060e58dee1 Content-Type: text/plain; charset="UTF-8" Hi Mark, Last night a number of disruptions in cliki.net availability were reported, a happens regularly lately -- and definitely during the weekly backup. So I checked the server this morning. The server had a load of 50 (>15 per cpu); 80% of that was going to Trac, however the access log trac.common-lisp.net_access.log was silent. Turns out all the traffic was on abcl.org. It was being spidered like crazy. Trac - or our server setup (or both) - really isn't up to that task, which is why there is only limited spidering allowed on trac.common-lisp.net. (Enough to spider "current" trunk of projects and their issues.) To restrict spidering, I've changed the robots.txt file at /project/armedbear/public_html/robots.txt. Could you please commit that robots.txt (or an even stricter one) to the repository that the site is being synced from? Thanks! Regards, Erik. --000000000000281f39060e58dee1 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr"><div>Hi Mark,</div><div><br></div><div><br></div><div>Last= night a number of disruptions in <a href=3D"http://cliki.net">cliki.net</a= > availability were reported, a happens regularly lately -- and definitely = during the weekly backup.</div><div><br></div><div>So I checked the server = this morning. The server had a load of 50 (>15 per cpu); 80% of that was= going to Trac, however the access log trac.common-lisp.net_access.log was = silent. Turns out all the traffic was on <a href=3D"http://abcl.org">abcl.o= rg</a>. It was being spidered like crazy. Trac - or our server setup (or bo= th) - really isn't up to that task, which is why there is only limited = spidering allowed on <a href=3D"http://trac.common-lisp.net">trac.common-li= sp.net</a>. (Enough to spider "current" trunk of projects and the= ir issues.)</div><div><br></div><div>To restrict spidering, I've change= d the robots.txt file at /project/armedbear/public_html/robots.txt.</div><d= iv><br></div><div>Could you please commit that robots.txt=C2=A0 (or an even= stricter one) to the repository that the site is being synced from?</div><= div><br></div><div><br></div><div>Thanks!</div><div><br></div><div>Regards,= </div><div><br></div><div>Erik.<br></div><div><br></div><div><br></div></di= v> --000000000000281f39060e58dee1--