Bug#949173: packages.debian.org: robots.txt doesn't actually block anything

Michael Lustfield <[email protected]>
Newsgroups gmane.linux.debian.devel.www
Message-ID <20200117235930.19df9392__4553.82464440248$1579327401$gmane$org@spike.lustfield.net>
On Sat, 18 Jan 2020 01:14:02 +0000
Paul Wise <[email protected]> wrote:

> On Fri, Jan 17, 2020 at 6:57 PM Adam D. Barratt wrote:
> 
> > which is effectively the same as allowing everything. "Disallow: /"
> > might be more logical, unless there is a desire / requirement to allow
> > crawling and indexing of (parts of) the site.  
> 
> I expect we want to allow crawling the site, all of the pages are
> public and most of them are useful for search engines to index.

+1 This is something I've found very helpful and convenient. I could survive
without those pages being indexed by search engines, but I'd prefer not.

Would it be helpful to disallow certain pages, such as */download?

This is random, but I noticed that the "Tags" link on package pages links to
debtags.alioth.debian.org/edit.html, which no longer exists.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.