Scraping sites with code

Julian Bond <julian_bond-/Fkc1E/MbsFWk0Htik3J/[email protected]>
Newsgroups gmane.network.syndication.discuss
Message-ID <[email protected]>
Somebody has just told me that Google News look for PHP in the user 
agent and return 403 forbidden if they find it. It seems likely that 
they also look for other auto-generated code names.

FWIW http://www.google.com/robots.txt also bans robots from all the main 
directories.

I wonder how many other sites that routinely get scraped to produce RSS 
do the same thing.

For those of you using gnews2rss.php, coding round this is easy and left 
as an exercise for the php programmer. Using Curl or looking at the 
documentation for fopen() will solve the problem.

Meanwhile isn't it about time Google produced RSS themselves as an 
alternate output format from Search and News Search? Way back in June a 
senior Google person told me it was coming soon.

-- 
Julian Bond Email&MSM: julian.bond-/Fkc1E/MbsFWk0Htik3J/[email protected]
Webmaster:              http://www.ecademy.com/
Personal WebLog:       http://www.voidstar.com/
M: +44 (0)77 5907 2173   T: +44 (0)192 0412 433

 

Your use of Yahoo! Groups is subject to http://docs.yahoo.com/info/terms/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.