Problems with 22000 plus file crawl with large file sizes
"James Garrett" <[email protected]> Fri, 7 Nov 2003 23:49:35 -1000
| Newsgroups | gmane.comp.web.perlfect-search |
|---|---|
| Message-ID | <003901c3a5dd$9e6572f0$8f474140@WorkStation> |
A couple of things, Set with a gig of ram and $LOW_MEMORY_INDEX =0; successful crawl took day and half but machine melted at point of copying hash to DBase and had to reboot. Any way to pick up from that point and finishing up copying the hash files without having to rerun ./indexer.pl? Have read through every archive searching for how to deal with large/numerous files and would like to use perlfect to crawl 80,000/plus files with max 6meg file sizes. Have tried numerous changes to conf.pl and have yet to have successful crawl exceeding 12,000. Kernel runs out of memory. Tests on smaller file count no probs. What could be happening that takes out a machine ? Setting $LOW_MEMORY_INDEX =1; is just so painfully slow. Found reference to $FLUSH_FREQUENCY =100; but couldn't fig how to patch ./indexer.pl and don't really know if that will help at all. Also am afraid temp files could grow hugh. Such a thing as a 2gig max file size in Linux? Perlfect ? Any light/advice/help/support/suggestions you can shed appreciated. James Garrett worldebooklibrary.com