Re: Wikipedia dump file processing shoot out
David Nicol <[email protected]>
| Newsgroups | gmane.comp.lang.perl.xml |
|---|---|
| Message-ID | <[email protected]> |
i would attempt to do this in two passes, the first to split the data into articles one per file, the second pass to process each article using a different tool. The second pass could be parallelized and could start as soon as the first single-article file has been created. -- intake, compression, power, exhaust, repeat. _______________________________________________ Perl-XML mailing list [email protected] To unsubscribe: http://listserv.ActiveState.com/mailman/mysubs