Re: Wikipedia dump file processing shoot out

David Nicol <[email protected]>
Newsgroups gmane.comp.lang.perl.xml
Message-ID <[email protected]>
i would attempt to do this in two passes, the first to split the data
into articles one per file, the second pass to process each article
using a different tool. The second pass could be parallelized and
could start as soon as the first single-article file has been created.


-- 
intake, compression, power, exhaust, repeat.
_______________________________________________
Perl-XML mailing list
[email protected]
To unsubscribe: http://listserv.ActiveState.com/mailman/mysubs
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.