Re: Wikipedia dump file processing shoot out
Petr Pajas <[email protected]>
| Newsgroups | gmane.comp.lang.perl.xml |
|---|---|
| Message-ID | <[email protected]> |
Also, have a look at XML::CompactTree::XS which also uses libxml2 XML reader but slurps a document into Perl data structures in one call to C. -- Petr 2009/12/10 Grant McLean <[email protected]>: > On Thu, 2009-12-10 at 15:39 -0500, Tod Harter wrote: >> Has anyone tried to identify WHY libxml/xslt running in perl is AT ALL >> slower than C? Obviously there is SOME going back and forth between >> perl and C there, but it kind of seems hideously more expensive than >> it should be. > > Tyler's benchmarks are using SAX. The SAX API involves passing Perl > hashrefs around. Constructing Perl hashrefs and strings for each parser > 'event' is expensive. > > The more common way to use libxml from Perl is the DOM approach with > XPath queries. In this mode, the DOM is a C structure and no Perl > structures are built until a query returns results to Perl. This is > *much* faster than SAX, but probably not feasible for a 22GB document. > > The other approach (which I've never used) is to use libxml's 'pull' API > exposed through XML::LibXML::Reader. When you call nextElement() to > fast-forward to the next element, all the action happens at the C layer. > Even better, when you find an interesting element, you can "instruct the > Reader to collect sub-document in form of a DOM tree" (quoted from the > docs, although I'm not sure which method you'd call to achieve it). > Once again, the DOM is a C structure and you can query it via XPath so > the number of Perl datastructures that need to be built is small. > > Cheers > Grant > > _______________________________________________ > Perl-XML mailing list > [email protected] > To unsubscribe: http://listserv.ActiveState.com/mailman/mysubs > > _______________________________________________ Perl-XML mailing list [email protected] To unsubscribe: http://listserv.ActiveState.com/mailman/mysubs