openxml full-text indexing
Juan Pablo Giménez <[email protected]>
| Newsgroups | gmane.comp.web.zope.plone.devel |
|---|---|
| Message-ID | <CAEAdVHtF6iMbse-zP0GrXE7+jYX2CXoHcX9nfFkzY_4ZRkXh6Q@mail.gmail.com> |
Hi, I was trying to enable full-text indexing for docx files (openxml) using Products.OpenXml and I hit this bug, http://plone.293351.n2.nabble.com/Unable-to-index-search-content-of-MS-office-2007-files-td7564104.html lxml rises that exception and digging into the problem seems to be a lxml bug https://bugs.launchpad.net/lxml/+bug/1185701 a quick and dirty solution is to patch openxmllib to not use etree.iterparse(), but lxml 3.2.2 implements a fix for that bug so I tried upgrading... everything seems to be working and I added cssselect to my buildout, because this versions of lxml splits cssselect into a new egg and this one is needed by diazo... So, anyone knows about some other issues than could be faced because of this upgrade? I know than openxml is not part of the core, but there is a ticket asking for better support for that kind of files, https://dev.plone.org/ticket/11091, should we try to upgrade lxml in the near future? I mean into the plone core... because openxml support is useless right now, and if we're a good intranet solution, we really need this... cheers, -- Juan Pablo Giménez skype & twitter: jpggimenez Simples Consultoria <http://simplesconsultoria.com.br/> ------------------------------------------------------------------------------ October Webinars: Code for Performance Free Intel webinars can help you accelerate application performance. Explore tips for MPI, OpenMP, advanced profiling, and more. Get the most from the latest Intel processors and coprocessors. See abstracts and register > http://pubads.g.doubleclick.net/gampad/clk?id=60135991&iu=/4140/ostg.clktrk _______________________________________________ Plone-developers mailing list Plone-developers-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org https://lists.sourceforge.net/lists/listinfo/plone-developers