Extracting Text
FDnC Red <[email protected]>
| Newsgroups | gmane.comp.java.lib.itext.general |
|---|---|
| Message-ID | <[email protected]> |
How in the world do I extract text from this PDF? It all comes out as gibberish non-printable characters, except when I choose "Copy With Formatting" in Acrobat. I've tried several different techniques on the web including iText SimpleExtractionStrategy, TopToBottomExtractionStrategy, LocationTextExtractionStrategyEx, and other tools as well. Adobe's "Copy With Formatting" somehow decodes the text in this PDF. Any ideas how to decode this text with iTextSharp? Thanks, Darren ------------------------------------------------------------------------------ Meet PCI DSS 3.0 Compliance Requirements with EventLog Analyzer Achieve PCI DSS 3.0 Compliant Status with Out-of-the-box PCI DSS Reports Are you Audit-Ready for PCI DSS 3.0 Compliance? Download White paper Comply to PCI DSS 3.0 Requirement 10 and 11.5 with EventLog Analyzer http://pubads.g.doubleclick.net/gampad/clk?id=154622311&iu=/4140/ostg.clktrk _______________________________________________ iText-questions mailing list [email protected] https://lists.sourceforge.net/lists/listinfo/itext-questions iText(R) is a registered trademark of 1T3XT BVBA. Many questions posted to this list can (and will) be answered with a reference to the iText book: http://www.itextpdf.com/book/ Please check the keywords list before you ask for examples: http://itextpdf.com/themes/keywords.php
pf000003.pdf
(application/pdf, 382.5 KB) - not displayed