Re: how to extract text reliably from a PDF file?
"Wald" <arnout.standaert@n-o_s-p-a-m.agr.kuleuven.ac.be> Thu, 22 Aug 2002 12:01:22 +0200
| Newsgroups | gmane.comp.printing.ghostscript.bugs |
|---|---|
| Organization | KULeuvenNet |
| Message-ID | <[email protected]> |
"Yaga" <[email protected]> wrote in message news:[email protected]... > Hi list, > > I want to use GhostScript GNU, to reliably extract text from a PDF file, > into a plain text file. > > The text file could contain Unicode or ASCII, depending on the PDF file. > > GhostScript seems very complicated. What is the technique? > > Please help me as I'm doing this as a projet for someone else and the > work has been delayed due to complications. You could check out the pdftotext program, it's part of the xpdf bundle. A search on google will give you the right address to download it. Regards, Wald