Re: how to extract text reliably from a PDF file?

"Wald" <arnout.standaert@n-o_s-p-a-m.agr.kuleuven.ac.be> Thu, 22 Aug 2002 12:01:22 +0200
Newsgroups gmane.comp.printing.ghostscript.bugs
Organization KULeuvenNet
Message-ID <[email protected]>
"Yaga" <[email protected]> wrote in message
news:[email protected]...
> Hi list,
>
> I want to use GhostScript GNU, to reliably extract text from a PDF file,
> into a plain text file.
>
> The text file could contain Unicode or ASCII, depending on the PDF file.
>
> GhostScript seems very complicated. What is the technique?
>
> Please help me as I'm doing this as a projet for someone else and the
> work has been delayed due to complications.

You could check out the pdftotext program, it's part of the xpdf bundle. A
search on google will give you the right address to download it.

Regards,
Wald