Re: reading pictures of text in pdf

Paul Merrell <[email protected]>
Newsgroups gmane.linux.redhat.blinux.general
Message-ID <CAJ1g4g9PVjTOox6b7ngak6HP1_OV7MnTjY8zRX5-t-NAS+DEMA@mail.gmail.com>
On Thu, Nov 12, 2015 at 4:10 PM, Brian Tew <[email protected]> wrote:
> Is there anything in linux that can convert a pdf file that is a picture of text
> into real actual plain text?

Assuming there's no DRM involved, tesseract-OCR is probably your best
bet. <https://code.google.com/p/tesseract-ocr/>. The source code has
moved to <https://github.com/tesseract-ocr> but the documentation
seems to still be on code.google.com.

Best regards,

Paul

-- 
[Notice not included in the above original message:  The U.S. National
Security Agency neither confirms nor denies that it intercepted this
message.]
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.