Re: New user and questions :-)

"Markus Hoenicka" <[email protected]>
Newsgroups gmane.text.refdb.general
Message-ID <[email protected]>
Janusz S. Bień writes:
 > There are probably some tools already for extracting metadata from
 > PDFs. I think such facility is in particular built into Greenstone
 > 
 >                 http://www.greenstone.org/cgi-bin/library
 > 
 > It is GPLed, so the relevant code can be reused.
 > 

Thanks for the pointer. They do seem to have some tools that seem
useful for this purpose, but they also mention the limited utility
with particular kinds of PDF files (e.g. PDFs of older articles that
contain scanned page images instead of text).

At least the newer PDFs all seem to contain a doi in the document
properties. These are accessible e.g. through a Perl API, see

http://search.cpan.org/~areibens/PDF-API2-0.51/lib/PDF/API2.pm

The $pdf->info function seems to return the metadata, with the doi
info usually in the title field. This may be a good starting point at
least for newer PDFs.

regards
Markus

-- 
Markus Hoenicka
[email protected]
(Spam-protected email: replace the quadrupeds with "mhoenicka")
http://www.mhoenicka.de



-------------------------------------------------------
This SF.Net email is sponsored by xPML, a groundbreaking scripting language
that extends applications into web and mobile media. Attend the live webcast
and join the prime developer group breaking into this new coding territory!
http://sel.as-us.falkag.net/sel?cmd=lnk&kid0944&bid$1720&dat1642
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.