Re: New user and questions :-)
"Markus Hoenicka" <[email protected]>
| Newsgroups | gmane.text.refdb.general |
|---|---|
| Message-ID | <[email protected]> |
Janusz S. Bień writes: > There are probably some tools already for extracting metadata from > PDFs. I think such facility is in particular built into Greenstone > > http://www.greenstone.org/cgi-bin/library > > It is GPLed, so the relevant code can be reused. > Thanks for the pointer. They do seem to have some tools that seem useful for this purpose, but they also mention the limited utility with particular kinds of PDF files (e.g. PDFs of older articles that contain scanned page images instead of text). At least the newer PDFs all seem to contain a doi in the document properties. These are accessible e.g. through a Perl API, see http://search.cpan.org/~areibens/PDF-API2-0.51/lib/PDF/API2.pm The $pdf->info function seems to return the metadata, with the doi info usually in the title field. This may be a good starting point at least for newer PDFs. regards Markus -- Markus Hoenicka [email protected] (Spam-protected email: replace the quadrupeds with "mhoenicka") http://www.mhoenicka.de ------------------------------------------------------- This SF.Net email is sponsored by xPML, a groundbreaking scripting language that extends applications into web and mobile media. Attend the live webcast and join the prime developer group breaking into this new coding territory! http://sel.as-us.falkag.net/sel?cmd=lnk&kid0944&bid$1720&dat1642