[EP-underground] Newby with questions
Denis <[email protected]> Mon, 01 Jul 2002 23:46:23 -0400
| Newsgroups | gmane.comp.web.eprints.general |
|---|---|
| Message-ID | <B9469A4E.7A4A%[email protected]> |
I just subscribed to this list with the hopes that someone will point me in the right direction - I am working with the nonprofit Leonard Peltier Defense Committee (http://www.freepeltier.org) Without going into too many details, one of the projects we must undertake is the analysis of a fairly large number of documents, generally of poor quality. These documents are typically obtained from US government agencies as the result of Freedom of Information Act (FOIA) pursuits. They range from fairly clean court documents to 30 year-old teletypes with black-marker redactions, multi-generation photocopies of same, lots of blemishes, marks, handwritten notes etc. (for a sample, see thisone from the SF Chronicle's expose of the FBI's war on the president of UC Berkeley: http://sfgate.com/news/special/pages/2002/campusfiles/documents/1a1.shtml ) The goal is to share the analysis among a group of experts widely separated in space, and to allow central indexing and archiving (of both the paper and digital documents). I am just picking up on the technology now, but I see I will need to 1. capture the docs 2. perform as best as possible OCR 3. database the docs - with various indices - doc#, page#, doc type, comments, reviewer(s), priority, date of doc, date of indexing, source location, etc. 4. Share the docs with the reviewers - through a secure server, or by porting the data via CD or some way then filing their comments into the database (ideally, the reviewers would not have to be connected to the internet for the entire time that they are reviewing) What I've seen on the OAI seem to fall into two categories - clean, dissertation-type docs that are very amenable to text searches, and historical archives, where very old handwritten documents are portrayed as images, with transcripts provided separately. I am sure I have missed quite a bit, and oversimplify here. I have a tight time frame, as well, and could relegate the indexing etc to a later date. I would like to begin the scanning/capture soon, though, so seek advice. I have seen and appreciate the GNU guidelines. I would hope that the work that we put into this project, which will ultimately lead to an open archive (although the material will be initially sequestered as it will hopefully enable court proceedings) can fall within the OAI/eprint guidelines. I would like the archive to be a model for social justice groups to begin putting their archives up on the web. I do need a fast (three weeks) solution to getting the capture phase done - I know there is some commercial software that could facilitate that, but if I am forced to do that, will the data be easily transferred to the eprints format? (I would probably get a decent sheet-feeding scanner and an OCR package from a company called ABBYY) Thanks for any advice. A group that does a lot of this is the National Security Archive, http://www.gwu.edu/~nsarchiv/ , but their tech person pointed me to the very expensive commercial products. I am sure they would be amenable to reviewing this sector of the digitization universe. I am very impressed with what I have seen through eprints/OAI Denis Moynihan Leonard Peltier Defense Committee Lawrence, KS 785-842-5774 [email protected]