[EP-underground] Newby with questions

Denis <[email protected]> Mon, 01 Jul 2002 23:46:23 -0400
Newsgroups gmane.comp.web.eprints.general
Message-ID <B9469A4E.7A4A%[email protected]>
I just subscribed to this list with the hopes that someone will point me in
the right direction -

I am working with the nonprofit Leonard Peltier Defense Committee
(http://www.freepeltier.org)

Without going into too many details, one of the projects we must undertake
is the analysis of a fairly large number of documents, generally of poor
quality.  These documents are typically obtained from US government agencies
as the result of Freedom of Information Act (FOIA) pursuits.  They range
from fairly clean court documents to 30 year-old teletypes with black-marker
redactions, multi-generation photocopies of same, lots of blemishes, marks,
handwritten notes etc. (for a sample, see thisone from the SF Chronicle's
expose of the FBI's war on the president of UC Berkeley:
http://sfgate.com/news/special/pages/2002/campusfiles/documents/1a1.shtml )

The goal is to share the analysis among a group of experts widely separated
in space, and to allow central indexing and archiving (of both the paper and
digital documents).

I am just picking up on the technology now, but I see I will need to
1. capture the docs
2. perform as best as possible OCR
3. database the docs - with various indices - doc#, page#, doc type,
comments, reviewer(s), priority, date of doc, date of indexing, source
location, etc.
4. Share the docs with the reviewers - through a secure server, or by
porting the data via CD or some way then filing their comments into the
database (ideally, the reviewers would not have to be connected to the
internet for the entire time that they are reviewing)

What I've seen on the OAI seem to fall into two categories - clean,
dissertation-type docs that are very amenable to text searches, and
historical archives, where very old handwritten documents are portrayed as
images, with transcripts provided separately.  I am sure I have missed quite
a bit, and oversimplify here.

I have a tight time frame, as well, and could relegate the indexing etc to a
later date.  I would like to begin the scanning/capture soon, though, so
seek advice.

I have seen and appreciate the GNU guidelines.  I would hope that the work
that we put into this project, which will ultimately lead to an open archive
(although the material will be initially sequestered as it will hopefully
enable court proceedings) can fall within the OAI/eprint guidelines.  I
would like the archive to be a model for social justice groups to begin
putting their archives up on the web.

I do need a fast (three weeks) solution to getting the capture phase done -
I know there is some commercial software that could facilitate that, but if
I am forced to do that, will the data be easily transferred to the eprints
format? (I would probably get a decent sheet-feeding scanner and an OCR
package from a company called ABBYY)

Thanks for any advice.  A group that does a lot of this is the National
Security Archive, http://www.gwu.edu/~nsarchiv/ , but their tech person
pointed me to the very expensive commercial products.  I am sure they would
be amenable to reviewing this sector of the digitization universe.  I am
very impressed with what I have seen through eprints/OAI

Denis Moynihan
Leonard Peltier Defense Committee
Lawrence, KS
785-842-5774
[email protected]