Re: [EP-underground] Newby with questions

ePrints Support <[email protected]> Tue, 2 Jul 2002 11:02:59 +0100
Newsgroups gmane.comp.web.eprints.general
Message-ID <[email protected]>
Hi and welcome, I'm Chris, the eprints technical guy,

I can't comment on scanning/ocr'ing free software - Last time I had to 
run a task like this we used Adobe Distiller and the scanners own software.
One thing I did learn which would apply whatever your software is that it's 
worth experimenting to get the settings right before you start. Example:
Our scanning software by default tried to identify graphics & text and do
"clever stuff" which slowed the process down about a factor of 4 and the net
effect was to make the resulting files huge.

The best results was to scan it and then convert it into PDF, performing an
OCR as part of the process. As part of the compression Acrobat can replace
the graphic of each word with a word rendered from a font (much smaller) but
this produced unsatisfactory results (for us). The most useful mode was to
perform the OCR and store it in the PDF for searching & cutting/pasting but
not messing with the image.

EPrints may well be helpful for stage 3 & 4. Records have 4 "states" - 
inbox (metadata still being entered by depositing user)
buffer (finished record awaiting approval)
archive (live and public on the web)
deletion (deleted records - we never throw anything away!)

Ideally you could have your scan-monkey enter the metadata & upload the 
document into eprints and then submit it for approval.

When records are in the approval buffer the documents are only available
to registered users who match a certain profile (by default that is user-type=
editor, user-type=admin OR user.id = eprint.depsitors_id) but this can be
changed.

Only "archive" and "deletion" are exported via OAI. If a record is formally
deleted it is helpful to tell the harvesters so they can remove it.

EPrints is not yet https friendly - which is to say you could run the WHOLE
site on an https server, but ideally you'd only run those parts which required
a username/password.

If you are running a service using eprints then you do not have to obey any
GNU guidelines. Only if you are distributing the changed software. And even
then you don't have to (in fact can't) distribute it as part of GNU, but it
must have the GPL.

Good Luck. 

Chris.

On Mon, Jul 01, 2002 at 11:46:23PM -0400, Denis wrote:
> I just subscribed to this list with the hopes that someone will point me in
> the right direction -
> 
> I am working with the nonprofit Leonard Peltier Defense Committee
> (http://www.freepeltier.org)
> 
> Without going into too many details, one of the projects we must undertake
> is the analysis of a fairly large number of documents, generally of poor
> quality.  These documents are typically obtained from US government agencies
> as the result of Freedom of Information Act (FOIA) pursuits.  They range
> from fairly clean court documents to 30 year-old teletypes with black-marker
> redactions, multi-generation photocopies of same, lots of blemishes, marks,
> handwritten notes etc. (for a sample, see thisone from the SF Chronicle's
> expose of the FBI's war on the president of UC Berkeley:
> http://sfgate.com/news/special/pages/2002/campusfiles/documents/1a1.shtml )
> 
> The goal is to share the analysis among a group of experts widely separated
> in space, and to allow central indexing and archiving (of both the paper and
> digital documents).
> 
> I am just picking up on the technology now, but I see I will need to
> 1. capture the docs
> 2. perform as best as possible OCR
> 3. database the docs - with various indices - doc#, page#, doc type,
> comments, reviewer(s), priority, date of doc, date of indexing, source
> location, etc.
> 4. Share the docs with the reviewers - through a secure server, or by
> porting the data via CD or some way then filing their comments into the
> database (ideally, the reviewers would not have to be connected to the
> internet for the entire time that they are reviewing)
> 
> What I've seen on the OAI seem to fall into two categories - clean,
> dissertation-type docs that are very amenable to text searches, and
> historical archives, where very old handwritten documents are portrayed as
> images, with transcripts provided separately.  I am sure I have missed quite
> a bit, and oversimplify here.
> 
> I have a tight time frame, as well, and could relegate the indexing etc to a
> later date.  I would like to begin the scanning/capture soon, though, so
> seek advice.
> 
> I have seen and appreciate the GNU guidelines.  I would hope that the work
> that we put into this project, which will ultimately lead to an open archive
> (although the material will be initially sequestered as it will hopefully
> enable court proceedings) can fall within the OAI/eprint guidelines.  I
> would like the archive to be a model for social justice groups to begin
> putting their archives up on the web.
> 
> I do need a fast (three weeks) solution to getting the capture phase done -
> I know there is some commercial software that could facilitate that, but if
> I am forced to do that, will the data be easily transferred to the eprints
> format? (I would probably get a decent sheet-feeding scanner and an OCR
> package from a company called ABBYY)
> 
> Thanks for any advice.  A group that does a lot of this is the National
> Security Archive, http://www.gwu.edu/~nsarchiv/ , but their tech person
> pointed me to the very expensive commercial products.  I am sure they would
> be amenable to reviewing this sector of the digitization universe.  I am
> very impressed with what I have seen through eprints/OAI
> 
> Denis Moynihan
> Leonard Peltier Defense Committee
> Lawrence, KS
> 785-842-5774
> [email protected]
> 

-- 

 Christopher Gutteridge                   [email protected]
 ePrints2 Coder, Support and Stuff        +44 23 8059 4833