Re: [EP-underground] Duplicate papers and generations of archives

Christopher Gutteridge <[email protected]> Tue, 24 May 2005 12:26:10 +0100
Newsgroups gmane.comp.web.eprints.general
Message-ID <[email protected]>
This is a question which will become more and more important as the archives
fill up.

My advice locally has always been that two copies are better than zero 
copies.

In our archive it is the joint responsibility of all local authors to 
ensure that the
record of their work is correct.

If there is a duplicate record in the local archive then the site admin 
can set
the "later version of" field on the less-acurate one, and then move it out
of the archive into the "deleted" buffer. This will cause it to be 
replaced by
a "newer version available" page.

Authors are required to submit work that was created while working at
our institution, they _may_ submit work created previously.

This is only my summary of our policy, and the policy will doubtless
evolve.

I'm interested to see what other points of view people have on this emerging
issue.

Ross A. Beyer wrote:

>Hello,
>
>I have some questions about duplicates and generations of archives.
>I can't say that I've read all of the open-access literature, and
>maybe there are already answers for my questions out there.  If so,
>please let me know.
>
>So I have some scenarios where a paper could get self-archived in
>two or more servers.  I'm curious about the technical ramifications
>and the etiquette.
>
>Duplicate Scenario 1:
>	Let's say that a self-archive server for my scientific
>	subject area gets created that anyone in my field can deposit
>	into.  This goes on for a few years and more people in my
>	discipline finally "get it".  They push to get their
>	institution to create an open archive.  Now they want to
>	set a good example, so they self-archive their papers into
>	their institutional archive, but they have already self-archived
>	them into the discipline archive.  There are now two servers
>	with the same information.  Or does etiquette indicate that
>	once you have self-archived your paper in a server, you
>	don't self-archive it again in another?
>
>Duplicate Scenario 2:
>	There is a paper with a big author list.  The third author is an 
>	open access advocate, and self-archives the paper in their 
>	institution's archive.  Some time later, the tenth author learns 
>	that his institution has an eprints server, and without knowing
>	that author 3 has already self-archived the paper, author 10
>	self-archives the paper at their own institution.  Again, the
>	same paper is now in more than one eprints server.  Or does only
>	the first author have the right to self-archive a paper?
>
>Duplicate Scenario 3:
>	I convince my institution to create a self-archiving eprints
>	server, and I self-archive my work there.  Then I get a new
>	job at another institution, and they also have (or I have
>	pushed to have them create) an eprints server.  Do I
>	re-self-archive my earlier work at the new institution's
>	server?  Naturally, this would duplicate the data.
>
>I suppose in some sense, it doesn't matter if more than one server
>has the same paper self-archived, it is up to whatever OAI
>data-harvester that ends up with the duplicate "copies" to decide
>what to do about them.  What I'm concerned about is the ultimate
>user of the data.  If someone searches for a particular article,
>are they going to get ten links to ten different eprint servers for
>the same paper?
>
>
>I also have a question about generations of archives.  Let's say I
>begin my own personal eprints server as a demonstration of concept.
>With it, I convince my department to start an eprints server.  Do
>I then shut down my personal server, and put all of those papers
>on the departmental server?  Similarly, the department server runs
>for a few years, and people know to go there for data, and then the
>university as a whole decides to run an eprints server.  Does the
>department then shut down their server, and put all of their data
>on the university's server?  I realize that there are clever things
>that can be done with redirects and URL rewriting on the server
>side to just redirect all URLs requests to the departmental server
>to the university's server, to cause that transition to be almost
>transparent, but still I'm curious about what you think.
>  
>

-- 
Christopher Gutteridge -- [email protected] -- +44 (0)23 8059 4833 
University of Southampton, School of Electronics and Computer Science

Chris is currently listening to: . Obligatory Random Quote: Shout     
Britain, raise a joyful shout, The Tyrant Tories all are out --       
Deluded Britains -- cease your din -- For lo -- the scoundrel Whigs   
are in. -- Hartley Coleridge (1796-1849)