Re: seeking links to TEI corpora

Matthew Davis <[email protected]>
Newsgroups gmane.text.tei.general
Message-ID <[email protected]>
Dear Matthew,

I don’t know that it’s what you’re looking for (it is still early days, there’s still a lot to transcribe and input, and I’m one person doing all the work), but I think my archive of Lydgate works may meet your criteria.  There’s a link to download the xml for each transformed html page, and the raw xml files are stored until an XML folder by work.  

The link is www.minorworksoflydgate.net <http://www.minorworksoflydgate.net/>.  Much of it is still behind a password as I’m hoping to have  a peer review done on it, but the items in the Clopton chantry chapel (http://www.minorworksoflydgate.net/Quis_Dabit/Clopton/ww_qd_1.html <http://www.minorworksoflydgate.net/Quis_Dabit/Clopton/ww_qd_1.html> and http://www.minorworksoflydgate.net/Testament/Clopton/sw_test_1.html <http://www.minorworksoflydgate.net/Testament/Clopton/sw_test_1.html>) are readily accessible since the transcriptions will be published in January.  If it’s what you’re looking for, send me a message off-list and I’ll give you the password credentials for the other items.

There’s also a section on the site, “About the Archive,” that articulates some of my thinking about site design, the decisions I made while encoding, etc.

All the best,
—Matt


> On Dec 19, 2016, at 9:13 AM, Lavin, Matthew J <[email protected]> wrote:
> 
> Apologies for any duplicates received due to cross-posting.
>  
> I am collecting links for publicly accessible, computable TEI (or other similar xml markup such as SGM, LMNL) files. In order to be included, archives/collections/datasets/corpora must have meet one of the two criteria:
>  
> Bulk download of raw xml (not html transformed)
> Xml fully accessible via predictable url structure (an example of this would be the Walk Whitman archive, which as a “raw xml” link on every transformed html page)
>  
> Please note that I am not interested in sample xml, only collections with some kind of curatorial or scholarly focus. Thank you all for any leads! 
>  
> Matthew Lavin
> Clinical Assistant Professor of English and Director of Digital Media Lab
> University of Pittsburgh
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.