Re: seeking links to TEI corpora
Matthew Davis <[email protected]>
| Newsgroups | gmane.text.tei.general |
|---|---|
| Message-ID | <[email protected]> |
Dear Matthew, I don’t know that it’s what you’re looking for (it is still early days, there’s still a lot to transcribe and input, and I’m one person doing all the work), but I think my archive of Lydgate works may meet your criteria. There’s a link to download the xml for each transformed html page, and the raw xml files are stored until an XML folder by work. The link is www.minorworksoflydgate.net <http://www.minorworksoflydgate.net/>. Much of it is still behind a password as I’m hoping to have a peer review done on it, but the items in the Clopton chantry chapel (http://www.minorworksoflydgate.net/Quis_Dabit/Clopton/ww_qd_1.html <http://www.minorworksoflydgate.net/Quis_Dabit/Clopton/ww_qd_1.html> and http://www.minorworksoflydgate.net/Testament/Clopton/sw_test_1.html <http://www.minorworksoflydgate.net/Testament/Clopton/sw_test_1.html>) are readily accessible since the transcriptions will be published in January. If it’s what you’re looking for, send me a message off-list and I’ll give you the password credentials for the other items. There’s also a section on the site, “About the Archive,” that articulates some of my thinking about site design, the decisions I made while encoding, etc. All the best, —Matt > On Dec 19, 2016, at 9:13 AM, Lavin, Matthew J <[email protected]> wrote: > > Apologies for any duplicates received due to cross-posting. > > I am collecting links for publicly accessible, computable TEI (or other similar xml markup such as SGM, LMNL) files. In order to be included, archives/collections/datasets/corpora must have meet one of the two criteria: > > Bulk download of raw xml (not html transformed) > Xml fully accessible via predictable url structure (an example of this would be the Walk Whitman archive, which as a “raw xml” link on every transformed html page) > > Please note that I am not interested in sample xml, only collections with some kind of curatorial or scholarly focus. Thank you all for any leads! > > Matthew Lavin > Clinical Assistant Professor of English and Director of Digital Media Lab > University of Pittsburgh >