Re: A question about encoding and text fixing platforms for undergraduate curators
"Sewell, David R. (drs2n)" <[email protected]> Thu, 27 Apr 2017 09:42:20 -0400
| Newsgroups | gmane.text.tei.general |
|---|---|
| Message-ID | <[email protected]> |
On Thu, 27 Apr 2017, Martin Holmes wrote: [...] > XML is not hard. Word, by contrast, is a concoction of frustrations, and > getting DOCX into decent TEI when you're done is horribly difficult. If you create a Word template with styles (whether built-in or custom) that are sufficient to define each block-level and inline element that needs to be expressed in XML, then it's certainly feasible to set up a workflow involving a translation tool like oXgarage. It mainly requires careful analysis of the result of oXgarage conversion in order to create a further XSLT transform to produce your desired final output. But in our experience (based on using such a workflow with distributed authors for a reference project), it's going to be almost inevitable that now and then the Word file you receive is going to have some kind of unexpected crud in it resulting from the user adding an unforeseen style, accidentally changing a format, or doing anything unpredictable. It is probably easier to teach people to use a consistent set of markup conventions in oXygen than to get them to use MS Word with 100% accuracy according to whatever template and style rules you want them to follow. David -- David Sewell Manager of Digital Initiatives The University of Virginia Press Email: [email protected] Tel: +1 434 924 9973 Web: http://www.upress.virginia.edu/rotunda