Fwd: XML 1.1 -- Please stab me with a dull knife and trample my dead body
Canadian Wilf <[email protected]> Fri, 7 Sep 2012 05:10:20 -0700
| Newsgroups | gmane.comp.java.jdom.general |
|---|---|
| Message-ID | <CAL8g3USfsG4UY3UzWddY=nQm2HHO8Dcz8ypd2AATC7DWFoUAKQ@mail.gmail.com> |
According to the xml 1.1 spec: 4 Physical Structures ... > [Definition: An *unparsed entity* is a resource whose contents may or may > not be text <http://www.w3.org/TR/xml11/#dt-text>, and if text, may be > other than XML. Each unparsed entity has an associated notation<http://www.w3.org/TR/xml11/#dt-notation>, > identified by name. Beyond a requirement that an XML processor make the > identifiers for the entity and notation available to the application, XML > places no constraints on the contents of unparsed entities.] AND Entities may be either parsed or unparsed. [Definition: The contents of a *parsed > entity* are referred to as its replacement text<http://www.w3.org/TR/xml11/#dt-repltext>; > this text <http://www.w3.org/TR/xml11/#dt-text> is considered an integral > part of the document.] [Definition: An *unparsed entity* is a resource whose contents may or may > not be text <http://www.w3.org/TR/xml11/#dt-text>, and if text, may be > other than XML. Each unparsed entity has an associated notation<http://www.w3.org/TR/xml11/#dt-notation>, > identified by name. Beyond a requirement that an XML processor make the > identifiers for the entity and notation available to the application, XML > places no constraints on the contents of unparsed entities.] > Parsed entities are invoked by name using entity references; unparsed > entities by name, given in the value of *ENTITY* or *ENTITIES* attributes. In the current JDOM version, Element method setText(string) and also addContent(CDATA) refuses text that contains illegal characters. It is treating the data provided as 'parsed' when it should by the spec be treating it as free content. I understand: 1) The xml 1.1 spec defines a parsed entity as its 'replacement text'. 2) Replacement text' would refer to the actual textual makeup of a serialized Element, not the data an Element holds in a Text content element Then, if the above is true, the current implementation is actually wrong to verify data. I propose that JDOM stop verifying data set as Element text and CDATA and leave it to the xerces (or whatever) to make sure the document is proper 1.1. Am I understanding everything correctly? Thoughts? ---------- Forwarded message ---------- From: Canadian Wilf <[email protected]> Date: Thu, Sep 6, 2012 at 9:52 PM Subject: XML 1.1 -- Please stab me with a dull knife and trample my dead body To: [email protected] Hi All, I just learned that in order to safely use JDOM2, I will need to sanitize my Element .setText(string) so that the parsed data does not contain verboten characters under the XML 1.1 spec. I have an ascii processor and it needs to be able to use xml as a document format. Unfortunately, not all ascii is allowed in an Element text. Stab me with a dull knife and trample my dead body. But ..... please please please don't make me sanitize all my data before putting it into XML Elements. 1) It makes my programming task much more cumbersome because I must ensure not to feed any of the new verboten and doomed ascii/UTF-8 characters to store as xml text. 2) No one uses xml 1.1, do they? 3) It slows down the parsing (a very small amount) with all the element text checking. Now that JDOM2 is xml 1.1 compatible, is there any turning back. Can this be undone? Does everyone understand that their software will bust if data provided as text is not adhering to the new standard? What about you? How do you deal with it when using the libraries? Wilf _______________________________________________ To control your jdom-interest membership: http://www.jdom.org/mailman/options/jdom-interest/[email protected]