Re: XPath 1.0 change proposal
"C. M. Sperberg-McQueen" <[email protected]> Thu, 14 Mar 2013 09:34:44 -0600
| Newsgroups | gmane.text.xml.xpath.general |
|---|---|
| Message-ID | <[email protected]> |
Thank you, James. I agree with some of what you say, and disagree with = some. On Mar 14, 2013, at 7:52 AM, James Clark wrote: >=20 >=20 > Although I appreciate Michael's work on formalizing the XPath 1.0 data = model, I do not think that at this stage a major rewrite of the XPath = 1.0 data model is a good idea. =20 I agree. =20 That's why I proposed small fixes to repair errors in the definition of = the data model, and not a major rewrite. =20 > I would suggest that, after nearly 14 years, an extremely conservative = policy should be adopted towards changes: changes should be made only = when there is a genuine error that is manifested in discrepancies = between implementations or inconsistencies between implementations and = the spec. The nature of the errors in the definition of the data model is that = they amount to discrepancies within the spec. The rest of XPath 1.0 assumes that = every=20 instance of the data model, as defined in section 5, will have certain = properties. It is the job of section 5 to ensure that that is so, and in the spec as = written=20 the properties in question are not in fact guaranteed. =20 These discrepancies are unlikely to show up in implementations of XPath = 1.0 as a whole, since implementors are likely to be guided by the = assumptions manifest elsewhere in the spec more than by the details of the data = model definition. They will, however, show up in any attempt to implement, or = formalize, the XPath 1.0 data model by itself. That is how I became aware of the = errors in the first place. > The change proposal claims that it was a goal of XPath 1.0 for the = data model be defined without dependencies on XML 1.0. I find this = claim bizarre given that XML 1.0 is referenced normatively and the data = data model definition is full of references to XML 1.0. If there was no intent to define the data model without dependencies on = XML 1.0, then at least half the text in section 5 is pointless, unnecessary repetition of things that = are obvious from the XML spec. The choice seems to be between reading the spec as having a coherent = goal which in some important details it failed to achieve, and reading the spec as given to garrulous = irrelevancies. > The change proposal seems to be claiming that XPath 1.0 is full of = bugs in need of correction because it does not meet a goal that it never = had. >=20 > The change proposal also claims that it is a goal of XPath 1.0 that = the data model be defined formally. This is clearly not the case. =20 By any mathematical standard, the prose of XPath 1.0 would count as = informal. But that is also true of the prose in the change proposal. Compared with other specs, I think the data model section of XPath 1.0 = is more explicit and formal than most. > XPath 1.0 does not make the slightest attempt to be formal. Rather it = aims to be succinct and readily understandable. The level of formality = in the data model definition is similar to that of the rest of the spec = and of companion specs (XML 1.0, XML Namespaces, XSLT 1.0). It is also = virtually impossible to be really rigorous about the construction of the = data model from the XML document, without specifying this in the XML = spec itself: for each syntax production the XML spec would need to = explain how to corresponding data model was constructed. I see no connection between the formality of an exposition and the title = of the document in which it appears. =10The XML spec is not formal about (for example) = identity criteria for elements, because nothing in the XML spec appeals to element identity = (at least, not=20 for the cases it leaves indeterminate). The XPath 1.0 spec does need to = be determinate about element node identity, and it seems bizarre to me to suggest that = it could not be more precise or careful without changes to the text of the XML 1.0 spec. >=20 > I am also not convinced that in many cases the proposed wording = changes are in fact improvements. If the WG does decide to go ahead = with this change, I can make some more detailed comments. But for the = moment, I would just mention a couple of points. >=20 > XPath 1.0 does not constrain the root node to have exactly one element = child. In the case where the data model is constructed from an XML = document, there will of course be exactly one child. But in other cases = (eg querying into a DOM DocumentFragment) it would be unhelpful to = impose such a restriction. (XPath 1.0 is generally fairly loose -- for = example, it does not define conformance -- so as to provide maximum = flexibility to referencing specs.) Thank you for this clarification. The XML spec seems, then, to be normative for the description of the = data model, except for the parts of it that don't apply. On this view, the XPath 1.0 spec is = readily understandable only for readers gifted with a certain degree of clairvoyance. >=20 > The reason why the spec uses terminology like "There is an element = node for every element" instead of referencing particular productions is = because of entity expansion. For example, given >=20 > <!DOCTYPE doc [ > <!ENTITY e "<x>foo</x>"> > ]> > <doc>&e;&e;</doc> >=20 > I am comfortable with saying (somewhat vaguely) that there are three = elements. =20 The problem is that there is nothing in the =10XML spec or the Infoset = spec that could be used to argue that there are three elements here, instead of two. Many = people are comfortable saying that there are three elements here, but a count = of two elements is equally compatible with the XML specification. > I am much less comfortable saying that there are three occurrences of = the "element" production (in fact, I would say it is clear that there = are only two occurrences of the "element" production). On the contrary; after entity expansion we have a sequence of character = types matching the document production of the XML spec, but we do not = necessarily have a sequence of character tokens matching the document production. = In the sequence of character types, there are clearly three occurrences of = strings (sequences of character types) matching the element production, even = though there are only two such string-types. That is the difference between a string = type and an occurrence of a string type.=20 --=20 **************************************************************** * C. M. Sperberg-McQueen, Black Mesa Technologies LLC * http://www.blackmesatech.com=20 * http://cmsmcq.com/mib =20 * http://balisage.net ****************************************************************