Re: XPath 1.0 change proposal

"C. M. Sperberg-McQueen" <[email protected]> Thu, 14 Mar 2013 09:34:44 -0600
Newsgroups gmane.text.xml.xpath.general
Message-ID <[email protected]>
Thank you, James.  I agree with some of what you say, and disagree with =
some.

On Mar 14, 2013, at 7:52 AM, James Clark wrote:

>=20
>=20
> Although I appreciate Michael's work on formalizing the XPath 1.0 data =
model, I do not think that at this stage a major rewrite of the XPath =
1.0 data model is a good idea. =20

I agree. =20

That's why I proposed small fixes to repair errors in the definition of =
the data model, and
not a major rewrite. =20

> I would suggest that, after nearly 14 years, an extremely conservative =
policy should be adopted towards changes: changes should be made only =
when there is a genuine error that is manifested in discrepancies =
between implementations or inconsistencies between implementations and =
the spec.

The nature of the errors in the definition of the data model is that =
they amount
to discrepancies within the spec.  The rest of XPath 1.0 assumes that =
every=20
instance of the data model, as defined in section 5, will have certain =
properties.
It is the job of section 5 to ensure that that is so, and in the spec as =
written=20
the properties in question are not in fact guaranteed. =20

These discrepancies are unlikely to show up in implementations of XPath =
1.0
as a whole, since implementors are likely to be guided by the =
assumptions
manifest elsewhere in the spec more than by the details of the data =
model
definition.  They will, however, show up in any attempt to implement, or =
formalize,
the XPath 1.0 data model by itself.  That is how I became aware of the =
errors
in the first place.

> The change proposal claims that it was a goal of XPath 1.0 for the =
data model be defined without dependencies on XML 1.0.  I find this =
claim bizarre given that XML 1.0 is referenced normatively and the data =
data model definition is full of references to XML 1.0.

If there was no intent to define the data model without dependencies on =
XML 1.0, then at least half the
text in section 5 is pointless, unnecessary repetition of things that =
are obvious from the XML spec.
The choice seems to be between reading the spec as having a coherent =
goal which in some important
details it failed to achieve, and reading the spec as given to garrulous =
irrelevancies.

> The change proposal seems to be claiming that XPath 1.0 is full of =
bugs in need of correction because it does not meet a goal that it never =
had.
>=20
> The change proposal also claims that it is a goal of XPath 1.0 that =
the data model be defined formally.  This is clearly not the case. =20

By any mathematical standard, the prose of XPath 1.0 would count as =
informal.  But that is also
true of the prose in the change proposal.

Compared with other specs, I think the data model section of XPath 1.0 =
is more explicit and
formal than most.

> XPath 1.0 does not make the slightest attempt to be formal.  Rather it =
aims to be succinct and readily understandable.  The level of formality =
in the data model definition is similar to that of the rest of the spec =
and of companion specs (XML 1.0, XML Namespaces, XSLT 1.0).  It is also =
virtually impossible to be really rigorous about the construction of the =
data model from the XML document, without specifying this in the XML =
spec itself: for each syntax production the XML spec would need to =
explain how to corresponding data model was constructed.

I see no connection between the formality of an exposition and the title =
of the document
in which it appears.  =10The XML spec is not formal about (for example) =
identity criteria for
elements, because nothing in the XML spec appeals to element identity =
(at least, not=20
for the cases it leaves indeterminate).  The XPath 1.0 spec does need to =
be determinate
about element node identity, and it seems bizarre to me to suggest that =
it could not be
more precise or careful without changes to the text of the XML 1.0 spec.

>=20
> I am also not convinced that in many cases the proposed wording =
changes are in fact improvements.  If the WG does decide to go ahead =
with this change, I can make some more detailed comments.  But for the =
moment, I would just mention a couple of points.
>=20
> XPath 1.0 does not constrain the root node to have exactly one element =
child. In the case where the data model is constructed from an XML =
document, there will of course be exactly one child.  But in other cases =
(eg querying into a DOM DocumentFragment) it would be unhelpful to =
impose such a restriction.  (XPath 1.0 is generally fairly loose -- for =
example, it does not define conformance -- so as to provide maximum =
flexibility to referencing specs.)

Thank you for this clarification.

The XML spec seems, then, to be normative for the description of the =
data model, except for the
parts of it that don't apply.  On this view, the XPath 1.0 spec is =
readily understandable only for
readers gifted with a certain degree of clairvoyance.

>=20
> The reason why the spec uses terminology like "There is an element =
node for every element" instead of referencing particular productions is =
because of entity expansion.  For example, given
>=20
> <!DOCTYPE doc [
> <!ENTITY e "<x>foo</x>">
> ]>
> <doc>&e;&e;</doc>
>=20
> I am comfortable with saying (somewhat vaguely) that there are three =
elements. =20

The problem is that there is nothing in the =10XML spec or the Infoset =
spec that could be
used to argue that there are three elements here, instead of two.  Many =
people
are comfortable saying that there are three elements here, but a count =
of two elements
is equally compatible with the XML specification.

> I am much less comfortable saying that there are three occurrences of =
the "element" production (in fact, I would say it is clear that there =
are only two occurrences of the "element" production).

On the contrary; after entity expansion we have a sequence of character =
types
matching the document production of the XML spec, but we do not =
necessarily
have a sequence of character tokens matching the document production.  =
In
the sequence of character types, there are clearly three occurrences of =
strings
(sequences of character types) matching the element production, even =
though there
are only two such string-types.  That is the difference between a string =
type and
an occurrence of a string type.=20

--=20
****************************************************************
* C. M. Sperberg-McQueen, Black Mesa Technologies LLC
* http://www.blackmesatech.com=20
* http://cmsmcq.com/mib                =20
* http://balisage.net
****************************************************************