Re: DOM Document.cloneNode(true), in-memory validation, and Attr.getSpecified()

"Jacob Kjome" <[email protected]>
Newsgroups gmane.text.xml.xerces-j.user
Message-ID <[email protected]>
I see.  So, once the attributes are set as specified during cloning, there's 
no going back, even after in-memory validation since it will have been as if 
the parser had found the attributes explicitly specified in the original XML.  
As such, the validation can't mark them as unspecified since, from all 
appearances, they were specified in the XML.  Correct?

What about UserDataHandlers?  Would it be possible, at parse time, to add them 
to Attr nodes that are unspecified so that when the clone occurs (setting 
these attributes to specified), the UserDataHandlers can reset these 
attributes back to unspecified?  I haven't looked into how that would work, 
but it would be nice to know whether to bother researching it or not.

The documentURI issue is much easier to deal with, since I can copy the value 
prior to the clone and reset it manually after the clone.

While I'm at it, I figure I'll lobby for application of a patch I submitted a 
while back related to cloning and Attr ID'ness...
https://issues.apache.org/jira/browse/XERCESJ-1597

The following issue also seems somewhat related...
https://issues.apache.org/jira/browse/XERCESJ-1430


Thanks,

Jake

On Wed, 28 Aug 2013 19:02:05 +0530
 Mukul Gandhi <[email protected]> wrote:
> Here are my opinion for these questions. Sorry I couldn't do any tests as
> yet.
> 
> On Wed, Aug 28, 2013 at 10:28 AM, Jacob Kjome <[email protected]> wrote:
> 
>> However, after cloning the document (Document = (Document)
>> doc.cloneNode(true)) all the equivalent cloned Attrs show up as
>> "Attr.getSpecified() == true".  Is that to be expected or a bug?
> 
> 
> I think this is an expected behavior. The cloneNode() documentation at,
> http://xerces.apache.org/xerces2-j/javadocs/api/org/w3c/dom/Node.htmlmentions
> 
> "In addition, clones of unspecified Attr nodes are specified". This should
> clarify this doubt I think.
> 
> And a side-question: should I expect the value of Document.getDocumentURI()
>> be preserved after cloning?  Because that value appears to be discarded.
>>
> 
> The cloneNode() spec doesn't mention any constraints wrt this aspect. So
> if, after a node clone a document URI is not preserved, I would say this is
> compliant to the spec of this method. Other than, going by the wordings of
> the spec, I think a clone is happening on an object representation of the
> original document and objects do not have a URI, therefore a clone
> operation not preserving the document URI looks acceptable to me.
> 
> 
>> I'd like to be able to preserve the "specified" state in the clone so that
>> when the document is serialized, I can avoid printing these attributes by
>> skipping those with "specified == false".  Seems to me this should be
>> possible.  If not, why not?
>>
> 
> I think, this is a sensible use case. But the default implementation
> provided, doesn't make it possible for you. This might be addressed by an
> external application logic. You may keep the "specified" state of
> attributes external to your main document (or a more challenging way may be
> to, read it directly from DTD), and during serialization use this extra
> information to affect the serialization in your desired way. Another
> approach may be to, let the serializer emit the unspecified attributes and
> then use this output with a DTD or an external metadata as a unit for your
> application.
> 
> 
> 
>>
>> --
>> Regards,
>> Mukul Gandhi <[email protected]>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.