Re: XML Schema classification help
Rolf Lear <[email protected]>
| Newsgroups | gmane.comp.java.jdom.general |
|---|---|
| Message-ID | <[email protected]> |
Hi Cliff. JDOM cannot generate an XSD for a document (interesting idea, but very complicated.... like, how would it set maxOccurs? ... and in your use case that would be significant....) The best I can suggest is that you will need to do a 'deep inspection' of the XML, create your own sort of 'fingerprint' for the document, and then use that. JDOM could possibly be useful because it makes the inspection part a whole lot easier than building a SAX ContentHandler, etc (but at the price of some speed and some memory). Once you have built the JDOM document you can run all sorts of functions on the data to create the 'fingerprint'. Again, this could potentially be done inside the database to be more efficient. Unfortunately (for you), this is not something that I think there is an easy, or preexisting solution for (nothing comes to mind). Also, as Michael says, you need to build up your 'taxanomical' (nice word, Michael) rules, and in a 'real world' instance, you should be namespace aware, etc. Again, JDOM can help with that.... but only as a part of a bigger solution. Rolf If you need to do 'deep inspection' of the XML to determine it's On 04/01/2012 4:00 PM, cliff palmer wrote: > Rolf it's a little more involved than reading the XSI refs. I need to > look at the nodes. > > <a> <b> <c> </c> </b> </a> and <a> <b> <c> </c> <c> </c> </b> </a> are > considered the same structure for what I am doing because repeating > nodes aren't considered a difference. > <a> <b> <c> </c> </b> </a> and <a> <b> <c> </c> </b> <b> <c> </c> <c> > </c> </b> </a> are considered the same structure for what I am doing > because repeating groups aren't considered a difference. > <a> <b> <c> </c> </b> </a> and <a> <b> </b> </a> are not the same > because <c> is missing in the 2nd case so it does not contain all the > elements as the 1st case. > > I had hoped (apparently it's just a hope) that JDOM could generate an > XSD from an XML DOM object. > > More ideas are welcome :) > > Cliff > > > On Wed, Jan 4, 2012 at 2:53 PM, Rolf Lear <[email protected] > <mailto:[email protected]>> wrote: > > Hi Cliff. > > I can't think of any magic 'short cut'.... and certainly, I do not > think JDOM will be the fastest/best way to 'classify' each document. > > Things you should consider though: > - Using a plain SAX Parser (xmlreader) with a clever 'Entity > Resolver' may help you to quickly access what external URL's > (probably XML Schemas) are needed to resolve the document > (although there is no concept of an 'order' of schemas). This > could help 'identify' the document. > - Cutting short the parser (throw a SAX exception) would speed > things up once you have entered the main part of the document > (startElement()) because you probably do not need to parse the > whole document, just the xsi schema-location references. > - Finally, depending on your database, you may already have a JRE > available in the database server ('big-brand databases mostly > already do, like DB2, Oracle, Sybase, etc.), in which case you can > build a 'clever' Java function that evaluates the document > *inside* the database, and avoid creating a lot of external > traffic.... for example, you may be able to create a custom > java-backed function 'xmlschemas()' which returns the list of > schemas in use in a document, and then you can do something like: > > select xmlschemas(xmldatacol) as schemas, count(*) from table > group by schemas > > Rolf > > > On 04/01/2012 2:11 PM, cliff palmer wrote: >> I need to examine XML documents contained in multiple columns in >> a database table with over a million rows and identify each of >> the different structures used for the XML data, producing a count >> if the number of instances that use each structure. >> >> I thought of using the SAXParser then creating a list of the XML >> headers in the order used and storing each unique list and >> accumulating a count based on matching an already encountered >> list object, but I am hoping there is a less cumbersome approach. >> >> I would appreciate any and all suggestions. >> >> Thanks! >> Cliff >> >> >> _______________________________________________ >> To control your jdom-interest membership: >> http://www.jdom.org/mailman/options/jdom-interest/[email protected] > > > _______________________________________________ > To control your jdom-interest membership: > http://www.jdom.org/mailman/options/jdom-interest/[email protected] > > _______________________________________________ To control your jdom-interest membership: http://www.jdom.org/mailman/options/jdom-interest/[email protected]