Re: XML Schema classification help
"New, Cecil (GE Aviation, US)" <[email protected]>
| Newsgroups | gmane.comp.java.jdom.general |
|---|---|
| Message-ID | <8E2ABA60D9C51C4894395E4D3B212814039511B6@CINMLVEM15.e2k.ad.ge.com> |
Oracle even supports an xquery function named "XMLQUERY" - so in the
SELECT clause you can put something like this in:
, XMLQUERY(
'for $foo in //tag/@name
where $tag = $tagdata/@name
return $ tagdata'
PASSING ... as "colname" RETURNING CONTENT
)
From: [email protected]
[mailto:[email protected]] On Behalf Of Rolf Lear
Sent: Wednesday, January 04, 2012 2:54 PM
To: [email protected]
Subject: Re: [jdom-interest] XML Schema classification help
Hi Cliff.
I can't think of any magic 'short cut'.... and certainly, I do not think
JDOM will be the fastest/best way to 'classify' each document.
Things you should consider though:
- Using a plain SAX Parser (xmlreader) with a clever 'Entity Resolver'
may help you to quickly access what external URL's (probably XML
Schemas) are needed to resolve the document (although there is no
concept of an 'order' of schemas). This could help 'identify' the
document.
- Cutting short the parser (throw a SAX exception) would speed things up
once you have entered the main part of the document (startElement())
because you probably do not need to parse the whole document, just the
xsi schema-location references.
- Finally, depending on your database, you may already have a JRE
available in the database server ('big-brand databases mostly already
do, like DB2, Oracle, Sybase, etc.), in which case you can build a
'clever' Java function that evaluates the document *inside* the
database, and avoid creating a lot of external traffic.... for example,
you may be able to create a custom java-backed function 'xmlschemas()'
which returns the list of schemas in use in a document, and then you can
do something like:
select xmlschemas(xmldatacol) as schemas, count(*) from table group by
schemas
Rolf
On 04/01/2012 2:11 PM, cliff palmer wrote:
I need to examine XML documents contained in multiple columns in a
database table with over a million rows and identify each of the
different structures used for the XML data, producing a count if the
number of instances that use each structure.
I thought of using the SAXParser then creating a list of the XML headers
in the order used and storing each unique list and accumulating a count
based on matching an already encountered list object, but I am hoping
there is a less cumbersome approach.
I would appreciate any and all suggestions.
Thanks!
Cliff
_______________________________________________
To control your jdom-interest membership:
http://www.jdom.org/mailman/options/jdom-interest/[email protected]
_______________________________________________
To control your jdom-interest membership:
http://www.jdom.org/mailman/options/jdom-interest/[email protected]