Re: How to exclude DTD and namespace during Sax (TagSoup) parsing in JDOM

Laurent Bihanic <[email protected]>
Newsgroups gmane.comp.java.jdom.general
Message-ID <[email protected]>
Hi Jack,

Some answers below,

Laurent

Jack Bush wrote:
> I am having difficulty parsing using Saxon and TagSoup parser on a 
> namespace html document. The relevant content of this document are as 
> follows:
> 
> <!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional/ /EN" 
> "http://www. w3.org/TR/ xhtml1/DTD/ xhtml1-transitio nal.dtd">
> 
> <html xmlns="http: //www.w3. org/1999/ xhtml">
> <head>
> <meta http-equiv=" Content-Type" content="text/ html; 8" />
> </head>
> <body>
>     <div id="container">
>         <div id="content">
>             <table class="sresults">
>                 <tr>
>                     <td>
>                         <a href="http:/ /www.abc. com/areas" title=" 
> Hollywood , CA "> hollywood </a>
>                     </td>
...

> Below is the relevant code snippets illustrates how I have attempted to 
> retrieve the contents (value of  <a>):
>              import java.util.*;
>              import org.jdom.*;
>              import org.jdom.xpath. *;
>              import org.saxpath. *;
>              import org.ccil.cowan. tagsoup.Parser;
> 
> ( 1 )       frInHtml = new FileReader(" C:\\Tmp\\ ABC.html" );
> ( 2 )       brInHtml = new BufferedReader( frInHtml) ;
> ( 3 )       SAXBuilder saxBuilder = new SAXBuilder(" org.ccil. 
> cowan.tagsoup. Parser");
> ( 4 )       org.jdom.Document jdomDocument = saxbuilder.build( brInHtml) ;
> ( 5 )       XPath xpath =  XPath.newInstance( "/ns:html/ ns:body/ns: 
> div[@id=' container' ]/ns:div[ @id='content' ]/ns:table[ 
> @class='sresults ']/ns:tr/ ns:td/ns: a");
> ( 6 )       xpath.addNamespace( "ns", "http://www. w3.org/1999/ xhtml");
...

>  I would like to achieve the following objectives if possible:
> 
>  ( i ) Exclude DTD and namespace in order to simplifying the parsing 
> process. How this could be done?

You can't exclude the DTD. It is required in case the HTML document uses 
entities not defined by XML such as &nbsp; &agrave; etc.

I don't think Tagsoup will let you remove the namespace. It does the best 
possible thing by making it the default namespace. But XPath does not support 
any default (i.e. unprefixed) namespace, you have to use a prefix.

> ( ii ) Failing to exlude DTD, how to change the lookup of a PUBLIC DTD 
> to a local SYSTEM one and include a local DTD for reference?

You could try attaching a SAX EntityResolver to SAXBuilder that will pass it 
on to Tagsoup. Tagsoup will use it to resolve the DOCTYPE.

>  I am running JDK 1.6.0_06, Netbeans 6.1, JDOM 1.1, Saxon6-5-5, Tagsoup 
> 1.2 on Windows XP platform.
> 
>  Any assistance would be appreciated.
> 
>  Thanks in advance,
> 
>  Jack
_______________________________________________
To control your jdom-interest membership:
http://www.jdom.org/mailman/options/jdom-interest/[email protected]
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.