Parsing HTML with SAX?

"Jakob Heinemann" <[email protected]> Sun, 21 Sep 2003 19:57:17 +0100
Newsgroups gmane.text.xml.sax.general
Message-ID <[email protected]>
Hello, 

just realized that sax parsing i fun, but it wont eat my old html-files. So my effors are either to convert all html to xml and do my data extraction from the xml, or try to overcome the initial limitations of SAX, being designed to fail for things like &nbsp and <link> tags with no terminal tag. Is there a way to have Sax parser be a bit more relaxed? I did try overriding the error and fatalerror, but it seems to be the wrong approach. Would it be possible to use something like JTidy together with Sax? 

Best Regards
/Jakob

-- 
__________________________________________________________
Sign-up for your own personalized E-mail at Mail.com
http://www.mail.com/?sr=signup

CareerBuilder.com has over 400,000 jobs. Be smarter about your job search
http://corp.mail.com/careers



-------------------------------------------------------
This sf.net email is sponsored by:ThinkGeek
Welcome to geek heaven.
http://thinkgeek.com/sf
_______________________________________________
List: [email protected]
See:  http://www.saxproject.org
https://lists.sourceforge.net/lists/listinfo/sax-users