Re: HTML Parser problem

"Shashank Kavishwar" <[email protected]> Thu, 10 Mar 2005 10:56:53 -0800
Newsgroups gmane.comp.lib.libwww
Message-ID <[email protected]>
Well, the '<' character can appear in HTML in no reference to a tag. For
e.g. I recently saw a signature as:

 

<')))><
Name
Title
Company
Direct Line
Mobile
e-mail:
web site
 

Shouldn't the '<' character, if its not an HTML tag be kept as it is
while parsing? Basically, I think the presence of a '<' character in the
HTML which is not part of a tag, should not break the parse function to
skip the rest of the HTML after the '<' character.

The Perl HTML::Parser has the capability to do this, and will parse the
HTML correctly.

 



FrontBridge introduces Message Archive and Secure Email. Get leading Enterprise Message Security services from FrontBridge. www.frontbridge.com.