Re: [patch] HTML parser bugfix

123 <[email protected]>
Newsgroups gmane.comp.web.dillo.devel
Message-ID <[email protected]>
On Thu, Jun 07, 2012 at 12:14:06AM +0100, Jeremy Henty wrote:
> 
> 123 wrote:
> 
> > On Sun, Jun 03, 2012 at 09:00:55PM +0100, Jeremy Henty wrote:
> 
> > > If you  don't detect and ignore  that extra double  quote you will
> > > break many pages that every other browser renders perfectly well.
> > 
> > Then there should be some logic for detecting double quotes.
> 
> I agree.  The hard question is: what logic?

IMO the right way to do it is to implement HTML Standard [1]. When
Html_write_raw is called, parser is in the Data state. When it returns
to Data state again, it is the end of token.

I have implemented standard comment/DOCTYPE parsing. It is incomplete,
EOF is not handled and DOCTYPE parsing is not changed. Patch is
attached. Next step is to rewrite tag parsing in standard way.

[1] http://www.whatwg.org/specs/web-apps/current-work/multipage/tokenization.html#tokenization
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.