RE: missing  

"Richard Pineger" <[email protected]>
Newsgroups gmane.comp.web.html-tidy.user
Message-ID <BA002C6073F74171841990CA6DDB1E5D@Thomas>
Good for you David. One of those threads was mine. We've been working around
it by preprocessing to replace nbsp and other important entities with a
marker phrase, "nbspmarker" etc. and post processing to put them back in.

BTW - my humble thanks to Tidy programmers. It is a great tool that has
helped as part of conversions of html to XHTML and DITA XML for XML content
management systems.

Richard

-----Original Message-----
From: [email protected] [mailto:[email protected]] On Behalf
Of ml
Sent: 06 November 2007 13:18
To: Bjoern Hoehrmann
Cc: [email protected]
Subject: Re: missing &nbsp;


With default encodings it will produce <meta...charset=us-ascii> and all
nonascii chars will be converted to entities. Both is wrong. Why
"quote-nbsp:true" doesn't take place?

David


Bjoern Hoehrmann napsal(a):
> * ml wrote:
>> I have following config:
>>
>> bare:false
>> indent:false
>> output-html:true
>> doctype:transitional
>> hide-comments:true
>> wrap:0
>> quote-nbsp:true
>> quote-marks:false
>> quote-ampersand:true
>> break-before-br:false
>> char-encoding:raw
>> input-encoding:raw
>> output-encoding:raw
> 
>> I found few threads discussing it but no solution. Tested versions 
>> are Ubuntu package "HTML Tidy for Linux/x86 released on 1 September 2005"
>> and Pecl PHP Tidy 1.2.
>>
>> Any suggestion to preven Tidy from removing "&nbsp;"? I need UTF-8 
>> input/output.
> 
> Then why do you specify 'raw'? Can you reproduce this problem with 
> only the default options?
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.