Re: reg: Unicode Comformance - - - - U+FFFD

"Mark Davis" <[email protected]> Mon, 17 Dec 2007 13:32:04 -0800
Newsgroups gmane.text.unicode.devel
Message-ID <[email protected]>
------=_Part_16835_20301618.1197927124937
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

This should really be directed to the ICU list.

ICU gives you a choice of whether to get an error and stop, or whether to
substitute a character (and what that is). So if you want to check for
actual FFFD characters in the input stream (as opposed to those that are
replacements for erroneous or missing sequences), you have the tools to do
that.

Mark

On Dec 17, 2007 2:31 AM, erra srikrishna <[email protected]> wrote:

> Hi all,
>
> i need a clarification regarding Replacement character U+FFFD.
>
> According to Unicode Conformance C4, C5 & C6,
>      If any non-characters, un-assigned & low or high surrogate codepoints
> are existed in unicode input then they should be skipped or Replaced with
> U+FFFD character
>
>      According to Unicode Conformance Clause C12a,
> *An y Unicode (UTF8, UTF16 & UTF32) *application* *should not accept
> ill-formed code unit sequences from its input. It should either signal an
> error or represent the code unit with a marker such as U+FFFD (REPLACEMENT
> CHARACTER).
>
> I am using IBM ICU and ICU uses FFFD as default replacement character. so
> i want to know if input itself contains U+FFFD character then how should we
> treat that character.
>
> I mean i want my application to return an error whenever above sequences
> are found and ICU by default replaces with FFFD. so here i am checking input
> for FFFD and concluding that some invalid sequence has occured that's why
> ICU replaced it with FFFD then generating error.
>
> But this will not be applicable for input actually with FFFD then in this
> what to do. whether to generate error or anything else. i didn't see any
> conformance clause specifying what should be done for FFFD.
>
> Here i am mainly convernec with UTF16 input.
>
> Thanks
>
> Regards
> Srikrishna Erra
>
>
> **
> *Krishna E*
>
> ------------------------------
> Now you can chat without downloading messenger. Click here<http://in.rd.yahoo.com/tagline_webmessenger_5/*http://in.messenger.yahoo.com/webmessengerpromo.php>to know how.




-- 
Mark

------=_Part_16835_20301618.1197927124937
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

This should really be directed to the ICU list.<br><br>ICU gives you a choice of whether to get an error and stop, or whether to substitute a character (and what that is). So if you want to check for actual FFFD characters in the input stream (as opposed to those that are replacements for erroneous or missing sequences), you have the tools to do that.
<br><br>Mark<br><br><div class="gmail_quote">On Dec 17, 2007 2:31 AM, erra srikrishna &lt;<a href="mailto:[email protected]">[email protected]</a>&gt; wrote:<br><blockquote class="gmail_quote" style="border-left: 1px solid rgb(204, 204, 204); margin: 0pt 0pt 0pt 0.8ex; padding-left: 1ex;">
<div>Hi all,</div>  <div>&nbsp;</div>  <div>i need a clarification regarding Replacement character U+FFFD.</div>  <div>&nbsp;</div>  <div>According to Unicode Conformance C4, C5 &amp; C6, </div>  <div>&nbsp;&nbsp;&nbsp;&nbsp; If any non-characters, un-assigned &amp; low or high surrogate codepoints are existed in unicode input then they should be skipped or Replaced with U+FFFD character
</div>  <div>&nbsp;</div>  <div>&nbsp;&nbsp;&nbsp;&nbsp; According to Unicode Conformance Clause C12a, </div>  <div style="margin: 0in 0in 0pt;"><font size="3"><font face="Times New Roman"><b>An y Unicode (UTF8, UTF16 &amp; UTF32)&nbsp;</b><span>application
<b> </b>should not accept ill-formed code unit sequences from its input. It should either signal an error or represent the code unit with a marker such as U+FFFD (REPLACEMENT CHARACTER).</span></font></font></div>  <div style="margin: 0in 0in 0pt;">
<font size="3"><font face="Times New Roman"><span></span></font></font>&nbsp;</div>  <div style="margin: 0in 0in 0pt;"><font size="3"><font face="Times New Roman"><span>I am using IBM ICU and ICU uses FFFD as default replacement character. so i want to know if input itself contains U+FFFD character then how should&nbsp;we treat that character. 
</span></font></font></div>  <div style="margin: 0in 0in 0pt;"><font size="3"><font face="Times New Roman"><span></span></font></font>&nbsp;</div>  <div style="margin: 0in 0in 0pt;"><font size="3"><font face="Times New Roman">
<span>I&nbsp;mean i want my application to return an error whenever above sequences are found and ICU by default replaces with FFFD. so here i am checking input for FFFD and concluding that some invalid sequence has occured that&#39;s why ICU replaced it with FFFD then generating error.
</span></font></font></div>  <div style="margin: 0in 0in 0pt;"><font size="3"><font face="Times New Roman"><span></span></font></font>&nbsp;</div>  <div style="margin: 0in 0in 0pt;"><font size="3"><font face="Times New Roman">
<span>But this will not be applicable for input actually with FFFD then in this what to do. whether to generate error or anything else.&nbsp;i didn&#39;t see any conformance clause
 specifying what should be done for FFFD.</span></font></font></div>  <div style="margin: 0in 0in 0pt;"><font size="3"><font face="Times New Roman"><span></span></font></font>&nbsp;</div>  <div style="margin: 0in 0in 0pt;"><font size="3">
<font face="Times New Roman"><span>Here i am mainly convernec with UTF16 input.</span></font></font></div>  <div style="margin: 0in 0in 0pt;"><font size="3"><font face="Times New Roman"><span></span></font></font>&nbsp;</div>  
<div style="margin: 0in 0in 0pt;"><font size="3"><font face="Times New Roman"><span>Thanks</span></font></font></div>  <div style="margin: 0in 0in 0pt;"><font size="3"><font face="Times New Roman"><span></span></font></font>
&nbsp;</div>  <div style="margin: 0in 0in 0pt;"><font size="3"><font face="Times New Roman"><span>Regards</span></font></font></div>  <div style="margin: 0in 0in 0pt;"><font size="3"><font face="Times New Roman"><span>Srikrishna Erra
</span></font></font></div><br><br><div>
<div><b><font color="#0000ff"></font></b>&nbsp;</div>
<div><b><font color="#0000ff">Krishna E</font></b></div><img></div><p> 


      </p><hr size="1"> Now you can chat without downloading messenger. <a href="http://in.rd.yahoo.com/tagline_webmessenger_5/*http://in.messenger.yahoo.com/webmessengerpromo.php" target="_blank">Click here</a> to know how.
</blockquote></div><br><br clear="all"><br>-- <br>Mark

------=_Part_16835_20301618.1197927124937--