reg: Unicode Comformance - - - - U+FFFD

erra srikrishna <[email protected]> Mon, 17 Dec 2007 10:31:30 +0000 (GMT)
Newsgroups gmane.text.unicode.devel
Message-ID <[email protected]>
--0-67051089-1197887490=:68309
Content-Type: text/plain; charset=iso-8859-1
Content-Transfer-Encoding: 8bit

Hi all,
   
  i need a clarification regarding Replacement character U+FFFD.
   
  According to Unicode Conformance C4, C5 & C6, 
       If any non-characters, un-assigned & low or high surrogate codepoints are existed in unicode input then they should be skipped or Replaced with U+FFFD character
   
       According to Unicode Conformance Clause C12a, 
  An y Unicode (UTF8, UTF16 & UTF32) application should not accept ill-formed code unit sequences from its input. It should either signal an error or represent the code unit with a marker such as U+FFFD (REPLACEMENT CHARACTER).
   
  I am using IBM ICU and ICU uses FFFD as default replacement character. so i want to know if input itself contains U+FFFD character then how should we treat that character. 
   
  I mean i want my application to return an error whenever above sequences are found and ICU by default replaces with FFFD. so here i am checking input for FFFD and concluding that some invalid sequence has occured that's why ICU replaced it with FFFD then generating error.
   
  But this will not be applicable for input actually with FFFD then in this what to do. whether to generate error or anything else. i didn't see any conformance clause specifying what should be done for FFFD.
   
  Here i am mainly convernec with UTF16 input.
   
  Thanks
   
  Regards
  Srikrishna Erra


 
Krishna E


       
---------------------------------
 Now you can chat without downloading messenger. Click here to know how.
--0-67051089-1197887490=:68309
Content-Type: text/html; charset=iso-8859-1
Content-Transfer-Encoding: 8bit

<div>Hi all,</div>  <div>&nbsp;</div>  <div>i need a clarification regarding Replacement character U+FFFD.</div>  <div>&nbsp;</div>  <div>According to Unicode Conformance C4, C5 &amp; C6, </div>  <div>&nbsp;&nbsp;&nbsp;&nbsp; If any non-characters, un-assigned &amp; low or high surrogate codepoints are existed in unicode input then they should be skipped or Replaced with U+FFFD character</div>  <div>&nbsp;</div>  <div>&nbsp;&nbsp;&nbsp;&nbsp; According to Unicode Conformance Clause C12a, </div>  <div class=MsoNormal style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list .5in"><FONT size=3><FONT face="Times New Roman"><B>An y Unicode (UTF8, UTF16 &amp; UTF32)&nbsp;</B><SPAN style="mso-bidi-font-weight: bold">application<B> </B>should not accept ill-formed code unit sequences 
 from its input. It should either signal an error or represent the code unit with a marker such as U+FFFD (REPLACEMENT CHARACTER).</SPAN></FONT></FONT></div>  <div class=MsoNormal
 style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list .5in"><FONT size=3><FONT face="Times New Roman"><SPAN style="mso-bidi-font-weight: bold"></SPAN></FONT></FONT>&nbsp;</div>  <div class=MsoNormal style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list .5in"><FONT size=3><FONT face="Times New Roman"><SPAN style="mso-bidi-font-weight: bold"><?xml:namespace prefix = o ns = "urn:schemas-microsoft-com:office:office" /><o:p>I am using IBM ICU and ICU uses FFFD as default replacement character. so i want to know if input itself contains U+FFFD character then how should&nbsp;we treat that character. </o:p></SPAN></FONT></FONT></div>  <div class=MsoNormal style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list .5in"><FONT size=3><FONT face="Times New 
 Roman"><SPAN style="mso-bidi-font-weight: bold"><o:p></o:p></SPAN></FONT></FONT>&nbsp;</div>  <div class=MsoNormal style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list
 .5in"><FONT size=3><FONT face="Times New Roman"><SPAN style="mso-bidi-font-weight: bold"><o:p>I&nbsp;mean i want my application to return an error whenever above sequences are found and ICU by default replaces with FFFD. so here i am checking input for FFFD and concluding that some invalid sequence has occured that's why ICU replaced it with FFFD then generating error.</o:p></SPAN></FONT></FONT></div>  <div class=MsoNormal style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list .5in"><FONT size=3><FONT face="Times New Roman"><SPAN style="mso-bidi-font-weight: bold"><o:p></o:p></SPAN></FONT></FONT>&nbsp;</div>  <div class=MsoNormal style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list .5in"><FONT size=3><FONT face="Times New Roman"><SPAN style="mso-bidi-font-w
 eight: bold"><o:p>But this will not be applicable for input actually with FFFD then in this what to do. whether to generate error or anything else.&nbsp;i didn't see any conformance clause
 specifying what should be done for FFFD.</o:p></SPAN></FONT></FONT></div>  <div class=MsoNormal style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list .5in"><FONT size=3><FONT face="Times New Roman"><SPAN style="mso-bidi-font-weight: bold"><o:p></o:p></SPAN></FONT></FONT>&nbsp;</div>  <div class=MsoNormal style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list .5in"><FONT size=3><FONT face="Times New Roman"><SPAN style="mso-bidi-font-weight: bold"><o:p>Here i am mainly convernec with UTF16 input.</o:p></SPAN></FONT></FONT></div>  <div class=MsoNormal style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list .5in"><FONT size=3><FONT face="Times New Roman"><SPAN style="mso-bidi-font-weight: bold"><o:p></o:p></SPAN></FONT></FONT>&nbsp;</div>  <div cla
 ss=MsoNormal style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list .5in"><FONT size=3><FONT face="Times New Roman"><SPAN style="mso-bidi-font-weight:
 bold"><o:p>Thanks</o:p></SPAN></FONT></FONT></div>  <div class=MsoNormal style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list .5in"><FONT size=3><FONT face="Times New Roman"><SPAN style="mso-bidi-font-weight: bold"><o:p></o:p></SPAN></FONT></FONT>&nbsp;</div>  <div class=MsoNormal style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list .5in"><FONT size=3><FONT face="Times New Roman"><SPAN style="mso-bidi-font-weight: bold"><o:p>Regards</o:p></SPAN></FONT></FONT></div>  <div class=MsoNormal style="MARGIN: 0in 0in 0pt; mso-list: l0 level1 lfo1; tab-stops: list .5in"><FONT size=3><FONT face="Times New Roman"><SPAN style="mso-bidi-font-weight: bold"><o:p>Srikrishna Erra</o:p></SPAN></FONT></FONT></div><BR><BR><DIV>
<DIV><STRONG><FONT color=#0000ff></FONT></STRONG>&nbsp;</DIV>
<DIV><STRONG><FONT color=#0000ff>Krishna E</FONT></STRONG></DIV><IMG src="http://us.i1.yimg.com/us.yimg.com/i/mesg/tsmileys2/26.gif"></DIV><p>&#32;


      <!--5--><hr size=1></hr> Now you can chat without downloading messenger. <a href="http://in.rd.yahoo.com/tagline_webmessenger_5/*http://in.messenger.yahoo.com/webmessengerpromo.php">Click here</a> to know how.
--0-67051089-1197887490=:68309--