Re: Odp: Pd: Missing legacy Arabic encoding

Philippe Verdy via Unicode <[email protected]> Wed, 6 May 2026 15:05:16 +0200
Newsgroups gmane.text.unicode.general
Message-ID <CAGa7JC3FnsRKBQcseO2zciKRnfo=Gcq4PfXQkHA-x4SpoW2Mqg@mail.gmail.com>
--000000000000689b4a065125d119
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

You actually don't need any new compatibility characters for Arabic
contextual forms, or for other contextual forms in other joining scripts
(like Adlam, or even Mongolian whichbis a LTR script).

You just have to prepend or append a ZWJ or ZWNJ formatting control to the
unified letter if you want to override its default contextual presentation
form.


Le mar. 5 mai 2026, 00:48, Asmus Freytag via Unicode <
[email protected]> a =C3=A9crit :

> The issue at hand is the distinction between a theoretical gap and
> real-life problem.
>
> You have demonstrated that there are specifications that, if chained in
> the right way, can lead to ambiguities or gaps in interchange.
>
> What we don't have is an actual use case with real-life consequences for =
a
> set of existing users, not hypothetical ones.
>
> When it comes to encoding decisions based on existing documents, there is
> a strong presumption that once sufficiently many documents exist that
> contain a character, that this character will be needed in digitizing the=
se
> documents, whether immediately, or eventually (e.g. in the case of future
> scholarly studies). Also, the texts themselves exist, barring accidents, =
in
> permanence. Therefore, it is justified to consider irrevocably allocating=
 a
> character that will map to this source in perpetuity, even though each
> encoded character carries a small cost for implementers.
>
> However, when it comes to legacy characters, there's an additional cost
> that is imposed, and that is based on the fact that characters that are
> encoded solely for compatibility will usually violate one or more of the
> other encoding principles, something that incrementally complicates the
> standard. Even for people who never intend to use that character.
>
> Therefore, the SEW is on solid ground when it demands not only a
> hypothetical scenario, but evidence of actual impact on actual users. Not
> only whether some application could invoke an API, but whether such
> applications exist and are used today to access documents encoded using t=
he
> legacy characters in a way that is compromised irreparably by not having =
an
> encoding for them.
>
> A./
>
>
> On 5/4/2026 10:24 AM, [email protected] via Unicode wrote:
>
> In UTC 187 Minutes, "Asmus Freytag noted that the fact that lists of
> things existed in the past does not make these things plain text. Ned
> Holbrook pointed out that the purported issue occurs in a closed system,
> not in public interchange.". However, the arguments in the proposal do
> not merely hinge on the encodings being lists of characters, but
> specifically points out methods to interchange text, including an example
> of copying terminal output and pasting to Notepad, where the copying
> invokes the mapping of the current terminal codepage to UCS-2 (as is
> CHAR_INFO compatible) and the pasting writes it into plain text. Win32 is
> also not a closed system, as Win32 can capture the tiles of the output of
> Windows 3.1 Arabic DOS/Win16 programs and Windows 95/98/ME Arabic
> DOS/Win16/Win32 programs, but Win32 can also interact with public text
> interchange systems by reading and writing to files and network. I'm not
> saying that Unicode absolutely must include those characters, but those
> kinds of misleading claims are causing users to misunderstand what the
> proposal is about, and I don't want Unicode to be relying on uninformed
> decisions to evaluate proposals.
>
>
> *Dnia 18 kwietnia 2026 13:36* [email protected] via Unicode
> <[email protected]> < [email protected] >
> <[email protected]> napisa=C5=82(a):
>
> The SEW subsequently explained that the actual reason is due to
> insufficient evidence of user community that would need to use the
> resulting mapping. Despite Win32 being a highly popular platform with
> plenty of backwards compatibility and native UCS-2 terminal support, the
> specific use cases of installing codepages into Windows NT and using
> terminal tiles from Windows 3.1/95/98/ME are not sufficiently documented,
> making it difficult for any user communities to form around it. So it see=
ms
> like the idea of standardizing legacy Arabic terminal BMP mappings is a
> dead end for now.
>
>
> *Dnia 17 kwietnia 2026 22:59* [email protected] via Unicode
> <[email protected]> < [email protected] >
> <[email protected]> napisa=C5=82(a):
>
> The Recommendations in L2/26-100 claim that Microsoft's documentation of
> legacy Arabic encodings is available at
> https://learn.microsoft.com/en-us/typography/legacy/legacy_arabic_fonts.
> However, that article only demonstrates two encodings of TrueType fonts,
> which are used in Windows 3.1 but are completely different from the eight
> terminal encodings. Unlike the TrueType encodings which represent interna=
l
> shaping mappings and are not used for text interchange, the terminal
> encodings have been demonstrated to be directly used in text interchange
> through int 10h and ReadConsoleOutputA/WriteConsoleOutputA as already
> demonstrated in L2/26-077. The Recommendations also claim that the propos=
al
> does not demonstrate any need for interchange or encoding, but the propos=
al
> actually demonstrated such a need due to the logical extension of the Win=
32
> terminal API to the functions ReadConsoleOutputW/WriteConsoleOutputW, whi=
ch
> are in Windows NT and may be used on the output of previously ran program=
s
> (including those that used the legacy Arabic terminal encodings), which
> given the CHAR_INFO structure, therefore implies a need for all the tiles
> to map to BMP for interchange. I'm not objecting to the SEW's conclusion =
of
> "Users are expected to use PUA.", which can indeed be used to provide a
> mapping even if not standardized, but the reasoning given was flawed.
>
>
> *Dnia 09 stycznia 2026 17:25* [email protected] < [email protected]
> > <[email protected]> napisa=C5=82(a):
>
> The following Win32 C code will output 256 characters in system console
> codepage into the character grid, capture those character tiles in UCS-2 =
if
> possible, and then output the current console codepage number.
>
>
> #include <windows.h>
> #include <stdio.h>
> int main(){
> HANDLE hConsole=3DGetStdHandle(STD_OUTPUT_HANDLE);
> CHAR_INFO screen[256];
> COORD size=3D{16,16,};
> COORD pos=3D{0,0,};
> SMALL_RECT rect=3D{0,0,15,15,};
> for(int i=3D0;i<256;i++){
> screen[i].Attributes=3D0xF0;
> screen[i].Char.AsciiChar=3Di;
> }
> WriteConsoleOutputA(hConsole,screen,size,pos,&rect);
> CHAR_INFO screenu[256];
> if(ReadConsoleOutputW(hConsole,screenu,size,pos,&rect)){
> for(int i=3D0;i<256;i++) printf("%04X ",screenu[i].Char.UnicodeChar);
> }
> else{
> printf("error %08X\n",GetLastError());
> }
> printf("codepage %u",GetConsoleOutputCP());
> }
>
> In most cases, whenever a legacy Win32 codepage is used, the application
> can run on Windows NT to capture the UCS-2 mapping of those character cel=
ls
> to the BMP (although for CJK codepages a more complex setup would be
> necessary due to thousands of fullwidth characters with 2-byte sequences)=
.
>
>
> However, in Arabic versions of Windows 9x (95/98/ME) the resulting
> character set has many presentation forms that are not in Unicode. This i=
s
> the result when running on Windows ME: https://i.imgur.com/QFm3SkI.png in
> 10=C3=9720 font, https://i.imgur.com/KUbLQ0A.png in 10=C3=9718 font (same=
 result
> also appears in Windows 95/98). 5=C3=9712, 7=C3=9712, 8=C3=9712, 10=C3=97=
18, 10=C3=9720, and 12=C3=9716
> bitmap fonts have been attested with that character set (VGAOEM.FON,
> 8514OEM.FON, DOSAPP.FON). The 10=C3=9720 font has slightly different mapp=
ing
> than the other sizes: 0x93 is =C3=B6 instead of =C3=B4, and 0x97 is missi=
ng (causing
> the following characters on the same line to be drawn at the wrong
> position). It also claims to be using codepage 720, but many characters
> differ from their CP720 mappings, including the bundled CP_720.NLS mappin=
gs
> (for example, =D9=80 (U+0640 ARABIC TATWEEL) is 0x95 in CP720, but in the
> console 0x95 is =D8=B4 instead, and the tatweel is at 0xFF). On Windows
> 9x, ReadConsoleOutputW is not supported so the UCS-2 mappings of the
> console character tiles cannot be captured (error 0x00000078
> ERROR_CALL_NOT_IMPLEMENTED).
>
>
> When that program runs on Arabic versions of Windows NT, the visual outpu=
t
> is of the CP437 character set if one of the bundled bitmap fonts is used =
(
> https://i.imgur.com/RxjtxMH.png), or the CP720 set if Lucida Console is
> used, with the Arabic letters either having glitchy font substitution (NT
> 4.0, NT 5.0/2000) or the .notdef glyph (NT 5.1/XP and up). In fact, it
> seems that the only Arabic bitmap fonts that occur in Windows NT are CP12=
56
> fonts, which are not used in terminals. So this appears to be one of thos=
e
> permanent Windows compatibility regressions that occured when Windows 9x
> ended, where the terminals can no longer render legacy Arabic text. Even =
if
> the user managed to use registry hacks to set the font to Courier New or
> Simplified Arabic Fixed, it would still use the CP720 mapping which is no=
t
> compatible with the Windows 9x set.
>
>
> It appears that in the Windows 9x Arabic terminal character set, 244
> characters (=E2=80=87=EF=BA=80=EF=BA=81=EF=BA=82=EF=BA=83=EF=BA=84=EF=BA=
=85=EF=BA=87=EF=BA=88=EF=BA=8A=EF=BA=8B=EF=BA=8D=EF=BA=8E=EF=BA=8F=EF=BA=91=
=EF=BA=93=E2=96=BA=E2=97=84=E2=86=95=EF=BA=95=C2=B6=C2=A7=EF=BA=97=EF=BA=99=
=E2=86=91=E2=86=93=E2=86=92=E2=86=90=EF=BA=9B=EF=B9=B0=E2=96=B2=E2=96=BC
> !"#$%&'()*+,-./0123456789:;<=3D>?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\]^_`abcdefg=
hijklmnopqrstuvwxyz{|}~=EF=BA=9D=EF=BA=9F=EF=BA=A1=C3=A9=C3=A2=EF=BA=A3=C3=
=A0=EF=BA=A5=C3=A7=C3=AA=C3=AB=C3=A8=C3=AF=C3=AE=EF=BA=A7=EF=BA=A9=EF=BA=AB=
=EF=BA=AD=EF=BA=AF=C3=B4=EF=BA=B3=C3=BB=C3=B9=EF=BA=B7=EF=BA=BB=C2=A3=EF=BA=
=BF=EF=BB=81=EF=BB=85=EF=BB=89=EF=BB=8A=EF=BB=8B=EF=BB=8C=EF=BB=8D=EF=BB=8E=
=EF=BB=8F=EF=BB=90=EF=BB=91=EF=BB=93=EF=BB=95=EF=BB=97=EF=BB=99=EF=BB=9B=C2=
=AB=C2=BB=EF=B9=B1=E2=96=92=EF=B9=B2=E2=94=82=E2=94=A4=EF=B9=B4=EF=B9=B6=EF=
=B9=B7=EF=B9=B8=D9=A0=D9=A1=D9=A2=D9=A3=EF=B9=B9=EF=B9=BA=E2=94=90=E2=94=94=
=E2=94=B4=E2=94=AC=E2=94=9C=E2=94=80=E2=94=BC=EF=B9=BB=EF=B9=BE=D9=A4=D9=A5=
=D9=A6=D9=A7=D9=A8=D9=A9=D8=8C=EF=B9=BF=EF=B1=9E=EF=B1=9F=EF=B1=A0=EF=B3=B2=
=EF=B1=A1=EF=B3=B3=EF=B1=A2=E2=94=98=E2=94=8C=D8=9B=D8=9F=C2=A4=EF=BB=9D=EF=
=BB=9F=EF=BB=A1=EF=BB=A3=EF=BB=A5=EF=BB=A7=C2=B5=EF=BB=A9=EF=BB=AB=EF=BB=AC=
=EF=BB=AD=EF=BB=AF=EF=BB=B0=EF=BB=B1=EF=BB=B2=EF=BB=B3=EF=B3=B4=EF=B9=BC=EF=
=B9=BD=EF=BA=B1=EF=BA=B5=EF=BA=B9=EF=BA=BD=EF=B9=B3=C2=B0=C2=B7=E2=96=A0=D9=
=80)
> are already in Unicode, but 12 characters are not in Unicode:
>
> =E2=80=A2 6 of them are pieces of lam-alef ligatures (0xDD, 0xDE, 0xF9, 0=
xFB,
> 0xFC, 0xFD)
>
> =E2=80=A2 2 of them are shadda with fathatan ligatures without or with ta=
tweel
> (0xD0, 0xD1)
>
> =E2=80=94 in some legacy Microsoft fonts, shadda with fathatan is mapped =
to
> private use U+E818
>
> =E2=80=A2 4 of them are disunifications of seen/sheen/sad/dad occuring ei=
ther with
> or without tail
>
> =E2=80=94 =EF=B9=B3 (U+FE73 ARABIC TAIL FRAGMENT) was originally encoded =
in Unicode 3.2
> for CP864 compatibility; in that codepage, the forms of seen/sheen/sad/da=
d
> attach to the tail fragment
>
> =E2=80=94 forms with included tail: 0x92, 0x95, 0x98, 0x8A
>
> =E2=80=94 forms without tail (attaching to tail fragment like in CP864): =
0xF3,
> 0xF4, 0xF5, 0xF6
>
>
> If someone tried to make a Win32 console implementation and tried to
> implement both Windows 9x Arabic terminal character set compatibility and
> wide string API (ReadConsoleOutputW) compatibility simultaneously, then
> they would run into the issue that there is currently no standardized
> mapping to handle that scenario. What should Windows 9x Arabic console
> compatible implementations do in that case?
>
>
>
>
>
>
>

--000000000000689b4a065125d119
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

<div dir=3D"auto">You actually don&#39;t need any new compatibility charact=
ers for Arabic contextual forms, or for other contextual forms in other joi=
ning scripts (like Adlam, or even Mongolian whichbis a LTR script).<div dir=
=3D"auto"><br></div><div dir=3D"auto">You just have to prepend or append a =
ZWJ or ZWNJ formatting control to the unified letter if you want to overrid=
e its default contextual presentation form.</div><div dir=3D"auto"><br></di=
v></div><br><div class=3D"gmail_quote gmail_quote_container"><div dir=3D"lt=
r" class=3D"gmail_attr">Le mar. 5 mai 2026, 00:48, Asmus Freytag via Unicod=
e &lt;<a href=3D"mailto:[email protected]">[email protected]<=
/a>&gt; a =C3=A9crit=C2=A0:<br></div><blockquote class=3D"gmail_quote" styl=
e=3D"margin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex"><u></u>

 =20
   =20
 =20
  <div>
    <div>The issue at hand is the distinction
      between a theoretical gap and real-life problem.</div>
    <div><br>
    </div>
    <div>You have demonstrated that there are
      specifications that, if chained in the right way, can lead to
      ambiguities or gaps in interchange.</div>
    <div><br>
    </div>
    <div>What we don&#39;t have is an actual use
      case with real-life consequences for a set of existing users, not
      hypothetical ones.</div>
    <div><br>
    </div>
    <div>When it comes to encoding decisions
      based on existing documents, there is a strong presumption that
      once sufficiently many documents exist that contain a character,
      that this character will be needed in digitizing these documents,
      whether immediately, or eventually (e.g. in the case of future
      scholarly studies). Also, the texts themselves exist, barring
      accidents, in permanence. Therefore, it is justified to consider
      irrevocably allocating a character that will map to this source in
      perpetuity, even though each encoded character carries a small
      cost for implementers.<br>
      <br>
      However, when it comes to legacy characters, there&#39;s an additiona=
l
      cost that is imposed, and that is based on the fact that
      characters that are encoded solely for compatibility will usually
      violate one or more of the other encoding principles, something
      that incrementally complicates the standard. Even for people who
      never intend to use that character.<br>
      <br>
      Therefore, the SEW is on solid ground when it demands not only a
      hypothetical scenario, but evidence of actual impact on actual
      users. Not only whether some application could invoke an API, but
      whether such applications exist and are used today to access
      documents encoded using the legacy characters in a way that is
      compromised irreparably by not having an encoding for them.</div>
    <div><br>
    </div>
    <div>A./</div>
    <div><br>
    </div>
    <div><br>
    </div>
    <div>On 5/4/2026 10:24 AM,
      <a href=3D"mailto:[email protected]" target=3D"_blank" rel=3D"nore=
ferrer">[email protected]</a> via Unicode wrote:<br>
    </div>
    <blockquote type=3D"cite">
     =20
      <p>In UTC 187 Minutes, &quot;<span style=3D"color:rgb(0,0,0);font-fam=
ily:&quot;DMCA Sans Serif 10.0 dev1&quot;;font-size:medium;font-style:norma=
l;font-variant-ligatures:normal;font-variant-caps:normal;font-weight:400;le=
tter-spacing:normal;text-align:start;text-indent:0px;text-transform:none;wh=
ite-space:normal;word-spacing:0px;text-decoration-style:initial;text-decora=
tion-color:initial;float:none;display:inline!important">Asmus
          Freytag noted that the fact that lists of things existed in
          the past does not make these things plain text. Ned Holbrook
          pointed out that the purported issue occurs in a closed
          system, not in public interchange.</span>&quot;. However, the
        arguments in the proposal do not merely hinge on the encodings
        being lists of characters, but specifically points out methods
        to interchange text, including an example of copying terminal
        output and pasting to Notepad, where the copying invokes the
        mapping of the current terminal codepage to UCS-2 (as is
        CHAR_INFO compatible) and the pasting writes it into plain text.
        Win32 is also not a closed system, as Win32 can capture the
        tiles of the output of Windows 3.1 Arabic DOS/Win16 programs and
        Windows 95/98/ME Arabic DOS/Win16/Win32 programs, but Win32 can
        also interact with public text interchange systems by reading
        and writing to files and network. I&#39;m not saying that Unicode
        absolutely must include those characters, but those kinds of
        misleading claims are causing users to misunderstand what the
        proposal is about, and I don&#39;t want Unicode to be relying on
        uninformed decisions to evaluate proposals.</p>
      <p><br>
      </p>
      <div>
        <blockquote style=3D"padding-top:12px">
          <p style=3D"padding-bottom:12px"><strong>Dnia 18 kwietnia 2026
              13:36</strong> <a href=3D"mailto:[email protected]" re=
l=3D"noopener noreferrer nofollow noreferrer" target=3D"_blank"><span style=
=3D"margin-left:4px">[email protected]
                via Unicode</span></a><span style=3D"margin-left:4px">
              <a href=3D"mailto:[email protected]" target=3D"_blank"=
 rel=3D"noreferrer">&lt; [email protected] &gt;</a></span> napisa=C5=
=82(a):</p>
          <div id=3D"m_5371084887079304582gwpbf2d884a">
            <div id=3D"m_5371084887079304582gwpbf2d884ah">
              <div>
                <p>The SEW subsequently explained that the actual reason
                  is due to insufficient evidence of user community that
                  would need to use the resulting mapping. Despite Win32
                  being a highly popular platform with plenty of
                  backwards compatibility and native UCS-2 terminal
                  support, the specific use cases of installing
                  codepages into Windows NT and using terminal tiles
                  from Windows 3.1/95/98/ME are not sufficiently
                  documented, making it difficult for any user
                  communities to form around it. So it seems like the
                  idea of standardizing legacy Arabic terminal BMP
                  mappings is a dead end for now.</p>
                <p><br>
                </p>
                <div>
                  <blockquote style=3D"padding-top:12px">
                    <p style=3D"padding-bottom:12px"><strong>Dnia 17
                        kwietnia 2026 22:59</strong> <a href=3D"mailto:unic=
[email protected]" rel=3D"noopener noreferrer nofollow noreferrer" targe=
t=3D"_blank"><span style=3D"margin-left:4px">[email protected]
                          via Unicode</span></a><span style=3D"margin-left:=
4px"> <a href=3D"mailto:[email protected]" target=3D"_blank" rel=3D"=
noreferrer">&lt;
                        [email protected] &gt;</a></span> napisa=C5=
=82(a):</p>
                    <div id=3D"m_5371084887079304582gwpbf2d884a_gwp2281a7f8=
">
                      <div id=3D"m_5371084887079304582gwpbf2d884a_gwp2281a7=
f8h">
                        <div>
                          <p>The Recommendations in L2/26-100 claim that
                            Microsoft&#39;s documentation of legacy Arabic
                            encodings is available at
                            <a href=3D"https://learn.microsoft.com/en-us/ty=
pography/legacy/legacy_arabic_fonts" target=3D"_blank" rel=3D"noreferrer">h=
ttps://learn.microsoft.com/en-us/typography/legacy/legacy_arabic_fonts</a>.
                            However, that article only demonstrates two
                            encodings of TrueType fonts, which are used
                            in Windows 3.1 but are completely different
                            from the eight terminal encodings. Unlike
                            the TrueType encodings which represent
                            internal shaping mappings and are not used
                            for text interchange, the terminal encodings
                            have been demonstrated to be directly used
                            in text interchange through int 10h and
                            ReadConsoleOutputA/WriteConsoleOutputA as
                            already demonstrated in L2/26-077. The
                            Recommendations also claim that the proposal
                            does not demonstrate any need for
                            interchange or encoding, but the proposal
                            actually demonstrated such a need due to the
                            logical extension of the Win32 terminal API
                            to the functions
                            ReadConsoleOutputW/WriteConsoleOutputW,
                            which are in Windows NT and may be used on
                            the output of previously ran programs
                            (including those that used the legacy Arabic
                            terminal encodings), which given the
                            CHAR_INFO structure, therefore implies a
                            need for all the tiles to map to BMP for
                            interchange. I&#39;m not objecting to the SEW&#=
39;s
                            conclusion of &quot;Users are expected to use
                            PUA.&quot;, which can indeed be used to provide=
 a
                            mapping even if not standardized, but the
                            reasoning given was flawed.</p>
                          <p><br>
                          </p>
                          <div>
                            <blockquote style=3D"padding-top:12px">
                              <p style=3D"padding-bottom:12px"><strong>Dnia
                                  09 stycznia 2026 17:25</strong> <span sty=
le=3D"margin-left:4px"><a href=3D"mailto:[email protected]" target=3D"_b=
lank" rel=3D"noreferrer">[email protected]</a>
                                  <a href=3D"mailto:[email protected]" t=
arget=3D"_blank" rel=3D"noreferrer">&lt; [email protected] &gt;</a></spa=
n>
                                napisa=C5=82(a):</p>
                              <div id=3D"m_5371084887079304582gwpbf2d884a_g=
wp2281a7f8_gwpa05276c7">
                                <div id=3D"m_5371084887079304582gwpbf2d884a=
_gwp2281a7f8_gwpa05276c7h">
                                  <div>
                                    <div id=3D"m_5371084887079304582gwpbf2d=
884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718">
                                      <div id=3D"m_5371084887079304582gwpbf=
2d884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718h">
                                        <div>
                                          <div id=3D"m_5371084887079304582g=
wpbf2d884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718_gwpa8b5f718">
                                            <div id=3D"m_537108488707930458=
2gwpbf2d884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718_gwpa8b5f718h">
                                              <div>
                                                <div id=3D"m_53710848870793=
04582gwpbf2d884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718_gwpa8b5f718_gwpa8b5f71=
8">
                                                  <div id=3D"m_537108488707=
9304582gwpbf2d884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718_gwpa8b5f718_gwpa8b5f=
718h">
                                                    <div>
                                                      <p>The following
                                                        Win32 C code
                                                        will output 256
                                                        characters in
                                                        system console
                                                        codepage into
                                                        the character
                                                        grid, capture
                                                        those character
                                                        tiles in UCS-2
                                                        if possible, and
                                                        then output the
                                                        current console
                                                        codepage number.<br=
>
                                                      </p>
                                                      <p><br>
                                                      </p>
                                                      <p>#include
                                                        &lt;windows.h&gt;<b=
r>
                                                        #include
                                                        &lt;stdio.h&gt;<br>
                                                        int main(){<br>
                                                        HANDLE
                                                        hConsole=3DGetStdHa=
ndle(STD_OUTPUT_HANDLE);<br>
                                                        CHAR_INFO
                                                        screen[256];<br>
                                                        COORD
                                                        size=3D{16,16,};<br=
>
                                                        COORD
                                                        pos=3D{0,0,};<br>
                                                        SMALL_RECT
                                                        rect=3D{0,0,15,15,}=
;<br>
                                                        for(int
                                                        i=3D0;i&lt;256;i++)=
{<br>
screen[i].Attributes=3D0xF0;<br>
screen[i].Char.AsciiChar=3Di;<br>
                                                        }<br>
WriteConsoleOutputA(hConsole,screen,size,pos,&amp;rect);<br>
                                                        CHAR_INFO
                                                        screenu[256];<br>
if(ReadConsoleOutputW(hConsole,screenu,size,pos,&amp;rect)){<br>
                                                        for(int
                                                        i=3D0;i&lt;256;i++)
                                                        printf(&quot;%04X
                                                        &quot;,screenu[i].C=
har.UnicodeChar);<br>
                                                        }<br>
                                                        else{<br>
                                                        printf(&quot;error
                                                        %08X\n&quot;,GetLas=
tError());<br>
                                                        }<br>
                                                        printf(&quot;codepa=
ge
%u&quot;,GetConsoleOutputCP());<br>
                                                        }<br>
                                                        <br>
                                                      </p>
                                                      <p>In most cases,
                                                        whenever a
                                                        legacy Win32
                                                        codepage is
                                                        used, the
                                                        application can
                                                        run on Windows
                                                        NT to capture
                                                        the UCS-2
                                                        mapping of those
                                                        character cells
                                                        to the BMP
                                                        (although for
                                                        CJK codepages a
                                                        more complex
                                                        setup would be
                                                        necessary due to
                                                        thousands of
                                                        fullwidth
                                                        characters with
                                                        2-byte
                                                        sequences).<br>
                                                      </p>
                                                      <p><br>
                                                      </p>
                                                      <p>However, in
                                                        Arabic versions
                                                        of Windows 9x
                                                        (95/98/ME) the
                                                        resulting
                                                        character set
                                                        has many
                                                        presentation
                                                        forms that are
                                                        not in Unicode.
                                                        This is the
                                                        result when
                                                        running on
                                                        Windows ME:=C2=A0<a=
 href=3D"https://i.imgur.com/QFm3SkI.png" rel=3D"noopener noreferrer norefe=
rrer" target=3D"_blank">https://i.imgur.com/QFm3SkI.png</a>=C2=A0in
                                                        10=C3=9720 font, <a=
 href=3D"https://i.imgur.com/KUbLQ0A.png" rel=3D"noopener noreferrer norefe=
rrer" target=3D"_blank">https://i.imgur.com/KUbLQ0A.png</a>=C2=A0in
                                                        10=C3=9718 font (sa=
me
                                                        result also
                                                        appears in
                                                        Windows 95/98).
                                                        5=C3=9712, 7=C3=971=
2,
                                                        8=C3=9712, 10=C3=97=
18,
                                                        10=C3=9720, and 12=
=C3=9716
                                                        bitmap fonts
                                                        have been
                                                        attested with
                                                        that character
                                                        set (VGAOEM.FON,
                                                        8514OEM.FON,
                                                        DOSAPP.FON). The
                                                        10=C3=9720 font has
                                                        slightly
                                                        different
                                                        mapping than the
                                                        other sizes:
                                                        0x93 is =C3=B6
                                                        instead of =C3=B4,
                                                        and 0x97 is
                                                        missing (causing
                                                        the following
                                                        characters on
                                                        the same line to
                                                        be drawn at the
                                                        wrong position).
                                                        It also claims
                                                        to be using
                                                        codepage 720,
                                                        but many
                                                        characters
                                                        differ from
                                                        their CP720
                                                        mappings,
                                                        including the
                                                        bundled=C2=A0CP_720=
.NLS
                                                        mappings (for
                                                        example, =D9=80
                                                        (U+0640 ARABIC
                                                        TATWEEL) is 0x95
                                                        in CP720, but in
                                                        the console 0x95
                                                        is =D8=B4 instead,
                                                        and the tatweel
                                                        is at 0xFF). On
                                                        Windows
                                                        9x,=C2=A0ReadConsol=
eOutputW
                                                        is not supported
                                                        so the UCS-2
                                                        mappings of the
                                                        console
                                                        character tiles
                                                        cannot be
                                                        captured (error
                                                        0x00000078
                                                        ERROR_CALL_NOT_IMPL=
EMENTED).<br>
                                                      </p>
                                                    </div>
                                                  </div>
                                                </div>
                                                <p><br>
                                                </p>
                                                <p>When that program
                                                  runs on Arabic
                                                  versions of Windows
                                                  NT, the visual output
                                                  is of the CP437
                                                  character set if one
                                                  of the bundled bitmap
                                                  fonts is used (<a href=3D=
"https://i.imgur.com/RxjtxMH.png" rel=3D"noopener noreferrer noreferrer" ta=
rget=3D"_blank">https://i.imgur.com/RxjtxMH.png</a>),
                                                  or the CP720 set if
                                                  Lucida Console is
                                                  used, with the Arabic
                                                  letters either having
                                                  glitchy font
                                                  substitution (NT 4.0,
                                                  NT 5.0/2000) or the
                                                  .notdef glyph (NT
                                                  5.1/XP and up). In
                                                  fact, it seems that
                                                  the only Arabic bitmap
                                                  fonts that occur in
                                                  Windows NT are CP1256
                                                  fonts, which are not
                                                  used in terminals. So
                                                  this appears to be one
                                                  of those permanent
                                                  Windows compatibility
                                                  regressions that
                                                  occured when Windows
                                                  9x ended, where the
                                                  terminals can no
                                                  longer render legacy
                                                  Arabic text. Even if
                                                  the user managed to
                                                  use registry hacks to
                                                  set the font to
                                                  Courier New or
                                                  Simplified Arabic
                                                  Fixed, it would still
                                                  use the CP720 mapping
                                                  which is not
                                                  compatible with the
                                                  Windows 9x set.<br>
                                                </p>
                                              </div>
                                            </div>
                                            <p><br>
                                            </p>
                                          </div>
                                        </div>
                                      </div>
                                    </div>
                                    <div id=3D"m_5371084887079304582gwpbf2d=
884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718">
                                      <div id=3D"m_5371084887079304582gwpbf=
2d884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718h">
                                        <div>
                                          <p>It appears that in the
                                            Windows 9x Arabic terminal
                                            character set, 244
                                            characters
                                            (=E2=80=87=EF=BA=80=EF=BA=81=EF=
=BA=82=EF=BA=83=EF=BA=84=EF=BA=85=EF=BA=87=EF=BA=88=EF=BA=8A=EF=BA=8B=EF=BA=
=8D=EF=BA=8E=EF=BA=8F=EF=BA=91=EF=BA=93=E2=96=BA=E2=97=84=E2=86=95=EF=BA=95=
=C2=B6=C2=A7=EF=BA=97=EF=BA=99=E2=86=91=E2=86=93=E2=86=92=E2=86=90=EF=BA=9B=
=EF=B9=B0=E2=96=B2=E2=96=BC
!&quot;#$%&amp;&#39;()*+,-./0123456789:;&lt;=3D&gt;?@ABCDEFGHIJKLMNOPQRSTUV=
WXYZ[\]^_`abcdefghijklmnopqrstuvwxyz{|}~=EF=BA=9D=EF=BA=9F=EF=BA=A1=C3=A9=
=C3=A2=EF=BA=A3=C3=A0=EF=BA=A5=C3=A7=C3=AA=C3=AB=C3=A8=C3=AF=C3=AE=EF=BA=A7=
=EF=BA=A9=EF=BA=AB=EF=BA=AD=EF=BA=AF=C3=B4=EF=BA=B3=C3=BB=C3=B9=EF=BA=B7=EF=
=BA=BB=C2=A3=EF=BA=BF=EF=BB=81=EF=BB=85=EF=BB=89=EF=BB=8A=EF=BB=8B=EF=BB=8C=
=EF=BB=8D=EF=BB=8E=EF=BB=8F=EF=BB=90=EF=BB=91=EF=BB=93=EF=BB=95=EF=BB=97=EF=
=BB=99=EF=BB=9B=C2=AB=C2=BB=EF=B9=B1=E2=96=92=EF=B9=B2=E2=94=82=E2=94=A4=EF=
=B9=B4=EF=B9=B6=EF=B9=B7=EF=B9=B8=D9=A0=D9=A1=D9=A2=D9=A3=EF=B9=B9=EF=B9=BA=
=E2=94=90=E2=94=94=E2=94=B4=E2=94=AC=E2=94=9C=E2=94=80=E2=94=BC=EF=B9=BB=EF=
=B9=BE=D9=A4=D9=A5=D9=A6=D9=A7=D9=A8=D9=A9=D8=8C=EF=B9=BF=EF=B1=9E=EF=B1=9F=
=EF=B1=A0=EF=B3=B2=EF=B1=A1=EF=B3=B3=EF=B1=A2=E2=94=98=E2=94=8C=D8=9B=D8=9F=
=C2=A4=EF=BB=9D=EF=BB=9F=EF=BB=A1=EF=BB=A3=EF=BB=A5=EF=BB=A7=C2=B5=EF=BB=A9=
=EF=BB=AB=EF=BB=AC=EF=BB=AD=EF=BB=AF=EF=BB=B0=EF=BB=B1=EF=BB=B2=EF=BB=B3=EF=
=B3=B4=EF=B9=BC=EF=B9=BD=EF=BA=B1=EF=BA=B5=EF=BA=B9=EF=BA=BD=EF=B9=B3=C2=B0=
=C2=B7=E2=96=A0=D9=80)
                                            are already in Unicode, but
                                            12 characters are not in
                                            Unicode:<br>
                                          </p>
                                          <p>=E2=80=A2 6 of them are pieces=
 of
                                            lam-alef ligatures (0xDD,
                                            0xDE, 0xF9, 0xFB, 0xFC,
                                            0xFD)<br>
                                          </p>
                                          <p>=E2=80=A2 2 of them are shadda=
 with
                                            fathatan ligatures without
                                            or with tatweel (0xD0, 0xD1)<br=
>
                                          </p>
                                          <p>=E2=80=94 in some legacy Micro=
soft
                                            fonts, shadda with fathatan
                                            is mapped to private use
                                            U+E818<br>
                                          </p>
                                          <p>=E2=80=A2 4 of them are
                                            disunifications of
                                            seen/sheen/sad/dad occuring
                                            either with or without tail<br>
                                          </p>
                                          <p>=E2=80=94=C2=A0=EF=B9=B3 (U+FE=
73 ARABIC TAIL
                                            FRAGMENT) was originally
                                            encoded in Unicode 3.2 for
                                            CP864 compatibility; in that
                                            codepage, the forms
                                            of=C2=A0seen/sheen/sad/dad atta=
ch
                                            to the tail fragment<br>
                                          </p>
                                          <p>=E2=80=94 forms with included
                                            tail:=C2=A00x92, 0x95, 0x98, 0x=
8A<br>
                                          </p>
                                          <p>=E2=80=94 forms without tail
                                            (attaching to tail fragment
                                            like in CP864):=C2=A00xF3, 0xF4=
,
                                            0xF5, 0xF6<br>
                                          </p>
                                        </div>
                                        <p><br>
                                        </p>
                                      </div>
                                      <p>If someone tried to make a
                                        Win32 console implementation and
                                        tried to implement both Windows
                                        9x Arabic terminal character set
                                        compatibility and wide string
                                        API (ReadConsoleOutputW)
                                        compatibility simultaneously,
                                        then they would run into the
                                        issue that there is currently no
                                        standardized mapping to handle
                                        that scenario. What should
                                        Windows 9x Arabic console
                                        compatible implementations do in
                                        that case?<br>
                                      </p>
                                    </div>
                                    <p><br>
                                    </p>
                                  </div>
                                </div>
                              </div>
                            </blockquote>
                          </div>
                          <p><br>
                          </p>
                        </div>
                      </div>
                    </div>
                  </blockquote>
                </div>
                <p><br>
                </p>
              </div>
            </div>
          </div>
        </blockquote>
      </div>
      <p><br>
      </p>
    </blockquote>
    <p><br>
    </p>
  </div>

</blockquote></div>

--000000000000689b4a065125d119--