RE: Odp: Pd: Missing legacy Arabic encoding

"[email protected] via Unicode" <[email protected]> Thu, 07 May 2026 18:02:57 +0200
Newsgroups gmane.text.unicode.general
Message-ID <[email protected]>
--2XKGPLQASURFHJVVTHQWBnhgwp
Content-Transfer-Encoding: quoted-printable
Content-Type: text/plain; charset=UTF-8

This is not of any interest, because the current Microsoft company is not e=
ven close ideologically to the past version of Microsoft that originally ma=
de the Arabic terminals, in fact they&#39;re not even ideologically compati=
ble with each other. There&#39;s no need to speculate on hypothetical votin=
g scenarios involving giant companies. What matters are logical arguments i=
nvolving encoding policy, and the prevailing reason is that there is no suf=
ficient evidence of user community that would need to interchange text from=
 those platforms into UCS-2 terminals. It is still necessary to debunk fals=
e arguments in order to prevent future decisions from being made incorrectl=
y, because even if they wouldn&#39;t affect the outcome of this proposal, t=
hey could still improperly influence the evaluation of future proposals.  D=
nia 07 maja 2026 17:36    Peter Constable via Unicode  &lt; [email protected]=
icode.org &gt;  napisa=C5=82(a): Just in case it might be of interest, if a=
 motion on this proposal were raised in a UTC meeting, I suspect Microsoft =
would vote against encoding. =C2=A0 =C2=A0 Peter =C2=A0 From:  Unicode &lt;=
[email protected]&gt;  On Behalf Of  [email protected] vi=
a Unicode  Sent:  May 6, 2026 11:09 AM  To:  Philippe Verdy via Unicode &lt=
;[email protected]&gt;; Philippe Verdy &lt;[email protected]&gt;  Sub=
ject:  Re: Odp: Pd: Missing legacy Arabic encoding  The ReadConsoleOutputW =
function, by definition, captures the tiles into the lpBuffer, which is a r=
andom access array of CHAR_INFO structure, whose horizontal and vertical si=
ze is specified in dwBufferSize. Since lpBuffer is random access, this impl=
ies one CHAR_INFO structure per character tile. The API therefore fundament=
ally imposes a strict memory layout that cannot be violated. The fact that =
some Unix-like terminals such as Windows Terminal may support features outs=
ide the scope of the CHAR_INFO structure for compatibility with ANSI escape=
 codes or WSL programs does not invalidate the compatibility considerations=
 for legacy DOS/Win16/Win32 programs that require all character tiles to fi=
t in the CHAR_INFO structure for random access, because 4 byte CHAR_INFO st=
ructure of Win32 is intended to be a fully backwards compatible extension o=
f the 2 byte VGA text mode tile structure of DOS/Win16.  Dnia 06 maja 2026 =
19:58  Philippe Verdy via Unicode  &lt;   [email protected]  &gt; na=
pisa=C5=82(a): Windows can use other ways to map 16-bit codes in its *legac=
y* Console buffer (using old CHAR_INFO structure), it can perfectly interna=
lly use compatibility characters, or PUAs of the BMP, and still present an =
API that exposes connforming sequences. You&#39;re talking about an old imp=
lementation that was built even=C2=A0 long before the Arabic script was ext=
ended (and newer scripts using contextual joining behaviors, that have neve=
r been part of the BMP, shcih as Adlam, and other scripts like Mongolian th=
at also may need such sequences with ZWJ/ZWNJ controls, or with other forma=
tting characters like those specific to Mongolian like FVS1...FVS4 and MVS,=
 or those common to many Bhramic scripts, that the *legacy* Console did not=
 support.  The *legacy* console was not built to support more than one plan=
e (including many CJK cgaracters). The newer console can!  Le=C2=A0mer. 6 m=
ai 2026 =C3=A0=C2=A018:37,   [email protected]  &lt;  piotrunio-2004@wp.=
pl &gt; a =C3=A9crit=C2=A0: Have you read the L2/26-077 proposal? Using ZWJ=
 or ZWNJ would not work for the compatibility purposes at all as already ex=
plained in the proposal. This is because ZWJ or ZWNJ would take the space o=
f one character tile in the CHAR_INFO structure. Suppose that you&#39;re tr=
ying to map 0xD0 from FP164 to a sequence of U+FE7C U+200D U+064B ( =EF=B9=
=BC=E2=80=8D=D9=8B ). The legacy application fills the 80=C3=9725 screen wi=
th all 0xD0 tiles. You subsequently try to capture the tiles with a Win32 p=
rogram by using ReadConsoleOutputA into an 80=C3=9725 buffer of 2000 tiles.=
 This succeeds and captures 0xD0 into all the tiles. You then try to captur=
e the tiles using ReadConsoleOutputW into an 80=C3=9725 buffer. Each sequen=
ce U+FE7C U+200D U+064B would take a sequence of three CHAR_INFO structures=
 to store, meaning 6000 such structures for the whole screen. But the 80=C3=
=9725 buffer has only room for 2000 instances of the structure (one per cha=
racter tile). Since CHAR_INFO stores 16-bit character code, by that same lo=
gic the compatibility characters would have to be in BMP for it to work. In=
 Windows 95 Vietnamese and Windows 95 Thai, there are instances where one c=
haracter tile takes multiple CHAR_INFO structures, causing visual width to =
be smaller than logical width, and when that happens, the remaining space a=
t the end of the line is left blank, allowing for CP1258/CP874 combining ch=
aracters in those systems to map 1:1 to their Unicode equivalents. Windows =
3.1/95/98/ME Arabic don&#39;t work that way and don&#39;t use combining cha=
racters or ZWJ sequences, so visual width is always equivalent to logical w=
idth, each character tile maps 1:1 to a CHAR_INFO structure and all charact=
ers may fill the entire line, which would be impossible if some of those ch=
aracters were mapped to composition sequences or non-BMP characters. Since =
there is currently no sufficient evidence of user community that would need=
 to use those mappings, there are no plans for those characters to be added=
 to Unicode, and therefore the only solution for the ReadConsoleOutputW to =
work properly in this case is to use agreed upon private use mappings for t=
hose compatibility characters.  Dnia 06 maja 2026 18:06  Philippe Verdy via=
 Unicode  &lt;   [email protected]  &gt; napisa=C5=82(a): You actual=
ly don&#39;t need any new compatibility characters for Arabic contextual fo=
rms, or for other contextual forms in other joining scripts (like Adlam, or=
 even Mongolian whichbis a LTR script).  You just have to prepend or append=
 a ZWJ or ZWNJ formatting control to the unified letter if you want to over=
ride its default contextual presentation form.   Le mar. 5 mai 2026, 00:48,=
 Asmus Freytag via Unicode &lt;  [email protected] &gt; a =C3=A9crit=
=C2=A0: The issue at hand is the distinction between a theoretical gap and =
real-life problem.  You have demonstrated that there are specifications tha=
t, if chained in the right way, can lead to ambiguities or gaps in intercha=
nge.  What we don&#39;t have is an actual use case with real-life consequen=
ces for a set of existing users, not hypothetical ones.  When it comes to e=
ncoding decisions based on existing documents, there is a strong presumptio=
n that once sufficiently many documents exist that contain a character, tha=
t this character will be needed in digitizing these documents, whether imme=
diately, or eventually (e.g. in the case of future scholarly studies). Also=
, the texts themselves exist, barring accidents, in permanence. Therefore, =
it is justified to consider irrevocably allocating a character that will ma=
p to this source in perpetuity, even though each encoded character carries =
a small cost for implementers.   However, when it comes to legacy character=
s, there&#39;s an additional cost that is imposed, and that is based on the=
 fact that characters that are encoded solely for compatibility will usuall=
y violate one or more of the other encoding principles, something that incr=
ementally complicates the standard. Even for people who never intend to use=
 that character.   Therefore, the SEW is on solid ground when it demands no=
t only a hypothetical scenario, but evidence of actual impact on actual use=
rs. Not only whether some application could invoke an API, but whether such=
 applications exist and are used today to access documents encoded using th=
e legacy characters in a way that is compromised irreparably by not having =
an encoding for them.  A./   On 5/4/2026 10:24 AM,   [email protected]  =
via Unicode wrote: In UTC 187 Minutes, &#34; Asmus Freytag noted that the f=
act that lists of things existed in the past does not make these things pla=
in text. Ned Holbrook pointed out that the purported issue occurs in a clos=
ed system, not in public interchange. &#34;. However, the arguments in the =
proposal do not merely hinge on the encodings being lists of characters, bu=
t specifically points out methods to interchange text, including an example=
 of copying terminal output and pasting to Notepad, where the copying invok=
es the mapping of the current terminal codepage to UCS-2 (as is CHAR_INFO c=
ompatible) and the pasting writes it into plain text. Win32 is also not a c=
losed system, as Win32 can capture the tiles of the output of Windows 3.1 A=
rabic DOS/Win16 programs and Windows 95/98/ME Arabic DOS/Win16/Win32 progra=
ms, but Win32 can also interact with public text interchange systems by rea=
ding and writing to files and network. I&#39;m not saying that Unicode abso=
lutely must include those characters, but those kinds of misleading claims =
are causing users to misunderstand what the proposal is about, and I don&#3=
9;t want Unicode to be relying on uninformed decisions to evaluate proposal=
s.  Dnia 18 kwietnia 2026 13:36  [email protected] via Unicode&lt; unico=
[email protected] &gt;  napisa=C5=82(a): The SEW subsequently explained t=
hat the actual reason is due to insufficient evidence of user community tha=
t would need to use the resulting mapping. Despite Win32 being a highly pop=
ular platform with plenty of backwards compatibility and native UCS-2 termi=
nal support, the specific use cases of installing codepages into Windows NT=
 and using terminal tiles from Windows 3.1/95/98/ME are not sufficiently do=
cumented, making it difficult for any user communities to form around it. S=
o it seems like the idea of standardizing legacy Arabic terminal BMP mappin=
gs is a dead end for now.  Dnia 17 kwietnia 2026 22:59  [email protected]=
l via Unicode&lt; [email protected] &gt;  napisa=C5=82(a): The Recom=
mendations in L2/26-100 claim that Microsoft&#39;s documentation of legacy =
Arabic encodings is available at  learn.microsoft.com https://learn.microso=
ft.com/en-us/typography/legacy/legacy_arabic_fonts . However, that article =
only demonstrates two encodings of TrueType fonts, which are used in Window=
s 3.1 but are completely different from the eight terminal encodings. Unlik=
e the TrueType encodings which represent internal shaping mappings and are =
not used for text interchange, the terminal encodings have been demonstrate=
d to be directly used in text interchange through int 10h and ReadConsoleOu=
tputA/WriteConsoleOutputA as already demonstrated in L2/26-077. The Recomme=
ndations also claim that the proposal does not demonstrate any need for int=
erchange or encoding, but the proposal actually demonstrated such a need du=
e to the logical extension of the Win32 terminal API to the functions ReadC=
onsoleOutputW/WriteConsoleOutputW, which are in Windows NT and may be used =
on the output of previously ran programs (including those that used the leg=
acy Arabic terminal encodings), which given the CHAR_INFO structure, theref=
ore implies a need for all the tiles to map to BMP for interchange. I&#39;m=
 not objecting to the SEW&#39;s conclusion of &#34;Users are expected to us=
e PUA.&#34;, which can indeed be used to provide a mapping even if not stan=
dardized, but the reasoning given was flawed.  Dnia 09 stycznia 2026 17:25 =
 [email protected]    &lt; [email protected] &gt;  napisa=C5=82(a): T=
he following Win32 C code will output 256 characters in system console code=
page into the character grid, capture those character tiles in UCS-2 if pos=
sible, and then output the current console codepage number.  #include &lt;w=
indows.h&gt;  #include &lt;stdio.h&gt;  int main(){  HANDLE hConsole=3DGetS=
tdHandle(STD_OUTPUT_HANDLE);  CHAR_INFO screen[256];  COORD size=3D{16,16,}=
;  COORD pos=3D{0,0,};  SMALL_RECT rect=3D{0,0,15,15,};  for(int i=3D0;i&lt=
;256;i++){  screen[i].Attributes=3D0xF0;  screen[i].Char.AsciiChar=3Di;  } =
 WriteConsoleOutputA(hConsole,screen,size,pos,&amp;rect);  CHAR_INFO screen=
u[256];  if(ReadConsoleOutputW(hConsole,screenu,size,pos,&amp;rect)){  for(=
int i=3D0;i&lt;256;i++) printf(&#34;%04X &#34;,screenu[i].Char.UnicodeChar)=
;  }  else{  printf(&#34;error %08X\n&#34;,GetLastError());  }  printf(&#34=
;codepage %u&#34;,GetConsoleOutputCP());  } In most cases, whenever a legac=
y Win32 codepage is used, the application can run on Windows NT to capture =
the UCS-2 mapping of those character cells to the BMP (although for CJK cod=
epages a more complex setup would be necessary due to thousands of fullwidt=
h characters with 2-byte sequences).  However, in Arabic versions of Window=
s 9x (95/98/ME) the resulting character set has many presentation forms tha=
t are not in Unicode. This is the result when running on Windows ME:=C2=A0 =
i.imgur.com https://i.imgur.com/QFm3SkI.png =C2=A0in 10=C3=9720 font,  i.im=
gur.com https://i.imgur.com/KUbLQ0A.png =C2=A0in 10=C3=9718 font (same resu=
lt also appears in Windows 95/98). 5=C3=9712, 7=C3=9712, 8=C3=9712, 10=C3=
=9718, 10=C3=9720, and 12=C3=9716 bitmap fonts have been attested with that=
 character set (VGAOEM.FON, 8514OEM.FON, DOSAPP.FON). The 10=C3=9720 font h=
as slightly different mapping than the other sizes: 0x93 is =C3=B6 instead =
of =C3=B4, and 0x97 is missing (causing the following characters on the sam=
e line to be drawn at the wrong position). It also claims to be using codep=
age 720, but many characters differ from their CP720 mappings, including th=
e bundled=C2=A0CP_720.NLS mappings (for example,  =D9=80  (U+0640 ARABIC TA=
TWEEL) is 0x95 in CP720, but in the console 0x95 is  =D8=B4  instead, and t=
he tatweel is at 0xFF). On Windows 9x,=C2=A0ReadConsoleOutputW is not suppo=
rted so the UCS-2 mappings of the console character tiles cannot be capture=
d (error 0x00000078 ERROR_CALL_NOT_IMPLEMENTED).  When that program runs on=
 Arabic versions of Windows NT, the visual output is of the CP437 character=
 set if one of the bundled bitmap fonts is used ( i.imgur.com https://i.img=
ur.com/RxjtxMH.png ), or the CP720 set if Lucida Console is used, with the =
Arabic letters either having glitchy font substitution (NT 4.0, NT 5.0/2000=
) or the .notdef glyph (NT 5.1/XP and up). In fact, it seems that the only =
Arabic bitmap fonts that occur in Windows NT are CP1256 fonts, which are no=
t used in terminals. So this appears to be one of those permanent Windows c=
ompatibility regressions that occured when Windows 9x ended, where the term=
inals can no longer render legacy Arabic text. Even if the user managed to =
use registry hacks to set the font to Courier New or Simplified Arabic Fixe=
d, it would still use the CP720 mapping which is not compatible with the Wi=
ndows 9x set.  It appears that in the Windows 9x Arabic terminal character =
set, 244 characters ( =E2=80=87 =EF=BA=80=EF=BA=81=EF=BA=82=EF=BA=83=EF=BA=
=84=EF=BA=85=EF=BA=87=EF=BA=88=EF=BA=8A=EF=BA=8B=EF=BA=8D=EF=BA=8E=EF=BA=8F=
=EF=BA=91=EF=BA=93=E2=96=BA=E2=97=84=E2=86=95=EF=BA=95=C2=B6=C2=A7=EF=BA=97=
=EF=BA=99=E2=86=91=E2=86=93=E2=86=92=E2=86=90=EF=BA=9B=EF=B9=B0 =E2=96=B2=
=E2=96=BC  !&#34;#$%&amp;&#39;()*+,-./0123456789:;&lt;=3D&gt;?@ABCDEFGHIJKL=
MNOPQRSTUVWXYZ[\]^_`abcdefghijklmnopqrstuvwxyz{|}~ =EF=BA=9D=EF=BA=9F=EF=BA=
=A1 =C3=A9=C3=A2 =EF=BA=A3 =C3=A0 =EF=BA=A5 =C3=A7=C3=AA=C3=AB=C3=A8=C3=AF=
=C3=AE =EF=BA=A7=EF=BA=A9=EF=BA=AB=EF=BA=AD=EF=BA=AF =C3=B4 =EF=BA=B3 =C3=
=BB=C3=B9 =EF=BA=B7=EF=BA=BB=C2=A3=EF=BA=BF=EF=BB=81=EF=BB=85=EF=BB=89=EF=
=BB=8A=EF=BB=8B=EF=BB=8C=EF=BB=8D=EF=BB=8E=EF=BB=8F=EF=BB=90=EF=BB=91=EF=BB=
=93=EF=BB=95=EF=BB=97=EF=BB=99=EF=BB=9B=C2=AB=C2=BB=EF=B9=B1=E2=96=92=EF=B9=
=B2=E2=94=82=E2=94=A4=EF=B9=B4=EF=B9=B6=EF=B9=B7=EF=B9=B8=D9=A0=D9=A1=D9=A2=
=D9=A3=EF=B9=B9=EF=B9=BA=E2=94=90=E2=94=94=E2=94=B4=E2=94=AC=E2=94=9C=E2=94=
=80=E2=94=BC=EF=B9=BB=EF=B9=BE=D9=A4=D9=A5=D9=A6=D9=A7=D9=A8=D9=A9=D8=8C=EF=
=B9=BF=EF=B1=9E=EF=B1=9F=EF=B1=A0=EF=B3=B2=EF=B1=A1=EF=B3=B3=EF=B1=A2=E2=94=
=98=E2=94=8C=D8=9B=D8=9F=C2 =C2=B5 =EF=BB=A9=EF=BB=AB=EF=BB=AC=EF=BB=AD=EF=
=BB=AF=EF=BB=B0=EF=BB=B1=EF=BB=B2=EF=BB=B3=EF=B3=B4=EF=B9=BC=EF=B9=BD=EF=BA=
=B1=EF=BA=B5=EF=BA=B9=EF=BA=BD=EF=B9=B3=C2=B0=C2=B7=E2=96=A0=D9=80 ) are al=
ready in Unicode, but 12 characters are not in Unicode: =E2=80=A2 6 of them=
 are pieces of lam-alef ligatures (0xDD, 0xDE, 0xF9, 0xFB, 0xFC, 0xFD) =E2=
=80=A2 2 of them are shadda with fathatan ligatures without or with tatweel=
 (0xD0, 0xD1) =E2=80=94 in some legacy Microsoft fonts, shadda with fathata=
n is mapped to private use U+E818 =E2=80=A2 4 of them are disunifications o=
f seen/sheen/sad/dad occuring either with or without tail =E2=80=94=C2=A0 =
=EF=B9=B3  (U+FE73 ARABIC TAIL FRAGMENT) was originally encoded in Unicode =
3.2 for CP864 compatibility; in that codepage, the forms of=C2=A0seen/sheen=
/sad/dad attach to the tail fragment =E2=80=94 forms with included tail:=C2=
=A00x92, 0x95, 0x98, 0x8A =E2=80=94 forms without tail (attaching to tail f=
ragment like in CP864):=C2=A00xF3, 0xF4, 0xF5, 0xF6  If someone tried to ma=
ke a Win32 console implementation and tried to implement both Windows 9x Ar=
abic terminal character set compatibility and wide string API (ReadConsoleO=
utputW) compatibility simultaneously, then they would run into the issue th=
at there is currently no standardized mapping to handle that scenario. What=
 should Windows 9x Arabic console compatible implementations do in that cas=
e?=0D

--2XKGPLQASURFHJVVTHQWBnhgwp
Content-Transfer-Encoding: quoted-printable
Content-Type: text/html; charset=UTF-8

<p>This is not of any interest, because the current Microsoft company is no=
t even close ideologically to the past version of Microsoft that originally=
 made the Arabic terminals, in fact they're not even ideologically compatib=
le with each other. There's no need to speculate on hypothetical voting sce=
narios involving giant companies. What matters are logical arguments involv=
ing encoding policy, and the prevailing reason is that there is no sufficie=
nt evidence of user community that would need to interchange text from thos=
e platforms into UCS-2 terminals. It is still necessary to debunk false arg=
uments in order to prevent future decisions from being made incorrectly, be=
cause even if they wouldn't affect the outcome of this proposal, they could=
 still improperly influence the evaluation of future proposals.</p><p><br><=
/p><div class=3D"nh_extra"><blockquote style=3D"padding-top: 12px;" class=
=3D"nh_qoute"><p style=3D"padding-bottom: 12px;"><strong>Dnia 07 maja 2026 =
17:36</strong> <a href=3D"mailto:[email protected]" rel=3D"noopener =
noreferrer nofollow" target=3D"_blank"><span style=3D"margin-left: 4px;">Pe=
ter Constable via Unicode</span></a><span style=3D"margin-left: 4px;"> &lt;=
 [email protected] &gt;</span> napisa=C5=82(a):</p><div id=3D"gwpb13=
52f8b"><div id=3D"gwpb1352f8bh"><style></style><div style=3D"overflow-wrap:=
 break-word;" lang=3D"EN-CA" class=3D"gwpb1352f8bb" data-message-body=3D"tr=
ue" data-color-mode=3D"light"><div class=3D"gwpb1352f8b_WordSection1"><p cl=
ass=3D"gwpb1352f8b_MsoNormal"><span style=3D"font-size: 11pt;">Just in case=
 it might be of interest, if a motion on this proposal were raised in a UTC=
 meeting, I suspect Microsoft would vote against encoding.</span></p><p cla=
ss=3D"gwpb1352f8b_MsoNormal"><span style=3D"font-size: 11pt;">&nbsp;</span>=
</p><p class=3D"gwpb1352f8b_MsoNormal"><span style=3D"font-size: 11pt;">&nb=
sp;</span></p><p class=3D"gwpb1352f8b_MsoNormal"><span style=3D"font-size: =
11pt;">Peter</span></p><p class=3D"gwpb1352f8b_MsoNormal"><span style=3D"fo=
nt-size: 11pt;">&nbsp;</span></p><div style=3D"border-right: none; border-b=
ottom: none; border-left: none; border-image: initial; border-top: 1pt soli=
d rgb(225, 225, 225); padding: 3pt 0cm 0cm;"><p class=3D"gwpb1352f8b_MsoNor=
mal"><span style=3D"font-size: 11pt; font-family: Calibri, sans-serif;" lan=
g=3D"EN-US"><strong>From:</strong> Unicode &lt;[email protected]=
.org&gt; <strong>On Behalf Of </strong>[email protected] via Unicode<br>=
<strong>Sent:</strong> May 6, 2026 11:09 AM<br><strong>To:</strong> Philipp=
e Verdy via Unicode &lt;[email protected]&gt;; Philippe Verdy &lt;ve=
[email protected]&gt;<br><strong>Subject:</strong> Re: Odp: Pd: Missing legacy=
 Arabic encoding</span></p></div><p class=3D"gwpb1352f8b_MsoNormal"><br></p=
><p>The ReadConsoleOutputW function, by definition, captures the tiles into=
 the lpBuffer, which is a random access array of CHAR_INFO structure, whose=
 horizontal and vertical size is specified in dwBufferSize. Since lpBuffer =
is random access, this implies one CHAR_INFO structure per character tile. =
The API therefore fundamentally imposes a strict memory layout that cannot =
be violated. The fact that some Unix-like terminals such as Windows Termina=
l may support features outside the scope of the CHAR_INFO structure for com=
patibility with ANSI escape codes or WSL programs does not invalidate the c=
ompatibility considerations for legacy DOS/Win16/Win32 programs that requir=
e all character tiles to fit in the CHAR_INFO structure for random access, =
because 4 byte CHAR_INFO structure of Win32 is intended to be a fully backw=
ards compatible extension of the 2 byte VGA text mode tile structure of DOS=
/Win16.</p><p><br></p><div><blockquote style=3D"margin-top: 5pt; margin-bot=
tom: 5pt;"><p><span style=3D"font-family: Aptos, sans-serif;"><strong>Dnia =
06 maja 2026 19:58</strong></span><a href=3D"mailto:[email protected]=
g" rel=3D"noopener noreferrer" target=3D"_blank">Philippe Verdy via Unicode=
</a> &lt; <a href=3D"mailto:[email protected]" rel=3D"noopener noref=
errer" target=3D"_blank">[email protected]</a> &gt; napisa=C5=82(a):=
</p><div id=3D"gwpb1352f8b_gwp6bd5644c"><div id=3D"gwpb1352f8b_gwp6bd5644ch=
"><div><div><p class=3D"gwpb1352f8b_MsoNormal">Windows can use other ways t=
o map 16-bit codes in its *legacy* Console buffer (using old CHAR_INFO stru=
cture), it can perfectly internally use compatibility characters, or PUAs o=
f the BMP, and still present an API that exposes connforming sequences. You=
're talking about an old implementation that was built even&nbsp; long befo=
re the Arabic script was extended (and newer scripts using contextual joini=
ng behaviors, that have never been part of the BMP, shcih as Adlam, and oth=
er scripts like Mongolian that also may need such sequences with ZWJ/ZWNJ c=
ontrols, or with other formatting characters like those specific to Mongoli=
an like FVS1...FVS4 and MVS, or those common to many Bhramic scripts, that =
the *legacy* Console did not support.<br>The *legacy* console was not built=
 to support more than one plane (including many CJK cgaracters). The newer =
console can!</p></div><p><br></p><div><div><p class=3D"gwpb1352f8b_MsoNorma=
l">Le&nbsp;mer. 6 mai 2026 =C3=A0&nbsp;18:37, <a href=3D"mailto:piotrunio-2=
[email protected]" rel=3D"noopener noreferrer" target=3D"_blank">piotrunio-2004@wp.=
pl</a> &lt;<a href=3D"mailto:[email protected]" rel=3D"noopener noreferr=
er" target=3D"_blank">[email protected]</a>&gt; a =C3=A9crit&nbsp;:</p><=
/div><blockquote style=3D"border-top: none; border-right: none; border-bott=
om: none; border-image: initial; border-left: 1pt solid rgb(204, 204, 204);=
 padding: 0cm 0cm 0cm 6pt; margin-left: 4.8pt; margin-right: 0cm;"><p>Have =
you read the L2/26-077 proposal? Using ZWJ or ZWNJ would not work for the c=
ompatibility purposes at all as already explained in the proposal. This is =
because ZWJ or ZWNJ would take the space of one character tile in the CHAR_=
INFO structure. Suppose that you're trying to map 0xD0 from FP164 to a sequ=
ence of U+FE7C U+200D U+064B (<span style=3D"font-family: &quot;Times New R=
oman&quot;, serif;" lang=3D"AR-SA" dir=3D"RTL">=EF=B9=BC=E2=80=8D=D9=8B</sp=
an>). The legacy application fills the 80=C3=9725 screen with all 0xD0 tile=
s. You subsequently try to capture the tiles with a Win32 program by using =
ReadConsoleOutputA into an 80=C3=9725 buffer of 2000 tiles. This succeeds a=
nd captures 0xD0 into all the tiles. You then try to capture the tiles usin=
g ReadConsoleOutputW into an 80=C3=9725 buffer. Each sequence U+FE7C U+200D=
 U+064B would take a sequence of three CHAR_INFO structures to store, meani=
ng 6000 such structures for the whole screen. But the 80=C3=9725 buffer has=
 only room for 2000 instances of the structure (one per character tile). Si=
nce CHAR_INFO stores 16-bit character code, by that same logic the compatib=
ility characters would have to be in BMP for it to work. In Windows 95 Viet=
namese and Windows 95 Thai, there are instances where one character tile ta=
kes multiple CHAR_INFO structures, causing visual width to be smaller than =
logical width, and when that happens, the remaining space at the end of the=
 line is left blank, allowing for CP1258/CP874 combining characters in thos=
e systems to map 1:1 to their Unicode equivalents. Windows 3.1/95/98/ME Ara=
bic don't work that way and don't use combining characters or ZWJ sequences=
, so visual width is always equivalent to logical width, each character til=
e maps 1:1 to a CHAR_INFO structure and all characters may fill the entire =
line, which would be impossible if some of those characters were mapped to =
composition sequences or non-BMP characters. Since there is currently no su=
fficient evidence of user community that would need to use those mappings, =
there are no plans for those characters to be added to Unicode, and therefo=
re the only solution for the ReadConsoleOutputW to work properly in this ca=
se is to use agreed upon private use mappings for those compatibility chara=
cters.</p><p><br></p><div><blockquote style=3D"margin-top: 5pt; margin-bott=
om: 5pt;"><p><span style=3D"font-family: Aptos, sans-serif;"><strong>Dnia 0=
6 maja 2026 18:06</strong></span><a href=3D"mailto:[email protected]=
" rel=3D"noopener noreferrer" target=3D"_blank">Philippe Verdy via Unicode<=
/a> &lt; <a href=3D"mailto:[email protected]" rel=3D"noopener norefe=
rrer" target=3D"_blank">[email protected]</a> &gt; napisa=C5=82(a):<=
/p><div id=3D"gwpb1352f8b_gwp6bd5644c_m_-2650882641749569987gwpb67625a5"><d=
iv id=3D"gwpb1352f8b_gwp6bd5644c_m_-2650882641749569987gwpb67625a5h"><div><=
div><p>You actually don't need any new compatibility characters for Arabic =
contextual forms, or for other contextual forms in other joining scripts (l=
ike Adlam, or even Mongolian whichbis a LTR script).</p><div><p class=3D"gw=
pb1352f8b_MsoNormal"><br></p></div><div><p class=3D"gwpb1352f8b_MsoNormal">=
You just have to prepend or append a ZWJ or ZWNJ formatting control to the =
unified letter if you want to override its default contextual presentation =
form.</p></div><div><p class=3D"gwpb1352f8b_MsoNormal"><br></p></div></div>=
<p><br></p><div><div><p class=3D"gwpb1352f8b_MsoNormal">Le mar. 5 mai 2026,=
 00:48, Asmus Freytag via Unicode &lt;<a href=3D"mailto:[email protected]=
e.org" rel=3D"noopener noreferrer" target=3D"_blank">[email protected]=
rg</a>&gt; a =C3=A9crit&nbsp;:</p></div><blockquote style=3D"border-top: no=
ne; border-right: none; border-bottom: none; border-image: initial; border-=
left: 1pt solid rgb(204, 204, 204); padding: 0cm 0cm 0cm 6pt; margin-left: =
4.8pt; margin-right: 0cm;"><div><p>The issue at hand is the distinction bet=
ween a theoretical gap and real-life problem.</p><p><br></p><p>You have dem=
onstrated that there are specifications that, if chained in the right way, =
can lead to ambiguities or gaps in interchange.</p><p><br></p><p>What we do=
n't have is an actual use case with real-life consequences for a set of exi=
sting users, not hypothetical ones.</p><p><br></p><p>When it comes to encod=
ing decisions based on existing documents, there is a strong presumption th=
at once sufficiently many documents exist that contain a character, that th=
is character will be needed in digitizing these documents, whether immediat=
ely, or eventually (e.g. in the case of future scholarly studies). Also, th=
e texts themselves exist, barring accidents, in permanence. Therefore, it i=
s justified to consider irrevocably allocating a character that will map to=
 this source in perpetuity, even though each encoded character carries a sm=
all cost for implementers.<br><br>However, when it comes to legacy characte=
rs, there's an additional cost that is imposed, and that is based on the fa=
ct that characters that are encoded solely for compatibility will usually v=
iolate one or more of the other encoding principles, something that increme=
ntally complicates the standard. Even for people who never intend to use th=
at character.<br><br>Therefore, the SEW is on solid ground when it demands =
not only a hypothetical scenario, but evidence of actual impact on actual u=
sers. Not only whether some application could invoke an API, but whether su=
ch applications exist and are used today to access documents encoded using =
the legacy characters in a way that is compromised irreparably by not havin=
g an encoding for them.</p><p><br></p><p>A./</p><p><br></p><p><br></p><p>On=
 5/4/2026 10:24 AM, <a href=3D"mailto:[email protected]" rel=3D"noopener=
 noreferrer" target=3D"_blank">[email protected]</a> via Unicode wrote:<=
/p><blockquote style=3D"margin-top: 5pt; margin-bottom: 5pt;"><p>In UTC 187=
 Minutes, "<span style=3D"font-size: 13.5pt; font-family: &quot;DMCA Sans S=
erif 10.0 dev1&quot;, serif; color: black;">Asmus Freytag noted that the fa=
ct that lists of things existed in the past does not make these things plai=
n text. Ned Holbrook pointed out that the purported issue occurs in a close=
d system, not in public interchange.</span>". However, the arguments in the=
 proposal do not merely hinge on the encodings being lists of characters, b=
ut specifically points out methods to interchange text, including an exampl=
e of copying terminal output and pasting to Notepad, where the copying invo=
kes the mapping of the current terminal codepage to UCS-2 (as is CHAR_INFO =
compatible) and the pasting writes it into plain text. Win32 is also not a =
closed system, as Win32 can capture the tiles of the output of Windows 3.1 =
Arabic DOS/Win16 programs and Windows 95/98/ME Arabic DOS/Win16/Win32 progr=
ams, but Win32 can also interact with public text interchange systems by re=
ading and writing to files and network. I'm not saying that Unicode absolut=
ely must include those characters, but those kinds of misleading claims are=
 causing users to misunderstand what the proposal is about, and I don't wan=
t Unicode to be relying on uninformed decisions to evaluate proposals.</p><=
p><br></p><div><blockquote style=3D"margin-top: 5pt; margin-bottom: 5pt;"><=
p><span style=3D"font-family: Aptos, sans-serif;"><strong>Dnia 18 kwietnia =
2026 13:36</strong></span><a href=3D"mailto:[email protected]" rel=
=3D"noopener noreferrer" target=3D"_blank">[email protected] via Unicode=
&lt; [email protected] &gt;</a> napisa=C5=82(a):</p><div id=3D"gwpb1=
352f8b_gwp6bd5644c_m_-2650882641749569987gwpb67625a5_m_5371084887079304582g=
wpbf2d884a"><div id=3D"gwpb1352f8b_gwp6bd5644c_m_-2650882641749569987gwpb67=
625a5_m_5371084887079304582gwpbf2d884ah"><div><p>The SEW subsequently expla=
ined that the actual reason is due to insufficient evidence of user communi=
ty that would need to use the resulting mapping. Despite Win32 being a high=
ly popular platform with plenty of backwards compatibility and native UCS-2=
 terminal support, the specific use cases of installing codepages into Wind=
ows NT and using terminal tiles from Windows 3.1/95/98/ME are not sufficien=
tly documented, making it difficult for any user communities to form around=
 it. So it seems like the idea of standardizing legacy Arabic terminal BMP =
mappings is a dead end for now.</p><p><br></p><div><blockquote style=3D"mar=
gin-top: 5pt; margin-bottom: 5pt;"><p><span style=3D"font-family: Aptos, sa=
ns-serif;"><strong>Dnia 17 kwietnia 2026 22:59</strong></span><a href=3D"ma=
ilto:[email protected]" rel=3D"noopener noreferrer" target=3D"_blank=
">[email protected] via Unicode&lt; [email protected] &gt;</a> na=
pisa=C5=82(a):</p><div id=3D"gwpb1352f8b_gwp6bd5644c_m_-2650882641749569987=
gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f8"><div id=3D"gwpb13=
52f8b_gwp6bd5644c_m_-2650882641749569987gwpb67625a5_m_5371084887079304582gw=
pbf2d884a_gwp2281a7f8h"><div><p>The Recommendations in L2/26-100 claim that=
 Microsoft's documentation of legacy Arabic encodings is available at <a hr=
ef=3D"https://learn.microsoft.com/en-us/typography/legacy/legacy_arabic_fon=
ts" =3D"" rel=3D"noopener noreferrer" target=3D"_blank">https://learn.micro=
soft.com/en-us/typography/legacy/legacy_arabic_fonts</a>. However, that art=
icle only demonstrates two encodings of TrueType fonts, which are used in W=
indows 3.1 but are completely different from the eight terminal encodings. =
Unlike the TrueType encodings which represent internal shaping mappings and=
 are not used for text interchange, the terminal encodings have been demons=
trated to be directly used in text interchange through int 10h and ReadCons=
oleOutputA/WriteConsoleOutputA as already demonstrated in L2/26-077. The Re=
commendations also claim that the proposal does not demonstrate any need fo=
r interchange or encoding, but the proposal actually demonstrated such a ne=
ed due to the logical extension of the Win32 terminal API to the functions =
ReadConsoleOutputW/WriteConsoleOutputW, which are in Windows NT and may be =
used on the output of previously ran programs (including those that used th=
e legacy Arabic terminal encodings), which given the CHAR_INFO structure, t=
herefore implies a need for all the tiles to map to BMP for interchange. I'=
m not objecting to the SEW's conclusion of "Users are expected to use PUA."=
, which can indeed be used to provide a mapping even if not standardized, b=
ut the reasoning given was flawed.</p><p><br></p><div><blockquote style=3D"=
margin-top: 5pt; margin-bottom: 5pt;"><p><span style=3D"font-family: Aptos,=
 sans-serif;"><strong>Dnia 09 stycznia 2026 17:25</strong></span><a href=3D=
"mailto:[email protected]" rel=3D"noopener noreferrer" target=3D"_blank"=
>[email protected]</a> <a href=3D"mailto:[email protected]" rel=3D"no=
opener noreferrer" target=3D"_blank">&lt; [email protected] &gt;</a> nap=
isa=C5=82(a):</p><div id=3D"gwpb1352f8b_gwp6bd5644c_m_-2650882641749569987g=
wpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f8_gwpa05276c7"><div i=
d=3D"gwpb1352f8b_gwp6bd5644c_m_-2650882641749569987gwpb67625a5_m_5371084887=
079304582gwpbf2d884a_gwp2281a7f8_gwpa05276c7h"><div><div id=3D"gwpb1352f8b_=
gwp6bd5644c_m_-2650882641749569987gwpb67625a5_m_5371084887079304582gwpbf2d8=
84a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718"><div id=3D"gwpb1352f8b_gwp6bd5644c=
_m_-2650882641749569987gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281=
a7f8_gwpa05276c7_gwpa8b5f718h"><div><div id=3D"gwpb1352f8b_gwp6bd5644c_m_-2=
650882641749569987gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f8_=
gwpa05276c7_gwpa8b5f718_gwpa8b5f718"><div id=3D"gwpb1352f8b_gwp6bd5644c_m_-=
2650882641749569987gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f8=
_gwpa05276c7_gwpa8b5f718_gwpa8b5f718h"><div><div id=3D"gwpb1352f8b_gwp6bd56=
44c_m_-2650882641749569987gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2=
281a7f8_gwpa05276c7_gwpa8b5f718_gwpa8b5f718_gwpa8b5f718"><div id=3D"gwpb135=
2f8b_gwp6bd5644c_m_-2650882641749569987gwpb67625a5_m_5371084887079304582gwp=
bf2d884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718_gwpa8b5f718_gwpa8b5f718h"><div=
><p>The following Win32 C code will output 256 characters in system console=
 codepage into the character grid, capture those character tiles in UCS-2 i=
f possible, and then output the current console codepage number.</p><p><br>=
</p><p style=3D"margin-bottom: 12pt;">#include &lt;windows.h&gt;<br>#includ=
e &lt;stdio.h&gt;<br>int main(){<br>HANDLE hConsole=3DGetStdHandle(STD_OUTP=
UT_HANDLE);<br>CHAR_INFO screen[256];<br>COORD size=3D{16,16,};<br>COORD po=
s=3D{0,0,};<br>SMALL_RECT rect=3D{0,0,15,15,};<br>for(int i=3D0;i&lt;256;i+=
+){<br>screen[i].Attributes=3D0xF0;<br>screen[i].Char.AsciiChar=3Di;<br>}<b=
r>WriteConsoleOutputA(hConsole,screen,size,pos,&amp;rect);<br>CHAR_INFO scr=
eenu[256];<br>if(ReadConsoleOutputW(hConsole,screenu,size,pos,&amp;rect)){<=
br>for(int i=3D0;i&lt;256;i++) printf("%04X ",screenu[i].Char.UnicodeChar);=
<br>}<br>else{<br>printf("error %08X\n",GetLastError());<br>}<br>printf("co=
depage %u",GetConsoleOutputCP());<br>}</p><p>In most cases, whenever a lega=
cy Win32 codepage is used, the application can run on Windows NT to capture=
 the UCS-2 mapping of those character cells to the BMP (although for CJK co=
depages a more complex setup would be necessary due to thousands of fullwid=
th characters with 2-byte sequences).</p><p><br></p><p>However, in Arabic v=
ersions of Windows 9x (95/98/ME) the resulting character set has many prese=
ntation forms that are not in Unicode. This is the result when running on W=
indows ME:&nbsp;<a href=3D"https://i.imgur.com/QFm3SkI.png" =3D"" rel=3D"no=
opener noreferrer" target=3D"_blank">https://i.imgur.com/QFm3SkI.png</a>&nb=
sp;in 10=C3=9720 font, <a href=3D"https://i.imgur.com/KUbLQ0A.png" =3D"" re=
l=3D"noopener noreferrer" target=3D"_blank">https://i.imgur.com/KUbLQ0A.png=
</a>&nbsp;in 10=C3=9718 font (same result also appears in Windows 95/98). 5=
=C3=9712, 7=C3=9712, 8=C3=9712, 10=C3=9718, 10=C3=9720, and 12=C3=9716 bitm=
ap fonts have been attested with that character set (VGAOEM.FON, 8514OEM.FO=
N, DOSAPP.FON). The 10=C3=9720 font has slightly different mapping than the=
 other sizes: 0x93 is =C3=B6 instead of =C3=B4, and 0x97 is missing (causin=
g the following characters on the same line to be drawn at the wrong positi=
on). It also claims to be using codepage 720, but many characters differ fr=
om their CP720 mappings, including the bundled&nbsp;CP_720.NLS mappings (fo=
r example, <span style=3D"font-family: &quot;Times New Roman&quot;, serif;"=
 lang=3D"AR-SA" dir=3D"RTL">=D9=80</span> (U+0640 ARABIC TATWEEL) is 0x95 i=
n CP720, but in the console 0x95 is <span style=3D"font-family: &quot;Times=
 New Roman&quot;, serif;" lang=3D"AR-SA" dir=3D"RTL">=D8=B4</span> instead,=
 and the tatweel is at 0xFF). On Windows 9x,&nbsp;ReadConsoleOutputW is not=
 supported so the UCS-2 mappings of the console character tiles cannot be c=
aptured (error 0x00000078 ERROR_CALL_NOT_IMPLEMENTED).</p></div></div></div=
><p><br></p><p>When that program runs on Arabic versions of Windows NT, the=
 visual output is of the CP437 character set if one of the bundled bitmap f=
onts is used (<a href=3D"https://i.imgur.com/RxjtxMH.png" =3D"" rel=3D"noop=
ener noreferrer" target=3D"_blank">https://i.imgur.com/RxjtxMH.png</a>), or=
 the CP720 set if Lucida Console is used, with the Arabic letters either ha=
ving glitchy font substitution (NT 4.0, NT 5.0/2000) or the .notdef glyph (=
NT 5.1/XP and up). In fact, it seems that the only Arabic bitmap fonts that=
 occur in Windows NT are CP1256 fonts, which are not used in terminals. So =
this appears to be one of those permanent Windows compatibility regressions=
 that occured when Windows 9x ended, where the terminals can no longer rend=
er legacy Arabic text. Even if the user managed to use registry hacks to se=
t the font to Courier New or Simplified Arabic Fixed, it would still use th=
e CP720 mapping which is not compatible with the Windows 9x set.</p></div><=
/div><p><br></p></div></div></div></div><div id=3D"gwpb1352f8b_gwp6bd5644c_=
m_-2650882641749569987gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a=
7f8_gwpa05276c7_gwpa8b5f718"><div id=3D"gwpb1352f8b_gwp6bd5644c_m_-26508826=
41749569987gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f8_gwpa052=
76c7_gwpa8b5f718h"><div><p>It appears that in the Windows 9x Arabic termina=
l character set, 244 characters (<span style=3D"font-family: Arial, sans-se=
rif;">=E2=80=87</span><span style=3D"font-family: &quot;Times New Roman&quo=
t;, serif;" lang=3D"AR-SA" dir=3D"RTL">=EF=BA=80=EF=BA=81=EF=BA=82=EF=BA=83=
=EF=BA=84=EF=BA=85=EF=BA=87=EF=BA=88=EF=BA=8A=EF=BA=8B=EF=BA=8D=EF=BA=8E=EF=
=BA=8F=EF=BA=91=EF=BA=93=E2=96=BA=E2=97=84=E2=86=95=EF=BA=95=C2=B6=C2=A7=EF=
=BA=97=EF=BA=99=E2=86=91=E2=86=93=E2=86=92=E2=86=90=EF=BA=9B=EF=B9=B0</span=
><span style=3D"font-family: Arial, sans-serif;">=E2=96=B2=E2=96=BC</span> =
!"#$%&amp;'()*+,-./0123456789:;&lt;=3D&gt;?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\]^_=
`abcdefghijklmnopqrstuvwxyz{|}~<span style=3D"font-family: &quot;Times New =
Roman&quot;, serif;" lang=3D"AR-SA" dir=3D"RTL">=EF=BA=9D=EF=BA=9F=EF=BA=A1=
</span>=C3=A9=C3=A2<span style=3D"font-family: &quot;Times New Roman&quot;,=
 serif;" lang=3D"AR-SA" dir=3D"RTL">=EF=BA=A3</span>=C3=A0<span style=3D"fo=
nt-family: &quot;Times New Roman&quot;, serif;" lang=3D"AR-SA" dir=3D"RTL">=
=EF=BA=A5</span>=C3=A7=C3=AA=C3=AB=C3=A8=C3=AF=C3=AE<span style=3D"font-fam=
ily: &quot;Times New Roman&quot;, serif;" lang=3D"AR-SA" dir=3D"RTL">=EF=BA=
=A7=EF=BA=A9=EF=BA=AB=EF=BA=AD=EF=BA=AF</span>=C3=B4<span style=3D"font-fam=
ily: &quot;Times New Roman&quot;, serif;" lang=3D"AR-SA" dir=3D"RTL">=EF=BA=
=B3</span>=C3=BB=C3=B9<span style=3D"font-family: &quot;Times New Roman&quo=
t;, serif;" lang=3D"AR-SA" dir=3D"RTL">=EF=BA=B7=EF=BA=BB=C2=A3=EF=BA=BF=EF=
=BB=81=EF=BB=85=EF=BB=89=EF=BB=8A=EF=BB=8B=EF=BB=8C=EF=BB=8D=EF=BB=8E=EF=BB=
=8F=EF=BB=90=EF=BB=91=EF=BB=93=EF=BB=95=EF=BB=97=EF=BB=99=EF=BB=9B=C2=AB=C2=
=BB=EF=B9=B1=E2=96=92=EF=B9=B2=E2=94=82=E2=94=A4=EF=B9=B4=EF=B9=B6=EF=B9=B7=
=EF=B9=B8=D9=A0=D9=A1=D9=A2=D9=A3=EF=B9=B9=EF=B9=BA=E2=94=90=E2=94=94=E2=94=
=B4=E2=94=AC=E2=94=9C=E2=94=80=E2=94=BC=EF=B9=BB=EF=B9=BE=D9=A4=D9=A5=D9=A6=
=D9=A7=D9=A8=D9=A9=D8=8C=EF=B9=BF=EF=B1=9E=EF=B1=9F=EF=B1=A0=EF=B3=B2=EF=B1=
=A1=EF=B3=B3=EF=B1=A2=E2=94=98=E2=94=8C=D8=9B=D8=9F=C2=A4=EF=BB=9D=EF=BB=9F=
=EF=BB=A1=EF=BB=A3=EF=BB=A5=EF=BB=A7</span>=C2=B5<span style=3D"font-family=
: &quot;Times New Roman&quot;, serif;" lang=3D"AR-SA" dir=3D"RTL">=EF=BB=A9=
=EF=BB=AB=EF=BB=AC=EF=BB=AD=EF=BB=AF=EF=BB=B0=EF=BB=B1=EF=BB=B2=EF=BB=B3=EF=
=B3=B4=EF=B9=BC=EF=B9=BD=EF=BA=B1=EF=BA=B5=EF=BA=B9=EF=BA=BD=EF=B9=B3=C2=B0=
=C2=B7=E2=96=A0=D9=80</span>) are already in Unicode, but 12 characters are=
 not in Unicode:</p><p>=E2=80=A2 6 of them are pieces of lam-alef ligatures=
 (0xDD, 0xDE, 0xF9, 0xFB, 0xFC, 0xFD)</p><p>=E2=80=A2 2 of them are shadda =
with fathatan ligatures without or with tatweel (0xD0, 0xD1)</p><p>=E2=80=
=94 in some legacy Microsoft fonts, shadda with fathatan is mapped to priva=
te use U+E818</p><p>=E2=80=A2 4 of them are disunifications of seen/sheen/s=
ad/dad occuring either with or without tail</p><p>=E2=80=94&nbsp;<span styl=
e=3D"font-family: &quot;Times New Roman&quot;, serif;" lang=3D"AR-SA" dir=
=3D"RTL">=EF=B9=B3</span> (U+FE73 ARABIC TAIL FRAGMENT) was originally enco=
ded in Unicode 3.2 for CP864 compatibility; in that codepage, the forms of&=
nbsp;seen/sheen/sad/dad attach to the tail fragment</p><p>=E2=80=94 forms w=
ith included tail:&nbsp;0x92, 0x95, 0x98, 0x8A</p><p>=E2=80=94 forms withou=
t tail (attaching to tail fragment like in CP864):&nbsp;0xF3, 0xF4, 0xF5, 0=
xF6</p></div><p><br></p></div><p>If someone tried to make a Win32 console i=
mplementation and tried to implement both Windows 9x Arabic terminal charac=
ter set compatibility and wide string API (ReadConsoleOutputW) compatibilit=
y simultaneously, then they would run into the issue that there is currentl=
y no standardized mapping to handle that scenario. What should Windows 9x A=
rabic console compatible implementations do in that case?</p></div><p><br><=
/p></div></div></div></blockquote></div><p><br></p></div></div></div></bloc=
kquote></div><p><br></p></div></div></div></blockquote></div><p><br></p></b=
lockquote><p><br></p></div></blockquote></div></div></div></div></blockquot=
e></div><p><br></p></blockquote></div></div></div></div></blockquote></div>=
<p><br></p></div></div></div></div></blockquote></div><p><br></p>
--2XKGPLQASURFHJVVTHQWBnhgwp--