Re: Odp: Pd: Missing legacy Arabic encoding

"[email protected] via Unicode" <[email protected]> Wed, 06 May 2026 20:08:33 +0200
Newsgroups gmane.text.unicode.general
Message-ID <[email protected]>
--2NLEUBNABWUYNAOYGQEGPnhgwp
Content-Transfer-Encoding: quoted-printable
Content-Type: text/plain; charset=UTF-8

The ReadConsoleOutputW function, by definition, captures the tiles into the=
 lpBuffer, which is a random access array of CHAR_INFO structure, whose hor=
izontal and vertical size is specified in dwBufferSize. Since lpBuffer is r=
andom access, this implies one CHAR_INFO structure per character tile. The =
API therefore fundamentally imposes a strict memory layout that cannot be v=
iolated. The fact that some Unix-like terminals such as Windows Terminal ma=
y support features outside the scope of the CHAR_INFO structure for compati=
bility with ANSI escape codes or WSL programs does not invalidate the compa=
tibility considerations for legacy DOS/Win16/Win32 programs that require al=
l character tiles to fit in the CHAR_INFO structure for random access, beca=
use 4 byte CHAR_INFO structure of Win32 is intended to be a fully backwards=
 compatible extension of the 2 byte VGA text mode tile structure of DOS/Win=
16.  Dnia 06 maja 2026 19:58    Philippe Verdy via Unicode  &lt; unicode@co=
rp.unicode.org &gt;  napisa=C5=82(a): Windows can use other ways to map 16-=
bit codes in its *legacy* Console buffer (using old CHAR_INFO structure), i=
t can perfectly internally use compatibility characters, or PUAs of the BMP=
, and still present an API that exposes connforming sequences. You&#39;re t=
alking about an old implementation that was built even=C2=A0 long before th=
e Arabic script was extended (and newer scripts using contextual joining be=
haviors, that have never been part of the BMP, shcih as Adlam, and other sc=
ripts like Mongolian that also may need such sequences with ZWJ/ZWNJ contro=
ls, or with other formatting characters like those specific to Mongolian li=
ke FVS1...FVS4 and MVS, or those common to many Bhramic scripts, that the *=
legacy* Console did not support.  The *legacy* console was not built to sup=
port more than one plane (including many CJK cgaracters). The newer console=
 can!  Le=C2=A0mer. 6 mai 2026 =C3=A0=C2=A018:37,   [email protected]  &=
lt;  [email protected] &gt; a =C3=A9crit=C2=A0:  Have you read the L2/26=
-077 proposal? Using ZWJ or ZWNJ would not work for the compatibility purpo=
ses at all as already explained in the proposal. This is because ZWJ or ZWN=
J would take the space of one character tile in the CHAR_INFO structure. Su=
ppose that you&#39;re trying to map 0xD0 from FP164 to a sequence of U+FE7C=
 U+200D U+064B (=EF=B9=BC=E2=80=8D=D9=8B). The legacy application fills the=
 80=C3=9725 screen with all 0xD0 tiles. You subsequently try to capture the=
 tiles with a Win32 program by using ReadConsoleOutputA into an 80=C3=9725 =
buffer of 2000 tiles. This succeeds and captures 0xD0 into all the tiles. Y=
ou then try to capture the tiles using ReadConsoleOutputW into an 80=C3=972=
5 buffer. Each sequence U+FE7C U+200D U+064B would take a sequence of three=
 CHAR_INFO structures to store, meaning 6000 such structures for the whole =
screen. But the 80=C3=9725 buffer has only room for 2000 instances of the s=
tructure (one per character tile). Since CHAR_INFO stores 16-bit character =
code, by that same logic the compatibility characters would have to be in B=
MP for it to work. In Windows 95 Vietnamese and Windows 95 Thai, there are =
instances where one character tile takes multiple CHAR_INFO structures, cau=
sing visual width to be smaller than logical width, and when that happens, =
the remaining space at the end of the line is left blank, allowing for CP12=
58/CP874 combining characters in those systems to map 1:1 to their Unicode =
equivalents. Windows 3.1/95/98/ME Arabic don&#39;t work that way and don&#3=
9;t use combining characters or ZWJ sequences, so visual width is always eq=
uivalent to logical width, each character tile maps 1:1 to a CHAR_INFO stru=
cture and all characters may fill the entire line, which would be impossibl=
e if some of those characters were mapped to composition sequences or non-B=
MP characters. Since there is currently no sufficient evidence of user comm=
unity that would need to use those mappings, there are no plans for those c=
haracters to be added to Unicode, and therefore the only solution for the R=
eadConsoleOutputW to work properly in this case is to use agreed upon priva=
te use mappings for those compatibility characters.  Dnia 06 maja 2026 18:0=
6    Philippe Verdy via Unicode  &lt;   [email protected]  &gt;  nap=
isa=C5=82(a): You actually don&#39;t need any new compatibility characters =
for Arabic contextual forms, or for other contextual forms in other joining=
 scripts (like Adlam, or even Mongolian whichbis a LTR script).  You just h=
ave to prepend or append a ZWJ or ZWNJ formatting control to the unified le=
tter if you want to override its default contextual presentation form.   Le=
 mar. 5 mai 2026, 00:48, Asmus Freytag via Unicode &lt;  [email protected]=
de.org &gt; a =C3=A9crit=C2=A0:  The issue at hand is the distinction betwe=
en a theoretical gap and real-life problem.  You have demonstrated that the=
re are specifications that, if chained in the right way, can lead to ambigu=
ities or gaps in interchange.  What we don&#39;t have is an actual use case=
 with real-life consequences for a set of existing users, not hypothetical =
ones.  When it comes to encoding decisions based on existing documents, the=
re is a strong presumption that once sufficiently many documents exist that=
 contain a character, that this character will be needed in digitizing thes=
e documents, whether immediately, or eventually (e.g. in the case of future=
 scholarly studies). Also, the texts themselves exist, barring accidents, i=
n permanence. Therefore, it is justified to consider irrevocably allocating=
 a character that will map to this source in perpetuity, even though each e=
ncoded character carries a small cost for implementers.   However, when it =
comes to legacy characters, there&#39;s an additional cost that is imposed,=
 and that is based on the fact that characters that are encoded solely for =
compatibility will usually violate one or more of the other encoding princi=
ples, something that incrementally complicates the standard. Even for peopl=
e who never intend to use that character.   Therefore, the SEW is on solid =
ground when it demands not only a hypothetical scenario, but evidence of ac=
tual impact on actual users. Not only whether some application could invoke=
 an API, but whether such applications exist and are used today to access d=
ocuments encoded using the legacy characters in a way that is compromised i=
rreparably by not having an encoding for them.  A./   On 5/4/2026 10:24 AM,=
   [email protected]  via Unicode wrote:  In UTC 187 Minutes, &#34; Asmu=
s Freytag noted that the fact that lists of things existed in the past does=
 not make these things plain text. Ned Holbrook pointed out that the purpor=
ted issue occurs in a closed system, not in public interchange. &#34;. Howe=
ver, the arguments in the proposal do not merely hinge on the encodings bei=
ng lists of characters, but specifically points out methods to interchange =
text, including an example of copying terminal output and pasting to Notepa=
d, where the copying invokes the mapping of the current terminal codepage t=
o UCS-2 (as is CHAR_INFO compatible) and the pasting writes it into plain t=
ext. Win32 is also not a closed system, as Win32 can capture the tiles of t=
he output of Windows 3.1 Arabic DOS/Win16 programs and Windows 95/98/ME Ara=
bic DOS/Win16/Win32 programs, but Win32 can also interact with public text =
interchange systems by reading and writing to files and network. I&#39;m no=
t saying that Unicode absolutely must include those characters, but those k=
inds of misleading claims are causing users to misunderstand what the propo=
sal is about, and I don&#39;t want Unicode to be relying on uninformed deci=
sions to evaluate proposals.  Dnia 18 kwietnia 2026 13:36    piotrunio-2004=
@wp.pl via Unicode    &lt; [email protected] &gt;  napisa=C5=82(a): =
The SEW subsequently explained that the actual reason is due to insufficien=
t evidence of user community that would need to use the resulting mapping. =
Despite Win32 being a highly popular platform with plenty of backwards comp=
atibility and native UCS-2 terminal support, the specific use cases of inst=
alling codepages into Windows NT and using terminal tiles from Windows 3.1/=
95/98/ME are not sufficiently documented, making it difficult for any user =
communities to form around it. So it seems like the idea of standardizing l=
egacy Arabic terminal BMP mappings is a dead end for now.  Dnia 17 kwietnia=
 2026 22:59    [email protected] via Unicode    &lt; [email protected]=
e.org &gt;  napisa=C5=82(a): The Recommendations in L2/26-100 claim that Mi=
crosoft&#39;s documentation of legacy Arabic encodings is available at  lea=
rn.microsoft.com https://learn.microsoft.com/en-us/typography/legacy/legacy=
_arabic_fonts . However, that article only demonstrates two encodings of Tr=
ueType fonts, which are used in Windows 3.1 but are completely different fr=
om the eight terminal encodings. Unlike the TrueType encodings which repres=
ent internal shaping mappings and are not used for text interchange, the te=
rminal encodings have been demonstrated to be directly used in text interch=
ange through int 10h and ReadConsoleOutputA/WriteConsoleOutputA as already =
demonstrated in L2/26-077. The Recommendations also claim that the proposal=
 does not demonstrate any need for interchange or encoding, but the proposa=
l actually demonstrated such a need due to the logical extension of the Win=
32 terminal API to the functions ReadConsoleOutputW/WriteConsoleOutputW, wh=
ich are in Windows NT and may be used on the output of previously ran progr=
ams (including those that used the legacy Arabic terminal encodings), which=
 given the CHAR_INFO structure, therefore implies a need for all the tiles =
to map to BMP for interchange. I&#39;m not objecting to the SEW&#39;s concl=
usion of &#34;Users are expected to use PUA.&#34;, which can indeed be used=
 to provide a mapping even if not standardized, but the reasoning given was=
 flawed.  Dnia 09 stycznia 2026 17:25    [email protected]    &lt; piotr=
[email protected] &gt;  napisa=C5=82(a): The following Win32 C code will outp=
ut 256 characters in system console codepage into the character grid, captu=
re those character tiles in UCS-2 if possible, and then output the current =
console codepage number.   #include &lt;windows.h&gt;  #include &lt;stdio.h=
&gt;  int main(){  HANDLE hConsole=3DGetStdHandle(STD_OUTPUT_HANDLE);  CHAR=
_INFO screen[256];  COORD size=3D{16,16,};  COORD pos=3D{0,0,};  SMALL_RECT=
 rect=3D{0,0,15,15,};  for(int i=3D0;i&lt;256;i++){  screen[i].Attributes=
=3D0xF0;  screen[i].Char.AsciiChar=3Di;  }  WriteConsoleOutputA(hConsole,sc=
reen,size,pos,&amp;rect);  CHAR_INFO screenu[256];  if(ReadConsoleOutputW(h=
Console,screenu,size,pos,&amp;rect)){  for(int i=3D0;i&lt;256;i++) printf(&=
#34;%04X &#34;,screenu[i].Char.UnicodeChar);  }  else{  printf(&#34;error %=
08X\n&#34;,GetLastError());  }  printf(&#34;codepage %u&#34;,GetConsoleOutp=
utCP());  }   In most cases, whenever a legacy Win32 codepage is used, the =
application can run on Windows NT to capture the UCS-2 mapping of those cha=
racter cells to the BMP (although for CJK codepages a more complex setup wo=
uld be necessary due to thousands of fullwidth characters with 2-byte seque=
nces).   However, in Arabic versions of Windows 9x (95/98/ME) the resulting=
 character set has many presentation forms that are not in Unicode. This is=
 the result when running on Windows ME:=C2=A0 i.imgur.com https://i.imgur.c=
om/QFm3SkI.png =C2=A0in 10=C3=9720 font,  i.imgur.com https://i.imgur.com/K=
UbLQ0A.png =C2=A0in 10=C3=9718 font (same result also appears in Windows 95=
/98). 5=C3=9712, 7=C3=9712, 8=C3=9712, 10=C3=9718, 10=C3=9720, and 12=C3=97=
16 bitmap fonts have been attested with that character set (VGAOEM.FON, 851=
4OEM.FON, DOSAPP.FON). The 10=C3=9720 font has slightly different mapping t=
han the other sizes: 0x93 is =C3=B6 instead of =C3=B4, and 0x97 is missing =
(causing the following characters on the same line to be drawn at the wrong=
 position). It also claims to be using codepage 720, but many characters di=
ffer from their CP720 mappings, including the bundled=C2=A0CP_720.NLS mappi=
ngs (for example, =D9=80 (U+0640 ARABIC TATWEEL) is 0x95 in CP720, but in t=
he console 0x95 is =D8=B4 instead, and the tatweel is at 0xFF). On Windows =
9x,=C2=A0ReadConsoleOutputW is not supported so the UCS-2 mappings of the c=
onsole character tiles cannot be captured (error 0x00000078 ERROR_CALL_NOT_=
IMPLEMENTED).   When that program runs on Arabic versions of Windows NT, th=
e visual output is of the CP437 character set if one of the bundled bitmap =
fonts is used ( i.imgur.com https://i.imgur.com/RxjtxMH.png ), or the CP720=
 set if Lucida Console is used, with the Arabic letters either having glitc=
hy font substitution (NT 4.0, NT 5.0/2000) or the .notdef glyph (NT 5.1/XP =
and up). In fact, it seems that the only Arabic bitmap fonts that occur in =
Windows NT are CP1256 fonts, which are not used in terminals. So this appea=
rs to be one of those permanent Windows compatibility regressions that occu=
red when Windows 9x ended, where the terminals can no longer render legacy =
Arabic text. Even if the user managed to use registry hacks to set the font=
 to Courier New or Simplified Arabic Fixed, it would still use the CP720 ma=
pping which is not compatible with the Windows 9x set.   It appears that in=
 the Windows 9x Arabic terminal character set, 244 characters (=E2=80=87=EF=
=BA=80=EF=BA=81=EF=BA=82=EF=BA=83=EF=BA=84=EF=BA=85=EF=BA=87=EF=BA=88=EF=BA=
=8A=EF=BA=8B=EF=BA=8D=EF=BA=8E=EF=BA=8F=EF=BA=91=EF=BA=93=E2=96=BA=E2=97=84=
=E2=86=95=EF=BA=95=C2=B6=C2=A7=EF=BA=97=EF=BA=99=E2=86=91=E2=86=93=E2=86=92=
=E2=86=90=EF=BA=9B=EF=B9=B0=E2=96=B2=E2=96=BC !&#34;#$%&amp;&#39;()*+,-./01=
23456789:;&lt;=3D&gt;?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\]^_`abcdefghijklmnopqrst=
uvwxyz{|}~=EF=BA=9D=EF=BA=9F=EF=BA=A1=C3=A9=C3=A2=EF=BA=A3=C3=A0=EF=BA=A5=
=C3=A7=C3=AA=C3=AB=C3=A8=C3=AF=C3=AE=EF=BA=A7=EF=BA=A9=EF=BA=AB=EF=BA=AD=EF=
=BA=AF=C3=B4=EF=BA=B3=C3=BB=C3=B9=EF=BA=B7=EF=BA=BB=C2=A3=EF=BA=BF=EF=BB=81=
=EF=BB=85=EF=BB=89=EF=BB=8A=EF=BB=8B=EF=BB=8C are already in Unicode, but 1=
2 characters are not in Unicode:  =E2=80=A2 6 of them are pieces of lam-ale=
f ligatures (0xDD, 0xDE, 0xF9, 0xFB, 0xFC, 0xFD)  =E2=80=A2 2 of them are s=
hadda with fathatan ligatures without or with tatweel (0xD0, 0xD1)  =E2=80=
=94 in some legacy Microsoft fonts, shadda with fathatan is mapped to priva=
te use U+E818  =E2=80=A2 4 of them are disunifications of seen/sheen/sad/da=
d occuring either with or without tail  =E2=80=94=C2=A0=EF=B9=B3 (U+FE73 AR=
ABIC TAIL FRAGMENT) was originally encoded in Unicode 3.2 for CP864 compati=
bility; in that codepage, the forms of=C2=A0seen/sheen/sad/dad attach to th=
e tail fragment  =E2=80=94 forms with included tail:=C2=A00x92, 0x95, 0x98,=
 0x8A  =E2=80=94 forms without tail (attaching to tail fragment like in CP8=
64):=C2=A00xF3, 0xF4, 0xF5, 0xF6   If someone tried to make a Win32 console=
 implementation and tried to implement both Windows 9x Arabic terminal char=
acter set compatibility and wide string API (ReadConsoleOutputW) compatibil=
ity simultaneously, then they would run into the issue that there is curren=
tly no standardized mapping to handle that scenario. What should Windows 9x=
 Arabic console compatible implementations do in that case?=0D

--2NLEUBNABWUYNAOYGQEGPnhgwp
Content-Transfer-Encoding: quoted-printable
Content-Type: text/html; charset=UTF-8

<p>The ReadConsoleOutputW function, by definition, captures the tiles into =
the lpBuffer, which is a random access array of CHAR_INFO structure, whose =
horizontal and vertical size is specified in dwBufferSize. Since lpBuffer i=
s random access, this implies one CHAR_INFO structure per character tile. T=
he API therefore fundamentally imposes a strict memory layout that cannot b=
e violated. The fact that some Unix-like terminals such as Windows Terminal=
 may support features outside the scope of the CHAR_INFO structure for comp=
atibility with ANSI escape codes or WSL programs does not invalidate the co=
mpatibility considerations for legacy DOS/Win16/Win32 programs that require=
 all character tiles to fit in the CHAR_INFO structure for random access, b=
ecause 4 byte CHAR_INFO structure of Win32 is intended to be a fully backwa=
rds compatible extension of the 2 byte VGA text mode tile structure of DOS/=
Win16.</p><p><br></p><div class=3D"nh_extra"><blockquote style=3D"padding-t=
op: 12px;" class=3D"nh_qoute"><p style=3D"padding-bottom: 12px;"><strong>Dn=
ia 06 maja 2026 19:58</strong> <a href=3D"mailto:[email protected]" =
rel=3D"noopener noreferrer nofollow" target=3D"_blank"><span style=3D"margi=
n-left: 4px;">Philippe Verdy via Unicode</span></a><span style=3D"margin-le=
ft: 4px;"> &lt; [email protected] &gt;</span> napisa=C5=82(a):</p><d=
iv id=3D"gwp6bd5644c"><div id=3D"gwp6bd5644ch"><div class=3D"gwp6bd5644cb" =
data-message-body=3D"true" data-color-mode=3D"light"><div dir=3D"ltr">Windo=
ws can use other ways to map 16-bit codes in its *legacy* Console buffer (u=
sing old CHAR_INFO structure), it can perfectly internally use compatibilit=
y characters, or PUAs of the BMP, and still present an API that exposes con=
nforming sequences. You're talking about an old implementation that was bui=
lt even&nbsp; long before the Arabic script was extended (and newer scripts=
 using contextual joining behaviors, that have never been part of the BMP, =
shcih as Adlam, and other scripts like Mongolian that also may need such se=
quences with ZWJ/ZWNJ controls, or with other formatting characters like th=
ose specific to Mongolian like FVS1...FVS4 and MVS, or those common to many=
 Bhramic scripts, that the *legacy* Console did not support.<br>The *legacy=
* console was not built to support more than one plane (including many CJK =
cgaracters). The newer console can!</div><p><br></p><div class=3D"gwp6bd564=
4c_gmail_quote gwp6bd5644c_gmail_quote_container"><div dir=3D"ltr" class=3D=
"gwp6bd5644c_gmail_attr">Le&nbsp;mer. 6 mai 2026 =C3=A0&nbsp;18:37, <a href=
=3D"mailto:[email protected]" rel=3D"noopener noreferrer" target=3D"_bla=
nk">[email protected]</a> &lt;<a href=3D"mailto:[email protected]" re=
l=3D"noopener noreferrer" target=3D"_blank">[email protected]</a>&gt; a =
=C3=A9crit&nbsp;:<br></div><blockquote style=3D"margin: 0px 0px 0px 0.8ex; =
border-left: 1px solid rgb(204, 204, 204); padding-left: 1ex;" class=3D"gwp=
6bd5644c_gmail_quote"><p>Have you read the L2/26-077 proposal? Using ZWJ or=
 ZWNJ would not work for the compatibility purposes at all as already expla=
ined in the proposal. This is because ZWJ or ZWNJ would take the space of o=
ne character tile in the CHAR_INFO structure. Suppose that you're trying to=
 map 0xD0 from FP164 to a sequence of U+FE7C U+200D U+064B (=EF=B9=BC=E2=80=
=8D=D9=8B). The legacy application fills the 80=C3=9725 screen with all 0xD=
0 tiles. You subsequently try to capture the tiles with a Win32 program by =
using ReadConsoleOutputA into an 80=C3=9725 buffer of 2000 tiles. This succ=
eeds and captures 0xD0 into all the tiles. You then try to capture the tile=
s using ReadConsoleOutputW into an 80=C3=9725 buffer. Each sequence U+FE7C =
U+200D U+064B would take a sequence of three CHAR_INFO structures to store,=
 meaning 6000 such structures for the whole screen. But the 80=C3=9725 buff=
er has only room for 2000 instances of the structure (one per character til=
e). Since CHAR_INFO stores 16-bit character code, by that same logic the co=
mpatibility characters would have to be in BMP for it to work. In Windows 9=
5 Vietnamese and Windows 95 Thai, there are instances where one character t=
ile takes multiple CHAR_INFO structures, causing visual width to be smaller=
 than logical width, and when that happens, the remaining space at the end =
of the line is left blank, allowing for CP1258/CP874 combining characters i=
n those systems to map 1:1 to their Unicode equivalents. Windows 3.1/95/98/=
ME Arabic don't work that way and don't use combining characters or ZWJ seq=
uences, so visual width is always equivalent to logical width, each charact=
er tile maps 1:1 to a CHAR_INFO structure and all characters may fill the e=
ntire line, which would be impossible if some of those characters were mapp=
ed to composition sequences or non-BMP characters. Since there is currently=
 no sufficient evidence of user community that would need to use those mapp=
ings, there are no plans for those characters to be added to Unicode, and t=
herefore the only solution for the ReadConsoleOutputW to work properly in t=
his case is to use agreed upon private use mappings for those compatibility=
 characters.</p><p><br></p><div><blockquote style=3D"padding-top: 12px;"><p=
 style=3D"padding-bottom: 12px;"><strong>Dnia 06 maja 2026 18:06</strong> <=
a href=3D"mailto:[email protected]" rel=3D"noopener noreferrer nofol=
low" target=3D"_blank"><span style=3D"margin-left: 4px;">Philippe Verdy via=
 Unicode</span></a><span style=3D"margin-left: 4px;"> &lt; </span><a href=
=3D"mailto:[email protected]" rel=3D"noopener noreferrer" target=3D"=
_blank"><span style=3D"margin-left: 4px;">[email protected]</span></=
a><span style=3D"margin-left: 4px;"> &gt;</span> napisa=C5=82(a):</p><div i=
d=3D"gwp6bd5644c_m_-2650882641749569987gwpb67625a5"><div id=3D"gwp6bd5644c_=
m_-2650882641749569987gwpb67625a5h"><div><div dir=3D"auto"><p>You actually =
don't need any new compatibility characters for Arabic contextual forms, or=
 for other contextual forms in other joining scripts (like Adlam, or even M=
ongolian whichbis a LTR script).</p><div dir=3D"auto"><br></div><div dir=3D=
"auto">You just have to prepend or append a ZWJ or ZWNJ formatting control =
to the unified letter if you want to override its default contextual presen=
tation form.</div><div dir=3D"auto"><br></div></div><p><br></p><div><div di=
r=3D"ltr">Le mar. 5 mai 2026, 00:48, Asmus Freytag via Unicode &lt;<a href=
=3D"mailto:[email protected]" rel=3D"noopener noreferrer" target=3D"=
_blank">[email protected]</a>&gt; a =C3=A9crit&nbsp;:<br></div><bloc=
kquote style=3D"margin: 0px 0px 0px 0.8ex; border-left: 1px solid rgb(204, =
204, 204); padding-left: 1ex;"><div><p>The issue at hand is the distinction=
 between a theoretical gap and real-life problem.</p><p><br></p><p>You have=
 demonstrated that there are specifications that, if chained in the right w=
ay, can lead to ambiguities or gaps in interchange.</p><p><br></p><p>What w=
e don't have is an actual use case with real-life consequences for a set of=
 existing users, not hypothetical ones.</p><p><br></p><p>When it comes to e=
ncoding decisions based on existing documents, there is a strong presumptio=
n that once sufficiently many documents exist that contain a character, tha=
t this character will be needed in digitizing these documents, whether imme=
diately, or eventually (e.g. in the case of future scholarly studies). Also=
, the texts themselves exist, barring accidents, in permanence. Therefore, =
it is justified to consider irrevocably allocating a character that will ma=
p to this source in perpetuity, even though each encoded character carries =
a small cost for implementers.<br><br>However, when it comes to legacy char=
acters, there's an additional cost that is imposed, and that is based on th=
e fact that characters that are encoded solely for compatibility will usual=
ly violate one or more of the other encoding principles, something that inc=
rementally complicates the standard. Even for people who never intend to us=
e that character.<br><br>Therefore, the SEW is on solid ground when it dema=
nds not only a hypothetical scenario, but evidence of actual impact on actu=
al users. Not only whether some application could invoke an API, but whethe=
r such applications exist and are used today to access documents encoded us=
ing the legacy characters in a way that is compromised irreparably by not h=
aving an encoding for them.</p><p><br></p><p>A./</p><p><br></p><p><br></p><=
p>On 5/4/2026 10:24 AM, <a href=3D"mailto:[email protected]" rel=3D"nore=
ferrer" target=3D"_blank">[email protected]</a> via Unicode wrote:<br></=
p><blockquote type=3D"cite"><p>In UTC 187 Minutes, "<span style=3D"color: r=
gb(0, 0, 0); font-family: &quot;DMCA Sans Serif 10.0 dev1&quot;; font-size:=
 medium; font-style: normal; font-variant-ligatures: normal; font-variant-c=
aps: normal; font-weight: 400; letter-spacing: normal; text-align: start; t=
ext-indent: 0px; text-transform: none; white-space: normal; word-spacing: 0=
px; text-decoration-style: initial; text-decoration-color: initial; float: =
none; display: inline;">Asmus Freytag noted that the fact that lists of thi=
ngs existed in the past does not make these things plain text. Ned Holbrook=
 pointed out that the purported issue occurs in a closed system, not in pub=
lic interchange.</span>". However, the arguments in the proposal do not mer=
ely hinge on the encodings being lists of characters, but specifically poin=
ts out methods to interchange text, including an example of copying termina=
l output and pasting to Notepad, where the copying invokes the mapping of t=
he current terminal codepage to UCS-2 (as is CHAR_INFO compatible) and the =
pasting writes it into plain text. Win32 is also not a closed system, as Wi=
n32 can capture the tiles of the output of Windows 3.1 Arabic DOS/Win16 pro=
grams and Windows 95/98/ME Arabic DOS/Win16/Win32 programs, but Win32 can a=
lso interact with public text interchange systems by reading and writing to=
 files and network. I'm not saying that Unicode absolutely must include tho=
se characters, but those kinds of misleading claims are causing users to mi=
sunderstand what the proposal is about, and I don't want Unicode to be rely=
ing on uninformed decisions to evaluate proposals.</p><p><br></p><div><bloc=
kquote style=3D"padding-top: 12px;"><p style=3D"padding-bottom: 12px;"><str=
ong>Dnia 18 kwietnia 2026 13:36</strong> <a href=3D"mailto:[email protected]=
code.org" rel=3D"noopener noreferrer nofollow noreferrer" target=3D"_blank"=
><span style=3D"margin-left: 4px;">[email protected] via Unicode</span><=
/a><span style=3D"margin-left: 4px;"> </span><a href=3D"mailto:unicode@corp=
.unicode.org" rel=3D"noreferrer" target=3D"_blank"><span style=3D"margin-le=
ft: 4px;">&lt; [email protected] &gt;</span></a> napisa=C5=82(a):</p=
><div id=3D"gwp6bd5644c_m_-2650882641749569987gwpb67625a5_m_537108488707930=
4582gwpbf2d884a"><div id=3D"gwp6bd5644c_m_-2650882641749569987gwpb67625a5_m=
_5371084887079304582gwpbf2d884ah"><div><p>The SEW subsequently explained th=
at the actual reason is due to insufficient evidence of user community that=
 would need to use the resulting mapping. Despite Win32 being a highly popu=
lar platform with plenty of backwards compatibility and native UCS-2 termin=
al support, the specific use cases of installing codepages into Windows NT =
and using terminal tiles from Windows 3.1/95/98/ME are not sufficiently doc=
umented, making it difficult for any user communities to form around it. So=
 it seems like the idea of standardizing legacy Arabic terminal BMP mapping=
s is a dead end for now.</p><p><br></p><div><blockquote style=3D"padding-to=
p: 12px;"><p style=3D"padding-bottom: 12px;"><strong>Dnia 17 kwietnia 2026 =
22:59</strong> <a href=3D"mailto:[email protected]" rel=3D"noopener =
noreferrer nofollow noreferrer" target=3D"_blank"><span style=3D"margin-lef=
t: 4px;">[email protected] via Unicode</span></a><span style=3D"margin-l=
eft: 4px;"> </span><a href=3D"mailto:[email protected]" rel=3D"noref=
errer" target=3D"_blank"><span style=3D"margin-left: 4px;">&lt; unicode@cor=
p.unicode.org &gt;</span></a> napisa=C5=82(a):</p><div id=3D"gwp6bd5644c_m_=
-2650882641749569987gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f=
8"><div id=3D"gwp6bd5644c_m_-2650882641749569987gwpb67625a5_m_5371084887079=
304582gwpbf2d884a_gwp2281a7f8h"><div><p>The Recommendations in L2/26-100 cl=
aim that Microsoft's documentation of legacy Arabic encodings is available =
at <a href=3D"https://learn.microsoft.com/en-us/typography/legacy/legacy_ar=
abic_fonts" =3D"" rel=3D"noreferrer" target=3D"_blank">https://learn.micros=
oft.com/en-us/typography/legacy/legacy_arabic_fonts</a>. However, that arti=
cle only demonstrates two encodings of TrueType fonts, which are used in Wi=
ndows 3.1 but are completely different from the eight terminal encodings. U=
nlike the TrueType encodings which represent internal shaping mappings and =
are not used for text interchange, the terminal encodings have been demonst=
rated to be directly used in text interchange through int 10h and ReadConso=
leOutputA/WriteConsoleOutputA as already demonstrated in L2/26-077. The Rec=
ommendations also claim that the proposal does not demonstrate any need for=
 interchange or encoding, but the proposal actually demonstrated such a nee=
d due to the logical extension of the Win32 terminal API to the functions R=
eadConsoleOutputW/WriteConsoleOutputW, which are in Windows NT and may be u=
sed on the output of previously ran programs (including those that used the=
 legacy Arabic terminal encodings), which given the CHAR_INFO structure, th=
erefore implies a need for all the tiles to map to BMP for interchange. I'm=
 not objecting to the SEW's conclusion of "Users are expected to use PUA.",=
 which can indeed be used to provide a mapping even if not standardized, bu=
t the reasoning given was flawed.</p><p><br></p><div><blockquote style=3D"p=
adding-top: 12px;"><p style=3D"padding-bottom: 12px;"><strong>Dnia 09 stycz=
nia 2026 17:25</strong> <a href=3D"mailto:[email protected]" rel=3D"nore=
ferrer" target=3D"_blank"><span style=3D"margin-left: 4px;">piotrunio-2004@=
wp.pl</span></a><span style=3D"margin-left: 4px;"> </span><a href=3D"mailto=
:[email protected]" rel=3D"noreferrer" target=3D"_blank"><span style=3D"=
margin-left: 4px;">&lt; [email protected] &gt;</span></a> napisa=C5=82(a=
):</p><div id=3D"gwp6bd5644c_m_-2650882641749569987gwpb67625a5_m_5371084887=
079304582gwpbf2d884a_gwp2281a7f8_gwpa05276c7"><div id=3D"gwp6bd5644c_m_-265=
0882641749569987gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f8_gw=
pa05276c7h"><div><div id=3D"gwp6bd5644c_m_-2650882641749569987gwpb67625a5_m=
_5371084887079304582gwpbf2d884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718"><div i=
d=3D"gwp6bd5644c_m_-2650882641749569987gwpb67625a5_m_5371084887079304582gwp=
bf2d884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718h"><div><div id=3D"gwp6bd5644c_=
m_-2650882641749569987gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a=
7f8_gwpa05276c7_gwpa8b5f718_gwpa8b5f718"><div id=3D"gwp6bd5644c_m_-26508826=
41749569987gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f8_gwpa052=
76c7_gwpa8b5f718_gwpa8b5f718h"><div><div id=3D"gwp6bd5644c_m_-2650882641749=
569987gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f8_gwpa05276c7_=
gwpa8b5f718_gwpa8b5f718_gwpa8b5f718"><div id=3D"gwp6bd5644c_m_-265088264174=
9569987gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f8_gwpa05276c7=
_gwpa8b5f718_gwpa8b5f718_gwpa8b5f718h"><div><p>The following Win32 C code w=
ill output 256 characters in system console codepage into the character gri=
d, capture those character tiles in UCS-2 if possible, and then output the =
current console codepage number.<br></p><p><br></p><p>#include &lt;windows.=
h&gt;<br>#include &lt;stdio.h&gt;<br>int main(){<br>HANDLE hConsole=3DGetSt=
dHandle(STD_OUTPUT_HANDLE);<br>CHAR_INFO screen[256];<br>COORD size=3D{16,1=
6,};<br>COORD pos=3D{0,0,};<br>SMALL_RECT rect=3D{0,0,15,15,};<br>for(int i=
=3D0;i&lt;256;i++){<br>screen[i].Attributes=3D0xF0;<br>screen[i].Char.Ascii=
Char=3Di;<br>}<br>WriteConsoleOutputA(hConsole,screen,size,pos,&amp;rect);<=
br>CHAR_INFO screenu[256];<br>if(ReadConsoleOutputW(hConsole,screenu,size,p=
os,&amp;rect)){<br>for(int i=3D0;i&lt;256;i++) printf("%04X ",screenu[i].Ch=
ar.UnicodeChar);<br>}<br>else{<br>printf("error %08X\n",GetLastError());<br=
>}<br>printf("codepage %u",GetConsoleOutputCP());<br>}<br><br></p><p>In mos=
t cases, whenever a legacy Win32 codepage is used, the application can run =
on Windows NT to capture the UCS-2 mapping of those character cells to the =
BMP (although for CJK codepages a more complex setup would be necessary due=
 to thousands of fullwidth characters with 2-byte sequences).<br></p><p><br=
></p><p>However, in Arabic versions of Windows 9x (95/98/ME) the resulting =
character set has many presentation forms that are not in Unicode. This is =
the result when running on Windows ME:&nbsp;<a href=3D"https://i.imgur.com/=
QFm3SkI.png" =3D"" rel=3D"noopener noreferrer noreferrer" target=3D"_blank"=
>https://i.imgur.com/QFm3SkI.png</a>&nbsp;in 10=C3=9720 font, <a href=3D"ht=
tps://i.imgur.com/KUbLQ0A.png" =3D"" rel=3D"noopener noreferrer noreferrer"=
 target=3D"_blank">https://i.imgur.com/KUbLQ0A.png</a>&nbsp;in 10=C3=9718 f=
ont (same result also appears in Windows 95/98). 5=C3=9712, 7=C3=9712, 8=C3=
=9712, 10=C3=9718, 10=C3=9720, and 12=C3=9716 bitmap fonts have been attest=
ed with that character set (VGAOEM.FON, 8514OEM.FON, DOSAPP.FON). The 10=C3=
=9720 font has slightly different mapping than the other sizes: 0x93 is =C3=
=B6 instead of =C3=B4, and 0x97 is missing (causing the following character=
s on the same line to be drawn at the wrong position). It also claims to be=
 using codepage 720, but many characters differ from their CP720 mappings, =
including the bundled&nbsp;CP_720.NLS mappings (for example, =D9=80 (U+0640=
 ARABIC TATWEEL) is 0x95 in CP720, but in the console 0x95 is =D8=B4 instea=
d, and the tatweel is at 0xFF). On Windows 9x,&nbsp;ReadConsoleOutputW is n=
ot supported so the UCS-2 mappings of the console character tiles cannot be=
 captured (error 0x00000078 ERROR_CALL_NOT_IMPLEMENTED).<br></p></div></div=
></div><p><br></p><p>When that program runs on Arabic versions of Windows N=
T, the visual output is of the CP437 character set if one of the bundled bi=
tmap fonts is used (<a href=3D"https://i.imgur.com/RxjtxMH.png" =3D"" rel=
=3D"noopener noreferrer noreferrer" target=3D"_blank">https://i.imgur.com/R=
xjtxMH.png</a>), or the CP720 set if Lucida Console is used, with the Arabi=
c letters either having glitchy font substitution (NT 4.0, NT 5.0/2000) or =
the .notdef glyph (NT 5.1/XP and up). In fact, it seems that the only Arabi=
c bitmap fonts that occur in Windows NT are CP1256 fonts, which are not use=
d in terminals. So this appears to be one of those permanent Windows compat=
ibility regressions that occured when Windows 9x ended, where the terminals=
 can no longer render legacy Arabic text. Even if the user managed to use r=
egistry hacks to set the font to Courier New or Simplified Arabic Fixed, it=
 would still use the CP720 mapping which is not compatible with the Windows=
 9x set.<br></p></div></div><p><br></p></div></div></div></div><div id=3D"g=
wp6bd5644c_m_-2650882641749569987gwpb67625a5_m_5371084887079304582gwpbf2d88=
4a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718"><div id=3D"gwp6bd5644c_m_-265088264=
1749569987gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f8_gwpa0527=
6c7_gwpa8b5f718h"><div><p>It appears that in the Windows 9x Arabic terminal=
 character set, 244 characters (=E2=80=87=EF=BA=80=EF=BA=81=EF=BA=82=EF=BA=
=83=EF=BA=84=EF=BA=85=EF=BA=87=EF=BA=88=EF=BA=8A=EF=BA=8B=EF=BA=8D=EF=BA=8E=
=EF=BA=8F=EF=BA=91=EF=BA=93=E2=96=BA=E2=97=84=E2=86=95=EF=BA=95=C2=B6=C2=A7=
=EF=BA=97=EF=BA=99=E2=86=91=E2=86=93=E2=86=92=E2=86=90=EF=BA=9B=EF=B9=B0=E2=
=96=B2=E2=96=BC !"#$%&amp;'()*+,-./0123456789:;&lt;=3D&gt;?@ABCDEFGHIJKLMNO=
PQRSTUVWXYZ[\]^_`abcdefghijklmnopqrstuvwxyz{|}~=EF=BA=9D=EF=BA=9F=EF=BA=A1=
=C3=A9=C3=A2=EF=BA=A3=C3=A0=EF=BA=A5=C3=A7=C3=AA=C3=AB=C3=A8=C3=AF=C3=AE=EF=
=BA=A7=EF=BA=A9=EF=BA=AB=EF=BA=AD=EF=BA=AF=C3=B4=EF=BA=B3=C3=BB=C3=B9=EF=BA=
=B7=EF=BA=BB=C2=A3=EF=BA=BF=EF=BB=81=EF=BB=85=EF=BB=89=EF=BB=8A=EF=BB=8B=EF=
=BB=8C=EF=BB=8D=EF=BB=8E=EF=BB=8F=EF=BB=90=EF=BB=91=EF=BB=93=EF=BB=95=EF=BB=
=97=EF=BB=99=EF=BB=9B=C2=AB=C2=BB=EF=B9=B1=E2=96=92=EF=B9=B2=E2=94=82=E2=94=
=A4=EF=B9=B4=EF=B9=B6=EF=B9=B7=EF=B9=B8=D9=A0=D9=A1=D9=A2=D9=A3=EF=B9=B9=EF=
=B9=BA=E2=94=90=E2=94=94=E2=94=B4=E2=94=AC=E2=94=9C=E2=94=80=E2=94=BC=EF=B9=
=BB=EF=B9=BE=D9=A4=D9=A5=D9=A6=D9=A7=D9=A8=D9=A9=D8=8C=EF=B9=BF=EF=B1=9E=EF=
=B1=9F=EF=B1=A0=EF=B3=B2=EF=B1=A1=EF=B3=B3=EF=B1=A2=E2=94=98=E2=94=8C=D8=9B=
=D8=9F=C2=A4=EF=BB=9D=EF=BB=9F=EF=BB=A1=EF=BB=A3=EF=BB=A5=EF=BB=A7=C2=B5=EF=
=BB=A9=EF=BB=AB=EF=BB=AC=EF=BB=AD=EF=BB=AF=EF=BB=B0=EF=BB=B1=EF=BB=B2=EF=BB=
=B3=EF=B3=B4=EF=B9=BC=EF=B9=BD=EF=BA=B1=EF=BA=B5=EF=BA=B9=EF=BA=BD=EF=B9=B3=
=C2=B0=C2=B7=E2=96=A0=D9=80) are already in Unicode, but 12 characters are =
not in Unicode:<br></p><p>=E2=80=A2 6 of them are pieces of lam-alef ligatu=
res (0xDD, 0xDE, 0xF9, 0xFB, 0xFC, 0xFD)<br></p><p>=E2=80=A2 2 of them are =
shadda with fathatan ligatures without or with tatweel (0xD0, 0xD1)<br></p>=
<p>=E2=80=94 in some legacy Microsoft fonts, shadda with fathatan is mapped=
 to private use U+E818<br></p><p>=E2=80=A2 4 of them are disunifications of=
 seen/sheen/sad/dad occuring either with or without tail<br></p><p>=E2=80=
=94&nbsp;=EF=B9=B3 (U+FE73 ARABIC TAIL FRAGMENT) was originally encoded in =
Unicode 3.2 for CP864 compatibility; in that codepage, the forms of&nbsp;se=
en/sheen/sad/dad attach to the tail fragment<br></p><p>=E2=80=94 forms with=
 included tail:&nbsp;0x92, 0x95, 0x98, 0x8A<br></p><p>=E2=80=94 forms witho=
ut tail (attaching to tail fragment like in CP864):&nbsp;0xF3, 0xF4, 0xF5, =
0xF6<br></p></div><p><br></p></div><p>If someone tried to make a Win32 cons=
ole implementation and tried to implement both Windows 9x Arabic terminal c=
haracter set compatibility and wide string API (ReadConsoleOutputW) compati=
bility simultaneously, then they would run into the issue that there is cur=
rently no standardized mapping to handle that scenario. What should Windows=
 9x Arabic console compatible implementations do in that case?<br></p></div=
><p><br></p></div></div></div></blockquote></div><p><br></p></div></div></d=
iv></blockquote></div><p><br></p></div></div></div></blockquote></div><p><b=
r></p></blockquote><p><br></p></div></blockquote></div></div></div></div></=
blockquote></div><p><br></p></blockquote></div></div></div></div></blockquo=
te></div><p><br></p>
--2NLEUBNABWUYNAOYGQEGPnhgwp--