| Newsgroups |
gmane.text.unicode.general |
| Message-ID |
<[email protected]> |
--2JBJUNXFBCKBRPRSPVHBWnhgwp
Content-Transfer-Encoding: quoted-printable
Content-Type: text/plain; charset=UTF-8
Have you read the L2/26-077 proposal? Using ZWJ or ZWNJ would not work for =
the compatibility purposes at all as already explained in the proposal. Thi=
s is because ZWJ or ZWNJ would take the space of one character tile in the =
CHAR_INFO structure. Suppose that you're trying to map 0xD0 from FP164 =
to a sequence of U+FE7C U+200D U+064B (=EF=B9=BC=E2=80=8D=D9=8B). The legac=
y application fills the 80=C3=9725 screen with all 0xD0 tiles. You subseque=
ntly try to capture the tiles with a Win32 program by using ReadConsoleOutp=
utA into an 80=C3=9725 buffer of 2000 tiles. This succeeds and captures 0xD=
0 into all the tiles. You then try to capture the tiles using ReadConsoleOu=
tputW into an 80=C3=9725 buffer. Each sequence U+FE7C U+200D U+064B would t=
ake a sequence of three CHAR_INFO structures to store, meaning 6000 such st=
ructures for the whole screen. But the 80=C3=9725 buffer has only room for =
2000 instances of the structure (one per character tile). Since CHAR_INFO s=
tores 16-bit character code, by that same logic the compatibility character=
s would have to be in BMP for it to work. In Windows 95 Vietnamese and Wind=
ows 95 Thai, there are instances where one character tile takes multiple CH=
AR_INFO structures, causing visual width to be smaller than logical width, =
and when that happens, the remaining space at the end of the line is left b=
lank, allowing for CP1258/CP874 combining characters in those systems to ma=
p 1:1 to their Unicode equivalents. Windows 3.1/95/98/ME Arabic don't w=
ork that way and don't use combining characters or ZWJ sequences, so vi=
sual width is always equivalent to logical width, each character tile maps =
1:1 to a CHAR_INFO structure and all characters may fill the entire line, w=
hich would be impossible if some of those characters were mapped to composi=
tion sequences or non-BMP characters. Since there is currently no sufficien=
t evidence of user community that would need to use those mappings, there a=
re no plans for those characters to be added to Unicode, and therefore the =
only solution for the ReadConsoleOutputW to work properly in this case is t=
o use agreed upon private use mappings for those compatibility characters. =
Dnia 06 maja 2026 18:06 Philippe Verdy via Unicode < [email protected]=
nicode.org > napisa=C5=82(a): You actually don't need any new compa=
tibility characters for Arabic contextual forms, or for other contextual fo=
rms in other joining scripts (like Adlam, or even Mongolian whichbis a LTR =
script). You just have to prepend or append a ZWJ or ZWNJ formatting contr=
ol to the unified letter if you want to override its default contextual pre=
sentation form. Le mar. 5 mai 2026, 00:48, Asmus Freytag via Unicode <=
[email protected] > a =C3=A9crit=C2=A0: The issue at hand is t=
he distinction between a theoretical gap and real-life problem. You have d=
emonstrated that there are specifications that, if chained in the right way=
, can lead to ambiguities or gaps in interchange. What we don't have i=
s an actual use case with real-life consequences for a set of existing user=
s, not hypothetical ones. When it comes to encoding decisions based on exi=
sting documents, there is a strong presumption that once sufficiently many =
documents exist that contain a character, that this character will be neede=
d in digitizing these documents, whether immediately, or eventually (e.g. i=
n the case of future scholarly studies). Also, the texts themselves exist, =
barring accidents, in permanence. Therefore, it is justified to consider ir=
revocably allocating a character that will map to this source in perpetuity=
, even though each encoded character carries a small cost for implementers.=
However, when it comes to legacy characters, there's an additional c=
ost that is imposed, and that is based on the fact that characters that are=
encoded solely for compatibility will usually violate one or more of the o=
ther encoding principles, something that incrementally complicates the stan=
dard. Even for people who never intend to use that character. Therefore, =
the SEW is on solid ground when it demands not only a hypothetical scenario=
, but evidence of actual impact on actual users. Not only whether some appl=
ication could invoke an API, but whether such applications exist and are us=
ed today to access documents encoded using the legacy characters in a way t=
hat is compromised irreparably by not having an encoding for them. A./ O=
n 5/4/2026 10:24 AM, [email protected] via Unicode wrote: In UTC 187=
Minutes, " Asmus Freytag noted that the fact that lists of things exis=
ted in the past does not make these things plain text. Ned Holbrook pointed=
out that the purported issue occurs in a closed system, not in public inte=
rchange. ". However, the arguments in the proposal do not merely hinge =
on the encodings being lists of characters, but specifically points out met=
hods to interchange text, including an example of copying terminal output a=
nd pasting to Notepad, where the copying invokes the mapping of the current=
terminal codepage to UCS-2 (as is CHAR_INFO compatible) and the pasting wr=
ites it into plain text. Win32 is also not a closed system, as Win32 can ca=
pture the tiles of the output of Windows 3.1 Arabic DOS/Win16 programs and =
Windows 95/98/ME Arabic DOS/Win16/Win32 programs, but Win32 can also intera=
ct with public text interchange systems by reading and writing to files and=
network. I'm not saying that Unicode absolutely must include those cha=
racters, but those kinds of misleading claims are causing users to misunder=
stand what the proposal is about, and I don't want Unicode to be relyin=
g on uninformed decisions to evaluate proposals. Dnia 18 kwietnia 2026 13:=
36 [email protected] via Unicode < [email protected] >=
; napisa=C5=82(a): The SEW subsequently explained that the actual reason i=
s due to insufficient evidence of user community that would need to use the=
resulting mapping. Despite Win32 being a highly popular platform with plen=
ty of backwards compatibility and native UCS-2 terminal support, the specif=
ic use cases of installing codepages into Windows NT and using terminal til=
es from Windows 3.1/95/98/ME are not sufficiently documented, making it dif=
ficult for any user communities to form around it. So it seems like the ide=
a of standardizing legacy Arabic terminal BMP mappings is a dead end for no=
w. Dnia 17 kwietnia 2026 22:59 [email protected] via Unicode <=
[email protected] > napisa=C5=82(a): The Recommendations in L2/=
26-100 claim that Microsoft's documentation of legacy Arabic encodings =
is available at learn.microsoft.com https://learn.microsoft.com/en-us/typo=
graphy/legacy/legacy_arabic_fonts . However, that article only demonstrates=
two encodings of TrueType fonts, which are used in Windows 3.1 but are com=
pletely different from the eight terminal encodings. Unlike the TrueType en=
codings which represent internal shaping mappings and are not used for text=
interchange, the terminal encodings have been demonstrated to be directly =
used in text interchange through int 10h and ReadConsoleOutputA/WriteConsol=
eOutputA as already demonstrated in L2/26-077. The Recommendations also cla=
im that the proposal does not demonstrate any need for interchange or encod=
ing, but the proposal actually demonstrated such a need due to the logical =
extension of the Win32 terminal API to the functions ReadConsoleOutputW/Wri=
teConsoleOutputW, which are in Windows NT and may be used on the output of =
previously ran programs (including those that used the legacy Arabic termin=
al encodings), which given the CHAR_INFO structure, therefore implies a nee=
d for all the tiles to map to BMP for interchange. I'm not objecting to=
the SEW's conclusion of "Users are expected to use PUA.", whic=
h can indeed be used to provide a mapping even if not standardized, but the=
reasoning given was flawed. Dnia 09 stycznia 2026 17:25 piotrunio-2004=
@wp.pl < [email protected] > napisa=C5=82(a): The following Wi=
n32 C code will output 256 characters in system console codepage into the c=
haracter grid, capture those character tiles in UCS-2 if possible, and then=
output the current console codepage number. #include <windows.h> =
#include <stdio.h> int main(){ HANDLE hConsole=3DGetStdHandle(STD_O=
UTPUT_HANDLE); CHAR_INFO screen[256]; COORD size=3D{16,16,}; COORD pos=
=3D{0,0,}; SMALL_RECT rect=3D{0,0,15,15,}; for(int i=3D0;i<256;i++){ =
screen[i].Attributes=3D0xF0; screen[i].Char.AsciiChar=3Di; } WriteConsol=
eOutputA(hConsole,screen,size,pos,&rect); CHAR_INFO screenu[256]; if(=
ReadConsoleOutputW(hConsole,screenu,size,pos,&rect)){ for(int i=3D0;i&=
lt;256;i++) printf("%04X ",screenu[i].Char.UnicodeChar); } else{ =
printf("error %08X\n",GetLastError()); } printf("codepage %u=
",GetConsoleOutputCP()); } In most cases, whenever a legacy Win32 co=
depage is used, the application can run on Windows NT to capture the UCS-2 =
mapping of those character cells to the BMP (although for CJK codepages a m=
ore complex setup would be necessary due to thousands of fullwidth characte=
rs with 2-byte sequences). However, in Arabic versions of Windows 9x (95/=
98/ME) the resulting character set has many presentation forms that are not=
in Unicode. This is the result when running on Windows ME:=C2=A0 i.imgur.c=
om https://i.imgur.com/QFm3SkI.png =C2=A0in 10=C3=9720 font, i.imgur.com h=
ttps://i.imgur.com/KUbLQ0A.png =C2=A0in 10=C3=9718 font (same result also a=
ppears in Windows 95/98). 5=C3=9712, 7=C3=9712, 8=C3=9712, 10=C3=9718, 10=
=C3=9720, and 12=C3=9716 bitmap fonts have been attested with that characte=
r set (VGAOEM.FON, 8514OEM.FON, DOSAPP.FON). The 10=C3=9720 font has slight=
ly different mapping than the other sizes: 0x93 is =C3=B6 instead of =C3=B4=
, and 0x97 is missing (causing the following characters on the same line to=
be drawn at the wrong position). It also claims to be using codepage 720, =
but many characters differ from their CP720 mappings, including the bundled=
=C2=A0CP_720.NLS mappings (for example, =D9=80 (U+0640 ARABIC TATWEEL) is 0=
x95 in CP720, but in the console 0x95 is =D8=B4 instead, and the tatweel is=
at 0xFF). On Windows 9x,=C2=A0ReadConsoleOutputW is not supported so the U=
CS-2 mappings of the console character tiles cannot be captured (error 0x00=
000078 ERROR_CALL_NOT_IMPLEMENTED). When that program runs on Arabic vers=
ions of Windows NT, the visual output is of the CP437 character set if one =
of the bundled bitmap fonts is used ( i.imgur.com https://i.imgur.com/Rxjtx=
MH.png ), or the CP720 set if Lucida Console is used, with the Arabic lette=
rs either having glitchy font substitution (NT 4.0, NT 5.0/2000) or the .no=
tdef glyph (NT 5.1/XP and up). In fact, it seems that the only Arabic bitma=
p fonts that occur in Windows NT are CP1256 fonts, which are not used in te=
rminals. So this appears to be one of those permanent Windows compatibility=
regressions that occured when Windows 9x ended, where the terminals can no=
longer render legacy Arabic text. Even if the user managed to use registry=
hacks to set the font to Courier New or Simplified Arabic Fixed, it would =
still use the CP720 mapping which is not compatible with the Windows 9x set=
. It appears that in the Windows 9x Arabic terminal character set, 244 ch=
aracters (=E2=80=87=EF=BA=80=EF=BA=81=EF=BA=82=EF=BA=83=EF=BA=84=EF=BA=85=
=EF=BA=87=EF=BA=88=EF=BA=8A=EF=BA=8B=EF=BA=8D=EF=BA=8E=EF=BA=8F=EF=BA=91=EF=
=BA=93=E2=96=BA=E2=97=84=E2=86=95=EF=BA=95=C2=B6=C2=A7=EF=BA=97=EF=BA=99=E2=
=86=91=E2=86=93=E2=86=92=E2=86=90=EF=BA=9B=EF=B9=B0=E2=96=B2=E2=96=BC !"=
;#$%&'()*+,-./0123456789:;<=3D>?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\=
]^_`abcdefghijklmnopqrstuvwxyz{|}~=EF=BA=9D=EF=BA=9F=EF=BA=A1=C3=A9=C3=A2=
=EF=BA=A3=C3=A0=EF=BA=A5=C3=A7=C3=AA=C3=AB=C3=A8=C3=AF=C3=AE=EF=BA=A7=EF=BA=
=A9=EF=BA=AB=EF=BA=AD=EF=BA=AF=C3=B4=EF=BA=B3=C3=BB=C3=B9=EF=BA=B7=EF=BA=BB=
=C2=A3=EF=BA=BF=EF=BB=81=EF=BB=85=EF=BB=89=EF=BB=8A=EF=BB=8B=EF=BB=8C are a=
lready in Unicode, but 12 characters are not in Unicode: =E2=80=A2 6 of th=
em are pieces of lam-alef ligatures (0xDD, 0xDE, 0xF9, 0xFB, 0xFC, 0xFD) =
=E2=80=A2 2 of them are shadda with fathatan ligatures without or with tatw=
eel (0xD0, 0xD1) =E2=80=94 in some legacy Microsoft fonts, shadda with fat=
hatan is mapped to private use U+E818 =E2=80=A2 4 of them are disunificati=
ons of seen/sheen/sad/dad occuring either with or without tail =E2=80=94=
=C2=A0=EF=B9=B3 (U+FE73 ARABIC TAIL FRAGMENT) was originally encoded in Uni=
code 3.2 for CP864 compatibility; in that codepage, the forms of=C2=A0seen/=
sheen/sad/dad attach to the tail fragment =E2=80=94 forms with included ta=
il:=C2=A00x92, 0x95, 0x98, 0x8A =E2=80=94 forms without tail (attaching to=
tail fragment like in CP864):=C2=A00xF3, 0xF4, 0xF5, 0xF6 If someone tri=
ed to make a Win32 console implementation and tried to implement both Windo=
ws 9x Arabic terminal character set compatibility and wide string API (Read=
ConsoleOutputW) compatibility simultaneously, then they would run into the =
issue that there is currently no standardized mapping to handle that scenar=
io. What should Windows 9x Arabic console compatible implementations do in =
that case?=0D
--2JBJUNXFBCKBRPRSPVHBWnhgwp
Content-Transfer-Encoding: quoted-printable
Content-Type: text/html; charset=UTF-8
<p>Have you read the L2/26-077 proposal? Using ZWJ or ZWNJ would not work f=
or the compatibility purposes at all as already explained in the proposal. =
This is because ZWJ or ZWNJ would take the space of one character tile in t=
he CHAR_INFO structure. Suppose that you're trying to map 0xD0 from FP164 t=
o a sequence of U+FE7C U+200D U+064B (=EF=B9=BC=E2=80=8D=D9=8B). The legacy=
application fills the 80=C3=9725 screen with all 0xD0 tiles. You subsequen=
tly try to capture the tiles with a Win32 program by using ReadConsoleOutpu=
tA into an 80=C3=9725 buffer of 2000 tiles. This succeeds and captures 0xD0=
into all the tiles. You then try to capture the tiles using ReadConsoleOut=
putW into an 80=C3=9725 buffer. Each sequence U+FE7C U+200D U+064B would ta=
ke a sequence of three CHAR_INFO structures to store, meaning 6000 such str=
uctures for the whole screen. But the 80=C3=9725 buffer has only room for 2=
000 instances of the structure (one per character tile). Since CHAR_INFO st=
ores 16-bit character code, by that same logic the compatibility characters=
would have to be in BMP for it to work. In Windows 95 Vietnamese and Windo=
ws 95 Thai, there are instances where one character tile takes multiple CHA=
R_INFO structures, causing visual width to be smaller than logical width, a=
nd when that happens, the remaining space at the end of the line is left bl=
ank, allowing for CP1258/CP874 combining characters in those systems to map=
1:1 to their Unicode equivalents. Windows 3.1/95/98/ME Arabic don't work t=
hat way and don't use combining characters or ZWJ sequences, so visual widt=
h is always equivalent to logical width, each character tile maps 1:1 to a =
CHAR_INFO structure and all characters may fill the entire line, which woul=
d be impossible if some of those characters were mapped to composition sequ=
ences or non-BMP characters. Since there is currently no sufficient evidenc=
e of user community that would need to use those mappings, there are no pla=
ns for those characters to be added to Unicode, and therefore the only solu=
tion for the ReadConsoleOutputW to work properly in this case is to use agr=
eed upon private use mappings for those compatibility characters.</p><p><br=
></p><div class=3D"nh_extra"><blockquote style=3D"padding-top: 12px;" class=
=3D"nh_qoute"><p style=3D"padding-bottom: 12px;"><strong>Dnia 06 maja 2026 =
18:06</strong> <a href=3D"mailto:[email protected]" rel=3D"noopener =
noreferrer nofollow" target=3D"_blank"><span style=3D"margin-left: 4px;">Ph=
ilippe Verdy via Unicode</span></a><span style=3D"margin-left: 4px;"> < =
[email protected] ></span> napisa=C5=82(a):</p><div id=3D"gwpb676=
25a5"><div id=3D"gwpb67625a5h"><div class=3D"gwpb67625a5b" data-message-bod=
y=3D"true" data-color-mode=3D"light"><div dir=3D"auto"><p>You actually don'=
t need any new compatibility characters for Arabic contextual forms, or for=
other contextual forms in other joining scripts (like Adlam, or even Mongo=
lian whichbis a LTR script).</p><div dir=3D"auto"><br></div><div dir=3D"aut=
o">You just have to prepend or append a ZWJ or ZWNJ formatting control to t=
he unified letter if you want to override its default contextual presentati=
on form.</div><div dir=3D"auto"><br></div></div><p><br></p><div class=3D"gw=
pb67625a5_gmail_quote gwpb67625a5_gmail_quote_container"><div dir=3D"ltr" c=
lass=3D"gwpb67625a5_gmail_attr">Le mar. 5 mai 2026, 00:48, Asmus Freytag vi=
a Unicode <<a href=3D"mailto:[email protected]" rel=3D"noopener n=
oreferrer" target=3D"_blank">[email protected]</a>> a =C3=A9crit&=
nbsp;:<br></div><blockquote style=3D"margin: 0px 0px 0px 0.8ex; border-left=
: 1px solid rgb(204, 204, 204); padding-left: 1ex;" class=3D"gwpb67625a5_gm=
ail_quote"><div><p>The issue at hand is the distinction between a theoretic=
al gap and real-life problem.</p><p><br></p><p>You have demonstrated that t=
here are specifications that, if chained in the right way, can lead to ambi=
guities or gaps in interchange.</p><p><br></p><p>What we don't have is an a=
ctual use case with real-life consequences for a set of existing users, not=
hypothetical ones.</p><p><br></p><p>When it comes to encoding decisions ba=
sed on existing documents, there is a strong presumption that once sufficie=
ntly many documents exist that contain a character, that this character wil=
l be needed in digitizing these documents, whether immediately, or eventual=
ly (e.g. in the case of future scholarly studies). Also, the texts themselv=
es exist, barring accidents, in permanence. Therefore, it is justified to c=
onsider irrevocably allocating a character that will map to this source in =
perpetuity, even though each encoded character carries a small cost for imp=
lementers.<br><br>However, when it comes to legacy characters, there's an a=
dditional cost that is imposed, and that is based on the fact that characte=
rs that are encoded solely for compatibility will usually violate one or mo=
re of the other encoding principles, something that incrementally complicat=
es the standard. Even for people who never intend to use that character.<br=
><br>Therefore, the SEW is on solid ground when it demands not only a hypot=
hetical scenario, but evidence of actual impact on actual users. Not only w=
hether some application could invoke an API, but whether such applications =
exist and are used today to access documents encoded using the legacy chara=
cters in a way that is compromised irreparably by not having an encoding fo=
r them.</p><p><br></p><p>A./</p><p><br></p><p><br></p><p>On 5/4/2026 10:24 =
AM, <a href=3D"mailto:[email protected]" rel=3D"noreferrer" target=3D"_b=
lank">[email protected]</a> via Unicode wrote:<br></p><blockquote type=
=3D"cite"><p>In UTC 187 Minutes, "<span style=3D"color: rgb(0, 0, 0); font-=
family: "DMCA Sans Serif 10.0 dev1"; font-size: medium; font-styl=
e: normal; font-variant-ligatures: normal; font-variant-caps: normal; font-=
weight: 400; letter-spacing: normal; text-align: start; text-indent: 0px; t=
ext-transform: none; white-space: normal; word-spacing: 0px; text-decoratio=
n-style: initial; text-decoration-color: initial; float: none; display: inl=
ine;">Asmus Freytag noted that the fact that lists of things existed in the=
past does not make these things plain text. Ned Holbrook pointed out that =
the purported issue occurs in a closed system, not in public interchange.</=
span>". However, the arguments in the proposal do not merely hinge on the e=
ncodings being lists of characters, but specifically points out methods to =
interchange text, including an example of copying terminal output and pasti=
ng to Notepad, where the copying invokes the mapping of the current termina=
l codepage to UCS-2 (as is CHAR_INFO compatible) and the pasting writes it =
into plain text. Win32 is also not a closed system, as Win32 can capture th=
e tiles of the output of Windows 3.1 Arabic DOS/Win16 programs and Windows =
95/98/ME Arabic DOS/Win16/Win32 programs, but Win32 can also interact with =
public text interchange systems by reading and writing to files and network=
. I'm not saying that Unicode absolutely must include those characters, but=
those kinds of misleading claims are causing users to misunderstand what t=
he proposal is about, and I don't want Unicode to be relying on uninformed =
decisions to evaluate proposals.</p><p><br></p><div><blockquote style=3D"pa=
dding-top: 12px;"><p style=3D"padding-bottom: 12px;"><strong>Dnia 18 kwietn=
ia 2026 13:36</strong> <a href=3D"mailto:[email protected]" rel=3D"n=
oopener noreferrer nofollow noreferrer" target=3D"_blank"><span style=3D"ma=
rgin-left: 4px;">[email protected] via Unicode</span></a><span style=3D"=
margin-left: 4px;"> </span><a href=3D"mailto:[email protected]" rel=
=3D"noreferrer" target=3D"_blank"><span style=3D"margin-left: 4px;">< un=
[email protected] ></span></a> napisa=C5=82(a):</p><div id=3D"gwpb6=
7625a5_m_5371084887079304582gwpbf2d884a"><div id=3D"gwpb67625a5_m_537108488=
7079304582gwpbf2d884ah"><div><p>The SEW subsequently explained that the act=
ual reason is due to insufficient evidence of user community that would nee=
d to use the resulting mapping. Despite Win32 being a highly popular platfo=
rm with plenty of backwards compatibility and native UCS-2 terminal support=
, the specific use cases of installing codepages into Windows NT and using =
terminal tiles from Windows 3.1/95/98/ME are not sufficiently documented, m=
aking it difficult for any user communities to form around it. So it seems =
like the idea of standardizing legacy Arabic terminal BMP mappings is a dea=
d end for now.</p><p><br></p><div><blockquote style=3D"padding-top: 12px;">=
<p style=3D"padding-bottom: 12px;"><strong>Dnia 17 kwietnia 2026 22:59</str=
ong> <a href=3D"mailto:[email protected]" rel=3D"noopener noreferrer=
nofollow noreferrer" target=3D"_blank"><span style=3D"margin-left: 4px;">p=
[email protected] via Unicode</span></a><span style=3D"margin-left: 4px;"=
> </span><a href=3D"mailto:[email protected]" rel=3D"noreferrer" tar=
get=3D"_blank"><span style=3D"margin-left: 4px;">< [email protected].=
org ></span></a> napisa=C5=82(a):</p><div id=3D"gwpb67625a5_m_5371084887=
079304582gwpbf2d884a_gwp2281a7f8"><div id=3D"gwpb67625a5_m_5371084887079304=
582gwpbf2d884a_gwp2281a7f8h"><div><p>The Recommendations in L2/26-100 claim=
that Microsoft's documentation of legacy Arabic encodings is available at =
<a href=3D"https://learn.microsoft.com/en-us/typography/legacy/legacy_arabi=
c_fonts" =3D"" rel=3D"noreferrer" target=3D"_blank">https://learn.microsoft=
.com/en-us/typography/legacy/legacy_arabic_fonts</a>. However, that article=
only demonstrates two encodings of TrueType fonts, which are used in Windo=
ws 3.1 but are completely different from the eight terminal encodings. Unli=
ke the TrueType encodings which represent internal shaping mappings and are=
not used for text interchange, the terminal encodings have been demonstrat=
ed to be directly used in text interchange through int 10h and ReadConsoleO=
utputA/WriteConsoleOutputA as already demonstrated in L2/26-077. The Recomm=
endations also claim that the proposal does not demonstrate any need for in=
terchange or encoding, but the proposal actually demonstrated such a need d=
ue to the logical extension of the Win32 terminal API to the functions Read=
ConsoleOutputW/WriteConsoleOutputW, which are in Windows NT and may be used=
on the output of previously ran programs (including those that used the le=
gacy Arabic terminal encodings), which given the CHAR_INFO structure, there=
fore implies a need for all the tiles to map to BMP for interchange. I'm no=
t objecting to the SEW's conclusion of "Users are expected to use PUA.", wh=
ich can indeed be used to provide a mapping even if not standardized, but t=
he reasoning given was flawed.</p><p><br></p><div><blockquote style=3D"padd=
ing-top: 12px;"><p style=3D"padding-bottom: 12px;"><strong>Dnia 09 stycznia=
2026 17:25</strong> <a href=3D"mailto:[email protected]" rel=3D"norefer=
rer" target=3D"_blank"><span style=3D"margin-left: 4px;">piotrunio-2004@wp.=
pl</span></a><span style=3D"margin-left: 4px;"> </span><a href=3D"mailto:pi=
[email protected]" rel=3D"noreferrer" target=3D"_blank"><span style=3D"mar=
gin-left: 4px;">< [email protected] ></span></a> napisa=C5=82(a):<=
/p><div id=3D"gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f8_gwpa=
05276c7"><div id=3D"gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f=
8_gwpa05276c7h"><div><div id=3D"gwpb67625a5_m_5371084887079304582gwpbf2d884=
a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718"><div id=3D"gwpb67625a5_m_53710848870=
79304582gwpbf2d884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718h"><div><div id=3D"g=
wpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a7f8_gwpa05276c7_gwpa8b5=
f718_gwpa8b5f718"><div id=3D"gwpb67625a5_m_5371084887079304582gwpbf2d884a_g=
wp2281a7f8_gwpa05276c7_gwpa8b5f718_gwpa8b5f718h"><div><div id=3D"gwpb67625a=
5_m_5371084887079304582gwpbf2d884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718_gwpa=
8b5f718_gwpa8b5f718"><div id=3D"gwpb67625a5_m_5371084887079304582gwpbf2d884=
a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718_gwpa8b5f718_gwpa8b5f718h"><div><p>The=
following Win32 C code will output 256 characters in system console codepa=
ge into the character grid, capture those character tiles in UCS-2 if possi=
ble, and then output the current console codepage number.<br></p><p><br></p=
><p>#include <windows.h><br>#include <stdio.h><br>int main(){<b=
r>HANDLE hConsole=3DGetStdHandle(STD_OUTPUT_HANDLE);<br>CHAR_INFO screen[25=
6];<br>COORD size=3D{16,16,};<br>COORD pos=3D{0,0,};<br>SMALL_RECT rect=3D{=
0,0,15,15,};<br>for(int i=3D0;i<256;i++){<br>screen[i].Attributes=3D0xF0=
;<br>screen[i].Char.AsciiChar=3Di;<br>}<br>WriteConsoleOutputA(hConsole,scr=
een,size,pos,&rect);<br>CHAR_INFO screenu[256];<br>if(ReadConsoleOutput=
W(hConsole,screenu,size,pos,&rect)){<br>for(int i=3D0;i<256;i++) pri=
ntf("%04X ",screenu[i].Char.UnicodeChar);<br>}<br>else{<br>printf("error %0=
8X\n",GetLastError());<br>}<br>printf("codepage %u",GetConsoleOutputCP());<=
br>}<br><br></p><p>In most cases, whenever a legacy Win32 codepage is used,=
the application can run on Windows NT to capture the UCS-2 mapping of thos=
e character cells to the BMP (although for CJK codepages a more complex set=
up would be necessary due to thousands of fullwidth characters with 2-byte =
sequences).<br></p><p><br></p><p>However, in Arabic versions of Windows 9x =
(95/98/ME) the resulting character set has many presentation forms that are=
not in Unicode. This is the result when running on Windows ME: <a hre=
f=3D"https://i.imgur.com/QFm3SkI.png" =3D"" rel=3D"noopener noreferrer nore=
ferrer" target=3D"_blank">https://i.imgur.com/QFm3SkI.png</a> in 10=C3=
=9720 font, <a href=3D"https://i.imgur.com/KUbLQ0A.png" =3D"" rel=3D"noopen=
er noreferrer noreferrer" target=3D"_blank">https://i.imgur.com/KUbLQ0A.png=
</a> in 10=C3=9718 font (same result also appears in Windows 95/98). 5=
=C3=9712, 7=C3=9712, 8=C3=9712, 10=C3=9718, 10=C3=9720, and 12=C3=9716 bitm=
ap fonts have been attested with that character set (VGAOEM.FON, 8514OEM.FO=
N, DOSAPP.FON). The 10=C3=9720 font has slightly different mapping than the=
other sizes: 0x93 is =C3=B6 instead of =C3=B4, and 0x97 is missing (causin=
g the following characters on the same line to be drawn at the wrong positi=
on). It also claims to be using codepage 720, but many characters differ fr=
om their CP720 mappings, including the bundled CP_720.NLS mappings (fo=
r example, =D9=80 (U+0640 ARABIC TATWEEL) is 0x95 in CP720, but in the cons=
ole 0x95 is =D8=B4 instead, and the tatweel is at 0xFF). On Windows 9x,&nbs=
p;ReadConsoleOutputW is not supported so the UCS-2 mappings of the console =
character tiles cannot be captured (error 0x00000078 ERROR_CALL_NOT_IMPLEME=
NTED).<br></p></div></div></div><p><br></p><p>When that program runs on Ara=
bic versions of Windows NT, the visual output is of the CP437 character set=
if one of the bundled bitmap fonts is used (<a href=3D"https://i.imgur.com=
/RxjtxMH.png" =3D"" rel=3D"noopener noreferrer noreferrer" target=3D"_blank=
">https://i.imgur.com/RxjtxMH.png</a>), or the CP720 set if Lucida Console =
is used, with the Arabic letters either having glitchy font substitution (N=
T 4.0, NT 5.0/2000) or the .notdef glyph (NT 5.1/XP and up). In fact, it se=
ems that the only Arabic bitmap fonts that occur in Windows NT are CP1256 f=
onts, which are not used in terminals. So this appears to be one of those p=
ermanent Windows compatibility regressions that occured when Windows 9x end=
ed, where the terminals can no longer render legacy Arabic text. Even if th=
e user managed to use registry hacks to set the font to Courier New or Simp=
lified Arabic Fixed, it would still use the CP720 mapping which is not comp=
atible with the Windows 9x set.<br></p></div></div><p><br></p></div></div><=
/div></div><div id=3D"gwpb67625a5_m_5371084887079304582gwpbf2d884a_gwp2281a=
7f8_gwpa05276c7_gwpa8b5f718"><div id=3D"gwpb67625a5_m_5371084887079304582gw=
pbf2d884a_gwp2281a7f8_gwpa05276c7_gwpa8b5f718h"><div><p>It appears that in =
the Windows 9x Arabic terminal character set, 244 characters (=E2=80=87=EF=
=BA=80=EF=BA=81=EF=BA=82=EF=BA=83=EF=BA=84=EF=BA=85=EF=BA=87=EF=BA=88=EF=BA=
=8A=EF=BA=8B=EF=BA=8D=EF=BA=8E=EF=BA=8F=EF=BA=91=EF=BA=93=E2=96=BA=E2=97=84=
=E2=86=95=EF=BA=95=C2=B6=C2=A7=EF=BA=97=EF=BA=99=E2=86=91=E2=86=93=E2=86=92=
=E2=86=90=EF=BA=9B=EF=B9=B0=E2=96=B2=E2=96=BC !"#$%&'()*+,-./0123456789=
:;<=3D>?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\]^_`abcdefghijklmnopqrstuvwxyz{|=
}~=EF=BA=9D=EF=BA=9F=EF=BA=A1=C3=A9=C3=A2=EF=BA=A3=C3=A0=EF=BA=A5=C3=A7=C3=
=AA=C3=AB=C3=A8=C3=AF=C3=AE=EF=BA=A7=EF=BA=A9=EF=BA=AB=EF=BA=AD=EF=BA=AF=C3=
=B4=EF=BA=B3=C3=BB=C3=B9=EF=BA=B7=EF=BA=BB=C2=A3=EF=BA=BF=EF=BB=81=EF=BB=85=
=EF=BB=89=EF=BB=8A=EF=BB=8B=EF=BB=8C=EF=BB=8D=EF=BB=8E=EF=BB=8F=EF=BB=90=EF=
=BB=91=EF=BB=93=EF=BB=95=EF=BB=97=EF=BB=99=EF=BB=9B=C2=AB=C2=BB=EF=B9=B1=E2=
=96=92=EF=B9=B2=E2=94=82=E2=94=A4=EF=B9=B4=EF=B9=B6=EF=B9=B7=EF=B9=B8=D9=A0=
=D9=A1=D9=A2=D9=A3=EF=B9=B9=EF=B9=BA=E2=94=90=E2=94=94=E2=94=B4=E2=94=AC=E2=
=94=9C=E2=94=80=E2=94=BC=EF=B9=BB=EF=B9=BE=D9=A4=D9=A5=D9=A6=D9=A7=D9=A8=D9=
=A9=D8=8C=EF=B9=BF=EF=B1=9E=EF=B1=9F=EF=B1=A0=EF=B3=B2=EF=B1=A1=EF=B3=B3=EF=
=B1=A2=E2=94=98=E2=94=8C=D8=9B=D8=9F=C2=A4=EF=BB=9D=EF=BB=9F=EF=BB=A1=EF=BB=
=A3=EF=BB=A5=EF=BB=A7=C2=B5=EF=BB=A9=EF=BB=AB=EF=BB=AC=EF=BB=AD=EF=BB=AF=EF=
=BB=B0=EF=BB=B1=EF=BB=B2=EF=BB=B3=EF=B3=B4=EF=B9=BC=EF=B9=BD=EF=BA=B1=EF=BA=
=B5=EF=BA=B9=EF=BA=BD=EF=B9=B3=C2=B0=C2=B7=E2=96=A0=D9=80) are already in U=
nicode, but 12 characters are not in Unicode:<br></p><p>=E2=80=A2 6 of them=
are pieces of lam-alef ligatures (0xDD, 0xDE, 0xF9, 0xFB, 0xFC, 0xFD)<br><=
/p><p>=E2=80=A2 2 of them are shadda with fathatan ligatures without or wit=
h tatweel (0xD0, 0xD1)<br></p><p>=E2=80=94 in some legacy Microsoft fonts, =
shadda with fathatan is mapped to private use U+E818<br></p><p>=E2=80=A2 4 =
of them are disunifications of seen/sheen/sad/dad occuring either with or w=
ithout tail<br></p><p>=E2=80=94 =EF=B9=B3 (U+FE73 ARABIC TAIL FRAGMENT=
) was originally encoded in Unicode 3.2 for CP864 compatibility; in that co=
depage, the forms of seen/sheen/sad/dad attach to the tail fragment<br=
></p><p>=E2=80=94 forms with included tail: 0x92, 0x95, 0x98, 0x8A<br>=
</p><p>=E2=80=94 forms without tail (attaching to tail fragment like in CP8=
64): 0xF3, 0xF4, 0xF5, 0xF6<br></p></div><p><br></p></div><p>If someon=
e tried to make a Win32 console implementation and tried to implement both =
Windows 9x Arabic terminal character set compatibility and wide string API =
(ReadConsoleOutputW) compatibility simultaneously, then they would run into=
the issue that there is currently no standardized mapping to handle that s=
cenario. What should Windows 9x Arabic console compatible implementations d=
o in that case?<br></p></div><p><br></p></div></div></div></blockquote></di=
v><p><br></p></div></div></div></blockquote></div><p><br></p></div></div></=
div></blockquote></div><p><br></p></blockquote><p><br></p></div></blockquot=
e></div></div></div></div></blockquote></div><p><br></p>
--2JBJUNXFBCKBRPRSPVHBWnhgwp--