| Newsgroups |
gmane.text.unicode.general |
| Message-ID |
<[email protected]> |
--2FPULRMRRFBYQSTATKAYKnhgwp
Content-Transfer-Encoding: quoted-printable
Content-Type: text/plain; charset=UTF-8
The Recommendations in L2/26-100 claim that Microsoft's documentation o=
f legacy Arabic encodings is available at https://learn.microsoft.com/en-u=
s/typography/legacy/legacy_arabic_fonts . However, that article only demons=
trates two encodings of TrueType fonts, which are used in Windows 3.1 but a=
re completely different from the eight terminal encodings. Unlike the TrueT=
ype encodings which represent internal shaping mappings and are not used fo=
r text interchange, the terminal encodings have been demonstrated to be dir=
ectly used in text interchange through int 10h and ReadConsoleOutputA/Write=
ConsoleOutputA as already demonstrated in L2/26-077. The Recommendations al=
so claim that the proposal does not demonstrate any need for interchange or=
encoding, but the proposal actually demonstrated such a need due to the lo=
gical extension of the Win32 terminal API to the functions ReadConsoleOutpu=
tW/WriteConsoleOutputW, which are in Windows NT and may be used on the outp=
ut of previously ran programs (including those that used the legacy Arabic =
terminal encodings), which given the CHAR_INFO structure, therefore implies=
a need for all the tiles to map to BMP for interchange. I'm not object=
ing to the SEW's conclusion of "Users are expected to use PUA."=
, which can indeed be used to provide a mapping even if not standardized, b=
ut the reasoning given was flawed. Dnia 09 stycznia 2026 17:25 piotrunio=
[email protected] < [email protected] > napisa=C5=82(a): The following =
Win32 C code will output 256 characters in system console codepage into the=
character grid, capture those character tiles in UCS-2 if possible, and th=
en output the current console codepage number. #include <windows.h>=
#include <stdio.h> int main(){ HANDLE hConsole=3DGetStdHandle(STD=
_OUTPUT_HANDLE); CHAR_INFO screen[256]; COORD size=3D{16,16,}; COORD pos=
=3D{0,0,}; SMALL_RECT rect=3D{0,0,15,15,}; for(int i=3D0;i<256;i++){ =
screen[i].Attributes=3D0xF0; screen[i].Char.AsciiChar=3Di; } WriteConsol=
eOutputA(hConsole,screen,size,pos,&rect); CHAR_INFO screenu[256]; if(=
ReadConsoleOutputW(hConsole,screenu,size,pos,&rect)){ for(int i=3D0;i&=
lt;256;i++) printf("%04X ",screenu[i].Char.UnicodeChar); } else{ =
printf("error %08X\n",GetLastError()); } printf("codepage %u=
",GetConsoleOutputCP()); } In most cases, whenever a legacy Win32 co=
depage is used, the application can run on Windows NT to capture the UCS-2 =
mapping of those character cells to the BMP (although for CJK codepages a m=
ore complex setup would be necessary due to thousands of fullwidth characte=
rs with 2-byte sequences). However, in Arabic versions of Windows 9x (95/=
98/ME) the resulting character set has many presentation forms that are not=
in Unicode. This is the result when running on Windows ME:=C2=A0 i.imgur.c=
om https://i.imgur.com/QFm3SkI.png =C2=A0in 10=C3=9720 font, i.imgur.com h=
ttps://i.imgur.com/KUbLQ0A.png =C2=A0in 10=C3=9718 font (same result also a=
ppears in Windows 95/98). 5=C3=9712, 7=C3=9712, 8=C3=9712, 10=C3=9718, 10=
=C3=9720, and 12=C3=9716 bitmap fonts have been attested with that characte=
r set (VGAOEM.FON, 8514OEM.FON, DOSAPP.FON). The 10=C3=9720 font has slight=
ly different mapping than the other sizes: 0x93 is =C3=B6 instead of =C3=B4=
, and 0x97 is missing (causing the following characters on the same line to=
be drawn at the wrong position). It also claims to be using codepage 720, =
but many characters differ from their CP720 mappings, including the bundled=
=C2=A0CP_720.NLS mappings (for example, =D9=80 (U+0640 ARABIC TATWEEL) is 0=
x95 in CP720, but in the console 0x95 is =D8=B4 instead, and the tatweel is=
at 0xFF). On Windows 9x,=C2=A0ReadConsoleOutputW is not supported so the U=
CS-2 mappings of the console character tiles cannot be captured (error 0x00=
000078 ERROR_CALL_NOT_IMPLEMENTED). When that program runs on Arabic vers=
ions of Windows NT, the visual output is of the CP437 character set if one =
of the bundled bitmap fonts is used ( i.imgur.com https://i.imgur.com/Rxjtx=
MH.png ), or the CP720 set if Lucida Console is used, with the Arabic lette=
rs either having glitchy font substitution (NT 4.0, NT 5.0/2000) or the .no=
tdef glyph (NT 5.1/XP and up). In fact, it seems that the only Arabic bitma=
p fonts that occur in Windows NT are CP1256 fonts, which are not used in te=
rminals. So this appears to be one of those permanent Windows compatibility=
regressions that occured when Windows 9x ended, where the terminals can no=
longer render legacy Arabic text. Even if the user managed to use registry=
hacks to set the font to Courier New or Simplified Arabic Fixed, it would =
still use the CP720 mapping which is not compatible with the Windows 9x set=
. It appears that in the Windows 9x Arabic terminal character set, 244 ch=
aracters (=E2=80=87=EF=BA=80=EF=BA=81=EF=BA=82=EF=BA=83=EF=BA=84=EF=BA=85=
=EF=BA=87=EF=BA=88=EF=BA=8A=EF=BA=8B=EF=BA=8D=EF=BA=8E=EF=BA=8F=EF=BA=91=EF=
=BA=93=E2=96=BA=E2=97=84=E2=86=95=EF=BA=95=C2=B6=C2=A7=EF=BA=97=EF=BA=99=E2=
=86=91=E2=86=93=E2=86=92=E2=86=90=EF=BA=9B=EF=B9=B0=E2=96=B2=E2=96=BC !"=
;#$%&'()*+,-./0123456789:;<=3D>?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\=
]^_`abcdefghijklmnopqrstuvwxyz{|}~=EF=BA=9D=EF=BA=9F=EF=BA=A1=C3=A9=C3=A2=
=EF=BA=A3=C3=A0=EF=BA=A5=C3=A7=C3=AA=C3=AB=C3=A8=C3=AF=C3=AE=EF=BA=A7=EF=BA=
=A9=EF=BA=AB=EF=BA=AD=EF=BA=AF=C3=B4=EF=BA=B3=C3=BB=C3=B9=EF=BA=B7=EF=BA=BB=
=C2=A3=EF=BA=BF=EF=BB=81=EF=BB=85=EF=BB=89=EF=BB=8A=EF=BB=8B=EF=BB=8C are a=
lready in Unicode, but 12 characters are not in Unicode: =E2=80=A2 6 of th=
em are pieces of lam-alef ligatures (0xDD, 0xDE, 0xF9, 0xFB, 0xFC, 0xFD) =
=E2=80=A2 2 of them are shadda with fathatan ligatures without or with tatw=
eel (0xD0, 0xD1) =E2=80=94 in some legacy Microsoft fonts, shadda with fat=
hatan is mapped to private use U+E818 =E2=80=A2 4 of them are disunificati=
ons of seen/sheen/sad/dad occuring either with or without tail =E2=80=94=
=C2=A0=EF=B9=B3 (U+FE73 ARABIC TAIL FRAGMENT) was originally encoded in Uni=
code 3.2 for CP864 compatibility; in that codepage, the forms of=C2=A0seen/=
sheen/sad/dad attach to the tail fragment =E2=80=94 forms with included ta=
il:=C2=A00x92, 0x95, 0x98, 0x8A =E2=80=94 forms without tail (attaching to=
tail fragment like in CP864):=C2=A00xF3, 0xF4, 0xF5, 0xF6 If someone tri=
ed to make a Win32 console implementation and tried to implement both Windo=
ws 9x Arabic terminal character set compatibility and wide string API (Read=
ConsoleOutputW) compatibility simultaneously, then they would run into the =
issue that there is currently no standardized mapping to handle that scenar=
io. What should Windows 9x Arabic console compatible implementations do in =
that case?=0D
--2FPULRMRRFBYQSTATKAYKnhgwp
Content-Transfer-Encoding: quoted-printable
Content-Type: text/html; charset=UTF-8
<p>The Recommendations in L2/26-100 claim that Microsoft's documentation of=
legacy Arabic encodings is available at <a href=3D"https://learn.microsoft=
.com/en=02us/typography/legacy/legacy_arabic_fonts" rel=3D"noopener norefer=
rer nofollow" target=3D"_blank">https://learn.microsoft.com/en-us/typograph=
y/legacy/legacy_arabic_fonts</a>. However, that article only demonstrates t=
wo encodings of TrueType fonts, which are used in Windows 3.1 but are compl=
etely different from the eight terminal encodings. Unlike the TrueType enco=
dings which represent internal shaping mappings and are not used for text i=
nterchange, the terminal encodings have been demonstrated to be directly us=
ed in text interchange through int 10h and ReadConsoleOutputA/WriteConsoleO=
utputA as already demonstrated in L2/26-077. The Recommendations also claim=
that the proposal does not demonstrate any need for interchange or encodin=
g, but the proposal actually demonstrated such a need due to the logical ex=
tension of the Win32 terminal API to the functions ReadConsoleOutputW/Write=
ConsoleOutputW, which are in Windows NT and may be used on the output of pr=
eviously ran programs (including those that used the legacy Arabic terminal=
encodings), which given the CHAR_INFO structure, therefore implies a need =
for all the tiles to map to BMP for interchange. I'm not objecting to the S=
EW's conclusion of "Users are expected to use PUA.", which can indeed be us=
ed to provide a mapping even if not standardized, but the reasoning given w=
as flawed.</p><p><br></p><div class=3D"nh_extra"><blockquote style=3D"paddi=
ng-top: 12px;" class=3D"nh_qoute"><p style=3D"padding-bottom: 12px;"><stron=
g>Dnia 09 stycznia 2026 17:25</strong> <span style=3D"margin-left: 4px;">pi=
[email protected] < [email protected] ></span> napisa=C5=82(a):</=
p><div id=3D"gwpa05276c7"><div id=3D"gwpa05276c7h"><div class=3D"gwpa05276c=
7b" data-message-body=3D"true" data-color-mode=3D"light"><div id=3D"gwpa052=
76c7_gwpa8b5f718"><div id=3D"gwpa05276c7_gwpa8b5f718h"><div data-color-mode=
=3D"light" class=3D"gwpa05276c7_gwpa8b5f718b" data-message-body=3D"true"><d=
iv id=3D"gwpa05276c7_gwpa8b5f718_gwpa8b5f718"><div id=3D"gwpa05276c7_gwpa8b=
5f718_gwpa8b5f718h"><div data-color-mode=3D"light" class=3D"gwpa05276c7_gwp=
a8b5f718_gwpa8b5f718b" data-message-body=3D"true"><div id=3D"gwpa05276c7_gw=
pa8b5f718_gwpa8b5f718_gwpa8b5f718"><div id=3D"gwpa05276c7_gwpa8b5f718_gwpa8=
b5f718_gwpa8b5f718h"><div data-color-mode=3D"light" class=3D"gwpa05276c7_gw=
pa8b5f718_gwpa8b5f718_gwpa8b5f718b" data-message-body=3D"true"><p>The follo=
wing Win32 C code will output 256 characters in system console codepage int=
o the character grid, capture those character tiles in UCS-2 if possible, a=
nd then output the current console codepage number.<br></p><p><br></p><p>#i=
nclude <windows.h><br>#include <stdio.h><br>int main(){<br>HAND=
LE hConsole=3DGetStdHandle(STD_OUTPUT_HANDLE);<br>CHAR_INFO screen[256];<br=
>COORD size=3D{16,16,};<br>COORD pos=3D{0,0,};<br>SMALL_RECT rect=3D{0,0,15=
,15,};<br>for(int i=3D0;i<256;i++){<br>screen[i].Attributes=3D0xF0;<br>s=
creen[i].Char.AsciiChar=3Di;<br>}<br>WriteConsoleOutputA(hConsole,screen,si=
ze,pos,&rect);<br>CHAR_INFO screenu[256];<br>if(ReadConsoleOutputW(hCon=
sole,screenu,size,pos,&rect)){<br>for(int i=3D0;i<256;i++) printf("%=
04X ",screenu[i].Char.UnicodeChar);<br>}<br>else{<br>printf("error %08X\n",=
GetLastError());<br>}<br>printf("codepage %u",GetConsoleOutputCP());<br>}<b=
r><br></p><p>In most cases, whenever a legacy Win32 codepage is used, the a=
pplication can run on Windows NT to capture the UCS-2 mapping of those char=
acter cells to the BMP (although for CJK codepages a more complex setup wou=
ld be necessary due to thousands of fullwidth characters with 2-byte sequen=
ces).<br></p><p><br></p><p>However, in Arabic versions of Windows 9x (95/98=
/ME) the resulting character set has many presentation forms that are not i=
n Unicode. This is the result when running on Windows ME: <a href=3D"h=
ttps://i.imgur.com/QFm3SkI.png" =3D"" rel=3D"noopener noreferrer" target=3D=
"_blank">https://i.imgur.com/QFm3SkI.png</a> in 10=C3=9720 font, <a hr=
ef=3D"https://i.imgur.com/KUbLQ0A.png" =3D"" rel=3D"noopener noreferrer" ta=
rget=3D"_blank">https://i.imgur.com/KUbLQ0A.png</a> in 10=C3=9718 font=
(same result also appears in Windows 95/98). 5=C3=9712, 7=C3=9712, 8=C3=97=
12, 10=C3=9718, 10=C3=9720, and 12=C3=9716 bitmap fonts have been attested =
with that character set (VGAOEM.FON, 8514OEM.FON, DOSAPP.FON). The 10=C3=97=
20 font has slightly different mapping than the other sizes: 0x93 is =C3=B6=
instead of =C3=B4, and 0x97 is missing (causing the following characters o=
n the same line to be drawn at the wrong position). It also claims to be us=
ing codepage 720, but many characters differ from their CP720 mappings, inc=
luding the bundled CP_720.NLS mappings (for example, =D9=80 (U+0640 AR=
ABIC TATWEEL) is 0x95 in CP720, but in the console 0x95 is =D8=B4 instead, =
and the tatweel is at 0xFF). On Windows 9x, ReadConsoleOutputW is not =
supported so the UCS-2 mappings of the console character tiles cannot be ca=
ptured (error 0x00000078 ERROR_CALL_NOT_IMPLEMENTED).<br></p></div></div></=
div><p><br></p><p>When that program runs on Arabic versions of Windows NT, =
the visual output is of the CP437 character set if one of the bundled bitma=
p fonts is used (<a href=3D"https://i.imgur.com/RxjtxMH.png" =3D"" rel=3D"n=
oopener noreferrer" target=3D"_blank">https://i.imgur.com/RxjtxMH.png</a>),=
or the CP720 set if Lucida Console is used, with the Arabic letters either=
having glitchy font substitution (NT 4.0, NT 5.0/2000) or the .notdef glyp=
h (NT 5.1/XP and up). In fact, it seems that the only Arabic bitmap fonts t=
hat occur in Windows NT are CP1256 fonts, which are not used in terminals. =
So this appears to be one of those permanent Windows compatibility regressi=
ons that occured when Windows 9x ended, where the terminals can no longer r=
ender legacy Arabic text. Even if the user managed to use registry hacks to=
set the font to Courier New or Simplified Arabic Fixed, it would still use=
the CP720 mapping which is not compatible with the Windows 9x set.<br></p>=
</div></div><p><br></p></div></div></div></div><div id=3D"gwpa05276c7_gwpa8=
b5f718"><div id=3D"gwpa05276c7_gwpa8b5f718h"><div data-color-mode=3D"light"=
class=3D"gwpa05276c7_gwpa8b5f718b" data-message-body=3D"true"><p>It appear=
s that in the Windows 9x Arabic terminal character set, 244 characters (=E2=
=80=87=EF=BA=80=EF=BA=81=EF=BA=82=EF=BA=83=EF=BA=84=EF=BA=85=EF=BA=87=EF=BA=
=88=EF=BA=8A=EF=BA=8B=EF=BA=8D=EF=BA=8E=EF=BA=8F=EF=BA=91=EF=BA=93=E2=96=BA=
=E2=97=84=E2=86=95=EF=BA=95=C2=B6=C2=A7=EF=BA=97=EF=BA=99=E2=86=91=E2=86=93=
=E2=86=92=E2=86=90=EF=BA=9B=EF=B9=B0=E2=96=B2=E2=96=BC !"#$%&'()*+,-./0=
123456789:;<=3D>?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\]^_`abcdefghijklmnopqrs=
tuvwxyz{|}~=EF=BA=9D=EF=BA=9F=EF=BA=A1=C3=A9=C3=A2=EF=BA=A3=C3=A0=EF=BA=A5=
=C3=A7=C3=AA=C3=AB=C3=A8=C3=AF=C3=AE=EF=BA=A7=EF=BA=A9=EF=BA=AB=EF=BA=AD=EF=
=BA=AF=C3=B4=EF=BA=B3=C3=BB=C3=B9=EF=BA=B7=EF=BA=BB=C2=A3=EF=BA=BF=EF=BB=81=
=EF=BB=85=EF=BB=89=EF=BB=8A=EF=BB=8B=EF=BB=8C=EF=BB=8D=EF=BB=8E=EF=BB=8F=EF=
=BB=90=EF=BB=91=EF=BB=93=EF=BB=95=EF=BB=97=EF=BB=99=EF=BB=9B=C2=AB=C2=BB=EF=
=B9=B1=E2=96=92=EF=B9=B2=E2=94=82=E2=94=A4=EF=B9=B4=EF=B9=B6=EF=B9=B7=EF=B9=
=B8=D9=A0=D9=A1=D9=A2=D9=A3=EF=B9=B9=EF=B9=BA=E2=94=90=E2=94=94=E2=94=B4=E2=
=94=AC=E2=94=9C=E2=94=80=E2=94=BC=EF=B9=BB=EF=B9=BE=D9=A4=D9=A5=D9=A6=D9=A7=
=D9=A8=D9=A9=D8=8C=EF=B9=BF=EF=B1=9E=EF=B1=9F=EF=B1=A0=EF=B3=B2=EF=B1=A1=EF=
=B3=B3=EF=B1=A2=E2=94=98=E2=94=8C=D8=9B=D8=9F=C2=A4=EF=BB=9D=EF=BB=9F=EF=BB=
=A1=EF=BB=A3=EF=BB=A5=EF=BB=A7=C2=B5=EF=BB=A9=EF=BB=AB=EF=BB=AC=EF=BB=AD=EF=
=BB=AF=EF=BB=B0=EF=BB=B1=EF=BB=B2=EF=BB=B3=EF=B3=B4=EF=B9=BC=EF=B9=BD=EF=BA=
=B1=EF=BA=B5=EF=BA=B9=EF=BA=BD=EF=B9=B3=C2=B0=C2=B7=E2=96=A0=D9=80) are alr=
eady in Unicode, but 12 characters are not in Unicode:<br></p><p>=E2=80=A2 =
6 of them are pieces of lam-alef ligatures (0xDD, 0xDE, 0xF9, 0xFB, 0xFC, 0=
xFD)<br></p><p>=E2=80=A2 2 of them are shadda with fathatan ligatures witho=
ut or with tatweel (0xD0, 0xD1)<br></p><p>=E2=80=94 in some legacy Microsof=
t fonts, shadda with fathatan is mapped to private use U+E818<br></p><p>=E2=
=80=A2 4 of them are disunifications of seen/sheen/sad/dad occuring either =
with or without tail<br></p><p>=E2=80=94 =EF=B9=B3 (U+FE73 ARABIC TAIL=
FRAGMENT) was originally encoded in Unicode 3.2 for CP864 compatibility; i=
n that codepage, the forms of seen/sheen/sad/dad attach to the tail fr=
agment<br></p><p>=E2=80=94 forms with included tail: 0x92, 0x95, 0x98,=
0x8A<br></p><p>=E2=80=94 forms without tail (attaching to tail fragment li=
ke in CP864): 0xF3, 0xF4, 0xF5, 0xF6<br></p></div><p><br></p></div><p>=
If someone tried to make a Win32 console implementation and tried to implem=
ent both Windows 9x Arabic terminal character set compatibility and wide st=
ring API (ReadConsoleOutputW) compatibility simultaneously, then they would=
run into the issue that there is currently no standardized mapping to hand=
le that scenario. What should Windows 9x Arabic console compatible implemen=
tations do in that case?<br></p></div><p><br></p></div></div></div></blockq=
uote></div><p></p>
--2FPULRMRRFBYQSTATKAYKnhgwp--