Re: Update List of Character Entity Names
Thomas Dickey <[email protected]>
| Newsgroups | gmane.comp.web.lynx.devel |
|---|---|
| Message-ID | <[email protected]> |
On Fri, Jan 10, 2025 at 09:56:12PM -0700, Brian Inglis wrote: > On 2025-01-09 14:34, Thomas Dickey wrote: > > On Thu, Jan 09, 2025 at 11:15:23AM -0700, Brian Inglis wrote: > > > Hi folks, > > > > > > Many sites are now using Character Entity Names defined under > > > > > > https://www.w3.org/TR/xml-entity-names/ > > > > https://www.w3.org/TR/xml-entity-names/#source > > https://www.w3.org/TR/xml-entity-names/bycodes.html > > https://www.w3.org/TR/xml-entity-names/byalpha.html > > the former is about 184KB, and the latter about 386KB, with a lot of HTML overhead. > As they have to index character name strings not just codepoint combos, they > probably need about an order of magnitude more space than compose data: > ~50KB source with lots of overhead actually ~8KB. I see. I'm expecting other issues with zero-width-whatever, but will (after current work on cdk & dialog) see about making a script to extract the data from bycodes.html -- Thomas E. Dickey <[email protected]> https://invisible-island.net
signature.asc
(application/pgp-signature, 659 B)
-----BEGIN PGP SIGNATURE----- iQGzBAABCgAdFiEEGYgtkt2kxADCLA1WzCr0RyFnvgMFAmeE1+EACgkQzCr0RyFn vgOZTgwA4ICUOYjnjjbQtx5r88HXYsRs7qWLCdLjLJtRddzomk3oZEicMK66c4Pc MNtX4LoLEjMeWR+hmjNOgXMyaAKoLPuQL26Xo5B04nXTbTUQhfk8mBLOjO7j+QYp IGvYlBZwW2DdGkfvFnEK/ntiX4MkA8sPPIs8ZEG3+Izct0E1w/vyjMM15IPe3WEN VZt37YISiTjJ7JEn1Mpm5LLxyAmTYzeZYpu+0Kr6VcBSr3ByREdFioKZrKCnmJY5 Q7u6ymvMnnEsL3DeS7C3MfpZxndv6YUjskm+QWfpwlt1nSnUsJgzH3NO62wNnQoa PKBn/meZtK+GpfmcK1TE0QTuBaTi58lJtU0hrjBUq0cN2LMSUfbHo8mIq0jGwNn6 q1Z/Q8eqeojf9cU4J+HINSIFTCrNTfBoC/VlW0v4UupNEpSD2Qf6D9gOHZ8ADHv2 QBnhJ5z9BDMXP24QLIFpsB8QncFKK8Lu5/yDQwsAvcWs8kp/e4nf3H+0AcPPdVEG Qfdj2SZ1 =4p3L -----END PGP SIGNATURE-----