note 102756 added to regexp.reference.unicode

[email protected]
Newsgroups php.notes
Message-ID <[email protected]>
There is a possibility to use \p{xx} and \P{xx} escape sequences with script names.

From http://www.pcre.org/pcre.txt

When PCRE is built with Unicode character property support, three addi-
tional escape sequences that match characters with specific  properties
are  available.   When not in UTF-8 mode, these sequences are of course
limited to testing characters whose codepoints are less than  256,  but
they do work in this mode.  The extra escape sequences are:

  \p{xx}   a character with the xx property
  \P{xx}   a character without the xx property
  \X       an extended Unicode sequence

The  property  names represented by xx above are limited to the Unicode
script names, the general category properties, "Any", which matches any
character   (including  newline),  and  some  special  PCRE  properties
(described in the next section).  Other Perl properties such as  "InMu-
sicalSymbols"  are  not  currently supported by PCRE. Note that \P{Any}
does not match any characters, so always causes a match failure.

Sets of Unicode characters are defined as belonging to certain scripts.
A  character from one of these sets can be matched using a script name.
For example:

  \p{Greek}
  \P{Han}

Those that are not part of an identified script are lumped together  as
"Common". The current list of scripts is:

Arabic, Armenian, Avestan, Balinese, Bamum, Bengali, Bopomofo, Braille,
Buginese, Buhid, Canadian_Aboriginal, Carian, Cham,  Cherokee,  Common,
Coptic,   Cuneiform,  Cypriot,  Cyrillic,  Deseret,  Devanagari,  Egyp-
tian_Hieroglyphs,  Ethiopic,  Georgian,  Glagolitic,   Gothic,   Greek,
Gujarati,  Gurmukhi,  Han,  Hangul,  Hanunoo,  Hebrew,  Hiragana, Impe-
rial_Aramaic, Inherited, Inscriptional_Pahlavi, Inscriptional_Parthian,
Javanese,  Kaithi, Kannada, Katakana, Kayah_Li, Kharoshthi, Khmer, Lao,
Latin,  Lepcha,  Limbu,  Linear_B,  Lisu,  Lycian,  Lydian,  Malayalam,
Meetei_Mayek,  Mongolian, Myanmar, New_Tai_Lue, Nko, Ogham, Old_Italic,
Old_Persian, Old_South_Arabian, Old_Turkic, Ol_Chiki,  Oriya,  Osmanya,
Phags_Pa,  Phoenician,  Rejang,  Runic, Samaritan, Saurashtra, Shavian,
Sinhala, Sundanese, Syloti_Nagri, Syriac,  Tagalog,  Tagbanwa,  Tai_Le,
Tai_Tham,  Tai_Viet,  Tamil,  Telugu,  Thaana, Thai, Tibetan, Tifinagh,
Ugaritic, Vai, Yi.

Each character has exactly one Unicode general category property, spec-
ified  by a two-letter abbreviation. For compatibility with Perl, nega-
tion can be specified by including a  circumflex  between  the  opening
brace  and  the  property  name.  For  example,  \p{^Lu} is the same as
\P{Lu}.

If only one letter is specified with \p or \P, it includes all the gen-
eral  category properties that start with that letter. In this case, in
the absence of negation, the curly brackets in the escape sequence  are
optional; these two examples have the same effect:

  \p{L}
  \pL
----
Server IP: 195.54.192.44
Probable Submitter: 127.0.0.1
----
Manual Page -- http://www.php.net/manual/en/regexp.reference.unicode.php
Edit        -- https://master.php.net/note/edit/102756
Del: integrated  -- https://master.php.net/note/delete/102756/integrated
Del: useless     -- https://master.php.net/note/delete/102756/useless
Del: bad code    -- https://master.php.net/note/delete/102756/bad+code
Del: spam        -- https://master.php.net/note/delete/102756/spam
Del: non-english -- https://master.php.net/note/delete/102756/non-english
Del: in docs     -- https://master.php.net/note/delete/102756/in+docs
Del: other reasons-- https://master.php.net/note/delete/102756
Reject      -- https://master.php.net/note/reject/102756
Search      -- https://master.php.net/manage/user-notes.php
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.