Skip to content

New sequences for Unicode groups and block ranges needed #43715

Description

@gmarketer
mannequin
BPO 1528154
Nosy @malemburg, @loewis, @terryjreedy, @devdanzin, @ezio-melotti
Dependencies
  • bpo-2636: Adding a new regex module (compatible with re)
  • Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

    Show more details

    GitHub fields:

    assignee = None
    closed_at = None
    created_at = <Date 2006-07-25.04:44:06.000>
    labels = ['expert-regex', 'type-feature', 'expert-unicode']
    title = 'New sequences for Unicode groups and block ranges needed'
    updated_at = <Date 2019-03-15.23:59:53.114>
    user = 'https://bugs.python.org/gmarketer'

    bugs.python.org fields:

    activity = <Date 2019-03-15.23:59:53.114>
    actor = 'BreamoreBoy'
    assignee = 'none'
    closed = False
    closed_date = None
    closer = None
    components = ['Regular Expressions', 'Unicode']
    creation = <Date 2006-07-25.04:44:06.000>
    creator = 'gmarketer'
    dependencies = ['2636']
    files = []
    hgrepos = []
    issue_num = 1528154
    keywords = []
    message_count = 15.0
    messages = ['54861', '54862', '54863', '54864', '54865', '54866', '84503', '84532', '84544', '100129', '130412', '185757', '185759', '221632', '221761']
    nosy_count = 8.0
    nosy_names = ['lemburg', 'loewis', 'effbot', 'terry.reedy', 'ajaksu2', 'gmarketer', 'ezio.melotti', 'mrabarnett']
    pr_nums = []
    priority = 'normal'
    resolution = None
    stage = 'test needed'
    status = 'open'
    superseder = None
    type = 'enhancement'
    url = 'https://bugs.python.org/issue1528154'
    versions = ['Python 3.4']

    Activity

    1. gmarketer commented on Jul 25, 2006

      gmarketermannequin
      MannequinAuthor

      The special sequences consist of "\" and another
      character need to be added to RE sintax to simplify the
      finding of several Unicode classes like:

      • All uppercase letters
      • All lowercase letters
    2. malemburg commented on Jul 25, 2006

      @malemburg
      Member

      Logged In: YES
      user_id=38388

      Could you make your request a little more specific ?

      We already have catregories in the re module, so adding a
      few more would be possible (patches are welcome !). However,
      we do need to know why you need them and whether there are
      other RE implementations that already have such special
      matching characters, e.g. the Perl RE implementation.

    3. gmarketer commented on Jul 26, 2006

      gmarketermannequin
      MannequinAuthor

      Logged In: YES
      user_id=1334865

      We need to process several strings in utf-8 and need to use
      regular expressions to match pattern, for ex.:
      r"[ANY_LANGUAGE_UPPERCASE_LETTER,0-9ANY_LANGUAGE_LOWERCASE_LETTER]+|NOT_ANY_LANGUAGE_CURRENCY"

      We don't know how to implement this logic by our hands.

      Also, I found this logic implemented in Microsoft dot NET
      regular expressions:

      \p{name} Matches any character in the named character
      class 'name'. Supported names are Unicode groups and block
      ranges. For example Ll, Nd, Z, IsGreek, IsBoxDrawing, and Sc
      (currency).

      \P{name} Matches text not included in the named
      character class 'name'.

      We need same logic in regular expressions.

    4. loewis commented on Sep 10, 2006

      loewismannequin
      Mannequin

      Logged In: YES
      user_id=21627

      If anything, I think Python should implement Unicode TR#18:

      http://www.unicode.org/unicode/reports/tr18/

      This does include the \p notation for property expressions,
      e.g. \p{Ll} or \p{East Asian Width:Narrow}.

      We currently don't include the Script property, so \p{Greek}
      could not be implemented (we can, of course, add support for
      the script property). I can't find anything in the report
      that makes \p{IsGreek} valid, so we shouldn't support it.

    5. effbot commented on Dec 4, 2006

      effbotmannequin
      Mannequin

      note that posix uses a special set syntax, [:name:], for this purpose:

      [:alnum:] [:cntrl:] [:lower:] [:space:]
      [:alpha:] [:digit:] [:print:] [:upper:]
      [:blank:] [:graph:] [:punct:] [:xdigit:]

      adding a new character escape will probably break more existing expressions, but no matter what syntax we chose, this is (micro-)PEP territory.

    6. effbot commented on Dec 4, 2006

      effbotmannequin
      Mannequin

      note that posix uses a special set syntax, [:name:], for this purpose:

      [:alnum:] [:cntrl:] [:lower:] [:space:]
      [:alpha:] [:digit:] [:print:] [:upper:]
      [:blank:] [:graph:] [:punct:] [:xdigit:]

      adding a new character escape will probably break more existing expressions, but no matter what syntax we chose, this is (micro-)PEP territory.

    7. devdanzin commented on Mar 30, 2009

      devdanzinmannequin
      Mannequin

      Has this been addressed for 2.6/3.0? Do the LOCALE and UNICODE constants
      cover this?

    8. loewis commented on Mar 30, 2009

      loewismannequin
      Mannequin

      No progress has been made. I still maintain that TR18 should be
      implemented.

      I'm not so sure whether the POSIX special groups should be provided. My
      understanding is that they originally were meant to integrate with the
      locale support, and change with locale. For Unicode, Annex C of TR18 makes
      a recommendation on how to provide the POSIX properties, and offers two
      alternative definitions: Standard Recommendation and POSIX Compatible.
      That alone tells me that it is best not to provide support for them:
      refuse the temptation to guess.

    9. mrabarnett commented on Mar 30, 2009

      mrabarnettmannequin
      Mannequin

      I implemented \p, \P and [:...:] for the simple categories (eg "Lu" and
      "upper", but not "IsGreek") in the work I did for issue bpo-2636.

    10. mrabarnett commented on Feb 25, 2010

      mrabarnettmannequin
      Mannequin

      \p{name} is supported for Unicode properties, scripts and blocks in my regex module (see issue bpo-2636).

      It also supports the POSIX set syntax, although I'm not sure that we really need to have 2 ways of doing it, eg \p{Alpha} and [[:Alpha:]].

    11. terryjreedy commented on Mar 9, 2011

      @terryjreedy
      Member

      Is there a practical issue left here? Mathew says his regex module does as requested, but adding that to the stdlib is a separate issue. Martin would like an implementation of Unicode TR18, but that is also another issue.

    12. terryjreedy commented on Apr 1, 2013

      @terryjreedy
      Member

      I am trying to decide if this issue still serves a purpose. It seems to be a request to add something to the existing re module. Fredrik semi-rejected the idea without a (micro)-pep. A python-ideas discussion is now another option. Matthew's regex implementation already has the feature, so this issue would be moot if it were ever part of the stdlib. But the fate of bpo-2636 is unclear. Rereading, it now seems that implementing the feature in the current re module using the TR18 syntax would be this issue, if someone were to do it. So I will not close yet.

    13. ezio-melotti commented on Apr 1, 2013

      @ezio-melotti
      Member

      We should really just include "regex" in 3.4.

    14. BreamoreBoy commented on Jun 26, 2014

      BreamoreBoymannequin
      Mannequin

      Is there an easy way to find out how many other issues have bpo-2636 as a dependency?

    15. ezio-melotti commented on Jun 28, 2014

      @ezio-melotti
      Member

      This seems to be the only one currently.
      Other issues might have closed in favor of bpo-2636 though.

    16. transferred this issue fromon Apr 10, 2022
    17. soreavis commented on Jul 19, 2026

      @soreavis
      Contributor

      \p{...}/\P{...} support landed for 3.16: the initial implementation (TR18 RL1.2 subset — general categories, binary properties, POSIX names) merged in June via #151969, with more property values in flight in #153023, all tracked under #95555. This request now lives there, so this issue looks closable as a duplicate of #95555.

      Quick check at current main (6df4993)
      import re, sys
      print(sys.version)
      for pat, s in [(r"\p{Lu}", "A"), (r"\p{Lu}", "a"), (r"\p{L}+", "abcÄ"), (r"\P{L}", "7")]:
          print(f"fullmatch({pat!r}, {s!r}) ->", bool(re.fullmatch(pat, s)))
      3.16.0a0 (remotes/upstream/HEAD:6df4993bb64, Jul 19 2026, 20:24:53) [Clang 21.0.0 (clang-2100.1.1.101)]
      fullmatch('\\p{Lu}', 'A') -> True
      fullmatch('\\p{Lu}', 'a') -> False
      fullmatch('\\p{L}+', 'abcÄ') -> True
      fullmatch('\\P{L}', '7') -> True
      
    18. terryjreedy commented on Jul 21, 2026

      @terryjreedy
      Member

      @serhiy Can you verify that the close as duplicate recommendation is correct?

    Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

    Metadata

    Metadata

    Assignees

    No one assigned

      Projects

      No projects

        Milestone

        No milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions