Skip to content

re should support \p{...} character properties #95555

Description

@LonYui

Feature or enhancement

Reproduced steps

re.compile('\p{script=Han}') # 

actual:
will throw error
expect:
pass

Info:

changed log
8/3 add more info . Final thx @JelleZijlstra give an helpful link

Linked PRs

Activity

  1. JelleZijlstra commented on Aug 2, 2022

    @JelleZijlstra
    Member

    Could you clarify why you think this should work? As far as I can tell \p has no special meaning in regex.

  2. ronaldoussoren commented on Aug 2, 2022

    @ronaldoussoren
    Contributor

    They may mean character properties as described here: https://www.unicode.org/reports/tr18/#property_syntax

    Perl has similar functionality.

  3. mrabarnett commented on Aug 2, 2022

    @mrabarnett

    Unicode properties are supported in the 'regex' package on PyPI here: https://pypi.org/project/regex/

  4. changed the title [-]re: 應該支援漢字正則表達式 \p{Han}[/-] [+]re: should support Chinese character regular expression \p{Han}[/+] on Aug 3, 2022
  5. corona10 commented on Aug 3, 2022

    @corona10
    Member

    @LonYui

    re: 應該支援漢字正則表達式 \p{Han}

    Please write the issue title in English if possible, some of the core devs are not familiar with Chinese.
    (Same reason why I do not write issue titles in Korean.)
    For reducing the cost of understanding your suggestion, I re-wrote the issue title in English.

  6. arhadthedev commented on Mar 4, 2023

    @arhadthedev
    Member

    Removing the pending label because @ronaldoussoren explained what the OP means.

  7. added
    stdlibStandard Library Python modules in the Lib/ directory
    and removed
    pendingThe issue will be closed if no feedback is provided
    on Mar 4, 2023
  8. changed the title [-]re: should support Chinese character regular expression \p{Han}[/-] [+]`re` should support `\p{...}` character properties[/+] on Mar 4, 2023
  9. arhadthedev commented on Mar 4, 2023

    @arhadthedev
    Member

    Note: currently re supports only \number, \a, \b, \f, \n, \N, \r, \t, \u, \U, \v, \x, and \\.

  10. ncoghlan commented on Jun 24, 2025

    @ncoghlan
    Contributor

    As a concrete example where this limitation can cause problems: while JSON Schema formally only supports a limited number of escape sequences, in practice schemas may end up containing patterns that use any of the defined JavaScript regex character classes.

    Python supports most of the same character classes, but \p/\P is a notable omission. Whether or not that can reasonably be worked around by using regex instead of re depends on how the affected regex is being processed (if it's in a third party library, it may not be straightforward to get it to use a different regex engine)

  11. added 4 commits that reference this issue on Jun 23, 2026
  12. added a commit that references this issue on Jul 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    stdlibStandard Library Python modules in the Lib/ directorytopic-regextype-featureA feature request or enhancement

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions