Repository navigation
[doc] Clarify exactly what \w matches in UNICODE mode #69929
Description
Activity
The
remodule documentation does not do a good job of explaining exactly what\wmatches. Quoting https://docs.python.org/3.5/library/re.html :\w
For Unicode (str) patterns:
Matches Unicode word characters; this includes most characters
that can be part of a word in any language, as well as numbers
and the underscore.Empirically, this appears to mean "everything in Unicode general categories L* and N*, plus U+005F (underscore)". That is a perfectly sensible definition and the documentation should state it in those terms. "Unicode word characters" could mean any number of different things; note for instance that UTS#18 gives a very different definition.
(Further reading: https://gist.github.com/zackw/3077f387591376c7bf67 plus links therefrom).
I would like to request also a clear explanation be given for the documentation in the 2.7 branch. From https://docs.python.org/2.7/library/re.html :
"\w ... If UNICODE is set, this will match the characters [0-9_] plus whatever is classified as alphanumeric in the Unicode character properties database"
This is ambiguous. Does it mean the "Alphabetic" property from UAX#44? Does it mean something else?
FWIW, the actual behavior of \w matching "everything in Unicode general categories L* and N*, plus U+005F (underscore)" is consistent across all versions I can conveniently test (2.7, 3.4, 3.5).
In 2.7, there are four characters in general category Nl that \w doesn't match, but I believe that is just a bug, not an intentional difference of behavior.
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Jan 4, 2016 It's too late for the 2.7 docs, but the current docs can still be updated.
- added3.9 (EOL)end of lifeend of life3.10 (EOL)end of lifeend of life3.11only security fixesonly security fixes
on Dec 1, 2021 - changed the title
[-]Clarify exactly what \w matches in UNICODE mode[/-][+][doc] Clarify exactly what \w matches in UNICODE mode[/+]on Dec 1, 2021 Would a change like this be accurate?
Matches Unicode word characters; this includes most alphanumeric characters as well as the underscore. In Unicode, alphanumeric characters are defined to be the general categories L + N (see https://unicode.org/reports/tr44/#General_Category_Values). If the :const:`ASCII` flag is used, only ``[a-zA-Z0-9_]`` is matched.
- added a commit that references this issue
on Dec 20, 2022 - added 4 commits that reference this issue
on Dec 20, 2022 - added a commit that references this issue
on Dec 28, 2022
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields:
Linked PRs