Bug report
Bug description:
Documented behaviour: The unicodedata documentation states: "Return the normal form form for the Unicode string unistr." Its ucd_3_2_0 entry states: "This is an object that has the same methods as the entire module, but uses the Unicode database version 3.2 instead". Unicode canonical ordering, referenced by the normalization documentation, does not reorder combining characters across a character with canonical combining class zero.
Expected: u.normalize('NFD', 'a\u1ab0\u0316') returns 'a\u1ab0\u0316'.
Actual: Returns 'a\u0316\u1ab0', moving U+0316 across a Unicode 3.2 starter.
import unicodedata as ud
u = ud.ucd_3_2_0
x, y = '\u1ab0', '\u0316'
s = 'a' + x + y
if not (u.category(x) == 'Cn' and u.combining(x) == 0
and u.decomposition(x) == '' and u.combining(y) > 0
and u.decomposition(y) == '' and u.decomposition('a') == ''
and u.combining('a') == 0):
print('REFUTATION REJECTED: input does not satisfy the premises')
else:
ref = []
for c in s:
ref.append(c)
i = len(ref) - 1
while i and 0 < u.combining(ref[i]) < u.combining(ref[i - 1]):
ref[i - 1], ref[i] = ref[i], ref[i - 1]
i -= 1
expected = ''.join(ref)
actual = u.normalize('NFD', s)
if actual != expected:
print('REFUTATION CONFIRMED:', repr(s), 'actual =', repr(actual),
'expected =', repr(expected))
else:
print('REFUTATION REJECTED: actual matches documented expectation')
Output on Python 3.14.6 (Windows-11-10.0.26220-SP0), standard library unicodedata:
REFUTATION CONFIRMED: 'a̖᪰' actual = 'a̖᪰' expected = 'a̖᪰'
This report was found and written by an automated property-testing tool I run (bugforge). The reproducer above was executed and its output is pasted unedited; no person reviewed the report before it was filed. The search script is in https://github.com/augusto-rehfeldt/bugforge-results/tree/main/unicodedata-20261003-064954-c2
CPython versions tested on:
3.14
Operating systems tested on:
Windows
Bug report
Bug description:
Documented behaviour: The unicodedata documentation states: "Return the normal form form for the Unicode string unistr." Its ucd_3_2_0 entry states: "This is an object that has the same methods as the entire module, but uses the Unicode database version 3.2 instead". Unicode canonical ordering, referenced by the normalization documentation, does not reorder combining characters across a character with canonical combining class zero.
Expected: u.normalize('NFD', 'a\u1ab0\u0316') returns 'a\u1ab0\u0316'.
Actual: Returns 'a\u0316\u1ab0', moving U+0316 across a Unicode 3.2 starter.
Output on Python 3.14.6 (Windows-11-10.0.26220-SP0), standard library
unicodedata:This report was found and written by an automated property-testing tool I run (bugforge). The reproducer above was executed and its output is pasted unedited; no person reviewed the report before it was filed. The search script is in https://github.com/augusto-rehfeldt/bugforge-results/tree/main/unicodedata-20261003-064954-c2
CPython versions tested on:
3.14
Operating systems tested on:
Windows