xml.dom.minidom.Node.normalize() merges adjacent text nodes with node.data = node.data + child.data, once per absorbed node. A run of n text nodes copies the growing string n times, so it's quadratic.
from xml.dom.minidom import Document
doc = Document()
root = doc.appendChild(doc.createElement('r'))
for _ in range(32_000):
root.appendChild(doc.createTextNode('x' * 48))
root.normalize() # ~450 ms; 8k nodes ~25 ms, 16k ~95 ms
The parsers already merge character data, so you only hit this with documents built through the DOM API (appendChild/createTextNode, etc.).
Fix: collect the pieces for each run and ''.join them once at the end. I'll open a PR.
Linked PRs
xml.dom.minidom.Node.normalize()merges adjacent text nodes withnode.data = node.data + child.data, once per absorbed node. A run of n text nodes copies the growing string n times, so it's quadratic.The parsers already merge character data, so you only hit this with documents built through the DOM API (appendChild/createTextNode, etc.).
Fix: collect the pieces for each run and
''.jointhem once at the end. I'll open a PR.Linked PRs