Repository navigation
HTMLParser differences from the HTML5 specification #135661
Copy link
Copy link
Open
3 / 33 of 3 issues completedLabels
3.10 (EOL)end of lifeend of life3.11only security fixesonly security fixes3.15bugs and security fixesbugs and security fixes3.9 (EOL)end of lifeend of lifestdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directorytype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or errortype-securityA security issueA security issue
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error3.13only security fixesonly security fixes3.14bugs and security fixesbugs and security fixes3.15bugs and security fixesbugs and security fixes
on Jun 18, 2025 - changed the title
[-]HTMLParser differences for the HTML5 specification[/-][+]HTMLParser differences from the HTML5 specification[/+]on Jun 18, 2025 - addedstdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directory
on Jun 19, 2025 Null character (U+0000) should not end the tag name.
I think we have some solutions if there is a \x00 at the end of a tag name:
- raise an Exception
- raise a warning and prase it without the last \x00
- raise a warning and prase it with the last \x00
- raise a warning and prase it with the last \x00 changed to \xfffd
- directly prase it with the last \x00 changed to \xfffd
- directly prase it without the last \x00
- added a commit that references this issue
on Jul 3, 2025 76 remaining items
All that remains is exposing the context-dependent CDATA parsing in some better way, but that's not a release blocker.
- added 4 commits that reference this issue
on Jul 4, 2026 #153028 provide complete fix of the CDATA issue on the stdlib side. I wonder if it can be backported.
Metadata
Metadata
Assignees
Labels
3.10 (EOL)end of lifeend of life3.11only security fixesonly security fixes3.15bugs and security fixesbugs and security fixes3.9 (EOL)end of lifeend of lifestdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directorytype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or errortype-securityA security issueA security issue
Projects
- StatusShow more project fieldsTodo
- StatusShow more project fieldsTodo
Bug report
Originally, the definition of the HTML format was not formally strict. It was similar to SGML and XML, but with a lot of looseness.
HTMLParsertried its best to parse anything that looked like HTML. But after creation of HTML5, its specification defines the parsing rules for HTML documents, whether they are syntactically correct or not. It is important to follow these rules, for security reasons.The current
HTMLParsermainly follows the HTML5 specification, but there are a number of differences:--!>should end the comment. gh-102555: Fix comment parsing in HTMLParser #135664-- >should not end the comment. gh-102555: Fix comment parsing in HTMLParser #135664<-->and<--->should be abnormally ended empty comments. gh-102555: Fix comment parsing in HTMLParser #135664] ]>and]] >should not end the CDATA section. gh-135661: Fix CDATA section parsing in HTMLParser #135665]]>and>).</and the tag name. E.g.</ script>should not end the script section. gh-135661: Fix parsing start and end tags in HTMLParser #135930\v) and non-ASCII whitespaces should not be recognized as whitespaces. The only whitespaces are\t\n\r\f. gh-135661: Fix parsing start and end tags in HTMLParser #135930\xfffd. I think we can leave this, because it is easy to do in pre-processing or post-processing, and they usually do not cause issues in Python.>. E.g.</script/foo=">"/>. gh-135661: Fix parsing start and end tags in HTMLParser #135930</script>does not match</ſcript>, andLINKdoes not matchLINK(the last letter is U+212A). gh-135661: Fix parsing start and end tags in HTMLParser #135930>in both start and end tags. E.g.<a foo=bar/ //>. gh-135661: Fix parsing start and end tags in HTMLParser #135930=separator between attribute name and value. E.g.<a foo==bar>should have attribute "foo" with value "=bar". gh-135661: Fix parsing start and end tags in HTMLParser #135930No whitespace should be acceptable between the=separator and attribute name and value. E.g.<a foo =bar>should have two attributes "foo" and "=bar", both with value None;<a foo= bar>should have two attributes: "foo" with value "" and "bar" with value None.This can cause security issues for some programs. If the program uses
HTMLParserto check the HTML input for dangerous code, it can miss some code. For example, "<!----!><script>...</script><!---->" is parsed by browsers as a script block surrounded by two comments, but the currentHTMLParserparses it as a single comment.Linked PRs