Repository navigation
Support other RAWTEXT and PLAINTEXT elements in HTMLParser #137836
Copy link
Copy link
Closed
Labels
stdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directorytype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or errortype-securityA security issueA security issue
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Aug 15, 2025 - changed the title
[-]Support other RAWMODE and PLAINTEXT elements in HTMLParser[/-][+]Support other RAWTEXT and PLAINTEXT elements in HTMLParser[/+]on Aug 15, 2025 - addedstdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directory
on Aug 15, 2025 - linked a pull request that will close this issuegh-137836: Support more RAWTEXT and PLAINTEXT elements in HTMLParser #137837
on Sep 19, 2025 - added a commit that references this issue
on Oct 31, 2025 4 remaining items
- added 6 commits that reference this issue
on Oct 31, 2025 - added 4 commits that reference this issue
on Oct 31, 2025 - removed a parent issue
on Oct 31, 2025 - added a commit that references this issue
on Nov 4, 2025
Metadata
Metadata
Assignees
Labels
stdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directorytype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or errortype-securityA security issueA security issue
Projects
- StatusShow more project fieldsDone
Bug report
HTMLParserinitially only supported RAWTEXT elements "style" and "script". Then support of RCDATA elements "title" and "textarea" was added in #118350. But there are more RAWTEXT elements: "xmp", "iframe", "noembed", and "noframes"."noscript" is also switches to the RAWTEXT mode if the scripting flag is enabled.
And the "plaintext" tag switches to the PLAINTEXT state from which there is no exit.
Support of other RAWTEXT elements can be enabled from the user code by adding them to
HTMLParser.CDATA_CONTENT_ELEMENTS(this can be done for separateHTMLParserinstance), but it would be better to support them by default. "plaintext" needs a special code.Linked PRs