Skip to content

Support other RAWTEXT and PLAINTEXT elements in HTMLParser #137836

Description

@serhiy-storchaka

Bug report

HTMLParser initially only supported RAWTEXT elements "style" and "script". Then support of RCDATA elements "title" and "textarea" was added in #118350. But there are more RAWTEXT elements: "xmp", "iframe", "noembed", and "noframes".

"noscript" is also switches to the RAWTEXT mode if the scripting flag is enabled.

And the "plaintext" tag switches to the PLAINTEXT state from which there is no exit.

Support of other RAWTEXT elements can be enabled from the user code by adding them to HTMLParser.CDATA_CONTENT_ELEMENTS (this can be done for separate HTMLParser instance), but it would be better to support them by default. "plaintext" needs a special code.

Linked PRs

Activity

  1. changed the title [-]Support other RAWMODE and PLAINTEXT elements in HTMLParser[/-] [+]Support other RAWTEXT and PLAINTEXT elements in HTMLParser[/+] on Aug 15, 2025
  2. added 2 commits that reference this issue on Aug 15, 2025
  3. added
    stdlibStandard Library Python modules in the Lib/ directory
    on Aug 15, 2025
  4. moved this from Todo to In Progress in html.parser issueson Aug 16, 2025
  5. added a commit that references this issue on Oct 31, 2025
  6. moved this from In Progress to Done in html.parser issueson Oct 31, 2025
  7. added a commit that references this issue on Oct 31, 2025
  8. 4 remaining items

  9. added 6 commits that reference this issue on Oct 31, 2025
  10. added 4 commits that reference this issue on Oct 31, 2025
  11. added a commit that references this issue on Dec 6, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

stdlibStandard Library Python modules in the Lib/ directorytype-bugAn unexpected behavior, bug, or errortype-securityA security issue

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions