Skip to content

Wrong attachement filename when mail mime header was too long #83221

Description

@manfred-kaiser
BPO 39040
Nosy @warsaw, @bitdancer, @maxking, @miss-islington, @manfred-kaiser
PRs
  • bpo-39040: added whitespaced to linesep_splitter in email.policy #17590
  • bpo-39040: Fix parsing of email mime headers with whitespace between encoded-words. #17620
  • [3.9] bpo-39040: Fix parsing of email mime headers with whitespace between encoded-words. (gh-17620) #20504
  • [3.8] bpo-39040: Fix parsing of email mime headers with whitespace between encoded-words. (gh-17620) #20505
  • [3.7] bpo-39040: Fix parsing of email mime headers with whitespace between encoded-words. (gh-17620) #20506
  • Files
  • testscript.py: Script to read filenames
  • error.eml: Mail with broken filename
  • original_mail_from_gmail.eml: downloaded mail from gmail (web interface). I only edited the the mail addresses for privacy reasons. Same problem with this mail
  • Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

    Show more details

    GitHub fields:

    assignee = None
    closed_at = None
    created_at = <Date 2019-12-13.16:59:51.619>
    labels = ['3.8', 'type-bug', '3.7', 'expert-email', '3.9']
    title = 'Wrong attachement filename when mail mime header was too long'
    updated_at = <Date 2020-05-29.11:43:51.599>
    user = 'https://github.com/manfred-kaiser'

    bugs.python.org fields:

    activity = <Date 2020-05-29.11:43:51.599>
    actor = 'miss-islington'
    assignee = 'none'
    closed = False
    closed_date = None
    closer = None
    components = ['email']
    creation = <Date 2019-12-13.16:59:51.619>
    creator = 'mkaiser'
    dependencies = []
    files = ['48775', '48776', '48777']
    hgrepos = []
    issue_num = 39040
    keywords = ['patch']
    message_count = 23.0
    messages = ['358343', '358345', '358350', '358351', '358352', '358378', '358384', '358393', '358395', '358396', '358458', '358460', '358463', '358500', '358530', '358573', '358608', '358853', '358857', '370276', '370298', '370299', '370300']
    nosy_count = 5.0
    nosy_names = ['barry', 'r.david.murray', 'maxking', 'miss-islington', 'mkaiser']
    pr_nums = ['17590', '17620', '20504', '20505', '20506']
    priority = 'normal'
    resolution = None
    stage = 'patch review'
    status = 'open'
    superseder = None
    type = 'behavior'
    url = 'https://bugs.python.org/issue39040'
    versions = ['Python 3.7', 'Python 3.8', 'Python 3.9']

    Activity

    1. manfred-kaiser commented on Dec 13, 2019

      manfred-kaisermannequin
      MannequinAuthor

      I'm working on a mailfilter in python and used the method "get_filename" of the "EmailMessage" class.

      In some cases a wrong filename was returned. The reason was, that the Content-Disposition Header had a line break and the following intention was interpreted as part of the filename.

      After fixing this bug, I was able to get the right filename.

      I had to change "linesep_splitter" in "email.policy" to match the intention.

      Old Value:

      linesep_splitter = re.compile(r'\n|\r')

      New Value:

      linesep_splitter = re.compile(r'\n\s+|\r\s+')
    2. changed the title [-]Wrong filename in when mime header was too long[/-] [+]Wrong filename in mail when mime header was too long[/+] on Dec 13, 2019
    3. changed the title [-]Wrong filename in when mime header was too long[/-] [+]Wrong filename in mail when mime header was too long[/+] on Dec 13, 2019
    4. changed the title [-]Wrong filename in mail when mime header was too long[/-] [+]Wrong attachement filename when mail mime header was too long[/+] on Dec 13, 2019
    5. changed the title [-]Wrong filename in mail when mime header was too long[/-] [+]Wrong attachement filename when mail mime header was too long[/+] on Dec 13, 2019
    6. bitdancer commented on Dec 13, 2019

      @bitdancer
      Member

      Thanks for the report. Can you provide an example that reproduces the problem?

      Per the RFC, lines may be broken before whitespace in certain places in certain headers, but that does not make the whitespace go away. Only the crlf sequence is removed when unfolding the header, per the RFC, so your proposed fix is incorrect. I suspect your example header is invalid, and the question will then become is there some sort of Postel-style error recovery we can and want to do in the function that parses the content-disposition header.

    7. manfred-kaiser commented on Dec 13, 2019

      manfred-kaisermannequin
      MannequinAuthor

      The original filename is "Schulbesuchsbestättigung.pdf", but when I use the method "get_filename" I got "Schulbesuchsbestättigung. pdf"

      I removed some headers from the mail for privacy reasons

    8. 13 remaining items

    9. maxking commented on Dec 17, 2019

      @maxking
      Contributor

      Thanks David! I applied the fixes as per your comments, can you please take another look?

    10. bitdancer commented on Dec 17, 2019

      @bitdancer
      Member

      One more tweak to the test and we'll be good to go.

    11. maxking commented on Dec 18, 2019

      @maxking
      Contributor

      Sure, fixed as per your comments in the PR.

    12. bitdancer commented on Dec 24, 2019

      @bitdancer
      Member

      I don't see the change to the test in the PR. Did you miss a push or is github doing something wonky with the review? (I haven't used github review in a while and I had forgetten how hard it is to use...)

    13. maxking commented on Dec 24, 2019

      @maxking
      Contributor

      I double checked, there should be 4 commits in the PR and last 2 have the changes that you asked for in the test case and NEWS entry.

      Your previous comment will point at the old diff, you might have to look at the full diff here: https://github.com/python/cpython/pull/17620/files or if you want, this is the diff for the 2 commits with the changes you requested: https://github.com/python/cpython/pull/17620/files/bf2cb76009d72869d9df6550b473b5818ceab311..016ceb3ef00b3b940993d35d539ce63d68437d4f

    14. bitdancer commented on May 29, 2020

      @bitdancer
      Member

      New changeset 21017ed by Abhilash Raj in branch 'master':
      bpo-39040: Fix parsing of email mime headers with whitespace between encoded-words. (gh-17620)
      21017ed

    15. miss-islington commented on May 29, 2020

      @miss-islington
      Contributor

      New changeset a6ae02d by Miss Islington (bot) in branch '3.9':
      bpo-39040: Fix parsing of email mime headers with whitespace between encoded-words. (gh-17620)
      a6ae02d

    16. miss-islington commented on May 29, 2020

      @miss-islington
      Contributor

      New changeset 6381ee0 by Miss Islington (bot) in branch '3.8':
      bpo-39040: Fix parsing of email mime headers with whitespace between encoded-words. (gh-17620)
      6381ee0

    17. miss-islington commented on May 29, 2020

      @miss-islington
      Contributor

      New changeset 5f977e0 by Miss Islington (bot) in branch '3.7':
      bpo-39040: Fix parsing of email mime headers with whitespace between encoded-words. (gh-17620)
      5f977e0

    18. transferred this issue fromon Apr 10, 2022
    19. jikamens commented on Jun 19, 2022

      @jikamens

      Strange that this hasn't been fixed when a fix was submitted literally years ago.

    20. bitdancer commented on Jun 20, 2022

      @bitdancer
      Member

      It looks like it was and the issue just wasn't closed.

    21. jikamens commented on Jun 20, 2022

      @jikamens

      @bitdancer It does not appear to have been fixed. I just ran into it with python 3.10.4:

      $ cat /tmp/testbug.py 
      #!/usr/bin/env python3
      
      import email
      
      message = email.message_from_string("""\
      MIME-Version: 1.0
      Content-Type: text/plain; name="This File Name Is
       Folded.pdf"
      Content-Disposition: attachment; filename="This File Name Is
       Folded.pdf"
      Content-Transfer-Encoding: 8bit
      
      Foo
      """)
      
      print(message.get_filename())
      $ python3 /tmp/testbug.py
      This File Name Is
       Folded.pdf
      $ python3 -V
      3.10.4
      
    22. bitdancer commented on Jun 21, 2022

      @bitdancer
      Member

      That's a different bug, and it is in the legacy API. The new API handles it correctly:

      rdmurray@pydev:~/cpython/master[master]>cat temp.py
      #!/usr/bin/env python3
      
      import email, email.policy
      
      message = email.message_from_string("""\
      MIME-Version: 1.0
      Content-Type: text/plain; name="This File Name Is
       Folded.pdf"
      Content-Disposition: attachment; filename="This File Name Is
       Folded.pdf"
      Content-Transfer-Encoding: 8bit
      
      Foo
      """, policy=email.policy.default)
      
      print(message.get_filename())
      rdmurray@pydev:~/cpython/master[master]>./python temp.py
      This File Name Is Folded.pdf
      rdmurray@pydev:~/cpython/master[master]>./python -V
      Python 3.12.0a0
      rdmurray@pydev:~/cpython/master[master]>
      

      I don't think this is worth trying to fix in the legacy api; the reason it happens there is probably a result of some deeply embedded assumptions in that code and may not be easy to fix. (What would be worth fixing is making the new API truly be the default policy, but I don't currently have time to shepherd that process).

    Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

    Metadata

    Metadata

    Assignees

    No one assigned

      Projects

      No projects

        Milestone

        No milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions