Skip to content

Inconsistency between quopri.decodestring(), email.quoprimime.decode() and binascii.a2b_qp() #62222

Description

@serhiy-storchaka
BPO 18022
Nosy @gvanrossum, @warsaw, @bitdancer, @jeremyhylton, @vadmium, @serhiy-storchaka

Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

Show more details

GitHub fields:

assignee = None
closed_at = None
created_at = <Date 2013-05-20.14:13:27.994>
labels = ['3.7', 'type-bug', 'library', 'expert-email']
title = 'Inconsistency between quopri.decodestring() and email.quoprimime.decode()'
updated_at = <Date 2018-05-15.11:59:26.100>
user = 'https://github.com/serhiy-storchaka'

bugs.python.org fields:

activity = <Date 2018-05-15.11:59:26.100>
actor = 'r.david.murray'
assignee = 'none'
closed = False
closed_date = None
closer = None
components = ['Library (Lib)', 'email']
creation = <Date 2013-05-20.14:13:27.994>
creator = 'serhiy.storchaka'
dependencies = []
files = []
hgrepos = []
issue_num = 18022
keywords = []
message_count = 12.0
messages = ['189663', '190874', '190876', '190889', '197665', '197715', '290521', '290525', '290533', '316563', '316622', '316642']
nosy_count = 6.0
nosy_names = ['gvanrossum', 'barry', 'r.david.murray', 'Jeremy.Hylton', 'martin.panter', 'serhiy.storchaka']
pr_nums = []
priority = 'normal'
resolution = None
stage = None
status = 'open'
superseder = None
type = 'behavior'
url = 'https://bugs.python.org/issue18022'
versions = ['Python 2.7', 'Python 3.5', 'Python 3.6', 'Python 3.7']

Linked PRs

Activity

  1. serhiy-storchaka commented on May 20, 2013

    @serhiy-storchaka
    MemberAuthor
    >>> import quopri, email.quoprimime
    >>> quopri.decodestring(b'==41')
    b'=41'
    >>> email.quoprimime.decode('==41')
    '=A'

    I don't see a rule about double '=' in RFC 1521-1522 or RFCs 2045-2047 and I think quopri is wrong.

    Other half of this bug (encoding '=' as '==') was fixed in 9bc52706d283.

  2. added
    stdlibStandard Library Python modules in the Lib/ directory
    type-bugAn unexpected behavior, bug, or error
    on May 20, 2013
  3. serhiy-storchaka commented on Jun 9, 2013

    @serhiy-storchaka
    MemberAuthor

    There are other inconsistencies. email.quoprimime.decode(), binascii.a2b_qp() and pure Python (by default binascii used) quopri.decodestring() returns different results for following data:

               quoprimime  binascii  quopri
    
     b'='      ''          b''       b'='
     b'=='     '='         b'='      b'=='
     b'= '     ''          b'= '     b'= '
     b'= \n'   ''          b'= \n'   b''
     b'=\r'    ''          b''       b'=\r'
     b'==41'   '=A'        b'=41'    b'=A'
    
  4. bitdancer commented on Jun 9, 2013

    @bitdancer
    Member

    Most of the variations represent different invalid-input recovery choices. I believe binascii's decoding of b'= \n' is incorrect, as is its decoding of b'==41'. quopri's decoding of b'=\r' is arguably incorrect as well, given that python generally supports universal line ends. Otherwise the decodings are all responses to erroneous input for which the behavior is not specified.

    That said, we ought to pick one error recovery scheme and implement it in all places, and IMO it shouldn't be exactly any of the ones we've got. Or better yet, use one common implementation. Untangling quopri is on my (too large) List of Things To Do :)

  5. serhiy-storchaka commented on Jun 10, 2013

    @serhiy-storchaka
    MemberAuthor

    Perl's MIME::QuotedPrint produces same result as pure Python quopri. konwert qp-8bit produces same result as binascii (except '==41' it decodes as '=A').

    RFC 2045 says:

    """A
    reasonable approach by a robust implementation might be
    to include the "=" character and the following
    character in the decoded data without any
    transformation and, if possible, indicate to the user
    that proper decoding was not possible at this point in
    the data.
    """

  6. serhiy-storchaka commented on Sep 13, 2013

    @serhiy-storchaka
    MemberAuthor

    So what scheme we will picked?

  7. bitdancer commented on Sep 14, 2013

    @bitdancer
    Member

    As I said, not exactly any of the above.

    I'll get back to this after I finish the new email code (which should happen before the end of the month). I need to take some time to look over the RFCs and real world examples and come up with the most appropriate rules.

  8. serhiy-storchaka commented on Mar 26, 2017

    @serhiy-storchaka
    MemberAuthor

    Ping.

  9. vadmium commented on Mar 26, 2017

    @vadmium
    Member

    The double equals "==" case for the “quopri” implementation in Python is now consistent with the others thanks to the fix in bpo-23681 (see also bpo-21511).

    According to bpo-20121, the quopri (Python) implementation only supports LF (\n) characters as line breaks, and the binascii (C) implementation also supports CRLF. So I agree that the whitespace-before-newline case "= \n" is a genuine bug (see bpo-16473). But the CR case "=\r" is not supported because neither quopri nor binascii support universal newlines or CR line breaks on their own.

  10. serhiy-storchaka commented on Mar 26, 2017

    @serhiy-storchaka
    MemberAuthor

    Thus currently the table of discrepancies looks as:

               quoprimime  binascii  quopri
    
     b'='      ''          b''       b'='
     b'= '     ''          b'= '     b'= '
     b'= \n'   ''          b'= \n'   b''
     b'=\r'    ''          b''       b'=\r'
     b'==41'   '=A'        b'=41'    b'=41'
    
  11. bitdancer commented on May 14, 2018

    @bitdancer
    Member

    OK, I've finally gotten around to looking at this. It looks like quopri and binascii are not stripping trailing whitespace.

                  quoprimime  binascii     quopri       preferred
    
     b'='         ''          b''          b'='         '='
     b'= '        ''          b'= '        b'= '        '='
     b'= \n'      ''          b'= \n'      b''                quoprimime  binascii  quopri
    
     b'='      ''          b''       b'='
     b'= '     ''          b'= '     b'= '
     b'= \n'   ''          b'= \n'   b''
     b'=\r'    ''          b''       b'=\r'
     b'==41'   '=A'        b'=41'    b'=41'    '=\n'
     b'=\r'       ''          b''          b'=\r'       '=\r'
     b'==41'      '=A'        b'=41'       b'=41'       '=A'
     b'= \n f\n'  ' f\n'      b'= \n f\n'  b'= \n f\n'  ' f\n'
    

    The RFC recommends that a trailing = be preserved, but that trailing whitespace be ignored. It doesn't speak directly to the ==41 case, but one can infer that the first = in the == pair is most likely to have "come from the source text" and not been encoded, while the =41 was an intentional encoding and so should be decoded.

    Now, that said, the actual behavior that our libraries have had for a long time is to treat the "last line" just like all other lines, and strip a trailing =. So I would be inclined to keep that behavior for backward compatibility reasons rather than change it to be more RFC compliant, given that we don't have any actual bug report related to it, and "fixing" it could break things. Given that, the current quoprimime behavior becomes the reference.

    However, backward compatibility concerns also arise around starting to strip trailing space in quopri and binascii. Maybe we only make that change in 3.8?

  12. serhiy-storchaka commented on May 15, 2018

    @serhiy-storchaka
    MemberAuthor

    Many thanks David! But sorry, your table confused me. I can't read it. Could you please reformat it?

  13. bitdancer commented on May 15, 2018

    @bitdancer
    Member

    I should have just deleted the table, actually.

    The only important info in it is that per RFC '=', '=\n', and '= \n' all ought to become '='. But I don't think we should make that change, I think we should continue to turn those into ''. So I consider the (current!) bwehavior of quoprimime to be the correct behavior.

    I also gave the example of '= \n foo\n', to show that quopri and binascii aren't stripping trailing blanks, as Martin noted in the other issue. They fold lines if they see '=\n', but not if they see '= \n', which is wrong per the (email!) RFC. I'm not clear if it is wrong for non-email uses of quopric, I haven't tried to research that.

  14. transferred this issue fromon Apr 10, 2022
  15. picnixz commented on Dec 6, 2024

    @picnixz
    Member

    Some real-world scenario where our test assumptions were wrong and this issue was (once again) found: https://github.com/python/cpython/actions/runs/12200589926/job/34037087571#step:22:214.

    IMO, whatever we choose, I'd like dec(enc(x)) == x. I don't know however how it could be disruptive to the email part =/

  16. added and removed on Dec 6, 2024
  17. changed the title [-]Inconsistency between quopri.decodestring() and email.quoprimime.decode()[/-] [+]Inconsistency between quopri.decodestring(), email.quoprimime.decode() and `binascii.a2b_qp()`[/+] on Dec 6, 2024
  18. serhiy-storchaka commented on Jun 30, 2026

    @serhiy-storchaka
    MemberAuthor

    Some current data on how the three functions, and a few external implementations, decode the ambiguous cases.

    CPython today

    quopri.decodestring() calls binascii.a2b_qp when binascii can be imported, and otherwise uses its own pure-Python loop in Lib/quopri.py. The two differ: the pure-Python loop strips trailing whitespace on \n-terminated lines, the binascii C path does not. In the table, "binascii" is therefore also what quopri.decodestring() returns by default; "quopri(pure)" is the binascii-unavailable fallback.

                quoprimime  binascii(=default quopri)  quopri(pure)
     b'='       ''          b''                        b'='
     b'=='      '='         b'='                       b'='
     b'= '      ''          b'= '                      b'= '
     b'= \n'    ''          b'= \n'                    b''
     b'=\r'     ''          b''                        b'=\r'
     b'==41'    '=A'        b'=41'                     b'=41'
     b'foo  \n' 'foo\n'     b'foo  \n'                 b'foo\n'
    

    (The pure path strips only on \n-terminated lines, so b'= ' at EOF stays b'= '.)

    Malformed = recovery (e.g. ==41)

    RFC 2045 §6.7 calls such a sequence "illegal" and gives only a non-normative suggestion: "A reasonable approach by a robust implementation might be to include the = character and the following character in the decoded data without any transformation." Observed results for ==41:

    behavior ==41 → implementations
    emit =, advance 1, rescan =A email.quoprimime; Thunderbird; Perl MIME::QuotedPrint 3.16; Go mime/quotedprintable; Node libqp; konwert
    emit =, advance 2 =41 binascii / quopri (the second = is dropped)
    drop the =, advance 1, rescan A Gmail (web)
    reject as an error — PHP quoted_printable_decode(); Apache commons-codec

    Verification: Perl run locally; Go/Node/PHP/commons-codec read from source; Gmail measured by APPENDing a hand-built QP message and reading the downloaded .eml plus the rendered body. (Gmail's "Show original" view is normalized — it silently rewrites malformed QP — so only the downloaded message reflects the stored bytes.)

    Trailing whitespace before a line end

    RFC 2045 §6.7, Rule #3: "when decoding a Quoted-Printable body, any trailing white space on a line must be deleted, as it will necessarily have been added by intermediate transport agents."

    • Strips it: email.quoprimime; quopri pure-Python path (on \n-terminated lines).
    • Keeps it: binascii.a2b_qp (hence default quopri); Gmail (measured); Thunderbird (measured; its mimeenc.cpp has a comment noting the non-compliance).

    Interior whitespace and encoded trailing whitespace (=20/=09) are handled identically by all of quoprimime/binascii/quopri.

  19. serhiy-storchaka commented on Jun 30, 2026

    @serhiy-storchaka
    MemberAuthor

    I think we can resolve this in steps, aiming for a single quoted-printable decoder instead of three slightly different ones.

    1. Fix b'==41' first. Here binascii/quopri are the outliers: they advance by two and silently drop the second =, giving =41, while quoprimime and everything else surveyed (Thunderbird, Perl, Go, Node, …) emit the stray = and re-scan, giving =A. I'd change binascii.a2b_qp to match. This likely resolves some of the other rows too, so it comes first.

    2. Then trailing whitespace. Either delete it unconditionally (per RFC 2045 §6.7, as quoprimime does), or add a flag to binascii/quopri to opt into stripping — keeping the non-stripping behavior available for compatibility with Thunderbird and Gmail, which also keep it.

    3. Then line ends. Decide whether end of input or a bare \r counts as a line end for the soft-break and trailing-whitespace rules (the b'=\r' and final-partial-line cases).

    4. Finally, consolidate. With binascii/quopri correct, quoprimime can defer to them. The pure-Python fallback in quopri can likely go too, since binascii is already required unconditionally (base64 depends on it). That leaves one implementation.

    I'll start with step 1.

  20. serhiy-storchaka commented on Jun 30, 2026

    @serhiy-storchaka
    MemberAuthor

    One caveat on email.quoprimime.decode, which we have been comparing against: the email module does not actually use it for body decoding. get_payload(decode=True) decodes via quopri, not via quoprimime, and quopri does not strip trailing whitespace — so the body path is non-compliant with RFC 2045, even though the package ships a compliant decoder (quoprimime.decode) that it bypasses.

    This also looks unintentional. The body path used to strip: it went through quopri.decodestring, which was pure Python and stripped trailing whitespace per line. Patch #462190 (16dc7f4, 2001) added a2b_qp to binascii and made quopri delegate to it for speed; the C decoder never replicated the stripping, so the body path silently became non-compliant and has stayed that way.

  21. added a commit that references this issue on Jun 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    extension-modulesC modules in the Modules dirstdlibStandard Library Python modules in the Lib/ directorytopic-emailtype-bugAn unexpected behavior, bug, or error

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions