Repository navigation
imaplib should support international mailbox names #49555
Description
Activity
The IMAP4rev1 specification allows for non-ASCII mailbox names using a
modified UTF-7 encoding (section 5.1.3 of RFC 2060 or 3501). However,
the imaplib routines taking a mailbox name just pass the string straight
through without any encoding.It would be useful if Python provided an encoder/decoder for the
modified UTF-7 encoding, and optionally if imaplib would perform the
encoding and decoding at the appropriate points.- addedstdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directory
on Feb 18, 2009 Can you provide a patch?
I'll have a go at implementing the algorithm. It looks like the
modifications to UTF-7 are large enough that you can't do a search and
replace on the output of the existing UTF-7 codec, so it'll probably
require new code.Would String2Mailbox and Mailbox2String utility functions be appropriate
here?IMAP4 UTF-7 is implemented in Twisted -
<http://twistedmatrix.com/trac/browser/trunk/twisted/mail/imap4.py#L5385\>,
<http://twistedmatrix.com/trac/browser/trunk/twisted/mail/test/test_imap.py#L58\>.
Feel free to re-use any of that code that would be helpful.I don't have a good understanding of imaplib; if you think it's
appropriate to provide the conversion through two functions, I trust you.The IMAP4rev1 specification allows for non-ASCII mailbox
names using a modified UTF-7 encodingUTF-7 already sounds like something horrible for me, but a *modified*
UTF-7 encoding is something a little bit more strange for me. Why not
reusing directly UTF-7.(sorry, it's an off topic dummy question)
UTF-7 already sounds like something horrible for me, but a *modified*
UTF-7 encoding is something a little bit more strange for me. Why not
reusing directly UTF-7.UTF-7 wasn't horrible for its time, but its time has very likely passed.
Alas, changing a standard like IMAP4 is so difficult, this mistake will
be with us for a long time to come.As for why IMAP4 uses a modified form of UTF-7, the RFC addresses this:
The purpose of these modifications is to correct the following
problems with UTF-7:1) UTF-7 uses the "+" character for shifting; this conflicts with the common use of "+" in mailbox names, in particular USENET newsgroup names. 2) UTF-7's encoding is BASE64 which uses the "/" character; this conflicts with the use of "/" as a popular hierarchy delimiter. 3) UTF-7 prohibits the unencoded usage of "\"; this conflicts with the use of "\" as a popular hierarchy delimiter. 4) UTF-7 prohibits the unencoded usage of "~"; this conflicts with the use of "~" in some servers as a home directory indicator. 5) UTF-7 permits multiple alternate forms to represent the same string; in particular, printable US-ASCII characters can be represented in encoded form.Whether you are convinced by these arguments or not is, of course,
entirely up to you. Note also, however, that the modified UTF-7 is not
mandated by the RFC:By convention, international mailbox names in IMAP4rev1 are specified
using a modified version of the UTF-7 encoding described in [UTF-7].
Modified UTF-7 may also be usable in servers that implement an
earlier version of this protocol.However, it seems stupid to say that the choice if encoding is only a
convention since there is no other way to communicate the choice of
encoding between client and server.twisted's code does not work good for "\t", "\r", "\n", those characters must encoded in modified base64 form according to RFC 3501.
So noone is working on this issue ATM?
There's a working implementation of this in PloneMailList.
http://svn.plone.org/svn/collective/mxmImapClient/trunk/imapUTF7.pyI've used the PloneMailList implementation in another project. It works well to add 'imap4-utf-7' as codec.
The twisted imap implementation seems to have been updated to properly support non-printable ASCII, but the twisted imap API is problematic for imaplib because twisted seems to expect its arguments to already be Python unicode.
So can we be specific about what kind of API change would satisfy this issue:
-
a number of API methods take one or more mailbox arguments. Of course, imaplib currently expects these to be ASCII, but what kind of argument should the methods take? UTF? Unicode? So would the library need a class property to describe an optional specified input encoding? Would it be expected to take Python unicode?
-
some methods, such as list and lsub, return mailbox names UTF-7 encoded and embedded in larger ASCII strings. Would imaplib be expected to alter the contents of these large strings and transform them into another other encoding (when a switch as described in 1) is active)?
-
Being bitten by this today.
Point 2 of cfraire message is a big issue.
What about leaving this problem to the library user simply providing two helper functions in the module to encode/decode mUTF-7?.
Or a new encoder/decoder in "codecs" module.
the twisted imap API is problematic for imaplib because twisted seems to expect its arguments to already be Python unicode.
Could you elaborate on this? As far as I can tell, it works fine:
>>> import twisted.mail.imap4 >>> print u"Hello, \N{SNOWMAN}".encode('imap4-utf-7') Hello, &JgM- >>> print b'Hello, &JgM-'.decode('imap4-utf-7') Hello, ☃ >>>What would you expect to work differently?
> the twisted imap API is problematic for imaplib because twisted seems to expect its arguments to already be Python unicode.
Could you elaborate on this? As far as I can tell, it works fine:twisted imap4-utf-7 seems to be improved in this 2 years. :-)
First step is to provide mUTF-7 in Python 3.5. Then we can try to update imaplib. I am specially worried about the points cfraire raises in http://bugs.python.org/issue5305#msg151859. Lets see.
> the twisted imap API is problematic for imaplib because twisted seems to expect its arguments to already be Python unicode.
Could you elaborate on this? As far as I can tell, it works fine:
I wasn't addressing encode/decode specifically. Both twisted and PloneMailList offer implementations with same encoding name, "imap4-utf-7".
I meant that it's difficult for the twisted API to inform what might be done for imaplib since twisted takes full unicode but imaplib expects only unicode-ASCII subset.
The first part of jamesh's original issue is just encoder/decoder, so either twisted or PloneMailList would seem to suffice. I was addressing jamesh's second part whether "optionally if imaplib would perform the encoding and decoding at the appropriate points."
Point 2 of my response seems the more difficult. imaplib list and lsub return str instances with ASCII + utf-7 stuffed together. (twisted avoids this by returning tuples of unicode, if I understand correctly).
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Jul 22, 2015 Ping.
ssu
PR #152703 makes imaplib quote mailbox names that contain spaces or other atom-special characters, per the RFC 3501 grammar (this had been broken since Python 3.0).
It does not add mUTF-7 (modified UTF-7) encoding, so genuinely non-ASCII mailbox names still require either an mUTF-7 codec (the dependency bpo-22598) or a connection with
UTF8=ACCEPTenabled. So #152703 is related and helps, but does not close this issue — the encoding part remains.PR #153391 covers the outbound side: a non-ASCII mailbox name passed to a command is now encoded as modified UTF-7 (RFC 3501, section 5.1.3), so it can be given as an ordinary
str.The inbound side is left as a follow-up.
LIST,LSUB,STATUS, etc. still return the server's raw bytes, so the caller decodes mailbox names itself. Decoding them automatically requires imaplib to parse the untagged responses, which is coming with a planned structured-results mode — there, mailbox-name fields will be decoded from modified UTF-7 (or UTF-8 underUTF8=ACCEPT) transparently.Oh wow, an issue opened in 2009. Better late than never!
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields:
Linked PRs