Skip to content

urllib.parse.urlparse output named tuples description is wrong for Python 3.9 and 3.10 #91708

Description

@dnicolodi

The documentation for urllib.parse.urlparse states that:

The return value is a named tuple, which means that its items can be accessed by index or as named attributes, which are:

Attribute Index Value Value if not present
...
params 3 No longer used always an empty string

However it seems that the documentation does not reflect reality:

>>> 
>>> import urllib.parse
>>> p = urllib.parse.urlparse('http://foo.test/test;param')
>>> p
ParseResult(scheme='http', netloc='foo.test', path='/test', params='param', query='', fragment='')
>>> p[3]
'param'

and the returned named tuple has a populated params field.

Activity

  1. domdfcoding commented on Apr 20, 2022

    @domdfcoding
    Contributor

    The example and table changed to what they are now in gh-29816, but I cannot see any related changes in https://github.com/python/cpython/blob/main/Lib/urllib/parse.py which would make urlparse always return an empty string for params. I think the change to the table in gh-29816 is incorrect.

    The example in the docs (urlparse("scheme://netloc/path;parameters?query#fragment")) shows correct output (params='') solely because "scheme" is not a scheme for which urlparse will parse the path params, per the list on line 59:

    uses_params = ['', 'ftp', 'hdl', 'prospero', 'http', 'imap',
    'https', 'shttp', 'rtsp', 'rtspu', 'sip', 'sips',
    'mms', 'sftp', 'tel']

    urlparse checks if the scheme is in that list; if it isn't params will be an empty string, but in other cases it will be parsed from the URL:

    cpython/Lib/urllib/parse.py

    Lines 389 to 392 in d7d7e6c

    if scheme in uses_params and ';' in url:
    url, params = _splitparams(url)
    else:
    params = ''

    I find it a little confusing that params is only parsed for known schemes, but it's probably too late now to change it without breakage. The docs would benefit from clarification of when it is and isn't parsed, with an example to demonstrate the parsing.

  2. xmo-odoo commented on Aug 29, 2022

    @xmo-odoo

    And if params should be de-emphasized, wouldn't it make sense to reorganise the documentation to promote urlsplit instead (despite its worse naming), and mark urlparse, urlunparse, and ParseResult* as deprecated?

    Also possibly urlunsplit, under the assumption that most users have stopped using raw tuples, and thus could just call geturl.

  3. added a commit that references this issue on Oct 7, 2022
  4. added 3 commits that reference this issue on Oct 7, 2022
  5. added 3 commits that reference this issue on Oct 7, 2022
  6. ambv commented on Oct 7, 2022

    @ambv
    Contributor

    Thanks! ✨ 🍰 ✨

  7. added a commit that references this issue on Oct 8, 2022
  8. added a commit that references this issue on Oct 11, 2022
  9. added a commit that references this issue on Oct 22, 2022
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    docsDocumentation in the Doc dir

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions