Repository navigation
Improve file URI ergonomics in urllib.request #125866
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Oct 23, 2024 - added 4 commits that reference this issue
on Oct 25, 2024 - addedstdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directory
on Oct 30, 2024 - added a commit that references this issue
on Oct 30, 2024 7 remaining items
@serhiy-storchaka I've revised the description as follows:
I request that we make
pathname2urlandurl2pathnameeasier to use:pathname2url()is made to accept an optional include_scheme argument that sticksfile:on the front when trueurl2pathname()is made to strip anyfile:prefix from its argument.
Does that seem like a good idea to you?
If we go for these changes, and the user provides a URL with a non-
file:scheme tourl2pathname(), what should happen in your opinion? (I'm erring towards raisingURLError.)Relatedly, what do you think should happen if the user provides a query or fragment in the URL? Currently it's returned as part of the file path. (This seems under-specified to me, but I'm leaning towards raising
URLErrorhere too.)Thank you
I wonder -- would not be better to add such functions in os.path? The implementation is clearly platform depending, and if we add support for platform with different path format (like old macpath), it will need a completely different implementation. In custom build for exotic platform it is easier to provide a separate os.path implementation than patch urllib.request.
Good question! It's something that's come up before:
The current Windows + POSIX implementations don't have much in common, but I think that will change when I address this issue and #123599. To solve them, I'm planning to use
urlsplit()to parse the URL, replacing parts of the janky implementation inurl2pathname(). Here's a rough draft of where I want to end up:def url2pathname(url): scheme, authority, path, query, fragment = urlsplit(url, scheme='file') if scheme != 'file': raise URLError(f'URL uses non-file scheme: {url!r}') if query: raise URLError(f'file URL has query: {url!r}') if fragment: raise URLError(f'file URL has fragment: {url!r}') if os.name == 'nt': # Windows-specific file URI quirks. if authority and authority != 'localhost': # e.g. file://server/share/file.txt path = '//' + authority + path elif path[:3] == '///': # e.g. file://///server/share/file.txt path = path[1:] else: if path[:1] == '/' and path[2:3] in ':|': # Skip past extra slash before DOS drive in URL path. path = path[1:] if path[1:2] == '|': # Older URLs use a pipe after a drive letter path = path.replace('|', ':', 1) path = path.replace('/', '\\') elif not _is_local_authority(authority): # POSIX only: reject URL if authority doesn't resolve to localhost. raise URLError(f'file URL has non-local authority: {url!r}') encoding = sys.getfilesystemencoding() errors = sys.getfilesystemencodeerrors() return unquote(path, encoding=encoding, errors=errors) def pathname2url(pathname, include_scheme=False): if os.name == 'nt': pathname = pathname.replace('\\', '/') encoding = sys.getfilesystemencoding() errors = sys.getfilesystemencodeerrors() prefix = 'file:' if include_scheme else '' drive, root, tail = os.path.splitroot(pathname) if drive: if drive[:4] == '//?/': drive = drive[4:] if drive[:4].upper() == 'UNC/': drive = '//' + drive[4:] if drive[1:] == ':': prefix += '///' drive = quote(drive, encoding=encoding, errors=errors, safe='/:') elif root: prefix += '//' tail = quote(tail, encoding=encoding, errors=errors) return prefix + drive + root + tail
There's clearly some OS-specific code in there, but there's also a fair amount of shared and URL-specific code which might feel out-of-place in
os.path.Besides, if we added functions in
os.pathwe'd need to either live with near-duplicate functionality inurllib, or deprecate theurllibfunctions (which isn't much fun for their users). I'm not ruling that out, but it seems unnecessarily if there's a reasonable path towards makingpathname2url()andurl2pathname()a bit more usable.FWIW, if we go with the above merger, then we can deprecate the oddball
nturl2pathmodule.- added 4 commits that reference this issue
on Apr 10, 2025
Feature or enhancement
I request that we make
pathname2urlandurl2pathnameeasier to use:pathname2url()is made to accept an optional include_scheme argument that sticksfile:on the front when trueurl2pathname()is made to strip anyfile:prefix from its argument.I think this would go a long way towards making these functions usable, and allow us to remove the scary "This does not accept/produce a complete URL" warnings from the docs.
Linked PRs
pathname2url()andurl2pathname()#125993pathname2url()andurl2pathname()(GH-125993) #126144pathname2url()andurl2pathname()(GH-125993) #126145nturl2pathmodule #131432