Repository navigation
Null bytes in url could cause some problems #39592
Description
Activity
url.parse()has a legacy status: https://nodejs.org/api/url.html#url_url_parse_urlstring_parsequerystring_slashesdenotehost
Why not just use the WHATWG URL API? It already classifies both urls as invalid ones.- addedurlIssues and PRs related to the legacy built-in url module.Issues and PRs related to the legacy built-in url module.whatwg-urlIssues and PRs related to the WHATWG URL implementation.Issues and PRs related to the WHATWG URL implementation.
on Aug 9, 2021 @nodejs/url
The second bug described here, where
new URL('a\0b')reports the input as'a'in the error message, originates in the C++ code. The null character gets treated as the end-of-string marker, but only in the error path. It gets encoded and handled correctly, for example, innew URL('http://example.com/a\0b'). (Safari and Firefox also handle that last one correctly, but Chrome throws. I believe Chrome is deviating from the spec, but even if I'm wrong about that, for the purposes of this conversation, it doesn't matter.)Possible solutions:
- Use
std::stringinstead ofchar *and then be very careful so we can preserve the NULL. (I suppose this approach might not work depending on the nature of how the string is handled elsewhere in the C++ code.) - Don't report the input string that caused the error back to the user. (Chrome and Safari both take this route.)
- Don't worry about it. Report the truncated string. (Firefox takes this approach.)
- Preserve the input in JavaScript and re-use it when an error is thrown from C++.
- Use
Chrome's behavior is indeed counter to the spec. https://crbug.com/1099721
4. Preserve the input in JavaScript and re-use it when an error is thrown from C++.
This is the approach used in #42263.
- added a commit that references this issue
on Mar 11, 2022 - Preserve the input in JavaScript and re-use it when an error is thrown from C++.
This is the approach used in #42263.
The WHATWG URL error message issue has been fixed. The legacy
url.parse()has several possible solutions. I think the best is to bail on parsing when there is a NULL character (and possibly other C0 characters) in the host (and possibly a few other places?) and return an object with onlypath/pathname/hrefvalues set to anything other thannull.The WHATWG URL error message issue has been fixed. The legacy
url.parse()has several possible solutions. I think the best is to bail on parsing when there is a NULL character (and possibly other C0 characters) in the host (and possibly a few other places?) and return an object with onlypath/pathname/hrefvalues set to anything other thannull.https://url.spec.whatwg.org/#host-miscellaneous
A forbidden host code point is U+0000 NULL, U+0009 TAB, U+000A LF, U+000D CR, U+0020 SPACE, U+0023 (#), U+002F (/), U+003A (:), U+003C (<), U+003E (>), U+003F (?), U+0040 (@), U+005B ([), U+005C (\), U+005D (]), U+005E (^), or U+007C (|).
A forbidden domain code point is a forbidden host code point, a C0 control, U+0025 (%), or U+007F DELETE.
Looks like
forbiddenHostCharsinurl.jsomitsNULLeven though it includes tab, lf, cr, space, etc. from the list above. It also skips checking on IPv6 hostnames. These seems like bugs that should be fixed. PR coming soon....- added a commit that references this issue
on Mar 12, 2022 - added a commit that references this issue
on Mar 14, 2022 - added a commit that references this issue
on Mar 21, 2022 - added 4 commits that reference this issue
on Apr 21, 2022 - added 2 commits that reference this issue
on Apr 25, 2022
Version
v16.6.0
Platform
Linux MAPLE 5.10.16.3-microsoft-standard-WSL2 #1 SMP Fri Apr 2 22:23:49 UTC 2021 x86_64 GNU/Linux
Subsystem
url
What steps will reproduce the bug?
There are two bugs about null byte:
And the error will be:
The error input is apprently truncated by the null byte.
How often does it reproduce? Is there a required condition?
I think this could only happen when attacker is trying to bypass some SSRF filter in some scenario, but I think it is almost unlikely to happen in realworld.
What is the expected behavior?
It should be invalid url, and http module shouldn't accept null byte.
What do you see instead?
Parsed successfully into a hostname with null byte.
Additional information
No response