Repository navigation
v20.0.0 ERR_SOCKET_CONNECTION_TIMEOUT when sending http request to some domains #47822
Description
Activity
cc @ShogunPanda - looks like
autoSelectFamilyfallout?- addedhttpIssues and PRs related to the http subsystem.Issues and PRs related to the http subsystem.netIssues and PRs related to the net subsystem.Issues and PRs related to the net subsystem.
on May 3, 2023 Yes, it seems to be the case.
@lpgera Can you please post here how is that endpoint resolved by your DNS?
These are the outputs from
dig:$ dig api.airvisual.com ; <<>> DiG 9.10.6 <<>> api.airvisual.com ;; global options: +cmd ;; Got answer: ;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 27494 ;; flags: qr rd ra; QUERY: 1, ANSWER: 2, AUTHORITY: 0, ADDITIONAL: 1 ;; OPT PSEUDOSECTION: ; EDNS: version: 0, flags:; udp: 1232 ;; QUESTION SECTION: ;api.airvisual.com. IN A ;; ANSWER SECTION: api.airvisual.com. 24 IN A 54.178.241.236 api.airvisual.com. 24 IN A 18.178.42.77 ;; Query time: 15 msec ;; SERVER: [[redacted]] ;; WHEN: Wed May 03 09:10:19 CEST 2023 ;; MSG SIZE rcvd: 78$ dig AAAA api.airvisual.com ; <<>> DiG 9.10.6 <<>> AAAA api.airvisual.com ;; global options: +cmd ;; Got answer: ;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 37500 ;; flags: qr rd ra; QUERY: 1, ANSWER: 0, AUTHORITY: 0, ADDITIONAL: 1 ;; OPT PSEUDOSECTION: ; EDNS: version: 0, flags:; udp: 1232 ;; QUESTION SECTION: ;api.airvisual.com. IN AAAA ;; Query time: 5 msec ;; SERVER: [[redacted]] ;; WHEN: Wed May 03 09:12:25 CEST 2023 ;; MSG SIZE rcvd: 46So it seems like there's no IPv6 address and there are two IPv4 addresses resolved, both of which match the TCP handshakes in the Wireshark dump.
I still see these
ERR_SOCKET_CONNECTION_TIMEOUTin v20.2.0. v18.16.0 is fine. Have no reproduction yet, but there issue seems to not be fully resolved. I'm on IPv4 only and the errors happens when doing a few parallel fetches, usually within 1-2 seconds.I'm seeing the same issue attempting to talk to https://api.linode.com/v4/linode/instance - both node-fetch and native fetch are affected. I'm on node v20.2.0 (Intel mac, homebrew)
api.linode.com curently resolves as follows
api.linode.com has address 72.14.191.203 api.linode.com has address 69.164.200.203 api.linode.com has address 72.14.180.203 api.linode.com has IPv6 address 2600:3c00::23 api.linode.com has IPv6 address 2600:3c00::33 api.linode.com has IPv6 address 2600:3c00::13as far as I can tell my ipv6 is currently broken (no dancing turtle on https://www.kame.net/) and if edit /etc/hosts to force an IPv4 addresses for api.linode.com the issue goes away
FYI, The fix for this (#47860) was not yet backported to v20, so that is why it's still happening.
Reacted by Brandon der Blätter and CoheeFor those waiting on a fix: the workaround is to add
dns-result-order=ipv4firstand/orno-network-family-autoselectionto node's command line options or the NODE_OPTIONS environment variable.Reacted by Hendri Pretorius, Santiago Lezica, Damilola Nifemi Adeyemi, Antoine M-P, Torch and VoidThis works:
const net = require("net"); // work around a node v20 bug: https://github.com/nodejs/node/issues/47822#issuecomment-1564708870 if (net.setDefaultAutoSelectFamily) { net.setDefaultAutoSelectFamily(false); }and it seems to prevent the
ERR_SOCKET_CONNECTION_TIMEOUTerror in Node v20.A more forward-compatible version (since this bug is fixed in
20.3.0):// Work around a node v20.0.0, v20.1.0, and v20.2.0 bug. The issue was fixed // in v20.3.0. // https://github.com/nodejs/node/issues/47822#issuecomment-1564708870 // Safe to remove once support for Node v20 is dropped. if ( // !process.env.IS_BROWSER && // uncomment this line if you use a bundler that sets env.IS_BROWSER during build time process.versions && // check for `node` in case we want to use this in "exotic" JS envs process.versions.node && process.versions.node.match(/20\.[0-2]\.0/) ) { require("net").setDefaultAutoSelectFamily(false); }Reacted by Abdullah AReacted by Cohee, Aviral Srivastava, Hendri Pretorius, champion18, Santiago Lezica, Lucas Morais Rodrigues, Abdullah A and Danny ZhangReacted by AzharkoivilaI am still facing this issue. I am on node v20.2.0. Is there any workaround for getting this to work?
I am still facing this issue. I am on node v20.2.0. Is there any workaround for getting this to work?
Oh sorry! Nevermind. the solution by @davidmurdoch works fine. Just including this code on the top fixed my issue.
I am still facing this issue. I am on node v20.2.0. Is there any workaround for getting this to work?
Oh sorry! Nevermind. the solution by @davidmurdoch works fine. Just including this code on the top fixed my issue.
where did you add this
7 remaining items
Your output above is helpful but slightly confusing. I saw you obfuscated the IP addresses but it seems like the client is connecting to the same IP several times, which should not happen.
Thank you for your time, @ShogunPanda.
In the example below I've carefully changed the ip adresses to not expose any sensitive information but to look just as real.
Below
192.168.15.137is local ipv4 address211.190.191.243is server ipv4 addressfa10::3c51:4b1b:3421:a2e1is local ipv6 address2a01:1a5::1d4is server ipv6 address
DNS response returns exactly one ipv4 address and one ipv6 address.
Apr 27 09:53:33 dnsmasq[37664]: forwarded s3.example.com to 1.1.1.1 Apr 27 09:53:33 dnsmasq[37664]: forwarded s3.example.com to 8.8.8.8 Apr 27 09:53:33 dnsmasq[37664]: query[AAAA] s3.example.com from 127.0.0.1 Apr 27 09:53:33 dnsmasq[37664]: forwarded s3.example.com to 1.1.1.1 Apr 27 09:53:33 dnsmasq[37664]: reply s3.example.com is 211.190.191.243 Apr 27 09:53:33 dnsmasq[37664]: reply s3.example.com is 2a01:1a5::1d4Executing node.js code :
01 at 0ms ipv4 [ SYN] 192.168.15.137 => 211.190.191.243 02 at 255ms ipv6 [ SYN] fa10::3c51:4b1b:3421:a2e1 => 2a01:1a5::1d4 03 at 281ms ipv4 [ SYN] 192.168.15.137 => 211.190.191.243 04 at 509ms ipv4 [SYN, ACK] 211.190.191.243 => 192.168.15.137 05 at 509ms ipv4 [ RST] 192.168.15.137 => 211.190.191.243 06 at 532ms ipv6 [ SYN] fa10::3c51:4b1b:3421:a2e1 => 2a01:1a5::1d4 07 at 650ms ipv4 [ SYN] 192.168.15.137 => 211.190.191.243 08 at 901ms ipv6 [ SYN] fa10::3c51:4b1b:3421:a2e1 => 2a01:1a5::1d4 09 at 1176ms ipv4 [SYN, ACK] 211.190.191.243 => 192.168.15.137 10 at 1176ms ipv4 [ RST] 192.168.15.137 => 211.190.191.243 11 at 1942ms ipv4 [SYN, ACK] 211.190.191.243 => 192.168.15.137 12 at 1942ms ipv4 [ RST] 192.168.15.137 => 211.190.191.243Multiple requests to the server happen due to client code receiving
ETIMEDOUTand firing a retry.Second request is fired at 281ms, indicating that
ETIMEDOUTalready happened on the first one. SYN+ACK for the very first request is received at 509ms, but it is already too late. Similarly at 1176ms SYN+ACK is received for the second ipv4 request, while third outgoing ipv4 request has been already fired at 650ms.However when
setDefaultAutoSelectFamily(false)is invoked, noETIMEDOUThappened even after 544ms from the request start time and then communication successfully continued.The code that triggers https request looks like this (constant values are hardcoded according to the values in debugger):
import { request, HttpResponse } from 'https'; const req = request(nodeHttpsOptions, res => { const response = new HttpResponse({ statusCode: res.statusCode || -1, reason: res.statusMessage, headers: res.headers, body: res, }); resolve({ response }); }); req.on('error', err => { // fall here with `ETIMEDOUT` // then trigger retry in a consuming code if (['ECONNRESET', 'EPIPE', 'ETIMEDOUT'].includes(err.code)) { reject(Object.assign(err, { name: 'TimeoutError' })); } else { reject(err); } }); // does not affect the behavior of the code req.on("socket", (socket) => { socket.setKeepAlive(true, 1000); });
Reproduced the behaviour is these node versions:
- 20.10.0
- 20.12.2
- 22.1.0
I see, I can follow it now.
Can you try two things:- Change the default timeout (via
setDefaultAutoSelectFamilyAttemptTimeout) to 100ms - Change the default timeout (via
setDefaultAutoSelectFamilyAttemptTimeout) to 500ms
And in both case post the sequence like you did above?
- Change the default timeout (via
I see, I can follow it now. Can you try two things:
1. Change the default timeout (via `setDefaultAutoSelectFamilyAttemptTimeout`) to 100ms 2. Change the default timeout (via `setDefaultAutoSelectFamilyAttemptTimeout`) to 500msAnd in both case post the sequence like you did above?
Assuming DNS response is the same as in previous example:
s3.example.com is 211.190.191.243 s3.example.com is 2a01:1a5::1d4Below are the results of calling
setDefaultAutoSelectFamilyAttemptTimeout(N)with different parameters.Autoselect timeout 50 (
ETIMEDOUT):01 at 0ms ipv4 [ SYN] 192.168.15.137 => 211.190.191.243 02 at 53ms ipv6 [ SYN] fa10::3c51:4b1b:3421:a2e1 => 2a01:1a5::1d4 03 at 126ms ipv4 [ SYN] 192.168.15.137 => 211.190.191.243 04 at 177ms ipv6 [ SYN] fa10::3c51:4b1b:3421:a2e1 => 2a01:1a5::1d4 05 at 365ms ipv4 [ SYN] 192.168.15.137 => 211.190.191.243 06 at 416ms ipv6 [ SYN] fa10::3c51:4b1b:3421:a2e1 => 2a01:1a5::1d4 07 at 439ms ipv4 [SYN, ACK] 211.190.191.243 => 192.168.15.137 08 at 439ms ipv4 [ RST] 192.168.15.137 => 211.190.191.243 09 at 542ms ipv4 [SYN, ACK] 211.190.191.243 => 192.168.15.137 10 at 542ms ipv4 [ RST] 192.168.15.137 => 211.190.191.243 11 at 859ms ipv4 [SYN, ACK] 211.190.191.243 => 192.168.15.137 12 at 859ms ipv4 [ RST] 192.168.15.137 => 211.190.191.243Autoselect timeout 100 (
ETIMEDOUT):01 at 0ms ipv4 [ SYN] 192.168.15.137 => 211.190.191.243 02 at 103ms ipv6 [ SYN] fa10::3c51:4b1b:3421:a2e1 => 2a01:1a5::1d4 03 at 196ms ipv4 [ SYN] 192.168.15.137 => 211.190.191.243 04 at 298ms ipv6 [ SYN] fa10::3c51:4b1b:3421:a2e1 => 2a01:1a5::1d4 05 at 506ms ipv4 [SYN, ACK] 211.190.191.243 => 192.168.15.137 06 at 506ms ipv4 [ RST] 192.168.15.137 => 211.190.191.243 07 at 677ms ipv4 [ SYN] 192.168.15.137 => 211.190.191.243 08 at 680ms ipv4 [SYN, ACK] 211.190.191.243 => 192.168.15.137 09 at 680ms ipv4 [ RST] 192.168.15.137 => 211.190.191.243 10 at 778ms ipv6 [ SYN] fa10::3c51:4b1b:3421:a2e1 => 2a01:1a5::1d4 11 at 1122ms ipv4 [SYN, ACK] 211.190.191.243 => 192.168.15.137 12 at 1122ms ipv4 [ RST] 192.168.15.137 => 211.190.191.243Autoselect timeout 150 (
ETIMEDOUT):01 at 0ms ipv4 [ SYN] 192.168.15.137 => 211.190.191.243 02 at 154ms ipv6 [ SYN] fa10::3c51:4b1b:3421:a2e1 => 2a01:1a5::1d4 03 at 194ms ipv4 [ SYN] 192.168.15.137 => 211.190.191.243 04 at 346ms ipv6 [ SYN] fa10::3c51:4b1b:3421:a2e1 => 2a01:1a5::1d4 05 at 422ms ipv4 [ SYN] 192.168.15.137 => 211.190.191.243 06 at 434ms ipv4 [SYN, ACK] 211.190.191.243 => 192.168.15.137 07 at 434ms ipv4 [ RST] 192.168.15.137 => 211.190.191.243 08 at 573ms ipv6 [ SYN] fa10::3c51:4b1b:3421:a2e1 => 2a01:1a5::1d4 09 at 638ms ipv4 [SYN, ACK] 211.190.191.243 => 192.168.15.137 10 at 638ms ipv4 [ RST] 192.168.15.137 => 211.190.191.243 11 at 952ms ipv4 [SYN, ACK] 211.190.191.243 => 192.168.15.137 12 at 952ms ipv4 [ RST] 192.168.15.137 => 211.190.191.243Autoselect timeout 500 (success):
01 at 0ms ipv4 [ SYN] 192.168.15.137 => 211.190.191.243 02 at 404ms ipv4 [SYN, ACK] 211.190.191.243 => 192.168.15.137 03 at 404ms ipv4 [ ACK] 192.168.15.137 => 211.190.191.243 04 at 406ms ipv4 [ACK, PSH] 192.168.15.137 => 211.190.191.243 05 at 906ms ipv4 [ ACK] 211.190.191.243 => 192.168.15.137 06 at 906ms ipv4 [ACK, PSH] 211.190.191.243 => 192.168.15.137 07 at 906ms ipv4 [ACK, PSH] 211.190.191.243 => 192.168.15.137 08 at 906ms ipv4 [ ACK] 192.168.15.137 => 211.190.191.243 09 at 906ms ipv4 [ ACK] 192.168.15.137 => 211.190.191.243 10 at 917ms ipv4 [ACK, PSH] 192.168.15.137 => 211.190.191.243 11 at 1418ms ipv4 [ACK, PSH] 211.190.191.243 => 192.168.15.137 12 at 1418ms ipv4 [ACK, PSH] 211.190.191.243 => 192.168.15.137What is confusing me here is that when receiving a timeout the system is somehow retrying it.
First of all, the last attempted record does not has the autoselect timeout applied so I don't see how the thing is relevant here.Second of all, an address is only attempted once.
How did you produce the output above? Which tool are you using?
If possible, can you please try without smithy?So in which Node 20 versions is this fixed?
It was said:
It was fixed in 20.3.0 by this PR: https://github.com/nodejs/node/pull/47860/files
But then
Reproduced the behaviour is these node versions:
20.10.0
20.12.2
22.1.0So was there a regression?
- added 2 commits that reference this issue
on Aug 23, 2024 I still had this problem in node v22.8.0 when trying to fetch data from an API of mine over the Cellular at sea 4G network on a cruise ship.
The
require("net").setDefaultAutoSelectFamily(false);workaround fixed it for me.Reacted by David MurdochI'm seeing this problem in v22.9.0 . Is require("net").setDefaultAutoSelectFamily(false) something that we can set globally?
@Apidcloud Yes, that's exactly its purpose
Reacted by Luís FernandesI am specifically seeing this in the following situation:
- Node 20.18.2 running in Docker on Debian in AWS ECS Fargate container
- Target is a Cloudflare proxy of a Go service in the same VPC.
The timeouts are intermittent. We see 2 ipv4 IPs and 2 ipv6 IPs being attempted. I am guessing the ipv6 ones fail in our VPC immediately and the ipv4 ones take 250ms to fail? The total elapsed time is 500ms. We're using the default 250ms timeout. When it fails, ipv6 fails wuth ENETUNREACH and ipv4 fails with ETIMEDOUT.
Locally, I have tried to "guess and check" my way around this by creating an agent like:
new Agent({ autoSelectFamily: false })and one like:
new Agent({ autoSelectFamily: true, autoSelectFamilyAttemptTimeout: 1000, })and then fetching with:
fetch('http://accounts.released.local:1122', { dispatcher: agent })But if I set NODE_DEBUG=net it still seems to use the same algorithm?
[app] [backend] NET 34953: connect: find host accounts.released.local [app] [backend] NET 34953: connect: dns options { family: undefined, hints: 1024 } [app] [backend] NET 34953: connect: autodetecting [app] [backend] NET 34953: connect/multiple: will try the following addresses [ { address: '::1', family: 6 }, { address: '127.0.0.1', family: 4 } ] [app] [backend] NET 34953: connect/multiple: attempting to connect to ::1:1122 (addressType: 6) [app] [backend] NET 34953: connect/multiple: setting the attempt timeout to 250 ms [app] [backend] NET 34953: connect/multiple: connection attempt to ::1:1122 completed with status -61 [app] [backend] NET 34953: connect/multiple: attempting to connect to 127.0.0.1:1122 (addressType: 4) [app] [backend] NET 34953: connect/multiple: connection attempt to 127.0.0.1:1122 completed with status 0 [app] [backend] NET 34953: afterConnectThe only difference in the logs is:
[app] [backend] NET 34953: connect/multiple: setting the attempt timeout to 250 msvs
[app] [backend] NET 35236: connect/multiple: setting the attempt timeout to 1000 msI have fixed this in nodejs/undici#4070 and also note for future travelers that
new Agent({ autoSelectFamily: false })currently does nothing, butnew Agent({ connect: { autoSelectFamily: false } })does the needful, even though it doesn't satisfy Typescript types.So now when I actually use
autoSelectFamily: falseI appropriately see it attempt only whatever comes first. I have personally also added an explicitfamily: 4because for my use case it is appropriate to only attempt ipv4 addresses.it seems like i am stuck with in issue in 22.14.0, which is LTS...
Reacted by Sick Leviathan and Mateusz Wiszniewskithis breaks node all over the world, in businesses, and half the world trying to connect to America and Europe. The developers deny this is a problem, re implemented multiple times.
Reacted by Mobin Askari, Mateusz Wiszniewski and Sanghee ParkAny update on that? I'm facing this issue on node 20. Downgrade to 18 helped but this is workaround as it's not LTS and it will block me in the near future.
Have anyone find out how to fix this?- added a commit that references this issue
on May 29, 2026
Version
v20.0.0
Platform
macOS 13.3.1, Linux linux 5.15.84-v8+ # 1613 SMP PREEMPT Thu Jan 5 12:03:08 GMT 2023 aarch64 GNU/Linux
Subsystem
http
What steps will reproduce the bug?
After switching to v20.0.0, I cannot send requests to
https://api.airvisual.com/v2anymore and I always get anERR_SOCKET_CONNECTION_TIMEOUTafter about one second.Reproduction 1:
Reproduction 2:
How often does it reproduce? Is there a required condition?
The issue is only present on v20.0.0, previous versions work fine.
What is the expected behavior? Why is that the expected behavior?
The API should respond with the following JSON body:
{"status":"fail","data":{"message":"incorrect_api_key"}}.What do you see instead?
Additional information
Tried a couple of other domains, such as example.com, google.com, cloudflare.com and the requests work fine.
Switching back to v19.8.1 makes the requests to api.airvisual.com work again, so the error seems to be node 20 specific.
I tried talking to the HTTP api instead of using HTTPS and I receive the same error in node.
I tried running
curlfrom every system on which I experience the error and thecurlcommand always succeeded.Tried looking into the network traffic with Wireshark and what I can see is that the node process starts connecting to both resolved IPv4 addresses, and then the connection is immediately closed by the client: